A/B Testing definition
A/B testing is a controlled experiment that compares two versions of a web page, app screen, email or feature by randomly showing each version to a different group of users and measuring which performs better on a chosen metric, such as sign-ups or purchases. Statistical analysis shows whether the difference is likely real rather than random chance.
How does A/B testing work?
Users are randomly assigned to version A, the current control, or version B, the variant, and each user consistently sees the same version. Because assignment is random, the groups are similar in every way except the change being tested, so a difference in results can be attributed to the change. A good test is planned before it starts, not interpreted creatively afterward.
- Write a hypothesis: changing X will improve Y because of Z.
- Choose one primary metric and guardrail metrics that must not get worse.
- Calculate the sample size from the baseline rate and the smallest effect worth detecting.
- Randomly split traffic and check the split is working correctly.
- Run for full weekly cycles until the planned sample is reached.
- Analyze the result, decide and document what was learned.
Example: testing a pricing page
A SaaS company suspects visitors leave its pricing page because they cannot tell which plan suits them. Version B adds a short "best for" line under each plan and highlights the most popular one. Half of visitors see each version for four weeks. Trial sign-ups rise in version B and the confidence interval excludes zero, while guardrails such as refund requests and support tickets stay flat, so the team ships the change and records the insight about plan clarity.
A/B testing statistics basics
Statistical significance estimates how likely a difference this large would be if the versions actually performed the same. Power is the chance of detecting a real effect of a given size, and the minimum detectable effect is the smallest improvement the test is designed to find. Low-traffic sites need large effects or long tests to reach reliable conclusions. Frequentist and Bayesian methods both work when used correctly.
Common errors include peeking at results daily and stopping when they look good, which inflates false positives; testing many metrics or variants without correction; and ignoring a sample ratio mismatch, where traffic is not split as intended, which usually signals a bug that invalidates the test.
A/B testing tools
Popular experimentation platforms include Optimizely, VWO and AB Tasty for websites, GrowthBook and Statsig for product teams, LaunchDarkly experiments built on feature flags, and Firebase A/B Testing for mobile apps. Google Optimize was discontinued in 2023, moving many marketers to other tools. Server-side testing through feature flags avoids flicker and works for backend changes such as pricing logic or recommendation algorithms.
When not to A/B test
Do not test obvious fixes such as broken buttons or slow pages; just fix them. Avoid testing when traffic is too low to detect realistic effects, and use qualitative research instead. Be careful with changes whose effects appear slowly, such as pricing or retention features, where short tests mislead. Multivariate tests and multi-armed bandits suit specific situations but need even more traffic or clear trade-offs. Nexzem sets up experimentation programs with proper tracking, feature flags and pre-registered analysis plans.