An A/B test shows two versions of something to different groups of real users at the same time and measures which performs better against a metric you chose in advance. The point is to replace an argument about what users will do with a measurement of what they did.
What makes a test trustworthy
Four things, and skipping any of them turns the result into a coin flip you have dressed up as evidence.
- One metric, chosen first
- Decide before you start what counts as winning. Choosing afterwards means you will find something that improved, because in any data set something always did.
- Random assignment
- Users go into groups at random, not by signup date or region or plan. Otherwise you are measuring the difference between the groups, not the change.
- Same time period
- Running A this week and B next week measures the week as much as the change.
- Enough traffic to see the difference
- The smaller the effect you are looking for, the more users you need. This is the constraint that stops most small products A/B testing anything meaningful.
The traffic problem, honestly
Detecting a change from a 4% conversion rate to 5% needs thousands of users per group. If your product has hundreds of weekly signups, an A/B test on that funnel will not reach a conclusion before the product has changed underneath it. That is not a reason to guess - it is a reason to use methods that work at your size: watching session recordings, talking to five users, or shipping the change and watching your error and feedback numbers.
What to do when you cannot test
Most teams cannot A/B test most decisions, and pretending otherwise leads to paralysis or to tests that are noise. The practical alternative is to make the decision with whatever evidence exists - error counts, support tickets, survey responses, a handful of user conversations - and to write down what you expected so you can tell afterwards whether you were right. Feature flags let you roll a change out to 10% first, which catches a disaster even when it cannot prove a small win.
Frequently asked questions
How long should a test run?
Until it reaches the sample size you calculated up front, and for at least one full weekly cycle so you are not measuring a Tuesday. Stopping the moment it looks good is the most common way to get a false result.
Can we test more than two versions?
Yes, that is an A/B/n test, but each extra version splits your traffic further and needs more of it. At most product sizes, two is what you can afford.
How Spectr helps
Prioritise with numbers instead of whoever sounds most certain
Real error and feedback numbers from PostHog and Sentry, quoted with a date in your specs
Related terms
Spend less of the week on documents
Spectr drafts the spec, the requirement document and the stories from your meetings. You review and approve.
Free plan available. 30-day free trial on Starter.