A/B testing has a credibility problem in small SaaS, and it’s deserved. Most tests run at small companies produce a number that looks like a result but isn’t one — because the sample was too small, the test was stopped when it looked good, or the change was too minor to detect. Here’s how to run tests that actually change what you believe.
First: do you have enough traffic to test at all?
This is the question that decides whether the rest of this article is relevant to you.
The intuition to internalize: detecting small differences requires a lot of data. Going from a 2% to a 3% conversion rate is a big business change but a small statistical one, and you’ll need thousands of visitors per variant to see it reliably. Detecting a jump from 2% to 4% needs far fewer.
Two practical consequences:
- If you have a few hundred visitors a month, don’t A/B test. You cannot detect anything short of an enormous change, and you’ll spend months learning nothing. Do qualitative work instead — watch five people use it, read your support inbox, ask why people didn’t sign up.
- Test big changes, not small ones. A different offer, a different headline promise, a restructured page. Button colors need traffic you don’t have and produce effects too small to matter anyway.
Use a sample size calculator before you start, not after. Enter your current conversion rate and the smallest lift that would actually change a decision. If the required sample is more traffic than you’ll get this quarter, don’t run the test.
Test one thing, and make it a real thing
Two rules that kill most bad tests:
One variable per test. If you change the headline, the image, and the CTA together and conversion rises, you’ve learned that something worked. You can’t ship the winner intelligently or apply the lesson elsewhere.
Make the change substantial. The point isn’t to be scientifically cautious about tiny variations — it’s to compare two genuinely different bets. “Free plan vs 14-day trial” is a test worth running. “Get Started” vs “Get started” is not.
Write the hypothesis down before you start
A test without a stated prediction turns into a hunt for a good-looking number. Write, in one sentence:
Because [observation], I believe [change] will [effect] for [audience], measured by [metric].
For example: because most trial signups never reach first value, I believe removing the credit card requirement will raise activated users per visitor, measured over four weeks.
This forces two useful things: a reason grounded in evidence, and a metric chosen in advance.
Choose the right metric — usually not the obvious one
Optimizing the immediate click is how you win a test and lose money. A headline that overpromises will beat an honest one on signups and lose on activation and retention.
- Measure the furthest-down-funnel outcome you can in a reasonable window: activated users, or paying customers.
- Watch a guardrail metric — if signups rise 20% but activation falls 30%, the variant lost.
- Decide the window in advance, and make it long enough to include your typical consideration cycle. Weekday and weekend traffic behave differently, so run in whole weeks.
Don’t stop early — this is the big one
The single most common way small teams fool themselves: checking daily and stopping the moment significance appears. Conversion rates fluctuate wildly in early samples, and if you keep looking, you will eventually see a “significant” difference that’s pure noise.
The discipline: decide the sample size and duration up front, then don’t act until you reach it. Look if you must — but don’t stop, and don’t ship, before the pre-agreed endpoint.
Related trap: running six tests and shipping the one that won. If you test enough things, one will look like a winner by chance alone.
Accept the boring outcome
Most tests come back inconclusive. That’s not a failed test — it’s a real result telling you the change didn’t matter enough to detect. The right response is to keep the simpler version and go test something bigger, not to re-run it hoping for a different answer.
Also worth internalizing: a “20% lift” from a small test is usually an overestimate. Effects measured in underpowered tests are systematically inflated, so expect the real-world result to be smaller than the test suggested.
What to test when you can’t A/B test
Most early-stage SaaS is in this position, and there’s plenty to do:
- Sequential testing — run version A for four weeks, B for four weeks, and treat the result as directional rather than proof. Imperfect (seasonality contaminates it), but better than nothing.
- Qualitative research. Five user sessions will tell you more about a broken signup flow than any test you can power.
- Ship the obvious fixes without testing. If a step is clearly redundant, remove it. You don’t need an experiment to justify deleting a form field nobody reads (friction removal).
- Paid traffic as a test harness. If you need a fast read on two messages, buying traffic to both is often quicker than waiting for organic volume — one of the genuinely good reasons to run ads early (paid vs organic).
Things worth testing, roughly ranked
- The offer — free plan vs trial, card vs no card. Biggest swings by a distance.
- The headline promise — not the wording, the claim.
- Page structure — what’s above the fold and in what order (landing page optimization).
- Onboarding path — what the user does first.
- Pricing presentation — how tiers are framed and ordered.
- Copy and creative variations.
- Visual details. Last, and only at real traffic volumes.
FAQ
How much traffic do I need for an A/B test?
Enough that the smallest lift you’d act on is detectable — often thousands of visitors per variant for typical SaaS conversion rates. Run a sample size calculation before starting; if the number exceeds your quarterly traffic, don’t run the test.
How long should an A/B test run?
Until it reaches the sample size you calculated in advance, and in whole weeks to avoid day-of-week effects. Stopping as soon as a result looks significant is the most common way tests produce false winners.
What should I do if I don’t have enough traffic to A/B test?
Use qualitative research — watch real users, read support tickets, ask people who didn’t convert — and ship obvious friction fixes without testing. Save experiments for large changes once volume supports them.
More conversion and growth guidance is in the growth & pricing hub.
