A/B testing guide

A/B testing guide: how to run tests you can trust

How to plan, run and read an A/B test without fooling yourself: whether you have enough traffic, what significance really means, what to test first, and the mistakes that turn noise into winners.
A · Control2.4%
Submit
B · Variant3.1%
Get my free quote

Illustrative. A lift only counts once the test reaches its planned sample.

What an A/B test is

You show two versions of the same page to comparable visitors at the same time: the one you have (A, the control) and one with a change (B, the variant). Whichever converts better, by a margin larger than chance would explain, wins. That is the whole idea. Everything else in this guide is about doing it without fooling yourself.

The method is borrowed from agricultural and medical trials, where Ronald Fisher formalised it in the 1920s. On the web it answers one question well: did this change cause more people to act? It does not tell you what to change. That comes from research: analytics, recordings, customer interviews, and an audit of the page.

Do you have enough traffic to test?

Most pages do not. The number of visitors you need depends on your current conversion rate and on how small a lift you want to be able to detect. Small lifts on low-converting pages need enormous samples. Try your own numbers:

Sample size calculator

2.5%
+20%
20,000

Visitors per variant

15,600

Weeks at your traffic

7

Short enough to test. Run it for at least two full weeks so every weekday is in the sample twice. Two variants, 95% confidence, 80% power.

If the answer is months, you have three honest options: test a bolder change that could move the rate by a lot, test on a page with more traffic, or make the change without a test and watch the trend. What you should not do is run the test for two weeks and read the result anyway.

The process, in seven steps

  1. Find the leak. In analytics, look for the step with the biggest drop: product page to cart, cart to checkout, landing page to form.
  2. Diagnose the page. An audit, a heatmap and a few session recordings tell you why people drop there. Guessing tells you nothing.
  3. Write a hypothesis. "Because the Add to cart button is third in visual weight, making it the darkest element will raise add-to-cart rate."
  4. Prioritise. Score ideas with ICE or PIE and start with high impact, high confidence, low effort.
  5. Plan the sample. Decide the metric, the sample size and the duration before you start, and write them down.
  6. Run it and do not peek. Every time you check and stop early you raise the chance of a false winner.
  7. Ship, log, repeat. Record what you tested, the result and what you learned, including the tests that lost.

Statistical significance, plainly

A 95% significance level means that if there were no real difference, you would see a result this strong only about one time in twenty. It does not mean there is a 95% chance B is better. Two rules keep you honest: decide the sample size in advance and stop only when you reach it, and run at least two full weeks so weekdays and weekends are both in the sample. Most testing tools handle the maths; your job is to not override it.

What to test first

Store pagesLead generation pages
Add to cart prominence and wordingNumber of form fields
Delivery cost and time on the product pageButton copy that says what you get
Reviews next to the priceHeadline that names who it is for
Guest checkout and express paymentProof beside the form
Free-delivery threshold messagingWhat happens after submission

Structural changes (what is visible, in what order, how much effort is asked) move conversion far more than colours and synonyms. For a longer list, see the 20 CRO techniques.

Ten mistakes that ruin tests

  • Stopping early because B is ahead after three days.
  • Peeking repeatedly and stopping at the first significant result.
  • Changing five things at once and not knowing which one worked.
  • Optimising clicks when the business needs orders or qualified leads.
  • Testing without a hypothesis, so a loss teaches nothing.
  • Too small a sample, which turns noise into winners.
  • Ignoring devices: a change can win on desktop and lose on mobile.
  • Testing during a sale and expecting the result to hold on a normal week.
  • Polishing details while the value proposition is unclear.
  • Not writing results down, so the team re-runs the same tests a year later.

Tools

Experimentation platforms such as VWO, Optimizely, AB Tasty, Convert and Kameleoon run the test, split traffic and do the statistics. GA4 reports the outcome but no longer runs experiments since Google Optimize was retired. Pick by the kind of test you need (client-side visual changes or server-side features), your traffic and your privacy requirements. See EyeCaptain vs VWO for how an audit and a testing platform fit together.

Where an audit fits

An A/B test answers "did this work?". An audit answers "what is worth trying?". EyeCaptain reads your page on desktop and mobile, measures what pulls attention and what competes with your call to action, and returns the changes ranked by expected impact, each with a before and after. On pages with enough traffic that becomes your testing backlog; on pages without it, those are the changes to ship first. Start with the free CRO audit.

Q  FAQ

Frequently asked questions

Until it reaches the sample size you planned before starting, and for at least two full weeks so every weekday appears twice. Stopping when one variant looks ahead is the most common way to crown a false winner.

See what your page is losing

One free audit, no card. If it does not tell you something you did not know, you will know that in five minutes.

No card · Result in under 5 minutes · Desktop and mobile