EyeCaptain
    A/B Test Hypothesis: A Template With Good and Bad Examples

    A/B Test Hypothesis: A Template With Good and Bad Examples

    Learn how to write an A/B test hypothesis with a simple template, five bad and good examples, metric choices, and the mistakes that waste test traffic.

    Dimitris Andreadakis

    7 min read
    202 views
    A/B test hypothesishypothesis templatehow to write a test hypothesisCRO hypothesis examplessplit test hypothesis

    A good A/B test hypothesis follows one template: "Because we observed X, we believe changing Y for audience Z will improve metric M." It names the evidence, the exact change, who sees it, and how success is measured. If any of the four parts is missing, the test can run but you will not learn much from the result.

    What is an A/B test hypothesis?

    An A/B test hypothesis is a written prediction that links a specific change to a specific outcome, along with the reason you expect it. It is written before the test is designed, and it decides what you build, who sees it, and what you measure.

    A hypothesis is not the same as an idea. "Try a green button" is an idea. A hypothesis explains why the button should change, for whom, and what number should move. That extra structure is what turns a test into learning, because when the result comes in you can say which belief it supported or broke.

    This article is about writing the hypothesis itself. For the statistics of reading results, see our posts on the peeking problem and on Bayesian vs frequentist testing.

    What is the best A/B test hypothesis template?

    The most useful template has four slots:

    Because we observed [X: evidence], we believe that changing [Y: the specific element and change] for [Z: the audience or page] will improve [M: the primary metric].

    Each slot does a job:

    • X, the observation, forces you to name evidence. Analytics, heatmaps, session recordings, user tests, support tickets, reviews, and audit findings all count. "The CEO thinks so" does not.
    • Y, the change, must be specific enough that a designer and developer could build it without asking questions.
    • Z, the audience, sets the scope: all visitors, mobile only, new visitors, paid traffic, a single product category.
    • M, the metric, is the one number that decides the test. Pick it before launch.

    Many teams add a fifth line, "We will know this is true when M changes, and we will watch guardrail metrics G." Guardrails are numbers that must not get worse, such as refund rate, average order value, or lead quality.

    EyeCaptain
    The four parts of a test hypothesis. Write it before the variant is designed.

    How do you write an A/B test hypothesis step by step?

    Start from evidence, not from the solution. The steps below use a hypothetical project management SaaS and its pricing page as the example.

    1. Write the observation first. Describe what you saw in one or two sentences, with the source. Example: "Session recordings show mobile visitors scrolling up and down between the plan cards and the feature table."
    2. Name the problem behind it. "Visitors cannot compare plans without leaving the cards."
    3. Propose one change that addresses that problem. "Add the three most important features as a short list inside each plan card."
    4. Define the audience. "Mobile visitors to the pricing page."
    5. Pick the primary metric closest to the change that still matters to the business. "Clicks on any plan's Start trial button" or, better if volume allows, "trial starts."
    6. List guardrails. "Paid conversion from trial must not drop."
    7. Write down what you will do with each outcome. If it wins, ship it and test the same pattern on desktop. If it loses, review recordings of the variant before the next idea.

    Step 7 is often skipped. If no result would change what you do next, the test is not worth the traffic.

    What do good and bad A/B test hypotheses look like?

    Bad hypotheses are vague about the change, missing evidence, or measured on a metric that cannot tell you anything. Here are five pairs, each fixed using the template.

    1. The button color hypothesis

    Bad: "A red button will convert better."

    Good: "Because a predictive heatmap and a scroll map show the beige Add to cart button gets little attention on mobile product pages, we believe changing it to a high contrast color and full width for mobile visitors will increase add to cart rate."

    The fix: the color itself was never the point. The evidence is about visibility, so the change targets visibility and the metric is the one closest to the button.

    2. The headline hypothesis

    Bad: "A better headline will increase conversions."

    Good: "Because support chats show new visitors often ask whether the tool works with Shopify, we believe changing the homepage headline from 'Smarter inventory' to 'Inventory forecasting for Shopify stores' for new visitors will increase trial signups."

    The fix: "better" is not a change. The good version states both headlines, the reason, and the audience.

    3. The social proof hypothesis

    Bad: "Adding testimonials builds trust."

    Good: "Because exit survey answers mention doubts about whether the course suits beginners, we believe adding two testimonials from beginner students directly under the Enroll button on the course page for all visitors will increase enrollments."

    The fix: the bad version is a general belief, not a testable prediction. The good one ties a specific kind of proof to a specific objection and places it where the decision happens.

    4. The form length hypothesis

    Bad: "Shorter forms convert better, so remove fields."

    EyeCaptainEyeCaptain
    97%
    Conversions Booster

    Your visitors leave without converting

    Most websites lose 97% of their traffic without a single conversion. Our AI scans 200+ CRO elements to find exactly where visitors drop off.

    Hero CTA missing

    Good: "Because form analytics show the highest abandonment at the 'Company size' and 'Budget' fields, we believe removing those two fields from the quote request form for paid search visitors will increase completed quote requests, while sales qualified lead rate stays the same."

    The fix: it names the fields and adds a guardrail. Removing qualification fields can raise volume and lower quality, so lead quality must be watched.

    5. The redesign hypothesis

    Bad: "The new product page design will improve the user experience."

    Good: "Because recordings show mobile visitors scroll past the long description and rarely reach reviews or shipping details, we believe moving the review summary and delivery estimate above the description on mobile product pages will increase add to cart rate."

    The fix: "user experience" is not a metric, and a full redesign changes too many things at once. If it wins, you will not know why. Test the change the evidence supports.

    EyeCaptain
    Write a hypothesis before designing variants. Start with the observed purchase hesitation. Change one clear element for a…

    How do you choose the right metric for a hypothesis?

    Choose the metric closest to the change that still reflects business value, and decide it before launch. A change to a product page button should be measured on add to cart rate, with orders as a secondary metric. A change to a checkout step should be measured on completed orders.

    Change locationGood primary metricUseful guardrail
    Product page CTA or layoutAdd to cart rateReturn rate, average order value
    Cart or checkout stepCompleted orders per checkout startAverage order value
    Pricing pageTrial or plan startsTrial to paid conversion
    Lead formCompleted form submissionsQualified lead rate
    Homepage heroClicks to key category or product pagesBounce rate, orders
    Booking flowCompleted bookingsCancellation rate

    Avoid vanity metrics like time on page or scroll depth as the primary metric. They can rise for good or bad reasons, so a win on them proves little.

    EyeCaptain
    Choose a metric close to the change. Add a guardrail for the business trade-off.

    What are the most common hypothesis mistakes?

    The most common mistake is writing the hypothesis after the test, to fit the result. Other frequent problems:

    • Bundling changes. A new headline, image, and button together can win, but you will not know which part mattered.
    • No evidence slot. If you cannot fill in "because we observed," you are testing a guess. Run research first, or a CRO audit, to find real problems.
    • Mismatched audience. Testing a mobile layout problem on all traffic dilutes the effect with desktop visitors who never had the problem.
    • Too many metrics. If you track ten metrics, one will move by chance. Pick one primary metric.
    • No plan for a loss. A losing test with a clear hypothesis still teaches you something. A losing test with a vague one teaches nothing.

    Key takeaways

    • Use the template: because we observed X, we believe changing Y for Z will improve M.
    • Start from evidence and write the hypothesis before anyone designs the variant.
    • Make the change specific, the audience explicit, and pick one primary metric plus guardrails.
    • Decide in advance what you will do if the test wins, loses, or shows no difference.

    For the full testing process, from research to rollout, see the A/B testing guide. A free CRO audit on EyeCaptain lists observed problems on one page, which you can use as the "because we observed" part of your next hypothesis.

    EyeCaptain
    Turn a worry into a test. Evidence, change, audience and metric belong together.

    Frequently asked questions

    What is the difference between a null and an alternative hypothesis in A/B testing?

    The null hypothesis says the change makes no difference to the metric. The alternative hypothesis says it does, in a stated direction. Your test result is evaluated against the null. The business hypothesis in this article, built from evidence, change, audience, and metric, is what you write first; the null and alternative are the statistical framing of that same idea.

    Can one A/B test have more than one hypothesis?

    One test should have one primary hypothesis tied to one primary metric. You can record secondary questions, such as whether the change affects mobile and desktop differently, but treat those as exploratory. If you have two separate ideas, run them as separate tests or as clearly separate variants, so each result can be traced back to its own reasoning.

    Where do good A/B test ideas come from?

    The best ideas come from evidence about where visitors struggle. Useful sources include funnel reports in analytics, heatmaps and session recordings, on-site surveys, customer support logs, product reviews, user tests, and expert audits of the page. Competitor pages can inspire solutions, but they do not tell you what your own visitors need.

    Should you write a hypothesis for small changes like copy tweaks?

    Yes, if you are going to test them. A one line hypothesis takes a minute and stops you from testing tweaks with no reason behind them. If a change is an obvious fix, such as a typo or a broken link, skip the test and the hypothesis, and simply ship it.

    Enjoyed this article?

    One short email a week with practical CRO and UX tips

    No spam. Unsubscribe anytime.

    Find what stops your visitors from converting

    EyeCaptain is an AI-powered CRO & UX audit tool that automatically scans your pages, identifies UX issues, and gives you actionable optimization suggestions to increase conversions. Try it for free, no card, no commitment.

    Free CRO Audit
    Share this article
    LinkedIn X

    About the author

    Founder of EyeCaptain. Writes about conversion rate optimization, UX analysis and the neuromarketing behind pages that actually sell.

    More articles by Dimitris Andreadakis