Skip to content
CRO

Ecommerce Hypothesis Testing: From Idea to Statistical Test

How to write ecommerce test hypotheses and read the statistics: null hypotheses, p-values, power, intervals, Bayesian results, revenue metrics and peeking.

Quick answer

Ecommerce hypothesis testing has two parts. First, a business hypothesis: because we saw specific evidence, we believe a particular change for a defined audience will improve a named metric. Second, a statistical test: assume no difference (the null hypothesis), collect enough data to detect the smallest effect worth caring about, and judge the result with an interval and a pre-set threshold. Understand what p-values do and don't mean, avoid peeking unless using sequential methods, and weigh practical as well as statistical significance.

Two Meanings of Hypothesis

In CRO, "hypothesis" is used for two things. The business hypothesis explains why a change should work, based on evidence. The statistical hypothesis is the formal setup that lets data decide whether an observed difference is likely to be real. Weak business hypotheses produce tests that can't teach anything even when statistically sound; weak statistics produce confident conclusions from noise. You need both.

This article connects them. For the overall test process, see A/B testing framework; for test ideas and basics, see ecommerce A/B testing.

Writing a Business Hypothesis

A useful hypothesis names its evidence, the change, the audience, the expected outcome and how it will be measured. The evidence is what separates a hypothesis from an opinion. It might come from analytics (drop-off at a step), research (users missing information in tests), support tickets, reviews or previous experiments.

PartWeakStrong
Evidence"We think the page is cluttered""Session recordings and a usability test show shoppers missing delivery information"
Change"Redesign the product page""Show the delivery date beside the add-to-cart button"
Audience"Everyone""Mobile visitors on product pages"
Outcome"More sales""Higher add-to-cart rate"
Measure"Conversion""Add-to-cart rate; guardrail: checkout abandonment"

From Business to Statistical Hypothesis

Once the business hypothesis is set, the statistical setup follows. The null hypothesis says the variant's add-to-cart rate equals the control's. The alternative says they differ (two-sided) or that the variant is higher (one-sided). The test then asks how likely the observed difference would be if the null were true.

Two-sided tests are the safer default because they detect harm as well as improvement. One-sided tests need less data but assume you'd never care about a negative effect, which is rarely true in ecommerce.

What P-Values Mean

A p-value is the probability of seeing a difference at least as large as the one observed, if there were really no difference. A p-value of 0.03 means that, if the change did nothing, you'd see a difference this big or bigger about 3% of the time.

It does not mean there's a 97% chance the variant is better, and it says nothing about how big the effect is. A tiny, unimportant effect can be highly significant with enough traffic; a large, important effect can be non-significant with too little. That's why intervals and practical significance matter.

Significance, Errors and Power

Every test balances two errors. A false positive (Type I error) is concluding there's an effect when there isn't; the significance level (commonly 5%) caps its rate for a single, correctly run test. A false negative (Type II error) is missing a real effect; power (commonly planned at 80%) is the chance of avoiding it for an effect of a given size.

Power depends on sample size, the baseline rate, the effect size and variance. Underpowered tests are common in ecommerce: they miss real effects and, when they do find significance, tend to overstate the effect. Planning the minimum detectable effect before the test avoids this.

TermPlain meaningTypical setting
Significance level (alpha)Accepted false positive rate5%
Power (1 − beta)Chance of detecting a real effect of the planned size80%
Minimum detectable effectSmallest effect the test is designed to detectSet from business value
Confidence levelCoverage of the interval method95%

Unsure how to read your test results?

ZSpace reviews experiment statistics and reporting so decisions reflect what the data can support.

Start a Project

Confidence Intervals

A confidence interval gives a range of effect sizes consistent with the data. "Add-to-cart rate +4% (95% CI: +1% to +7%)" tells you the likely direction and size. An interval that spans zero (−2% to +6%) means the test can't rule out no effect. Report intervals with every result; they make practical significance visible and prevent overconfident claims about exact lifts.

Frequentist and Bayesian Approaches

Many testing tools report Bayesian results such as "probability to be best". These are valid when the method and priors are sound, and they're often easier to explain. They don't remove the need for planning: stopping as soon as a probability crosses a threshold can still lead to poor decisions, and priors influence results with small samples.

FrequentistBayesian
OutputP-value, confidence intervalPosterior probability, credible interval
Question answeredHow surprising is this if there's no effect?How likely is each effect given data and prior?
PlanningFixed sample or sequential designStopping rules still needed
Explaining to stakeholdersOften misreadOften more intuitive
Main riskPeeking, misreading p-valuesUnexamined priors, early stopping

Peeking and Sequential Testing

Fixed-horizon tests assume you look at the result once, at the planned end. If you check daily and stop when p drops below 0.05, the real false positive rate is far higher than 5%, because random fluctuations cross the line at some point in many tests.

Sequential testing methods adjust for repeated looks, letting you stop early when evidence is strong while keeping error rates controlled. Some testing tools offer them. If yours doesn't, set the sample size in advance and decide only at the end, while still monitoring for bugs.

Testing Revenue Metrics

Conversion rate is a proportion with fairly predictable variance. Revenue per visitor and average order value are not: most visitors spend nothing, and a few spend a lot, so a single large order can swing results. Revenue tests need larger samples. Common approaches include capping extreme values at a high percentile (decided before the test), bootstrapping intervals, or using conversion as the primary metric with revenue as secondary. State the method in the test plan.

Multiple Comparisons

Every extra metric, segment or variant is another chance for a false positive. Test five segments at 5% significance and there's a good chance one will look significant by chance. Declare a small number of segments and metrics in advance, apply corrections when testing many variants, and treat unexpected segment findings as hypotheses for new tests.

Worked Example

Illustrative numbers, not client data: control has 20,000 visitors and 1,000 add to carts (5.0%); the variant has 20,000 visitors and 1,080 (5.4%). The relative difference is 8%. A two-proportion z-test gives a p-value of about 0.07, and the 95% interval for the absolute difference runs from roughly −0.04 to +0.84 percentage points. Under a 5% threshold, the result is inconclusive: the data is consistent with no effect and with a meaningful one. With the planned sample reached, the team records it as inconclusive and weighs a bolder version of the change.

Two-proportion z-test (sketch)
p1, n1 = 1000/20000, 20000
p2, n2 = 1080/20000, 20000
p = (1000 + 1080) / (n1 + n2)
se = sqrt(p * (1 - p) * (1/n1 + 1/n2))
z = (p2 - p1) / se          # about 1.81
p_value = 2 * (1 - Phi(abs(z)))   # about 0.07

Practical Significance

A statistically significant result still needs a business judgement. Is the effect large enough to justify the cost of building and maintaining the change? Does it hold across devices and markets? Are there risks (returns, brand, accessibility) that outweigh it? Decision rules in the test plan should include a practical threshold, not only a statistical one. See experiment prioritization.

Choosing the Minimum Detectable Effect

The minimum detectable effect (MDE) is a business decision, not a statistical one. Ask: what's the smallest improvement that would justify building and maintaining this change? For a cheap copy change, a small effect may be worth it; for an expensive feature, only a larger one. Then check whether traffic allows detecting that effect in reasonable time. If not, make the change bolder, move to a higher-traffic page or accept that the test can only detect large effects.

SituationImplication
Required MDE is large relative to realistic effectsTest likely inconclusive; rethink
High-traffic template, small realistic effectFeasible with patience
Low traffic, bold changeFeasible for large effects only
Expensive change, uncertain valueSet a higher MDE; test before building fully

Variance Reduction

Some testing platforms support variance reduction techniques that use pre-experiment data (such as a visitor's previous behaviour) to reduce noise, allowing smaller samples for the same power. CUPED is a widely used example. These methods are valid when implemented correctly, but they add complexity; use them through tools that support them rather than building ad hoc versions.

Explaining Results to Stakeholders

Report results in plain language: the estimated effect, the range it could plausibly be in, whether guardrails held, and the decision. For example: "Add-to-cart rate was about 4% higher with the variant, likely somewhere between 1% and 7%. Returns and speed were unchanged. We're shipping it." Avoid saying a variant is "95% likely to win" unless your method actually produces that probability, and avoid precise lift figures presented without intervals. See dashboard design.

Common Mistakes

  • Hypotheses without evidence
  • Reading p-values as the probability the variant wins
  • Underpowered tests with no minimum detectable effect
  • Peeking with fixed-horizon statistics
  • Reporting lifts without intervals
  • Revenue tests without a plan for outliers
  • Finding winners in unplanned segments

Ready to make test results trustworthy?

Talk to ZSpace about experimentation reviews, hypothesis-led design and test implementation.

Start a Project

Conclusion

Good hypothesis testing pairs a business hypothesis grounded in evidence with statistics used correctly: planned power, intervals over bare p-values, no peeking without sequential methods and careful handling of revenue metrics. Related: experimentation mistakes and experimentation program.

FAQ

Common questions

A statement linking evidence to a change and an expected outcome: because we observed X, we believe changing Y for audience Z will improve metric M. It makes tests purposeful and results interpretable.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.