Checkout Conversion A/B Test
Did the new checkout actually convert better?
An end-to-end analysis of an e-commerce checkout experiment: current flow vs. a streamlined redesign, with conversion as the single primary metric. The hypothesis, sample size, and ship rule were fixed before looking at any outcome; the data was audited for assignment integrity and sample-ratio mismatch before any test ran. The redesign converted no better — the 95% CI on the difference is [−0.38, +0.09] pp, and even its best case falls far short of the +1 pp worth shipping for. A well-founded “no” that saves building something that doesn’t work.
How it works
Change in checkout conversion from the redesign, across 286,690 users. The 95% interval contains zero, and even its best case sits far below the bar worth shipping for. Recommendation: do not ship.
Decide the rule before the data
Before looking at a single outcome, the experiment was pre-registered: one primary metric (conversion per user), a two-sided test at α = 0.05, and a minimum effect worth shipping of +1 pp — 12.02% → 13.02%, about +8.3% relative.
At 80% power, detecting that effect needs ~17,210 users per group. The test ended with ~143,300 — so heavily over-powered that even a trivial +0.2 pp would read as significant. That’s why the ship rule demands both conditions below, not just a small p-value.
Trust the data before testing it
3,893 rows had a group that disagreed with the page actually shown — and every one of them belonged to a user who appeared twice. Dropping those users whole removed all contamination at once.
| Drop log | Rows |
|---|---|
| Raw rows | 294,478 |
| Duplicated users — 3,894 users seen twice, removed whole | −7,788 |
| Inconsistent group / page left after that | −0 |
| Clean, one row per user | 286,690 |
With p = 0.846 there is no sample-ratio mismatch: randomization held, so whatever the test says next can be trusted.
What would it have taken to ship?
The bar is the 95% confidence interval on the difference. Move the lift and the sample size to see when a result becomes significant — and when it’s significant but still too small to be worth it.
- 95% CI
- [−0.38, +0.09] pp
- excludes 0?
- no
- lift ≥ +1 pp?
- no
The interval contains zero — the flows can’t be told apart.
Not inconclusive — informative
The interval [−0.38, +0.09] ppcontains zero, so the two checkouts can’t be told apart. But it says more than that: its most favorable end is 11× smaller than the +1 pp worth shipping for.
With this much data, a real effect of any meaningful size would have shown up. Its absence is the answer.
No early bump that fades
A pooled average can hide a novelty effect — a redesign that wins for a few days and then decays. Hover the chart to read each day of the 23-day window.
view as table
| Day | Control | Treatment | Diff (pp) | Users |
|---|---|---|---|---|
| Jan 2 | 12.58% | 11.92% | −0.66 | 5,631 |
| Jan 3 | 11.36% | 11.41% | +0.04 | 13,025 |
| Jan 4 | 12.14% | 11.68% | −0.47 | 12,929 |
| Jan 5 | 12.32% | 11.44% | −0.89 | 12,748 |
| Jan 6 | 11.51% | 12.36% | +0.85 | 13,168 |
| Jan 7 | 12.01% | 11.64% | −0.37 | 13,013 |
| Jan 8 | 11.89% | 12.10% | +0.21 | 13,231 |
| Jan 9 | 11.93% | 11.84% | −0.09 | 13,064 |
| Jan 10 | 11.30% | 12.66% | +1.36 | 13,184 |
| Jan 11 | 11.77% | 11.53% | −0.24 | 13,183 |
| Jan 12 | 12.17% | 12.19% | +0.02 | 12,992 |
| Jan 13 | 11.66% | 11.14% | −0.52 | 12,893 |
| Jan 14 | 12.59% | 12.01% | −0.58 | 12,983 |
| Jan 15 | 12.06% | 11.32% | −0.74 | 13,081 |
| Jan 16 | 12.19% | 11.96% | −0.23 | 12,948 |
| Jan 17 | 12.33% | 12.67% | +0.34 | 12,976 |
| Jan 18 | 12.38% | 12.39% | +0.01 | 12,915 |
| Jan 19 | 12.02% | 11.61% | −0.41 | 12,968 |
| Jan 20 | 11.47% | 11.78% | +0.31 | 13,051 |
| Jan 21 | 12.55% | 11.61% | −0.94 | 13,139 |
| Jan 22 | 11.94% | 11.77% | −0.16 | 13,059 |
| Jan 23 | 12.65% | 12.09% | −0.56 | 13,164 |
| Jan 24 | 11.79% | 11.98% | +0.19 | 7,345 |
The lines cross again and again with no sustained gap and no early treatment lead. The daily swings are sampling noise; the null result is stable over time, not an artifact of averaging.
My contribution
Three takeaways
Framing note: the public dataset is technically a landing-page test (old vs. new page). Its structure — assigned group, variant shown, converted yes/no — is identical to a checkout test, so it’s framed that way for the narrative. The statistics are unchanged.