Back to projects
shippedlive demo

Checkout Conversion A/B Test

Did the new checkout actually convert better?

An end-to-end analysis of an e-commerce checkout experiment: current flow vs. a streamlined redesign, with conversion as the single primary metric. The hypothesis, sample size, and ship rule were fixed before looking at any outcome; the data was audited for assignment integrity and sample-ratio mismatch before any test ran. The redesign converted no better — the 95% CI on the difference is [−0.38, +0.09] pp, and even its best case falls far short of the +1 pp worth shipping for. A well-founded “no” that saves building something that doesn’t work.

Users286,690
Observed lift−0.14 pp
Decisiondo not ship
PythonpandasstatsmodelsSciPymatplotlib
how.it.works

How it works

Pre-register+1 pphypothesis · metric · ship bar
Power17,210users per group required
Clean294k → 287k−3,894 duplicated users
SRMp = 0.84650/50 split holds
Test−0.14 ppz-test · 95% CI on the diff
Decideno shipfails both conditions
−0.14 pp

Change in checkout conversion from the redesign, across 286,690 users. The 95% interval contains zero, and even its best case sits far below the bar worth shipping for. Recommendation: do not ship.

[−0.38, +0.09]95% CI on the differencepp · contains 0
0.23p-valuetwo-sided z-test
8.3×the required sample143k vs 17k per group
01

Decide the rule before the data

Before looking at a single outcome, the experiment was pre-registered: one primary metric (conversion per user), a two-sided test at α = 0.05, and a minimum effect worth shipping of +1 pp — 12.02% → 13.02%, about +8.3% relative.

At 80% power, detecting that effect needs ~17,210 users per group. The test ended with ~143,300 — so heavily over-powered that even a trivial +0.2 pp would read as significant. That’s why the ship rule demands both conditions below, not just a small p-value.

ship only if
1Statistically significant — the 95% CI on the difference excludes 0✕ fails
and
2Practically significant — the lift is at least +1 pp✕ fails
02

Trust the data before testing it

3,893 rows had a group that disagreed with the page actually shown — and every one of them belonged to a user who appeared twice. Dropping those users whole removed all contamination at once.

Drop logRows
Raw rows294,478
Duplicated users — 3,894 users seen twice, removed whole−7,788
Inconsistent group / page left after that−0
Clean, one row per user286,690
sample-ratio check
control143,293treatment143,397
χ² = 0.038 · p = 0.846 · 50.02% treatment

With p = 0.846 there is no sample-ratio mismatch: randomization held, so whatever the test says next can be trusted.

03

What would it have taken to ship?

The bar is the 95% confidence interval on the difference. Move the lift and the sample size to see when a result becomes significant — and when it’s significant but still too small to be worth it.

observed lift−0.14 pp
users per group
95% CI
[−0.38, +0.09] pp
excludes 0?
no
lift ≥ +1 pp?
no
✕ Do not ship

The interval contains zero — the flows can’t be told apart.

no effect+1 pp ship bar
−3−2−10+1+2+3
difference in conversion, treatment − control (pp) · 95% CI
04

Not inconclusive — informative

The interval [−0.38, +0.09] ppcontains zero, so the two checkouts can’t be told apart. But it says more than that: its most favorable end is 11× smaller than the +1 pp worth shipping for.

With this much data, a real effect of any meaningful size would have shown up. Its absence is the answer.

+0.09 ppthe best case compatible with the data — against a bar of +1 pp. No plausible outcome justifies the redesign.
05

No early bump that fades

A pooled average can hide a novelty effect — a redesign that wins for a few days and then decays. Hover the chart to read each day of the 23-day window.

control · current checkout treatment · redesign
11.0%11.5%12.0%12.5%13.0%Jan 2Jan 6Jan 10Jan 14Jan 18Jan 22Jan 24controltreatment
view as table
DayControlTreatmentDiff (pp)Users
Jan 212.58%11.92%−0.665,631
Jan 311.36%11.41%+0.0413,025
Jan 412.14%11.68%−0.4712,929
Jan 512.32%11.44%−0.8912,748
Jan 611.51%12.36%+0.8513,168
Jan 712.01%11.64%−0.3713,013
Jan 811.89%12.10%+0.2113,231
Jan 911.93%11.84%−0.0913,064
Jan 1011.30%12.66%+1.3613,184
Jan 1111.77%11.53%−0.2413,183
Jan 1212.17%12.19%+0.0212,992
Jan 1311.66%11.14%−0.5212,893
Jan 1412.59%12.01%−0.5812,983
Jan 1512.06%11.32%−0.7413,081
Jan 1612.19%11.96%−0.2312,948
Jan 1712.33%12.67%+0.3412,976
Jan 1812.38%12.39%+0.0112,915
Jan 1912.02%11.61%−0.4112,968
Jan 2011.47%11.78%+0.3113,051
Jan 2112.55%11.61%−0.9413,139
Jan 2211.94%11.77%−0.1613,059
Jan 2312.65%12.09%−0.5613,164
Jan 2411.79%11.98%+0.197,345

The lines cross again and again with no sustained gap and no early treatment lead. The daily swings are sampling noise; the null result is stable over time, not an artifact of averaging.

06

My contribution

01Pre-registrationhypothesis, metric, ship rule fixed first
02Power analysissample size from a business MDE
03Integrity auditcontaminated users dropped, logged
04SRM testχ² check on the 50/50 split
05Inferencez-test + CI on the difference
06Novelty checkday-by-day stability across 23 days
07

Three takeaways

StatisticsThe interval is the headlineA p-value says “not significant.” The CI says how big the effect could plausibly be — and here, even its best case is worthless.
DesignDecide the rule before the dataThe +1 pp bar was set before any outcome was seen, so the conclusion can’t be bent toward a convenient answer.
BusinessA null is still a decisionWith 8.3× the required sample, this is not “inconclusive.” It is a clear no that saves building a redesign that doesn’t pay.

Framing note: the public dataset is technically a landing-page test (old vs. new page). Its structure — assigned group, variant shown, converted yes/no — is identical to a checkout test, so it’s framed that way for the narrative. The statistics are unchanged.