DECISION SYSTEMS TOOLKIT · EXAMPLE 02

Was 8 more sign-ups a real result?

A 14-day trial versus a 30-day trial. The honest answer took three questions, not one.

Worked example · proportion comparison + sample size Sign-ups counted on one fixed calendar date
Both groups measured on the same date, after every trial had ended — otherwise the 30-day version simply had 16 more days. 600 people per group

14-day trial

43

signed up, out of 600  ·  7.2%

30-day trial

51

signed up, out of 600  ·  8.5%

The 30-day version won by 8 people. Written as a percentage, that's an 18.6% lift — a number that moves through a team fast. So: is 8 people a real difference, or is it luck?

8 sits inside a much wider range of "could be anything" The data is consistent with the 30-day version being worse by about 10 people, or better by nearly 27. An 8-person edge doesn't rule either one out.

The part worth sitting with: this test was never able to answer the question. With 600 people per group, the smallest gap it could reliably spot was about 28 people. The team was hoping to find 8.

That limit wasn't set by the result. It was set the day someone chose to run 600 people per group.

Catching a gap this small takes roughly 6,400 people per group — about eleven times what they ran. Small differences are expensive to prove, and the cost climbs far faster than most people expect.

The statistics behind those numbers
Conversion, 14-day vs 30-day7.17% vs 8.50%
Difference+1.33 pp  (+18.6% relative)  ≈ 8 people
p-value (pooled z-test)0.39
95% CI on the difference (Newcombe)−1.7 to +4.4 pp  ≈ −10 to +27 people
Minimum detectable effect at n=600≈ 66% relative lift (7.17% → 11.9%)  ≈ 28 people
n required to detect +18.6%≈ 6,400 per group

The people-count range is the percentage-point confidence interval converted to raw counts (pp ÷ 100 × 600 per group), which is exact when both groups are the same size. Assumptions: 95% significance, 80% power, two-sided.

One design note that matters more than the statistics: a 30-day trial reaches its end 16 days later than a 14-day trial. If sign-ups were counted at each group's own trial end, the longer arm would have had more elapsed time, and the comparison would measure duration rather than trial length. Both arms here are measured on a single calendar date, late enough that every trial in both arms has finished. Full source code is on GitHub.

A/B TEST TOOLKIT · WHICH TABS

Mean comparison
Proportion comparison
Sample size
Statistical power

Signed-up-or-not is a proportion, so that tab handles the first question. The second one — could this test have seen it? — comes from Sample size, and it's worth running before the test, not after.

gracege.com/tools/ab-test-calculator →

Figures in this example are illustrative and generated for teaching purposes. They do not describe any real company.
This tool is provided for reference only and does not constitute professional statistical or business advice.