The business question
COW Ice Cream has 400+ stores. Take-home 4-packs have better margins and can drive repeat purchases, but only about 12% of transactions include one.
The promotion
At checkout, the register prompts the cashier to offer the 4-pack. No external advertising, email, app push, or social campaign.
What changes — and what doesn't?
A clean experiment changes one thing on purpose and keeps the important alternatives as similar as possible.
Promotion ON vs OFF
The checkout offer is shown in promotion stores and not shown in control stores.
Everything we can reasonably control
Same product, same four-week window, same script, same table card, same training, and no other promotions during the test.
Why test stores, not individual shoppers?
The checkout card is visible to everyone in a store. You cannot randomly show it to one shopper and hide it from the next without changing the experience. So the promotion is assigned at the store level.
We select 16 stores and make 8 matched pairs. Within each pair, one store is randomly assigned to promotion and the other to control. Stores are matched on store type, recent transaction volume, and recent 4-pack attach rate.
What are we measuring?
4-pack attach rate
The percentage of transactions that include at least one 4-pack.
One primary metric is decided before the results are seen.
Things we still watch
First, look at the overall result
Across all 16 stores, each group had about 16,000 transactions during the four-week test.
1,944 4-pack transactions out of 16,000.
1,951 4-pack transactions out of 16,000.
Seven more 4-packs.
16,000 transactions in each group. The difference is only 7 purchases.
The overall two-proportion test gives p = 0.91. There is no evidence here that the promotion changed the overall attach rate.
Then someone asks: "What about the store types?"
So the team looks at the eight store-type comparisons separately.
| Store type | Control | Promotion | Difference | p-value |
|---|---|---|---|---|
| Downtown | 11.9% | 12.55% | +13 | 0.55 |
| Suburban | 12.3% | 11.7% | −12 | 0.57 |
| Mall | 12.0% | 12.7% | +14 | 0.52 |
| Roadside | 12.5% | 11.95% | −11 | 0.61 |
| College town | 11.8% | 12.55% | +15 | 0.49 |
| Coastal | 12.2% | 11.5% | −14 | 0.50 |
| Airport | 11.9% | 14.3% | +48 | 0.024 |
| Small town | 12.6% | 10.3% | −46 | 0.022 |
Airport: +48
14.3% vs 11.9%. p = 0.024.
"Travelers want something to take home." The story sounds plausible.
Small town: −46
10.3% vs 12.6%. p = 0.022.
Almost the mirror image of the airport result.
The promotion actually did nothing.
In this teaching example, the true effect is zero in every store type.
The airport result and the small-town result are random fluctuations. The numbers are real. The apparent "winner" is not.
Eight chances to get fooled
A single test at the 5% significance level has a 5% chance of producing a false alarm when there is no real effect. But we didn't run one test. We looked at eight.
chance of seeing at least one false "winner" across eight independent tests when the true effect is zero everywhere.
Roughly one in three such experiments would produce at least one false alarm under these assumptions.
And seeing two is unusual — but not impossible
In this example, both Airport and Small town crossed the ordinary 0.05 threshold. The chance of seeing two or more false positives across eight independent tests is about 5.7%.
For readers who want the statistical detail
What should happen next?
Don't roll out to airports because Airport "won."
The result is a signal worth investigating, not proof that the promotion works better there.
Write down the question before the next test.
If Airport is genuinely important, make it a planned hypothesis and design a test specifically for it.
Keep the overall business result in view.
The nationwide test was essentially flat. A segment result should not quietly replace the primary decision metric.
Who should ask what?
The order matters. First, the person presenting the analysis should be able to answer the basic questions. Then the decision-maker can interrogate the evidence without needing to become a statistician.
"What did we decide to measure before the test?"
One primary metric. Not the metric that happened to look best afterward.
"What else did we agree to keep an eye on?"
Margin, basket value, cannibalization, traffic, and execution issues are guardrails — not substitutes for the primary metric.
"How many store types did we examine?"
Show the full search space, not only the airport chart.
"Which findings were hypotheses before the test, and which were discovered afterward?"
That distinction changes how strongly a result should be presented.
The slide shows you the winner. It doesn't show you how many stores you looked at before you found it.
One honest limitation
This is a teaching example, not a production experiment plan.
The promotion is assigned at the store level, while the simplified calculations above treat transactions as independent observations. A production analysis should account for the fact that transactions within the same store are related.
This test runs in September. A September result should not automatically be assumed to apply to July or another seasonal period.
Most importantly, this example deliberately sets the true promotion effect to zero everywhere so that the multiple-comparisons problem is easy to see. Real experiments are messier: genuine effects, store differences, execution differences, and random noise can all coexist.