DECISION SYSTEMS TOOLKIT · EXAMPLE 03

Would You Trust the Store That "Worked"?

Eight store types test the same promotion. One looks like a winner. Another looks like a loser. The promotion actually did nothing anywhere. So what happened?

Worked example · store-level promotion testPrimary outcome: 4-pack add-on rate
Illustrative national chain · fictional data · designed for teachingPromotion ran for 4 weeks in September

The business question

COW Ice Cream has 400+ stores. Take-home 4-packs have better margins and can drive repeat purchases, but only about 12% of transactions include one.

Should we run a promotion to increase 4-pack purchases nationwide?

The promotion

Product
Take-home 4-pack only
$12
Regular price
Same product
$15
Campaign
4 weeks in September
28d

At checkout, the register prompts the cashier to offer the 4-pack. No external advertising, email, app push, or social campaign.

What changes — and what doesn't?

A clean experiment changes one thing on purpose and keeps the important alternatives as similar as possible.

Changes

Promotion ON vs OFF

The checkout offer is shown in promotion stores and not shown in control stores.

Held constant

Everything we can reasonably control

Same product, same four-week window, same script, same table card, same training, and no other promotions during the test.

Why test stores, not individual shoppers?

The checkout card is visible to everyone in a store. You cannot randomly show it to one shopper and hide it from the next without changing the experience. So the promotion is assigned at the store level.

We select 16 stores and make 8 matched pairs. Within each pair, one store is randomly assigned to promotion and the other to control. Stores are matched on store type, recent transaction volume, and recent 4-pack attach rate.

What are we measuring?

Primary metric

4-pack attach rate

The percentage of transactions that include at least one 4-pack.

4-pack attach rate = transactions with ≥1 4-pack ÷ total transactions
Baseline
About one in eight transactions
12%

One primary metric is decided before the results are seen.

Guardrails

Things we still watch

Single-scoop sales / transaction
Cannibalization
Average order value
Discount impact
Gross profit / transaction
Business economics
Daily transaction count
Unexpected traffic change

First, look at the overall result

Across all 16 stores, each group had about 16,000 transactions during the four-week test.

Control
12.15%

1,944 4-pack transactions out of 16,000.

Promotion
12.19%

1,951 4-pack transactions out of 16,000.

Seven more 4-packs.

16,000 transactions in each group. The difference is only 7 purchases.

The overall two-proportion test gives p = 0.91. There is no evidence here that the promotion changed the overall attach rate.

Then someone asks: "What about the store types?"

So the team looks at the eight store-type comparisons separately.

Store typeControlPromotionDifferencep-value
Downtown11.9%12.55%+130.55
Suburban12.3%11.7%−120.57
Mall12.0%12.7%+140.52
Roadside12.5%11.95%−110.61
College town11.8%12.55%+150.49
Coastal12.2%11.5%−140.50
Airport11.9%14.3%+480.024
Small town12.6%10.3%−460.022
Looks like a winner

Airport: +48

14.3% vs 11.9%. p = 0.024.

"Travelers want something to take home." The story sounds plausible.

Looks like a loser

Small town: −46

10.3% vs 12.6%. p = 0.022.

Almost the mirror image of the airport result.

The twist

The promotion actually did nothing.

In this teaching example, the true effect is zero in every store type.

The airport result and the small-town result are random fluctuations. The numbers are real. The apparent "winner" is not.

Eight chances to get fooled

A single test at the 5% significance level has a 5% chance of producing a false alarm when there is no real effect. But we didn't run one test. We looked at eight.

≈ 34%

chance of seeing at least one false "winner" across eight independent tests when the true effect is zero everywhere.

1 − (0.95)8 ≈ 0.34

Roughly one in three such experiments would produce at least one false alarm under these assumptions.

And seeing two is unusual — but not impossible

In this example, both Airport and Small town crossed the ordinary 0.05 threshold. The chance of seeing two or more false positives across eight independent tests is about 5.7%.

For readers who want the statistical detail
A simple Bonferroni adjustment would divide the 0.05 threshold by 8: 0.05 ÷ 8 = 0.00625. Neither 0.024 nor 0.022 clears that stricter threshold. This is a teaching example; a production experiment should choose an appropriate multiple-testing strategy in advance.

What should happen next?

Don't roll out to airports because Airport "won."

The result is a signal worth investigating, not proof that the promotion works better there.

Write down the question before the next test.

If Airport is genuinely important, make it a planned hypothesis and design a test specifically for it.

Keep the overall business result in view.

The nationwide test was essentially flat. A segment result should not quietly replace the primary decision metric.

Who should ask what?

The order matters. First, the person presenting the analysis should be able to answer the basic questions. Then the decision-maker can interrogate the evidence without needing to become a statistician.

For the person presenting the result

"What did we decide to measure before the test?"

One primary metric. Not the metric that happened to look best afterward.

"What else did we agree to keep an eye on?"

Margin, basket value, cannibalization, traffic, and execution issues are guardrails — not substitutes for the primary metric.

"How many store types did we examine?"

Show the full search space, not only the airport chart.

"Which findings were hypotheses before the test, and which were discovered afterward?"

That distinction changes how strongly a result should be presented.

For the decision-maker
1 · "How many places did you look?"Eight store types means eight chances for noise to look interesting.
2 · "Was Airport a hypothesis before we saw the result?"If not, treat it as a lead for the next test, not proof from this one.
3 · "What was the overall result?"The national result was essentially flat: +7 boxes, p = 0.91.
4 · "What would have made us roll out?"If there was no decision rule before the test, it is easy to invent one afterward.
5 · "What would make us stop or test again?"A good analysis should make the next decision clearer, not merely give the current slide a winner.

The slide shows you the winner. It doesn't show you how many stores you looked at before you found it.

One honest limitation

This is a teaching example, not a production experiment plan.

The promotion is assigned at the store level, while the simplified calculations above treat transactions as independent observations. A production analysis should account for the fact that transactions within the same store are related.

This test runs in September. A September result should not automatically be assumed to apply to July or another seasonal period.

Most importantly, this example deliberately sets the true promotion effect to zero everywhere so that the multiple-comparisons problem is easy to see. Real experiments are messier: genuine effects, store differences, execution differences, and random noise can all coexist.

A/B TEST TOOLKIT · WHICH TAB

Mean comparison
Proportion comparison
Sample size
Statistical power

Purchased-or-not at each store is a proportion, so that tab handles each individual test correctly on its own. What it can't do — yet — is warn you when you're running many of them at once and reading them side by side. That judgment call is still yours.

gracege.com/tools/ab-test-calculator →