When the Experiment Answers the Wrong Question

September 2026  |  Applied Economics  |  Experimentation  |  ← Back to Blog

A statistically significant result at small scale doesn't guarantee the same result at full rollout. This post explains why the conditions an experiment can change once the treatment reaches everyone, and what that means for how you launch.


The Result Was Real. The Question Was Narrower Than It Looked.

You ran the experiment. It worked. The result is statistically significant, well-powered, exactly what a clean test is supposed to produce. You still hesitate to ship it. That hesitation is worth taking seriously, because a valid experiment can still answer a narrower question than the one the business actually needs answered.

An experiment estimates a treatment effect under the market conditions that existed while it ran: a specific point in time, a specific competitive environment, a specific slice of users. Those conditions are not incidental background noise. They are part of what the estimate actually describes. Roll the same treatment out to everyone, and those conditions themselves change. The number you measured was correct.

"The result you measured was correct. It just described a different world than the one your full rollout lands in."


Four Ways the Conditions Change at Scale

Elasticity shifts. A price change tested on a small, randomly selected subset of users affects a correspondingly small share of aggregate demand. Roll the same change out to everyone, and the market-wide response can differ from what you observed in the experiment. The responsiveness you measured on the treated subset does not necessarily describe how demand responds when the entire market adjusts to the change.

Competitors react. A pricing or product change affecting a small, randomized slice of users is easy for a competitor to miss entirely, or to write off as noise. The same change rolled out to an entire market is a visible, public move, and it can trigger a competitive response that never existed during the experiment itself.

Supply adapts. Drivers, sellers, or suppliers on the other side of a marketplace respond to demand signals. A change affecting a random subset of buyers barely registers as a shift in aggregate demand. The same change affecting every buyer at once is a real, system-wide shift, and supply-side behavior, where sellers show up, what they price, how much they list, can adapt to it in ways the experiment never observed.

User composition shifts. The population exposed to a treatment during a small-scale test is not necessarily the same population exposed to it at full rollout. A promotion tested on existing active users may, at scale, pull in a very different type of new or marginal user, one whose response to the same treatment looks nothing like the original experimental population's.


Why This Isn't an Argument Against Experimentation

None of this means experiments are unreliable or that the result you measured was wrong. The result was correct, for the conditions it was measured under. The mistake is treating a valid small-scale estimate as if it were automatically a valid full-scale one, when the conditions that produced it were never designed to persist at scale in the first place.


What Actually Follows From This

A statistically significant result at small scale doesn't guarantee the same result at full scale, because the conditions themselves change once everyone is affected. This is why a holdout group matters even after you ship. It is the one remaining piece of the population still experiencing the original conditions, giving you a live comparison point rather than a single before-and-after snapshot. It is why post-launch monitoring isn't optional. The real answer to whether the effect holds at scale only becomes available once the rollout is actually large enough to change the conditions being measured. And it is why the strongest experiment designs build a ramp plan in from the start, not as a cautious afterthought, but as the actual mechanism for answering the one question a single small-scale experiment structurally cannot answer on its own: does this hold at scale.


Further Reading

Kohavi, R., Tang, D., and Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press.


← Back to Blog