Why A/B Tests Break in Marketplaces, and How Switchback Experiments Fix It

August 2026  |  Applied Economics  |  Causal Inference  |  ← Back to Blog

A simple A/B test assumes one user's treatment does not affect another user's outcome. In a marketplace with shared resources, that assumption can fail. This post explains the violation behind that failure, how switchback experiments address it, and where the fix still leaves gaps worth knowing about.


When Treating One User Changes Another User's Outcome

Say you run a surge pricing experiment on one group of users. The treatment does not just change their behavior. It changes what happens to the control group too. In a two-sided marketplace, supply is shared. If a price change pulls supply toward the treated group, that supply is no longer available for the control group. The control group's experience changes, not because of anything assigned to them, but because of something assigned to someone else.


The Assumption That Breaks: SUTVA

This is a violation of SUTVA, the Stable Unit Treatment Value Assumption. SUTVA says a unit's outcome should depend only on the treatment it receives, not on what treatment anyone else receives. A simple A/B test relies on the assumption that SUTVA holds.

"The fix is not a bigger sample or a longer experiment. It is a different unit of randomization."

In a marketplace with shared resources, SUTVA can fail by design, not by accident. Riders share a pool of drivers. Buyers share a pool of sellers. Diners share a pool of delivery couriers. Whenever two experimental units draw from the same underlying supply, treating one changes what is available to the other. More sample size does not fix that. A longer run time does not fix it either. The problem is structural interference between units, not a precision problem that more data would solve.


The Fix: Randomize the Market, Not the User

A switchback experiment solves this by changing the unit of randomization. Instead of assigning treatment or control to individual users, you assign it to a time-region window, a specific geography during a specific time period. Everyone interacting within that window, on both sides of the market, experiences the same condition, either treatment or control. You are not randomizing across people who share a supply pool anymore. You are randomizing across the windows themselves. The interference that would otherwise contaminate the comparison stays inside a single window instead of leaking between a treated group and a control group operating in the same place at the same time.


Contains the Interference, Does Not Fully Eliminate It

This is the caveat worth knowing before treating switchback as a complete fix. A switchback design contains interference within each window, but it does not eliminate interference at the boundaries between windows. Supply can still physically move. A driver near the edge of a surging region can relocate toward it, which means a neighboring region assigned to control can lose supply because of a treatment condition next door, not because of anything happening inside its own window. Switchback does not eliminate interference. It contains it within the experimental design.


Choosing the Window Size: A Real Tradeoff

Window size is not a free parameter. Smaller windows give you more assignment opportunities and potentially more statistical power. But they also increase carryover, a treatment effect from one window can bleed into the next before the system resets. Larger windows reduce carryover, but they leave you with fewer assignment opportunities and potentially less power. There's no single right size. It depends on how fast the system responds and how much carryover you can live with given your sample.

A practical way to handle this is to build in washout periods, short buffer windows between a treatment period and a control period, specifically to let carryover effects from the prior assignment dissipate before the next period's outcomes are measured. Historical data on how quickly the system responds to a similar change can help estimate how long that carryover actually lasts, which in turn informs how short a window can safely be.


Analyzing the Results

Switchback data has a panel structure, the same regions observed repeatedly over time. One natural analysis approach is a panel regression with region and time fixed effects, with inference accounting for the randomization scheme and the repeated observations within regions over time. That's functionally related to the same underlying logic as difference-in-differences, applied to a panel where treatment assignment switches back and forth within each region over time rather than happening once. The dependence structure in the data follows directly from what was actually randomized. The model needs to respect that structure rather than treating individual observations as independent.


Why This Matters Beyond the Design

A simple A/B test can still run and produce a number in a marketplace. That number just is not credible once outcomes spill across units. Switchback experiments are not a workaround for a minor inconvenience. They are the appropriate response to a genuine structural feature of two-sided markets. The two sides depend on each other, and any credible experiment has to account for that dependence rather than assume it away.


Further Reading

Bojinov, I., Simchi-Levi, D., and Zhao, J. (2023). Design and Analysis of Switchback Experiments. Management Science, 69(7), 3759-3777.

Rubin, D. B. (1980). Randomization Analysis of Experimental Data: The Fisher Randomization Test Comment. Journal of the American Statistical Association, 75(371), 591-593.


← Back to Blog