Choosing the Right Experiment for Your Marketplace

September 2026  |  Applied Economics  |  Experimentation  |  ← Back to Blog

Deciding to run an experiment is the easy part. This post walks through how the structure of a marketplace, specifically where interference actually lives, should determine the randomization design, and what to do when no randomized design provides enough power to support a meaningful decision.


The Decision Comes First, the Design Is Harder

You've decided to run an experiment. That decision was the easy part. The harder part is choosing an experimental design that fits the marketplace you're actually studying. Marketplaces almost always have interference, one person's treatment can affect another person's outcome. Where you randomize matters because of this. The level of randomization determines whether your treatment and control groups stay meaningfully distinct, or just look clean on the surface.

"Containing interference is different from pretending it doesn't exist."


The First Question: Can You Contain the Interference

Clustering at the region, network, or region-time level can reduce interference substantially. None of these designs necessarily eliminate it. Some crossover between conditions can still happen, a unit assigned to control ends up interacting with someone in treatment anyway. The goal of clustering isn't zero contamination, it's shrinking contamination to a level where the treatment contrast is still interpretable.

This connects to a broader idea in the causal inference literature sometimes called partial interference. The standard setup behind many experiments, often expressed through SUTVA, the Stable Unit Treatment Value Assumption, assumes that one unit's outcome does not depend on the treatment assigned to anyone else. Partial interference relaxes that, it allows interference to exist within a cluster but assumes it's negligible across clusters. That's a weaker, more realistic assumption for a marketplace, and it's what actually makes cluster-based experiments identifiable in practice.


The Second Question: Where Does the Interaction Actually Happen

If interactions stay mostly within a region, region-level randomization is usually enough. A driver and rider matching pool that rarely crosses city lines doesn't need anything more complicated than randomizing at the city level.

If those interactions shift with supply and demand over time, region-level randomization alone isn't enough. Consider a rideshare marketplace: drivers and riders interact within a geographic market, but that market's equilibrium changes continuously throughout the day, morning commute, midday dip, evening surge. A region assigned to treatment in the morning and control in the evening would blur exactly the contrast you're trying to measure. This is why region-time randomization, switchback design, exists as its own category, the relevant unit of interference isn't the region alone, it's the region at a specific point in time.

If interaction follows relationships rather than geography, a social network where treated users affect their friends, or a marketplace where buyers and sellers have ongoing, repeat relationships, network-level randomization is what actually captures it. Two users on opposite sides of the country can still be tightly linked if one is a regular buyer from the other's shop. Geography tells you nothing useful about that interference structure, only the relationship graph does.


The Practical Constraint: Statistical Power

Sometimes clustering leaves you with too few independent units to detect an effect reliably, too few regions, too few distinct networks. Statistical power in a clustered design depends heavily on the number of clusters, not just the number of individual units inside them. Ten thousand users spread across only four regions does not give you the same inferential information as ten thousand independently randomized users. For the treatment contrast, you effectively have only four independent clusters.

In that case, forcing a randomized design doesn't solve the problem. It may simply produce an estimate with too much uncertainty to support a meaningful decision, a confidence interval wide enough to be useless, even if the point estimate looks encouraging.


When Randomization Genuinely Isn't Feasible

Depending on the setting, a quasi-experimental approach such as difference-in-differences or synthetic control may provide a more credible way to use the markets and time periods you actually have. But that comes with its own identification assumptions, and they aren't free.

Difference-in-differences requires parallel trends, that the treated and control markets would have moved together absent the intervention, an assumption that has to be argued for, not assumed by default. Synthetic control requires a good pre-period fit, that a weighted combination of untreated markets can genuinely track the treated market's pre-intervention trajectory closely enough to trust its projection forward. Neither method is a free substitute for randomization, they're a different set of assumptions, chosen because they're more defensible than trying to force a randomized design without enough power to support it.


The Real Decision Rule

Choosing to run an experiment is just the beginning. Designing one that fits how your marketplace actually works, and what your data can actually support, is where the real work starts. The correct experimental unit isn't the customer, the seller, the region, or the network by convention or by default, it's whichever one actually contains where the interference lives. When the unit doesn't match where the interference actually lives, a statistically clean-looking experiment can still end up measuring something other than what it appears to measure.


Further Reading

Eckles, D., Karrer, B., and Bakshy, E. (2017). Design and Analysis of Experiments in Networks: Reducing Bias from Interference. Journal of Causal Inference, 5(1), 1-23.

Abadie, A. (2021). Using Synthetic Controls: Feasibility, Data Requirements, and Methodological Aspects. Journal of Economic Literature, 59(2), 391-425.


← Back to Blog