When There Is No Comparable Control: How Synthetic Controls Build One
August 2026 | Applied Economics | Causal Inference | ← Back to Blog
Not every causal question has an obvious control group. This post explains how synthetic control methods build a credible counterfactual out of a weighted combination of untreated units, and how to validate whether that counterfactual can actually be trusted.
When Difference-in-Differences Runs Out of Road
Difference-in-Differences depends on finding a comparable control group, a unit that would have followed a similar trajectory to the treated unit absent the intervention. In many settings this is available. In others, it is not.
A state changes its regulation on a marketplace. A city rolls out a new platform policy. A single large firm undergoes a merger. In cases like these, there is often exactly one treated unit and no single comparable unit that would make a credible DiD control. The treated market may simply be too different from every other market to have a credible single comparison group. Forcing a DiD comparison in this setting means picking a control that does not actually resemble the treated unit, which undermines the credibility of the entire exercise before any estimation even begins.
Building a Counterfactual Instead of Finding One
Synthetic control, developed by Abadie, Diamond, and Hainmueller, addresses this by constructing a counterfactual rather than searching for one. Instead of a single comparable unit, the method combines several untreated units, the donor pool, into a weighted average designed to resemble the treated unit as closely as possible before the intervention.
"A credible counterfactual does not always come from a single comparable market. Sometimes it comes from combining several markets that, together, behave like the treated market did before the intervention."
The weights are estimated, not chosen arbitrarily. They are selected, typically through a constrained optimization, to minimize the gap between the treated unit and the weighted combination of donor units across a set of predictor variables and pre-intervention outcomes. Each donor unit's weight is non-negative, and the weights sum to one, meaning the synthetic control is a genuine weighted average, not an extrapolation beyond the range of the observed data. This restriction, often referred to as the convex hull condition, is part of what gives the method its discipline. The synthetic control cannot claim to represent a combination that lies outside what the actual donor units show.
Validating the Fit: More Than Just a Visual Check
The credibility of a synthetic control estimate rests entirely on how well the synthetic unit tracks the treated unit before the intervention. There are three common ways to check this, each adding a different kind of evidence.
Visual inspection. Plot the treated unit and the synthetic control together over the pre-intervention period. If the two lines track closely before treatment, that is a first, informal signal of credibility. If they diverge meaningfully, the counterfactual should not be trusted regardless of what the post-treatment gap shows.
Pre-period RMSPE. Root mean squared prediction error quantifies the visual check numerically, measuring the average gap between the treated unit and the synthetic control across the pre-intervention period. A low pre-period RMSPE indicates a tight fit. This number also becomes the basis for the third check.
Placebo tests. The same synthetic control method is applied to every unit in the donor pool, as if each one had been treated instead of the actual treated unit. This produces a distribution of placebo "effects," estimated treatment effects for units that were never actually treated. If the real treated unit's effect is large relative to this distribution of placebo effects, that is evidence the estimated effect is not simply an artifact of noise or a coincidentally good pre-period fit. Units with poor pre-treatment fit are often excluded from the placebo comparison because an unreliable synthetic control cannot produce a meaningful placebo estimate.
Donor Pool Selection Matters as Much as the Weights
A common mistake is treating donor pool selection as a formality. In practice, which units are eligible to enter the donor pool is itself a judgment call with real consequences. Donor units should plausibly be unaffected by the treatment, similar enough in underlying structure to be a meaningful comparison, and observed over a long enough pre-period to allow a meaningful fit to be constructed. Including donor units that were themselves partially exposed to the treatment, directly or through spillover, contaminates the counterfactual in much the same way that using an already-treated unit as a DiD control does.
Where the Method Runs Into Real Limits
Synthetic control is not a universal fix for missing comparison groups, and it carries its own specific limitations worth knowing before relying on it.
The treated unit must be inside the donor pool's range, not an outlier. Because weights are constrained to be non-negative and sum to one, the synthetic control can only represent combinations that lie within the convex hull of the donor units. If the treated unit is structurally extreme relative to every available donor, for example, an unusually large market with no comparably sized peers, no combination of donors can replicate its pre-intervention trajectory well, and the method will simply fail to produce a credible fit rather than silently producing a misleading one.
A short pre-treatment period limits how much can be verified. The entire credibility case rests on demonstrating a close pre-period fit. With only a few pre-treatment periods available, it becomes hard to know whether a synthetic control that fits well is genuinely well-matched or just fits well by chance.
Inference is different from standard regression. Synthetic control does not naturally produce a standard error or a p-value in the way a regression coefficient does. Credibility instead comes from the placebo distribution described above, which is a valid but less familiar form of inference, and one that requires a donor pool large enough to generate a meaningful distribution of placebo effects in the first place.
None of these limitations make the method unusable. They do mean that a synthetic control estimate should be presented alongside its pre-period fit statistics and placebo results, not as a standalone number, since the credibility of the estimate is inseparable from how well those diagnostics hold up.
Why This Matters Beyond the Method
Synthetic control is particularly useful in exactly the settings where economic consulting, policy evaluation, and marketplace analysis most often operate: a single state passes a new law, one large platform changes its policy, one firm undergoes a merger. These are precisely the cases where a standard comparison group does not exist by default, and where a skeptical stakeholder will reasonably ask why the chosen comparison is credible. That is ultimately why synthetic control matters. The method is valuable not because it produces a number, but because it produces a defensible counterfactual.
Further Reading
Abadie, A., Diamond, A., and Hainmueller, J. (2010). Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California's Tobacco Control Program. Journal of the American Statistical Association, 105(490), 493-505.
Abadie, A. (2021). Using Synthetic Controls: Feasibility, Data Requirements, and Methodological Aspects. Journal of Economic Literature, 59(2), 391-425.