Difference-in-Differences Is a Synthetic Control With Uniform Weights

Econometric Theory
Causal Inference
Author

Jared Greathouse

Published

September 12, 2026

This is a follow-up to my last excerpt, and (like that one) a preview of my forthcoming book on synthetic control methods.

“I think we may be In a different book, on a different page You said you are different but you’re the same” — Jhené Aiko, “stranger”


After the last post, where I showed that the arithmetic mean is really the solution to an optimization problem, somebody asked me a good question: why do people keep saying difference-in-differences (DID) is just synthetic control (SCM) with uniform weights? It is a fair thing to be confused about. The two are usually taught in different classes, with different notation, and DID in particular is almost never presented as an optimization problem at all — it is presented as a table with four cells and a subtraction, justified by parallel trends.

So let me start where everyone already is, with parallel trends. Then I will rebuild the whole thing from one axiom — the closest number to a sequence — until DID and synthetic control fall out as the same optimization problem asked under two different restrictions. At the end I will run it on Proposition 99 in Python, so you can watch the weights and the intercept do their jobs.

First principles: the closest number to a sequence

Everything below is built out of one idea, and it is the idea from the last post. Given a list of numbers, which single number best stands in for the whole list?

We have to say what “best” means, and the answer is distance: the best stand-in is the number sitting closest to all of them at once, where the cost of being far away is the squared difference. Written down, that is

\[ \operatorname{mean}(x_1,\dots,x_n) \defeq \operatorname*{argmin}_{c\,\in\,\mathbb{R}} \sum_{i=1}^n (x_i - c)^2 = \frac{1}{n}\sum_{i=1}^n x_i . \]

Read the middle piece out loud: “the value of \(c\) that makes the total squared distance as small as possible.” Solving it is one derivative, \(-2\sum_i (x_i-c)=0\), so \(c=\frac1n\sum_i x_i\), and since the second derivative \(2n\) is positive, that really is the bottom of the bowl.

Two special cases carry the rest of this post:

  • \(n=1\), the law of identity. The closest number to the one-item list \((x)\) is \(x\) itself. There is nothing to compromise between.
  • \(n=2\). The closest number to \((2,4)\) is \(3\).

And notice what the mean already is: a weighted sum, \(\sum_i w_i x_i\), with every weight equal to \(w_i = 1/n\). The weights are there from the very beginning, even in the humblest average. That will matter.

More than one sequence, and what to call them

Last time we had one list of numbers. Now we have several, and they need names before we can do anything with them.

A sequence here just means one unit’s path over time: a row of numbers. We observe \(N\) units over \(T\) periods, and we write

\[ y_{jt} \;=\; \text{the outcome of unit } j \text{ at time } t . \]

The first subscript says which unit, the second says when. Unit \(1\) is the treated unit, and its sequence is

\[ \mathbf{y}_1 = (y_{11},\, y_{12},\, \dots,\, y_{1T}). \]

Every other unit is a control. We collect their labels in the set \(\mathcal{N}_0\), which has \(N_0\) members, and each control \(j \in \mathcal{N}_0\) has a sequence of its very own,

\[ \mathbf{y}_j = (y_{j1},\, y_{j2},\, \dots,\, y_{jT}), \qquad j \in \mathcal{N}_0 . \]

So there is not one sequence now; there are \(N_0 + 1\) of them stacked on top of each other. Let us make that physical. Take three units and three periods, with unit \(1\) treated and units \(2\) and \(3\) as controls:

\(t=1\) \(t=2\) \(t=3\)
treated \(\mathbf{y}_1\) 16 19 18
control \(\mathbf{y}_2\) 10 12 11
control \(\mathbf{y}_3\) 12 16 15

Here \(N_0 = 2\), \(\mathcal{N}_0 = \{2,3\}\), and \(T = 3\). Reading the table: \(y_{12} = 19\) is the treated unit in period \(2\), and \(y_{31} = 12\) is control \(3\) in period \(1\). Each row is one of the sequences we just named.

Averaging down a column

There are two directions you can average a table like this, and only one of them is what we want. Averaging along a row would summarize one unit across time. Averaging down a column summarizes the controls at one moment, and that is the object DID is built from.

So freeze a time \(t\) and look only at the controls in that column. That column is a little list of numbers, \((y_{jt})_{j \in \mathcal{N}_0}\) — and a list of numbers is exactly what the axiom knows how to handle. Apply it:

\[ \bar y_{0t} = \operatorname*{argmin}_{c}\sum_{j\in\mathcal N_0}(y_{jt}-c)^2 = \frac{1}{N_0}\sum_{j\in\mathcal N_0} y_{jt}. \]

The bar means “averaged,” the \(0\) means “over the controls,” and the \(t\) says we did it at one particular time. Doing that to each column in turn gives \(\bar y_{01} = \frac{10+12}{2} = 11\), then \(\bar y_{02} = \frac{12+16}{2} = 14\), then \(\bar y_{03} = \frac{11+15}{2} = 13\). Those three numbers are themselves a sequence, so we name it like everything else:

\[ \bar{\mathbf{y}}_0 = (11,\, 14,\, 13). \]

This is the synthetic stand-in for the control group: not a real unit, just the column-by-column closest number.

The gap, and the average difference

We now have two sequences we care about, the treated path \(\mathbf{y}_1\) and the control average \(\bar{\mathbf{y}}_0\). Subtract them period by period and you get a third sequence — the one this whole post is really about:

\[ g_t \defeq y_{1t} - \bar y_{0t}, \qquad \mathbf{g} = (g_1,\, g_2,\, \dots,\, g_T). \]

On our numbers:

\(t=1\) \(t=2\) \(t=3\)
treated \(\mathbf{y}_1\) 16 19 18
control average \(\bar{\mathbf{y}}_0\) 11 14 13
gap \(\mathbf{g}\) 5 5 5

Look at what happened. Both lines move around — the controls go \(11 \to 14 \to 13\), up then down, and the treated unit follows them \(16 \to 19 \to 18\). Neither path is flat. But the distance between them never budges. It is \(5\), three times in a row. That is the sticky number from the top of this post, sitting in a table.

Now ask the axiom one more time, of this new sequence. What single number best stands in for the gaps? Same question, same answer — their mean:

\[ \beta^\ast = \operatorname*{argmin}_{\beta\,\in\,\mathbb{R}} \sum_{t=1}^{T_0}(g_t-\beta)^2 = \frac{1}{T_0}\sum_{t=1}^{T_0} g_t \;=\; \frac{5+5+5}{3} \;=\; 5 . \]

Here \(T_0\) counts the periods before any treatment; in the table all three are pre-treatment. So \(\beta^\ast\) is “the average difference” between the treated unit and its control average: one number, standing in for a whole sequence of differences.

Where day 1 comes in

Suppose you only ever got to see day 1. Then your gap sequence is the one-item list \((g_1) = (5)\), and by the law of identity its closest number is \(5\) itself. The day-1 difference is the estimate, because with one observation there is nothing to average.

Add more pre-treatment days and nothing changes in kind, only in degree: the closest number to the gap sequence generalizes from “the difference on day 1” to “the average difference across the days.” So \(\beta^\ast\) is not new machinery. It is the same closest-number question, asked of the difference sequence, and it collapses back to the day-1 difference when there is only one day.

That single number is the level — the “5.”

The weights were hiding in the control average

Here is the part that connects all of this to synthetic control, and it has been sitting in plain sight. Look again at how we built the control average in the first column:

\[ \bar y_{01} = 11 = \tfrac12 (10) + \tfrac12 (12) = w_2\, y_{21} + w_3\, y_{31}, \qquad w_2 = w_3 = \tfrac12 = \tfrac{1}{N_0}. \]

That is a weighted sum, exactly as the axiom warned. Taking “the average” was never a neutral act: it chose weights of \(1/N_0\) for every control, and it chose them before anyone looked at the data. Put general weights in their place and the gap becomes

\[ g_t = y_{1t} - \sum_{j\in\mathcal N_0} w_j\, y_{jt}, \qquad w_j \ge 0, \quad \sum_{j\in\mathcal N_0} w_j = 1 . \]

The two constraints keep the result an honest average: no negative shares, and the shares add to one. That set of allowed weights has a name, the simplex, and it is the same set the arithmetic mean has been living on all along.

Our three little sequences can already show why the choice matters. With equal weights the gap was \((5,5,5)\), flat as a board. Put all the weight on control \(2\) instead, \(w_2=1, w_3=0\), and the control path becomes \((10,12,11)\), so the gap becomes

\[ (16-10,\; 19-12,\; 18-11) = (6,\, 7,\, 7), \]

which is not flat. Put all the weight on control \(3\) and the gap is \((4,3,3)\) — also not flat. Hold onto that last one; we meet it again in a moment.

So “the gap is constant” was never a fact about the data alone. It was a fact about the weights we happened to use.

One problem, two knobs

Step back and look at what we have actually been doing all along. We are trying to reproduce the treated unit’s pre-treatment path out of two ingredients: a weighted combination of the controls, and a level shift. Written as one optimization problem, with both ingredients left as unknowns,

\[ \min_{\mathbf w \,\in\, \mathcal W,\;\; \beta \,\in\, \mathcal B}\; \sum_{t \le T_0}\Big(\, y_{1t} \;-\; \beta \;-\; \sum_{j\in\mathcal N_0} w_j\, y_{jt} \,\Big)^{2} . \]

That is the whole apparatus. There are exactly two knobs: the weights \(\mathbf w\), and the intercept \(\beta\). And the two methods we have been talking about are nothing more than two different decisions about which knob you are allowed to turn:

weights \(\mathcal W\) intercept \(\mathcal B\)
classic synthetic control the simplex: \(w_j \ge 0,\ \sum_j w_j = 1\) \(\{0\}\)
difference-in-differences the single point \(w_j = 1/N_0\) \(\mathbb{R}\)

Read that table twice, because it is the punchline. Each method frees exactly what the other one freezes. Synthetic control hunts over every allowed weighting but forbids itself a level shift. Difference-in-differences refuses to hunt at all, fixing the weights at the dead centre of the simplex, but lets the level go wherever it likes. Nothing else separates them: same objective, same data, same donors — different permissions.

Solving it, both ways

We can just solve the thing, on our own three sequences, and watch each restriction do its work.

Difference-in-differences first. The weights are nailed down at \(w_2=w_3=\tfrac12\), so \(\sum_j w_j y_{jt}\) is the control average \((11,14,13)\) and the only unknown left is \(\beta\). The problem collapses to the one we already solved, \(\min_{\beta} \sum_{t\le T_0} (g_t - \beta)^2\) with \(\mathbf g = (5,5,5)\), whose solution is \(\beta^\ast = 5\) and whose minimized total is \(0\). A perfect fit: \((11,14,13) + 5 = (16,19,18)\), exactly the treated path.

Now classic synthetic control. Here \(\beta\) is pinned at \(0\) and the weights are free on the simplex. With two donors the simplex is just a line segment, so there is only one unknown: put \(w_2 = w\) and \(w_3 = 1-w\) with \(0 \le w \le 1\). The synthetic path is \(w\,\mathbf y_2 + (1-w)\,\mathbf y_3\), which period by period is \((12 - 2w,\; 16 - 4w,\; 15 - 4w)\), and subtracting it from the treated path \((16,19,18)\) leaves residuals \((4 + 2w,\; 3 + 4w,\; 3 + 4w)\). Square them and add:

\[ \text{SSE}(w) = (4+2w)^2 + 2\,(3+4w)^2 = 36w^2 + 64w + 34 . \]

Its derivative is \(72w + 64\), positive everywhere on \([0,1]\) — the function climbs the whole way — so the smallest value sits at the left-hand end, \(w=0\). Synthetic control therefore puts all the weight on control \(3\), and its best achievable error is \(\text{SSE}(0) = 34\), with residuals \((4,3,3)\). There is that \((4,3,3)\) again, exactly as promised.

Why one wins so badly here

Line them up:

best fit it can reach total squared error
classic synthetic control (\(\beta = 0\)) \(w = (0,1)\) \(34\)
difference-in-differences (\(w = 1/N_0\)) \(\beta^\ast = 5\) \(0\)
neither knob freed — \(75\)

Synthetic control, allowed to search every weighting there is, cannot get its error below \(34\). Difference-in-differences, allowed to search nothing at all but handed one intercept, drives it to \(0\).

The reason is visible in the original table. At every period the treated unit sits above both donors: \(16\) over \((10,12)\), then \(19\) over \((12,16)\), then \(18\) over \((11,15)\). A weighted average with non-negative weights summing to one can never exceed the largest thing it is averaging, so no allowed \(\mathbf w\) can ever reach the treated path. The treated unit lies outside the donors’ convex hull, and without an intercept synthetic control has no way to climb to it. Every weighting it tries is too low; the best it can manage is to pick the highest donor and eat the remaining gap as error.

The intercept is exactly the tool for that job. It does not care about shape at all — it just lifts. And since the gap here was a constant \(5\), lifting by \(5\) was all the fitting that was ever required.

So the contest is not “which method is better.” It is which knob the data needs. When the treated unit sits outside the hull but drifts in parallel with it, you need the level, and DID hands it to you. When the treated unit sits inside the hull but drifts differently from the simple average, you need the weights, and synthetic control hands you those.

Two honest caveats before we go to real data. This example was built so the gap is exactly constant, which is to say it was built with DID’s assumption true by construction — that is why it wins so cleanly. Real data is rarely so obliging, and we are about to watch California fail this exact test. And nothing stops you from freeing both knobs at once, weights on the simplex and an intercept in \(\mathbb{R}\), which is a real and useful estimator in its own right. The point of the table is not that two cells exhaust the world. It is that the two most famous estimators in this literature are one optimization problem, asked under two different sets of permissions.

Proposition 99, with uniform weights and an intercept

Let me show the weights and the intercept in the flesh. I will use Abadie, Diamond and Hainmueller’s canonical Proposition 99 setup: California is treated by the 1989 anti-tobacco law, there are 38 control states, and we observe per-capita cigarette sales from 1970 to 2000. To follow along you’ll want the dataprep helper from my mlsynth package:

pip install -U git+https://github.com/jgreathouse9/mlsynth.git

We load the panel and hand it to dataprep, which returns the treated vector y, the donor_matrix of the 38 controls, and the number of pre-treatment periods.

import numpy as np
import pandas as pd
from mlsynth.utils.datautils import dataprep

base_url = "https://raw.githubusercontent.com/jgreathouse9/mlsynth/refs/heads/main/"
url = base_url + "basedata/smoking_data.csv"
data = pd.read_csv(url)

prepped = dataprep(
    data,
    data.columns[0],   # unit   : state
    data.columns[1],   # time   : year
    data.columns[2],   # outcome: cigsale
    data.columns[-1],  # treat  : Proposition 99
)

Now the estimator, exactly as the algebra prescribes. The uniform weights \(1/N_0\) are applied by averaging the donor matrix across its columns; the intercept \(\beta^\ast\) is the mean pre-treatment gap; the counterfactual is the shifted donor mean; and the ATT is the average post-treatment gap between California and that counterfactual.

y      = prepped["y"]
T0     = prepped["pre_periods"]
donors = prepped["donor_matrix"]
N0     = donors.shape[1]

# Uniform weights: every donor gets 1/N0. These are the SCM simplex weights,
# planted at the centroid.
w = np.full(N0, 1.0 / N0)
donor_mean = donors @ w                       # equivalently donors.mean(axis=1)

# The one free parameter: the intercept beta*.
beta = np.mean(y[:T0] - donor_mean[:T0])

yhat = donor_mean + beta                       # the DID counterfactual
ATT  = np.mean(y[T0:] - yhat[T0:])

print(f"weights sum to      {w.sum():.4f}, min weight {w.min():.4f}  (on the simplex)")
print(f"intercept  beta*  =  {beta:.4f}")
print(f"ATT              =  {ATT:.4f}")
weights sum to      1.0000, min weight 0.0263  (on the simplex)
intercept  beta*  =  -14.3590
ATT              =  -27.3491

The weights sum to one and are all positive, the intercept comes out to \(\beta^\ast \approx -14.36\) — California smoked about 14 fewer packs per capita than the average control state before the law — and the ATT is about \(-27.35\) packs per capita. Here is the picture.

Figure 2: Difference-in-Differences on Proposition 99: uniform donor weights plus an intercept.

The counterfactual does not hug California especially well before 1989 — it undershoots for the first few years and overshoots later. That misfit is the whole reason SCM exists, and it is parallel trends failing in front of you. The uniform weights force the counterfactual to be the average control state, shifted to California’s level; the intercept can only slide that line up and down, it cannot bend it. If California’s pre-period gap is not flat, no choice of \(\beta\) will make it flat, and the gap you see before 1989 is a preview of the bias you cannot see after it.

None of this hangs on my hand-rolled arithmetic. The diff-diff package — a small, dedicated DID estimator, a cousin to my own mlsynth — computes the same effect from the plain \(2 \times 2\) table, once we hand it the two indicators DID actually runs on: a treated-group flag (California, every year) and a post flag (every state, 1989 on).

from diff_diff import DifferenceInDifferences

# DID's two indicators, rebuilt from the panel: a treated-group flag
# (California in every year) and a post flag (every state, 1989 on).
did_df = data.copy()
treated_states = data.loc[data["Proposition 99"] == 1, "state"].unique()
post_years     = data.loc[data["Proposition 99"] == 1, "year"].unique()
did_df["treated"] = data["state"].isin(treated_states).astype(int)
did_df["post"]    = data["year"].isin(post_years).astype(int)

result = DifferenceInDifferences().fit(
    did_df, outcome="cigsale", treatment="treated", post="post"
)
print(result)
DiDResults(ATT=-27.3491***, SE=4.5522, p=0.0000)

Same ATT — about \(-27.35\) — now carrying a standard error and a \(p\)-value, from a routine that never heard of a weight vector or a simplex. It agrees because it is the same estimator under another name: the hand-rolled intercept \(\beta^\ast\) is the average pre-period gap, and the \(2 \times 2\) DID is that same gap differenced across the treatment date.

So what?

Now, you may be saying, “Hey Jared, this is a lot of algebra to conclude that a mean is a mean.” But the payoff is a clean way to see what you are actually assuming when you reach for DID, and why the choice between DID and SCM is a choice about which knob you are allowed to turn.

Both estimators answer the same question — what would California have done without Prop 99? — by averaging control states. DID commits to equal weights before it looks at the data; SCM lets the pre-treatment fit choose the weights. Neither fabricates a “fake” California or grafts together a “Frankenstein” unit, any more than the national average income fabricates a fake person. They both summarize real control units with weights that live on the same simplex. The difference is only whether those weights are handed to you as an axiom or estimated as a decision.

And the reason that decision matters shows up the moment you write untreated outcomes as a factor model, \(y_{jt}(0) = a_j + \mathbf{b}_j^\top \mathbf{f}_t + u_{jt}\). Averaging the donor pool averages its factor loadings \(\bar{\mathbf{b}}_0\), and — as I show in the book — the DID counterfactual is unbiased exactly when the treated unit’s loading equals that average, \(\mathbf{b}_1 = \bar{\mathbf{b}}_0\). That is parallel trends restated one more time: the leftover gap is \((\mathbf{b}_1 - \bar{\mathbf{b}}_0)^\top (\mathbf{f}_t - \bar{\mathbf{f}})\), a term that moves with \(t\) whenever the treated loading differs from the donor average — a gap no intercept can flatten. Uniform weights only reach California’s factor exposure if California happens to sit at the donor pool’s center of gravity. When it does not — and a state with California’s tobacco culture usually does not — you want the freedom to move the weights off the centroid, toward the donors that actually share its exposure. That freedom is synthetic control, and DID is the center of it you get by refusing to use it.

So the answer to the question I was asked is: DID is a synthetic control whose weight vector you filled in yourself, with \(1/N_0\) in every slot, before the data had a chance to weigh in — and which, having given up the weights, keeps the intercept that classic synthetic control throws away. Everything else — the loss, the simplex, the gap, the counterfactual, and parallel trends itself — is shared. You are, as ever, never not averaging. The only question is who gets to pick the weights.

What Is An Average?

Econometric Theory
Statistics

What is a Synthetic Control?

Econometrics
Causal Inference

Synthetic Controls With More Than One Outcome

Causal Inference
Econometrics

Forward Selected Synthetic Control

Machine Learning
Econometrics
No matching items