Difference-in-Differences Is a Synthetic Control With Uniform Weights

Published:

This is a follow-up to my last excerpt, and (like that one) a preview of my forthcoming book on synthetic control methods. ::: {.epigraph} > "I think we may be > In a different book, on a different page > You said you are different but you're the same" > — [*Jhené Aiko*](https://genius.com/Jhene-aiko-stranger-lyrics), "stranger" ::: --- After the [last post](averagesexcp.html), where I showed that the arithmetic mean is really the solution to an optimization problem, somebody asked me a good question: *why do people keep saying difference-in-differences (DID) is just synthetic control (SCM) with uniform weights?* It is a fair thing to be confused about. The two are usually taught in different classes, with different notation, and DID in particular is almost never presented as an optimization problem at all — it is presented as a table with four cells and a subtraction, justified by parallel trends. So let me start where everyone already is, with parallel trends. Then I will rebuild the whole thing from one axiom — the closest number to a sequence — until DID and synthetic control fall out as the *same* optimization problem asked under two different restrictions. At the end I will run it on Proposition 99 in Python, so you can watch the weights and the intercept do their jobs. ## Start with parallel trends Almost everyone who has met DID has met parallel trends. The setup is two groups, one treated and one not, observed before and after some intervention. DID takes the *change* in the treated group's outcome and subtracts the *change* in the control group's outcome. If, absent treatment, the two groups' average outcomes would have moved in parallel, then the control group's change stands in for the change the treated group would have had anyway, and the leftover — the difference of the two differences — is the treatment effect. What is parallel trends really saying, though? Set the formula aside for a second. Suppose that on the very first day we look, our treated unit sits about 5 units above the control average. Not because of treatment — nothing has happened yet — but simply because that is where it starts: maybe it is a richer state, a bigger market, a heavier-smoking population. Parallel trends is the claim that this gap of 5 is *sticky*. On the next day, absent treatment, the treated unit is again about 5 above the controls, in expectation. Ten periods later, still about 5. The two lines may climb, fall, and zig-zag with the business cycle however they like — but the *distance between them* stays parked at 5, period after period, for as long as no treatment intervenes. Sit with what that does and does not allow. It does not ask the treated and control groups to be at the same level — they are 5 apart the entire time, and that is fine. It asks only that the 5 never moves. So the treatment, when it finally arrives, is whatever makes the gap *depart* from 5. If the gap opens up to 11 after the policy, the treatment did the extra 6; the 5 was always going to be there. That one sticky number is the whole ballgame — and, as we are about to see, it is exactly the number the optimization solves for. @fig-pta-cartoon draws both worlds. On the left, the two lines wander all over — up, down, with the cycle — but the distance between them never leaves 5. That is parallel trends. On the right, the two start 5 apart and then drift; the gap grows to 15 for reasons that have nothing to do with any treatment. Both panels are the untreated world, before anyone has been treated. DID is only trustworthy in the world on the left. ```{python} #| label: fig-pta-cartoon #| fig-cap: "What parallel trends assumes, absent treatment: the gap between treated and control is a single sticky number. Left, it stays at 5 (parallel trends holds); right, it drifts off 5 (parallel trends fails)." #| echo: False import numpy as np import matplotlib.pyplot as plt t = np.arange(0, 20) common = 10 + 3 * np.sin(t / 3.0) # a shared, wandering trend control = common treated_hold = control + 5 # constant gap of 5 treated_fail = control + 5 + 0.55 * t # gap drifts away from 5 fig, axes = plt.subplots(1, 2, figsize=(12, 4.5), sharey=True) for ax, treated, title in [ (axes[0], treated_hold, "Parallel trends holds: the gap stays at 5"), (axes[1], treated_fail, "Parallel trends fails: the gap drifts"), ]: ax.plot(t, control, color="gray", lw=2, label="control average") ax.plot(t, treated, color="black", lw=2, label="treated (absent treatment)") for xi in (0, -1): # annotate the gap at the ends c, d = control[xi], treated[xi] ax.annotate("", xy=(t[xi], d), xytext=(t[xi], c), arrowprops=dict(arrowstyle="<->", color="crimson", lw=1.5)) ax.text(t[xi] + 0.4, (c + d) / 2, f"{d - c:.0f}", color="crimson", fontsize=12, va="center") ax.set_title(title, fontsize=11) ax.set_xlabel("Time") axes[0].set_ylabel("Untreated outcome") axes[0].legend(loc="upper left", frameon=False, fontsize=9) fig.tight_layout() plt.show() ``` Write it for a single treated unit and a donor pool $\mathcal{N}_0$ of $N_0$ never-treated units. Let $y_{1t}$ be the treated unit and $\bar{y}_{0t}$ the control-group average at time $t$. Parallel trends says the expected gap between them is one and the same constant over time, treatment aside: $$ \mathbb{E}\!\left[\, y_{1t}(0) - \bar{y}_{0t}(0) \,\right] = \beta, \qquad \forall \, t, $$ the same $\beta$ before *and* after the intervention. This is the assumption that the average untreated outcome evolves in parallel for the two groups; everything DID reports is downstream of it. Here is the thing I want to flag before we go further, because it is the crux of the question I was asked. In the way DID is almost always taught and applied, this is a **large-$N$** statement. You have many controls, and $\bar{y}_{0t}$ is not really "38 states averaged with weight $1/38$" in anyone's mind — it is a sample estimate of a population expectation, $\mathbb{E}[y_t(0) \mid D = 0]$, the untreated mean in the control population. As the control group grows, the sample average converges to that expectation, and nobody stops to notice that an expectation is itself a weighting. In the [Baker, Callaway, Cunningham, Goodman-Bacon and Sant'Anna (2026)](https://doi.org/10.1257/jel.20251650) practitioner's guide the point is explicit from the first equations: the DID estimand is written with a *weighted* expectation $\mathbb{E}_\omega[\cdot]$, and the weights, they note, enter "as part of the definition of the causal parameter" — before you have estimated anything. Large-$N$ asymptotics just let you fold them into the word "expectation" and stop looking. We do not usually need to treat the weights as objects. That does not mean they are not there. ## First principles: the closest number to a sequence Everything below is built out of one idea, and it is the idea from the [last post](averagesexcp.html). Given a list of numbers, which *single* number best stands in for the whole list? We have to say what "best" means, and the answer is *distance*: the best stand-in is the number sitting closest to all of them at once, where the cost of being far away is the squared difference. Written down, that is $$ \operatorname{mean}(x_1,\dots,x_n) \defeq \operatorname*{argmin}_{c\,\in\,\mathbb{R}} \sum_{i=1}^n (x_i - c)^2 = \frac{1}{n}\sum_{i=1}^n x_i . $$ Read the middle piece out loud: "the value of $c$ that makes the total squared distance as small as possible." Solving it is one derivative, $-2\sum_i (x_i-c)=0$, so $c=\frac1n\sum_i x_i$, and since the second derivative $2n$ is positive, that really is the bottom of the bowl. Two special cases carry the rest of this post: - $n=1$, the law of identity. The closest number to the one-item list $(x)$ is $x$ itself. There is nothing to compromise between. - $n=2$. The closest number to $(2,4)$ is $3$. And notice what the mean already is: a *weighted sum*, $\sum_i w_i x_i$, with every weight equal to $w_i = 1/n$. The weights are there from the very beginning, even in the humblest average. That will matter. ## More than one sequence, and what to call them Last time we had one list of numbers. Now we have several, and they need names before we can do anything with them. A *sequence* here just means one unit's path over time: a row of numbers. We observe $N$ units over $T$ periods, and we write $$ y_{jt} \;=\; \text{the outcome of unit } j \text{ at time } t . $$ The first subscript says *which unit*, the second says *when*. Unit $1$ is the treated unit, and its sequence is $$ \mathbf{y}_1 = (y_{11},\, y_{12},\, \dots,\, y_{1T}). $$ Every other unit is a control. We collect their labels in the set $\mathcal{N}_0$, which has $N_0$ members, and each control $j \in \mathcal{N}_0$ has a sequence of its very own, $$ \mathbf{y}_j = (y_{j1},\, y_{j2},\, \dots,\, y_{jT}), \qquad j \in \mathcal{N}_0 . $$ So there is not one sequence now; there are $N_0 + 1$ of them stacked on top of each other. Let us make that physical. Take three units and three periods, with unit $1$ treated and units $2$ and $3$ as controls: | | $t=1$ | $t=2$ | $t=3$ | |:--|:--:|:--:|:--:| | treated $\mathbf{y}_1$ | 16 | 19 | 18 | | control $\mathbf{y}_2$ | 10 | 12 | 11 | | control $\mathbf{y}_3$ | 12 | 16 | 15 | Here $N_0 = 2$, $\mathcal{N}_0 = \{2,3\}$, and $T = 3$. Reading the table: $y_{12} = 19$ is the treated unit in period $2$, and $y_{31} = 12$ is control $3$ in period $1$. Each *row* is one of the sequences we just named. ### Averaging down a column There are two directions you can average a table like this, and only one of them is what we want. Averaging *along a row* would summarize one unit across time. Averaging *down a column* summarizes the controls at one moment, and that is the object DID is built from. So freeze a time $t$ and look only at the controls in that column. That column is a little list of numbers, $(y_{jt})_{j \in \mathcal{N}_0}$ — and a list of numbers is exactly what the axiom knows how to handle. Apply it: $$ \bar y_{0t} = \operatorname*{argmin}_{c}\sum_{j\in\mathcal N_0}(y_{jt}-c)^2 = \frac{1}{N_0}\sum_{j\in\mathcal N_0} y_{jt}. $$ The bar means "averaged," the $0$ means "over the controls," and the $t$ says we did it at one particular time. Doing that to each column in turn gives $\bar y_{01} = \frac{10+12}{2} = 11$, then $\bar y_{02} = \frac{12+16}{2} = 14$, then $\bar y_{03} = \frac{11+15}{2} = 13$. Those three numbers are themselves a sequence, so we name it like everything else: $$ \bar{\mathbf{y}}_0 = (11,\, 14,\, 13). $$ This is the *synthetic* stand-in for the control group: not a real unit, just the column-by-column closest number. ## The gap, and the average difference We now have two sequences we care about, the treated path $\mathbf{y}_1$ and the control average $\bar{\mathbf{y}}_0$. Subtract them period by period and you get a third sequence — the one this whole post is really about: $$ g_t \defeq y_{1t} - \bar y_{0t}, \qquad \mathbf{g} = (g_1,\, g_2,\, \dots,\, g_T). $$ On our numbers: | | $t=1$ | $t=2$ | $t=3$ | |:--|:--:|:--:|:--:| | treated $\mathbf{y}_1$ | 16 | 19 | 18 | | control average $\bar{\mathbf{y}}_0$ | 11 | 14 | 13 | | gap $\mathbf{g}$ | 5 | 5 | 5 | Look at what happened. Both lines move around — the controls go $11 \to 14 \to 13$, up then down, and the treated unit follows them $16 \to 19 \to 18$. Neither path is flat. But the *distance between them* never budges. It is $5$, three times in a row. That is the sticky number from the top of this post, sitting in a table. Now ask the axiom one more time, of this new sequence. What single number best stands in for the gaps? Same question, same answer — their mean: $$ \beta^\ast = \operatorname*{argmin}_{\beta\,\in\,\mathbb{R}} \sum_{t=1}^{T_0}(g_t-\beta)^2 = \frac{1}{T_0}\sum_{t=1}^{T_0} g_t \;=\; \frac{5+5+5}{3} \;=\; 5 . $$ Here $T_0$ counts the periods *before* any treatment; in the table all three are pre-treatment. So $\beta^\ast$ is "the average difference" between the treated unit and its control average: one number, standing in for a whole sequence of differences. ### Where day 1 comes in Suppose you only ever got to see day 1. Then your gap sequence is the one-item list $(g_1) = (5)$, and by the law of identity its closest number is $5$ itself. The day-1 difference *is* the estimate, because with one observation there is nothing to average. Add more pre-treatment days and nothing changes in kind, only in degree: the closest number to the gap sequence generalizes from "the difference on day 1" to "the average difference across the days." So $\beta^\ast$ is not new machinery. It is the same closest-number question, asked of the difference sequence, and it collapses back to the day-1 difference when there is only one day. That single number is the level — the "5." ## The weights were hiding in the control average Here is the part that connects all of this to synthetic control, and it has been sitting in plain sight. Look again at how we built the control average in the first column: $$ \bar y_{01} = 11 = \tfrac12 (10) + \tfrac12 (12) = w_2\, y_{21} + w_3\, y_{31}, \qquad w_2 = w_3 = \tfrac12 = \tfrac{1}{N_0}. $$ That is a weighted sum, exactly as the axiom warned. Taking "the average" was never a neutral act: it *chose* weights of $1/N_0$ for every control, and it chose them before anyone looked at the data. Put general weights in their place and the gap becomes $$ g_t = y_{1t} - \sum_{j\in\mathcal N_0} w_j\, y_{jt}, \qquad w_j \ge 0, \quad \sum_{j\in\mathcal N_0} w_j = 1 . $$ The two constraints keep the result an honest average: no negative shares, and the shares add to one. That set of allowed weights has a name, the *simplex*, and it is the same set the arithmetic mean has been living on all along. Our three little sequences can already show why the choice matters. With equal weights the gap was $(5,5,5)$, flat as a board. Put *all* the weight on control $2$ instead, $w_2=1, w_3=0$, and the control path becomes $(10,12,11)$, so the gap becomes $$ (16-10,\; 19-12,\; 18-11) = (6,\, 7,\, 7), $$ which is not flat. Put all the weight on control $3$ and the gap is $(4,3,3)$ — also not flat. Hold onto that last one; we meet it again in a moment. So "the gap is constant" was never a fact about the data alone. It was a fact about *the weights we happened to use.* ## One problem, two knobs Step back and look at what we have actually been doing all along. We are trying to reproduce the treated unit's pre-treatment path out of two ingredients: a weighted combination of the controls, and a level shift. Written as one optimization problem, with both ingredients left as unknowns, $$ \min_{\mathbf w \,\in\, \mathcal W,\;\; \beta \,\in\, \mathcal B}\; \sum_{t \le T_0}\Big(\, y_{1t} \;-\; \beta \;-\; \sum_{j\in\mathcal N_0} w_j\, y_{jt} \,\Big)^{2} . $$ That is the whole apparatus. There are exactly two knobs: the weights $\mathbf w$, and the intercept $\beta$. And the two methods we have been talking about are nothing more than two different decisions about which knob you are allowed to turn: | | weights $\mathcal W$ | intercept $\mathcal B$ | |:--|:--|:--| | classic synthetic control | the simplex: $w_j \ge 0,\ \sum_j w_j = 1$ | $\{0\}$ | | difference-in-differences | the single point $w_j = 1/N_0$ | $\mathbb{R}$ | Read that table twice, because it is the punchline. Each method frees exactly what the other one freezes. Synthetic control hunts over every allowed weighting but forbids itself a level shift. Difference-in-differences refuses to hunt at all, fixing the weights at the dead centre of the simplex, but lets the level go wherever it likes. Nothing else separates them: same objective, same data, same donors — different permissions. ### What parallel trends is doing in this picture Freezing the weights is cheap; the price is paid later, and parallel trends is the bill. Because DID never gets to shape the gap, it has to *assume* the gap it was handed behaves. In the language of this optimization, the assumption is that the constant fitting the pre-treatment gaps is the same constant that would fit the post-treatment gaps: $$ \operatorname*{argmin}_{\beta}\; \mathbb{E}\!\left[\sum_{t\le T_0}(g_t-\beta)^2\right] = \operatorname*{argmin}_{\beta}\; \mathbb{E}\!\left[\sum_{t>T_0}(g_t-\beta)^2\right] = \beta^\ast . $$ We only ever get to *solve* the left-hand one. The gaps after treatment contain the very counterfactual we are trying to recover, so they are not ours to use. Parallel trends is precisely the permission slip that lets us carry $\beta^\ast$ across the treatment date and use it on the right. Grant it and the counterfactual writes itself — take the control average and lift it by that one number, $\hat y_{1t}(0) = \bar y_{0t} + \beta^\ast$ — and the treatment effect is however far the gap *departs* from its sticky level, $$ \hat\tau_t = g_t - \beta^\ast = y_{1t} - \big(\bar y_{0t} + \beta^\ast\big). $$ Let the policy land in a fourth period. Suppose the controls come in at $y_{24} = 13$ and $y_{34} = 17$, so $\bar y_{04} = 15$, while the treated unit is observed at $y_{14} = 26$. The counterfactual is $15 + 5 = 20$, the observed $26$ is $6$ above it, and equivalently the gap jumped from $5$ to $11$: $\hat\tau_4 = 11 - 5 = 6$. The $5$ was always going to be there. The $6$ is the treatment. And now you can see what the weights have to do with parallel trends, which is the question I was actually asked. The gap carries $\mathbf w$ inside it. Change the weights and you change the gap, and you change whether its expectation is constant. DID fixes $\mathbf w = 1/N_0$ and then *hopes* the resulting gap is flat — a hope about a gap it never got to shape. Synthetic control *chooses* $\mathbf w$ to make the pre-period gap as flat as it can, buying in-sample parallel trends by construction instead of assuming it. ### Solving it, both ways We can just solve the thing, on our own three sequences, and watch each restriction do its work. Difference-in-differences first. The weights are nailed down at $w_2=w_3=\tfrac12$, so $\sum_j w_j y_{jt}$ is the control average $(11,14,13)$ and the only unknown left is $\beta$. The problem collapses to the one we already solved, $\min_{\beta} \sum_{t\le T_0} (g_t - \beta)^2$ with $\mathbf g = (5,5,5)$, whose solution is $\beta^\ast = 5$ and whose minimized total is $0$. A perfect fit: $(11,14,13) + 5 = (16,19,18)$, exactly the treated path. Now classic synthetic control. Here $\beta$ is pinned at $0$ and the weights are free on the simplex. With two donors the simplex is just a line segment, so there is only one unknown: put $w_2 = w$ and $w_3 = 1-w$ with $0 \le w \le 1$. The synthetic path is $w\,\mathbf y_2 + (1-w)\,\mathbf y_3$, which period by period is $(12 - 2w,\; 16 - 4w,\; 15 - 4w)$, and subtracting it from the treated path $(16,19,18)$ leaves residuals $(4 + 2w,\; 3 + 4w,\; 3 + 4w)$. Square them and add: $$ \text{SSE}(w) = (4+2w)^2 + 2\,(3+4w)^2 = 36w^2 + 64w + 34 . $$ Its derivative is $72w + 64$, positive everywhere on $[0,1]$ — the function climbs the whole way — so the smallest value sits at the left-hand end, $w=0$. Synthetic control therefore puts *all* the weight on control $3$, and its best achievable error is $\text{SSE}(0) = 34$, with residuals $(4,3,3)$. There is that $(4,3,3)$ again, exactly as promised. ### Why one wins so badly here Line them up: | | best fit it can reach | total squared error | |:--|:--|:--:| | classic synthetic control ($\beta = 0$) | $w = (0,1)$ | $34$ | | difference-in-differences ($w = 1/N_0$) | $\beta^\ast = 5$ | $0$ | | neither knob freed | — | $75$ | Synthetic control, allowed to search every weighting there is, cannot get its error below $34$. Difference-in-differences, allowed to search nothing at all but handed one intercept, drives it to $0$. The reason is visible in the original table. At every period the treated unit sits above *both* donors: $16$ over $(10,12)$, then $19$ over $(12,16)$, then $18$ over $(11,15)$. A weighted average with non-negative weights summing to one can never exceed the largest thing it is averaging, so no allowed $\mathbf w$ can ever reach the treated path. The treated unit lies outside the donors' convex hull, and without an intercept synthetic control has no way to climb to it. Every weighting it tries is too low; the best it can manage is to pick the highest donor and eat the remaining gap as error. The intercept is exactly the tool for that job. It does not care about shape at all — it just lifts. And since the gap here was a constant $5$, lifting by $5$ was all the fitting that was ever required. So the contest is not "which method is better." It is which knob the data needs. When the treated unit sits outside the hull but drifts in parallel with it, you need the level, and DID hands it to you. When the treated unit sits inside the hull but drifts differently from the simple average, you need the weights, and synthetic control hands you those. Two honest caveats before we go to real data. This example was built so the gap is exactly constant, which is to say it was built with DID's assumption true by construction — that is why it wins so cleanly. Real data is rarely so obliging, and we are about to watch California fail this exact test. And nothing stops you from freeing both knobs at once, weights on the simplex *and* an intercept in $\mathbb{R}$, which is a real and useful estimator in its own right. The point of the table is not that two cells exhaust the world. It is that the two most famous estimators in this literature are one optimization problem, asked under two different sets of permissions. ## Proposition 99, with uniform weights and an intercept Let me show the weights and the intercept in the flesh. I will use Abadie, Diamond and Hainmueller's canonical Proposition 99 setup: California is treated by the 1989 anti-tobacco law, there are 38 control states, and we observe per-capita cigarette sales from 1970 to 2000. To follow along you'll want the `dataprep` helper from my `mlsynth` package: ```{shell} pip install -U git+https://github.com/jgreathouse9/mlsynth.git ``` We load the panel and hand it to `dataprep`, which returns the treated vector `y`, the `donor_matrix` of the 38 controls, and the number of pre-treatment periods. ```{python} import numpy as np import pandas as pd from mlsynth.utils.datautils import dataprep base_url = "https://raw.githubusercontent.com/jgreathouse9/mlsynth/refs/heads/main/" url = base_url + "basedata/smoking_data.csv" data = pd.read_csv(url) prepped = dataprep( data, data.columns[0], # unit : state data.columns[1], # time : year data.columns[2], # outcome: cigsale data.columns[-1], # treat : Proposition 99 ) ``` Now the estimator, exactly as the algebra prescribes. The uniform weights $1/N_0$ are applied by averaging the donor matrix across its columns; the intercept $\beta^\ast$ is the mean pre-treatment gap; the counterfactual is the shifted donor mean; and the ATT is the average post-treatment gap between California and that counterfactual. ```{python} y = prepped["y"] T0 = prepped["pre_periods"] donors = prepped["donor_matrix"] N0 = donors.shape[1] # Uniform weights: every donor gets 1/N0. These are the SCM simplex weights, # planted at the centroid. w = np.full(N0, 1.0 / N0) donor_mean = donors @ w # equivalently donors.mean(axis=1) # The one free parameter: the intercept beta*. beta = np.mean(y[:T0] - donor_mean[:T0]) yhat = donor_mean + beta # the DID counterfactual ATT = np.mean(y[T0:] - yhat[T0:]) print(f"weights sum to {w.sum():.4f}, min weight {w.min():.4f} (on the simplex)") print(f"intercept beta* = {beta:.4f}") print(f"ATT = {ATT:.4f}") ``` The weights sum to one and are all positive, the intercept comes out to $\beta^\ast \approx -14.36$ — California smoked about 14 fewer packs per capita than the average control state before the law — and the ATT is about $-27.35$ packs per capita. Here is the picture. ```{python} #| label: fig-did-cali #| fig-cap: "Difference-in-Differences on Proposition 99: uniform donor weights plus an intercept." #| echo: False import matplotlib.pyplot as plt T = len(y) time = np.arange(1970, 1970 + T) plt.plot(time, y, label="California", linewidth=2, color="black") plt.plot(time, yhat, "--", label="DID counterfactual", linewidth=2, color="blue") plt.axvline(time[T0 - 1], color="black", linestyle="--", label="Proposition 99") plt.text(1970 + T0 + 1, min(y), f"ATT = {ATT:.2f}", fontsize=12, color="red") plt.xlabel("Year") plt.ylabel("Cigarette Sales Per Capita") plt.legend() plt.show() ``` The counterfactual does not hug California especially well before 1989 — it undershoots for the first few years and overshoots later. That misfit is the whole reason SCM exists, and it is parallel trends failing in front of you. The uniform weights force the counterfactual to be the *average* control state, shifted to California's level; the intercept can only slide that line up and down, it cannot bend it. If California's pre-period gap is not flat, no choice of $\beta$ will make it flat, and the gap you see before 1989 is a preview of the bias you cannot see after it. None of this hangs on my hand-rolled arithmetic. The [`diff-diff`](https://github.com/igerber/diff-diff) package — a small, dedicated DID estimator, a cousin to my own `mlsynth` — computes the same effect from the plain $2 \times 2$ table, once we hand it the two indicators DID actually runs on: a treated-group flag (California, every year) and a post flag (every state, 1989 on). ```{python} from diff_diff import DifferenceInDifferences # DID's two indicators, rebuilt from the panel: a treated-group flag # (California in every year) and a post flag (every state, 1989 on). did_df = data.copy() treated_states = data.loc[data["Proposition 99"] == 1, "state"].unique() post_years = data.loc[data["Proposition 99"] == 1, "year"].unique() did_df["treated"] = data["state"].isin(treated_states).astype(int) did_df["post"] = data["year"].isin(post_years).astype(int) result = DifferenceInDifferences().fit( did_df, outcome="cigsale", treatment="treated", post="post" ) print(result) ``` Same ATT — about $-27.35$ — now carrying a standard error and a $p$-value, from a routine that never heard of a weight vector or a simplex. It agrees because it is the same estimator under another name: the hand-rolled intercept $\beta^\ast$ is the average pre-period gap, and the $2 \times 2$ DID is that same gap differenced across the treatment date. ## So what? Now, you may be saying, "Hey Jared, this is a lot of algebra to conclude that a mean is a mean." But the payoff is a clean way to see what you are actually assuming when you reach for DID, and why the choice between DID and SCM is a choice about which knob you are allowed to turn. Both estimators answer the same question — what would California have done without Prop 99? — by averaging control states. DID commits to equal weights before it looks at the data; SCM lets the pre-treatment fit choose the weights. Neither fabricates a "fake" California or grafts together a "Frankenstein" unit, any more than the national average income fabricates a fake person. They both summarize real control units with weights that live on the same simplex. The difference is only whether those weights are handed to you as an axiom or estimated as a decision. And the reason that decision matters shows up the moment you write untreated outcomes as a factor model, $y_{jt}(0) = a_j + \mathbf{b}_j^\top \mathbf{f}_t + u_{jt}$. Averaging the donor pool averages its factor loadings $\bar{\mathbf{b}}_0$, and — as I show in the book — the DID counterfactual is unbiased exactly when the treated unit's loading equals that average, $\mathbf{b}_1 = \bar{\mathbf{b}}_0$. That is parallel trends restated one more time: the leftover gap is $(\mathbf{b}_1 - \bar{\mathbf{b}}_0)^\top (\mathbf{f}_t - \bar{\mathbf{f}})$, a term that *moves with $t$* whenever the treated loading differs from the donor average — a gap no intercept can flatten. Uniform weights only reach California's factor exposure if California happens to sit at the donor pool's center of gravity. When it does not — and a state with California's tobacco culture usually does not — you want the freedom to move the weights off the centroid, toward the donors that actually share its exposure. That freedom *is* synthetic control, and DID is the center of it you get by refusing to use it. So the answer to the question I was asked is: DID is a synthetic control whose weight vector you filled in yourself, with $1/N_0$ in every slot, before the data had a chance to weigh in — and which, having given up the weights, keeps the intercept that classic synthetic control throws away. Everything else — the loss, the simplex, the gap, the counterfactual, and parallel trends itself — is shared. You are, as ever, never *not* averaging. The only question is who gets to pick the weights.