The Iterative Synthetic Control Method

Econometrics
Author

Jared Greathouse

Published

July 9, 2025

This will be a short blog post. I’ve spent the last two months doing a little industry work in the marketing realm, and in the meantime I made some substantial changes to mlsynth. This blog post simply shows one of the new key features I have implemented.

Chances are if you’re reading this, you know what I mean by the notion of SUTVA, or the stable unit treatment value assumption. It is the idea that if we care about the causal impact of a treatment on one unit, but other units are affected by the treatment or otherwise experience a similar treatment, that this exposure confounds our treatment effect with respect to the original unit we do care about. Say we wish to study the impact of German Reunification on West Germany’s GDP. We know West Germany was exposed, but what about neighboring nations like Austria or France? What if the reunification had regional effects? Analysts therefore have a problem: Austria and France may be very similar to West Germany, and therefore informative of West Germany’s counterfactual, but we are concerned they are exposed or affected by the main treatment of interest. What do we do? Before, researchers would need to drop these units or argue for their inclusion/exclusion, despite them being treated. Now, we do not need to do that, as SCM has a few approaches that deal with spillovers (this post covers just one). The approach, called iterative synthetic controls, is deceptively simple.

Suppose Austria and France are partly treated. Step one of iSCM is to estimate a synthetic control for Austria, including France as a donor but excluding West Germany. The SCM that does the cleaning is handled internally by mlsynth’s spillover-aware estimator; the precise details are not really important here, but you may read the docs should you like. We then take the model’s prediction for Austria and replace Austria’s contaminated post-treatment outcomes with the synthetic values the model expects absent any spillover.

Next, we do France: using the cleaned up Austria as a donor, we estimate the synthetic control for France, using the now-cleaned up Austria and the remaining donor pool units. As before, we replace the values for the original France (in the original dataset) with the new synthetic France. We now have cleaned up our two donors that may be exposed to the treatment.

Now, with these two cleaned donors, we estimate the counterfactual for West Germany, with our 14 totally unexposed donors and the two now cleaned up donors that were once partially exposed.

Estimation in Python

“But Jared!”, you will say, this seems like a lot of looping and lots of donor tracking. Well fear not, that is what mlsynth is for. In order to get these results, you need Python (3.9 or greater) and mlsynth, which you may install from the Github repo. You’ll need the most recent version.

pip install -U git+https://github.com/jgreathouse9/mlsynth.git

First we estimate the original model with CLUSTERSC, the baseline that treats Austria and France as ordinary donors.

import pandas as pd
from mlsynth import CLUSTERSC

# Load the reunification dataset
url = "https://raw.githubusercontent.com/jgreathouse9/mlsynth/main/basedata/german_reunification.csv"
df = pd.read_csv(url)

# Define configuration for CLUSTERSC
config = {
    "df": df,
    "outcome": "gdp",          # per capita GDP
    "treat": "Reunification",     # binary treatment indicator
    "unitid": "country",          # country name
    "time": "year",               # time variable
    "display_graphs": True,       # display counterfactual plots
    "save": False,
    "counterfactual_color": ["red", "blue"],
    "Frequentist": True, "method": "BOTH"
}

originalresult = CLUSTERSC(config).fit()
findfont: Failed to find font weight medium for DejaVu Sans, now using 400.
findfont: Failed to find font weight medium for DejaVu Sans, now using 400.
findfont: Failed to find font weight medium for DejaVu Sans, now using 400.

These are the original results. Now we see how sensitive the results are to adjusting for spillover effects. For that we hand the same problem to SPILLSYNTH, the spillover-aware estimator, and ask it for the iterative method.

from mlsynth import SPILLSYNTH
from IPython.display import display, Markdown

spill_config = {
    "df": df,
    "outcome": "gdp",
    "treat": "Reunification",
    "unitid": "country",
    "time": "year",
    "method": "iterative",                 # the waterfall donor-cleaning method
    "affected_units": ["Austria", "France"],  # units we suspect are partly treated
    "display_graphs": True,
}

iSCM = SPILLSYNTH(spill_config).fit()
fit = iSCM.iterative

effects = pd.DataFrame(
    [
        {"Quantity": "Naive SCM ATT",          "Value": f"{iSCM.att_scm:,.0f}"},
        {"Quantity": "Spillover-adjusted ATT", "Value": f"{iSCM.att:,.0f}"},
        {"Quantity": "Pre-treatment RMSE",     "Value": f"{fit.pre_rmspe:,.1f}"},
    ]
)
spillovers = pd.DataFrame(
    [{"Cleaned donor": u, "Spillover effect": f"{v:,.0f}"}
     for u, v in fit.spillover_att.items()]
)

display(Markdown("### Iterative SCM\n\n" + effects.to_markdown(index=False)))
display(Markdown("#### Post-period spillover carried by each cleaned donor\n\n"
                 + spillovers.to_markdown(index=False)))

Iterative SCM

Quantity Value
Naive SCM ATT -1306
Spillover-adjusted ATT -1384
Pre-treatment RMSE 60.9

Post-period spillover carried by each cleaned donor

Cleaned donor Spillover effect
Austria -261
France 172

You pass the same config you would give any mlsynth estimator, set method="iterative", and hand it a list of the units you believe are partly treated through affected_units. From there the algorithm handles the donor cleaning under the hood and returns a single spillover-adjusted fit, with the naive synthetic control kept alongside for comparison. The naive fit — the one that trusts Austria and France as ordinary donors — puts the ATT at about -1,306. Once we clean both donors, the adjusted ATT moves to about -1,384: a slightly larger drop, not a smaller one. The reason is in the per-donor spillover column. Austria carried a negative post-reunification spillover of roughly -261 and France a positive one of about +172, so netting them out leaves West Germany looking a little worse off than the naive fit suggested, not better. Clean only Austria and the ATT is about -1,390; clean only France and it is about -1,295. The pre-treatment RMSE sits at about 61 either way, because this waterfall reuses the pre-period weights and only rewrites the affected donors’ post-treatment values — the fit you validated before the intervention is the fit you keep, so the donor weights do not move; what moves is the counterfactual after 1990. Either way the headline is unchanged: reunification cost West German per-capita GDP a great deal, and that conclusion survives the spillover adjustment. Of course, the key aspect of this procedure is still knowing which units are likely to have spillover effects in the first place.

Comments

So, this is not the only way to do this. There are plenty of other methods that people have developed for this purpose too. I likely will not program all of thse myself into mlsynth, but others who are so inclined are welcome to assist in the effort! in the future, I’ll also allow you to switch between the options for each kind of spillover management (the inclusive method versus the iterative method, for example). But, now you know how to use this for your own work. As usual, comments or suggestions are always appreciated.

What Is An Average?

Econometric Theory
Statistics

What is a Synthetic Control?

Econometrics
Causal Inference

Synthetic Controls With More Than One Outcome

Causal Inference
Econometrics

Forward Selected Synthetic Control

Machine Learning
Econometrics
No matching items