Written and maintained by CASRAI Editorial Board
Last updated
Difference-in-differences (DiD) earns its causal claim from a single untestable assumption: that in the absence of treatment, the treated and comparison groups would have moved in parallel. Everything else — the two-way regression, the event-study plot, the clustered standard errors — is machinery built on top of that one claim. The two things that most often go wrong are that the assumption is asserted rather than evidenced, and that the estimator used to exploit it silently stops estimating the parameter you think it does the moment units adopt treatment at different times.
This page treats both. The first half sets out what a pre-trend plot, an event-study specification and each family of placebo test actually establish, and — more usefully — what they do not. The second half deals with the staggered-adoption problem: with variable treatment timing and treatment effects that change over time, the standard two-way fixed effects (TWFE) regression uses already-treated units as controls and can return an estimate with the wrong sign while every underlying effect is positive. A seeded simulation below produces exactly that, along with an event-study plot showing large apparent pre-trends in data where parallel trends holds exactly by construction.
DiD is one of the remedies for time-varying omitted-variable bias catalogued in the three sources of endogeneity and the remedy that matches each, which shows with its own simulation why unit fixed effects alone remove only time-invariant confounding. This page is the downstream detail: what it costs to make the DiD version of that remedy credible.
What the design actually assumes
Write potential outcomes for unit i in period t as Y(1) and Y(0), in the standard potential-outcomes framing of a causal question. DiD identifies the average treatment effect on the treated (ATT) under four conditions, not one:
- Parallel counterfactual trends. The average change in Y(0) between the pre-period and the post-period is the same for the treated group as for the comparison group. Note what this is a statement about: the untreated potential outcome of the treated group, in the post-period, which is never observed.
- No anticipation. Units do not change behaviour in the pre-period in response to a treatment they know is coming. A policy announced two years before it binds violates this even if the parallel-trends assumption itself is fine.
- No spillover between groups. The comparison group’s outcomes are unaffected by the treated group’s treatment. Geographically adjacent comparison units and firms in the same product market routinely fail this.
- Stable composition. The units being compared do not change in a way correlated with treatment — differential attrition converts the design into a selection problem that no amount of differencing fixes.
Parallel trends is also scale-dependent. Trends that are parallel in levels are generally not parallel in logs, and vice versa. Choosing the functional form after seeing which one produces flat pre-trends is a specification search, not a diagnostic; the choice should follow from what the outcome means, and be stated before the plot is drawn.
Evidencing parallel trends: what each diagnostic does and does not establish
The assumption concerns unobserved counterfactual trends in the post-period. Every available diagnostic concerns observed trends in the pre-period. That gap is not a technicality — it is the entire reason these tests are weaker evidence than they are usually presented as being.
The pre-trend plot
Plotting group means by period before treatment is the minimum. It is genuinely informative: a visible divergence before treatment is close to dispositive evidence against the design, and no amount of subsequent estimation repairs it. What it cannot do is establish the converse. Two series that moved together for four pre-periods may diverge in the fifth for reasons unrelated to treatment, and a plot is silent on that.
Two practical failures are common. First, plotting levels when the assumption is about changes — two series can sit far apart and still be perfectly parallel, and eyeballing the gap invites the wrong conclusion. Second, plotting raw means when the specification includes covariates, so the picture shown is not the picture the estimator uses.
The event-study specification
The formal version replaces the single treatment indicator with a set of relative-time (event-time) indicators, one for each period before and after adoption, with one pre-period — conventionally the period immediately before treatment — omitted as the reference. The pre-period coefficients are the pre-trend test; the post-period coefficients are the dynamic treatment effect path.
Three details decide whether the plot means anything:
- The reference period is a normalisation, not a neutral choice. Every coefficient is a difference relative to it. If the omitted period is itself anomalous, the whole path shifts. Report which period is omitted.
- Endpoint bins. Far leads and lags are usually estimated off a shrinking set of cohorts. Binning the endpoints is standard; leaving them unbinned produces the noisy, wide-interval extremes that dominate the visual impression of a plot for no substantive reason.
- Joint, not pointwise, testing. Five pre-period coefficients each individually insignificant is not the same as a joint test of the null that all five are zero, and the joint test is the one that matches the claim being made.
Why passing a pre-trend test is not proof
This is the point most tutorials skip, and it has two distinct halves, both established in Roth’s Pretest with Caution (American Economic Review: Insights 4(3), 2022, 305–322).
First, conventional pre-trend tests may have low power. The paper analyses this in theory and in simulations calibrated to a survey of recent papers in leading economics journals, and reports that these limitations are important in practice. A pre-trend test that fails to reject is frequently a test that could not have rejected a violation large enough to overturn the headline result. Reporting “we tested for pre-trends and found none” without reporting the magnitude of violation the test could actually have detected is reporting the absence of evidence as evidence of absence.
Second — and this is the less intuitive half — conditioning the analysis on the result of a pretest can distort estimation and inference, potentially exacerbating the bias of point estimates and under-coverage of confidence intervals. Selecting the sample, the specification or the paper itself on having passed a pre-trend screen changes the sampling distribution of what survives. In the presence of a real underlying violation, the surviving estimates are drawn from the tail where noise happened to mask it, and their confidence intervals cover at less than the nominal rate. The pretest does not merely fail to help; run as a gate, it can make things worse.
There is also a mirror-image failure that Roth’s result does not cover and that catches people from the other direction: an event-study plot can show large, statistically significant apparent pre-trends when parallel trends holds exactly. That is a property of the estimator, not the data, and it is demonstrated in panel D of the simulation below.
Placebo tests, one at a time
“We ran placebo tests” describes at least four different procedures that establish four different things. They are not interchangeable, and only two of them are evidence about parallel trends at all.
| Test | What it does | What it establishes | What it does not |
|---|---|---|---|
| In-time placebo — assign a fake treatment date entirely within the pre-period | Runs the full specification on pre-treatment data only | That the pre-period is stable under the estimator you are actually using, covariates and all — stronger than eyeballing a plot | Anything about the post-period. It inherits the same low power as the pre-trend test, and it is subject to the same pretesting distortion |
| In-space placebo — assign treatment at random to untreated units, many times | Builds an empirical distribution of the estimate under a known null of no effect | Whether your point estimate is large relative to what the design produces by chance; a randomisation-inference p-value that does not rely on asymptotics | It is an inference device, not evidence about parallel trends. A design with a real confound can pass it comfortably |
| Placebo outcome — an outcome the treatment cannot plausibly affect | Runs the same design on the unaffected outcome | Genuine evidence about confounding: a treated-group shock that moves the placebo outcome shows up as a non-zero effect where none can exist | Only detects confounders that move that outcome. Choosing the placebo outcome after seeing the results reduces it to decoration |
| Triple difference (DDD) — a within-treated-group comparison unaffected by the policy | Differences out shocks common to the treated group | Robustness to a specific, named class of confounder — one hitting the whole treated group equally | It does not remove the assumption; it replaces parallel trends in levels with parallel trends in differences across the third dimension, which can be the harder claim |
A useful discipline: before running any of these, write down the confounding story it is meant to rule out. A placebo test with no named alternative behind it cannot fail informatively, and cannot succeed informatively either. The same logic applies to balance diagnostics after propensity score matching — good balance on what you measured is not evidence about what you did not.
Bounding the violation instead of testing it
The more honest framing, and increasingly the expected one in economics, is to stop treating parallel trends as a hypothesis to pass and start treating it as a quantity to bound. Rambachan and Roth’s A More Credible Approach to Parallel Trends (Review of Economic Studies 90(5), 2023, 2555–2591) develops inference that remains valid when parallel trends holds only approximately, by restricting how large the post-treatment violation could be — for example, no larger than some multiple of the largest violation observed in the pre-period — and reporting the resulting identified set rather than a point estimate.
The reportable output is a breakdown value: how big a violation of parallel trends would have to be before the conclusion changes sign or loses significance. That is a number a reader can argue with on substantive grounds, which a binary “the pre-trend test passed” is not. If you report nothing else from this section, report that.
The staggered-adoption problem
Everything above assumes a clean 2×2: one treated group, one comparison group, one treatment date. Most real applications are not that. States adopt a policy in different years; hospitals roll out a protocol ward by ward; funders phase in a mandate. The reflex is to keep the same regression and let the treatment indicator switch on at each unit’s own adoption date, with unit and period fixed effects absorbing the rest. That specification does not do what it appears to do.
What TWFE actually estimates
Goodman-Bacon’s Difference-in-differences with variation in treatment timing (Journal of Econometrics 225(2), 2021, 254–277) derives the answer: the staggered TWFE estimator is a weighted average of all possible two-group/two-period DiD estimators in the data. Those 2×2 comparisons fall into three kinds:
- Each treated cohort against the never-treated units — clean.
- An earlier-adopting cohort against a later-adopting cohort, over the window while the later cohort is still untreated — also clean; the later cohort is a legitimate not-yet-treated comparison.
- A later-adopting cohort against an already-treated earlier cohort, over the window after the earlier cohort has adopted. This is the problem term.
In the third comparison, the “control” group’s outcome is already moving because of its own treatment. What the estimator differences out is not a counterfactual trend but the ongoing evolution of the earlier cohort’s treatment effect. Goodman-Bacon’s stated result is precise on the condition: the estimand averages treatment effect heterogeneity, and it is biased when effects change over time.
Two things about the weights are worth stating carefully, because they are routinely conflated. In Goodman-Bacon’s decomposition the weights attach to 2×2 estimators; they are non-negative and sum to one, and they are driven by group sizes and by how much treatment variance each comparison contributes — which means the comparisons with the most balanced timing get the most weight, regardless of whether they are the clean ones. The negative weights people refer to are from a different decomposition, below, and attach to something else.
Negative weights on the treatment effects themselves
De Chaisemartin and D’Haultfœuille’s Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects (American Economic Review 110(9), 2020, 2964–2996) decomposes the same regression differently. Their result, in the paper’s own words: linear regressions with period and group fixed effects “estimate weighted sums of the average treatment effects (ATE) in each group and period, with weights that may be negative. Due to the negative weights, the linear regression coefficient may for instance be negative while all the ATEs are positive.”
That is the claim to take literally, and it is worth being precise about what it does and does not say. It does not say TWFE is generally wrong-signed, or that negative weights always appear, or that the bias is usually large. It says the weights may be negative, and that when they are, sign reversal is possible rather than merely conceivable. Whether it happens in a particular dataset is an empirical question about that dataset’s timing structure and treatment shares — the paper’s companion diagnostic computes the weights TWFE attaches to each group-period ATE in your own data, so you can count how many are negative and how much total weight sits on them before deciding whether it matters. The paper also proposes an alternative estimator that avoids the issue.
When TWFE is fine — the precise conditions
Overstating this is as unhelpful as ignoring it. TWFE is not broken. It is unbiased for the ATT in each of these cases:
- Common treatment timing. All treated units adopt in the same period. There are then no already-treated comparisons to make, the decomposition collapses to the single clean 2×2, and TWFE is the canonical DiD estimator. Nothing in this section applies.
- Staggered timing with homogeneous, static effects. If the treatment effect is constant across cohorts and constant over time since adoption, every 2×2 comparison — the forbidden ones included — targets the same number, so a weighted average of them recovers it. The already-treated control group’s outcome is shifted by a constant that differences out cleanly.
- Staggered timing where no unit is ever used as a control after its own adoption — for example, if the estimation window is trimmed so that only not-yet-treated units serve as comparisons.
The second case is the one people implicitly rely on, and it is a strong restriction. It requires that the effect neither grows, decays, nor differs by cohort. Effects that ramp up over the first few periods are the norm rather than the exception in policy evaluation, and a ramp is exactly what breaks it. If your own event-study plot shows a rising post-treatment path — the thing authors usually present as evidence the design is working — you have shown that this condition fails.
Contaminated event-study coefficients
The dynamic version of the regression is affected too, and in a way that undermines the pre-trend test built on it. Sun and Abraham’s Estimating dynamic treatment effects in event studies with heterogeneous treatment effects (Journal of Econometrics 225(2), 2021, 175–199) shows that with variation in treatment timing, “the coefficient on a given lead or lag can be contaminated by effects from other periods, and apparent pretrends can arise solely from treatment effects heterogeneity.” They propose an interaction-weighted estimator that is free of the contamination.
The practical consequence is severe: the standard pre-trend test can reject in data where parallel trends holds perfectly, purely because the estimator smears post-treatment effects from one cohort into the pre-period coefficients of another. An author who then “fixes” the design in response — dropping units, adding unit-specific trends, changing the window — is repairing a problem that does not exist while leaving the one that does.
A simulation where TWFE returns the wrong sign
The following was run in R 4.6.1 with set.seed(20260826). Every figure printed below is real output from the code shown; nothing is illustrative. The data-generating process is deliberately benign — parallel trends holds exactly by construction, because the time trend is common to every unit, and every true treatment effect is positive. The only complications are staggered timing and an effect that ramps up after adoption.
set.seed(20260826)
TT <- 20
n_early <- 200; n_late <- 200; n_never <- 20
g_early <- 4; g_late <- 15; ramp <- 0.6
n <- n_early + n_late + n_never
g <- c(rep(g_early, n_early), rep(g_late, n_late), rep(Inf, n_never))
id <- rep(1:n, each = TT)
tt <- rep(1:TT, times = n)
gi <- g[id]
alpha <- rep(rnorm(n, sd = 1), each = TT) # unit fixed effect
trend <- 0.10 * tt # COMMON to every unit: parallel trends holds exactly
e <- tt - gi # event time
tau <- ifelse(is.finite(gi) & e >= 0, ramp * (e + 1), 0) # every true effect is POSITIVE
y <- alpha + trend + tau + rnorm(n * TT, sd = 0.5)
D <- as.numeric(is.finite(gi) & tt >= gi)
d <- data.frame(id, tt, gi, e, D, y, tau)
demean2 <- function(v, i, s) v - ave(v, i) - ave(v, s) + mean(v)
true_att <- mean(d$tau[d$D == 1])
twfe <- coef(lm(demean2(d$y, d$id, d$tt) ~ demean2(d$D, d$id, d$tt) - 1))[[1]]
Two adopting cohorts of 200 units each, adopting in period 4 and period 15, plus a small never-treated group of 20, observed over 20 periods. The two-way within transformation is algebraically identical to including unit and period dummies on this balanced panel, and avoids fitting 420 dummy variables.
A. TWFE AGAINST THE TRUTH
true ATT on the treated : 4.539
two-way fixed effects : -0.318
The true average effect on the treated is +4.539. The regression returns −0.318. Not attenuated — reversed. Every individual treatment effect in this dataset is positive, parallel trends holds exactly, there is no confounding of any kind, and the standard specification reports a negative effect.
This is not a lucky seed. Re-drawing only the noise 500 times, holding the design fixed, gives a mean TWFE estimate of −0.303 with a standard deviation of 0.021, and the estimate is negative in 500 of 500 replications — the largest value drawn was −0.245. The sign reversal is a property of the design, not of sampling variability, and a conventional standard error would report it as precisely estimated.
Where the sign comes from
Goodman-Bacon’s decomposition says the estimate is a weighted average of 2×2 comparisons. Computing the four of them directly shows which one is responsible:
never <- !is.finite(d$gi); earl <- d$gi == g_early; late <- d$gi == g_late
win <- function(grp, a, b) mean(d$y[grp & d$tt >= a & d$tt <= b])
bac <- function(trt, ctl, pa, pb, qa, qb)
(win(trt, qa, qb) - win(trt, pa, pb)) - (win(ctl, qa, qb) - win(ctl, pa, pb))
k1 <- bac(earl, never, 1, g_early - 1, g_early, TT)
k2 <- bac(late, never, 1, g_late - 1, g_late, TT)
k3 <- bac(earl, late, 1, g_early - 1, g_early, g_late - 1)
k4 <- bac(late, earl, g_early, g_late - 1, g_late, TT)
B. THE 2x2 BUILDING BLOCKS TWFE IS AVERAGING
early vs never-treated : 5.425
late vs never-treated : 2.089
early vs late, before late adopts : 3.613
late vs early, AFTER early adopts : -3.034 <- forbidden
true effect in that last window : 2.100
The three clean comparisons are all positive, as they should be. The fourth — the late cohort measured against the early cohort after the early cohort has already adopted — returns −3.034 where the true effect over that window is +2.100. The mechanism is transparent once seen: over that window the early cohort’s own effect is still ramping, rising faster than the late cohort’s newly-started effect, so subtracting the “control” group’s change subtracts a treatment effect. Because Goodman-Bacon’s weights are non-negative and sum to one, the overall estimate must lie between −3.034 and 5.425; the timing structure puts enough weight on the forbidden term to drag it below zero.
An estimator that recovers the truth
The fix in Callaway and Sant’Anna’s Difference-in-Differences with multiple time periods (Journal of Econometrics 225(2), 2021, 200–230) is to stop asking the regression for one number and instead identify a group-time average treatment effect ATT(g,t) for each adopting cohort g and each period t, each estimated against a comparison group that is not yet treated (or never treated), with the period immediately before that cohort’s adoption as the baseline. Aggregation into a single summary is then a separate, explicit choice.
att_gt <- function(gg, t) {
base <- gg - 1; trt <- d$gi == gg
(win(trt, t, t) - win(trt, base, base)) - (win(never, t, t) - win(never, base, base))
}
rows <- do.call(rbind, lapply(c(g_early, g_late), function(gg)
do.call(rbind, lapply(gg:TT, function(t)
data.frame(g = gg, t = t, e = t - gg, att = att_gt(gg, t),
truth = ramp * (t - gg + 1), nunits = sum(g == gg))))))
cs <- sum(rows$nunits / sum(rows$nunits) * rows$att)
C. GROUP-TIME ATT(g,t), NEVER-TREATED COMPARISON GROUP
g t e ATT(g,t) truth
4 4 0 0.786 0.600
4 5 1 1.291 1.200
4 6 2 1.917 1.800
15 15 0 0.465 0.600
15 16 1 1.376 1.200
15 17 2 1.796 1.800
aggregated (group-size weighted) : 4.649 true ATT 4.539
The aggregate is 4.649 against a truth of 4.539 — a 2.4 percent overshoot, well inside sampling noise given that the comparison group here is only 20 units. The same data, the same assumption, a different estimator: the sign reversal disappears entirely because no already-treated unit is ever used as a control.
The event study, and the pre-trends that are not there
The last panel runs the dynamic specification — relative-time indicators with period −1 omitted and the endpoints binned — alongside the cohort-specific effects at the same event times.
D. EVENT STUDY: DYNAMIC TWFE vs COHORT-SPECIFIC, AGAINST THE TRUTH
e dynamic TWFE cohort-specific truth
-6 3.238 -0.053 0.000
-5 2.993 0.177 0.000
-4 2.329 0.160 0.000
-3 0.545 0.087 0.000
-2 0.290 0.107 0.000
0 0.357 0.625 0.600
1 0.594 1.333 1.200
2 0.944 1.857 1.800
3 1.320 2.471 2.400
4 1.650 3.046 3.000
5 3.438 3.667 3.600
The pre-period coefficients from the dynamic TWFE regression are 3.238, 2.993 and 2.329 at event times −6, −5 and −4 — large, and in a dataset where the pre-treatment trend is identical for every unit by construction. This is Sun and Abraham’s result made concrete: apparent pre-trends arising solely from treatment-effect heterogeneity being smeared across relative periods. A researcher looking at this plot would conclude the parallel-trends assumption had failed, and would be wrong. The cohort-specific column, computed the same way as panel C, sits at essentially zero throughout the pre-period, as the truth requires.
The post-treatment columns make the second point. The dynamic TWFE path (0.357, 0.594, 0.944, 1.320, 1.650) understates the true path (0.600, 1.200, 1.800, 2.400, 3.000) by a growing margin, so it also misrepresents the shape of the effect — it makes a linear ramp look like a slow, flattening response. The cohort-specific estimates track the truth closely at every horizon.
What this simulation does and does not show, honestly
- It is a constructed case, not a claim about typical magnitudes. It was tuned to make the sign reversal visible: a long panel, a large gap between adoption dates, a strongly ramping effect, and a deliberately small never-treated group. Milder versions of the same structure produce attenuation rather than reversal — which is harder to notice and therefore arguably worse.
- The modern estimators here are hand-implemented in base R following the Callaway–Sant’Anna group-time logic and the Sun–Abraham cohort-by-relative-period logic. The
didandfixestpackages were not available in this environment, so the published estimators’ doubly-robust identification, covariate conditioning, and simultaneous-inference bootstrap are not reproduced. What is shown is the identification idea, not those packages’ output. Use the packages in real work; do not use the code above as an estimator. - No standard errors are reported for the group-time estimates. Correct inference for them is not a by-product of the point estimates and is one of the main things the published implementations provide.
- The panel is balanced and complete. Unbalanced panels, treatment reversal (units switching off again) and continuous or multi-valued treatments each raise further issues that none of the above addresses.
Choosing an estimator
| Situation | Estimator | What it needs |
|---|---|---|
| One treatment date, two groups | Standard TWFE / the 2×2 DiD | Nothing further. The staggered-adoption literature does not apply |
| Staggered adoption, absorbing treatment, want group-time effects and a defensible aggregate | Callaway & Sant’Anna (2021) group-time ATT | A never-treated or not-yet-treated comparison group; a choice of aggregation scheme, stated in advance |
| Staggered adoption, want an uncontaminated dynamic event-study path | Sun & Abraham (2021) interaction-weighted estimator | Cohort shares by relative period; a clean reference period |
| Staggered adoption, efficiency matters, long pre-period available | Borusyak, Jaravel & Spiess (2024) imputation estimator | Fitting unit and period effects on untreated observations only, then imputing untreated potential outcomes |
| Treatment switches on and off, or is non-binary | De Chaisemartin & D’Haultfœuille (2020) and their subsequent work | Comparisons restricted to units whose treatment actually changes between adjacent periods |
| Parallel trends is doubtful rather than clearly violated | Rambachan & Roth (2023) sensitivity analysis, layered on any of the above | An explicit restriction on the size of the possible violation, and a reported breakdown value |
Two practical notes. In most applied settings the modern estimators and TWFE give similar answers, and reporting both — with the difference explained — is more informative than reporting either alone. And a large gap between them is itself a diagnostic: it means the forbidden comparisons carry real weight in your data, which is worth reporting rather than resolving silently in favour of whichever number you prefer.
What to report
- The adoption date of every cohort, and the size of each — including the never-treated group, or a statement that there is none.
- The estimator, named, with the comparison group it uses (never-treated, not-yet-treated, or both) and the reference period.
- An event-study plot from an estimator that is robust to staggered adoption if timing varies, with binned endpoints and the reference period labelled.
- A joint test of the pre-period coefficients, not a set of pointwise intervals — and an honest statement of what magnitude of violation that test had power to detect.
- A breakdown value: the size of parallel-trends violation that would overturn the conclusion.
- The placebo tests actually run, each with the confounding story it was meant to rule out named in advance.
- The clustering level for standard errors — the unit of treatment assignment, not the unit of observation — and the number of clusters. With few clusters, a wild cluster bootstrap rather than the analytic standard error.
- Whether the functional form (levels or logs) was chosen before or after seeing the pre-trends.
Pre-specifying the estimator, the comparison group and the event window in a registered protocol removes most of the room in which the choices above become results-dependent. For DiD specifically, the pre-trend plot is the highest-risk moment: it arrives before the main estimate, it invites specification changes, and Roth’s second result is precisely about what those changes do to inference.
Frequently asked questions
What is the parallel trends assumption in difference-in-differences?
That in the absence of treatment, the average outcome of the treated group would have changed by the same amount as the average outcome of the comparison group. It is a statement about the treated group’s unobserved untreated potential outcome in the post-period, which is why it cannot be tested directly — only supported indirectly, by evidence from the pre-period and by bounding how large a violation is plausible.
Does passing a pre-trend test prove parallel trends holds?
No. Roth (2022) establishes two limitations: conventional pre-trend tests may have low power, so failing to reject is often uninformative about violations large enough to matter; and conditioning the analysis on having passed such a pretest can distort estimation and inference, potentially exacerbating bias and producing confidence intervals that under-cover. A pre-trend test is necessary-ish evidence, not sufficient evidence. Report the magnitude the test could have detected, and a breakdown value.
Can two-way fixed effects really give the wrong sign?
Yes. De Chaisemartin and D’Haultfœuille (2020) show TWFE estimates weighted sums of the average treatment effects in each group and period with weights that may be negative, and that consequently the regression coefficient may be negative while all the underlying treatment effects are positive. The simulation on this page produces exactly that: a true ATT of +4.539 and a TWFE estimate of −0.318, negative in 500 of 500 replications, in data with no confounding whatsoever.
When is two-way fixed effects still unbiased with staggered adoption?
When treatment effects are homogeneous — constant across cohorts and constant over time since adoption — every 2×2 comparison in Goodman-Bacon’s decomposition targets the same parameter, so the weighted average recovers it. It is also unbiased whenever treatment timing is common to all treated units, because there are then no already-treated comparisons at all. The restriction that usually fails is the static one: an effect that ramps up or decays after adoption breaks it, and a rising event-study path is direct evidence that it has.
Why are already-treated units a problem as controls?
Because a control group is supposed to supply the counterfactual trend, and an already-treated group’s outcome is moving partly because of its own treatment. Differencing it out removes a treatment effect along with the trend. In panel B above, that single comparison returns −3.034 where the true effect over the window is +2.100.
What is the Goodman-Bacon decomposition?
An exact algebraic result showing that the staggered TWFE estimate is a weighted average of all possible two-group, two-period DiD estimators available in the data, with weights determined by group sizes and treatment-variance shares. It is diagnostic rather than corrective: it tells you how much weight your particular dataset places on the already-treated comparisons, which is what decides whether the problem is severe or cosmetic in your case.
Are Goodman-Bacon’s negative weights the same as de Chaisemartin and D’Haultfœuille’s?
No, and conflating them causes confusion. Goodman-Bacon’s weights attach to 2×2 estimators and are non-negative, summing to one; the bias there comes from the already-treated comparisons being bad estimates, not from negative weights. De Chaisemartin and D’Haultfœuille decompose the same coefficient into a weighted sum of group-period average treatment effects, and it is those weights that may be negative. Both results describe the same underlying failure from different angles.
Can an event-study plot show pre-trends when parallel trends actually holds?
Yes, and this is Sun and Abraham’s (2021) result: with variation in treatment timing, a lead or lag coefficient can be contaminated by effects from other relative periods, so apparent pre-trends can arise solely from treatment-effect heterogeneity. Panel D above shows pre-period coefficients above 3.0 in data with a perfectly common time trend. Using a robust estimator for the event-study plot is not optional if timing varies — the plot is otherwise not a test of anything.
Do I need a never-treated group to use these methods?
Not necessarily. Callaway and Sant’Anna’s approach can use not-yet-treated units as the comparison group, which works whenever some cohorts adopt later. What is not identified without a never-treated group is the effect on the last-adopting cohort after it adopts, because at that point no untreated comparison exists. If every unit eventually adopts, say so and report which cohort-periods are outside the identified set.
How does DiD relate to fixed effects and to instrumental variables?
All three address omitted-variable bias, but at different scopes. Unit fixed effects remove only time-invariant confounding — demonstrated with a simulation in the guide to endogeneity sources and remedies, where a pooled OLS estimate of 2.510 falls only to 1.727 against a true effect of 1.0 once a time-varying confounder is present. DiD adds the second difference and can handle a time-varying confounder, provided it affects both groups equally. Instrumental variables makes no parallel-trends claim at all but requires an exclusion restriction that is equally untestable. Choosing between them is a question about which untestable assumption you can defend from institutional facts.
Sources
- Goodman-Bacon, A. (2021). “Difference-in-differences with variation in treatment timing.” Journal of Econometrics 225(2), 254–277. Working paper version: NBER Working Paper 25018.
- de Chaisemartin, C., and D’Haultfœuille, X. (2020). “Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects.” American Economic Review 110(9), 2964–2996.
- Callaway, B., and Sant’Anna, P. H. C. (2021). “Difference-in-Differences with multiple time periods.” Journal of Econometrics 225(2), 200–230.
- Sun, L., and Abraham, S. (2021). “Estimating dynamic treatment effects in event studies with heterogeneous treatment effects.” Journal of Econometrics 225(2), 175–199.
- Roth, J. (2022). “Pretest with Caution: Event-Study Estimates after Testing for Parallel Trends.” American Economic Review: Insights 4(3), 305–322.
- Rambachan, A., and Roth, J. (2023). “A More Credible Approach to Parallel Trends.” Review of Economic Studies 90(5), 2555–2591.
- Borusyak, K., Jaravel, X., and Spiess, J. (2024). “Revisiting Event-Study Designs: Robust and Efficient Estimation.” Review of Economic Studies 91(6), 3253–3285.
- Simulation output above was generated in R 4.6.1 with
set.seed(20260826)using base R only; the code shown reproduces the printed figures exactly. Bibliographic details for every citation above were verified against the publisher record; the de Chaisemartin–D’Haultfœuille, Roth, Goodman-Bacon, Callaway–Sant’Anna and Sun–Abraham statements quoted are taken from the papers’ own published abstracts.








