📈 Economics · Graduate · ECON 410

Econometrics

A graduate course in econometrics organised around one question: what would have to be true for this number to be a causal effect? It starts with potential outcomes, writes the gap between a treated and an untreated average as an equation, and shows exactly which term is the causal effect and which is selection bias. Least squares is then derived from the first-order conditions and computed by…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 17 lessons. The full lesson text is below so you can read it right here.

Module 1: The Causal Question and the Least Squares Machine

Write a causal claim in potential-outcomes notation, split an observed difference into effect and bias, then derive least squares and compute one by hand before asking what its assumptions are worth.

Going to Hospital Makes You Sicker: Potential Outcomes and Selection Bias

  • Write an individual causal effect, the ATE, the ATT and the ATU in potential-outcomes notation.
  • Decompose an observed difference in means into a treatment effect plus a selection bias term, and verify the decomposition on a dataset where both potential outcomes are visible.
  • State what random assignment does to the bias term, and what SUTVA rules out.

In the 2005 National Health Interview Survey, respondents were asked to rate their own health on a five-point scale, and separately whether they had spent a night in hospital in the previous twelve months. The people who had been in hospital averaged about 3.2. The people who had not averaged about 3.9. The gap is roughly 0.7 points, and with tens of thousands of respondents its standard error is about 0.03, so it is nowhere near noise.

Read literally, the number says that going to hospital makes you sicker. Nobody believes that, and everybody knows why: the people who go to hospital were already sick. But knowing why is not the same as knowing where the sentence goes wrong. By the end of this lesson you will be able to write the hospital comparison as an equation with two terms, point at the term that is a causal effect, point at the term that is not, and say exactly what would have to be true for the second term to vanish.

Two outcomes for every person, one of which you never see

Fix a person, call them i, and a treatment, call it D, coded 1 for treated and 0 for not. Attached to that person are two numbers. Yi(1) is the outcome that would occur if they were treated. Yi(0) is the outcome that would occur if they were not. These are called potential outcomes, and the framework built on them is the Rubin causal model, developed by Donald Rubin from earlier work by Jerzy Neyman on agricultural experiments.

Both numbers exist for every person. That is a substantive claim, not a notation: it says the question "what would have happened to this person otherwise" has an answer, whether or not anyone can observe it. The individual causal effect is the difference:

deltai = Yi(1) - Yi(0)

What you actually observe is one of the two, selected by the treatment the person received:

Yi = Di Yi(1) + (1 - Di) Yi(0)

Substitute Di = 1 and you get Yi(1); substitute 0 and you get Yi(0). The other number is not merely unmeasured. It did not happen. Paul Holland named this the fundamental problem of causal inference in 1986: deltai is never observed for anyone, so no amount of data collection recovers it.

What can be recovered are averages, and three are worth naming. The average treatment effect is ATE = E[Y(1) - Y(0)], over everybody. The average treatment effect on the treated is ATT = E[Y(1) - Y(0) | D = 1]. The average treatment effect on the untreated is ATU = E[Y(1) - Y(0) | D = 0]. These coincide only when the effect is identical for everyone, and much of the second half of this course is about which of the three a given design actually delivers.

The equation the hospital number is hiding

What you can compute from data is the difference in observed averages between treated and untreated. Substitute the observation rule.

E[Y | D = 1] - E[Y | D = 0] = E[Y(1) | D = 1] - E[Y(0) | D = 0]

Now the trick that carries the framework: add and subtract E[Y(0) | D = 1], the average outcome the treated group would have had untreated. That quantity is unobservable, which is precisely why adding and subtracting it is legal and useful.

= (E[Y(1) | D = 1] - E[Y(0) | D = 1]) + (E[Y(0) | D = 1] - E[Y(0) | D = 0])

The first bracket is the ATT: the same group, compared with itself under the two treatments. The second bracket is selection bias: the difference between the two groups in what would have happened to them anyway. So

naive difference = ATT + selection bias

Now the hospital number reads properly. E[Y(0) | D = 1] is the health the hospitalised group would have reported without hospitalisation, and it is low, because they were ill. E[Y(0) | D = 0] is the health of the non-hospitalised group, and it is higher. The selection bias term is large and negative. It is large enough to swamp a positive ATT and leave the observed difference negative. Nothing about the treatment has been measured. What has been measured is who goes to hospital.

The point: every observational comparison you will ever compute equals a causal effect plus a bias term, and the entire discipline consists of designs that force the bias term to zero or bound it.

Ten units with both columns visible

The table below is a fiction in one specific way: it shows both potential outcomes for all ten units, which no dataset ever does. Everything else is honest arithmetic.

UnitY(0)Y(1)EffectDObserved Y
15105110
248418
369319
47136113
535215
6810208
7910109
810133010
911121011
1012153012

Work the four quantities. Units 1 to 5 are treated, so ATT = (5 + 4 + 3 + 6 + 2) / 5 = 4. Units 6 to 10 are not, so ATU = (2 + 1 + 3 + 1 + 3) / 5 = 2. Pooling, ATE = 30 / 10 = 3. Every unit is helped; the smallest individual effect is 1.

Now compute what an analyst with real data would compute. The observed mean among the treated is 45 / 5 = 9; among the untreated it is 50 / 5 = 10. The naive difference is -1. The treatment helps every single unit, and the data say it hurts.

The decomposition accounts for the whole gap. Selection bias is E[Y(0) | D = 1] - E[Y(0) | D = 0] = (5 + 4 + 6 + 7 + 3) / 5 - 10 = 5 - 10 = -5. And indeed ATT + selection bias = 4 + (-5) = -1, exactly the naive difference. The units that got treated were the ones with the worst untreated prospects, by five whole points, and that single fact reverses the sign of the answer.

A third term, hiding inside the second

The two-term version is stated in terms of the ATT. If you want the ATE instead, one more term appears. Let pi = P(D = 1) be the treated share. Since ATE = pi (ATT) + (1 - pi)(ATU), rearranging gives ATT = ATE + (1 - pi)(ATT - ATU), and substituting into the decomposition:

naive difference = ATE + selection bias + (1 - pi)(ATT - ATU)

The last term is heterogeneous treatment effect bias: nonzero whenever the people who took the treatment gained more, or less, from it than the people who did not. Check it on the table: pi = 0.5, so the term is 0.5 (4 - 2) = 1, and 3 + (-5) + 1 = -1. The books balance. When you meet the local average treatment effect in Module 4 you will be meeting this term in a different costume: a design can be perfectly clean and still deliver an average over a subpopulation you did not choose.

What randomisation actually does

Suppose D is assigned by a coin flip, independently of everything about the person. Then D is independent of the pair (Y(0), Y(1)), which is written (Y(0), Y(1)) independent of D and called ignorability or unconfoundedness. Independence means conditioning on D changes nothing, so

E[Y(0) | D = 1] = E[Y(0) | D = 0] = E[Y(0)]

The bias term is a difference between two things that are now equal, so it is zero. The same argument applied to Y(1) makes ATT = ATU = ATE, so the third term vanishes too. A randomised experiment does one job: it makes the untreated group a valid stand-in for the treated group's counterfactual.

Note what randomisation does not promise. It does not promise that the two groups are balanced in any particular trial; it promises that the assignment mechanism is independent of potential outcomes, so the bias term has expectation zero across repetitions. In a single trial with forty people, one group can easily end up older, sicker or richer. That distinction returns in Lesson 17, where Angus Deaton and Nancy Cartwright build a serious case against treating any one randomised trial as self-evidently authoritative.

SUTVA: the assumption nobody writes down

The notation Yi(1) smuggles in two claims. First, that unit i's outcome depends on unit i's treatment and nobody else's: no interference between units. Second, that there is only one version of the treatment, so Yi(1) is a single well-defined number. Together these are the stable unit treatment value assumption, SUTVA, named by Rubin.

Both fail routinely. A job training programme whose graduates take the scarce jobs a control-group member would otherwise have got violates no-interference: the control outcome depends on other units' treatment, and the experiment measures displacement as well as training. A vaccine trial in a small town violates it through herd immunity, in the opposite direction. "Attending college" is not one treatment but hundreds. When you read a paper and cannot say what Yi(1) means for a specific person, the paper has not yet asked a causal question.

Common misconceptions

  • "Selection bias means the sample was not representative." Two problems share the name. Unrepresentative sampling damages external validity. Selection bias in the sense above damages internal validity: the comparison inside your own sample is not a causal effect, no matter how many people you interview. The hospital comparison would still be wrong with the entire United States in the sample.
  • "With enough control variables the bias goes away." Controls can shrink the bias term when the right variables are available and measured. Whether they do is an assumption about unobservables, not a fact about your regression, and Lessons 6 to 8 show controls making bias larger as often as smaller.
  • "If the treatment helps everybody, the naive comparison must at least get the sign right." The table above is a counterexample built to order. Every unit gains, and the comparison is negative.
  • "Potential outcomes are just notation for the same old regressions." The notation forces you to name the counterfactual before choosing an estimator, which is why it catches questions with no answer. Ask what the potential outcome is for being born in 1955 and the framework tells you, correctly, that there is no manipulation and so no effect to estimate.

Where this leaves us

Two potential outcomes per unit, one never observed. Three averages worth distinguishing, and no reason for them to agree. One decomposition, naive = ATT + selection bias, extended to naive = ATE + selection bias + (1 - pi)(ATT - ATU) when effects vary, checked line by line on ten units where the treatment helped everyone and the data said it hurt. Randomisation earns its status by zeroing the bias term in expectation, not by balancing any particular trial. Why this matters: from here on, judging a study means asking what its design does to that bias term, and never asking whether its regression was estimated correctly.

The next lesson sets the framework aside and builds the tool that will do the estimating. Least squares is not a causal method; it is an arithmetic operation on a cloud of points. Knowing exactly what that arithmetic does, and what it does not, is the precondition for everything after it.

Sources

  1. Cunningham, S. (2021). Potential outcomes and randomization. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  2. Huntington-Klein, N. (2022). Treatment effects. The Effect: An Introduction to Research Design and Causality. Chapman and Hall. theeffectbook.net
  3. Wikipedia contributors. (n.d.). Rubin causal model. Wikipedia. en.wikipedia.org
  4. Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association, 81(396), 945-960.
  5. Angrist, J. D., & Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist's Companion, Chapter 1. Princeton University Press.
  6. Imbens, G. W., & Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction, Chapters 1-3. Cambridge University Press.
Key terms
Potential outcome
The value the outcome would take for a given unit under a specified treatment, written Y(1) or Y(0); only one is ever observed.
Fundamental problem of causal inference
Holland's observation that an individual causal effect is never observable, because it is a difference between two outcomes only one of which occurs.
ATE
Average treatment effect, E[Y(1) - Y(0)] over the whole population.
ATT
Average treatment effect on the treated, E[Y(1) - Y(0) | D = 1]; the quantity most quasi-experimental designs actually estimate.
Selection bias
E[Y(0) | D = 1] - E[Y(0) | D = 0]: how the treated and untreated groups would have differed even with no treatment.
Ignorability
Independence of treatment assignment from the potential outcomes, which sets the selection bias term to zero.
SUTVA
Stable unit treatment value assumption: no interference between units, and only one version of each treatment.
Heterogeneous treatment effect bias
The term (1 - pi)(ATT - ATU) that separates the ATT from the ATE when the treatment helps different people by different amounts.

Where -1.5 Comes From: Least Squares Derived and Computed by Hand

  • Derive the bivariate OLS slope and intercept from the first-order conditions of the sum of squared residuals.
  • Compute a slope, intercept, fitted values, residuals, R-squared, standard error, t-statistic and confidence interval by hand on an eight-row dataset.
  • State what the two normal equations force to be true of the residuals, and why R-squared says nothing about causality.

Before you read: write down, in one line, what you would have to believe about eight school districts for a fitted slope of class size on test scores to be the effect of class size. Keep the line. The last section of this lesson checks it.

Here are eight school districts. The smallest classes have 14 students per teacher, the largest 26. Average test scores run from 681 to 701. Fit a straight line through those sixteen numbers by ordinary least squares and the slope comes out at exactly -1.5 points per additional student per teacher, with a t-statistic of -7.35 and an R-squared of 0.90. This lesson derives that -1.5, computes every number behind it by hand, and then says carefully what it is a fact about.

The problem OLS actually solves

You have n pairs (xi, yi) and want the line y = b0 + b1x that fits best. Best has to be defined before it can be found. Least squares defines it as the line minimising the sum of squared vertical distances:

S(b0, b1) = sumi (yi - b0 - b1xi)2

Squaring is a choice, not a law. It punishes a residual of 4 sixteen times as hard as a residual of 1, which makes the fit sensitive to outliers; minimising absolute deviations gives a different, more robust line. Squares win because the estimator comes out as a smooth closed form, and because of the optimality result in the next lesson.

Minimise by setting both partial derivatives to zero. Differentiating with respect to b0:

dS/db0 = -2 sumi (yi - b0 - b1xi) = 0, so sumi ei = 0

Differentiating with respect to b1:

dS/db1 = -2 sumi xi(yi - b0 - b1xi) = 0, so sumi xi ei = 0

These two are the normal equations, and they are worth reading as statements rather than as steps. The first says the residuals sum to zero, so the line passes through the point of means (xbar, ybar). The second says the residuals are uncorrelated with the regressor: whatever x contained about y, OLS has taken all of it in a linear sense. Both are mechanical consequences of the arithmetic, and neither is evidence of anything about the world.

Solve them. The first gives b0 = ybar - b1 xbar. Substituting into the second, and writing deviations as dxi = xi - xbar and dyi = yi - ybar, collapses everything to

b1 = sumi dxi dyi / sumi dxi2 = Sxy / Sxx

Key idea: the slope is a covariance divided by a variance. Everything OLS ever does, including multiple regression, instrumental variables and fixed effects, is a version of that ratio applied to some transformed version of the data.

Eight districts, every column shown

Here is the dataset, x the students per teacher and y the district's average test score. Work down the columns yourself; the point of a table this small is that nothing hides inside software.

Districtxydxdydx times dydx squareddy squared
114701-611-6636121
216695-45-201625
318691-21-241
42069101001
5206890-1001
6226892-1-241
7246814-9-361681
8266836-7-423649
Sum160552000-168112280

The means are xbar = 160 / 8 = 20 and ybar = 5520 / 8 = 690. The two sums that matter are Sxy = -168 and Sxx = 112. So

b1 = -168 / 112 = -1.5 and b0 = 690 - (-1.5)(20) = 690 + 30 = 720

The fitted line is yhat = 720 - 1.5x. Evaluate it at each x and subtract to get residuals.

Districtxyyhatex times ee squared
1147016992284
216695696-1-161
318691693-2-364
4206916901201
520689690-1-201
6226896872444
724681684-3-729
8266836812524
Sum0028

Both normal equations check out on the page: the residuals sum to zero, and so does sum xiei. That is the cheapest audit of a hand computation there is.

How much of the variation the line explains, and why that is not the question

Three sums of squares partition the outcome. TSS = sum dy2 = 280, SSR = sum e2 = 28, and the explained sum is the difference, ESS = 252, which you can also get as b12 Sxx = 2.25 times 112. Then

R2 = ESS / TSS = 252 / 280 = 0.90

The line accounts for ninety percent of the variation in district scores. That is a fact about fit and nothing else. In Lesson 6 these same eight districts produce a class-size coefficient of -0.5 once one further variable is added. A high R-squared and a badly biased coefficient live together comfortably.

Note also the identity b1 = r (sy / sx). Here sx = sqrt(112/7) = 4, sy = sqrt(280/7) = 6.3246 and r = -168 / sqrt(112 times 280) = -0.9487, so -0.9487 times 1.5811 = -1.5. Correlation and slope carry the same information on different scales.

The standard error, from residuals to a confidence interval

The slope is an estimate, so it has a sampling distribution. Under the assumptions of the next lesson its variance is Var(b1) = sigma2 / Sxx. You do not know sigma2, so estimate it from the residuals, dividing by n - k with k = 2 estimated coefficients:

s2 = SSR / (n - 2) = 28 / 6 = 4.6667, so s = 2.1602

SE(b1) = s / sqrt(Sxx) = 2.1602 / sqrt(112) = 2.1602 / 10.5830 = 0.2041

The t-statistic against a null of zero is t = -1.5 / 0.2041 = -7.35 on n - 2 = 6 degrees of freedom. The two-sided 5 percent critical value from the t distribution with 6 degrees of freedom is 2.447, so the 95 percent interval is

-1.5 plus or minus 2.447 times 0.2041 = -1.5 plus or minus 0.50, that is [-2.00, -1.00]

Note where Sxx sits: in the denominator. Spread the x values further apart and the slope is estimated more precisely, which is why an experiment assigning extreme doses beats one assigning a narrow band. Remember: precision comes from variation in the regressor, not from sample size alone.

OLS as a weighted average of the outcomes

One rewriting will matter repeatedly. Because sum dxi ybar = ybar sum dxi = 0, the numerator sum dxi dyi equals sum dxi yi. Therefore

b1 = sumi wi yi where wi = dxi / Sxx

The weights depend only on x: here -6/112, -4/112, -2/112, 0, 0, 2/112, 4/112, 6/112. Two consequences follow. OLS is linear in y, which is the property the Gauss-Markov theorem needs. And districts 4 and 5, sitting exactly at the mean class size, contribute nothing to the slope. Observations far from xbar carry the estimate, which is what leverage means and why one distant point can move a regression a long way.

What the -1.5 is a fact about

It is a fact about these eight districts and this arithmetic: among them, districts with one more student per teacher scored 1.5 points lower on average. Nothing in the derivation mentioned causes. Districts with large classes are not a randomly chosen half of the districts; they are typically poorer, and poorer districts would have scored lower at any class size. That is E[Y(0) | large classes] < E[Y(0) | small classes], Lesson 1's selection bias term, sitting inside the -1.5 with no label on it. Check the line you wrote before reading: to treat -1.5 as the effect of class size you must believe class size is unrelated to everything else that moves scores. Lesson 6 puts a number on how false that is.

Common misconceptions

  • "A high R-squared means the model is right." R-squared measures fit, and fit is a property of the sample and the functional form. Our R-squared of 0.90 sits on a coefficient that is off by a factor of three. Conversely, a well-identified experiment with a genuinely small effect can report an R-squared near 0.01 and still be the best causal evidence available.
  • "The residuals being uncorrelated with x shows the model has no omitted variables." Zero correlation between e and x is forced by the second normal equation. It holds by construction in every OLS regression ever run, including badly misspecified ones, so it can never be evidence.
  • "Least squares assumes the errors are normal." The estimator itself assumes nothing; it is a minimiser. Normality is used only to justify exact small-sample t and F distributions.
  • "More data always tightens the estimate." The standard error depends on Sxx, which is n times the variance of x. Adding a thousand districts that all have exactly 20 students per teacher raises n and leaves Sxx untouched.

What to carry forward

Least squares minimises squared vertical distances; the two first-order conditions are the normal equations; and they deliver b1 = Sxy / Sxx and b0 = ybar - b1 xbar. On eight districts that is -168 / 112 = -1.5 and an intercept of 720, with SSR = 28, TSS = 280, R-squared = 0.90, SE = 0.2041, t = -7.35 and a 95 percent interval of [-2.00, -1.00]. The slope is a weighted average of the outcomes with weights proportional to xi - xbar, which makes it linear in y and gives distant observations the most influence. And it is, so far, a description of eight districts.

The next lesson asks what has to be true of the errors for that description to also be an unbiased estimate of something, and what each of those conditions buys. Two of them are cheap, one is the whole ball game, and one buys nothing you need.

Sources

  1. Hansen, B. E. (2022). Least squares regression. Econometrics, Chapters 3-4. Princeton University Press. users.ssc.wisc.edu
  2. Wikipedia contributors. (n.d.). Ordinary least squares. Wikipedia. en.wikipedia.org
  3. MIT OpenCourseWare. (2007). 14.32 Econometrics, Spring 2007. Massachusetts Institute of Technology. ocw.mit.edu
  4. Wooldridge, J. M. (2019). Introductory Econometrics: A Modern Approach (7th ed.), Chapter 2. Cengage.
  5. Stock, J. H., & Watson, M. W. (2019). Introduction to Econometrics (4th ed.), Chapters 4-5. Pearson.
Key terms
Normal equations
The two first-order conditions of least squares: the residuals sum to zero and are uncorrelated with the regressor.
Residual
e = y - yhat, the vertical distance from a data point to the fitted line; not the same object as the unobserved error.
TSS, ESS, SSR
Total, explained and residual sums of squares, satisfying TSS = ESS + SSR for a regression with an intercept.
R-squared
ESS divided by TSS: the share of the outcome's variation the fitted line reproduces. A measure of fit, never of identification.
Standard error of the slope
s divided by the square root of Sxx, where s squared is SSR/(n-2); it falls as the regressor's spread rises.
Leverage
The influence an observation has on the slope, proportional to how far its x value sits from the mean of x.
Point of means
The pair (xbar, ybar), through which every OLS line with an intercept must pass.

Five Assumptions, Four Prices: What Gauss-Markov Actually Buys

  • State the Gauss-Markov assumptions and prove that OLS is unbiased under the zero conditional mean assumption.
  • Rank the assumptions by what their failure costs, distinguishing failures that break the coefficient from failures that only break the standard error.
  • Explain why E[u|x] = 0 is not testable from the residuals and must be argued from design.

In 1821 Carl Friedrich Gauss published a proof that among all estimators that are linear in the outcome and unbiased, least squares has the smallest variance. Andrey Markov restated it in 1900, and the result carries both names. Notice what the sentence does not contain. It says nothing about causality. It says nothing about normal errors. It does not say the fitted line is right. It says that within a narrow class of estimators, and given a list of conditions, this one is the most precise.

That list of conditions is the subject of this lesson, and the useful way to hold it is not as five things you hope are true. It is as five things with wildly different prices. One of them, if it fails, means your coefficient is estimating the wrong number and no amount of data or clever standard errors will save it. Two of them, if they fail, cost you nothing but a corrected standard error. One is a definitional requirement. One is not on the list at all and is often mistaken for being on it.

The five conditions

Write the model as yi = beta0 + beta1xi + ui, where ui is everything other than x that moves y. The error u is not the residual e. The residual is what your fitted line leaves over and is computable; the error is a feature of the world and is not.

  • A1 Linearity in parameters. The model is linear in beta, though x may enter as a log, a square or an interaction.
  • A2 Random sampling. The observations are independent draws from the population.
  • A3 Variation in the regressor. Sxx > 0; with several regressors, no exact linear dependence among them.
  • A4 Zero conditional mean. E[ui | xi] = 0. Knowing x tells you nothing about the average of everything else.
  • A5 Homoskedasticity. Var(ui | xi) = sigma2, the same at every value of x.

A sixth, normality of u, is sometimes appended. It is not needed for unbiasedness, not needed for Gauss-Markov, and not needed for large-sample inference. It buys exact t and F distributions in small samples, and that is all it buys.

Unbiasedness, in six lines

Recall from Lesson 2 that b1 = sum wiyi with wi = dxi/Sxx. Two properties of those weights do all the work. First, sum wi = 0, because the deviations sum to zero. Second, sum wixi = 1, because sum dxixi = sum dxi(dxi + xbar) = Sxx. Substitute the model:

b1 = sum wi(beta0 + beta1xi + ui) = beta0(0) + beta1(1) + sum wiui = beta1 + sum wiui

The estimate equals the truth plus a weighted sum of the errors. Take the expectation conditional on the x values. The weights are functions of x alone, so they pass through, and A4 kills each term:

E[b1 | x] = beta1 + sum wiE[ui | x] = beta1

That is the entire proof, and A4 is the only assumption it used beyond A1 and A3. What matters here: unbiasedness rests on one condition, and it is a condition about unobservables. A2 and A5 have not appeared yet.

They appear next, in the variance. Under A2 the errors are independent, so the variance of the sum is the sum of the variances, and under A5 each of those is sigma2:

Var(b1 | x) = sum wi2 Var(ui | x) = sigma2 sum wi2 = sigma2 / Sxx

because sum wi2 = sum dxi2/Sxx2 = 1/Sxx. That is the formula behind the 0.2041 you computed in Lesson 2, and it is the formula that stops being correct the moment A2 or A5 fails.

The theorem itself follows in a line. Let any competing linear unbiased estimator have weights ci = wi + di. Unbiasedness for all beta forces sum di = 0 and sum dixi = 0, which also makes sum widi = 0. So its variance is sigma2(sum wi2 + sum di2), which exceeds sigma2/Sxx unless every di is zero. OLS is the best linear unbiased estimator, BLUE.

The price list

AssumptionWhat it buysIf it failsRepairable?
A1 LinearityA slope interpretable as a constant marginal effectThe fit is still the best linear approximation to the true conditional mean, but the coefficient answers a different questionYes: logs, polynomials, splines, interactions
A2 Random samplingIndependence, hence the variance formula and valid t-testsCoefficient is fine; standard errors can be far too smallYes: clustered or autocorrelation-robust standard errors
A3 Variation in xThe estimator exists and is uniqueNothing to estimate; software drops the variableDefinitional: get variation or drop the question
A4 E[u|x] = 0Unbiasedness and consistencyThe coefficient converges to the wrong number, by exactly Cov(x,u)/Var(x)No. Only a research design repairs it
A5 HomoskedasticityEfficiency, and the simple variance formulaCoefficient is fine; classical standard errors are wrong in either directionYes: heteroskedasticity-robust standard errors
Normality (not required)Exact small-sample t and FNothing, once n is moderateYes: large-sample approximations

Read down the third column. Four rows describe inconveniences. One row describes the failure of the enterprise. The upshot: everything from Module 3 onward exists to defend A4, and everything in Module 2 exists to repair A2 and A5.

What A4 says in the language of Lesson 1

Suppose A4 fails because x is correlated with something in u. Take the plim of the estimator using b1 = beta1 + sum wiui and the law of large numbers:

plim b1 = beta1 + Cov(x, u) / Var(x)

With more data the estimate converges, but not to beta1. It converges to beta1 plus a term you cannot see, in a direction set by the sign of the covariance. In the eight-district data of Lesson 2, income sits inside u and is negatively correlated with class size. In Lesson 6 that term will turn out to be exactly -1.0, which is why -0.5 is measured as -1.5.

This connects directly to the potential-outcomes decomposition. Coding x as a binary treatment D, the statement E[u|D] = 0 is the statement that E[Y(0)|D=1] = E[Y(0)|D=0]. A4 and "no selection bias" are the same sentence in two dialects.

One refinement, because it matters later. A4 is mean independence. A weaker condition, Cov(x, u) = 0, is enough for consistency but not for unbiasedness in finite samples, and it is what instrumental variables arguments typically deliver. In panel and time-series settings a stronger version is required: strict exogeneity, E[uit | xi1, ..., xiT] = 0, which rules out the past outcome influencing future treatment. That distinction is what makes a lagged dependent variable dangerous inside a fixed effects model, as Lesson 11 will show.

Why you cannot test the assumption that matters

Every diagnostic you can run uses residuals, and the residuals were constructed to satisfy sum xiei = 0. Plot e against x for a hopelessly confounded regression and you will find no pattern, because the second normal equation removed it. A Breusch-Pagan or White test can detect heteroskedasticity, since A5 concerns the spread of the residuals rather than their mean. A Durbin-Watson statistic can detect serial correlation. Nothing detects A4 failure, because the quantity in question, E[u|x], involves the unobserved error rather than the residual.

This is the structural reason econometrics turned into a discipline about research design rather than about estimators. If the decisive assumption were testable, the field would be a catalogue of tests. It is not, so the field is a catalogue of designs, each of which makes A4 credible by construction rather than by assertion: lotteries, thresholds, policy changes, and instruments.

Common misconceptions

  • "Robust standard errors fix a biased coefficient." They fix the second moment of the estimator's sampling distribution. A biased coefficient is a problem with the first moment. Robust standard errors around a wrong number give you a precise wrong number.
  • "Unbiased means close to the truth." Unbiasedness says the estimator is right on average across repeated samples. A single estimate from a single sample can be far away, and an unbiased estimator with a huge variance is often worse in practice than a slightly biased one with a small variance, which is the whole logic of ridge regression and shrinkage.
  • "BLUE means best." Best within the class of linear unbiased estimators, given A1 to A5. Drop unbiasedness and shrinkage estimators beat OLS on mean squared error. Drop linearity and, under non-normal errors, median regression can beat it. The theorem is a statement about a competition with a restricted entry list.
  • "A good R-squared or a passing residual plot supports the causal reading." Both are computed from residuals, which are orthogonal to x by construction. They cannot speak to A4 even in principle.

The short version

Five assumptions, and they are not equals. A3 makes the estimator exist. A1 makes the coefficient mean what you want it to mean. A2 and A5 make the standard error correct and are repairable with the tools of the next two lessons. A4, E[u|x] = 0, makes the coefficient an estimate of the parameter rather than of the parameter plus Cov(x,u)/Var(x), is the same claim as "no selection bias" in potential-outcomes language, and cannot be tested with residuals because the residuals are built to satisfy the wrong condition. Gauss-Markov gives you BLUE, a genuinely useful optimality result inside a restricted class, and no causal content whatever.

Module 2 takes A2 and A5 in turn: what a standard error is estimating, why heteroskedasticity-robust standard errors are the default, and why in panel data clustering is not a refinement but a requirement.

Sources

  1. Wikipedia contributors. (n.d.). Gauss-Markov theorem. Wikipedia. en.wikipedia.org
  2. Hansen, B. E. (2022). The algebra of least squares and least squares regression. Econometrics, Chapters 3-5. Princeton University Press. users.ssc.wisc.edu
  3. Huntington-Klein, N. (2022). Regression and its assumptions. The Effect: An Introduction to Research Design and Causality. Chapman and Hall. theeffectbook.net
  4. Wooldridge, J. M. (2019). Introductory Econometrics: A Modern Approach (7th ed.), Chapters 2-3 and Appendix E. Cengage.
  5. Angrist, J. D., & Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist's Companion, Chapter 3. Princeton University Press.
Key terms
Error versus residual
The error u is the unobserved deviation in the population model; the residual e is what a fitted line leaves over in a sample. Diagnostics see only the second.
Zero conditional mean (A4)
E[u|x] = 0: the regressor carries no information about the average of everything omitted. The assumption that makes OLS unbiased.
Homoskedasticity (A5)
Constant error variance across values of x. Buys efficiency and the simple variance formula, not unbiasedness.
BLUE
Best linear unbiased estimator: minimum variance among estimators that are linear in y and unbiased, given A1 to A5.
Strict exogeneity
E[u_it | x_i1, ..., x_iT] = 0: the error in one period is unrelated to the regressor in every period, past and future. Required by fixed effects.
Asymptotic bias
plim b1 - beta1 = Cov(x,u)/Var(x): the number OLS converges to in excess of the truth when A4 fails.
Mean independence
E[u|x] = 0, a stronger condition than zero covariance, which is why unbiasedness is a stronger claim than consistency.

Module 2: Inference, and the Errors That Break It

What a standard error is estimating, why the robust version is the default and is not always larger, and why in panel data clustering stops being a refinement and becomes a requirement.

Three Standard Errors from One Regression: Robust Inference Debugged

  • Derive the sandwich variance formula and compute HC0 and HC1 standard errors by hand on an eight-row regression.
  • Explain why a heteroskedasticity-robust standard error is sometimes smaller than the classical one, and what that means in small samples.
  • Identify the settings where heteroskedasticity is present by construction, and state what robust standard errors do not fix.

Here is a claim that is wrong in an instructive way: robust standard errors are the cautious choice, because they are bigger. Take the eight districts from Lesson 2. The classical standard error of the slope is 0.2041. The Eicker-Huber-White estimator, computed below with no degrees-of-freedom correction, gives 0.1956, which is smaller. The version most software reports by default gives 0.2259, which is larger. Three numbers, one regression, one dataset. Working out which to believe, and why they differ, is the content of this lesson.

What the number is supposed to be

A standard error is the estimated standard deviation of an estimator's sampling distribution. Imagine drawing a fresh sample of eight districts from the same population, refitting, and recording the slope; do it ten thousand times and you get a distribution of slopes centred, under A4, on the true beta1. The standard error estimates the spread of that distribution from the one sample you actually have.

Lesson 3 derived it. Since b1 = beta1 + sum wiui with wi = dxi/Sxx, and independent observations make the variance of the sum the sum of the variances,

Var(b1 | x) = sumi wi2 sigmai2 = (sumi dxi2 sigmai2) / Sxx2

Nothing has been assumed here except independence. Every observation is allowed its own error variance sigmai2. The classical formula appears only when you impose A5, setting every sigmai2 equal to a common sigma2, which pulls it out of the sum and leaves sigma2/Sxx. That step is the whole of the homoskedasticity assumption, and it is doing more than it looks: it says a district with 26 students per teacher has exactly the same amount of unexplained variation in its score as a district with 14.

The sandwich, and why it is called that

Halbert White's 1980 paper made the obvious move. You cannot observe sigmai2, but you have one draw from a distribution with that variance, namely ei. Substituting ei2 for sigmai2 gives

VarHC0(b1) = (sumi dxi2 ei2) / Sxx2

A single squared residual is a hopeless estimate of a single variance. The estimator works anyway, because what is needed is not each sigmai2 but the weighted sum, and the sum of many noisy pieces converges. In matrix notation the general form is (X'X)-1 (X' diag(e2) X) (X'X)-1, a meat term between two identical bread terms, which is where the sandwich name comes from. Under homoskedasticity the meat collapses to sigma2X'X and the sandwich reduces to the classical formula, which is why robust standard errors cost you nothing asymptotically even when A5 holds.

The three numbers, computed

From Lesson 2: Sxx = 112, SSR = 28, n = 8, and the residuals were 2, -1, -2, 1, -1, 2, -3, 2.

Districtdxedx squarede squareddx squared times e squared
1-62364144
2-4-116116
3-2-24416
401010
50-1010
6224416
74-3169144
862364144
Sum11228480

Classical. s2 = 28/6 = 4.6667, so Var = 4.6667/112 = 0.041667 and SE = 0.2041. The t-statistic is -7.35.

HC0. Var = 480 / 1122 = 480/12544 = 0.038265, so SE = 0.1956 and t = -7.67.

HC1. Multiply HC0 by n/(n-k) = 8/6 = 1.3333: Var = 0.051020, SE = 0.2259, t = -6.64. This is what Stata's robust option reports and what R's sandwich package gives with type HC1.

Now debug the original claim. HC0 came out below the classical number for a reason that is arithmetic rather than deep: the classical formula divides SSR by n - 2 = 6, inflating the average squared residual from 3.5 to 4.667, while HC0 uses the raw squared residuals with no correction at all. With n = 8, that correction is a 33 percent difference in the variance and swamps everything else. Once HC1 restores the same correction, the robust standard error is larger, as the pattern in the data suggests it should be: the biggest squared residual, 9, sits at dx = 4, a high-leverage point where the classical formula assumed only average noise.

In short: robust standard errors are not automatically bigger, and in samples this small the difference between HC0 and HC1 is not a detail. HC2 and HC3, which divide each squared residual by 1 - hi or its square using the leverage hi, correct further and are the ones to prefer below roughly fifty observations.

Where heteroskedasticity is not a possibility but a certainty

Three settings guarantee it, so in these cases testing for it is a waste of a paragraph.

  • Binary outcomes. Fit a linear probability model and the error is 1 - p(x) with probability p(x) and -p(x) otherwise, so Var(u|x) = p(x)(1 - p(x)). At p = 0.5 that is 0.25; at p = 0.05 it is 0.0475, five times smaller. Heteroskedasticity is built into the model definition.
  • Grouped or averaged data. If each row is a state average over ng people, its sampling variance is proportional to 1/ng. Wyoming's average is far noisier than California's, and a regression across states that ignores this treats them as equally informative.
  • Anything measured in money against anything measured in money. The spread of household spending around its conditional mean grows with income. Taking logs often helps, precisely because it converts a multiplicative spread into an additive one.

You can test formally. Breusch-Pagan regresses e2 on the regressors; White's test adds squares and cross-products. Both are fine and both are usually beside the point, because the modern default is to report robust standard errors regardless. Under A5 they are consistent and merely a little less efficient in finite samples; without A5 they are the only ones that mean anything. Testing first and then choosing also has a subtler cost: the reported standard error is then conditional on a pre-test outcome, and its true coverage is not the nominal 95 percent.

What robust standard errors are not

Return to the eight districts. Whichever of the three numbers you pick, the estimate is -1.5, and Lesson 6 will show the causal parameter is -0.5. The confidence interval built from HC1 is -1.5 plus or minus 2.447 times 0.2259, that is [-2.05, -0.95], and it excludes the truth. It excludes the truth with 95 percent nominal confidence and it would go on excluding it with a million districts, tighter and tighter.

Bottom line: a standard error describes how much an estimate would jump around across repeated samples from the same flawed design. It has nothing to say about whether the design was flawed. Robust inference is a repair to A5 and A2, and the word robust names a narrow kind of robustness.

One efficiency footnote, since heteroskedasticity does cost you something real. Under A5 failure OLS remains unbiased but stops being BLUE, and weighted least squares, which divides each observation by its own error standard deviation, is efficient if you know the variance function. In practice you rarely do. Guessing it wrong leaves you unbiased but no longer efficient, and you have added an assumption for nothing, which is why weighting fell out of fashion for anything but survey design weights and known group sizes.

Common misconceptions

  • "Robust standard errors are always larger, so reporting them is conservative." They can be smaller, as HC0 is here. Report them because they are consistent under a weaker assumption, not because they are safer.
  • "Heteroskedasticity biases the coefficient." It does not touch the coefficient. It invalidates the classical variance formula, which is a different and more repairable problem.
  • "Robust standard errors handle any dependence in the data." The sandwich derived above assumed independent observations. Correlation within a state, a school or a household is untouched by it, and the next lesson shows how badly.
  • "With robust standard errors you no longer need to think about the model." The interval computed above is tight, robust, and wrong. Precision and validity are separate properties, and only the first is under the software's control.

Putting it together

The general variance of the OLS slope is sum dxi2 sigmai2 divided by Sxx2; imposing a common sigma2 collapses it to the classical formula, and substituting ei2 gives the sandwich. On the eight districts that is 0.2041 classical, 0.1956 for HC0 and 0.2259 for HC1, with the gap between the two robust versions coming entirely from a degrees-of-freedom correction that matters enormously at n = 8 and hardly at all at n = 8000. Heteroskedasticity is guaranteed in linear probability models, in grouped data and in most money-on-money regressions, so report robust standard errors by default and skip the test. And none of this touches the estimate itself: a robust standard error around a confounded coefficient buys you a precisely stated wrong answer.

The sandwich assumed independence across observations. The next lesson removes that assumption and finds that the cost of ignoring it is not a few percent on a standard error but a rejection rate of 45 percent on a null that is true by construction.

Sources

  1. Wikipedia contributors. (n.d.). Heteroskedasticity-consistent standard errors. Wikipedia. en.wikipedia.org
  2. Hansen, B. E. (2022). Asymptotic theory and covariance matrix estimation. Econometrics, Chapter 7. Princeton University Press. users.ssc.wisc.edu
  3. Cameron, A. C., & Miller, D. L. (2015). A practitioner's guide to cluster-robust inference. Journal of Human Resources, 50(2), 317-372. cameron.econ.ucdavis.edu
  4. White, H. (1980). A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity. Econometrica, 48(4), 817-838.
  5. Angrist, J. D., & Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist's Companion, Chapter 8. Princeton University Press.
Key terms
Standard error
The estimated standard deviation of an estimator's sampling distribution; it describes variability across hypothetical repeated samples, not correctness.
Sandwich estimator
A variance estimator of the form bread times meat times bread, where the meat uses squared residuals rather than a single common variance.
HC0
White's original heteroskedasticity-consistent variance, using raw squared residuals with no small-sample correction.
HC1
HC0 multiplied by n/(n-k), the degrees-of-freedom correction most software applies by default.
Linear probability model
OLS with a binary outcome; its error variance is p(1-p), so heteroskedasticity is present by construction.
Weighted least squares
GLS dividing each observation by its error standard deviation; efficient if the variance function is known, and pointless if it is guessed.
Breusch-Pagan test
A test for heteroskedasticity that regresses squared residuals on the regressors; largely superseded by simply reporting robust standard errors.

Forty-Five Percent: Placebo Laws and the Case for Clustering

  • Describe the Bertrand, Duflo and Mullainathan placebo experiment and explain why serial correlation produced a 45 percent rejection rate.
  • Compute a Moulton factor and an effective sample size from a group size and an intra-class correlation.
  • Choose a clustering level from the level of treatment assignment, and name the remedies available when the number of clusters is small.

Around 2001, Marianne Bertrand, Esther Duflo and Sendhil Mullainathan ran an experiment on econometrics itself. They took the Current Population Survey's outgoing rotation groups, kept women aged 25 to 50 from 1979 to 1999, and computed log weekly earnings for each of the fifty states in each of the twenty-one years. Then they invented laws. For each replication a state was drawn at random and a year between 1985 and 1995 was drawn at random, and a variable was coded 1 for that state from that year onward and 0 everywhere else. No such law existed. The variable was noise with a particular shape.

They then ran the regression an applied economist would run: log earnings on the fake law, with state fixed effects, year fixed effects, and the usual individual controls, and tested the coefficient at the 5 percent level. A correct procedure rejects a true null 5 percent of the time. Their procedure rejected it in about 45 percent of replications. Nine times out of twenty, a law that never happened produced a statistically significant effect on wages.

Why the test broke

Two facts about the data combine badly. First, state wage series are highly persistent: a state above trend in 1988 is above trend in 1989, so the regression errors within a state are strongly autocorrelated over time. Second, the fake treatment is persistent by construction: it switches on once and stays on, so the regressor is autocorrelated too. Neither alone is fatal. Together they are.

The variance formula from Lesson 4 assumed independence across observations, which let the variance of a sum become the sum of the variances. When observations within a state move together, the cross terms are positive and do not vanish. The formula therefore leaves them out, understates the variance, and produces a standard error that is too small, a t-statistic that is too large, and rejections that arrive far too often.

The intuition worth keeping is about information. A state contributes 21 rows to the dataset, but if its deviations from trend are nearly the same number 21 times, it has contributed nearly one observation's worth of independent evidence about how much state wage series bounce around. Counting it as 21 is the error.

The Moulton factor, computed

Brent Moulton gave the correction in 1986. If observations fall into equal groups of size m, and the regressor and the errors have within-group correlations rhox and rhoe, then

true SE / naive SE = sqrt(1 + (m - 1) rhox rhoe)

Work an example that matches the placebo study. Fifty states, twenty-one years each, so m = 21. A policy variable that is constant within a state at any point in the comparison has rhox at or near 1. Suppose the within-state correlation of the errors is rhoe = 0.4, which is unremarkable for wage series. Then

sqrt(1 + 20 times 0.4) = sqrt(9) = 3

The honest standard error is three times the reported one. A t-statistic of 6 is really a t-statistic of 2; a t-statistic of 3 is really a t-statistic of 1, and the paper that reported it has found nothing. Run the same arithmetic with 30 observations per cluster and rhoe = 0.2 and you get sqrt(1 + 29 times 0.2) = sqrt(6.8) = 2.61. The factor is large even when the correlation looks small, because it is multiplied by the group size.

The same fact in different units: the effective sample size is Gm / (1 + (m - 1) rho). With 50 states, 21 years and rho = 0.4, that is 1050 / 9 = 117. A dataset of 1,050 state-years carries roughly the information of 117 independent observations, and the software reporting a standard error based on 1,050 is off by a factor of three. Worth holding on to: the effective sample size is governed by the number of groups, not the number of rows.

What the cluster-robust estimator does

The repair keeps the sandwich of Lesson 4 and changes the meat. Instead of summing squared residuals observation by observation, sum the products within each cluster first, then across clusters:

Varcluster(b) = (X'X)-1 [ sumg (Xg' eg)(Xg' eg)' ] (X'X)-1

Read what that permits. Within a cluster, the correlation structure is left completely unrestricted: any pattern of correlation across the 21 years of one state is absorbed, and no model of it needs to be specified. Across clusters, independence is still assumed. Cluster-robust standard errors are therefore a trade, buying freedom within groups by asserting independence between them.

Applied to the placebo laws, clustering at the state level brought the rejection rate back to roughly the nominal 5 percent. So did the alternative Bertrand and coauthors also recommend: collapse each state's series into a single pre-treatment average and a single post-treatment average, which destroys the serial correlation problem by destroying the time series. With fifty states, either works.

What to cluster on

The default answer, from Abadie, Athey, Imbens and Wooldridge, is the level at which treatment is assigned. If a law is passed by a state, cluster by state. If a curriculum is adopted by a school, cluster by school, even if the outcome is measured on individual pupils. If a randomised trial randomises villages, cluster by village.

Two errors follow from misunderstanding that rule. Clustering too finely, for example by state-year when the policy varies at the state level, does not fix the problem the placebo experiment exposed, because the correlation being ignored is exactly the correlation across years within a state. Clustering too coarsely, for example by region when treatment varies within region, throws away power for no gain and can leave you with too few clusters for the asymptotics to work. The instinct that clustering more coarsely is always safer is wrong in both directions.

Few clusters, and what to do instead

Cluster-robust standard errors are consistent as the number of clusters G grows, not as the number of observations grows. With G small they are biased downward, sometimes badly, and the usual normal or large-sample critical values compound the error. Common practice treats 40 to 50 clusters as comfortable, though the threshold depends on how unbalanced the clusters are: a design with 40 clusters of which one contains half the data behaves much worse than 40 equal ones.

SituationProblemPractical remedy
50 states, 21 yearsSerial correlation within stateCluster by state, or collapse to pre and post averages
6 treated and 6 control regionsCluster-robust variance badly biased downWild cluster bootstrap; randomisation inference
One treated unitNo cluster-level variation to work withSynthetic control with permutation inference (Lesson 15)
Clusters of wildly unequal sizeEffective number of clusters far below GReport the effective number; use bootstrap critical values

The wild cluster bootstrap of Cameron, Gelbach and Miller is the standard fix. It resamples by flipping the sign of each cluster's residual vector under the null, refits, and builds the distribution of the t-statistic empirically rather than assuming it is normal. Randomisation inference, which reassigns the treatment across clusters in every possible or a large random subset of ways and asks where the actual estimate falls in the resulting distribution, is the other, and it returns in Lesson 15.

Common misconceptions

  • "Clustering is a robustness check." In panel data with a policy that switches on at the group level it is a requirement. The placebo experiment is the evidence: without it, a procedure with a nominal 5 percent size has an actual size near 45 percent.
  • "More clusters is always safer." Clustering at a finer level than the treatment leaves the relevant correlation uncorrected. Clustering at a coarser level than necessary costs power and can push you below the number of clusters the asymptotics need.
  • "Adding state and year fixed effects handles the correlation." Fixed effects remove the group means. They do nothing about the correlation of the deviations around those means, which is precisely what the placebo regressions already included and precisely what failed.
  • "Clustering fixes the estimate." The point estimate does not move at all. Clustering changes only the second moment, so a design that would have delivered a biased coefficient still delivers it, now with an honest interval around the wrong number.

What you now know

A fake law, randomly assigned to a random state in a random year, was found to be statistically significant 45 percent of the time by a regression specification in ordinary use. The cause was serial correlation in both the outcome and the treatment indicator, which the independence assumption behind the classical and heteroskedasticity-robust variance formulas cannot accommodate. The Moulton factor sqrt(1 + (m - 1) rhox rhoe) puts a number on the damage, tripling the standard error at m = 21 and rhoe = 0.4, and the effective sample size falls from 1,050 to 117. Cluster-robust standard errors permit arbitrary correlation within groups and assume independence between them; cluster at the level treatment is assigned; and below about forty clusters use a wild cluster bootstrap or randomisation inference instead of the asymptotic formula.

The core of it: the two lessons of this module both concern the second moment, and both leave the first moment exactly where they found it. Module 3 turns to the first moment, which is where the causal question lives.

Sources

  1. Bertrand, M., Duflo, E., & Mullainathan, S. (2004). How much should we trust differences-in-differences estimates? Quarterly Journal of Economics, 119(1), 249-275. NBER Working Paper 8841. nber.org
  2. Abadie, A., Athey, S., Imbens, G. W., & Wooldridge, J. M. (2023). When should you adjust standard errors for clustering? Quarterly Journal of Economics, 138(1), 1-35. nber.org
  3. Cameron, A. C., & Miller, D. L. (2015). A practitioner's guide to cluster-robust inference. Journal of Human Resources, 50(2), 317-372. cameron.econ.ucdavis.edu
  4. Moulton, B. R. (1986). Random group effects and the precision of regression estimates. Journal of Econometrics, 32(3), 385-397.
  5. Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. Review of Economics and Statistics, 90(3), 414-427.
Key terms
Serial correlation
Correlation of a series with its own past values; in panels it makes successive observations on one unit far from independent.
Moulton factor
sqrt(1 + (m-1) rho_x rho_e): the ratio of the correct standard error to the one computed under independence.
Intra-class correlation
The share of an outcome's variance that lies between groups rather than within them; the rho in the Moulton factor.
Effective sample size
Gm/(1 + (m-1)rho): the number of independent observations a clustered dataset is really worth.
Cluster-robust variance
A sandwich estimator whose meat sums within-cluster score vectors, allowing arbitrary correlation inside clusters and none between them.
Wild cluster bootstrap
A resampling procedure that flips the sign of each cluster's residual vector under the null to build critical values when clusters are few.
Randomisation inference
Inference by reassigning treatment across units many times and locating the actual estimate in the resulting distribution.

Module 3: Controls, Good and Bad

The omitted variable bias formula worked to a number, the conditional independence assumption and what matching can and cannot deliver, and the controls that make an estimate worse.

From -1.5 to -0.5: The Omitted Variable Bias Formula, Worked

  • Derive the omitted variable bias formula and apply it as an exact identity linking a short and a long regression.
  • Compute a two-regressor OLS fit by hand from the sums of cross-products and verify the bias decomposition to the decimal.
  • Sign and bound the bias from an omitted variable using only knowledge of two directions.

Add one column to the eight districts of Lesson 2 and the class-size coefficient moves from -1.5 to -0.5. The column is median household income, and the R-squared barely moves: 0.900 becomes 0.929. Two thirds of the estimated effect of class size was never about class size, and the fit statistic gave no warning at all. This lesson is that computation, done fully, followed by the formula that predicts the answer before you run anything.

Two regressions, one relationship

Suppose the true model has two regressors:

y = beta0 + beta1x + beta2w + u, with E[u | x, w] = 0

and you fit the short regression, leaving w out:

y = alpha0 + alpha1x + v, where v = beta2w + u

The short regression's error contains w. If w is correlated with x, assumption A4 fails for the short regression, and Lesson 3 already told you the size of the damage: Cov(x, v)/Var(x). Expand it. Since Cov(x, u) = 0,

alpha1 = beta1 + beta2 Cov(x, w)/Var(x) = beta1 + beta2 delta

where delta is the slope from regressing the omitted variable on the included one, called the auxiliary regression. That is the omitted variable bias formula, and it is worth memorising in words: the short coefficient equals the long coefficient plus the effect of the omitted variable times its regression on the included one.

Two things about it are easy to miss. First, it is not only a statement about probability limits. Run all three regressions in a finite sample by OLS and the identity a1 = b1 + b2d holds exactly, to the last decimal, as an algebraic fact. Second, nothing in the derivation requires w to be a cause of anything. The formula is arithmetic about three regressions; whether the long coefficient deserves a causal reading is a separate question that A4 in the long regression has to answer.

The eight districts, with income restored

Here is the same dataset with median household income w, in thousands of dollars, added.

Districtxwydxdwdy
11476701-61411
21668695-465
31864691-221
42064691021
5206468902-1
622566892-6-1
724526814-10-9
826526836-10-7

The means are xbar = 20, wbar = 62, ybar = 690. Five sums of cross-products carry everything:

  • Sxx = 36 + 16 + 4 + 0 + 0 + 4 + 16 + 36 = 112
  • Sww = 196 + 36 + 4 + 4 + 4 + 36 + 100 + 100 = 480
  • Sxw = -84 - 24 - 4 + 0 + 0 - 12 - 40 - 60 = -224
  • Sxy = -66 - 20 - 2 + 0 + 0 - 2 - 36 - 42 = -168
  • Swy = 154 + 30 + 2 + 2 - 2 + 6 + 90 + 70 = 352

Step 1, the auxiliary regression. Regress w on x: d = Sxw/Sxx = -224/112 = -2, with intercept 62 - (-2)(20) = 102. Districts with one more student per teacher have median incomes two thousand dollars lower.

Step 2, the long regression. The two normal equations in deviation form are

Sxxb1 + Sxwb2 = Sxy and Sxwb1 + Swwb2 = Swy

which here are 112b1 - 224b2 = -168 and -224b1 + 480b2 = 352. From the first, b1 = -1.5 + 2b2. Substitute into the second:

-224(-1.5 + 2b2) + 480b2 = 336 - 448b2 + 480b2 = 336 + 32b2 = 352

so b2 = 0.5 and b1 = -1.5 + 1.0 = -0.5. Holding income constant, one more student per teacher costs half a point. Each extra thousand dollars of median income adds half a point.

Step 3, the identity. a1 = b1 + b2d = -0.5 + (0.5)(-2) = -0.5 - 1.0 = -1.5. The short coefficient reappears exactly. The bias is -1.0, twice the size of the parameter it is contaminating, and it is negative because a positive effect of income is being transmitted through a negative correlation between income and class size.

Step 4, the fit. In the long regression ESS = b1Sxy + b2Swy = (-0.5)(-168) + (0.5)(352) = 84 + 176 = 260, so SSR = 280 - 260 = 20 and R2 = 260/280 = 0.929. Adding the variable that cut the coefficient by two thirds raised the R-squared by three points. So what?: the fit statistic is nearly blind to the thing that matters most.

Signing the bias when you cannot measure the omitted variable

The formula's practical value is that both of its inputs can often be signed from what you already know, even when w is unmeasured. You need the direction in which w moves y, and the direction in which w moves with x.

Corr(x, w) positiveCorr(x, w) negative
Effect of w on y positiveUpward bias: short coefficient too highDownward bias: short coefficient too low
Effect of w on y negativeDownward biasUpward bias

The class-size case is the bottom-left, or equivalently the top-right: income raises scores and falls with class size, so the estimate is pushed away from zero and made to look larger in magnitude than it is. Note the word "upward" always means toward positive, not toward larger in absolute value, and mixing those up reverses the conclusion.

The classic application is the return to schooling. Regress log wages on years of education and omit ability. If ability raises wages (beta2 > 0) and is positively correlated with schooling (delta > 0), the return to schooling is overstated. That prediction is so intuitive it went unchallenged for years, and then the twins and instrumental-variable literature that David Card surveys in 1999 found that estimated returns often go up rather than down when ability is better controlled, which means either the ability story is wrong or measurement error and heterogeneous returns are pulling in the opposite direction. A sign prediction is a hypothesis, not a result.

Proxies, partial controls, and coefficient stability

You rarely have w. You often have something correlated with it. Controlling for a proxy removes the part of the bias that runs through the proxy and leaves the rest, so it moves the estimate in the right direction without arriving. Under some conditions a poor proxy can even increase the bias, which is one reason "we controlled for a rich set of covariates" is a claim rather than an argument.

A common defence is stability: the coefficient hardly moved when controls were added, so the remaining bias must be small. That inference is only valid jointly with the movement in R-squared. Emily Oster's formalisation makes the point sharply: if adding controls barely changes the coefficient and barely changes the R-squared, the controls were uninformative and you have learned nothing about the unobservables. If the coefficient barely moves while the R-squared jumps, the controls explained a lot without disturbing the estimate, which is genuinely reassuring. Our worked example is the third case, the coefficient moving a great deal while the R-squared barely does, and that is the pattern that should worry you most.

Common misconceptions

  • "Omitted variable bias always attenuates the estimate toward zero." It moves the estimate in the direction of beta2delta, which here made the coefficient three times too large. The measurement-error case, where classical error in a regressor does attenuate, is a different phenomenon that gets confused with this one.
  • "If a coefficient is still significant after adding controls, it is causal." Significance is about the standard error. The omitted variable bias formula concerns the coefficient. Adding controls can leave a precisely estimated, strongly significant, badly biased number.
  • "Add every variable you have, to be safe." The next two lessons show two ways this backfires: controlling for a variable on the causal path removes part of the effect you want, and controlling for a common consequence of treatment and outcome manufactures a correlation that was not there.
  • "A high R-squared means little is omitted." The short regression here has an R-squared of 0.900 and a coefficient off by a factor of three; the long one has 0.929. Three points of fit separated a wrong answer from a right one.

The takeaway

The short regression coefficient equals the long one plus beta2delta, exactly, in sample as well as in the limit. On eight districts, with income omitted, the auxiliary slope is -2, the income coefficient is 0.5, the bias is -1.0, and the true -0.5 is reported as -1.5 while the R-squared moves from 0.929 to 0.900. When the omitted variable is unmeasured, sign the bias from two directions and use the table; when a proxy is available, expect partial correction; and treat coefficient stability as evidence only when it is accompanied by a large movement in explained variation.

Key idea: a control is not a virtue in itself. It is a claim that a specific back-door path has been closed, and the next lesson makes that claim precise enough to test against a diagram.

Sources

  1. Wikipedia contributors. (n.d.). Omitted-variable bias. Wikipedia. en.wikipedia.org
  2. Huntington-Klein, N. (2022). Causal diagrams and statistical adjustment. The Effect: An Introduction to Research Design and Causality. Chapman and Hall. theeffectbook.net
  3. Cunningham, S. (2021). Probability and regression review. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  4. Oster, E. (2019). Unobservable selection and coefficient stability: Theory and evidence. Journal of Business and Economic Statistics, 37(2), 187-204.
  5. Card, D. (1999). The causal effect of education on earnings. In O. Ashenfelter & D. Card (Eds.), Handbook of Labor Economics (Vol. 3A, pp. 1801-1863). Elsevier.
Key terms
Short regression
The regression that omits a relevant variable, so the omitted variable lives inside its error term.
Long regression
The regression including both the variable of interest and the potential confounder.
Auxiliary regression
The regression of the omitted variable on the included one; its slope delta is the second factor in the bias formula.
Omitted variable bias
beta2 times delta: the effect of the omitted variable on the outcome, multiplied by its regression on the included regressor.
Upward bias
Bias toward positive values, which for a negative true coefficient makes the estimate smaller in magnitude, not larger.
Proxy variable
An imperfect stand-in for an unmeasured confounder; it removes part of the bias and can occasionally worsen it.
Coefficient stability
The argument that an estimate surviving added controls is credible; informative only when the added controls move the R-squared substantially.

LaLonde's Challenge: Conditional Independence, Matching, and What It Cannot Buy

  • State the conditional independence and overlap assumptions and explain why together they identify the ATT.
  • Compute a subclassification estimate of the ATT and ATE from a stratified table and compare it with the naive difference.
  • Argue both sides of the dispute over whether propensity score methods solved LaLonde's problem, and name the evidence that would settle it.

In 1986 Robert LaLonde did something unusual with an experiment: he threw away its control group. The National Supported Work Demonstration had randomly assigned disadvantaged workers to a subsidised job training programme in the mid-1970s, so an unimpeachable estimate of its effect on 1978 earnings existed. LaLonde kept the treated men, discarded the randomised controls, and replaced them with comparison groups constructed from the Current Population Survey and the Panel Study of Income Dynamics. Then he ran the non-experimental estimators of the day and asked whether any of them recovered the experimental answer.

They did not. The experimental benchmark was a modest positive gain of a few hundred dollars. The non-experimental estimates scattered across a range that included large negative numbers and positive numbers several times the benchmark, and they moved as much with the estimator as with the data. The paper was read as a demolition, and for a decade "LaLonde" was shorthand for the view that observational studies of programme effects could not be trusted.

Then a group of statisticians said: you used the wrong estimators. That disagreement is still live, and it is the subject of this lesson.

The assumption that would make it work

Everything in this lesson rests on a strengthening of Lesson 1's ignorability. Instead of treatment being independent of potential outcomes outright, it is independent within cells defined by covariates:

(Y(0), Y(1)) independent of D, conditional on X

This is the conditional independence assumption, also called unconfoundedness or selection on observables. In words: among people with the same values of X, who got treated is as good as random. It buys you the ability to compute the effect within each cell, where selection bias is zero by assumption, and then average up.

It is not enough by itself. You also need overlap, or common support:

0 < P(D = 1 | X) < 1 for all X in the population of interest

If some cell contains only treated units, there is nothing to compare them with, and no estimator can conjure a counterfactual for it. Rosenbaum and Rubin called the pair strong ignorability in 1983. The point: conditional independence is an assumption about unobservables and is untestable; overlap is a property of your data and is checkable, and checking it is the single most useful thing you can do before matching anything.

Subclassification, computed

Take a study of 200 treated and 300 control units, sorted into three strata by a single covariate.

StratumTreated NTreated mean YControl NControl mean YWithin-cell difference
Low2012.018010.0+2.0
Middle5016.010015.0+1.0
High13022.02022.5-0.5

The naive comparison first. Treated mean: (20(12) + 50(16) + 130(22))/200 = 3900/200 = 19.5. Control mean: (180(10) + 100(15) + 20(22.5))/300 = 3750/300 = 12.5. The difference is +7.0, and almost all of it comes from the fact that treated units are concentrated in the high stratum where everybody scores around 22.

Now average the within-cell differences, weighting by the treated counts, which gives the ATT:

ATT = (20(2.0) + 50(1.0) + 130(-0.5))/200 = (40 + 50 - 65)/200 = 25/200 = 0.125

Weighting instead by the total counts in each stratum, 200, 150 and 150, gives the ATE:

ATE = (200(2.0) + 150(1.0) + 150(-0.5))/500 = (400 + 150 - 75)/500 = 0.95

Three numbers from one table: 7.0, 0.125 and 0.95. The naive figure is fifty times the ATT. And the ATT and ATE differ by a factor of seven, because the treatment helps the low stratum and slightly hurts the high one, and the treated are concentrated in the high one. Which number you want depends on the policy question: the ATT answers "was this programme worth running for the people who took it", the ATE answers "should it be extended to everybody".

Why the propensity score exists

Subclassification is exact and transparent with one covariate. With ten binary covariates there are 1,024 cells, and with realistic sample sizes most will be empty of treated units, control units, or both. This is the curse of dimensionality, and it is why LaLonde's era used regression, which imposes a linear functional form to fill the gaps.

Rosenbaum and Rubin's 1983 theorem provides the alternative. Define the propensity score p(X) = P(D = 1 | X). Their result: if treatment is strongly ignorable given X, it is strongly ignorable given the single number p(X). Conditioning on one scalar suffices; the 1,024 cells collapse to a single dimension along which you can stratify, match or weight.

The estimators built on it come in three families. Matching pairs each treated unit with one or several control units of similar score. Stratification cuts the score into blocks and applies the arithmetic above. Inverse probability weighting reweights every observation by the inverse of its probability of receiving the treatment it got:

ATE = E[ DY/p(X) ] - E[ (1 - D)Y/(1 - p(X)) ]

Read the weights and you can see the failure mode. A treated unit with p(X) = 0.02 receives a weight of 50 and can dominate the estimate on its own. This is the overlap condition making itself felt numerically: as the score approaches 0 or 1, the variance explodes. Trimming units with extreme scores is standard practice and is also, quietly, a change of estimand, since you are now estimating an effect for the trimmed population.

One procedural warning that catches nearly everyone. The propensity score model is not a prediction problem, and its pseudo R-squared is not a measure of success. The diagnostic is covariate balance: after matching or weighting, are the standardised differences in each covariate between treated and comparison groups small? A score model that predicts treatment brilliantly and leaves an unbalanced covariate has failed; a mediocre model that balances everything has succeeded.

The dispute

The case that matching solved it. Rajeev Dehejia and Sadek Wahba, in 1999, went back to LaLonde's data, restricted attention to the subsample for which two years of pre-treatment earnings were observed, estimated propensity scores, and matched. Their non-experimental estimates landed close to the experimental benchmark, across several matching estimators, where LaLonde's regressions had not. The argument they made is specific and strong: the failures of 1986 were failures of functional form and of comparing units with no overlap, and propensity score methods fix exactly those two things by refusing to extrapolate. Later within-study comparisons in education and public health, reviewed by Thomas Cook and coauthors in 2008, found the same pattern: non-experimental estimates match experimental benchmarks reasonably well when the comparison group is drawn locally, when pre-treatment outcomes are measured the same way in both groups, and when the selection process is understood well enough to model.

The case that it did not. Jeffrey Smith and Petra Todd, in 2005, re-examined the same data and reported that the Dehejia and Wahba results were fragile: change the subsample, change the set of covariates in the score model, or change the matching estimator, and the estimates moved substantially, sometimes away from the benchmark. Their reading is that the good results came from choices validated against an experimental answer that was already known, which is not a procedure available to anyone doing real work. Gary King and Richard Nielsen add a separate objection: matching on the propensity score approximates a completely randomised experiment rather than a fully blocked one, and pruning observations along a scalar can increase imbalance on the covariates themselves, so the method can move the estimate in the wrong direction while its own diagnostics improve.

What would settle it. Not more argument about the NSW data, which has been re-analysed past exhaustion. The relevant evidence is the accumulating set of within-study comparisons in which an experimental benchmark exists and a non-experimental estimate is constructed blind to it, with the design choices registered in advance. The pattern in that literature is not that a method works or fails, but that data features do: local comparison groups, pre-treatment outcomes measured identically, and a documented assignment process. On that reading both sides are right about different things. Matching fixed the functional-form and overlap failures that LaLonde exposed, and it did nothing whatsoever about selection on unobservables, which is the failure that actually kills observational estimates.

Common misconceptions

  • "Matching is a causal identification strategy." It is a weighting scheme. It changes how the comparison is computed, not what has to be assumed. Conditional independence carries the identification, and matching, regression and weighting are three ways of imposing the same assumption.
  • "Good balance means the estimate is unbiased." Balance is achievable only on covariates you observe. The confounder that ruins the study is by definition not in the table you are balancing.
  • "A high pseudo R-squared means a good propensity model." A score of 0.99 for treated units means no comparable controls exist, which is an overlap failure, not a success. Judge the model by balance and by the overlap of the score distributions.
  • "Regression and matching answer the same question." Regression extrapolates through regions with no comparable units and weights strata by the variance of treatment within them; matching refuses to extrapolate and weights by treated counts. On data with poor overlap the two can differ substantially, and the difference is diagnostic rather than cosmetic.

Summing up

Conditional independence plus overlap identifies the ATT, and subclassification computes it: on the worked table, a naive difference of 7.0 becomes an ATT of 0.125 and an ATE of 0.95. The propensity score reduces conditioning on many covariates to conditioning on one, which is what makes matching, stratification and inverse probability weighting practical, at the cost of exploding variance where scores approach 0 or 1. The diagnostic is covariate balance, never the fit of the score model. And LaLonde's challenge, forty years on, has a two-part answer that neither side of the dispute owns outright.

Remember: the assumption that breaks every estimator in this lesson is a single one. If assignment to treatment depends on anything that also affects the outcome and is absent from X, all of these estimates are biased, and no balance table, overlap plot or specification check in the toolkit will reveal it. The next lesson shows that adding more variables to X is not a way out, and can make things worse.

Sources

  1. Cunningham, S. (2021). Matching and subclassification. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  2. Huntington-Klein, N. (2022). Matching. The Effect: An Introduction to Research Design and Causality. Chapman and Hall. theeffectbook.net
  3. Wikipedia contributors. (n.d.). Propensity score matching. Wikipedia. en.wikipedia.org
  4. LaLonde, R. J. (1986). Evaluating the econometric evaluations of training programs with experimental data. American Economic Review, 76(4), 604-620.
  5. Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1), 41-55.
  6. Smith, J., & Todd, P. (2005). Does matching overcome LaLonde's critique of nonexperimental estimators? Journal of Econometrics, 125(1-2), 305-353.
Key terms
Conditional independence assumption
Treatment is independent of the potential outcomes within cells defined by covariates X; also called unconfoundedness or selection on observables.
Overlap
Every covariate cell contains both treated and untreated units, so 0 < p(X) < 1; checkable in the data, unlike conditional independence.
Strong ignorability
Conditional independence together with overlap, the pair Rosenbaum and Rubin showed sufficient for identification.
Propensity score
p(X) = P(D=1|X); conditioning on this single number is as good as conditioning on all of X when treatment is strongly ignorable.
Subclassification
Splitting the sample into covariate strata, computing the effect within each, and averaging with weights chosen for the estimand.
Inverse probability weighting
Reweighting each unit by the reciprocal of its probability of receiving the treatment it actually got.
Covariate balance
The similarity of covariate distributions between treated and comparison groups after matching or weighting; the correct diagnostic for a score model.
Within-study comparison
A design that estimates the same effect experimentally and non-experimentally on the same population, to test whether the observational method recovers the benchmark.

Manufacturing a Correlation: Bad Controls and Collider Bias

  • Distinguish a confounder, a mediator and a collider, and state the rule for each.
  • Compute the correlation that conditioning on a collider manufactures between two independent variables.
  • Show, on an exact numerical example, how controlling for a mediator can reverse the sign of an estimated effect.

In 1946 Joseph Berkson, a statistician at the Mayo Clinic, published a two-page note about hospital records. Suppose two diseases occur independently in the general population. Suppose each one, on its own, makes a person more likely to be admitted to hospital. Then among hospital patients the two diseases will appear negatively associated, because a patient in the ward with neither cause of admission is rare, so having one disease is evidence against having the other. Nothing about the diseases changed. The sample did.

Berkson's note is the oldest clean statement of a mistake that this course's toolkit makes very easy to commit. Lesson 6 ended with a slogan that sounds sensible: control for the confounders. The reasoning fails as soon as it is generalised to "control for everything you have", and this lesson traces exactly where.

Three jobs a third variable can hold

Draw the variable of interest D, the outcome Y, and a third variable. There are three distinct structures, and the correct action is different in each.

StructurePictureNameWhat to do
Common causeD is caused by W, and Y is caused by WConfounder, or forkControl it. Leaving it out is Lesson 6's omitted variable bias
On the pathD causes M, and M causes YMediator, or chainDo not control it if you want the total effect
Common effectD causes C, and Y causes CColliderNever control it, or any descendant of it

Judea Pearl's directed acyclic graph notation exists to make this decision mechanical. Draw every variable you believe matters, draw an arrow for every causal claim you are willing to defend, and then apply the back-door criterion: a set of controls is sufficient if it blocks every path from D to Y that begins with an arrow into D, without opening any path through a collider. The graph does not discover the truth. It forces you to write your assumptions down where a reader can disagree with them, which is most of the value.

Conditioning on a collider, arithmetic

Here is a correlation of exactly -0.5 manufactured out of independence. Take 100 aspiring actors. Fifty are talented, decided by a coin flip. Independently, fifty are good-looking, decided by another. Independence means the four combinations each contain 25 people.

Casting directors hire anyone who has at least one of the two qualities. That is a collider: talent causes being cast, looks cause being cast. Now look at the 75 actors who were cast.

Good-lookingNot good-lookingRow total
Talented252550
Not talented250 cast (25 rejected)25
Column total502575

Among the cast, the probability of being good-looking given talent is 25/50 = 0.50. Given no talent, it is 25/25 = 1.00. Talent is now evidence against looks. Put a number on it. Coding both variables 0 or 1 and working within the 75 who were cast:

E[T] = 50/75 = 2/3, E[B] = 50/75 = 2/3, E[TB] = 25/75 = 1/3

Cov(T, B) = 1/3 - (2/3)(2/3) = 3/9 - 4/9 = -1/9

Var(T) = Var(B) = (2/3)(1/3) = 2/9, so Corr(T, B) = (-1/9)/(2/9) = -0.5

Two independent variables, a correlation of -0.5, and no data were falsified. The sample was defined by their common effect. What matters here: the bias is created by conditioning, which means it appears both when you add a collider to a regression and when your sample was assembled by one.

The second case is the dangerous one, because it is invisible in the specification. Studies of hospitalised patients condition on admission. Studies of wages condition on employment, and both education and unobserved ability affect whether someone works. Surveys condition on response. Studies of professional athletes condition on selection into the league, which is why height and skill are famously negatively correlated among NBA players and positively correlated in the population. The birth-weight paradox, in which maternal smoking appears protective among low-birth-weight infants, was resolved by Hernandez-Diaz, Schisterman and Hernan as exactly this: birth weight is a collider between smoking and the unmeasured causes of infant mortality.

Controlling for a mediator, with a sign flip

The collider case is dramatic. The mediator case is more common and more insidious, because the variable looks like an obviously relevant control. Take an exactly specified world:

M = 4D + U and Y = 2D + 3M + 5U

where D is randomly assigned and U is an unobserved variable affecting both the mediator and the outcome. Substituting the first equation into the second gives the reduced form:

Y = 2D + 3(4D + U) + 5U = 14D + 8U

So the total effect of D on Y is 14, and because D is randomised and independent of U, a regression of Y on D alone recovers 14. The direct effect, the part not running through M, is 2 by construction.

Now do the thing that feels careful: add M as a control, to isolate the direct effect. Rewrite U = M - 4D and substitute into the reduced form:

Y = 14D + 8(M - 4D) = -18D + 8M

Since Y is now an exact linear function of D and M, the regression returns those coefficients exactly. The estimated effect of D is -18. The total effect is +14. The direct effect is +2. Controlling for the mediator has produced a number that is not any of the quantities of interest and has the wrong sign.

Two separate things went wrong at once, and it is worth separating them. First, conditioning on M removes the 12 units of effect that travel through the mediator, which is intended if you genuinely want the direct effect. Second, M is a collider on the path D to M and U to M, so conditioning on it induces a negative association between D and U, and since U raises Y, that manufactures an extra negative bias. The first effect takes 14 down to 2; the second takes 2 down to -18. This is why estimating direct effects honestly requires assumptions about U that the regression cannot supply.

Debugging a real specification

Consider the classic labour economics regression: log wages on years of schooling. Someone proposes adding occupation dummies, on the grounds that occupation is obviously relevant. Trace it. Schooling causes occupation; occupation causes wages. Occupation is a mediator, and the largest channel by which education pays is that it changes what job you can hold. Comparing a graduate and a non-graduate within the same occupation deliberately switches off the main effect being estimated. Angrist and Pischke call these bad controls, and the diagnostic they give is simple: a control that could itself have been an outcome of the treatment is suspect.

The natural repair, "only control for variables measured before treatment", is a good rule and not a theorem. The counterexample is M-bias. Suppose an unobserved U1 causes both a pre-treatment variable Z and the treatment D, and an unobserved U2 causes both Z and the outcome Y. With Z left alone, there is no open path from D to Y. Control for Z, which is a collider between U1 and U2, and you open one. The variable was measured before treatment, is correlated with treatment and with the outcome, and adding it creates bias out of nothing.

Common misconceptions

  • "Adding a control can only reduce bias." The collider computation above adds a control to a correctly specified model and manufactures a correlation of -0.5. There is no general result that says more controls are safer, and the belief that there is causes real damage.
  • "A variable correlated with both treatment and outcome must be a confounder." Colliders are correlated with both by construction, since both cause them. Correlation patterns do not identify the structure; only causal assumptions do.
  • "Controlling for the mediator gives the direct effect." It gives the direct effect only when nothing unobserved causes both the mediator and the outcome. Whenever such a variable exists, the mediator is also a collider and the estimate is biased, as the -18 shows.
  • "Sample selection is a separate problem from control variables." They are the same operation. Selecting a sample on a variable is conditioning on it, so a sample built on a collider carries collider bias whatever the regression does afterwards.

Looking back

Three structures, three rules: control forks, leave mediators alone when the total effect is wanted, and never condition on a collider or its descendants. Conditioning on the common effect of two independent variables produced a correlation of exactly -0.5 in a table of 75 actors. Controlling for a mediator that shares an unobserved cause with the outcome turned a total effect of +14 and a direct effect of +2 into an estimate of -18. Sample selection is conditioning by another route, which is why hospital samples, employed samples and survey respondents carry the same risk as a badly chosen regressor.

Bottom line: the decision to include a variable is a causal claim, and the only way to make it defensible is to write the graph down. Module 4 stops trying to fix the comparison by choosing controls and starts changing the source of variation instead.

Sources

  1. Wikipedia contributors. (n.d.). Collider (statistics). Wikipedia. en.wikipedia.org
  2. Cunningham, S. (2021). Directed acyclical graphs. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  3. Huntington-Klein, N. (2022). Causal diagrams. The Effect: An Introduction to Research Design and Causality. Chapman and Hall. theeffectbook.net
  4. Elwert, F., & Winship, C. (2014). Endogenous selection bias: The problem of conditioning on a collider variable. Annual Review of Sociology, 40, 31-53.
  5. Hernandez-Diaz, S., Schisterman, E. F., & Hernan, M. A. (2006). The birth weight paradox uncovered? American Journal of Epidemiology, 164(11), 1115-1120.
  6. Angrist, J. D., & Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist's Companion, Section 3.2.3 on bad controls. Princeton University Press.
Key terms
Confounder
A common cause of treatment and outcome; the variable whose omission produces the bias of Lesson 6, and which should be controlled.
Mediator
A variable on the causal path from treatment to outcome; controlling for it removes the part of the effect that travels through it.
Collider
A common effect of two variables; conditioning on it induces association between causes that were independent.
Berkson's paradox
The negative association between independent conditions observed in a sample selected on either of them, named for a 1946 note on hospital records.
Back-door criterion
Pearl's rule: a control set suffices if it blocks every path into the treatment without opening a path through a collider.
Bad control
A control variable that could itself have been affected by the treatment; including it changes the estimand and often biases it.
M-bias
Bias created by controlling for a pre-treatment variable that is a collider between two unobserved causes, showing that pre-treatment is not a sufficient safeguard.
Direct effect
The part of a treatment's effect not travelling through a specified mediator; identified only under assumptions about unobserved causes of the mediator.

Module 4: Instruments and Panels

Finding variation that was not chosen by the people you are studying: the Wald estimator on a lottery, what LATE means and when instruments are too weak to use, and what fixed effects can and cannot absorb.

Capsule 001, September 14: Instrumental Variables and the Draft Lottery

  • State the three conditions an instrument must satisfy and identify which of them is untestable.
  • Compute a Wald estimate from a two-by-two table of means and interpret it as a ratio of reduced form to first stage.
  • Quantify how a violation of the exclusion restriction is amplified by a weak first stage.

Before you read: write down one sentence explaining why you cannot learn the effect of military service on earnings by comparing veterans with non-veterans. You wrote the general form of the answer in Lesson 1; this lesson is about what to do next.

On 1 December 1969, in a room at Selective Service headquarters in Washington, 366 blue plastic capsules were drawn from a glass jar. Each held a date. The first capsule out was September 14, so every man born on that date received random sequence number 001. The last was June 8, number 366. For the 1970 draft, the call went down the list to number 195; men born in 1950 with a lower number were draft eligible, men above it were not. Two subsequent lotteries set ceilings of 125 for the 1951 cohort and 95 for the 1952 cohort.

Two decades later Joshua Angrist matched those numbers to Social Security earnings records and asked what military service had cost the men who served. The answer, for white men, was a loss of roughly 15 percent of annual earnings, persisting into the early 1980s. The important part is not the number. It is that a jar of capsules made the number knowable at all.

Why the direct comparison fails, and what the lottery replaces

Comparing veterans with non-veterans is Lesson 1's hospital problem in uniform. Enlistment is chosen; those who volunteer differ in schooling, family background, local labour market conditions and preferences, and every one of those differences sits in E[Y(0) | veteran] - E[Y(0) | non-veteran]. Controls do not fix it, because the relevant differences are exactly the ones nobody measured.

The lottery supplies something else: a variable that shifts the probability of service without being chosen by anyone. Call it Z, coded 1 for draft eligible and 0 otherwise. It is not the treatment. Plenty of eligible men never served, through medical deferment, enrolment, or conscientious objection, and plenty of ineligible men volunteered. What Z does is move D, and that is enough.

Three conditions, one of them untestable

A variable Z is a valid instrument for D in an equation for Y if it satisfies three requirements.

  • Relevance. Cov(Z, D) is not 0. Draft eligibility must actually change the probability of service. This is testable: run the first-stage regression and look.
  • Independence. Z is as good as randomly assigned with respect to the potential outcomes. The lottery draw satisfies this by construction, which is the entire reason for using it.
  • Exclusion restriction. Z affects Y only through D. A low lottery number must change earnings only by changing whether a man served, and by no other route.

The third is not testable and never will be. It is a claim about a channel that, if it existed, would be invisible in the data. And it is genuinely arguable here. A man with a low number could enrol in university to obtain a deferment, and that extra education would change his earnings without any military service. He could take a job in a draft-exempt occupation, or emigrate. Angrist discusses these channels explicitly, and his defence is empirical rather than logical: the schooling response to draft eligibility was small in the cohorts studied. Why this matters: an instrument is defended, never proved, and the defence is always a substantive argument about the world.

The Wald estimator

Write the model Y = alpha + beta D + u, where u is correlated with D so OLS fails. Take covariances with Z:

Cov(Z, Y) = beta Cov(Z, D) + Cov(Z, u)

Independence and exclusion together make Cov(Z, u) = 0. Rearranging,

betaIV = Cov(Z, Y) / Cov(Z, D)

and when Z is binary this reduces to a ratio of two differences in means, the Wald estimator:

betaWald = (E[Y|Z=1] - E[Y|Z=0]) / (E[D|Z=1] - E[D|Z=0])

The numerator is the reduced form: the effect of eligibility on earnings. The denominator is the first stage: the effect of eligibility on service. Read the formula as a scaling argument. Eligibility moved only some men into service, so the effect of eligibility on earnings is a diluted version of the effect of service on earnings, diluted exactly by the fraction whose service status the lottery changed. Dividing undoes the dilution.

Working the arithmetic

Take a clean illustrative table first, with round numbers rather than Angrist's, so every step is checkable. Two thousand men, half eligible.

GroupNServedVeteran rateMean 1981 earnings
Draft eligible (Z = 1)10003500.35015,880
Not eligible (Z = 0)10001900.19016,300
Difference0.160-420

betaWald = -420 / 0.160 = -2,625 dollars per year of veteran status. Notice that the reduced-form difference, 420 dollars, is small, and the estimated effect on those actually affected is six times larger. That factor of six is 1/0.160, and it will matter enormously in a moment.

Angrist's own magnitudes, rounded from his tables, are of the same shape: for white men born in 1950 the first stage is about 0.16, and the effect of eligibility on annual earnings in the early 1980s is on the order of 400 to 500 dollars in 1978 dollars, giving a Wald estimate near 2,700 dollars, roughly 15 percent of that cohort's mean annual earnings.

Two-stage least squares, and why it is the same thing

With covariates or a non-binary instrument, the Wald ratio generalises to two-stage least squares. Stage one regresses D on Z and any covariates, producing fitted values Dhat. Stage two regresses Y on Dhat. The logic is that Dhat contains only the part of D that the instrument explains, and that part is uncorrelated with u by assumption. With a single binary instrument and no covariates, 2SLS returns the Wald estimate exactly.

One practical warning. Do not run the two stages by hand and report the second-stage standard errors. They treat Dhat as data rather than as an estimate and are too small. Use a command that computes IV standard errors properly. A useful approximation is that the IV standard error equals roughly the OLS standard error divided by the correlation between Z and D, so a weak first stage costs precision in direct proportion.

How a small first stage amplifies everything

Here is the calculation that should govern how you read any IV paper. Suppose the exclusion restriction is slightly violated: eligibility has a direct effect on earnings of gamma dollars, not through service. The reduced form is then beta times firststage + gamma, and the Wald estimator returns

beta + gamma / firststage

The bias is not gamma. It is gamma divided by the first stage. With a first stage of 0.16, a direct effect worth only 50 dollars a year, far too small to detect and easy to dismiss, produces a bias of 50/0.16 = 312 dollars in an estimate of -2,625, which is 12 percent of the answer. Halve the first stage to 0.08 and the same undetectable 50 dollars produces a bias of 625 dollars.

In short: the weaker the instrument, the more a small violation of exclusion is magnified, which is why "the first stage is significant" is a completely inadequate defence and why the next lesson is about the strength of instruments.

Common misconceptions

  • "A random instrument makes the estimate valid." Randomness delivers independence, one of three conditions. Exclusion is separate, and a perfectly random instrument with a second channel to the outcome gives a biased estimate.
  • "The instrument must be correlated with the outcome." It must be correlated with the treatment. Its correlation with the outcome is the reduced form, which is the numerator, and if the treatment effect is zero the correct reduced form is zero too.
  • "Draft eligibility is the treatment." It is not. The reduced-form effect of eligibility, about 420 dollars in the table, is the effect of being eligible on the whole cohort; the 2,625 is the effect of service on the men whose service status eligibility changed. Confusing them understates the treatment effect by a factor of the first stage.
  • "IV fixes measurement error and omitted variables at once." IV addresses any correlation between the regressor and the error, which does include classical measurement error. But it does so only under the exclusion restriction, and it replaces one untestable assumption with another rather than removing the need for assumptions.

Recap

A lottery drawn from a jar in December 1969 created variation in military service that nobody chose, and that variation is what makes the causal question answerable. An instrument needs relevance, which you can test; independence, which the lottery supplies; and the exclusion restriction, which you must argue and can never verify. The Wald estimator divides the reduced form by the first stage, giving -2,625 dollars in the worked table and something near -2,700 in Angrist's data, about 15 percent of annual earnings. Two-stage least squares generalises it, and its standard errors must come from an IV routine rather than from a hand-run second stage. And the factor 1/firststage that converts the reduced form into the estimate converts any violation of exclusion into bias at the same rate.

The next lesson asks two questions this one left open. Whose effect is 2,625 dollars an estimate of, given that most men in the sample had their service status unchanged by the lottery? And how strong does a first stage have to be before that division stops being dangerous?

Sources

  1. Angrist, J. D., & Krueger, A. B. (2001). Instrumental variables and the search for identification: From supply and demand to natural experiments. Journal of Economic Perspectives, 15(4), 69-85. aeaweb.org
  2. Cunningham, S. (2021). Instrumental variables. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  3. Wikipedia contributors. (n.d.). Draft lottery (1969). Wikipedia. en.wikipedia.org
  4. Angrist, J. D. (1990). Lifetime earnings and the Vietnam era draft lottery: Evidence from Social Security administrative records. American Economic Review, 80(3), 313-336.
  5. Angrist, J. D., & Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist's Companion, Chapter 4. Princeton University Press.
Key terms
Instrument
A variable that shifts the treatment, is as good as randomly assigned, and affects the outcome only through the treatment.
Relevance
Non-zero covariance between instrument and treatment; the one instrument condition that is directly testable.
Exclusion restriction
The requirement that the instrument affect the outcome through no channel other than the treatment; untestable and always argued substantively.
First stage
The regression of the treatment on the instrument, and the denominator of the Wald ratio.
Reduced form
The regression of the outcome on the instrument, and the numerator of the Wald ratio.
Wald estimator
The ratio of the reduced-form difference in means to the first-stage difference in means, for a binary instrument.
Two-stage least squares
The general IV estimator: fit the treatment on instruments and covariates, then regress the outcome on the fitted values.

Whose Effect Is It? LATE, Monotonicity, and Instruments Too Weak to Use

  • Classify a population into always-takers, never-takers, compliers and defiers, and compute the LATE from their effects and shares.
  • State the monotonicity assumption and explain what the Wald estimator becomes when it fails.
  • Diagnose a weak instrument from a first-stage F statistic and describe the inference problems it causes.

The Wald estimate at the end of the last lesson was -2,625 dollars. Here is the question that estimate does not answer: whose 2,625 dollars? Most men in that sample had their service status decided by something other than the lottery. Some volunteered whatever their number. Some obtained a deferment whatever their number. For all of them, the lottery moved nothing, and an estimator built entirely on the variation the lottery created cannot have learned anything about them.

Guido Imbens and Joshua Angrist answered this in 1994, and the answer is more restrictive and more useful than "the effect of military service".

Four kinds of person

With a binary instrument and a binary treatment, every individual has two potential treatment statuses: D(1), what they would do if Z = 1, and D(0), what they would do if Z = 0. Four combinations exist.

TypeD(0)D(1)In the draft lotteryDoes the instrument move them?
Always-taker11Volunteers, whatever their numberNo
Never-taker00Deferred or unfit, whatever their numberNo
Complier01Served because their number came upYes
Defier10Would serve only if not calledYes, backwards

Monotonicity is the assumption that defiers do not exist: shifting the instrument on can push people into treatment but never out of it. In the lottery it is easy to defend, since being called up does not make anyone less likely to serve. In other designs it is much harder, and it is often assumed silently.

Under independence, exclusion, relevance and monotonicity, Imbens and Angrist proved that the Wald estimator identifies the local average treatment effect: the average causal effect among compliers, and nobody else.

A population where LATE and ATE have opposite signs

Take 1,000 people, with the instrument assigned by a fair coin.

TypeNumberShareAverage effect of treatment
Always-takers2000.20+10
Compliers3000.30-5
Never-takers5000.50+2

The average treatment effect over everybody is

ATE = 0.20(10) + 0.30(-5) + 0.50(2) = 2 - 1.5 + 1 = +1.5

Now compute what IV returns. Among those with Z = 1, the treated are the always-takers plus the compliers, so P(D=1|Z=1) = 0.50. Among those with Z = 0, only always-takers are treated, so P(D=1|Z=0) = 0.20. The first stage is 0.30, which is exactly the complier share, and that is not a coincidence: the first stage always equals the share of compliers.

For the reduced form, note that the instrument changes nothing for always-takers or never-takers, by definition, so their outcomes contribute identically to both arms and cancel. Only compliers move, and only 30 percent of the sample are compliers:

E[Y|Z=1] - E[Y|Z=0] = 0.30 times (-5) = -1.5

and therefore betaWald = -1.5 / 0.30 = -5, which is the LATE exactly. The true ATE is +1.5. A flawless instrument, a valid design, correct arithmetic, and the sign is the opposite of the population average effect. The upshot: IV does not estimate a badly measured ATE. It estimates a different quantity precisely.

The shares of the other two groups are also recoverable: always-takers are P(D=1|Z=0) = 0.20 and never-takers are P(D=0|Z=1) = 0.50, and the three add to one.

Who the compliers are, and why it matters

You cannot point at an individual and say "complier", because that requires knowing both D(1) and D(0) for one person, which is Lesson 1's fundamental problem again. But you can compute the distribution of any covariate among compliers, since for a binary covariate X the complier share within X = 1 is just the first stage computed in that subgroup, and comparing it with the overall first stage tells you whether compliers are over-represented there.

This matters because it converts an abstract caveat into a concrete description. Draft lottery compliers are men who would not have volunteered and could not obtain a deferment, which skews toward less educated men from poorer backgrounds. Compulsory schooling law compliers are people who would have left school at the earliest legal moment. Distance-to-college compliers are people for whom a nearby campus was decisive. Three instruments for "the return to schooling", three different populations, and no reason at all for the three LATEs to agree. When two IV papers on the same question report different numbers, the first hypothesis should be that they estimated different parameters, not that one of them is wrong.

When monotonicity fails

If defiers exist, the Wald estimator becomes a weighted average of complier and defier effects in which the defiers receive negative weight, and the result need not lie between the largest and smallest individual effects at all. It can fall outside the range of every effect in the population.

The setting where this bites hardest is judge or examiner assignment, a popular instrument in which cases are randomly allocated to decision makers of differing strictness. A judge who is strict on average may be unusually lenient toward one category of defendant, and for that category the instrument runs backwards. Monotonicity must then hold judge by judge and subgroup by subgroup, which is a much stronger claim than "this judge convicts more often overall".

Weak instruments

Everything above assumed the first stage is real. When it is small relative to its noise, two separate problems appear, and only one of them is fixed by a larger sample.

Finite-sample bias. Two-stage least squares is biased toward the OLS estimate, and the size of that bias relative to OLS bias is approximately 1/F, where F is the first-stage F statistic. At F = 10 the 2SLS estimate carries about a tenth of the OLS bias, which is the origin of Staiger and Stock's familiar rule of thumb. At F = 1 it carries essentially all of it, and the IV estimate is a laundered version of the OLS estimate it was supposed to improve on.

Broken inference. Conventional IV confidence intervals assume the first stage is well determined. With a weak instrument the sampling distribution of the t-ratio is not approximately normal, so a nominal 95 percent interval can have far worse coverage. David Lee, Justin McCrary, Marcelo Moreira and Jack Porter showed that screening on F > 10 and then applying the usual 1.96 critical value does not deliver a 5 percent test; to justify 1.96 after such a screen you need a first stage F above roughly 105, and below that their tF procedure supplies a larger critical value. Anderson-Rubin confidence sets are valid whatever the instrument strength, and can be empty or infinite, which is an honest report rather than a failure.

The demonstration everybody remembers is John Bound, David Jaeger and Regina Baker's in 1995. Angrist and Krueger had used quarter of birth, interacted with state and year, as 180 instruments for schooling. Bound and coauthors replaced quarter of birth with numbers drawn from a random number generator, which have no first stage at all, and re-ran the procedure. The estimates and standard errors that came out looked much like the published ones. With many weak instruments the first stage fits the endogenous variable's own noise, and 2SLS collapses onto OLS while continuing to report respectable-looking output.

SymptomDiagnosisResponse
First-stage F near 22SLS carries about half the OLS biasReport the reduced form; consider abandoning the design
F of 12, t-ratio of 2.1Passes the old rule, fails the tF thresholdUse tF or Anderson-Rubin critical values
IV estimate far larger than OLS with a small first stagePossible exclusion violation amplified by 1/firststageArgue the channel explicitly; bound it
Many instruments, high pooled FOverfitting of the first stageLIML or jackknife IV; report the just-identified estimate

Common misconceptions

  • "LATE is the ATE for a slightly restricted population." The worked table has LATE at -5 and ATE at +1.5. There is no general reason for them to be close, or even to share a sign.
  • "A significant first stage means the instrument is strong enough." Significance is a statement about zero. The relevant question is how large F is, because both the finite-sample bias and the amplification of exclusion violations scale with its reciprocal.
  • "Adding more instruments improves precision." It improves the pooled first stage mechanically and worsens the finite-sample bias, which is exactly the Bound, Jaeger and Baker result. Just-identified IV with one strong instrument is usually preferable.
  • "Monotonicity is automatic." It is easy to defend for a draft lottery and hard to defend for judges, examiners, doctors or caseworkers, where the instrument's strictness can reverse for identifiable subgroups.

What to remember

Under monotonicity, IV identifies the average effect among compliers, the people whose treatment status the instrument actually changed. The first stage equals the complier share, always-takers are P(D=1|Z=0), never-takers are P(D=0|Z=1), and on the worked population the LATE is -5 while the ATE is +1.5. Different instruments for the same treatment define different complier groups, so disagreeing IV estimates may both be right. Defiers destroy the interpretation by entering with negative weight. And a weak first stage does two things at once: it drags 2SLS toward OLS at a rate of about 1/F, and it breaks the normal approximation that conventional confidence intervals depend on.

The core of it: report the first stage, report the reduced form, and describe the compliers. A paper that gives you only the second-stage coefficient has withheld the three things you need to judge it.

Sources

  1. Wikipedia contributors. (n.d.). Local average treatment effect. Wikipedia. en.wikipedia.org
  2. Bound, J., Jaeger, D. A., & Baker, R. (1995). Problems with instrumental variables estimation when the correlation between the instruments and the endogenous explanatory variable is weak. Journal of the American Statistical Association, 90(430), 443-450. NBER Technical Working Paper 137. nber.org
  3. Lee, D. S., McCrary, J., Moreira, M. J., & Porter, J. (2022). Valid t-ratio inference for IV. American Economic Review, 112(10), 3260-3290. aeaweb.org
  4. Imbens, G. W., & Angrist, J. D. (1994). Identification and estimation of local average treatment effects. Econometrica, 62(2), 467-475.
  5. Staiger, D., & Stock, J. H. (1997). Instrumental variables regression with weak instruments. Econometrica, 65(3), 557-586.
  6. Imbens, G. W. (2010). Better LATE than nothing: Some comments on Deaton (2009) and Heckman and Urzua (2009). Journal of Economic Literature, 48(2), 399-423. aeaweb.org
Key terms
Complier
A unit that takes the treatment when the instrument is on and not when it is off; the only type IV learns anything about.
Always-taker
A unit treated regardless of the instrument; its share equals P(D=1|Z=0).
Never-taker
A unit untreated regardless of the instrument; its share equals P(D=0|Z=1).
Defier
A unit that takes the treatment only when the instrument is off; ruled out by monotonicity and, if present, entering the estimate with negative weight.
Monotonicity
The assumption that the instrument moves everyone in the same direction, so no defiers exist.
LATE
Local average treatment effect: the average effect among compliers, which is what IV identifies under the four conditions.
First-stage F
The F statistic on the excluded instruments; 2SLS finite-sample bias relative to OLS bias is roughly its reciprocal.
Anderson-Rubin confidence set
An inference procedure valid regardless of instrument strength, which may return an empty or unbounded set.

Three Cities, Two Answers: Fixed Effects and What They Cannot Absorb

  • Compute a pooled OLS slope and a within-unit slope on the same panel and explain why they differ.
  • Show that the within estimator, least squares dummy variables and first differencing coincide when T = 2.
  • List what fixed effects absorb, what they cost, and which confounders survive them.

Three cities, each observed twice. City A employs 1 then 2 police officers per thousand residents, and records 10 then 9 crimes per thousand. City B goes from 4 to 5 officers and from 20 to 19 crimes. City C goes from 7 to 8 officers and from 30 to 29 crimes. In every city, more police went with less crime, by exactly one crime per extra officer.

Pool the six observations and run the regression. The slope is +3.16. More police, more crime, strongly and unambiguously, in a dataset in which every single city shows the opposite.

The arithmetic of the contradiction

Pooling, the means are xbar = 27/6 = 4.5 and ybar = 117/6 = 19.5. The deviations are dx = -3.5, -2.5, -0.5, 0.5, 2.5, 3.5 and dy = -9.5, -10.5, 0.5, -0.5, 10.5, 9.5, giving

Sxy = 33.25 + 26.25 - 0.25 - 0.25 + 26.25 + 33.25 = 118.5 and Sxx = 12.25 + 6.25 + 0.25 + 0.25 + 6.25 + 12.25 = 37.5

so the pooled slope is 118.5/37.5 = 3.16. The estimate is dominated by the comparison between cities, where the big city has both more police and more crime, and this between-city variation has nothing to do with what happens when a city hires an officer.

Now subtract each city's own mean from its own observations. City A has xbar = 1.5 and ybar = 9.5, so its deviations are dx = -0.5, +0.5 and dy = +0.5, -0.5. Cities B and C give exactly the same pair, because each moved by the same amounts. Summing over all six rows:

Sxy = 3 times [(-0.5)(0.5) + (0.5)(-0.5)] = -1.5 and Sxx = 3 times [0.25 + 0.25] = 1.5

The within estimator is -1.5/1.5 = -1.0, which is what every city in the data actually did. Remember: a panel contains two kinds of variation, and they can carry opposite signs; choosing an estimator is choosing which one to use.

The model behind the transformation

Write

yit = beta xit + ai + uit

where ai is everything about unit i that does not change over the sample: its size, its geography, its institutions, its policing culture, the whole set of things that make a large city both heavily policed and high-crime. Pooled OLS puts ai in the error and is biased by exactly the Lesson 6 formula whenever ai is correlated with xit.

The within transformation subtracts each unit's time average from every variable. Since ai is constant within the unit, it equals its own average and vanishes:

yit - ybari = beta (xit - xbari) + (uit - ubari)

The ai is gone without ever being measured, named or modelled, which is the whole appeal of fixed effects.

Three estimators do this and are algebraically identical when T = 2. The within estimator demeans as above. Least squares dummy variables puts one dummy per unit into the regression, and the Frisch-Waugh-Lovell theorem says that is the same computation. First differencing subtracts consecutive periods; here dx = 1 and dy = -1 for every city, so the slope is -3/3 = -1.0, matching. For T greater than 2 within and first differences diverge, and which is preferable depends on the serial correlation of uit: first differencing wins when the errors are close to a random walk, within wins when they are close to independent.

What fixed effects buy, and at what price

Absorbed without measurementNot absorbed
Any unit characteristic constant over the sample window: geography, founding institutions, a person's fixed abilityAny confounder that changes over time: a local recession, a change in reporting practice, a new mayor
With time effects added, any shock common to all units in a period: a national recession, a federal lawReverse causality: crime rising this year causing police hiring next year
The level differences that dominate cross-sectional comparisonsUnit-specific trends, unless separately included at further cost in variation
Anticipation: units responding before the treatment arrives

The costs are real and are often not stated.

Time-invariant regressors disappear. Sex, race, country of birth, distance to the coast: all collinear with the unit dummies, all dropped. If your research question is about one of them, fixed effects is not available.

Measurement error gets worse, sometimes catastrophically. Suppose the true regressor is persistent within a unit and the measurement error is not. Take a within-unit variance of the true variable of 1 with autocorrelation 0.9, and a classical measurement error variance of 0.25 in each period. In levels, if the total variance of the true regressor is 9 because most of it is between units, the reliability ratio is 9/9.25 = 0.97 and attenuation is negligible. In first differences, the true signal has variance 2(1)(1 - 0.9) = 0.2 while the error has variance 2(0.25) = 0.5, so the reliability ratio falls to 0.2/0.7 = 0.29. The same data, differenced, attenuates the coefficient toward zero by a factor of three. Zvi Griliches and Jerry Hausman made this point in 1986, and it is the standard explanation for panel estimates of the return to schooling coming in far below cross-sectional ones.

Only switchers contribute. A unit whose regressor never changes has zero within-variation and drops out of the estimate entirely, however many rows it occupies. Reporting the number of switching units is as important as reporting the sample size.

Strict exogeneity is required, not just contemporaneous exogeneity. The within transformation involves ubari, which contains errors from every period, so the error in the transformed equation is mechanically related to regressors in all periods. Fixed effects therefore needs E[uit | xi1, ..., xiT, ai] = 0. Feedback breaks it: if this year's crime affects next year's police budget, the police variable in later periods responds to earlier shocks and the estimator is inconsistent. Including a lagged dependent variable guarantees the violation, and the resulting downward bias, described by Stephen Nickell in 1981, is of order 1/T, so it is severe in a three-period panel and negligible in a fifty-period one. Dynamic panel estimators in the Arellano-Bond family instrument the differenced lag with deeper lags to escape it.

Fixed or random effects

The random effects estimator treats ai as a draw from a distribution uncorrelated with the regressors and uses generalised least squares. When that assumption is true, random effects is more efficient and keeps time-invariant regressors in the model. When it is false, random effects is biased in precisely the way fixed effects exists to avoid.

In applied causal work the assumption is almost never plausible, since unobserved unit heterogeneity being uncorrelated with the treatment is close to assuming the problem away. The Hausman test compares the two estimates and rejects when they differ by more than sampling noise, but a failure to reject is weak evidence, especially in small samples: it can mean the assumption holds, or that the test lacks power.

Common misconceptions

  • "Fixed effects control for everything unobserved." They control for everything unobserved and constant. A confounder that moves passes straight through, and the confounders in most interesting questions move.
  • "Adding unit dummies is different from demeaning." They are the same estimator, by Frisch-Waugh-Lovell. Least squares dummy variables is just a computationally wasteful way to demean.
  • "A larger panel makes the lagged dependent variable problem go away." More units does not help; only more periods does, because the Nickell bias is of order 1/T and does not depend on N.
  • "Fixed effects estimates are more reliable because they are usually smaller." They are often smaller partly because measurement error attenuates them harder. A shrinking coefficient is evidence about the transformation as much as about confounding.

Pulling it together

On three cities, pooled OLS gives +3.16 and the within estimator gives -1.0, and only the second uses the variation that answers the question. The within transformation removes any time-invariant unit characteristic without measuring it; unit dummies and, when T = 2, first differencing do the same thing. The price is that time-invariant regressors cannot be estimated, that classical measurement error is amplified, in the worked example from a reliability of 0.97 to 0.29, that only units whose regressor changes contribute anything, and that strict exogeneity is required, which feedback and lagged dependent variables violate.

Worth holding on to: the assumption that breaks fixed effects is stated in one line. If anything that changes over time affects both the treatment and the outcome, fixed effects removes none of it. The next module builds designs that attack exactly that residue, by using the timing of a policy rather than the identity of a unit.

Sources

  1. Huntington-Klein, N. (2022). Fixed effects. The Effect: An Introduction to Research Design and Causality. Chapman and Hall. theeffectbook.net
  2. Cunningham, S. (2021). Panel data. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  3. Wikipedia contributors. (n.d.). Fixed effects model. Wikipedia. en.wikipedia.org
  4. Nickell, S. (1981). Biases in dynamic models with fixed effects. Econometrica, 49(6), 1417-1426.
  5. Griliches, Z., & Hausman, J. A. (1986). Errors in variables in panel data. Journal of Econometrics, 31(1), 93-118.
  6. Wooldridge, J. M. (2010). Econometric Analysis of Cross Section and Panel Data (2nd ed.), Chapters 10-11. MIT Press.
Key terms
Unobserved effect model
y_it = beta x_it + a_i + u_it, where a_i captures every time-invariant characteristic of unit i, observed or not.
Within estimator
OLS on unit-demeaned data; it removes a_i without measuring it and uses only variation over time inside each unit.
Least squares dummy variables
Including one dummy per unit; identical to the within estimator by the Frisch-Waugh-Lovell theorem.
First differencing
Regressing changes on changes; identical to within when T = 2 and preferable when the errors resemble a random walk.
Strict exogeneity
The error in any period is unrelated to the regressor in every period; required by fixed effects and violated by feedback.
Nickell bias
The downward bias in a dynamic fixed effects model, of order 1/T, cured by longer panels or by dynamic panel instruments.
Switcher
A unit whose regressor changes over the sample; units that never switch contribute nothing to a fixed effects estimate.
Random effects
A GLS estimator assuming a_i is uncorrelated with the regressors, which is the assumption fixed effects exists to avoid.

Module 5: Designs Built from Timing and Thresholds

Four designs that manufacture a counterfactual out of when a policy arrived or where a rule cut: difference in differences, event studies and staggered adoption, regression discontinuity, and synthetic control.

Newark and Easton: Difference in Differences on the New Jersey Minimum Wage

  • Compute a difference in differences estimate from a two-by-two table of means and reproduce it as a regression with an interaction term.
  • State the parallel trends assumption in potential-outcomes notation and explain why it is untestable.
  • Identify the standard threats to a two-period difference in differences, including functional form dependence.

On 1 April 1992 New Jersey raised its minimum wage from 4.25 dollars an hour to 5.05. Pennsylvania left its wage at the federal floor of 4.25. A Burger King in Newark and a Burger King in Easton, twelve miles apart across the Delaware, now faced different labour costs for the same job.

David Card and Alan Krueger had surveyed 410 fast food restaurants in the two states in February and March 1992, before the increase, and went back to almost all of them in November and December. Their outcome was full-time equivalent employment: full-time workers plus managers plus half the part-timers. The standard prediction from a competitive labour market model is that New Jersey employment should fall relative to Pennsylvania's. It rose.

The four numbers

Before (Feb-Mar 1992)After (Nov-Dec 1992)Change
New Jersey (treated)20.4421.03+0.59
Pennsylvania (control)23.3321.17-2.16
Difference in differences+2.75

The estimate is 0.59 - (-2.16) = 2.75 full-time equivalent workers per restaurant. Card and Krueger report 2.76, computed from unrounded means; the difference is rounding in the published table, not a discrepancy, and it is worth noticing that a third of a percent of arithmetic slack sits between what you can reproduce from a table and what the authors computed from the data.

Now see why neither simpler comparison would do. The before-and-after comparison within New Jersey gives +0.59, but 1992 was a recession year and fast food employment was falling nationally, so a positive change is not evidence of a positive effect. The cross-sectional comparison after the rise gives 21.03 - 21.17 = -0.14, but the two states differed by 2.89 workers per restaurant beforehand, so a small gap afterwards is not evidence of anything either. Each comparison contains a bias; differencing them removes the bias each one carries, provided the biases are the same.

The assumption, stated properly

In potential-outcomes notation, the estimand is

DiD = (E[Ypost|T] - E[Ypre|T]) - (E[Ypost|C] - E[Ypre|C])

Add and subtract the untreated potential outcome for the treated group after treatment, E[Y(0)post|T], which is the thing nobody observes:

DiD = ATT + { (E[Y(0)post|T] - E[Y(0)pre|T]) - (E[Y(0)post|C] - E[Y(0)pre|C]) }

The brace is the difference between what would have happened to New Jersey without the wage rise and what did happen to Pennsylvania. Parallel trends is the assumption that this brace is zero. Key idea: difference in differences does not assume the two groups are alike. It assumes they would have moved alike.

The levels in the table make the distinction concrete. Pennsylvania restaurants employed nearly three more workers than New Jersey ones before anything happened, and the design does not care, because the unit fixed effect implicit in the first difference removes it. What the design cannot tolerate is a reason for New Jersey's counterfactual path to differ from Pennsylvania's, and that is a statement about an unobservable, so it cannot be tested. Pre-trends can be inspected when more than one pre-period exists, which is the subject of the next lesson, but a matching pre-trend is evidence, not proof.

The same estimate as a regression

Run, over all restaurant-period observations,

Y = alpha + gamma NJ + lambda Post + delta (NJ times Post) + e

and the four coefficients reproduce the four cell means exactly. The intercept is the Pennsylvania pre-period mean, alpha = 23.33. The state dummy is the pre-period gap, gamma = 20.44 - 23.33 = -2.89. The period dummy is the control group's change, lambda = -2.16. And the interaction is the answer, delta = 2.75. Check the fourth cell: 23.33 - 2.89 - 2.16 + 2.75 = 21.03, which is the New Jersey post-period mean.

Writing it as a regression buys three things: covariates can be added, standard errors come out of the machinery, and the design generalises immediately to many units and many periods by replacing the two dummies with unit and time fixed effects. The last of those turns out to be far more dangerous than it looks, which is Lesson 13.

Functional form is part of the assumption

Parallel trends in levels and parallel trends in logs are different assumptions, and they cannot both hold unless the trend is zero. Take a clean case. Suppose in the absence of any policy both states' employment would have grown by 10 percent. New Jersey starts at 20 and would reach 22, a change of +2. Pennsylvania starts at 25 and would reach 27.5, a change of +2.5. In levels the counterfactual trends differ by 0.5, so a levels difference in differences would report -0.5 as a treatment effect when the truth is zero. In logs both changes are 0.0953, the brace is zero, and the log specification is correct.

Reverse the example, with both states growing by two workers, and levels are parallel while logs are not. There is no neutral choice. So what?: reporting a difference in differences in both levels and logs is not a robustness check, it is a report of which assumption you are willing to defend, and if the two disagree you owe the reader an argument rather than a footnote.

What people argued about afterwards

The finding was contested immediately, and the shape of the dispute is instructive because it was not about the method. David Neumark and William Wascher, publishing in 2000, collected payroll records from a partly overlapping set of New Jersey and Pennsylvania restaurants and reported that employment fell in New Jersey relative to Pennsylvania, by around 4 percent. Their objection was to the data: a telephone survey asking managers how many people worked at the restaurant is a noisier and possibly biased instrument than a payroll ledger.

Card and Krueger replied in the same issue using a third source, the Bureau of Labor Statistics ES-202 administrative payroll files covering essentially the entire fast food industry in both states, and again found no negative relative employment effect in New Jersey. The methodological framework was common ground throughout. The argument was about measurement, sample construction, and whose data covered which restaurants, which is what a productive disagreement in applied economics looks like.

Threats to a two-period design

  • The Ashenfelter dip. Units often enter a programme after a temporary bad spell, so their pre-period is unusually low and the post-period recovery is mistaken for an effect. Named for Orley Ashenfelter's observation in training programme evaluations.
  • Anticipation. New Jersey's increase was legislated before it took effect. Restaurants that adjusted hiring in advance contaminate the pre-period, which biases the estimate toward zero if the adjustment was negative.
  • Spillovers. The control group must be untreated. Pennsylvania restaurants near the border compete for the same workers and customers as New Jersey ones, so a wage rise on one side can move employment on the other, in which case the control is not a clean counterfactual.
  • Compositional change. The two waves must describe the same population. If restaurants that closed are missing from the second wave and closure correlates with treatment, the survivors' averages are compared on a shifted basis.
  • Inference. With two groups and two periods, clustering has nothing to work with, and the relevant sampling variation is across restaurants. In the multi-period, multi-state versions this becomes exactly the problem of Lesson 5, and the placebo law experiment was run on precisely this kind of specification.

Common misconceptions

  • "Difference in differences requires the groups to be similar." It requires their untreated paths to be parallel. Pennsylvania restaurants were nearly three workers larger and the design is untroubled by it.
  • "A matching pre-trend proves parallel trends." The assumption concerns the treated group's counterfactual path after treatment. Pre-period agreement is supporting evidence about a period in which nothing had happened yet.
  • "The result showed minimum wages raise employment." The estimate of +2.75 has a standard error of about 1.4 in the original paper. The defensible reading is that a moderate increase did not produce the substantial employment loss the standard model predicted, which is a different and more careful claim.
  • "Levels and logs are just a scaling choice." They encode different parallel trends assumptions, and a design that gives different answers in the two has told you something important about itself.

Where this leaves us

Four numbers, 20.44 and 21.03 against 23.33 and 21.17, give 0.59 - (-2.16) = 2.75, and the same estimate falls out as the interaction coefficient in a regression whose other three coefficients reproduce the remaining cell means. The identifying assumption is parallel counterfactual trends, not similarity of levels, and it is untestable in a two-period design and sensitive to whether you write the outcome in levels or in logs. The threats are anticipation, spillovers across a border, compositional change, and selection into treatment after a dip. And the substantive argument that followed was about measurement rather than about identification.

Two periods and two groups is the clean case. The next lesson takes the same regression, adds many states adopting at different times, and shows that the interaction coefficient stops meaning what it meant here.

Sources

  1. Card, D., & Krueger, A. B. (1994). Minimum wages and employment: A case study of the fast-food industry in New Jersey and Pennsylvania. American Economic Review, 84(4), 772-793. NBER Working Paper 4509. nber.org
  2. Huntington-Klein, N. (2022). Difference-in-differences. The Effect: An Introduction to Research Design and Causality. Chapman and Hall. theeffectbook.net
  3. Neumark, D., & Wascher, W. (2000). Minimum wages and employment: A case study of the fast-food industry in New Jersey and Pennsylvania: Comment. American Economic Review, 90(5), 1362-1396. aeaweb.org
  4. Card, D., & Krueger, A. B. (2000). Minimum wages and employment: A case study of the fast-food industry in New Jersey and Pennsylvania: Reply. American Economic Review, 90(5), 1397-1420. aeaweb.org
  5. Angrist, J. D., & Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist's Companion, Chapter 5. Princeton University Press.
Key terms
Difference in differences
The change in the treated group's outcome minus the change in the control group's, over the same two periods.
Parallel trends
The assumption that the treated group's untreated potential outcome would have moved like the control group's; untestable because it concerns a counterfactual.
Full-time equivalent employment
Card and Krueger's outcome: full-time workers plus managers plus half of part-time workers, per restaurant.
Interaction coefficient
The coefficient on treated times post in the regression form of DiD; it equals the two-by-two estimate exactly.
Ashenfelter dip
A temporary pre-treatment decline in the outcome among units that then select into treatment, mimicking a positive effect.
Anticipation
Behavioural response before a legislated policy takes effect, which contaminates the pre-period and biases the estimate.
Spillover
Treatment affecting the control group, which violates SUTVA and makes the control an invalid counterfactual.

The Forbidden Comparison: Event Studies and Staggered Adoption

  • Read an event study plot, choosing a reference period and interpreting leads and lags correctly.
  • Show on a six-observation panel how two-way fixed effects can return zero when every treatment effect is positive.
  • Name the estimators that avoid already-treated controls and state when plain two-way fixed effects is still safe.

Two units, three periods, and a treatment that helps both. Unit E is treated from period 2 and gains 1 in that period and 3 in the next, as the effect builds. Unit L is treated from period 3 and gains 1. Neither would have changed at all without treatment; both sit at 10.

UnitPeriod 1Period 2Period 3Treated from
E101113Period 2
L101011Period 3

The average effect across the three treated unit-periods is (1 + 3 + 1)/3 = 1.667. Now run the regression everybody runs: outcome on a treatment indicator, with unit fixed effects and period fixed effects. The coefficient is exactly 0.

Where the zero comes from

Two-way demeaning makes it visible. For the outcome, the unit means are 11.333 and 10.333, the period means are 10, 10.5 and 12, and the grand mean is 10.833. Subtracting unit mean and period mean and adding back the grand mean gives

Ytilde = -0.5, 0, 0.5 for unit E and 0.5, 0, -0.5 for unit L

Do the same to the treatment indicator, whose unit means are 2/3 and 1/3, period means 0, 0.5 and 1, and grand mean 0.5:

Dtilde = -0.1667, 0.3333, -0.1667 for E and 0.1667, -0.3333, 0.1667 for L

The coefficient is sum(Dtilde times Ytilde) / sum(Dtilde2). The numerator is 0.0833 + 0 - 0.0833 + 0.0833 + 0 - 0.0833 = 0. The denominator is 0.3333. The estimate is 0/0.3333 = 0.

Nothing went wrong with the arithmetic. The specification did what it was asked. What it was asked is the problem.

The forbidden comparison

Andrew Goodman-Bacon proved in 2021 that a two-way fixed effects estimate with staggered timing is a weighted average of every possible two-by-two difference in differences that the data allow. In this tiny panel there are exactly two.

Comparison one, periods 1 to 2. Unit E switches on; unit L is not yet treated and serves as a clean control. (11 - 10) - (10 - 10) = +1, which is E's period-2 effect exactly.

Comparison two, periods 2 to 3. Unit L switches on. The only available control is unit E, which is already treated. (11 - 10) - (13 - 11) = 1 - 2 = -1. The true effect on L is +1, and the comparison returns -1.

The second comparison is not measuring L's treatment effect against a stable counterfactual. It is measuring L's effect minus E's change in effect, and E's effect was growing from 1 to 3. Because the two comparisons carry equal weight here, the estimator averages +1 and -1 and reports 0. What matters here: the problem is not staggering by itself and not heterogeneity by itself; it is treatment effects that change over time being subtracted as if they were a control group's trend.

The general statement, from Clement de Chaisemartin and Xavier D'Haultfoeuille in 2020, is that the two-way fixed effects coefficient is a weighted average of unit-period treatment effects in which some weights are negative. With enough negative weight, the estimate can fall outside the range of every individual effect in the data, which is what has happened above.

The event study, and how to read it

The repair begins by refusing to compress the effect into one number. An event study specification replaces the single treatment dummy with a set of indicators for time relative to treatment:

Yit = ai + lt + sumk betak 1[t - gi = k] + eit

where gi is unit i's adoption period and one value of k must be omitted as the reference, conventionally k = -1, the period just before treatment. Three practical rules follow.

  • Every coefficient is relative to the omitted period. Moving the reference from -1 to -3 shifts the entire plotted path vertically. A plot without a stated reference period cannot be read.
  • Bin the endpoints. Very distant leads and lags are estimated off few units and are noisy; grouping everything beyond, say, five periods into a single indicator keeps the picture honest.
  • Leads are a diagnostic, not a result. Coefficients before treatment should be near zero if parallel trends holds, and if they are not the design is in trouble.

That last rule needs a serious caveat, made precise by Jonathan Roth. Pre-trend tests are typically underpowered: they routinely fail to reject differential trends large enough to change the conclusion. Worse, conditioning on passing the test distorts the estimator, because you are selecting samples in which the pre-period noise happened to look flat, and that selection can bias the post-treatment coefficients in a predictable direction. Reporting "we tested for pre-trends and found none" is therefore weaker evidence than it sounds, and Roth's recommendation is to report the magnitude of pre-trend the test could actually have detected.

Estimators that refuse the forbidden comparison

EstimatorComparison group usedWhat it reports
Two-way fixed effectsEverything, including already-treated unitsOne number, with possibly negative weights
Callaway and Sant'Anna (2021)Never-treated, or not-yet-treated, units onlyGroup-time effects ATT(g,t), aggregated as you choose
Sun and Abraham (2021)Never-treated or last-treated cohortInteraction-weighted event study coefficients
de Chaisemartin and D'Haultfoeuille (2020)Units whose treatment did not changeAn estimator for switchers, plus a negative-weight diagnostic
Imputation estimatorsUntreated observations, used to fit the counterfactualPredicted Y(0) for every treated cell, then averaged

The shared idea is simple once you have seen the -1 above: never use an already-treated unit as a control. Brantly Callaway and Pedro Sant'Anna's approach makes it explicit by estimating a separate ATT(g, t) for each adoption cohort g and each period t, always against units that have not yet been treated, and then leaving the aggregation to the researcher: an overall average, an average by cohort, or an event-study path.

Apply that to the toy panel. ATT(g=2, t=2) = +1 using L as a not-yet-treated control. ATT(g=2, t=3) has no clean control left, so it is not identified from these two units and is simply not reported. ATT(g=3, t=3) also has no untreated comparison. The honest output is one identified number, +1, and an explicit statement that the rest requires a never-treated group. Two-way fixed effects, by contrast, reported 0 without complaint.

When plain two-way fixed effects is still fine

Two conditions each rescue it. If every treated unit adopts at the same time, there is no already-treated unit to serve as a control, so no forbidden comparison exists and the estimator is the Lesson 12 design with more units. If treatment effects are constant across units and across time, then the growing-effect mechanism that produced the -1 is absent and the negative weights, though present, multiply identical numbers and cancel.

Neither condition is testable in the strong sense, but the first is a fact about your data that you can simply look up, and the diagnostic for the second is the event study path: an effect that climbs after adoption is exactly the pattern that breaks the estimator.

Common misconceptions

  • "Negative weights are a small technical correction." In the worked panel they moved the estimate from 1.667 to 0. Published re-analyses have found sign reversals in real applications.
  • "Adding unit and time fixed effects generalises difference in differences safely." It generalises the regression, not the design. With staggered timing the coefficient no longer corresponds to the two-by-two logic that justified it.
  • "Flat pre-trends validate the design." They are consistent with it. Roth's point is that the test often lacks the power to see violations large enough to matter, and that selecting on it introduces its own bias.
  • "The new estimators give you more." They usually give you less: fewer comparisons, wider intervals, and sometimes a refusal to report a cell at all. That is the correct response to a design that never contained the information.

The short version

Two-way fixed effects with staggered adoption is a weighted average of every available two-by-two comparison, and some of those use already-treated units as controls. When effects grow over time, those comparisons subtract the growth and enter with negative weight. On a six-observation panel where the true average effect is 1.667 and every individual effect is positive, the estimator returns exactly 0. Event study specifications expose the dynamics that cause the problem, provided you state the reference period, bin the endpoints, and resist over-reading an underpowered pre-trend test. The Callaway and Sant'Anna, Sun and Abraham, de Chaisemartin and D'Haultfoeuille, and imputation estimators all restrict comparisons to units that are not yet treated.

Bottom line: the assumption that breaks staggered difference in differences is not parallel trends alone. It is parallel trends plus the requirement that already-treated units have stopped changing, and the second half is almost never true.

Sources

  1. Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. Journal of Econometrics, 225(2), 254-277. NBER Working Paper 25018. nber.org
  2. Callaway, B., & Sant'Anna, P. H. C. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics, 225(2), 200-230. arxiv.org
  3. Roth, J. (2022). Pretest with caution: Event-study estimates after testing for parallel trends. American Economic Review: Insights, 4(3), 305-322. aeaweb.org
  4. de Chaisemartin, C., & D'Haultfoeuille, X. (2020). Two-way fixed effects estimators with heterogeneous treatment effects. American Economic Review, 110(9), 2964-2996. aeaweb.org
  5. Sun, L., & Abraham, S. (2021). Estimating dynamic treatment effects in event studies with heterogeneous treatment effects. Journal of Econometrics, 225(2), 175-199. arxiv.org
  6. Roth, J., Sant'Anna, P. H. C., Bilinski, A., & Poe, J. (2023). What's trending in difference-in-differences? A synthesis of the recent econometrics literature. Journal of Econometrics, 235(2), 2218-2244.
Key terms
Staggered adoption
A design in which units begin treatment at different dates, so already-treated units are available as comparison groups.
Forbidden comparison
A two-by-two difference in differences in which an already-treated unit serves as the control, subtracting its changing treatment effect as if it were a trend.
Goodman-Bacon decomposition
The result that a two-way fixed effects estimate is a weighted average of all available two-by-two comparisons, with computable weights.
Negative weights
Weights on some unit-period treatment effects that are less than zero, allowing the estimate to fall outside the range of all true effects.
Event study specification
A regression replacing the treatment dummy with indicators for time relative to adoption, with one period omitted as reference.
ATT(g,t)
Callaway and Sant'Anna's group-time average treatment effect: the effect in period t for the cohort adopting in period g, estimated against not-yet-treated units.
Pre-trend test
A test that pre-treatment event study coefficients are zero; underpowered in typical designs, and distorting when conditioned upon.

Winning by 0.1 Percent: Regression Discontinuity and the Incumbency Advantage

  • Estimate a sharp regression discontinuity by fitting local linear regressions on each side of a cutoff and differencing the intercepts.
  • Explain the bias-variance tradeoff in bandwidth choice and show numerically how the estimate responds to widening the window.
  • Apply the McCrary density test and covariate balance checks, and state the assumption whose failure destroys the design.

Before you read: write down what you think separates a US House candidate who wins by one tenth of a percentage point from one who loses by the same margin. Keep your list; it is the argument the whole design rests on.

David Lee assembled every US House election from 1948 to 1998 and lined them up by the Democratic candidate's margin of victory. On one side of zero the Democrat holds the seat going into the next election; on the other side the Republican does. Everything else about a district decided by 0.1 percent, its demography, its partisan history, the quality of its candidates, is essentially the same on both sides of that line, because nobody can control an election outcome to within a tenth of a percentage point.

What Lee found is that winning a seat raises the party's probability of winning the next election by something in the region of 40 to 45 percentage points, and its vote share in the next election by about 8 points. The incumbency advantage had been estimated for decades by comparing incumbents with non-incumbents, a comparison hopelessly contaminated by the fact that incumbents are, on average, the stronger candidates. The threshold removed that.

The design, stated formally

A regression discontinuity design needs a running variable X, a cutoff c, and a rule that assigns treatment based on which side of c a unit falls. In the sharp case, D = 1[X >= c] exactly.

The identifying assumption is continuity: both E[Y(0) | X = x] and E[Y(1) | X = x] are continuous at x = c. In words, everything other than the treatment that affects the outcome changes smoothly through the threshold, so any jump in the observed outcome at c must be the treatment. Under continuity,

ATE at the cutoff = limx to c from above E[Y|X=x] - limx to c from below E[Y|X=x]

Note what this does not assume. It does not assume the running variable is unrelated to the outcome; in Lee's data, a bigger margin this time genuinely predicts a bigger margin next time, and that relationship is estimated rather than assumed away. It assumes only that the relationship does not jump at the threshold for any reason except the treatment.

Estimating the jump, by hand

Suppose you bin the data by margin and compute the mean of the outcome, the probability that the Democrat wins the next election, in each bin. Within five points of the cutoff you observe:

Margin (bin centre)-5-3-1+1+3+5
Mean outcome0.280.320.360.800.840.88

The crudest estimate compares the two bins nearest the cutoff: 0.80 - 0.36 = 0.44. That treats the outcome as flat within 5 points on either side, which the table plainly contradicts, and the error is in the direction you would expect: the left-hand bin at -1 sits above the left-hand trend extrapolated to zero, and the right-hand bin at +1 sits below it.

Local linear regression does it properly. On the left, the three points lie on a line with slope (0.36 - 0.28)/4 = 0.02 per point, so the fitted value at the cutoff is 0.36 + 0.02(1) = 0.38. On the right, the slope is also 0.02, and the fitted value at the cutoff is 0.80 - 0.02(1) = 0.78. The estimated jump is

0.78 - 0.38 = 0.40

The crude estimate of 0.44 overstates it by a tenth. The point: a regression discontinuity estimate is the difference between two extrapolated intercepts, not a difference between two averages, and the extrapolation is short but never zero.

What happens when you widen the window

Now include bins out to nine points. The outcome curves away from the line as you move from the cutoff.

Margin-9-7-5-3-1+1+3+5+7+9
Mean outcome0.180.230.280.320.360.800.840.880.900.91

Fit a line to the five left-hand points. Their mean margin is -5 and their mean outcome is 1.37/5 = 0.274. With deviations -4, -2, 0, 2, 4, the cross-product sum is 0.376 + 0.088 + 0 + 0.092 + 0.344 = 0.900 and the squared-deviation sum is 40, so the slope is 0.0225 and the intercept at the cutoff is 0.274 + 0.0225(5) = 0.3865.

On the right, the mean margin is 5 and the mean outcome is 4.33/5 = 0.866. The cross-product sum is 0.264 + 0.052 + 0 + 0.068 + 0.176 = 0.560, giving a slope of 0.014 and an intercept of 0.866 - 0.014(5) = 0.796. The estimated jump is 0.796 - 0.3865 = 0.41.

Three numbers from one dataset: 0.44 from the nearest bins, 0.40 at a bandwidth of 5, 0.41 at a bandwidth of 9. And for contrast, the difference in means over the whole wide window, 0.866 - 0.274 = 0.592, which is not an RD estimate at all and is nearly half again too large.

Bandwidth is a bias-variance tradeoff and nothing else. Narrow it and the linear approximation gets better, so bias falls, while the number of observations falls and variance rises. Widen it and the reverse. Guido Imbens and Karthik Kalyanaraman derived a bandwidth that minimises mean squared error, and Sebastian Calonico, Matias Cattaneo and Rocio Titiunik supply confidence intervals that correct for the bias the optimal bandwidth deliberately leaves in, which the conventional interval ignores. Reporting the estimate across a range of bandwidths is standard, and an estimate that moves a great deal across that range is telling you something.

Resist the temptation to fit a high-order global polynomial instead. Andrew Gelman and Guido Imbens showed that global polynomials of order four or above give weights to distant observations that no one would choose deliberately, and can produce dramatic artefacts at the boundary. Local linear or local quadratic within a chosen bandwidth is the standard.

Fuzzy discontinuities

Sometimes crossing the threshold changes the probability of treatment without determining it. Scoring above a cutoff makes a scholarship available but not compulsory; a poverty index above a line qualifies a household for a programme it may decline. Then the design is instrumental variables with 1[X >= c] as the instrument, and the estimate is the Wald ratio:

fuzzy RD estimate = (jump in E[Y|X] at c) / (jump in E[D|X] at c)

Everything from Lesson 10 applies. The estimate is a LATE, defined on units whose treatment status the threshold changed, and a small jump in treatment probability in the denominator amplifies any violation of the exclusion restriction, which here means any other consequence of crossing the line.

The test that can kill the design

Continuity fails when units can precisely manipulate their position relative to the cutoff. If a household knows the poverty index formula and can adjust an input, or a teacher can nudge a borderline exam mark, the units just above and just below the line are no longer comparable: they differ by who chose to cross it.

Justin McCrary's 2008 test looks for the fingerprint. If manipulation occurred, the density of the running variable should jump at the cutoff, with a pile-up on the favourable side and a hole on the other. Estimate the density separately on each side and test whether the two limits agree. A large jump is close to disqualifying.

Three further checks belong in any RD paper. Covariate balance: run the same discontinuity estimate with pre-determined covariates as outcomes; they should show no jump. Placebo cutoffs: repeat the estimation at fake thresholds where nothing happens, and confirm no jumps appear. Donut estimates: drop observations immediately adjacent to the cutoff and re-estimate, which guards against heaping and rounding at the threshold itself.

Close elections have been argued over on exactly these grounds. Devin Caughey and Jasjeet Sekhon reported in 2011 that in post-war US House elections, bare winners looked systematically different from bare losers on pre-determined covariates, which suggests the margin was not beyond manipulation. Andrew Eggers and coauthors then examined some 40,000 close elections across many countries and offices in 2015 and found the imbalance to be specific to that one setting rather than a general property of close-election designs. The episode is a good model of how a design gets audited: a specific claim, a specific test, and a much larger replication.

Common misconceptions

  • "RD estimates the average effect of the treatment." It estimates the effect at the cutoff, for units near it. Whether the effect of winning a seat by a whisker resembles the effect of winning comfortably is a separate question the design cannot answer.
  • "More data far from the cutoff improves the estimate." Observations far away enter only through the extrapolation and can bias it. The information lives near the threshold, which is why RD studies are often underpowered despite enormous samples.
  • "Controlling for the running variable in a regression is the same thing." Including X as a covariate over the whole range imposes one functional form globally and uses distant observations to identify the jump. RD deliberately restricts to a neighbourhood and fits each side separately.
  • "Manipulation of the running variable always invalidates the design." Only precise manipulation does. Units that influence their score imprecisely, with noise they cannot control, still land on either side effectively at random near the threshold, which is exactly the argument for close elections.

What to carry forward

A rule that assigns treatment at a threshold creates a comparison as good as random in a neighbourhood of that threshold, provided potential outcomes are continuous there. The estimate is the difference between two locally fitted intercepts: 0.40 at a bandwidth of 5 in the worked data, 0.41 at a bandwidth of 9, against 0.44 from the nearest bins and 0.592 from a naive difference in means. Bandwidth trades bias against variance and should be reported across a range. Fuzzy designs are IV with the threshold as instrument, and inherit LATE and weak instrument concerns. The McCrary density test, covariate balance at the cutoff, placebo thresholds and donut estimates are the audit.

In short: the assumption that destroys a regression discontinuity is precise manipulation of the running variable. If units can place themselves on the side of the line they prefer, the two sides differ by choice rather than by chance, and no bandwidth will repair it.

Sources

  1. Lee, D. S. (2008). Randomized experiments from non-random selection in U.S. House elections. Journal of Econometrics, 142(2), 675-697. NBER Working Paper 8441. nber.org
  2. Lee, D. S., & Lemieux, T. (2010). Regression discontinuity designs in economics. Journal of Economic Literature, 48(2), 281-355. aeaweb.org
  3. Cunningham, S. (2021). Regression discontinuity. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  4. Imbens, G. W., & Kalyanaraman, K. (2012). Optimal bandwidth choice for the regression discontinuity estimator. Review of Economic Studies, 79(3), 933-959. NBER Working Paper 14726. nber.org
  5. McCrary, J. (2008). Manipulation of the running variable in the regression discontinuity design: A density test. Journal of Econometrics, 142(2), 698-714.
  6. Eggers, A. C., Fowler, A., Hainmueller, J., Hall, A. B., & Snyder, J. M. (2015). On the validity of the regression discontinuity design for estimating electoral effects. American Journal of Political Science, 59(1), 259-274.
Key terms
Running variable
The continuous variable whose value relative to a cutoff determines treatment; also called the forcing or assignment variable.
Sharp RD
A design where treatment is a deterministic step function of the running variable at the cutoff.
Fuzzy RD
A design where the cutoff changes the probability of treatment; estimated as IV, giving a LATE at the threshold.
Continuity assumption
Both potential outcome functions are continuous in the running variable at the cutoff, so any jump must be the treatment.
Bandwidth
The window of the running variable used in local estimation; narrower reduces bias and raises variance.
McCrary density test
A test for a jump in the density of the running variable at the cutoff, the signature of precise manipulation.
Donut RD
Re-estimating after dropping observations immediately adjacent to the cutoff, to guard against heaping and rounding.
Local linear regression
Fitting a separate straight line on each side of the cutoff within the bandwidth and differencing the intercepts at the threshold.

Building California: Synthetic Control When You Have One Treated Unit

  • Compute synthetic control weights by hand from a pre-treatment system of equations and verify them on periods that were not used to fit them.
  • State the synthetic control optimisation, explain why the weights are restricted to the simplex, and say what that restriction rules out.
  • Run and interpret a permutation test based on the ratio of post-treatment to pre-treatment fit, and name the assumption whose failure destroys the estimate.

On 8 November 1988, 58.17 percent of Californian voters approved Proposition 99. From 1 January 1989 every packet of cigarettes sold in the state carried an extra 25 cents of excise tax, with the revenue earmarked for anti-tobacco advertising and health programmes. Sales per head, already drifting down through the 1980s, fell faster afterwards.

How much of that fall was the proposition? Bring the last three lessons to the question and each dies at the first step. No cutoff, so no regression discontinuity. No instrument for living in California. Difference in differences needs a control state, and the instant you name one the question is why that one: Utah smokes far less to begin with, Nevada sells enormously to visitors, New York raised its own tobacco tax in the same window. What Abadie, Diamond and Hainmueller published in 2010 was a way to stop choosing a control and start building one.

Choosing the control chooses the answer

Strip the problem to six years and four units. T is treated at the start of year 5; A, B and C never are. The numbers are packs sold per head per year.

UnitYear 1Year 2Year 3Year 4Year 5Year 6
T (treated from year 5)1009692908276
A140132124120116112
B807876757473
C100100100100100100

Each of these is a defensible difference in differences, comparing T's change from year 4 to year 6 with a control's over the same span.

  • Against A: (76 - 90) - (112 - 120) = -14 + 8 = -6.
  • Against B: (76 - 90) - (73 - 75) = -14 + 2 = -12.
  • Against C: (76 - 90) - (100 - 100) = -14.
  • Against the simple average of all three, which falls from 295/3 = 98.33 to 285/3 = 95: -14 + 3.33 = -10.67.

Four answers spanning more than a factor of two, and nothing in the arithmetic says which to publish. The pre-treatment check from Lesson 12 disqualifies all four: over years 1 to 4, T falls by 10 while A falls by 20, B by 5, C by 0 and the average by 8.33. Nothing here moves in parallel with T before treatment, so nothing here has earned the right to stand in for T after it.

Worth holding on to: when the treated unit is a single region, the choice of control is not a step before the estimate. It is the estimate.

Solving for the weights instead of picking a state

The move is to allow a weighted average of the donors and let the pre-treatment data choose the weights. Write wA, wB, wC, non-negative and summing to 1, and ask for the combination whose pre-treatment path matches T's. Two of the four years pin them down. Impose the match in year 1 and year 4, substituting wC = 1 - wA - wB.

Year 1: 140wA + 80wB + 100(1 - wA - wB) = 100, which simplifies to 40wA - 20wB = 0, so wB = 2wA.

Year 4: 120wA + 75wB + 100(1 - wA - wB) = 90, which simplifies to 20wA - 25wB = -10. Substituting wB = 2wA gives -30wA = -10.

So wA = 1/3, wB = 2/3, wC = 0.

Now the step that makes the method worth trusting. Those weights were fixed by years 1 and 4 alone; years 2 and 3 had no vote. Apply them anyway: (1/3)(132) + (2/3)(78) = 96, and (1/3)(124) + (2/3)(76) = 92. Both land exactly on T. The synthetic unit reproduces the treated unit in two periods it was never fitted to, which is out-of-sample performance inside the pre-treatment window and the closest thing this design has to a test of itself.

SeriesYear 1Year 2Year 3Year 4Year 5Year 6
T (actual)1009692908276
Synthetic T1009692908886
Gap0000-6-10

The estimated effect is 6 packs in year 5 and 10 in year 6, against the -6, -12, -14 and -10.67 from the four single-control comparisons. And look at unit C: it matched T's level exactly in year 1 and got zero weight, because matching a level in one year is not matching a path across four.

The general statement, and what the constraints buy

You have J + 1 units. Unit 1 is treated from period T0 + 1; units 2 to J + 1 are the donor pool. Let X1 hold pre-treatment characteristics for the treated unit and X0 the same for the donors. The synthetic control method chooses W = (w2, ..., wJ+1) to minimise

(X1 - X0W)' V (X1 - X0W) subject to wj >= 0 for all j and sumj wj = 1

and then reports, for each post-treatment period, effectt = Y1t - sumj wj Yjt. The matrix V says how much each predictor counts, and is itself chosen by an outer optimisation: pick the V whose resulting weights minimise prediction error over the pre-treatment outcomes.

The two constraints are not housekeeping. Non-negativity and the sum-to-one rule confine the synthetic unit to the convex hull of the donors, so the counterfactual interpolates and never extrapolates. An unrestricted regression would cheerfully put -0.8 on Utah and 1.9 on Nevada, building a state that smokes a negative number of packs. The simplex forbids that, and makes the answer sparse: most donors get exactly zero, so the weights can be printed and argued about. If the treated unit sits outside the hull, no convex combination can match it, the fit will be visibly poor, and Abadie's advice is to stop rather than report the gap anyway.

Why this is not difference in differences with cleverer weights

Suppose untreated outcomes follow Yit(0) = deltat + thetatZi + lambdatmui + eit, where mui is an unobserved characteristic of unit i and lambdat is how much it matters in period t. Difference in differences is the special case with lambdat constant: the unobservable may differ across units, but its influence must not change over time, which is parallel trends in another notation. Synthetic control relaxes exactly that. A combination matching the treated unit across many periods carrying different lambdat has matched mui itself, and Abadie, Diamond and Hainmueller bound the bias by a term shrinking as the pre-period lengthens relative to the transitory shocks eit.

So what?: matching four pre-treatment years with three donors is curve fitting; matching nineteen years with a handful of donors out of thirty-eight is evidence.

California, with the real numbers

Abadie, Diamond and Hainmueller start from the fifty states, drop California, then prune for contamination. Massachusetts, Arizona, Oregon and Florida ran large tobacco programmes of their own; Alaska, Hawaii, Maryland, Michigan, New Jersey, New York and Washington raised cigarette taxes by 50 cents a pack or more. All eleven go, leaving 38 donors. The predictors are retail price, income per head, the share aged 15 to 24, beer consumption per head, and lagged sales in 1975, 1980 and 1988, matched over 1970 to 1988.

Donor stateWeight
Utah0.334
Nevada0.234
Montana0.199
Colorado0.164
Connecticut0.069
The other 33 states0.000

Five states carry the whole counterfactual. Synthetic California tracks the real one closely from 1970 to 1988, separates after 1989, and by 2000 the gap has widened to roughly 26 packs a head a year, averaging around 20 a year across 1989 to 2000. Notice how much more you can say about this than about a regression coefficient: you can name the five states, look at the pre-1989 fit, and ask whether Utah and Nevada plausibly stand in for California.

Inference when the sample size is one

There is no standard error here in the Module 2 sense. Abadie and coauthors run a permutation test instead: pretend in turn that each of the 38 donors passed Proposition 99 in 1989, build each a synthetic control from the remaining states, and record the gap.

The statistic is not the raw gap. A state the method fitted badly before 1989 will show a large gap after it for reasons unrelated to tobacco. So the statistic is the ratio of post-treatment mean squared prediction error to pre-treatment mean squared prediction error, which asks how much worse the synthetic control became relative to how good it was. California's ratio is the largest of all 39, and under a null of no effect the treated unit's rank among 39 exchangeable units is uniform, so the one-sided p-value is 1/39 = 0.026.

Two further checks belong alongside it. An in-time placebo re-runs the estimation pretending treatment arrived in 1980, fitting only on 1970 to 1979; a gap opening in 1981 says the method manufactures gaps. A leave-one-out run drops each weighted donor in turn.

Common misconceptions

  • "A perfect pre-treatment fit proves the design works." With many donors and few periods you can fit anything, noise included. Matching four periods with thirty-eight donors is guaranteed, and guarantees nothing. What matters is the ratio: many periods, few donors, a fit holding where the weights were not fitted.
  • "It is difference in differences with better weights." The weights are the visible difference; the assumption is the real one. Difference in differences needs the unobservable's influence to hold constant over time. Synthetic control lets it vary, and pays by demanding a match on the whole pre-treatment path.
  • "The permutation p-value is an ordinary p-value." It is a rank, so its resolution is bounded below by 1/(J + 1): with ten donors the smallest value obtainable is 1/11 = 0.091, whatever the effect size, and reporting p below 0.01 from a study with 20 donors is arithmetically impossible.
  • "More donors are always better." A donor that was itself treated can receive weight precisely because it happens to fit the pre-period. Abadie and coauthors deleted eleven states before any weight was computed, on substantive grounds, and that ordering matters.

What you now know

When one unit is treated and no control is obviously right, construct one: a convex combination of untreated units whose pre-treatment path reproduces the treated unit's. In the worked table, weights of 1/3 and 2/3 were fixed by two years and then reproduced two more, giving effects of 6 and 10 packs where four ordinary difference in differences comparisons had given -6, -12, -14 and -10.67. In California, five of 38 pruned donors carry the counterfactual, the gap reaches about 26 packs a head by 2000, and the permutation rank is one in 39.

The point: the assumption that destroys a synthetic control estimate is that the weighted donors would have gone on tracking the treated unit had nothing happened. It fails if the pre-treatment match was a match on noise rather than on the factors driving the outcome, if a heavily weighted donor took a shock of its own at the same moment, or if the treated unit moved in anticipation of the policy.

Sources

  1. Abadie, A., Diamond, A., & Hainmueller, J. (2010). Synthetic control methods for comparative case studies: Estimating the effect of California's tobacco control program. Journal of the American Statistical Association, 105(490), 493-505. NBER Technical Working Paper 335. nber.org
  2. Abadie, A. (2021). Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature, 59(2), 391-425. aeaweb.org
  3. Cunningham, S. (2021). Synthetic control. Causal Inference: The Mixtape. Yale University Press. mixtape.scunning.com
  4. Abadie, A., & Gardeazabal, J. (2003). The economic costs of conflict: A case study of the Basque Country. American Economic Review, 93(1), 113-132.
Key terms
Donor pool
The set of untreated units eligible to receive weight; pruned on substantive grounds before any weight is computed.
Synthetic control unit
A weighted average of donor units whose pre-treatment path reproduces the treated unit's, used as the counterfactual.
Convex hull restriction
Weights are non-negative and sum to one, so the counterfactual interpolates between donors and never extrapolates beyond them.
Predictor weights V
The diagonal matrix saying how much each pre-treatment characteristic counts, chosen by an outer optimisation over pre-treatment fit.
MSPE ratio
Post-treatment mean squared prediction error divided by pre-treatment mean squared prediction error; the statistic used in the permutation test.
Permutation test
Re-running the estimation pretending each untreated unit was treated, and ranking the real unit's statistic among the placebos.
In-time placebo
Re-estimating with a fake treatment date inside the pre-treatment period, to check the method does not manufacture gaps.
Interactive fixed effects
A model where an unobserved unit characteristic has a time-varying influence; the structure synthetic control can handle and fixed effects cannot.

Module 6: Prediction, Machine Learning, and the Argument About Evidence

Two problems that look alike and are not, what an algorithm can and cannot be asked to settle, and the forty-year dispute over whether research design repaired empirical economics or narrowed it.

Rain Dances and Umbrellas: Prediction, Causal Inference, and Where Machine Learning Fits

  • Separate a prediction problem from a causal problem using the decomposition of a policy payoff, and say which quantity each requires.
  • Show numerically why a regularised coefficient cannot be read as a treatment effect, and why a variable set chosen for predictive accuracy is not a set of confounders.
  • State what double machine learning does through orthogonalisation and cross-fitting, and what it leaves entirely to the analyst.

Regress test scores on class size across the eight districts of Lesson 2 and the slope is -1.5, with an R-squared of 0.900. Add median household income and the slope becomes -0.5 while the R-squared rises to 0.929. Both numbers came out of the same eight rows. If your job is to forecast next year's score for a ninth district, the second model is simply better and neither slope interests you. If your job is to decide whether to hire teachers, the fit statistic was never evidence about anything, and the only number that matters is the one that moved.

These are not two views of one problem. They are two problems, and most of the confusion about machine learning in economics comes from running them together.

The derivative that separates them

Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan and Ziad Obermeyer put the distinction as two droughts. A policymaker deciding whether to hire a rain dancer has to know whether dancing causes rain. A policymaker deciding whether to carry an umbrella to the ceremony has only to know whether it will rain. Write the payoff as pi(X0, Y), where X0 is the decision and Y the outcome, and differentiate:

d pi / d X0 = (partial pi / partial X0) + (partial pi / partial Y)(dY / dX0)

The second term carries a causal object, dY/dX0. When the decision does not move the outcome, that term is zero and the whole problem reduces to knowing Y accurately. Their leading example is joint replacement surgery under Medicare: the operation pays off over years, so a patient who will not live those years gains little, and identifying such patients in advance requires no causal estimate at all. Whether the operation works is a causal question that has already been answered elsewhere; whom to operate on is a forecast.

The boundary is less tidy than it looks. Deciding which defendants to release on bail seems to need only a forecast of who will fail to appear, but that outcome is never observed for anyone a judge detains, so the training data are labelled by the very decisions the algorithm is meant to improve. That selective-labels problem is causal in structure even though the tool is a predictor.

Side by side

QuestionPrediction problemCausal problem
What is wantedAn accurate guess of YHow Y responds to an intervention on D
Can success be measured?Yes, on held-out dataNo; it rests on an assumption you argue for
How the model is chosenCross-validationBy design, not by fit
Two highly correlated predictorsInterchangeable and harmlessA live question about which is the confounder
Chief enemyOverfittingOmitted variables, bad controls, selection
What a high R-squared showsThat the model predictsNothing
Effect of shrinkageLowers mean squared errorBiases the parameter of interest
More data buysA better fitA tighter interval around the same bias

The upshot: every row of that table is a place where a habit trained on one problem does damage on the other.

What a penalty does to a coefficient

Take the eight districts again. Their cross-products were Sxy = -168 and Sxx = 112, so the least squares slope is -168/112 = -1.5. Ridge regression adds a penalty on the squared coefficient, and in the bivariate case the estimator has a closed form:

b(lambda) = Sxy / (Sxx + lambda)

Penalty lambda028112448
Estimated slope-1.500-1.200-0.750-0.300

Read that table next to Lesson 6. There the slope fell from -1.5 to -0.5 because a genuine confounder was restored. Here it falls to -0.75 because a penalty was tuned by cross-validation to forecast well. The two coefficients are similar sizes and mean completely different things, and nothing in the printed number tells them apart. Regularisation buys lower prediction error by accepting bias, which is an excellent trade when the target is a forecast and a disqualifying one when the target is an effect.

The lasso raises a second problem. It sets coefficients to exactly zero, so it performs variable selection, and it is tempting to read the surviving variables as the ones that matter. Sendhil Mullainathan and Jann Spiess show why that reading fails: when predictors are correlated, refitting on a different sample keeps a noticeably different set of variables while predicting about as well. For forecasting this is harmless, since the substitutes carry the same information. For a causal claim it is fatal, because which variable is in the model is precisely the confounding question, and the algorithm decided it on predictive grounds.

Double machine learning, stated

There is a real problem worth solving here. Suppose

Y = D theta + g(X) + U and D = m(X) + V, with E[U | X, D] = 0 and E[V | X] = 0

where theta is the parameter you want and g and m are nuisance functions of possibly many controls whose shape you do not know. Fitting g with a flexible learner and then regressing Y on D and the fitted values does not work: the learner is regularised, its error is biased rather than merely noisy, and that bias passes straight into theta, at a rate that does not vanish fast enough for the usual asymptotics.

Victor Chernozhukov and coauthors repair it with two devices. The first is orthogonalisation. Predict Y from X and predict D from X, each with whatever learner fits best, take the residuals, and regress the residual of Y on the residual of D. This is the Frisch-Waugh-Lovell partialling out of Lesson 6 with machine learning in the partialling step, and the resulting moment condition is insensitive, to first order, to small errors in either nuisance estimate. The second is cross-fitting. Split the sample, estimate the nuisance functions on one part, evaluate the score on the other, then swap and average, so no observation helps fit the function that is later used to residualise it. Together these give an estimate of theta that converges at the usual root-n rate and carries usable confidence intervals, even though nothing about g or m was assumed in advance.

Be exact about what has been bought. The condition E[U | X, D] = 0 is the conditional independence assumption of Lesson 7, and it is assumed here, not tested. Double machine learning frees you from having to guess whether the control enters linearly, quadratically or through an interaction. It does not tell you whether the controls you have are the controls you need, and it will happily deliver a beautifully orthogonalised estimate of a badly confounded parameter.

Why this matters: the flexible part of a modern causal estimator is the functional form of the controls. The assumption underneath is the same one you have been defending since Module 3.

Where the tools genuinely earn their place

  • Selecting controls from many candidates. Post-double-selection runs a lasso of the outcome on the candidates and a second of the treatment on the candidates, and keeps the union. Selecting once, on the outcome alone, drops variables that matter for treatment and reintroduces bias.
  • First stages with many instruments. When dozens of weak instruments are available, a regularised first stage avoids the many-instruments bias of Lesson 10 while retaining power.
  • Heterogeneous effects. Random forest methods adapted for causal estimation search for subgroups whose effects differ, with sample splitting so the discovery does not contaminate the inference.
  • Propensity scores and balance. Estimating the score of Lesson 7 with a flexible learner, then checking balance rather than fit, is a defensible use.
  • Genuine prediction policy problems. Triage, targeting, inspection scheduling: decisions that need a good forecast and no causal estimate at all.

Common misconceptions

  • "A model that predicts well has the right causal structure." The collider of Lesson 8 predicts beautifully and is causally poisonous. Predictive accuracy is a statement about a joint distribution, and every design in Modules 4 and 5 exists because joint distributions do not answer intervention questions.
  • "Machine learning can find the confounders." It finds predictors. A mediator predicts the outcome very well and must stay out; a collider predicts it well and must stay out. The distinction is causal knowledge the data do not contain.
  • "Double machine learning solves identification." It solves a nuisance-estimation problem inside an assumption. Change the assumption and the estimate changes with it.
  • "Big data makes the bias small." From Lesson 3, plim b = beta + Cov(x,u)/Var(x), a quantity with no n in it. A million observations give you a very precise interval centred on the wrong number, and precision reported without identification is the most confident way to be wrong.

Summing up

A prediction problem wants an accurate guess and can be scored on held-out data; a causal problem wants a response to an intervention and can never be scored at all. Cross-validation answers the first and is silent on the second. Ridge on the eight districts moved the slope from -1.5 to -0.75 with no confounder anywhere in sight, and the lasso's selected variables shift across samples without any loss of predictive accuracy. Double machine learning makes flexible learners safe for a causal parameter through orthogonalisation and cross-fitting, and leaves the identifying assumption exactly where it was.

Bottom line: the assumption that, if false, breaks a double machine learning estimate is unconfoundedness given the controls you happened to collect. No algorithm has ever supplied it, and the more flexible the fitting, the easier it is to mistake a good fit for one.

Sources

  1. Kleinberg, J., Ludwig, J., Mullainathan, S., & Obermeyer, Z. (2015). Prediction policy problems. American Economic Review, 105(5), 491-495. aeaweb.org
  2. Mullainathan, S., & Spiess, J. (2017). Machine learning: An applied econometric approach. Journal of Economic Perspectives, 31(2), 87-106. aeaweb.org
  3. Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., & Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters. Econometrics Journal, 21(1), C1-C68. arxiv.org
  4. Belloni, A., Chernozhukov, V., & Hansen, C. (2014). High-dimensional methods and inference on structural and treatment effects. Journal of Economic Perspectives, 28(2), 29-50. aeaweb.org
  5. Athey, S., & Imbens, G. W. (2019). Machine learning methods that economists should know about. Annual Review of Economics, 11, 685-725.
Key terms
Prediction problem
A decision whose payoff depends on the outcome but not on how the decision changes it; scorable on held-out data.
Regularisation
Penalising coefficient size to lower prediction error, at the cost of biasing every coefficient toward zero.
Cross-validation
Choosing a model by its performance on data held out of the fit; a criterion for prediction, not for identification.
Partially linear model
Y = D theta + g(X) + U with g unknown; the parameter theta is of interest and g is a nuisance function.
Orthogonalisation
Residualising outcome and treatment on the controls separately, so the estimate is first-order insensitive to nuisance estimation error.
Cross-fitting
Estimating nuisance functions on one part of the sample and evaluating the score on another, to keep own-observation overfitting out of the estimate.
Post-double-selection
Selecting controls by lasso on the outcome and again on the treatment, then keeping the union of both sets.
Selective labels
Training data whose outcomes are observed only for units a past decision maker chose to release or treat.

Taking the Con Out: The Credibility Revolution as a Live Dispute

  • Reconstruct Leamer's 1983 demonstration of specification fragility and explain what extreme bounds analysis proposed in its place.
  • State the strongest version of the design-based case and the strongest version of the structural and external-validity objections, with the evidence each rests on.
  • Identify what the two positions agree on, what evidence would move the dispute, and which questions the revolution has not settled.

Does an execution deter murder? In 1983 Edward Leamer took the published cross-state regressions on that question and refitted them under different sets of control variables, each set defensible and each standing for a different view of what drives homicide: offenders weighing costs, poverty and age structure, retribution rather than deterrence. The estimated number of murders deterred by one execution swung from strongly positive to negative across those specifications. The sign of the answer was a function of the analyst.

Leamer's line about the state of the field is quoted more often than his remedy: "hardly anyone takes anyone else's data analysis seriously". The remedy was extreme bounds analysis, which asks the author to define the space of specifications a reasonable reader might accept and to report the range of estimates across it, rather than the one that survived an unreported search. If the bounds contain zero, say so.

Extreme bounds analysis did not become standard practice. Something else did.

What research design changed, according to its advocates

In 2010 Joshua Angrist and Jorn-Steffen Pischke argued in the Journal of Economic Perspectives that the con had largely been taken out, and not by sensitivity analysis. It had been taken out by refusing to run Leamer's regression at all.

Their evidence is the machinery of this course. When Vietnam draft numbers are drawn from a drum, the comparison does not depend on which controls the author admits, because the instrument was assigned by a lottery. When New Jersey raises its minimum wage and Pennsylvania does not, the counterfactual is a place rather than a covariate list. When a candidate wins by a tenth of a point, the comparison group is chosen by the electorate. Each design has an assumption, and the assumption is stated as a sentence about the world rather than buried in a specification: the lottery number affects earnings only through service, the two states would have moved together, potential outcomes are continuous at the cutoff. Those sentences can be argued about by people who cannot read a regression table, which is the point.

Angrist and Pischke also pointed to volume. Randomised evaluations went from rare to routine in development and education economics; administrative datasets replaced small surveys; the share of empirical papers in leading journals resting on an explicit design rose steeply. In 2021 the Nobel committee awarded the prize in economic sciences to David Card for empirical work in labour economics and to Angrist and Guido Imbens for the methodology of causal inference, which is as close to institutional ratification as a research programme gets.

The core of it: the claim is not that designs are assumption-free. It is that they move the assumption from a place where it is invisible to a place where it can be attacked.

The case that the revolution narrowed the field

The same 2010 symposium carried three objections, and they are not variations on one complaint.

Christopher Sims argued that the questions that matter most in macroeconomics admit no quasi-experiment. There is no control country for a monetary regime, no lottery assigning fiscal consolidations. If the only admissible evidence comes from designs, then the effect of an interest rate decision is simply not studied, and declining to answer an important question is not neutrality; it is a decision about what economics is for. Assumptions do not disappear when you refuse to write a model, they migrate into what you choose to estimate.

Michael Keane pressed a different point. A local average treatment effect is a number attached to a population of compliers under a policy that was actually tried. Policy usually asks about something that has not been tried: a tax credit at a rate never used, a programme scaled from a district to a nation. Getting from an estimated effect to that counterfactual requires a model of behaviour whether or not the analyst admits to having one, and a structural model at least writes its assumptions down where they can be checked. Aviv Nevo and Michael Whinston added that design-based purism can become its own dogma, since credibility is a property of an argument rather than of a technique.

Leamer himself replied in the same issue and did not concede. His view was that the sensitivity problem had been relocated rather than solved: choices about samples, bandwidths, clustering levels, functional forms and which of several designs to report still leave an author with room, and the room is rarely mapped for the reader.

What randomisation does not do

The sharpest recent statement of the sceptical case is Angus Deaton and Nancy Cartwright's 2018 paper, and its force comes from conceding the easy points first. Randomisation is genuinely valuable and genuinely does remove selection bias in expectation. The trouble is what people take that to mean.

First, balance in expectation is not balance. Randomisation guarantees that, averaged over many hypothetical repetitions, treated and control groups match on every covariate, observed or not. It guarantees nothing about the single trial you ran. In a small trial, one confounder being badly imbalanced is ordinary, and the standard error does not know which covariate it was.

Second, the estimate is a mean. When outcomes are skewed, which is normal for earnings, medical costs and firm revenue, a mean can be driven by a handful of observations, and the conventional standard error behaves badly in exactly those cases. Alwyn Young reanalysed dozens of published experimental papers using randomisation inference and found that a substantial share of individually significant treatment effects were no longer significant when the test respected the actual assignment mechanism.

Third, and most consequential, a trial tells you what happened in one population at one scale under one implementation. Transporting it requires knowing why the effect occurred and whether the causal structure holds elsewhere, which is theory. Deaton and Cartwright's conclusion is not that trials are worthless but that they earn their keep only inside an argument about mechanism. External validity is not a footnote to be added at the end; it is the reason anyone outside the study site cares.

Where the two camps actually agree

Both sides want the assumption stated and testable. Both agree that a coefficient without an identifying argument is decoration. Both agree that the specification search Leamer described was real and that it produced literatures nobody could believe. The disagreement is over where the assumption should sit: inside a stated design and a narrow population, or inside a written model and a broader question. Framed that way, the case for doing both is not a compromise but a research strategy, and it is what the strongest applied work now does: identify a parameter with a design, then use it to discipline a model that can answer the counterfactual the design cannot reach.

What would move the dispute

Three kinds of evidence bear on it, and only some have been gathered.

Evidence on whether design-based results replicate has begun to arrive, and it is mixed by method. Abel Brodeur, Nikolai Cook and Anthony Heyes examined roughly twenty thousand hypothesis tests in leading journals and found the distribution of test statistics bunching just past conventional significance thresholds, with instrumental variables and difference in differences noticeably worse than randomised trials and regression discontinuity. That is a finding about the credibility revolution's own tools, produced by its own standards.

Evidence on transportability would settle more. When the same intervention is run in several populations, do the effects agree, and are the disagreements predicted by theory? Where this has been done, the effects vary in ways a mechanism can often explain, which supports Deaton and Cartwright's framing more than it undermines it.

Evidence on out-of-sample structural prediction would settle most of all. Fit a structural model on non-experimental data, use it to predict the result of an experiment it never saw, and compare. This is the cleanest test available of whether written-down assumptions buy anything, and it is still done far too rarely to draw conclusions from.

Common misconceptions

  • "The credibility revolution replaced assumptions with data." It replaced one kind of assumption with another. Exclusion restrictions, parallel trends and continuity at a cutoff are assumptions, and none of them is testable in the way a coefficient is.
  • "Leamer was refuted." His demonstration stands. What changed is that his remedy lost to a different one. His 2010 reply argues the specification problem simply moved to bandwidths, clustering and design choice, and the p-hacking evidence gives that view some support.
  • "Structural modelling means assuming your answer." A structural model states its assumptions in a form a reader can dispute and can be tested against data it was not fitted to. The complaint that it assumes too much is a complaint about particular models, not about the practice.
  • "A randomised trial is the top of an evidence hierarchy." It is the best available answer to one question in one place. A well-designed trial in a setting unlike yours can be worse guidance than a careful observational study in your own, and Deaton and Cartwright's argument is precisely that ranking methods rather than arguments gets this backwards.

Looking back

Leamer showed in 1983 that a defensible change of controls could flip the sign of a published deterrence estimate, and proposed reporting bounds over the space of defensible specifications. The profession instead adopted research designs, and this course is the result: lotteries, borders, thresholds, timing and donor pools, each carrying an assumption stated in words. The objections are serious and distinct. Sims: the important macroeconomic questions have no design. Keane: a local effect does not deliver an untried counterfactual without a model. Deaton and Cartwright: randomisation balances in expectation, not in your trial, and a mean from one site does not travel on its own.

In short: what the revolution settled is that an empirical claim owes the reader a stated identifying assumption. What it left open is which questions can be asked at all, and whether the assumptions have merely moved somewhere quieter.

Sources

  1. Leamer, E. E. (1983). Let's take the con out of econometrics. American Economic Review, 73(1), 31-43.
  2. Angrist, J. D., & Pischke, J.-S. (2010). The credibility revolution in empirical economics: How better research design is taking the con out of econometrics. Journal of Economic Perspectives, 24(2), 3-30. aeaweb.org
  3. Leamer, E. E. (2010). Tantalus on the road to Asymptopia. Journal of Economic Perspectives, 24(2), 31-46. aeaweb.org
  4. Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science and Medicine, 210, 2-21. NBER Working Paper 22595. nber.org
  5. Brodeur, A., Cook, N., & Heyes, A. (2020). Methods matter: p-hacking and publication bias in causal analysis in economics. American Economic Review, 110(11), 3634-3660.
Key terms
Specification search
Fitting many models and reporting the one that supports the conclusion, leaving the reader unable to judge the evidence.
Extreme bounds analysis
Leamer's proposal to report the range of estimates over all specifications a reasonable reader would accept.
Credibility revolution
The shift toward research designs whose identifying assumption can be stated as a claim about the world and disputed.
Structural model
An explicit model of behaviour whose parameters can generate counterfactuals for policies never observed.
External validity
Whether an estimate obtained in one population, scale and implementation applies in another.
Randomisation inference
Testing by re-running the assignment mechanism itself rather than relying on an asymptotic standard error.
p-hacking
Adjusting analytic choices until a test statistic crosses a conventional threshold, visible as bunching just past that threshold.

Open the interactive version with quizzes and progress →