👥 Sociology · Undergraduate · SOC 350

Social Inequality & Stratification

Stratification is the study of how societies distribute income, wealth, status, health, and opportunity, and of the machinery that keeps those distributions stable across generations. This course starts with measurement, because almost every public argument about inequality is really an argument about a statistic: the Gini coefficient and the Lorenz curve it comes from, percentile ratios, top…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 15 lessons. The full lesson text is below so you can read it right here.

Module 1: Measuring Inequality

The instruments before the arguments: the Lorenz curve and the Gini coefficient worked by hand, percentile ratios and top shares, the very different shapes of the income and wealth distributions, and the long fight over where to draw a poverty line.

Gini, Lorenz, and What a Single Number Costs You

  • Build a Lorenz curve from raw incomes and compute a Gini coefficient by hand.
  • Explain why two societies with visibly different shapes can share almost the same Gini, and which measures separate them.
  • Choose the right measure for a given question among the Gini, percentile ratios, the Palma, and top shares.
  • State the four decisions hidden inside any published inequality statistic: which income, whose income, over what period, and before or after what.

Two numbers, one scale

The World Bank's most recent household survey estimate for South Africa puts its income Gini coefficient at 0.63, measured on 2014 data. That is the highest figure in its database. Slovakia, measured on the same scale with the same instrument, comes in around 0.24. The United States Census Bureau's Gini for household money income was 0.397 in 1967, and has run between roughly 0.48 and 0.49 since the mid-2010s.

Those numbers are doing an enormous amount of work in public argument, and almost nobody who cites them can say what they are. That is today's job. By the end of this lesson you will have built one from scratch, on paper, and you will know exactly what it throws away.

Start with the intuition. A Gini coefficient is a number between 0 and 1. Zero means every person or household gets exactly the same income. One means a single person gets everything and everyone else gets nothing. Real societies live between about 0.22 and about 0.65. Nothing in the definition tells you whether 0.45 is good or bad; that is a separate argument, and we will not confuse the two.

Building the curve, one villager at a time

Forget national statistics for twenty minutes. Here is a village of ten adults, with annual incomes in thousands, already sorted from poorest to richest:

10, 15, 20, 25, 30, 35, 45, 60, 90, 170

Total village income is 500. Now do the only calculation that matters: walk up the line from the poorest person and keep a running total of how much of the village's income you have accounted for.

People counted (poorest first)Share of populationIncome so farShare of total income
110%102%
220%255%
330%459%
440%7014%
550%10020%
660%13527%
770%18036%
880%24048%
990%33066%
10100%500100%

Plot the third column against the first, population share on the horizontal axis, income share on the vertical. That plotted line is the Lorenz curve, named after Max Lorenz, who proposed it in 1905 while still a graduate student. If everyone earned the same, the bottom 40 percent would hold 40 percent of the income and the curve would be the diagonal straight line from corner to corner. The further the curve sags below that diagonal, the more unequal the distribution.

Look at what the table already tells you. The bottom half of this village holds a fifth of its income. The top person alone holds 34 percent. You have not computed anything yet and you already know more than a headline would give you.

Key idea: The Lorenz curve is just a cumulative running total of income against a cumulative running total of people. Every inequality measure in this course is a way of summarising how far that curve sags below the diagonal.

Turning the curve into a number

The Gini coefficient is the area between the diagonal and the Lorenz curve, divided by the whole triangle under the diagonal. You do not have to measure areas by hand. For a set of n incomes sorted from smallest to largest, there is an exact formula:

G = (2 * sum of (rank i times income i)) / (n * total income) - (n + 1) / n

Run it on the village. Multiply each income by its rank and add: 1 times 10, plus 2 times 15, plus 3 times 20, and so on up to 10 times 170. That sum is 3,865. Then 2 times 3,865 is 7,730, divided by 10 times 500 is 5,000, giving 1.546. Subtract 11 divided by 10, which is 1.1. The village Gini is 0.446.

That is a real calculation on real arithmetic, and it lands within a few hundredths of the United States figure, which is worth sitting with. A ten-person village where the richest person earns seventeen times the poorest is roughly as unequal, by this measure, as the American household income distribution.

Now the problem

Here is a second village, also ten people, also 500 in total income:

14, 14, 14, 14, 14, 58, 58, 58, 128, 128

This is a different society. Five people are stuck at the bottom on identical incomes with nothing separating them. Three sit in a middle tier. Two are rich. There is no ladder; there are three landings with nothing in between. Run the formula. The rank-weighted sum is 3,860, so the Gini is 7,720 divided by 5,000, minus 1.1, which is 0.444.

The two villages differ by two thousandths. A Gini coefficient cannot tell them apart, and no amount of decimal places will help, because the information was destroyed when the curve was collapsed into an area.

The point: A Gini is a summary, and every summary is a loss of information. Distributions with completely different shapes can share a Gini, so a Gini alone cannot answer a question about shape.

Measures that see what the Gini misses

Three families of measures do the work the Gini cannot.

Percentile ratios compare two named points in the distribution and ignore everything else, which is exactly their virtue. The P90/P10 ratio divides the income at the 90th percentile by the income at the 10th. In our first village, taking the ninth and second people as stand-ins, that is 90 over 15, or 6.0. In the second village it is 128 over 14, or 9.1. The measure separates them immediately. The P50/P10 ratio, which asks how far the poor are below the median, is even sharper: 30 over 15 is 2.0 in the first village, and 14 over 14 is exactly 1.0 in the second, because there is no distance at all between the poor and the middle of the bottom tier.

Choose the ratio that matches your question. If you care about the distance between the poor and the middle, use P50/P10 and the top of the distribution becomes irrelevant to your answer. If you care about the pull-away at the top, use P90/P50. This is a feature. A measure that is sensitive to everything is sensitive to nothing in particular.

The Palma ratio, named after the economist Gabriel Palma, divides the income share of the richest 10 percent by the income share of the poorest 40 percent. Palma's observation was empirical: across countries, the share held by the middle five deciles is remarkably stable at around half of national income, so almost all the cross-national variation is a tug of war between the top tenth and the bottom four tenths. The Palma puts that contest directly in the numerator and denominator instead of averaging it away.

Top shares abandon summary statistics entirely and just report a fact: the percentage of national income going to the top 10 percent, the top 1 percent, or the top 0.1 percent. Thomas Piketty and Emmanuel Saez rebuilt these series from income tax records rather than surveys, which matters more than it sounds, because household surveys systematically miss the very top. Nobody samples enough billionaires, and the ones you do reach may decline to answer. On the tax-record series, the top 1 percent share of United States pre-tax income was around 10 percent in the late 1970s and has run near 19 percent in recent years. That is the single most cited number in the whole field.

What matters here: Percentile ratios answer questions about specific parts of the distribution, the Palma isolates the top-against-bottom contest, and top shares from tax data see a tail that surveys cannot reach. Report at least two of these together and you have told the truth; report only the Gini and you have not.

The four decisions hidden in every published figure

Before you compare two inequality statistics, check that they are the same statistic. Four choices sit underneath every one of them, and different agencies make them differently.

Which income? Market income means wages, self-employment earnings, and returns on capital. Gross income adds cash transfers such as pensions and unemployment benefit. Disposable income subtracts direct taxes. Some series then add the estimated value of in-kind benefits such as public health care. The United States market Gini and the United States disposable Gini differ by roughly a tenth, which is larger than most of the differences people argue about.

Whose income? Per household, per person, or per equivalised person? A household of four on 80,000 is not four times as well off as a person living alone on 20,000, because rent, heating, and a refrigerator are shared. Statistical agencies handle this with an equivalence scale. The OECD-modified scale gives the first adult a weight of 1, each additional adult 0.5, and each child 0.3, so a couple with two children needs 2.1 times a single person's income to reach the same standard of living. Change the scale and the measured inequality changes with it.

Over what period? An annual snapshot counts a medical resident earning 65,000 and a retired professor drawing down savings as if their positions were permanent. Lifetime inequality is always lower than annual inequality, because some of what a single year records is simply people being at different ages.

Before or after what? Census money income excludes capital gains and non-cash benefits. The Congressional Budget Office's series includes both and adjusts for taxes. Two honest agencies can therefore publish different Ginis for the same country in the same year without either being wrong.

Reading a cross-national table

With those cautions in place, here is roughly where rich countries sit on disposable household income, using OECD-comparable estimates from the late 2010s. Round numbers, because the exact figures move a point or two between data vintages.

CountryGini, disposable incomeRough character
Slovak Republicabout 0.24Among the lowest recorded anywhere
Denmarkabout 0.27Nordic pattern: compressed wages plus large transfers
Franceabout 0.29High market inequality, heavily reduced by taxes and transfers
Germanyabout 0.30Middle of the European range
United Kingdomabout 0.35High for Europe, well below the United States
United Statesabout 0.39Highest in the G7 on this measure
South Africaabout 0.63The highest in the World Bank series

Two things in that table repay attention. First, the gap between France and the United States on disposable income is far larger than the gap between them on market income, which is a preview of the welfare state lesson late in this course. Second, South Africa is not just an outlier, it is off the scale of European variation entirely, which is what a century of legally enforced racial stratification looks like when it is written into a single statistic.

Common misconceptions

  • A Gini of 0.45 means the top 45 percent hold everything. The Gini has no such reading. It is a ratio of areas, and its value has no direct percentage interpretation of that kind.
  • A rising Gini means the poor got poorer. It means the distribution spread out. Every income can rise while the Gini rises, if the top rises faster.
  • Two countries with the same Gini have the same inequality. Our two villages differ by two thousandths of a Gini and are structurally different societies. Always ask for a percentile ratio alongside.
  • Survey data captures the top accurately. It does not. Top incomes are undersampled and underreported, which is why tax-record series from Piketty and Saez show higher concentration than surveys do.
  • The Gini is a measure of poverty. It is a measure of dispersion. A country can reduce poverty and see its Gini rise, if incomes at the very top rise faster still.

Putting it together

  • The Lorenz curve plots cumulative income share against cumulative population share; the Gini is the area between that curve and the equality diagonal, expressed as a fraction of the triangle.
  • You can compute a Gini exactly from sorted incomes with the rank-weighted formula; our worked village gave 0.446, close to the United States household figure.
  • Two distributions with different shapes can share a Gini to within thousandths, so a single summary number cannot answer a question about structure.
  • Percentile ratios such as P90/P10 and P50/P10, the Palma, and top income shares each answer a different question, and top shares from tax records see a tail surveys miss.
  • Every published figure embeds four choices: which income concept, which unit and equivalence scale, what time period, and which taxes and transfers are counted. Compare like with like or do not compare.

Sources

  1. U.S. Census Bureau. (n.d.). Income inequality. census.gov
  2. World Bank. (n.d.). Gini index (World Bank estimate). World Development Indicators. data.worldbank.org
  3. Our World in Data. (n.d.). Economic inequality. ourworldindata.org
  4. World Inequality Lab. (n.d.). World Inequality Database. wid.world
  5. Gini, C. (1912). Variabilita e mutabilita: Contributo allo studio delle distribuzioni e delle relazioni statistiche. Bologna: Tipografia di Paolo Cuppini.
Key terms
Lorenz curve
A plot of cumulative income share against cumulative population share, sorted from poorest to richest; the diagonal represents perfect equality.
Gini coefficient
The area between the Lorenz curve and the equality diagonal as a fraction of the whole triangle, running from 0 at perfect equality to 1 at total concentration.
Percentile ratio
The income at one percentile divided by the income at another, such as P90/P10, which measures distance between two named points and ignores the rest.
Palma ratio
The income share of the top 10 percent divided by the share of the bottom 40 percent, built on the observation that the middle five deciles hold a stable share.
Top income share
The percentage of total income received by a top group such as the top 1 percent, usually estimated from tax records because surveys undersample the very rich.
Equivalence scale
A set of weights that converts household income into a per-person figure while allowing for shared costs; the OECD-modified scale uses 1, 0.5 for further adults, and 0.3 per child.
Market income
Income from wages, self-employment, and capital before any government transfer is added or any direct tax is subtracted.
Disposable income
Income after cash transfers have been added and direct taxes subtracted, which is the concept most cross-national inequality comparisons use.

Income and Wealth: Two Distributions That Do Not Match

  • Distinguish income as a flow from wealth as a stock, and explain why the two distributions have different shapes.
  • Read the 2022 Survey of Consumer Finances figures for median and mean family net worth and say what the gap between them reveals.
  • Describe how portfolio composition changes across the wealth distribution and why that changes who gains from asset booms.
  • Compare the survey, capitalisation, and estate-multiplier methods for measuring top wealth, and explain the disagreement they produce.

One dataset, two numbers, a factor of five

In the 2022 Survey of Consumer Finances, the Federal Reserve's triennial deep survey of American household balance sheets, the median family had a net worth of 192,900 dollars. The mean was 1,063,700 dollars. Same families, same year, same questionnaire.

Those two numbers are the whole lesson in miniature. A mean five and a half times the median can only happen if the distribution has a long, thin, extremely heavy tail. Move the richest few families out of the sample and the mean collapses; the median barely twitches. When a politician says the average American family has a million dollars, they are not lying about the arithmetic. They are using the statistic that a skewed distribution most rewards.

Income is skewed too, but far less. The median-to-mean gap for household income runs closer to a third than to a factor of five. Something is different about wealth, and this lesson is about what.

Stocks and flows

The distinction is the one physics teaches with a bathtub. Income is a flow: it is measured per unit of time, and the sentence is incomplete without one. Fifty thousand dollars a year. Wealth, or net worth, is a stock: everything you own minus everything you owe, measured at an instant. It needs a date, not a duration.

The two are connected but not the same, and the connection runs mostly one way. Saving out of income accumulates into wealth. Wealth then generates income in the form of interest, dividends, rents, and capital gains. But the accumulation is slow and the leaks are many, so a person's position in the income distribution predicts their position in the wealth distribution only loosely.

Two people illustrate the gap. A thirty-one-year-old physician earning 260,000 dollars with 310,000 dollars of student debt and no house has a strongly negative net worth and sits in the top few percent of the income distribution. A retired schoolteacher with a paid-off house worth 400,000 dollars, a pension, and 120,000 dollars in savings has an income near the national median and net worth in the top third. Any classification that treats income position as class position will place both of them wrongly.

Key idea: Income is a flow measured over a period and wealth is a stock measured at a date. Wealth accumulates out of income slowly and unevenly, so the two distributions have different shapes and rank people differently.

How much more concentrated is wealth?

Substantially, and consistently, in every country that measures it. Here are the American figures side by side, drawn from the Survey of Consumer Finances and the Federal Reserve's Distributional Financial Accounts, which push the survey results forward quarterly.

IncomeWealth
What it measuresFlow, per yearStock, at a date
Can it be negative?Rarely, for business lossesYes, and commonly: roughly one family in ten owes more than it owns
Approximate Gini, United StatesAbout 0.39 on disposable incomeAround 0.85
Top 1 percent shareNear 19 percent of pre-tax incomeClose to a third of household net worth
Bottom 50 percent shareRoughly a seventh of incomeAbout 2.5 percent of net worth
Best data sourceTax records and the CPSSurvey of Consumer Finances, estate tax files, capitalised tax data

Read the bottom-50-percent row twice. Half of American families hold, between them, something in the region of one fortieth of the country's household net worth. That is not a statement about who is poor; most of those families work and many own cars, phones, and modest savings. It is a statement about how little of the national balance sheet sits below the midpoint.

Why the shapes differ: four mechanisms

Wealth concentrates more than income for reasons that compound on each other.

Saving rates rise with income. A family spending everything it earns accumulates nothing regardless of how long it works. A family saving a fifth of a large income accumulates quickly. The gap in saving rates is not a moral fact about thrift; it follows from the arithmetic of fixed costs against income.

Returns compound. A dollar saved at twenty-five and left in a diversified portfolio for forty years is a very different object from a dollar saved at fifty-five. Compounding rewards early accumulation, which rewards families who did not need to spend their twenties repaying debt.

Inheritance transmits stocks directly. Income has to be earned each year by the person who receives it. Wealth can be handed over intact. Even modest transfers matter enormously at the moment they arrive, because a 20,000 dollar gift toward a deposit converts a renter into an owner, and ownership then compounds.

Wealth absorbs shocks, and its absence transmits them. A family with three months of expenses in the bank treats a broken transmission as an annoyance. A family without it treats the same event as a job loss, because the car was how they got to work. Michael Sherraden's argument in Assets and the Poor was precisely this: assets do work that income does not, because they change what a household can survive.

Why this matters: The concentration of wealth is not simply the concentration of income compounded. Differential saving, compounding, inheritance, and the buffering function of assets each add concentration on top of whatever the income distribution already contains.

What people actually own, by position

Portfolio composition changes so sharply across the distribution that the same asset boom makes very different things happen to different families.

For a family in the middle of the American wealth distribution, the balance sheet is dominated by two items: the equity in the home and a retirement account. Vehicles come third. For families in the top 1 percent, the dominant items are business equity and directly held corporate equity, with the primary residence a small fraction of the total.

The consequence is mechanical. When house prices rise sharply, middle wealth rises and the wealth distribution compresses slightly. When the stock market runs ahead of housing, wealth at the top rises much faster and concentration increases. The 2019 to 2022 period is a clean illustration: median family net worth rose by about 37 percent in real terms, the largest three-year rise the survey has recorded, driven largely by house prices and by pandemic-era transfers, while the very top also gained on equities. A single measure moving up can conceal quite different things happening at different points.

Debt composition differs too. Federal student loan balances in the United States total roughly 1.6 trillion dollars, and they sit disproportionately on households in their twenties and thirties who have not yet accumulated offsetting assets. That is why the age profile of net worth is so steep and why cross-sectional wealth statistics always mix genuine inequality with ordinary life-cycle position.

How anyone knows: three methods that disagree

Measuring the top of the wealth distribution is genuinely hard, and the honest state of the field is a live methodological dispute.

Surveys such as the SCF ask families directly. To reach the top, the SCF draws a supplementary sample from statistical tax records, deliberately oversampling wealthy households, then reweights. It still excludes the Forbes 400 by design, since including a handful of billionaires would make the estimates hostage to whether one of them happened to answer.

Capitalisation works backwards from tax returns. If a return reports 40,000 dollars of dividend income and the economy-wide dividend yield is 2 percent, you infer roughly 2 million dollars of underlying equity. Emmanuel Saez and Gabriel Zucman built their long American wealth series this way. The method's weakness is its assumption that everyone in a given asset class earns the same rate of return.

Estate multipliers use estate tax records: take the wealth of people who died in a year, and scale it up by the inverse of the mortality rate for their age and sex to estimate the wealth of the living.

These methods do not agree at the very top. Saez and Zucman's capitalisation estimates put the top 0.1 percent share of household wealth near 20 percent in recent years. Matthew Smith, Owen Zidar, and Eric Zwick, allowing for the fact that the rich hold more fixed-income assets at higher yields and that a great deal of top wealth is private business equity valued differently, estimate the same share nearer 15 percent. Both teams are serious, both published in leading journals, and the disagreement has not resolved.

Worth holding on to: The top of the wealth distribution is estimated, not observed. Different defensible methods put the top 0.1 percent share between roughly 15 and 20 percent, and a source that quotes one figure without noting the range is telling you less than it appears to.

Common misconceptions

  • High income and high wealth are the same thing. A high-earning doctor with student debt can have negative net worth while a median-income retiree with a paid-off house sits in the top third of wealth.
  • The average American family has about a million dollars. The mean is 1,063,700 dollars and the median is 192,900. On a heavily skewed distribution the mean describes almost nobody.
  • Negative net worth means destitution. It usually means a young household holding student or car debt before it has accumulated assets, which is why age composition matters when reading wealth statistics.
  • Rising asset prices benefit everyone proportionally. Portfolio composition differs by position, so a housing boom and an equity boom move the distribution in opposite directions.
  • Wealth statistics are precisely measured. The very top is inferred from tax data by methods whose assumptions differ, and the resulting estimates differ by several percentage points of the national total.

What to carry forward

  • Income is a flow and wealth is a stock; the two rank households differently, and neither substitutes for the other.
  • The 2022 SCF put median family net worth at 192,900 dollars against a mean of 1,063,700, a gap that is the signature of extreme right skew.
  • Wealth is far more concentrated than income everywhere it is measured: a United States wealth Gini near 0.85 against a disposable income Gini near 0.39.
  • Differential saving, compounding, inheritance, and the buffering role of assets each add concentration beyond what income inequality alone would produce.
  • Middle wealth is housing and retirement accounts; top wealth is business and corporate equity, so different asset booms move the distribution in different directions.
  • Top wealth shares are estimated by survey, capitalisation, and estate-multiplier methods that currently disagree by roughly five percentage points at the top 0.1 percent.

Sources

  1. Board of Governors of the Federal Reserve System. (2023). Changes in U.S. family finances from 2019 to 2022: Evidence from the Survey of Consumer Finances. federalreserve.gov
  2. Board of Governors of the Federal Reserve System. (n.d.). Distributional Financial Accounts. federalreserve.gov
  3. World Inequality Lab. (n.d.). World Inequality Database: United States. wid.world
  4. Our World in Data. (n.d.). Economic inequality. ourworldindata.org
  5. Sherraden, M. (1991). Assets and the poor: A new American welfare policy. Armonk, NY: M. E. Sharpe.
Key terms
Net worth
Everything a household owns minus everything it owes, measured at a point in time; it can be zero or negative.
Stock and flow
A stock is a quantity held at an instant, such as wealth; a flow is a quantity per unit of time, such as annual income.
Survey of Consumer Finances
The Federal Reserve's triennial survey of American household balance sheets, which oversamples wealthy households using statistical tax records.
Capitalisation method
Inferring wealth from tax-reported capital income by dividing it by an assumed rate of return for each asset class.
Estate multiplier method
Estimating the wealth of the living by scaling up the recorded wealth of people who died, using age and sex specific mortality rates.
Portfolio composition
The mix of assets a household holds; middle wealth is dominated by housing and retirement accounts, top wealth by business and corporate equity.
Right skew
A distribution with a long upper tail, in which the mean sits well above the median because extreme high values pull the average up.
Asset buffer
The capacity of savings to absorb an income or expense shock, which changes what events a household can survive without cascading loss.

Where to Draw the Line: Poverty Measurement and Its Disputes

  • Explain how Mollie Orshansky built the American official poverty threshold and which of her assumptions have never been updated.
  • Contrast the Official Poverty Measure with the Supplemental Poverty Measure on income concept, threshold construction, and geography.
  • Interpret the 2021 to 2022 change in the child Supplemental Poverty Measure and what it demonstrates about policy.
  • Compare absolute, relative, and consumption-based approaches to poverty, and state what each one can and cannot settle.

A food budget, multiplied by three

In 1963 and 1964 Mollie Orshansky, an economist at the Social Security Administration, needed a number. She took the Department of Agriculture's economy food plan, the cheapest of four diets the USDA costed out, and multiplied it by three. The multiplier came from a 1955 household survey finding that the average American family spent about a third of its after-tax income on food. Multiply the minimum food budget by three and you have, roughly, the minimum budget for everything.

That calculation is still the basis of the American official poverty line. It has been updated for inflation every year since, and in no other way. It does not vary by state or city. It uses pre-tax cash income, so it cannot see the Earned Income Tax Credit, SNAP, or a housing voucher. And the one-third assumption it rests on is long dead: American families now spend closer to an eighth of their budgets on food, so a multiplier of three understates the non-food budget substantially.

In 2022 the official threshold was a little under 30,000 dollars for a family of four, and the official poverty rate was 11.5 percent. This lesson is about why almost no researcher thinks that is the right number, and why they cannot agree on what should replace it.

The two American measures, side by side

Since 2011 the Census Bureau has published a second measure alongside the official one. The Supplemental Poverty Measure grew out of a 1995 National Academy of Sciences panel chaired around the work of Constance Citro and Robert Michael, which laid out what a modern threshold should include. Its design differs from Orshansky's at almost every joint.

Design choiceOfficial Poverty MeasureSupplemental Poverty Measure
Threshold basis1963 economy food plan times three, inflation-adjustedRecent household spending on food, clothing, shelter, and utilities, plus a small margin
Updated howBy the CPI onlyBy moving expenditure data, so it rises with living standards
GeographyOne national thresholdAdjusted for local housing costs
Housing tenureIgnoredSeparate thresholds for owners with a mortgage, owners without, and renters
Income countedPre-tax cash onlyPost-tax cash plus SNAP, housing subsidies, school meals, and refundable credits
Expenses subtractedNoneOut-of-pocket medical spending, work expenses, childcare, child support paid
Family unitRelated by blood, marriage, or adoptionWider: includes cohabiting partners and unrelated children in the household

Key idea: The official measure asks whether pre-tax cash income clears a 1963 food budget times three. The supplemental measure asks whether resources after taxes, transfers, and unavoidable expenses clear a threshold built from what households currently spend on necessities in that local housing market. These are different questions, and they have different answers.

The natural experiment of 2021

Two consecutive years show what the SPM can see and the OPM cannot.

In March 2021 the American Rescue Plan Act temporarily expanded the Child Tax Credit: it raised the amount, extended it to families with little or no earnings, and paid part of it monthly. On the official measure, none of this exists, because a refundable tax credit is not pre-tax cash income.

On the supplemental measure it is enormous. The child SPM poverty rate fell to 5.2 percent in 2021, the lowest ever recorded. The expansion then expired at the end of that year. In 2022 the child SPM rate was 12.4 percent, more than double the previous year. The overall SPM rate rose from 7.8 percent to 12.4 percent over the same twelve months.

Take the measurement lesson before the policy one. The official rate barely moved across those years. If you had only that series, the largest single-year change in American child poverty ever recorded would have been invisible to you. A measure that excludes the main instrument of anti-poverty policy cannot evaluate anti-poverty policy.

The policy reading is genuinely contested, and both halves of the argument are worth stating properly. The case for the expansion points to the measured halving of child poverty, the absence of any detectable large employment response in the studies conducted during 2021, and the evidence that families spent the money on food, rent, and school costs. The case against argues that the 2021 window was too short and too confounded by other pandemic supports to identify long-run labour supply effects, and that a permanent unconditional payment would eventually reduce work incentives in a way that a six-month emergency measure would not reveal. That objection is about a counterfactual, which is exactly the kind of claim a natural experiment of this length cannot settle.

Absolute against relative

Both American measures are, at bottom, absolute: a fixed basket, priced. Most of Europe uses a relative standard instead. Eurostat's at-risk-of-poverty threshold is 60 percent of national median equivalised disposable income, and the share of the EU population below it runs in the mid-teens. The OECD's standard relative line is 50 percent of the national median.

The two conceptions answer different questions, and each has an uncomfortable implication that its advocates must accept.

A relative line is really a measure of inequality in the lower half of the distribution. If every income in a country doubled overnight, relative poverty would not change at all. Critics say that is absurd. Defenders answer that poverty is partly about participation in a society, and that a household cannot function in a country where everyone else has internet access, a phone, and a bank account if it has none of them. Adam Smith made this argument in 1776 with linen shirts: a shirt is not a physical necessity, but in a country where a labourer would be ashamed to appear in public without one, its absence is a real deprivation.

An absolute line has the opposite problem. Hold it fixed for sixty years and it eventually describes a standard of living nobody recognises. The American official line is doing exactly this: it is a 1963 standard, carried forward on the CPI, in a country whose median income has roughly doubled in real terms since.

The point: Absolute lines measure whether a fixed basket is affordable and relative lines measure distance from the middle of one's own society. Neither is wrong; they answer different questions, and the choice of measure often determines the answer to the policy question before the data arrives.

The income-against-consumption dispute

A third camp argues both American measures are built on the wrong variable. Bruce Meyer and James Sullivan have argued for years that consumption, not income, is the better indicator of material living standards, on two grounds.

The first is measurement. When researchers link survey responses to administrative records, they find large underreporting of transfers: substantial fractions of SNAP and cash assistance dollars that agencies record as paid do not appear in survey answers. If income surveys miss transfer income, they overstate poverty and understate its decline.

The second is economic. Households smooth consumption across bad years by drawing on savings, borrowing, or help from relatives. A year of low income need not mean a year of low consumption, and the consumption measure captures the difference.

The counter-argument is that consumption is measured with its own large errors, that the Consumer Expenditure Survey underreports too, and that borrowing to maintain consumption is not the same as being out of poverty, since the debt remains. There is also a subtler point: a household that maintains consumption by exhausting its savings has become more fragile, and no consumption measure records that.

Notice what kind of dispute this is. It is not a disagreement about values. It is a disagreement about which measurement error is larger, and it is in principle resolvable by better linked data.

Hardship measures, which sidestep the thresholds

A different strategy abandons dollar thresholds and asks directly about outcomes. The USDA's household food security survey asks whether, in the past year, the household ran short of money for food, cut meal sizes, or went without eating. It reported 13.5 percent of American households as food insecure at some point during 2023.

Hardship indicators of this kind have a real advantage: they measure the thing everyone actually cares about rather than a proxy for it, and they do not require anyone to agree on a threshold. Their disadvantage is that they are hard to compare over long stretches, because the questions change and because what counts as hardship shifts with expectations.

The global line

The World Bank maintains an international extreme poverty line, converted between countries using purchasing power parity exchange rates rather than market rates. It was 1.90 dollars a day at 2011 prices, revised to 2.15 dollars at 2017 prices in 2022, and revised again in 2025 to 3.00 dollars at 2021 prices. Each revision changes the headline count without anything changing in anyone's life, which is a useful reminder of what these numbers are.

The trend, on any of these lines, is one of the largest social facts of the last forty years. The share of the world's population below the international extreme poverty line fell from around 38 percent in 1990 to under a tenth by 2019, driven above all by China and later India.

Critics of the line, including Sanjay Reddy and Thomas Pogge, argue that PPP conversions are built from consumption baskets weighted toward things poor people do not buy, so the line does not represent a consistent standard of living across countries, and that a line this low describes survival rather than adequacy. Defenders reply that the line was never meant to define adequacy, only to track the very bottom consistently, and that the trend is robust to reasonable alternative lines even if the level is not.

Common misconceptions

  • The official poverty line reflects what it costs to live today. It reflects a 1963 food budget multiplied by three and adjusted for inflation, with no geographic variation and no account of taxes, transfers, or medical costs.
  • The Supplemental Poverty Measure is always higher than the official one. It was lower in 2021, when refundable credits it counts and the official measure ignores pushed millions above the threshold.
  • Relative poverty measures are just badly designed absolute ones. They are answering a different question, about distance from the social middle, and they behave differently by design.
  • Survey income data accurately records benefits received. Linked administrative records show substantial underreporting of SNAP and cash assistance, which biases income-based poverty upward.
  • Global poverty counts are direct observations. They depend on purchasing power parity conversions and on a line that has been redrawn three times since 2015, each time changing the count.

Where this leaves us

  • The American official measure is Orshansky's 1963 economy food plan times three, inflation-adjusted, pre-tax, and national; its central assumption about food shares is decades out of date.
  • The Supplemental Poverty Measure uses current spending on necessities, adjusts for local housing costs and tenure, counts taxes and non-cash benefits, and subtracts medical and work expenses.
  • Child SPM poverty fell to 5.2 percent in 2021 under the expanded Child Tax Credit and rose to 12.4 percent in 2022 after it expired, a change the official measure could not see.
  • Absolute lines price a fixed basket; relative lines such as the EU 60 percent of median threshold measure distance from the social middle; the choice frequently decides the policy answer.
  • Meyer and Sullivan argue consumption is the better variable because income surveys underreport transfers and because households smooth spending, and the counter-argument is that consumption data has its own errors and that borrowing is not escape.
  • The World Bank international line has moved from 1.90 to 2.15 to 3.00 dollars a day as price bases were updated, while the long trend of falling extreme poverty holds under any of them.

Sources

  1. U.S. Census Bureau. (n.d.). Poverty. census.gov
  2. U.S. Census Bureau. (n.d.). Supplemental Poverty Measure. census.gov
  3. World Bank. (n.d.). Poverty. worldbank.org
  4. Our World in Data. (n.d.). Poverty. ourworldindata.org
  5. Citro, C. F., and Michael, R. T. (Eds.). (1995). Measuring poverty: A new approach. Washington, DC: National Academy Press.
Key terms
Orshansky threshold
The 1963 poverty line built by multiplying the USDA economy food plan by three, which remains the basis of the American official measure.
Official Poverty Measure
The United States poverty statistic based on pre-tax cash income against a national inflation-adjusted threshold, with no geographic or tenure adjustment.
Supplemental Poverty Measure
A Census measure using post-tax resources including non-cash benefits, thresholds from current spending on necessities, and adjustments for local housing costs and tenure.
Absolute poverty line
A threshold defined by the cost of a fixed basket of goods, which does not move when general living standards rise.
Relative poverty line
A threshold defined as a fraction of the national median income, such as the EU at-risk-of-poverty line at 60 percent of the median.
Purchasing power parity
An exchange rate constructed so that a common basket of goods costs the same across countries, used to convert international poverty lines.
Consumption poverty
Poverty measured by what a household spends rather than what it receives, on the argument that spending better reflects material living standards.
Food insecurity
A hardship measure recording whether a household lacked money for adequate food during the year, reported by the USDA rather than derived from a dollar threshold.

Module 2: What Class Means

The concept that organises the field, argued over rather than defined: Marx against Weber and the traditions that grew from each, the schemas survey research actually uses, and then the occupational structure and the labour market institutions that produce the distribution in the first place.

Marx, Weber, and the Argument About What a Class Is

  • State Marx's definition of class and the specific predictions it generates, and identify which of them the twentieth century did not confirm.
  • Explain Weber's three-dimensional scheme of class, status, and party, and use it on cases where the three come apart.
  • Compare the schemas of Wright, Goldthorpe, and Bourdieu on what each one takes to be the underlying mechanism.
  • Evaluate the Great British Class Survey's seven-class model and the methodological objections raised against it.

Seven classes, announced on a Wednesday

In April 2013 Mike Savage and a team of sociologists published an analysis of the Great British Class Survey in the journal Sociology. The BBC web survey behind it had drawn 161,400 responses, and the paper argued that Britain no longer had three classes but seven: an elite, an established middle class, a technical middle class, new affluent workers, a traditional working class, emergent service workers, and a precariat. The BBC put a class calculator online. Several million people used it.

Two things about that episode are worth holding. The first is that a British national broadcaster thought class was interesting enough to build a quiz around, which tells you the concept has not gone anywhere. The second is that within a year the paper had drawn serious methodological objections, principally that a self-selected web sample skews enormously toward the highly educated and that the seven groups fell out of the researchers' choice of clustering method rather than out of the social world.

That is the shape of the whole field. Everyone agrees class matters. Nobody agrees on what it is. This lesson lays out the four or five serious answers, and shows what each one predicts that the others do not.

Marx: class is a relationship, not a rank

Start with the answer that everything else is a response to.

For Marx, class is not a position on a ladder of income or prestige. It is a relationship to the means of production: to the factories, land, machines, and capital with which things are made. In industrial capitalism this generates two principal classes. The bourgeoisie owns the means of production and buys labour power. The proletariat owns nothing but its labour power and must sell it to live. Between and beside them sit the petty bourgeoisie, who own means of production but work them themselves, such as a shopkeeper or an independent tradesman.

The mechanism binding them is exploitation, which in Marx is a technical term rather than an insult. A worker paid enough to reproduce their labour power produces more value in a day than that wage represents. The difference, surplus value, is appropriated by the owner. Profit, on this account, is not a reward for risk or organisation but a transfer.

Marx also distinguished a class in itself from a class for itself: a group that objectively shares a position in production, against a group that recognises the shared position and acts on it. The gap between the two is where the whole subsequent Marxist literature on consciousness, ideology, and false consciousness lives.

Now the honest scorecard. Marx predicted polarisation into two camps, the disappearance of intermediate strata, deepening immiseration of workers, and the eventual emergence of revolutionary class consciousness in the most industrialised countries. What happened instead was the growth of an enormous salaried middle stratum, real wage gains across the twentieth century in the industrial democracies, and revolutions that occurred mainly in agrarian societies rather than advanced industrial ones. Those are failed predictions, and they are why the twentieth-century Marxist literature is largely an attempt to repair the framework rather than restate it.

What survived is substantial: the insistence that class is about relations of production rather than consumption or taste, the analysis of how ownership converts into power inside the workplace, and the observation that the labour market is not a meeting of equals.

Key idea: Marx defines class by relationship to the means of production and explains inequality through the extraction of surplus value. His descriptive apparatus outlived his predictions, which polarisation, immiseration, and revolutionary consciousness in industrial societies did not confirm.

Weber: three dimensions, not one

Max Weber's answer, written a generation later, accepts that economic position matters and denies that it is the only thing that does. He separates three orders that can vary independently.

Class for Weber is market situation: what you bring to markets and what it commands. Crucially this includes not only property but marketable skill and credentials, which lets him distinguish an unskilled labourer, a skilled engineer, and a rentier without collapsing them.

Status, Stand in the German, is social honour: the esteem attached to a way of life, expressed through consumption, manners, education, and who may marry whom. Status groups are communities in a way that classes need not be.

Party is organised power: the capacity to act collectively to influence decisions, through parties, unions, professional associations, or lobbies.

The point of the scheme is the cases where the dimensions separate. An impoverished aristocrat has high status and little class position. A newly rich businessman may command enormous market power and be excluded from institutions his employees' children attend. A union with modest members can hold serious party power. A professor of classics has high status, moderate class position, and negligible party power. Weber's framework describes all four without strain, and Marx's cannot.

Weber also gives us life chances, the probability of obtaining the things a society values, which remains the working definition of what stratification research measures. That is his most quietly influential contribution.

What came after: four repairs

Four twentieth-century programmes tried to make class operational.

Erik Olin Wright asked the question that most embarrasses a two-class model: where does a middle manager go? She does not own the firm, so she is not bourgeois. She exercises authority over others and is paid far above a production worker, so calling her proletarian empties the word. Wright's answer was that such positions occupy contradictory class locations, sharing features of more than one class. His mature scheme treated three assets as generating class position: ownership of means of production, organisational authority, and scarce skills or credentials. This preserves the Marxist emphasis on exploitation while accounting for a salaried middle that Marx did not expect.

John Goldthorpe built the schema that European official statistics actually use, based not on prestige or income but on the employment relationship. Some workers hold a labour contract: effort is monitored, pay tracks hours or output, and the relationship is short-horizon. Others hold what Goldthorpe called a service relationship: they are hard to monitor, are trusted with discretion, and are compensated with salary increments, career prospects, and job security. The distinction is not about how nice the job is; it is about the employer's monitoring problem. The resulting Erikson-Goldthorpe-Portocarero classes underpin the British NS-SEC and the European ESeC.

Pierre Bourdieu argued that economic capital is only one of the currencies. Cultural capital is competence in the legitimate culture, embodied in knowing which fork, which composer, which register of speech, and objectified in credentials. Social capital is the resources available through networks. These convert into one another, imperfectly, and the conversion rates are the real subject. His concept of habitus, a set of durable dispositions instilled by early environment, explains why people from different origins make systematically different choices while experiencing those choices as free preferences.

The measurement tradition mostly ignored theory and asked people to rank occupations. Prestige studies from the 1940s onward found rankings that are extraordinarily stable across decades and across countries: physicians and judges near the top, refuse collectors and shoe shiners near the bottom, with very high correlations between countries as different as Japan and Brazil. Duncan's socioeconomic index and later the international ISEI turned those rankings into continuous scores that regression can use.

TraditionWhat defines classCore mechanismHow it is measured
MarxRelation to means of productionExploitation, extraction of surplus valueOwnership and employment status
WeberMarket situation, plus separate status and party ordersLife chances set by what one brings to marketsProperty, skill, credentials, honour, organisation
WrightOwnership, authority, and credentials togetherExploitation with contradictory locationsTwelve-category schema from survey items
GoldthorpeType of employment contractThe employer's monitoring problemEGP classes, operationalised as NS-SEC and ESeC
BourdieuVolume and composition of economic, cultural, and social capitalConversion between capitals, reproduced through habitusCorrespondence analysis of tastes, credentials, networks

What matters here: These are not five names for one idea. Each locates the engine of class somewhere different: in production for Marx, in markets for Weber, in authority and credentials for Wright, in the contract for Goldthorpe, and in transmissible cultural competence for Bourdieu. Which one you adopt determines what you will measure and therefore what you will find.

Back to the seven classes

The Great British Class Survey was Bourdieu operationalised. Savage and colleagues measured economic capital, cultural capital, and social capital separately, then clustered respondents. That is why their classes are not a ladder: the technical middle class is high on economic capital and thin on social connections, while emergent service workers are low on economic capital and high on emerging cultural capital, meaning contemporary rather than highbrow taste.

The objections were serious. The web sample of 161,400 was self-selected and heavily overrepresented the educated and affluent, so the researchers weighted it against a much smaller nationally representative survey of about a thousand respondents. Critics including Colin Mills argued that the number of classes fell out of a modelling choice rather than a discovery, that a different clustering specification would produce a different number, and that the resulting categories predict outcomes less well than the older employment-relations schema they were meant to supersede.

The honest verdict is mixed rather than dismissive. The survey demonstrated that cultural and social resources are measurably distinct from income and that people sort on them. It did not demonstrate that seven is the right number of classes, and no clustering exercise could.

Does class still explain anything?

In 1996 Jan Pakulski and Malcolm Waters published The Death of Class, arguing that class had been displaced by identities based on consumption, culture, and lifestyle. Others answered with data. Goldthorpe and Wright's response, in different idioms, was the same: class origin continues to predict educational attainment, occupational destination, health, and mortality with effect sizes that have not collapsed, whatever people say about their own identity.

Subjective identification is genuinely strange, and worth knowing before you interpret survey answers. In American polls a large majority place themselves in the middle class or upper middle class, including many people in the bottom quarter of the income distribution and many in the top few percent. Self-report and structural position are different variables. When a study says class predicts an outcome, always check which of the two it measured.

Common misconceptions

  • Class means income bracket. Not in any of the major traditions. Marx defines it by ownership, Weber by market situation, Goldthorpe by contract type, Bourdieu by the composition of capitals. Income is a consequence, not the definition.
  • Exploitation in Marx is a moral accusation. It is a technical claim about the appropriation of surplus value. You can accept the mechanism and reject the moral conclusion, or the reverse.
  • Weber refuted Marx. Weber added dimensions rather than deleting one. Market situation remains central to his account; he denied it was the only thing structuring life chances.
  • Prestige rankings are culturally arbitrary. They are strikingly stable across time and across very different societies, which is one of the more robust findings in the field.
  • If most people call themselves middle class, class has stopped operating. Subjective identification and structural position are separate variables, and the structural one keeps predicting mortality, schooling, and mobility.

Pulling it together

  • Marx defines class by relation to the means of production, with exploitation as the mechanism; his descriptive apparatus survived, his predictions of polarisation and revolutionary consciousness in industrial societies did not.
  • Weber separates class, status, and party, which lets him describe the impoverished aristocrat and the excluded new money that a one-dimensional scheme cannot.
  • Wright's contradictory class locations add authority and credentials to ownership, in order to place the salaried manager.
  • Goldthorpe builds class from the employment contract and the employer's monitoring problem, and his schema is what European official statistics use.
  • Bourdieu adds cultural and social capital and the conversion between them, reproduced through habitus.
  • The Great British Class Survey operationalised Bourdieu on 161,400 self-selected web responses and produced seven classes; the finding that cultural and social resources are distinct survived the criticism, the specific number of classes did not.

Sources

  1. Savage, M., Devine, F., Cunningham, N., Taylor, M., Li, Y., Hjellbrekke, J., Le Roux, B., Friedman, S., and Miles, A. (2013). A new model of social class? Findings from the BBC's Great British Class Survey experiment. Sociology, 47(2), 219-250. doi.org
  2. Britannica. (2025). Social class. In Encyclopaedia Britannica. britannica.com
  3. Britannica. (2025). Max Weber. In Encyclopaedia Britannica. britannica.com
  4. Bourdieu, P. (1984). Distinction: A social critique of the judgement of taste (R. Nice, Trans.). Cambridge, MA: Harvard University Press. (Original work published 1979.)
  5. Wright, E. O. (1997). Class counts: Comparative studies in class analysis. Cambridge: Cambridge University Press.
Key terms
Means of production
The factories, land, machinery, and capital used to produce goods; in Marx, ownership of them defines class position.
Surplus value
The difference between the value a worker produces and the wage paid, appropriated by the owner; the technical core of Marx's exploitation.
Class for itself
A group that recognises its shared position in production and acts collectively on it, as distinct from a class in itself that merely occupies that position.
Market situation
Weber's term for what a person brings to markets, including property, skill, and credentials, and what it commands in exchange.
Status group
A community sharing social honour and a way of life, which may be aligned with or independent of economic class position.
Contradictory class location
Wright's term for positions such as middle management that share features of more than one class because they combine authority without ownership.
Service relationship
Goldthorpe's contract type in which discretion, salary increments, and career prospects substitute for close monitoring of effort.
Cultural capital
Competence in the legitimate culture, held as embodied disposition, as objects, and as credentials, and convertible into economic advantage.
Habitus
Bourdieu's term for durable dispositions instilled by early social environment, which produce systematically patterned choices experienced as free preferences.
Life chances
Weber's term for the probability of obtaining valued goods and outcomes, which remains the working object of stratification research.

The Occupational Structure and the Hollowing of the Middle

  • Describe the long shift in the American occupational structure and give the employment figures behind it.
  • Explain job polarisation and the routine-task account that predicts it, and identify what that account leaves unexplained.
  • Assess how unions, minimum wages, and employer wage-setting power shape the wage distribution, with the measured magnitudes.
  • Explain why a growing share of earnings inequality now sits between firms rather than within them.

Two employment series, forty years apart

United States manufacturing employment peaked in mid-1979 at roughly 19.5 million jobs. By 2019 it was about 12.8 million. Over the same period, employment in health care and social assistance rose from under 7 million to more than 20 million. Those two lines crossing is the single most important thing that happened to the American occupational structure in living memory, and almost everything else in this lesson follows from it.

The move is not simply from making things to serving people. It is a move between kinds of employment relationship, kinds of wage-setting institution, and kinds of career. A 1975 assembly job at a unionised plant and a 2025 home health aide job pay differently, are governed differently, and offer different odds of a raise. Understanding stratification means understanding what people actually do all day and what determines what that work pays.

What Americans do

The Bureau of Labor Statistics counts employment and wages occupation by occupation in its Occupational Employment and Wage Statistics programme. The largest single occupations in the United States are not what most people guess. Retail salespersons, fast food and counter workers, cashiers, home health and personal care aides, and registered nurses each employ in the region of three to four million people. Software developers, by comparison, number well under two million.

Two consequences follow immediately. First, when someone describes the modern economy as knowledge work, they are describing a real and well-paid sector that employs a minority. Second, the fastest-growing large occupation in the United States is personal and home care aide, which is also among the lowest paid. Growth and pay are not the same axis, and a lot of confused public argument comes from assuming they move together.

Key idea: The occupational structure is dominated by service occupations employing millions at low wages, not by the professional and technical work that dominates public discussion of the economy. Any account of inequality has to start from what most people actually do.

Polarisation: growth at both ends, hollowing in the middle

The most robust empirical finding about rich-country labour markets over the past forty years is that employment growth has been U-shaped across the wage distribution. Maarten Goos and Alan Manning named it in a 2007 paper on Britain, with the phrase lousy and lovely jobs: strong growth in high-paid professional and managerial work, strong growth in low-paid personal service work, and decline in the middle. The pattern has since been documented across most of western Europe and the United States.

The leading explanation is the routine-task model developed by David Autor, Frank Levy, and Richard Murnane. Their argument is that computers do not substitute for workers, they substitute for tasks, and specifically for tasks that can be written down as explicit rules. That description fits an enormous share of the mid-wage jobs that used to anchor the middle of the distribution: bookkeeping, filing, telephone switching, machine operation to a fixed specification, routine assembly.

It does not fit the jobs above or below. Above, non-routine analytic and interpersonal work such as diagnosis, litigation, engineering design, and management is complemented by computing rather than replaced by it. Below, non-routine manual work such as cleaning a hotel room, cutting hair, or helping a frail adult out of a bath requires physical dexterity and situational judgement that has proved much harder to automate than office arithmetic.

So the middle hollows. Displaced mid-skill workers compete for jobs at the bottom, which pushes wages there down further even as employment there grows.

What the model does not explain is worth naming, because a theory that explains everything explains nothing. It does not explain why polarisation arrived at different times in different countries with similar technology, and it does not explain why the wage effects were so much larger in the United States than in, say, Germany, which adopted the same machines. That gap is where institutions enter.

Institutions: unions, floors, and employer power

Technology is a shock. Institutions determine who absorbs it.

Unions. The Current Population Survey has measured union membership on a comparable basis since 1983, when 20.1 percent of American wage and salary workers were members. The rate has fallen to around 10 percent, and in the private sector to about 6 percent. That decline matters for the distribution in three separate ways. Unions raise members' wages, with the union wage premium usually estimated in the range of 10 to 20 percent. They compress wages within the workplace, because collective agreements narrow differentials between grades. And they exert a threat effect on non-union employers in the same labour market who raise wages to avoid organising drives. Work by Henry Farber, Daniel Herbst, Ilyana Kuziemko, and Suresh Naidu, assembling union household data back to 1936, found that union households have consistently had higher incomes and that union density and inequality track each other inversely across the century.

Minimum wages. The United States federal minimum wage has been 7.25 dollars an hour since July 2009, the longest period without an increase since the standard was created in 1938. Many states and cities set higher floors, which is why the federal figure now binds in a shrinking part of the country. The economics here changed in the 1990s. The textbook model said a binding minimum wage must reduce employment. David Card and Alan Krueger's 1994 comparison of fast food restaurants in New Jersey, which raised its minimum, and neighbouring Pennsylvania, which did not, found no employment loss. Decades of subsequent work have produced a range of estimates rather than a consensus, but the modern position is that moderate increases from low levels produce small or negligible employment effects while raising earnings at the bottom, and that very large increases relative to local median wages are a genuinely open question.

Monopsony. The reason a minimum wage can raise pay without cutting jobs is the concept that has reorganised labour economics in the past decade. In a competitive labour market, an employer faces a going wage and can hire as many workers as it wants at that wage. In reality employers have wage-setting power, because workers face search costs, geographic constraints, non-compete clauses, and a small number of local employers in their occupation. Under employer wage-setting power, pay sits below the value of what a worker produces, and a wage floor can raise pay and employment simultaneously.

The upshot: The same technological shock produces different wage distributions in different countries because unions, wage floors, and the degree of employer wage-setting power determine how the shock is shared. Institutions are not a footnote to the technology story; they are half of it.

Where inequality moved: between firms, not within them

One finding has quietly reorganised how researchers think about earnings inequality. Jae Song, David Price, Fatih Guvenen, Nicholas Bloom, and Till von Wachter analysed United States earnings records covering millions of workers from 1978 to 2013 and asked a simple question: is the growing gap between high and low earners happening inside workplaces, or between them?

The answer was overwhelmingly between them. The dispersion of average pay across firms grew enormously, while inequality within a given firm grew much less. High earners increasingly work alongside other high earners, and low earners alongside other low earners.

David Weil's account of the fissured workplace supplies the mechanism. Over three decades, large firms shed functions that were once in-house. The cleaners, security guards, cafeteria staff, and call centre workers who used to be employees of a profitable corporation, sharing its pay scales and its benefits, became employees of contractors competing on price. Nothing about the work changed. The employment relationship did, and with it the link between a firm's profitability and what the person cleaning its floors is paid.

This also explains a puzzle in the data. Studies of pay within occupations often find modest inequality growth, while economy-wide inequality grew sharply. Both are true, because a large part of the growth happened by moving jobs across organisational boundaries rather than by changing pay within a given workplace.

Segments, ladders, and credentials

Two older ideas remain useful for reading all of this.

Peter Doeringer and Michael Piore's 1971 account of internal labour markets described firms in which entry occurs at a small number of ports, and higher positions are filled by promotion from within under administrative rules rather than by market competition. Where such ladders exist, a first job is a career; where they have been dismantled, it is a job. Dual labour market theory extended this into a distinction between a primary segment with stability, training, and progression, and a secondary segment with high turnover and no ladder. Mobility between the two is the thing to watch.

Credentialism is the other. Randall Collins argued in The Credential Society that educational requirements often rise faster than the skill content of jobs, functioning as a sorting and exclusion device rather than a productivity measure. The test is straightforward and worth applying: when a job that required a high school diploma in 1990 requires a bachelor's degree in 2025, ask whether the work changed. Sometimes it did. Often the applicant pool changed, and the employer raised the bar because it could.

Common misconceptions

  • The middle disappeared because factories closed. Manufacturing decline is part of it, but polarisation also hollowed mid-wage clerical and administrative work, which no factory closure touches.
  • Automation destroys jobs. The routine-task account says it reallocates them, destroying mid-wage routine work while complementing analytic work and leaving non-routine manual work in place.
  • Unions only help their own members. Threat effects raise non-union wages in the same labour market, and collective agreements compress differentials, so the distributional effect is wider than the membership.
  • A minimum wage must reduce employment. That follows only under a competitive labour market. Under employer wage-setting power a floor can raise pay and employment together, and the empirical record from moderate increases shows small effects.
  • Growing inequality means bosses raised their own pay inside firms. Most of the measured growth in United States earnings dispersion since 1978 occurred between firms rather than within them.

The short version

  • United States manufacturing employment fell from about 19.5 million in 1979 to about 12.8 million in 2019 while health care and social assistance rose from under 7 million to over 20 million.
  • The largest occupations are retail, food service, and care work, each employing millions at low wages; the fastest-growing large occupation is also among the lowest paid.
  • Employment growth has been U-shaped, with the routine-task model explaining the hollowing of mid-wage work that could be written down as explicit rules.
  • The same technology produced different outcomes in different countries because unions, minimum wages, and employer wage-setting power determine who absorbs a shock.
  • Union membership fell from 20.1 percent of wage and salary workers in 1983 to around 10 percent, with a union wage premium usually estimated at 10 to 20 percent plus compression and threat effects.
  • Most of the growth in American earnings dispersion since 1978 occurred between firms, consistent with Weil's account of functions being contracted out of high-paying organisations.

Sources

  1. U.S. Bureau of Labor Statistics. (n.d.). Occupational Employment and Wage Statistics. bls.gov
  2. U.S. Bureau of Labor Statistics. (n.d.). Labor Force Statistics from the Current Population Survey. bls.gov
  3. U.S. Department of Labor. (n.d.). Minimum wage. dol.gov
  4. Goos, M., and Manning, A. (2007). Lousy and lovely jobs: The rising polarization of work in Britain. Review of Economics and Statistics, 89(1), 118-133. doi.org
  5. Weil, D. (2014). The fissured workplace: Why work became so bad for so many and what can be done to improve it. Cambridge, MA: Harvard University Press.
  6. Doeringer, P. B., and Piore, M. J. (1971). Internal labor markets and manpower analysis. Lexington, MA: Heath.
Key terms
Job polarisation
Employment growth concentrated at the top and bottom of the wage distribution with decline in the middle, documented across rich countries since the 1980s.
Routine task
Work that can be specified as explicit rules, which is therefore substitutable by computing; the central category in the Autor, Levy, and Murnane account.
Union wage premium
The pay advantage of union members over comparable non-members, usually estimated in the range of 10 to 20 percent.
Threat effect
The wage increase non-union employers grant to reduce the likelihood that their workforce organises, extending union influence beyond membership.
Monopsony
Employer wage-setting power arising from search costs, geography, and limited local competition, under which pay sits below the value of what a worker produces.
Fissured workplace
Weil's term for the contracting out of functions once performed in-house, which severs the link between a firm's profitability and the pay of the people doing its work.
Internal labour market
A firm in which entry occurs at limited ports and higher positions are filled by promotion under administrative rules rather than open competition.
Credentialism
The tendency of educational requirements to rise faster than the skill content of jobs, functioning as a sorting device rather than a measure of productivity.

Module 3: Mobility and Education

Whether the ladder moves people, and what school does about it: intergenerational elasticities and rank-rank slopes read correctly, the Great Gatsby Curve and what it can support, and education examined as both an escalator and a sorting device.

What an Intergenerational Elasticity of 0.4 Does Not Mean

  • Define the intergenerational elasticity and the rank-rank slope, and state precisely what each coefficient reports.
  • Diagnose the two most common misreadings of a mobility coefficient and explain exactly where the reasoning fails.
  • Distinguish absolute from relative mobility and interpret the decline in the share of American children out-earning their parents.
  • Evaluate the Great Gatsby Curve and the evidence on place effects, saying what each design can and cannot establish.

Two cities, one country

In 2014 Raj Chetty, Nathaniel Hendren, Patrick Kline, and Emmanuel Saez published estimates built from around 40 million United States tax records linking parents to children. A child raised by parents at the 25th percentile of the income distribution had, on average, a 7.5 percent chance of reaching the top fifth as an adult. But the national average concealed enormous variation. In the San Jose commuting zone the figure was 12.9 percent. In Charlotte it was 4.4 percent. Same country, same decade, a threefold difference.

This lesson is built around a mistake, because the mistake is more instructive than the correct answer. The mistake is what nearly everyone does with the single most quoted number in mobility research: the intergenerational elasticity.

The number, and the wrong reading

The intergenerational elasticity of income, usually written IGE, comes from a regression of the logarithm of a child's adult income on the logarithm of their parents' income. For the United States, estimates typically land between 0.4 and 0.6. For Denmark they run near 0.15, for Canada near 0.2, for the United Kingdom near 0.5.

Here is the reading you will encounter constantly, in newspapers and in student essays: an IGE of 0.4 means that 40 percent of a person's income is determined by their parents, and 60 percent by their own efforts.

That is wrong, and tracing exactly where it goes wrong teaches you most of what you need to know about the literature.

The first failure is the confusion of an elasticity with a share. An elasticity relates proportional changes. An IGE of 0.4 says that comparing two families whose incomes differ by 10 percent, the children's incomes are expected to differ by about 4 percent. It says nothing at all about how much of the variation in child income is accounted for by parental income. The quantity that does say that is the R-squared of the regression, and it is typically far lower than the coefficient, because parental income leaves the great majority of the variance in child outcomes unexplained.

The second failure is the assumption that whatever is not parents must be effort. The residual contains luck, health, the local labour market, the state of the business cycle in the year you graduated, discrimination, measurement error in both incomes, and effort, in no known proportion. Calling the residual merit is a value judgement wearing a statistic's clothes.

The third failure is subtler and matters for reading the literature. The IGE is very sensitive to how parental income is measured. Use a single year of parental income and you have measured permanent income with error, and classical measurement error in a regressor pulls the coefficient toward zero. Gary Solon and, separately, Bhashkar Mazumder showed that averaging parental income over many years raises United States estimates substantially, with Mazumder's multi-year estimates reaching around 0.6. A large part of the apparent disagreement in this literature is not disagreement about mobility but about the number of years averaged.

The upshot: An IGE is an elasticity, not a share of variance and not a share of destiny. It also moves with the number of years of parental income you average over, so two studies reporting 0.35 and 0.6 for the same country may both be right about different quantities.

The rank-rank slope, and why Chetty prefers it

Log-income regressions have practical problems. What do you do with a zero income, whose logarithm is undefined? What do you do with a hedge fund manager whose income makes the estimate hostage to a handful of observations?

The rank-based alternative sidesteps both. Rank every parent in their cohort from 0 to 100, rank every child in theirs, and regress child rank on parent rank. The coefficient is the rank-rank slope. For the United States it is about 0.34, which reads cleanly: children whose parents were 10 percentiles apart end up, on average, about 3.4 percentiles apart.

The rank measure is not better in every way. It is by construction relative, so if everyone's real income rises it registers nothing, and it cannot tell you anything about how far apart the rungs are. A society could have perfect rank mobility and enormous gaps between the top and bottom rungs, and the coefficient would report zero problem. That is the difference between mobility and inequality, and it is the reason this course treats them as separate lessons.

Absolute mobility: the question people actually mean

When someone asks whether their children will do better than they did, they are not asking about ranks. They are asking about absolute mobility: will my child, in real terms, earn more than I did at the same age?

Chetty and colleagues answered that in 2017 by combining tax data with historical Census data. Among Americans born in 1940, about 90 percent were earning more at age 30 than their parents had at the same age. Among those born in 1984, the figure was about 50 percent. A coin flip.

Then they did the decomposition that makes the paper important. They asked how much of the decline was caused by slower economic growth, and how much by the growth being distributed more unequally. Simulating the 1940 cohort's experience under later growth rates but the earlier distribution, and vice versa, they concluded that the more unequal distribution of growth accounted for the larger part of the fall. Returning to mid-century growth rates alone would not restore mid-century absolute mobility, because the gains would not reach the same people.

Key idea: Relative mobility measures whether people change rank. Absolute mobility measures whether they out-earn their parents in real terms. The American decline in absolute mobility, from about 90 percent to about 50 percent across those cohorts, was driven more by how growth was distributed than by how much there was.

The Great Gatsby Curve, and how much weight it carries

Plot each country's income inequality on the horizontal axis and its intergenerational elasticity on the vertical, and the points line up: more unequal countries have less mobility. Miles Corak assembled the version that became famous, and Alan Krueger gave it the name in a 2012 speech.

The correlation is real and it replicates. Now the cautions, which are not quibbles.

The sample is a dozen or so countries, which is a very small n for a cross-sectional relationship. The two axes are measured with different instruments in different countries and different years. And the relationship is a correlation across countries, which cannot by itself establish that inequality causes immobility. Three other stories fit the same scatter: immobility could cause inequality, both could be produced by some third factor such as the structure of educational institutions, or countries could differ in ways that generate both.

There is a mechanism that makes the causal story plausible, which is why the curve is taken seriously rather than dismissed. If the rungs of the ladder are further apart, the resources a parent can invest in a child differ more, and the consequence of falling differs more. But plausibility is not identification, and anyone who tells you the Great Gatsby Curve proves inequality causes immobility has skipped a step.

Place: from correlation to something closer to cause

The geographic variation in the opening of this lesson raises the obvious question. Is Charlotte bad for poor children, or do different kinds of family end up in Charlotte?

Chetty and Hendren answered with a design that gets much closer to an answer. They studied five million families who moved between commuting zones, and asked whether a child's adult outcome improved in proportion to the number of childhood years spent in the better place. It did, roughly linearly: each additional year of exposure to a higher-mobility area shifted the child's expected outcome by a consistent increment, and children who moved when young gained more than their older siblings who moved at the same time.

The sibling comparison is what makes this convincing. It compares children within the same family, moving on the same date, differing only in age at the move. Whatever made the family the kind of family that moved is held constant.

Descriptively, high-mobility places share five correlates in the original work: less residential segregation, less income inequality, better primary schools, more two-parent families, and greater social capital. Those are correlations, and Chetty's team said so. Later work has pushed on one of them directly: a 2022 analysis of Facebook friendship data covering tens of millions of users found that economic connectedness, the extent to which low-income people have high-income friends, was the neighbourhood characteristic most strongly associated with upward mobility.

Siblings, surnames, and the long run

Two other designs are worth knowing because they measure things a parent-child regression cannot.

Sibling correlations ask how much of the variation in adult outcomes is shared between brothers or sisters. That captures everything siblings have in common: parental income, but also neighbourhood, schools, genes, and family culture. American sibling correlations in earnings run around 0.4 to 0.5, meaningfully higher than the parent-child elasticity, which tells you that family background operates through more channels than income.

Surname studies take the long view. Gregory Clark tracked rare surnames associated with elite status through records of Oxford and Cambridge admissions, medical registers, and probate records across centuries, and argued that underlying social status regresses to the mean far more slowly than single-generation estimates suggest, with a persistence coefficient nearer 0.7 to 0.75 per generation in every society he examined. If that is right, conventional mobility estimates understate persistence badly, because income in a single year is a noisy indicator of an underlying status that transmits much more reliably.

The objections are substantial. Rare elite surnames are a selected group; regression to the mean operates differently at the extremes; and the method assumes a latent status variable that is never directly measured. Treat Clark's estimates as a serious challenge that has not been generally accepted rather than as an established finding.

Common misconceptions

  • An IGE of 0.4 means 40 percent of your income comes from your parents. It is an elasticity relating proportional differences, and the share of variance explained is much smaller.
  • Whatever the IGE does not explain is individual effort. The residual contains luck, health, local labour markets, discrimination, and measurement error, in unknown proportions.
  • Different countries' IGEs are directly comparable. They are sensitive to how many years of parental income are averaged, and studies differ on this, which explains much of the apparent disagreement.
  • High rank mobility means a fair society. Rank mobility says nothing about how far apart the rungs are; a society can shuffle people efficiently between very unequal positions.
  • The Great Gatsby Curve proves that inequality causes immobility. It is a cross-country correlation on about a dozen observations, consistent with several causal stories.

What you now know

  • A child of parents at the 25th percentile had a 7.5 percent chance of reaching the top fifth nationally, ranging from 4.4 percent in Charlotte to 12.9 percent in San Jose.
  • The IGE is an elasticity, not a share of variance, and it is attenuated by measuring parental income over too few years.
  • The rank-rank slope of about 0.34 in the United States is robust to zero and extreme incomes but is purely relative and blind to the distance between rungs.
  • The share of Americans out-earning their parents at 30 fell from about 90 percent for the 1940 cohort to about 50 percent for the 1984 cohort, driven mainly by the distribution of growth rather than its rate.
  • The Great Gatsby Curve is a real correlation across roughly a dozen countries and cannot on its own identify a causal direction.
  • The exposure design comparing siblings who moved at different ages provides much stronger evidence that place itself matters, with economic connectedness the strongest neighbourhood correlate identified so far.

Sources

  1. Chetty, R., Hendren, N., Kline, P., and Saez, E. (2014). Where is the land of opportunity? The geography of intergenerational mobility in the United States. Quarterly Journal of Economics, 129(4), 1553-1623. doi.org
  2. Chetty, R., Grusky, D., Hell, M., Hendren, N., Manduca, R., and Narang, J. (2017). The fading American dream: Trends in absolute income mobility since 1940. Science, 356(6336), 398-406. doi.org
  3. Corak, M. (2013). Income inequality, equality of opportunity, and intergenerational mobility. Journal of Economic Perspectives, 27(3), 79-102. doi.org
  4. Opportunity Insights. (n.d.). Research and data on economic opportunity. Harvard University. opportunityinsights.org
  5. Clark, G. (2014). The son also rises: Surnames and the history of social mobility. Princeton, NJ: Princeton University Press.
Key terms
Intergenerational elasticity
The coefficient from regressing log child income on log parent income, reporting the proportional difference in child income associated with a proportional difference in parent income.
Rank-rank slope
The coefficient from regressing a child's percentile rank on a parent's percentile rank, robust to zero and extreme incomes but purely relative.
Absolute mobility
Whether children out-earn their parents in real terms at the same age, as distinct from whether they change position in the distribution.
Attenuation bias
The pull of a regression coefficient toward zero when the explanatory variable is measured with error, which is why single-year parental income understates the IGE.
Great Gatsby Curve
The cross-country correlation between income inequality and intergenerational persistence, which is real but rests on roughly a dozen observations.
Exposure design
Estimating a place effect from how much a child's outcome improves per year of childhood spent in a new area, using siblings of different ages who moved on the same date.
Economic connectedness
The extent to which people with low incomes have friendships with people of high income, the neighbourhood measure most strongly associated with upward mobility in the social capital study.
Sibling correlation
The share of variance in adult outcomes shared between siblings, which captures all common family and neighbourhood influences rather than parental income alone.

Education: Escalator, Sorting Machine, or Both

  • Summarise the Coleman Report's central finding and the two ways it is routinely misread.
  • Contrast human capital and signalling explanations of the education wage premium and identify evidence that distinguishes them.
  • Interpret the seasonal comparison research showing when achievement gaps grow, and what it implies about schools.
  • Describe stratification within higher education, including selective-college returns and the Ivy-Plus admissions findings.

Six hundred thousand students, one uncomfortable result

In July 1966 James Coleman delivered a report Congress had commissioned under the Civil Rights Act. His team had surveyed around 570,000 students and 60,000 teachers in some 4,000 schools, the largest social science study conducted in the United States to that point. Congress expected a catalogue of unequal school resources that would explain unequal achievement.

What the report found instead was that measured differences in school resources, per-pupil spending, library books, laboratory equipment, teacher credentials, explained far less of the variation in student achievement than differences in family background did. The strongest school-level correlate of a student's achievement was the social composition of the student body, not the budget.

That finding has been misread in opposite directions for sixty years, and both misreadings are worth naming before we go anywhere else.

The first misreading is that schools do not matter. The report compared schools that existed in 1965 on the resources measured in 1965; it could not evaluate schools better than any then operating, and it did not measure teaching practice, which turns out to matter enormously. The second misreading is that Coleman showed poverty is destiny. He showed that variation in the measured inputs was weakly associated with variation in outcomes, which is a claim about the range of the data, not about what is possible.

Why the wage premium exists: two accounts

Start with the fact that any theory has to explain. The Bureau of Labor Statistics tracks median usual weekly earnings by educational attainment. Workers with a bachelor's degree earn roughly 1,500 dollars a week against roughly 900 for those whose highest qualification is a high school diploma, and the unemployment rate for the degree holders is consistently about half. Estimates from the Mincer earnings equation put the average return at something like 8 to 10 percent additional earnings per year of schooling.

Two theories explain that premium, and they have different policy implications, so distinguishing them is not an academic exercise.

Human capital theory, developed by Gary Becker and Jacob Mincer, says schooling makes people more productive. Skills are acquired, employers pay for productivity, and the premium is the return on an investment. If this is the whole story, expanding education raises national output and individual earnings together.

Signalling theory, formalised by Michael Spence in 1973, says schooling reveals productivity that already existed. Education is costly, and it is less costly for the able, so completing it credibly signals ability to employers who cannot observe it directly. If this is the whole story, expanding education raises credentials without raising output, and the private return exceeds the social return.

What separates them empirically? The best-known test is the sheepskin effect: earnings jump discontinuously at degree completion, more than the extra year of study alone would predict. A student who completes three and a half years of a four-year degree has nearly all the human capital and none of the certificate, and earns substantially less. That is difficult to explain on pure human capital grounds and easy to explain on signalling grounds.

The honest position is that both operate. Almost nobody in the field holds either pure version. What is genuinely contested is the mix, and the mix varies by field: a nursing qualification transmits verifiable clinical skill, while a generalist degree in a non-vocational subject does considerably more signalling.

What matters here: The education wage premium is real and large, but its existence does not by itself tell you whether education creates the productivity it is paid for or reveals it. Sheepskin effects at degree completion are the clearest evidence that signalling is doing part of the work.

When do gaps grow? The seasonal evidence

Here is a study design that changes how you read everything else. If schools generate inequality, achievement gaps should widen fastest while school is in session. If schools compensate for inequality, gaps should widen fastest when school is out.

Douglas Downey, Paul von Hippel, and Beckett Broh used data that tested children in both autumn and spring, which allows the school year and the summer to be separated. The pattern they found was that socioeconomic gaps in achievement grew faster during summers than during school terms. Later work using the same seasonal logic has generally supported the direction of that finding, though the size varies with the test and the cohort.

Read that carefully, because it cuts against intuition on both sides. It does not say schools are excellent. It says schools are, on the whole, an equalising environment relative to what children experience when they are not in them, and that a large share of the gap that appears by the end of schooling was assembled outside it.

The corollary is that much of what looks like school failure is the accumulated effect of unequal environments that schools partially offset rather than create. It is also why the achievement gap is already large at kindergarten entry, before schooling has had any opportunity to act.

Gaps that grew, and the one that did not

Sean Reardon assembled a long series of achievement gaps by family income across nineteen nationally representative studies. His central finding was that the gap in test scores between children from families at the 90th percentile of income and the 10th grew by around 40 percent between the cohort born in the mid-1970s and the cohort born around 2001, and that this income gap is now larger than the Black-white achievement gap, which narrowed substantially over the same period.

The two trends together are the point. A society can reduce one axis of educational inequality while another widens, and reporting only the improving one or only the worsening one misdescribes what happened.

Sorting inside the system

Mass higher education did not eliminate educational stratification; it relocated it. When almost everyone can go to college, the question stops being whether and becomes where.

The American system is steeply tiered, from open-access community colleges through regional public universities to a small set of highly selective private institutions. The tiers differ in completion rates, in resources per student, and in the networks they confer.

So does attending a selective institution cause higher earnings? Stacy Dale and Alan Krueger designed the study that gets closest to an answer. They compared students who were admitted to the same set of selective colleges but chose to attend institutions of different selectivity. Among students with the same admission portfolio, attending the more selective college produced little or no earnings advantage on average. The exceptions matter: for Black and Hispanic students and for students whose parents had less education, attending the more selective institution did raise earnings.

That pattern has a clean interpretation. Where a student already has family networks and information, the institution adds little that the family was not supplying. Where the student does not, the institution supplies it, and the effect appears.

A separate line of work asks who gets in. Analysing admissions and tax records for the Ivy-Plus institutions, Chetty, Deming, and Friedman found that applicants from families in the top 1 percent of the income distribution were substantially more likely to be admitted than applicants from middle-income families with the same standardised test scores, with the gap driven by legacy preferences, recruited athletics, and the weighting of non-academic ratings that correlate with attending a private secondary school. At the same time, institutions such as the City University of New York and the California State system enrol large numbers of students from the bottom of the distribution and move a substantial fraction of them to the top, which is a different kind of institutional success entirely and one that selectivity rankings do not measure.

The reproduction argument

Set against the escalator account is a tradition that treats schooling primarily as a mechanism of reproduction.

Samuel Bowles and Herbert Gintis argued in Schooling in Capitalist America that schools mainly transmit the dispositions each class position requires: punctuality and rule-following in schools serving working-class neighbourhoods, initiative and internalised standards in schools serving professional ones. Bourdieu's version, developed with Jean-Claude Passeron, is that schools reward a cultural capital that middle-class children arrive already holding, then present the resulting advantage as individual merit, which converts inherited advantage into legitimate achievement.

The strongest evidence for the reproduction view is the durability of the income achievement gradient, its presence before schooling begins, and the finding that students of similar measured ability from different backgrounds end up in different places in the system. The strongest evidence against a pure reproduction account is that education does raise earnings for those who obtain it, including for those from the bottom, and that the specific institutions with the highest mobility rates are neither elite nor especially resourced.

Common misconceptions

  • The Coleman Report proved schools do not matter. It found that the school resources measured in 1965 explained less variance than family background, which is a statement about that range of inputs, not about the potential of schooling.
  • The earnings premium proves education raises productivity. The premium is consistent with both human capital and signalling, and sheepskin effects at completion indicate signalling is part of the answer.
  • Achievement gaps open because of what happens in classrooms. Seasonal comparison research finds socioeconomic gaps growing faster in summer than in term, and gaps are already sizeable at kindergarten entry.
  • Getting into a more selective college raises anyone's earnings. Among students with the same admission portfolio the average effect is close to zero, with real gains concentrated among minority and first-generation students.
  • Educational expansion by itself equalises. When participation rises, stratification moves from whether you attend to where you attend and whether you complete.

Summing up

  • Coleman's 1966 survey of some 570,000 students found family background and student-body composition more strongly associated with achievement than measured school resources.
  • The education wage premium is large, roughly 1,500 dollars a week for bachelor's degree holders against about 900 for high school completers, and both human capital and signalling contribute to it.
  • Sheepskin effects, the discontinuous earnings jump at degree completion, are the clearest single piece of evidence for signalling.
  • Seasonal comparison research finds socioeconomic achievement gaps widening faster during summers than during school terms, implying schools compress rather than create them.
  • The income achievement gap grew by roughly 40 percent across cohorts born from the mid-1970s to 2001 while the Black-white gap narrowed.
  • Selective college attendance has little average earnings effect among similarly admitted students but real effects for minority and first-generation students, and admissions at the most selective institutions favour the top 1 percent at equal test scores.

Sources

  1. National Center for Education Statistics. (n.d.). Data and reports on the condition of education. U.S. Department of Education. nces.ed.gov
  2. U.S. Bureau of Labor Statistics. (n.d.). Education pays: Earnings and unemployment rates by educational attainment. bls.gov
  3. Dale, S. B., and Krueger, A. B. (2002). Estimating the payoff to attending a more selective college. Quarterly Journal of Economics, 117(4), 1491-1527. doi.org
  4. Opportunity Insights. (n.d.). Research on higher education and economic mobility. Harvard University. opportunityinsights.org
  5. Coleman, J. S., Campbell, E. Q., Hobson, C. J., McPartland, J., Mood, A. M., Weinfeld, F. D., and York, R. L. (1966). Equality of educational opportunity. Washington, DC: U.S. Government Printing Office.
  6. Bowles, S., and Gintis, H. (1976). Schooling in capitalist America: Educational reform and the contradictions of economic life. New York: Basic Books.
Key terms
Coleman Report
The 1966 federal survey of some 570,000 students that found family background and student-body composition more predictive of achievement than measured school resources.
Human capital theory
The account in which schooling raises productivity, so the wage premium is a return on an investment in skill.
Signalling theory
Spence's account in which education reveals pre-existing productivity to employers who cannot observe it, because completion is less costly for the able.
Sheepskin effect
The discontinuous jump in earnings at degree completion, larger than the additional study time alone would predict, taken as evidence for signalling.
Seasonal comparison
A design that tests students in autumn and spring so that learning during the school year can be separated from change over the summer.
Income achievement gap
The difference in measured achievement between children of high-income and low-income families, which grew by around 40 percent across cohorts born from the mid-1970s to 2001.
Mobility rate of a college
The share of a college's students who come from the bottom of the income distribution and reach the top, which ranks institutions very differently from selectivity.
Educational reproduction
The argument that schooling transmits and legitimates existing class advantage, by rewarding cultural competences that advantaged children arrive already holding.

Module 4: Ascribed Divisions

Three stratification systems that operate alongside class and interact with it: race and ethnicity, gender and the division of paid and unpaid work, and the neighbourhood, treated as a variable with measurable causal effects rather than as a backdrop.

Race and Ethnicity as a Stratification System

  • Compare the racial income gap with the racial wealth gap and explain why the second is several times larger.
  • Summarise the audit study evidence on hiring discrimination and state what the design can and cannot establish.
  • Distinguish taste-based from statistical discrimination and identify which policies each implies.
  • Explain why panethnic categories conceal within-group variation, using measured examples.

Two ratios that are not the same size

In the 2022 Survey of Consumer Finances, the median white non-Hispanic family held 285,000 dollars in net worth. The median Black family held 44,900 dollars. The median Hispanic family held 61,600 dollars.

Now hold that beside the income figures. Census data for roughly the same period put median household income for non-Hispanic white households near 81,000 dollars and for Black households near 53,000. That is a ratio of about 1.5 to 1.

The wealth ratio is over 6 to 1. The income ratio is about 1.5 to 1. Any explanation of racial stratification that works only on wages is explaining the smaller of the two gaps, and the arithmetic of why they differ is where this lesson starts.

Why the wealth gap is the larger one

Recall the mechanism from the wealth lesson: net worth accumulates out of saving, inheritance, and compounding returns over decades, so it records history in a way that annual income does not. Four specific channels do most of the work here.

Timing of homeownership. Home equity is the dominant asset for households in the middle of the American wealth distribution. Federal mortgage insurance programmes from the 1930s through the 1960s used underwriting maps that graded neighbourhoods, marking those with Black residents as ineligible, and the resulting mortgages built white suburban equity during the largest housing appreciation in American history. A family that bought in 1955 has had seventy years of compounding. A family excluded until the Fair Housing Act of 1968, and in practice for years afterward, has not.

Inheritance and inter vivos transfers. Wealth passes forward. Families whose grandparents accumulated equity can supply a deposit; families whose grandparents were excluded cannot. The transfer need not be large to be decisive, because it converts renting into owning, and ownership then compounds.

Homeownership rates today. The Census housing surveys put the white homeownership rate near three quarters and the Black rate near 45 percent, a gap that has been remarkably durable across decades.

Differential returns on the same asset. Houses in majority-Black neighbourhoods have historically appreciated less and been appraised lower, so identical saving behaviour produces unequal accumulation.

Key idea: Income is a flow that reflects this year's labour market. Wealth is a stock that reflects a century of who could buy, borrow, and inherit. That is why the racial wealth ratio is roughly four times the racial income ratio, and why closing wage gaps alone would close the wealth gap only very slowly.

Measuring discrimination directly

A wage gap is not evidence of discrimination, because the people being compared differ in many ways. The field's solution is the audit study, in which the researcher controls the applicant.

The best-known is Marianne Bertrand and Sendhil Mullainathan's 2004 experiment. They sent about 5,000 fictitious resumes to help-wanted advertisements in Boston and Chicago, randomly assigning names that were distinctively white-sounding or distinctively African American sounding to otherwise identical documents. Resumes with white-sounding names received callbacks at roughly one in ten; those with African American sounding names at roughly one in fifteen. That is about 50 percent more callbacks for the same credentials, and the gap was equivalent to the effect of eight additional years of experience.

Devah Pager's 2003 study added a second dimension. She sent matched pairs of real testers to apply for entry-level jobs in Milwaukee, with half of each pair reporting a felony drug conviction. White applicants without a record received callbacks at roughly 34 percent, white applicants with a record at 17 percent, Black applicants without a record at 14 percent, and Black applicants with a record at 5 percent. Read across that table once. A white applicant with a felony conviction did about as well as a Black applicant with a clean record.

Has this improved? Lincoln Quillian and colleagues answered by meta-analysis in 2017, pooling every field experiment on hiring discrimination conducted in the United States between 1989 and 2015. They found no change in the level of discrimination against Black applicants over that quarter century. The estimated decline for Hispanic applicants was modest and not precisely estimated.

Be clear about what audit studies do and do not establish. They establish causation at a specific decision point, because the only thing randomised is the signal of race. They do not measure the size of discrimination in the whole labour market, because they observe callbacks rather than hires and wages, and because they sample the kinds of jobs that advertise publicly. They also cannot detect discrimination that occurs before an application is submitted, such as who hears about the opening at all.

Two theories of why an employer discriminates

Gary Becker's model treats discrimination as a taste: the employer has a preference against hiring from a group and is willing to pay for it, in the form of hiring less productive workers or paying higher wages to preferred ones. Becker's prediction is that competition erodes this, because a firm indulging a costly taste is at a disadvantage against one that does not.

Edmund Phelps and Kenneth Arrow proposed statistical discrimination instead. Here the employer has no preference at all, only imperfect information. Facing an applicant whose individual productivity is unobservable, the employer uses group averages as a predictor. No animus is required, and competition does not erode it, because using available information is profitable rather than costly.

The distinction matters because the two imply different remedies. If discrimination is taste-based, more competition and antidiscrimination enforcement help. If it is statistical, the remedy is better individual information: credentials, verified records, structured assessment, work samples.

Distinguishing them empirically is hard. One useful test is what happens when better individual information is supplied. Some evidence from criminal background check policies points in an uncomfortable direction: when employers were prohibited from asking about criminal records, callbacks to young Black men without records in some settings fell rather than rose, consistent with employers substituting group-level statistical inference for the individual information they had lost. That result is contested and context-dependent, but it illustrates the point that a policy aimed at one mechanism can worsen outcomes if the other mechanism is operating.

The point: Taste-based and statistical discrimination produce similar-looking gaps and require different remedies. A policy that removes information can reduce taste-based discrimination while increasing statistical discrimination in the same market.

What has and has not converged

The Black-white earnings gap among men narrowed substantially between about 1940 and the mid-1970s, driven by migration out of the rural South, the dismantling of legal segregation, and rising relative educational attainment. It then stopped narrowing.

Patrick Bayer and Kerwin Kofi Charles made a sharper point by measuring position rather than ratios. Asking where the median Black man sits in the white male earnings distribution, they found that by the middle of the 2010s his position was no better than in 1950. Part of that is composition: as the earnings distribution spread out, holding one's relative position required faster absolute progress. Part of it is the large fraction of Black men with zero recorded earnings, which ratio-based comparisons of workers exclude by construction and which position-based comparisons capture.

That is a good example of why measurement choice is not a technicality. Two defensible statistics about the same population point in different directions, and the reason is that they are counting different people.

Ethnicity is not race, and panethnic labels hide a lot

Race, as used in American statistics, is a small set of official categories. Ethnicity is narrower and often self-defined by national origin, language, and descent. The two come apart constantly, and the aggregate categories used in official data conceal enormous variation.

The clearest case is the Asian American category. Pew Research Center analyses have found income inequality within the Asian American population to be higher than within any other major racial or ethnic group in the United States, because a single label covers Indian American households with very high median incomes and Hmong, Burmese, and Bhutanese American households with poverty rates several times the national average. Reporting an Asian American median therefore describes almost no one.

Immigrant selectivity explains part of the variation. People who migrate are not random draws from their origin country, and visa systems select further. Nigerian Americans have among the highest rates of graduate degree attainment of any national origin group in the United States, which reflects who was able to emigrate under skill-based criteria rather than any general property of national origin.

American Indian and Alaska Native populations occupy a different position again, because tribal citizenship is a political status arising from treaties rather than a racial category, and poverty rates on many reservations run far above the national figure.

Common misconceptions

  • The racial wealth gap is just the income gap accumulated. The wealth ratio is roughly four times the income ratio, because inheritance, homeownership timing, and differential asset appreciation add concentration on top of wage differences.
  • Audit studies measure how much discrimination exists in the economy. They measure a randomised effect at the callback stage in advertised jobs, which is a strong causal claim about a narrow moment.
  • Statistical discrimination is not really discrimination. It imposes the same costs on the individual, and it is unlawful in most employment contexts regardless of whether animus was present.
  • Discrimination in hiring has steadily declined. The meta-analysis of field experiments from 1989 to 2015 found no decline in discrimination against Black applicants.
  • Panethnic statistics describe their groups. The Asian American category spans the highest and among the lowest household incomes in the country, so its median describes a population that does not exist.

What to carry forward

  • Median family net worth in the 2022 SCF was 285,000 dollars for white non-Hispanic families, 61,600 for Hispanic families, and 44,900 for Black families, a far wider gap than the income ratio of roughly 1.5 to 1.
  • Homeownership timing, exclusion from federally insured mortgages, inheritance, and differential appreciation explain why wealth gaps exceed wage gaps.
  • Bertrand and Mullainathan found about 50 percent more callbacks for white-sounding names on identical resumes; Pager found a white applicant with a felony record fared about as well as a Black applicant without one.
  • Quillian and colleagues found no decline in discrimination against Black applicants across field experiments from 1989 to 2015.
  • Taste-based and statistical discrimination look alike in the data and imply different remedies, and removing individual information can increase the statistical kind.
  • Panethnic categories such as Asian American conceal the widest within-group income inequality in the country, and immigrant selectivity explains much of the variation between national origin groups.

Sources

  1. Board of Governors of the Federal Reserve System. (2023). Changes in U.S. family finances from 2019 to 2022: Evidence from the Survey of Consumer Finances. federalreserve.gov
  2. Bertrand, M., and Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991-1013. doi.org
  3. Quillian, L., Pager, D., Hexel, O., and Midtboen, A. H. (2017). Meta-analysis of field experiments shows no change in racial discrimination in hiring over time. Proceedings of the National Academy of Sciences, 114(41), 10870-10875. doi.org
  4. Pager, D. (2003). The mark of a criminal record. American Journal of Sociology, 108(5), 937-975. doi.org
  5. Pew Research Center. (n.d.). Social and demographic trends. pewresearch.org
Key terms
Racial wealth gap
The difference in net worth between racial groups, which in the 2022 SCF ran at roughly six to one between white and Black median families.
Audit study
A field experiment in which otherwise identical applications differ only in a randomly assigned signal of group membership, allowing a causal estimate at that decision point.
Correspondence test
An audit conducted on paper, sending matched fictitious resumes rather than sending human testers, which permits much larger samples.
Taste-based discrimination
Becker's model in which an employer holds a preference against a group and bears a cost to indulge it, which competition should erode.
Statistical discrimination
Phelps and Arrow's model in which employers use group averages to predict unobserved individual productivity, requiring no animus and not eroded by competition.
Immigrant selectivity
The fact that migrants are not a random sample of their origin population, and are further selected by visa systems, which shapes group outcomes on arrival.
Panethnic category
An official grouping such as Asian American or Hispanic that spans many national origins with widely different outcomes, so that its aggregate statistics describe no actual subgroup well.
Positional comparison
Measuring where a group's median member sits within another group's whole distribution, which can show stagnation where ratio comparisons show progress.

Gender, Paid Work, and the Work That Is Not Paid

  • Interpret the gender pay gap correctly, distinguishing the raw gap from the adjusted residual and saying what each one is evidence of.
  • Explain the child penalty design and what it establishes that a cross-sectional wage comparison cannot.
  • State Goldin's account of nonlinear pay for hours and use it to predict which occupations have large and small gaps.
  • Quantify the division of unpaid care work and explain how it interacts with occupational choice and earnings.

The graph that turns at one moment

Using Danish administrative records covering the whole population from 1980 to 2013, Henrik Kleven, Camille Landais, and Jakob Sogaard did something simple and devastating. They plotted men's and women's earnings year by year around the birth of a first child.

Before the birth, the two lines move together. At the birth, the women's line drops. It partially recovers, then flattens, and ten years later it sits about 20 percent below where the pre-birth trend says it should be. The men's line does not move at all.

They called it the child penalty, and its importance is in the design rather than the size. Every worker in the study is compared to herself, before and after a dated event, with the trajectory of men experiencing the same event as the control. There is no need to argue about whether women choose different fields or negotiate differently, because the same women are on both sides of the comparison.

Denmark, moreover, is not a hard case. It has universal subsidised childcare, long paid leave available to both parents, and one of the smallest raw pay gaps in the world. The penalty appears anyway.

Reading the pay gap without embarrassing yourself

The headline figure most people encounter is that women earn about 83 cents for every dollar men earn, from the Bureau of Labor Statistics comparison of median usual weekly earnings for full-time wage and salary workers. Census figures for annual earnings among full-time year-round workers give a similar ratio.

Three things about that number are routinely misstated, in both directions.

First, it is not a comparison of a woman and a man in the same job. It compares all full-time working women to all full-time working men, across different occupations, industries, hours, and experience.

Second, restricting to full-time workers understates the total earnings gap, because women are considerably more likely to work part-time. A comparison including all workers produces a wider gap.

Third, adjusting for everything measurable does not make the gap disappear. Francine Blau and Lawrence Kahn's comprehensive review traced how the composition of the gap changed: education used to explain part of it and now works the other way, since women in recent cohorts have higher average attainment than men. What remains explained is mostly occupation, industry, and experience. What remains unexplained after all of those adjustments has been estimated at something on the order of eight percentage points.

Now the interpretive trap, which is the important part. The unexplained residual is not a measure of discrimination, and the explained portion is not innocent. If women are in lower-paid occupations partly because of steering, expectations, and constrained hours, then controlling for occupation controls away part of the phenomenon you are studying. A regression coefficient does not know which of its controls are causes and which are consequences. You have to decide that, and say so.

What matters here: The raw gap and the adjusted gap answer different questions. The raw gap asks how differently men and women are paid in this economy. The adjusted gap asks how differently they are paid conditional on where they ended up. Controlling for occupation can absorb the very mechanism under investigation.

Goldin: the shape of the pay schedule

Claudia Goldin's account, laid out in her 2014 presidential address to the American Economic Association, explains why the gap persists most stubbornly among the highly educated and within occupations rather than between them.

Her argument is about the shape of the relationship between hours and pay. In some occupations, pay is roughly linear in hours: someone working 40 hours earns about half of what someone working 80 hours earns, and workers are substitutable, so one can hand off to another without loss. In others, pay is nonlinear: the person available at 9pm, who has carried the client relationship from the start, and who is never handed off, earns disproportionately more than twice the person working half those hours.

The prediction follows immediately. Where pay is linear and workers are substitutable, the gender gap should be small. Where it is nonlinear, the gap should be large, because whoever takes on the greater share of care work cannot supply the specific hours the schedule rewards.

Pharmacy is Goldin's exhibit. The occupation moved from independent proprietorships to chains and hospitals, pharmacists became genuinely substitutable, information systems allowed one to take over from another mid-prescription, and the gender earnings gap within the occupation became one of the smallest of any profession. Law, corporate management, and finance sit at the other end, and have among the largest within-occupation gaps.

The policy implication is not what people expect. Goldin's claim is that the residual gap is not primarily about employers valuing women less; it is about a pay structure that rewards particular hours at a premium, combined with an unequal division of the work that makes those hours impossible.

The unpaid half of the economy

Which brings us to the work that no statistical agency counts as work.

The International Labour Organization's global assessment found that women perform 76.2 percent of the total hours of unpaid care work worldwide, about 3.2 times the amount men perform. The ratio varies by region and narrows in richer countries, but it does not approach parity anywhere that has been measured.

American time diary data from the Bureau of Labor Statistics tells the same story with more granularity. Women spend more time on household activities and on the physical care of children, and the gap persists among couples where both partners work full-time. It narrowed substantially through the 1970s and 1980s and has narrowed much more slowly since.

Two features of this work matter for stratification. It is invisible in GDP, which counts a paid nanny and not a parent doing the same tasks, so national accounts record a rise in output when care shifts from home to market even if the same hours are worked. And it is time-inflexible in a specific way: a child must be collected at 3:30 whether or not the client call runs late. That inflexibility interacts precisely with Goldin's nonlinear pay schedules, which is why the two literatures are really one.

Michelle Budig and Paula England estimated a motherhood wage penalty of roughly 7 percent per child in American data, of which only part could be explained by lost experience and reduced hours. Some of the residual reflects employers' expectations of mothers, and some reflects effort constraints that are real rather than imagined.

Segregation, and the devaluation problem

Occupational sex segregation fell substantially between 1970 and 1990 as women entered law, medicine, and management, and then largely stopped falling. A large share of the American workforce still works in occupations that are strongly male or strongly female.

The obvious explanation is that female-dominated occupations happen to be lower paid. Asher Levanon, Paula England, and Paul Allison tested the causal direction with fixed-effects models on census data across decades, asking what happens to an occupation's pay after it feminises. They found evidence of devaluation: as the share of women in an occupation rose, its median pay fell relative to other occupations, controlling for skill demands and education. The direction of causation runs at least partly from composition to pay, not only from pay to composition.

Recreation and design work provide the standard illustrations, but the cleanest way to think about it is the counterexample: computer programming was substantially female in its early decades, and its pay and prestige rose as it masculinised.

What has converged, and what stalled

American women's labour force participation rose from roughly a third in 1950 to about 60 percent by 2000. Then it plateaued, and in the two decades since it has moved little. Over the same period participation continued rising in most other rich countries, several of which overtook the United States.

Blau and Kahn's cross-national comparison attributes a substantial share of that relative decline to policy: the United States is unusual among rich countries in having no statutory paid parental leave and limited public childcare, and countries that expanded both saw participation rise. The same analysis notes the trade-off honestly: long leave entitlements and part-time protections raise participation but are associated with women being less likely to reach management positions, apparently because long absences interrupt exactly the continuous attachment that nonlinear pay schedules reward.

The upshot: Policies that raise female employment and policies that raise female representation at the top are not the same policies, and some evidence suggests they trade against each other. That is a genuine tension rather than a failure of design.

Common misconceptions

  • The 83 cent figure compares a man and a woman doing the same job. It compares medians across all full-time workers, and restricting to full-time workers actually narrows rather than widens the total gap.
  • Once you control for occupation the gap disappears. It shrinks to roughly eight percentage points and does not vanish, and controlling for occupation absorbs part of the mechanism being studied.
  • The gap is closing steadily. It narrowed sharply through the 1980s and 1990s and has moved slowly since, and American female labour force participation has been roughly flat since 2000.
  • Generous family policy straightforwardly helps women. It raises participation and is associated with lower representation in management, a trade-off visible in cross-national data.
  • Female-dominated occupations are low paid because they require less skill. Fixed-effects work finds that pay falls after an occupation feminises, controlling for skill demands, which points at devaluation rather than composition alone.

Where this leaves us

  • The Danish child penalty design shows women's earnings falling about 20 percent below trend a decade after a first birth while men's do not move, in a country with strong family policy.
  • The raw gap of roughly 83 cents on the dollar and the adjusted residual of roughly eight percentage points answer different questions, and controlling for occupation can absorb the mechanism under study.
  • Goldin's account attributes much of the remaining gap to nonlinear pay for specific long hours, predicting small gaps where workers are substitutable, as in pharmacy, and large ones in law and finance.
  • Women perform about 76 percent of unpaid care work globally, roughly 3.2 times men's hours, and this work is time-inflexible in exactly the way that nonlinear schedules punish.
  • Budig and England estimated a motherhood wage penalty near 7 percent per child, only partly explained by lost experience and hours.
  • Occupational sex segregation fell to 1990 and then stalled, and evidence indicates pay falls after an occupation feminises rather than only the reverse.

Sources

  1. Kleven, H., Landais, C., and Sogaard, J. E. (2019). Children and gender inequality: Evidence from Denmark. American Economic Journal: Applied Economics, 11(4), 181-209. doi.org
  2. Goldin, C. (2014). A grand gender convergence: Its last chapter. American Economic Review, 104(4), 1091-1119. doi.org
  3. Blau, F. D., and Kahn, L. M. (2017). The gender wage gap: Extent, trends, and explanations. Journal of Economic Literature, 55(3), 789-865. doi.org
  4. U.S. Bureau of Labor Statistics. (n.d.). American Time Use Survey. bls.gov
  5. International Labour Organization. (2018). Care work and care jobs for the future of decent work. Geneva: ILO. ilo.org
Key terms
Child penalty
The persistent shortfall in a mother's earnings relative to her own pre-birth trajectory, estimated at about 20 percent ten years after a first birth in Denmark, with no equivalent for fathers.
Raw pay gap
The unadjusted difference in median earnings between men and women across all workers in scope, without controls for occupation, hours, or experience.
Adjusted pay gap
The difference remaining after controlling for measured characteristics, which answers a narrower question and can absorb part of the mechanism under study.
Nonlinear pay for hours
Goldin's term for pay schedules in which working twice the hours earns more than twice the pay, because specific availability and continuity command a premium.
Occupational substitutability
The degree to which one worker can take over another's task without loss, which Goldin identifies as the condition for small within-occupation gender gaps.
Unpaid care work
Domestic labour and care of children and dependent adults performed without payment, which women supply at roughly 3.2 times the rate men do worldwide.
Motherhood wage penalty
The earnings loss associated with each child, estimated by Budig and England at about 7 percent per child in American data.
Occupational devaluation
The finding that an occupation's relative pay falls after its share of women rises, controlling for skill demands, indicating causation from composition to pay.

Neighbourhoods, Segregation, and the Effects of Place

  • Compute and interpret the index of dissimilarity, and distinguish evenness from exposure as dimensions of segregation.
  • Explain Schelling's tipping model and what it shows about the relationship between individual preferences and aggregate patterns.
  • Summarise the Moving to Opportunity results, including why the early and later findings differed.
  • Distinguish selection from causation in neighbourhood effects research and name the designs that separate them.

A number that fell by a quarter and then stopped

The index of dissimilarity for Black and white residents across American metropolitan areas averaged roughly 0.79 in 1970. By the 2010s it had fallen to the mid-0.50s, where it has largely stayed. Hispanic and Asian segregation over the same period moved much less, and in some fast-growing metropolitan areas rose slightly as immigration built new concentrations.

Interpret the index before you interpret the trend. Dissimilarity runs from 0 to 1 and answers one specific question: what fraction of one group would have to move to a different neighbourhood for the two groups to be spread evenly across the metropolitan area? A value of 0.55 means that 55 percent of Black residents, or equivalently 55 percent of white residents, would have to relocate to produce an even distribution. Values above 0.6 are conventionally called high segregation.

So the decline is real, and the remaining level is large. Both halves of that sentence are usually dropped by whichever side is talking.

Two dimensions that are not the same

Dissimilarity measures evenness: how the groups are spread relative to each other. It has a blind spot, and it is a serious one. Dissimilarity is unaffected by group size, so a city where a tiny minority is perfectly evenly spread and a city where a large minority is perfectly evenly spread score identically, even though the daily experience of living in them differs completely.

The complement is exposure, or its inverse, isolation: what share of your neighbours belong to your own group, on average. Exposure does depend on group size. A Black resident of a metropolitan area that is 5 percent Black cannot have a majority-Black neighbourhood without extraordinary concentration; a Black resident of a metropolitan area that is 40 percent Black may live in one with only moderate unevenness.

Douglas Massey and Nancy Denton set out five distinct dimensions, adding concentration, centralization, and clustering, and argued that Black segregation in a number of American cities was severe on all five simultaneously, a pattern they called hypersegregation. The point of using several measures is that a city can improve on one and not the others.

Key idea: Evenness and exposure are different questions. Dissimilarity ignores group size and exposure does not, so reporting one without the other tells you either how spread out groups are or what people actually experience, but not both.

How the pattern was made

Three families of cause operate, and they are not alternatives.

Law and public administration. Racially restrictive covenants were written into deeds and were enforceable in American courts until 1948. The Federal Housing Administration's underwriting practice from the 1930s graded neighbourhoods and treated the presence of Black residents as a risk warranting denial of mortgage insurance, which channelled decades of federally supported lending toward new all-white suburbs. Urban renewal and interstate highway routing cleared and divided Black neighbourhoods on a large scale during the 1950s and 1960s. Exclusionary zoning, which mandates minimum lot sizes and prohibits multifamily housing, continues to operate in the present and does not mention race at all.

Private discrimination. Steering by agents, blockbusting, and unequal lending terms operated at the scale of individual transactions. Paired testing conducted for the Department of Housing and Urban Development has repeatedly found that minority testers are shown fewer available units than white testers with identical stated finances.

Sorting and preferences. Households sort on income as well as race, and Sean Reardon and Kendra Bischoff documented a rise in income segregation over recent decades, with a growing share of families living in either affluent or poor neighbourhoods rather than mixed ones.

Schelling's uncomfortable model

In 1971 Thomas Schelling published a model that every student of this subject should be able to reconstruct on a chessboard.

Place two kinds of counter on a grid at random. Give every counter one rule: if fewer than one third of my neighbours are like me, I move to the nearest square where at least one third are. Note what the rule is not. Nobody wants segregation. Every counter is content to be in a two-thirds minority.

Run it. The board segregates, thoroughly, in a few rounds. A counter moves because its own local threshold is violated; its departure pushes a neighbouring counter below its threshold; that one moves; and the cascade runs until large uniform blocks form.

The lesson is not that preferences do not matter. It is that the aggregate pattern cannot be read backwards to infer the preferences that produced it. A society can be sharply segregated with no one preferring segregation, which means that observing segregation does not tell you that anyone wanted it, and equally that eliminating hostile preferences would not automatically eliminate the pattern.

Schelling's model is often misused as an argument that segregation is natural and policy is futile. That does not follow either. The model shows that thresholds produce cascades; it says nothing about where the thresholds came from, and a great deal of American housing law was designed to create them.

Does the neighbourhood do anything, or does it just collect people?

Here is the hard causal question. Poor neighbourhoods have worse outcomes. Do they produce those outcomes, or do people with worse prospects end up in them?

In 1994 the Department of Housing and Urban Development launched an experiment designed to answer exactly this. Moving to Opportunity enrolled about 4,600 families living in high-poverty public housing in Baltimore, Boston, Chicago, Los Angeles, and New York, and randomly assigned them to three groups: a voucher usable only in a low-poverty neighbourhood plus counselling, an unrestricted voucher, or a control group receiving neither.

The interim evaluation, published about a decade in, was a disappointment to its designers. Adults who moved showed no significant gains in employment or earnings. They did show substantial improvements in mental health, in obesity and diabetes markers, and in feelings of safety, which are real outcomes and were not what the programme had been sold on.

Then in 2016 Raj Chetty, Nathaniel Hendren, and Lawrence Katz linked the participants to tax records and looked at the children, now grown. The pattern was age-dependent and sharp. Children who moved to a low-poverty neighbourhood before age 13 earned about 31 percent more in their mid-twenties than the control group, roughly 3,477 dollars a year, and were more likely to attend college and less likely to be single parents. Children who moved as teenagers did slightly worse than controls, consistent with the disruption of a move outweighing a short exposure.

That result reframed the whole literature. The effect of a neighbourhood is not a fixed quantity; it is roughly proportional to childhood years spent in it, which is why an evaluation measuring adults a few years after a move found nothing and an evaluation measuring children twenty years later found a great deal.

Why this matters: Randomisation plus a long enough follow-up turned an apparent null result into one of the largest measured effects in social policy. Both findings came from the same experiment. The difference was who was measured and for how long.

What the neighbourhood is made of

Robert Sampson's Chicago research programme asked what varies between neighbourhoods beyond poverty rates. His central construct is collective efficacy: the combination of social cohesion among residents and their willingness to intervene for the common good, measured by asking whether neighbours would act if children were skipping school or spraying graffiti. Collective efficacy predicts lower violence, and it does so after controlling for poverty, residential stability, and racial composition.

Sampson's follow-up work, tracking Chicago neighbourhoods across decades, found something further: the relative ranking of neighbourhoods is remarkably durable. Places that were disadvantaged in the 1960s were disadvantaged in the 2000s, through changes of population, industry, and policy. Neighbourhood inequality is not a snapshot of current sorting; it is a structure that persists across the people passing through it.

What follows for policy

If place matters causally and mainly through childhood exposure, then two policy families follow and they are in tension. Mobility programmes help families move to better places. Place-based investment improves the neighbourhood people already live in.

One recent finding sharpened the mobility side considerably. A randomised experiment in Seattle and King County by Peter Bergman and colleagues offered voucher recipients customised search assistance, help with landlord relationships, and short-term financial support. The share of families using their voucher in a high-opportunity neighbourhood rose from about 15 percent in the control group to about 53 percent among those offered the services. The vouchers had been available all along; the barrier was search, information, and landlord willingness rather than the subsidy.

Notice what this does to a common assumption. The gap between having a voucher and using it well was not a preference for staying; it was friction, and friction is addressable.

Common misconceptions

  • Segregation is mostly a legacy that is fading. The dissimilarity index fell substantially to about 1990 and then largely stopped, and exclusionary zoning operates in the present without mentioning race.
  • A low dissimilarity index means a group is not isolated. Dissimilarity ignores group size; exposure and isolation measure what residents actually experience.
  • Schelling's model shows segregation is natural and policy cannot help. It shows that thresholds cascade; where the thresholds came from is a separate question with a substantially legal answer.
  • Moving to Opportunity showed that neighbourhoods do not matter. The adult economic null was real, and the same experiment showed large gains for children who moved young, plus mental health gains for adults.
  • Families with vouchers who stay in poor neighbourhoods are revealing a preference. A randomised trial that removed search and landlord frictions raised high-opportunity use from about 15 percent to about 53 percent.

The takeaway

  • Black-white dissimilarity fell from roughly 0.79 in 1970 to the mid-0.50s and then plateaued, which is both a real decline and a high remaining level.
  • Evenness and exposure are distinct dimensions; dissimilarity is insensitive to group size and isolation is not, and Massey and Denton's hypersegregation means severity on several dimensions at once.
  • Restrictive covenants, FHA underwriting, urban renewal, highway routing, and exclusionary zoning are documented public causes, alongside private steering and lending discrimination and income sorting.
  • Schelling showed that mild individual thresholds cascade into near-total segregation, so aggregate patterns cannot be read backward to infer preferences.
  • Moving to Opportunity found no adult earnings effect but substantial adult mental health gains, and children who moved before age 13 later earned about 31 percent more.
  • Collective efficacy predicts lower violence net of poverty and composition, and neighbourhood rankings persist across decades and across the people who live in them.

Sources

  1. Chetty, R., Hendren, N., and Katz, L. F. (2016). The effects of exposure to better neighborhoods on children: New evidence from the Moving to Opportunity experiment. American Economic Review, 106(4), 855-902. doi.org
  2. Schelling, T. C. (1971). Dynamic models of segregation. Journal of Mathematical Sociology, 1(2), 143-186. doi.org
  3. U.S. Department of Housing and Urban Development. (n.d.). HUD User: Policy development and research. huduser.gov
  4. U.S. Census Bureau. (n.d.). Housing. census.gov
  5. Massey, D. S., and Denton, N. A. (1993). American apartheid: Segregation and the making of the underclass. Cambridge, MA: Harvard University Press.
  6. Sampson, R. J. (2012). Great American city: Chicago and the enduring neighborhood effect. Chicago: University of Chicago Press.
Key terms
Index of dissimilarity
The share of one group that would have to move neighbourhoods for two groups to be evenly distributed across a metropolitan area, running from 0 to 1.
Exposure and isolation
Measures of the share of one's neighbours belonging to one's own or another group, which unlike dissimilarity depend on group size.
Hypersegregation
Massey and Denton's term for severe segregation on several dimensions at once, including evenness, exposure, concentration, centralization, and clustering.
Exclusionary zoning
Land use rules such as minimum lot sizes and prohibitions on multifamily housing that restrict who can afford to live in a jurisdiction without referring to race.
Schelling tipping model
A simulation showing that mild individual thresholds for same-group neighbours cascade into near-total segregation, so patterns cannot be read back to preferences.
Moving to Opportunity
The 1994 randomised housing voucher experiment across five cities that produced a null adult earnings result and large gains for children who moved young.
Collective efficacy
Sampson's measure combining social cohesion with willingness to intervene for the common good, which predicts lower violence net of poverty and composition.
Search friction
The information, time, and landlord barriers that stop a voucher holder using a subsidy in a high-opportunity area, shown to be addressable by the Seattle experiment.

Module 5: Consequences and the Top

What stratification does to bodies, and who sits at the summit: the health gradient and the mechanisms proposed for it, then elites, the top one percent, and the specific ways large fortunes are made and kept.

The Health Gradient: Why Rank Predicts How Long You Live

  • State the measured association between income and life expectancy in the United States and explain why it is a gradient rather than a threshold.
  • Explain what the Whitehall studies established that a comparison of rich and poor countries cannot.
  • Apply fundamental cause theory to predict when a health gradient will widen.
  • Weigh social causation against health selection as explanations of the association.

Fourteen and a half years

In 2016 Raj Chetty and colleagues published in JAMA an analysis linking around 1.4 billion United States tax records to Social Security death records. Their headline estimate: among men, life expectancy at age 40 differed by 14.6 years between the top 1 percent and the bottom 1 percent of the income distribution. Among women the gap was 10.1 years.

Fourteen years is not a subtle finding. It is larger than the life expectancy cost of a lifetime of smoking, and it is a difference between groups within one wealthy country with universal emergency care and a very large medical system.

The first thing to establish is the shape of the relationship, because the shape rules out most of the easy explanations.

It is a gradient, not a cliff

If the association were driven by destitution, you would expect a threshold: very poor people would have much worse health, and above some level of adequacy the association would flatten. That is not what the data show. Each step up the distribution is associated with better health and longer life, all the way to the top. The person at the 80th percentile lives longer, on average, than the person at the 60th, who lives longer than the person at the 40th.

That pattern is called the social gradient in health, and it is one of the most reproducible findings in epidemiology. It appears in every country with the data to measure it, in every era measured, and for causes of death as different as cardiovascular disease, several cancers, respiratory illness, and injury.

The gradient is what makes the phenomenon interesting. Absolute deprivation cannot explain why a senior manager outlives a middle manager who outlives a junior one, when none of the three is deprived of anything material.

Whitehall: the study that removed the obvious explanations

Michael Marmot's studies of British civil servants are the classic demonstration, and their design is the reason.

Every participant was employed. Every participant had access to the National Health Service, free at the point of use. Nobody was in poverty. None worked in a hazardous industry. The population was, in the ways usually blamed for health inequality, remarkably homogeneous.

And mortality still stepped down the grade hierarchy. In the original Whitehall cohort, men in the lowest employment grades had coronary heart disease mortality several times that of men in the highest, with each intervening grade falling in order. The second study, Whitehall II, extended the design to women and to a wider set of outcomes and found the gradient again.

Marmot's interpretation, which he calls the status syndrome, emphasises two features of low-grade work: low control over the pace and content of one's tasks, and limited opportunity for full social participation. The physiological pathway proposed is chronic activation of stress responses, producing what researchers call allostatic load, the cumulative wear from repeated or sustained stress activation, expressed in blood pressure, cortisol regulation, inflammatory markers, and metabolic function.

Key idea: Whitehall matters because it holds constant income adequacy, health care access, and occupational hazard, and the gradient survives all three. Whatever produces it is not solely material deprivation or lack of medical care.

Four families of mechanism

No single pathway accounts for the gradient, and the honest position is that several operate together.

MechanismWhat it proposesIts strongest evidenceWhat it struggles with
MaterialExposure to hazards, poor housing, inadequate nutrition, unaffordable careLarge effects at the bottom of the distributionCannot explain gradients among comfortable professionals
BehaviouralSmoking, drinking, diet, exercise differ by positionBehaviours do differ steeply and are strong risk factorsAdjusting for them reduces but does not remove the gradient; behaviours are themselves stratified
PsychosocialChronic stress from low control, insecurity, and subordinate statusWhitehall gradient net of material factors; work control findingsStress is hard to measure directly, and the pathway is inferred more than observed
Early lifeConditions before birth and in childhood set adult disease riskBirth weight associations with adult cardiovascular diseaseSeparating prenatal effects from the continuing environment is difficult

The behavioural row deserves a caution, because it is where casual discussion of this topic usually stops. Smoking is more common at lower positions, and smoking causes disease. But treating that as an explanation just moves the question: why is smoking stratified? Cigarette prices are the same for everyone. The answer involves marketing history, stress, social networks, and the differential uptake of health information, all of which are themselves patterned by position.

Fundamental cause theory, and a prediction you can test

Bruce Link and Jo Phelan proposed a framework that explains something otherwise puzzling: the specific diseases responsible for the gradient keep changing, and the gradient does not.

In 1900 the leading killers were infectious. Today they are chronic. If the gradient were caused by any particular disease mechanism, it should have disappeared when that disease was controlled. It did not; it simply reattached to whatever the current major causes are.

Their explanation is that socioeconomic position is a fundamental cause: it commands flexible resources, money, knowledge, power, prestige, and useful social connections, which can be deployed against whatever the leading health risks happen to be at the time. Change the risks and the resources are redeployed. The gradient regenerates.

This yields a sharp and uncomfortable prediction. When an effective new preventive or treatment technology appears, socioeconomic gradients in the relevant disease should widen, because those with more resources adopt it first and use it more consistently. The comparison used to test this is between diseases where prevention is possible and diseases where it is not: for cancers with effective screening or known preventable causes, the socioeconomic gradient in mortality is steeper than for cancers with neither. A treatment can improve everyone's health and increase inequality at the same time.

The upshot: Fundamental cause theory predicts that health innovation without attention to distribution widens gaps. This is not an argument against innovation; it is an argument for expecting the widening and planning for it.

Does health cause status, or status cause health?

The reverse-causation worry is legitimate and has a name: health selection. Sick people earn less, are less likely to be promoted, and may drift down the occupational ladder. If so, the gradient could be illness producing low status rather than low status producing illness.

Selection is real and does contribute. The evidence that it is not the whole story is several-fold. In Whitehall, employment grade was established years before the mortality was observed, and grade at entry predicted later disease. Studies that adjust for baseline health still find gradients. And natural experiments in which position changes for reasons unrelated to health, such as a plant closure or a lottery win, show health effects in the expected direction.

The best-supported summary is that causation runs mainly from position to health, with selection contributing a real but smaller share, and with the two reinforcing each other over a life course.

The American case, and where it went wrong

United States life expectancy at birth was about 78.8 years in 2019. It fell to 76.4 by 2021 during the pandemic, and recovered to about 78.4 by 2023. Across the same decades, the country has spent a larger share of its national income on health care than any other, and its life expectancy has sat below that of comparable rich countries. Spending and outcomes are not the same axis.

Two American findings are worth carrying separately.

Anne Case and Angus Deaton documented rising midlife mortality among white non-Hispanic Americans beginning around 1999, at a time when mortality was falling in every other rich country and among other American groups. The rise was concentrated among those without a bachelor's degree and was driven by drug overdose, suicide, and alcohol-related liver disease, which they named deaths of despair. Their proposed explanation is the cumulative disadvantage of a deteriorating labour market for people without degrees, operating through family formation, community institutions, and pain.

The second is geographic. In the Chetty analysis, life expectancy for people in the bottom income quartile varied by roughly four to five years across American commuting zones. What predicted the good places was not the local supply of medical care and not local income inequality, but a cluster of local characteristics associated with health behaviours, including smoking and obesity rates and local public health policy. Where a low-income person lives changes how long they live, by an amount comparable to major clinical risk factors.

Common misconceptions

  • Health inequality is about poverty. It is a gradient running the full length of the distribution, and it appears among comfortable professionals in Whitehall who differ only in rank.
  • Universal health care would remove the gradient. British civil servants all had NHS access, and the gradient was fully present.
  • The gradient is explained by smoking and diet. Adjusting for behaviours reduces it without removing it, and the behaviours are themselves stratified by position.
  • Medical advances reduce health inequality. Fundamental cause theory predicts and evidence supports the opposite in the short run, since those with more resources adopt effective new prevention first.
  • The association is just sick people becoming poor. Health selection is real, but grade measured years before disease predicts later mortality, and shocks to position produce health effects in the expected direction.

Putting it together

  • Life expectancy at 40 in the United States differed by 14.6 years for men and 10.1 for women between the top and bottom 1 percent of income.
  • The association is a gradient at every step of the distribution, not a threshold effect of deprivation.
  • Whitehall removed poverty, health care access, and occupational hazard as explanations, and the stepwise mortality gradient by employment grade remained.
  • Material, behavioural, psychosocial, and early life mechanisms all contribute, and behaviours cannot be the terminal explanation because behaviours are themselves stratified.
  • Fundamental cause theory explains why the gradient survives changes in which diseases kill people, and predicts that effective new prevention widens gaps before it narrows them.
  • American life expectancy fell from 78.8 in 2019 to 76.4 in 2021 and recovered to about 78.4 in 2023, and midlife deaths of despair rose among non-degree-holding white Americans from around 1999.

Sources

  1. Chetty, R., Stepner, M., Abraham, S., Lin, S., Scuderi, B., Turner, N., Bergeron, A., and Cutler, D. (2016). The association between income and life expectancy in the United States, 2001-2014. JAMA, 315(16), 1750-1766. doi.org
  2. Marmot, M. G., Smith, G. D., Stansfeld, S., Patel, C., North, F., Head, J., White, I., Brunner, E., and Feeney, A. (1991). Health inequalities among British civil servants: The Whitehall II study. The Lancet, 337(8754), 1387-1393. doi.org
  3. Case, A., and Deaton, A. (2015). Rising morbidity and mortality in midlife among white non-Hispanic Americans in the 21st century. Proceedings of the National Academy of Sciences, 112(49), 15078-15083. doi.org
  4. National Center for Health Statistics. (n.d.). Health statistics and data. Centers for Disease Control and Prevention. cdc.gov
  5. Link, B. G., and Phelan, J. (1995). Social conditions as fundamental causes of disease. Journal of Health and Social Behavior, 35(extra issue), 80-94.
Key terms
Social gradient in health
The finding that health and life expectancy improve at every step up the socioeconomic distribution rather than only above a threshold of adequacy.
Whitehall studies
Marmot's cohort studies of British civil servants, in which mortality stepped down the employment grade hierarchy among people who were all employed and all covered by the NHS.
Allostatic load
The cumulative physiological wear produced by repeated or sustained stress activation, measured through blood pressure, cortisol regulation, inflammation, and metabolic markers.
Fundamental cause
Link and Phelan's concept of a social condition that commands flexible resources deployable against whatever the current health risks are, so its association with health regenerates as diseases change.
Health selection
Reverse causation in which illness reduces earnings and occupational position, contributing to but not accounting for the observed gradient.
Deaths of despair
Case and Deaton's term for the rise since around 1999 in midlife American mortality from drug overdose, suicide, and alcohol-related liver disease, concentrated among people without a bachelor's degree.
Social causation
The direction of effect in which socioeconomic position produces health outcomes, supported by grade measured before disease onset and by shocks to position.
Preventability gradient
The observation that socioeconomic differences in mortality are steeper for diseases with effective prevention or screening than for those without.

Elites, the Top One Percent, and How Fortunes Are Made

  • Describe the change in top income shares in the United States since 1980 and identify who occupies the top of the distribution.
  • Compare managerial power, superstar, and scale explanations of executive compensation.
  • State Piketty's r greater than g argument precisely and summarise the strongest objections to it.
  • Assess the evidence on elite political influence, including the main criticisms of the studies most often cited.

Two shares that crossed

Using distributional national accounts that allocate all of national income to individuals, Thomas Piketty, Emmanuel Saez, and Gabriel Zucman estimated that the top 1 percent of American adults received about 12 percent of pre-tax national income around 1980 and roughly 20 percent by the mid-2010s. Over the same period the bottom 50 percent's share fell from about 20 percent to roughly 12 percent.

Read those four numbers together. The top hundredth and the bottom half traded places in the national accounts. That is a larger structural change than almost anything else this course measures, and it happened inside forty years.

This lesson asks three questions about the group at the top. Who are they? Where does the money come from? And does concentration at that scale convert into political power?

Who is in the top 1 percent

Start by dispelling the image. The top 1 percent of American households by income is a group of roughly 1.3 million households, and entry requires a household income in the region of half a million dollars or more, depending on the year and the measure. The great majority of them work.

Jon Bakija, Adam Cole, and Bradley Heim used tax data with occupation codes to establish the composition. Executives, managers, and supervisors of non-financial firms make up the largest single block. Financial professionals, including those in banking, asset management, and private equity, are the next largest and are heavily overrepresented at the very top. Physicians, lawyers, and owners of closely held businesses fill much of the rest. In their analysis, executives and financial professionals together accounted for a majority of the top 0.1 percent.

Two implications follow. First, most of the American top is a working top, not a leisured rentier class; salaries, bonuses, and business income dominate. Second, the top is occupationally concentrated in a way that matters for policy, because rules governing corporate pay and financial sector compensation reach a large share of it.

Matthew Smith, Danny Yagan, Owen Zidar, and Eric Zwick complicated that picture usefully. Much top income arrives as profit from pass-through businesses, partnerships and closely held corporations, whose profits are a blend of returns to the owner's own labour and returns to capital. When such an owner retires or dies, the firm's profits typically fall sharply, which suggests a large share of that income was compensation for the person's work rather than a pure return on assets. The old distinction between capitalists and workers does not cut cleanly through this group.

Key idea: The American top 1 percent is dominated by corporate executives, financial professionals, and owner-managers of closely held businesses, and much of its income blends returns to labour and capital in a way that resists the traditional category scheme.

Why executives are paid what they are: three accounts

The Economic Policy Institute's series on chief executive compensation at large American firms puts the ratio to typical worker compensation at roughly 20 to 1 in the mid-1960s and in the region of 300 to 1 in recent years, using a measure that includes realised stock options. However the ratio is constructed, its direction and magnitude are not seriously disputed. The explanation is.

The scale account, associated with Xavier Gabaix and Augustin Landier, is the least dramatic and fits a lot of data. If chief executive talent is scarce and a small difference in ability produces a large difference in outcomes at a large firm, then pay should rise with firm size. Firms grew enormously over this period. On this account executive pay tracks the market value of the firms being run, and the rise is a pricing consequence rather than a governance failure.

The superstar account generalises this. Sherwin Rosen's 1981 analysis of superstar markets showed that when technology allows one performer to serve a vast audience at low marginal cost, small differences in perceived quality produce enormous differences in reward. It was written about opera singers and comedians. It applies to software, finance, and the management of global firms, all of which have seen the reach of a single decision-maker expand dramatically.

The managerial power account, argued by Lucian Bebchuk and Jesse Fried, says the pay-setting process is not arms-length. Boards are frequently populated by people the chief executive influenced the selection of, compensation consultants are retained by the company, and benchmarking against peer medians produces a ratchet, because no board will announce that its chief executive is below average. On this account a substantial part of the rise reflects rent extraction rather than the price of scarce talent.

The evidence does not cleanly select one. Pay does track firm size, supporting the scale account. Pay also rises with luck outside management's control, such as oil prices for oil firms, which is hard to reconcile with pure pricing of talent and easy to reconcile with weak governance. The honest summary is that both mechanisms operate and their relative sizes are unresolved.

Piketty's inequality, stated precisely

Thomas Piketty's Capital in the Twenty-First Century made an argument that is frequently repeated and rarely stated correctly, so here it is carefully.

Let r be the average rate of return on capital and g the growth rate of the economy. Piketty's historical claim is that r has usually exceeded g, with the mid-twentieth century a striking exception produced by two wars, the Depression, inflation, and high taxation. His analytical claim is that when r exceeds g, wealth accumulated in the past grows faster than income and output, so the ratio of capital to national income rises and, because capital ownership is very concentrated, wealth inequality tends to rise with it. He calls this the central contradiction of capitalism, and forecasts a return to a patrimonial society in which inherited wealth dominates unless taxation intervenes.

Now the objections, which are substantial and come from economists broadly sympathetic to the concern.

Matthew Rognlie showed that once depreciation is netted out, the long-run rise in the capital share of income in rich countries is almost entirely accounted for by housing, not by machines, equipment, or intellectual property. If the story is really about housing scarcity in expensive cities, then the remedy is land use and construction policy rather than a global wealth tax, and the mechanism is quite different from the one Piketty describes.

A second objection is technical: for a rising capital-to-income ratio to raise the capital share of income, the elasticity of substitution between capital and labour must exceed one, and most empirical estimates put it below one. If that is right, more capital relative to income lowers the return enough to reduce, not raise, capital's share.

A third is about mechanism. Much of the observed rise in top incomes in the United States has been labour income for working executives, which the r greater than g framework does not directly explain.

What matters here: Piketty's r greater than g is a claim about the tendency of accumulated wealth to outgrow income, not a claim that inequality rises automatically. Rognlie's housing result and the elasticity objection are the two strongest challenges, and both are about mechanism rather than about whether concentration has increased.

Inheritance and the long memory of capital

Piketty's French series on inheritance flows is the clearest long-run measurement available. The annual flow of inherited wealth, as a share of national income, was very high in the nineteenth century, collapsed in the middle of the twentieth as wars and taxation destroyed and redistributed private fortunes, and has been rising since. That trajectory is a reminder that the mid-century compression which many people treat as normal was in fact unusual, produced by specific catastrophes and specific policies.

Tax policy is part of the same story. The top marginal federal income tax rate in the United States was 91 percent through much of the 1950s, 70 percent until 1981, and was cut to 28 percent by the 1986 reform before settling in the high thirties. Effective rates paid were always far below the statutory top rate, because of deductions and because the rate applied only to income above a very high threshold. But the direction is not in doubt, and Saez and Zucman's work links the fall in top rates to the rise in top pre-tax incomes, on the argument that lower rates increase the incentive to bargain hard for compensation.

Gabriel Zucman's separate work on offshore wealth estimates that a substantial share of global household financial wealth, on the order of 8 percent, sits in tax havens and is therefore missing from national statistics. Because that wealth belongs overwhelmingly to the top, standard measures understate concentration.

Does money buy policy?

The most cited study on this is Martin Gilens and Benjamin Page's 2014 analysis, which assembled roughly 1,800 policy questions on which national survey data recorded the preferences of median-income and affluent Americans, and asked which preferences predicted the policy outcome. Their conclusion was that the preferences of economic elites and organised business groups had substantial independent effects, while the estimated independent effect of average citizens' preferences was near zero.

That result deserves both its attention and its criticism.

The criticism is methodological and serious. On the great majority of the policy questions, the preferences of median-income and affluent Americans agree. Where two predictors are highly correlated, separating their independent effects requires the disagreement cases to do all the work, and those cases are a modest subset. Critics including Peter Enns and Omar Bashir have argued that the confidence intervals around the near-zero estimate for average citizens are wide enough that a substantial effect cannot be ruled out, and that the finding is better read as elite preferences prevailing when preferences diverge, which is a narrower claim.

The older sociological literature makes a structural argument that does not depend on that regression at all. C. Wright Mills's The Power Elite traced overlapping leadership across corporations, the military, and the executive branch, with movement between them and shared schooling and clubs. G. William Domhoff extended this with network analyses of interlocking directorates and policy planning organisations. Their claim is not that elites always get their way; it is that the set of options considered is shaped before any public preference is measured, which is a proposition a survey-based design cannot test.

Common misconceptions

  • The top 1 percent is mostly inherited money. In the United States it is predominantly a working top of executives, financial professionals, and owner-managers, though inherited wealth matters far more at the very top of the wealth distribution than of the income distribution.
  • Executive pay rose because of a governance failure alone. Firm size grew enormously and pay tracks size, which supports a pricing story; pay also rises with luck outside management's control, which supports the governance story. Both operate.
  • Piketty says inequality automatically rises. His claim is conditional and about the tendency of accumulated wealth to outgrow income when r exceeds g, and it explicitly treats the mid-century compression as produced by war and taxation.
  • Rognlie refuted Piketty. He showed that the measured rise in capital's share is dominated by housing, which redirects the mechanism and the policy remedy rather than denying the concentration.
  • Gilens and Page proved ordinary voters have no influence. Elite and median preferences coincide on most issues, so the estimate rests on the minority of divergent cases, and the confidence intervals are wide.

Looking back

  • The top 1 percent's share of American pre-tax national income rose from about 12 percent around 1980 to roughly 20 percent by the mid-2010s, while the bottom half's share fell from about 20 percent to about 12 percent.
  • The top is occupationally concentrated in corporate executives, financial professionals, and owner-managers, and much of its income blends labour and capital returns through pass-through businesses.
  • Executive pay ratios rose from roughly 20 to 1 in the mid-1960s to the region of 300 to 1, with scale, superstar, and managerial power accounts each explaining part of it.
  • Piketty's r greater than g predicts accumulated wealth outgrowing income; Rognlie's housing decomposition and the capital-labour substitution elasticity are the strongest objections.
  • Top marginal United States income tax rates fell from 91 percent in the 1950s to 28 percent after 1986, and an estimated 8 percent of global household financial wealth sits offshore and outside official statistics.
  • Gilens and Page found elite preferences predicting policy outcomes where preferences diverge, a real finding whose strength is limited by how rarely elite and median preferences actually differ.

Sources

  1. Piketty, T., Saez, E., and Zucman, G. (2018). Distributional national accounts: Methods and estimates for the United States. Quarterly Journal of Economics, 133(2), 553-609. doi.org
  2. Gilens, M., and Page, B. I. (2014). Testing theories of American politics: Elites, interest groups, and average citizens. Perspectives on Politics, 12(3), 564-581. doi.org
  3. World Inequality Lab. (n.d.). World Inequality Database. wid.world
  4. Internal Revenue Service. (n.d.). Tax statistics. irs.gov
  5. Piketty, T. (2014). Capital in the twenty-first century (A. Goldhammer, Trans.). Cambridge, MA: Harvard University Press.
  6. Mills, C. W. (1956). The power elite. New York: Oxford University Press.
Key terms
Distributional national accounts
A method that allocates the whole of national income to individuals, allowing top shares to be measured consistently with macroeconomic totals.
Working rich
Top earners whose income comes primarily from salaries, bonuses, and active business profits rather than from passive returns on inherited assets.
Pass-through business income
Profits of partnerships and closely held corporations taxed on the owner's return, which blend compensation for the owner's labour with returns to capital.
Superstar market
Rosen's model in which technology lets one performer serve a vast audience, so small quality differences generate enormous reward differences.
Managerial power hypothesis
Bebchuk and Fried's argument that executive pay is set through processes that are not arms-length, producing rent extraction alongside genuine pricing of talent.
r greater than g
Piketty's condition in which the return on capital exceeds economic growth, so accumulated wealth grows faster than income and the capital to income ratio rises.
Rognlie critique
The finding that the long-run rise in the net capital share of income in rich countries is accounted for almost entirely by housing rather than by productive capital.
Offshore wealth
Household financial assets held in tax havens, estimated at around 8 percent of the global total, which are missing from official statistics and belong overwhelmingly to the top.

Module 6: Causes, Policy, and the Argument

What states do about the distribution and why they differ, the dispute over whether globalisation or technology produced the wage curve of the last forty years, and the philosophical argument between equality of opportunity and equality of outcome given at full strength on both sides.

Welfare States: How Much Redistribution, and Why Countries Differ

  • Compare market and disposable income inequality across countries and quantify how much taxes and transfers change the distribution.
  • Summarise Esping-Andersen's three regime types and the main criticisms of the typology.
  • Explain the paradox of redistribution and the political mechanism proposed for it.
  • Distinguish visible spending from tax expenditures and describe how the latter redistributes.

Two countries that start in the same place

Before any tax is collected or any benefit paid, France's Gini coefficient for household income sits around 0.52. The United States, measured the same way on market income, sits around 0.51. The two countries generate their pre-government distributions at almost the same level of inequality.

After taxes and transfers, France is near 0.29 and the United States near 0.39.

Hold that comparison, because it disposes of the most common story told about why the two countries differ. The difference is not that French markets produce more equal outcomes. It is what happens next.

Measuring what a state does

The standard measure is the redistributive effect: the percentage reduction from the market income Gini to the disposable income Gini. On that measure the United States reduces inequality by roughly a fifth to a quarter, while France, Belgium, Finland, and several others reduce it by around two fifths.

Two components produce the reduction, and it is important to keep them apart.

Transfers do most of the work almost everywhere. Pensions, unemployment benefit, disability payments, child benefits, and social assistance are large flows aimed at people with low market income, and they move the distribution substantially.

Taxes do less than people expect. A progressive income tax reduces inequality, but income tax is only part of the revenue system, and payroll and consumption taxes are usually flat or regressive. In cross-national comparisons, the bulk of measured redistribution runs through the spending side.

There is also a large component that these Gini comparisons miss entirely: in-kind services. Public health care, subsidised childcare, and free education transfer real resources without appearing as income. A British household and an American household with identical disposable incomes are not equally well off if one faces no medical bills. Analyses that impute the value of in-kind services find that they reduce measured inequality further, and by more in countries with large public service sectors.

Key idea: Rich countries differ far less in the inequality their markets produce than in how much of it their governments undo. Most measured redistribution runs through transfers and services rather than through the tax schedule.

Esping-Andersen's three worlds

The organising typology in this field is Gosta Esping-Andersen's, published in 1990. He argued that welfare states do not lie on a single scale from small to large; they cluster into qualitatively different regimes, distinguished by how much they free people from dependence on selling their labour, a property he called decommodification, and by what pattern of stratification they themselves create.

RegimeTypical casesLogicCharacteristic effect
LiberalUnited States, United Kingdom, Canada, AustraliaModest, often means-tested benefits; the market is the default providerLow decommodification; a sharp line between recipients and taxpayers
Conservative or corporatistGermany, France, Austria, ItalySocial insurance tied to employment status, with the family assumed as carerPreserves status differences across occupational groups; historically weak support for maternal employment
Social democraticSweden, Denmark, NorwayUniversal benefits and extensive public services at a middle-class standardHigh decommodification; broad political coalitions because the middle class are beneficiaries too

The typology has been criticised productively rather than discarded. Feminist scholars including Ann Orloff and Jane Lewis pointed out that decommodification measures freedom from the labour market and says nothing about freedom from unpaid care obligations, and proposed defamilialisation as the missing dimension: the extent to which a person can maintain an independent household without depending on family. On that axis the conservative regimes look very different, because their generosity to male breadwinners coexisted with policies that kept women at home.

Others argued that Southern European countries with strong family obligation and weak public provision, and East Asian systems built around employment and family, do not fit any of the three. Esping-Andersen accepted some of these amendments in later work.

The paradox of redistribution

Here is a result that reverses an intuition almost everyone holds.

If a government has a fixed budget to reduce poverty, targeting seems obviously more efficient than universality. Why send a child benefit to a barrister? Give everything to those who need it and the same money goes further.

Walter Korpi and Joakim Palme tested this across countries and found the opposite. Countries with the most narrowly targeted, means-tested systems achieved less poverty reduction and less inequality reduction than countries with broad, universal, earnings-related systems. They called it the paradox of redistribution: the more you target benefits at the poor, the less redistribution you achieve.

The mechanism is political rather than arithmetic. Universal programmes have universal constituencies. When the middle class receives the pension, the child benefit, and the health service, it defends them at the ballot box and tolerates the taxes that fund them, so budgets are large. Narrowly targeted programmes serve a politically weak minority, are easier to cut, are held to lower quality standards, and carry stigma that depresses take-up. A perfectly targeted programme with a small and shrinking budget redistributes less than a leaky universal one with a large and defended budget.

Two honest qualifications. Later work by Ive Marx and colleagues found the paradox has weakened in recent decades, as some countries have combined large budgets with more targeting within them. And the mechanism depends on political conditions that vary; it is a claim about coalitions, not a law.

The point: The efficiency of a transfer system cannot be evaluated by looking at the transfer alone, because the design determines the size and durability of the budget. Targeting is efficient per dollar and often produces fewer dollars.

The welfare state you cannot see

Comparisons of social spending as a share of GDP put the United States well below France, Finland, or Denmark. Those comparisons are correct and incomplete, because a great deal of American social policy operates through the tax code rather than through spending.

Christopher Howard's term for this is the hidden welfare state. The exclusion of employer-paid health insurance premiums from taxable income, the mortgage interest deduction, and the tax treatment of retirement saving are all enormous programmes that never appear as expenditure. They are counted as revenue not collected.

Their distributional profile is the opposite of ordinary transfers. A deduction is worth more to a household in a higher marginal tax bracket, and it is worth nothing to a household with no tax liability. A subsidy for mortgage interest goes to owners rather than renters, and rises with the size of the mortgage. So the visible American welfare state flows downward and the hidden one flows upward, and comparing only the visible half misdescribes what the state is doing.

The Earned Income Tax Credit is the important exception: it is delivered through the tax system and is strongly targeted at low-income working families with children, which is why it appears in the Supplemental Poverty Measure and not in the official one.

Does redistribution cost growth?

The classic statement of the worry is Arthur Okun's leaky bucket: money carried from rich to poor arrives diminished, because taxes distort work and investment decisions and administration consumes resources. Okun's own position was that the bucket leaks and the transfer is often still worth making, which is a more careful position than either side usually attributes to him.

Peter Lindert examined the historical record across two centuries and posed what he called the free lunch puzzle: the countries that built the largest welfare states did not grow more slowly than those that did not. His explanation is that large welfare states were built with tax and spending designs that limit distortion, financing generous benefits with broad consumption and payroll taxes rather than with punishing marginal rates on capital, and spending on things that raise productivity such as health and childcare that lets parents work.

The counter-case is not silly and should be stated properly. Cross-country growth comparisons are confounded by dozens of factors, the Nordic countries were unusually institutionally capable before their welfare states were large, and specific programmes clearly do reduce labour supply at the margin, most visibly in early retirement and disability pathways. Nobody serious claims that transfers have no behavioural effects. The claim is about the size of those effects relative to the redistribution achieved, which is a quantitative question rather than a matter of principle.

Common misconceptions

  • European markets produce more equal outcomes than American ones. Market income Ginis are similar; the difference is almost entirely in taxes and transfers.
  • Redistribution is mostly done by progressive taxes. Transfers and public services do most of the work, and consumption and payroll taxes are usually flat or regressive.
  • Targeting benefits at the poor maximises poverty reduction. Korpi and Palme found the opposite across countries, because universal programmes command coalitions that sustain much larger budgets.
  • The United States has a small welfare state. It has a modest visible one and a very large hidden one delivered through tax expenditures, which flows up the distribution rather than down.
  • Large welfare states must reduce growth. The historical record shows no such relationship, though specific programmes do have measurable labour supply effects at the margin.

What to remember

  • France and the United States have similar market income Ginis near 0.5 and very different disposable income Ginis, roughly 0.29 against 0.39.
  • Redistributive effect is the reduction from market to disposable Gini: about a fifth to a quarter in the United States and around two fifths in several European countries.
  • Transfers and in-kind services do most of the redistributing; tax schedules do less than commonly assumed.
  • Esping-Andersen's liberal, conservative, and social democratic regimes differ in decommodification and in the stratification they themselves create; defamilialisation was the main missing dimension.
  • The paradox of redistribution holds that narrowly targeted systems redistribute less, because universal programmes sustain larger and better-defended budgets.
  • Tax expenditures such as the employer health exclusion and the mortgage interest deduction constitute a hidden welfare state whose benefits rise with income.

Sources

  1. Organisation for Economic Co-operation and Development. (n.d.). Inequality and income distribution. oecd.org
  2. Luxembourg Income Study. (n.d.). LIS cross-national data center. lisdatacenter.org
  3. Our World in Data. (n.d.). Government spending. ourworldindata.org
  4. Korpi, W., and Palme, J. (1998). The paradox of redistribution and strategies of equality: Welfare state institutions, inequality, and poverty in the Western countries. American Sociological Review, 63(5), 661-687. DOI 10.2307/2657333.
  5. Esping-Andersen, G. (1990). The three worlds of welfare capitalism. Princeton, NJ: Princeton University Press.
  6. Howard, C. (1997). The hidden welfare state: Tax expenditures and social policy in the United States. Princeton, NJ: Princeton University Press.
Key terms
Redistributive effect
The percentage reduction from a country's market income Gini to its disposable income Gini, which measures how much taxes and transfers change the distribution.
Decommodification
Esping-Andersen's measure of how far a welfare state allows a person to maintain a livelihood without depending on selling their labour.
Defamilialisation
The dimension proposed by feminist critics measuring how far a person can maintain an independent household without depending on family support or unpaid care.
Paradox of redistribution
Korpi and Palme's finding that narrowly targeted benefit systems achieve less poverty and inequality reduction than universal ones, because of the political coalitions each sustains.
Tax expenditure
Revenue forgone through a deduction, exclusion, or credit, which functions as a spending programme but appears nowhere in expenditure statistics.
Hidden welfare state
Howard's term for social policy delivered through the tax code, whose benefits generally rise with a household's marginal tax rate and therefore with income.
In-kind transfer
A public service such as health care, childcare, or education supplied directly rather than as cash, which raises real living standards without appearing in income statistics.
Leaky bucket
Okun's image for the losses incurred when resources are transferred from rich to poor, which he regarded as real and often worth bearing.

Globalisation Against Technology: Two Explanations for One Curve

  • State the skill-biased technological change account and the race between education and technology, and identify what each predicts.
  • Summarise the China shock estimates and explain what a local labour market design can and cannot show about aggregate effects.
  • Evaluate the timing and polarisation objections to the technology account.
  • Explain why cross-country comparison is the sharpest available test between the two explanations.

A quarter of the decline

In 2013 David Autor, David Dorn, and Gordon Hanson published an estimate that has shaped the argument ever since. Comparing American local labour markets according to how exposed their industry mix was to rising Chinese import competition, they concluded that this competition accounted for roughly a quarter of the decline in United States manufacturing employment between 1990 and 2007.

That single number arrived in the middle of a long-running dispute. For twenty years before it, most economists had held that trade explained only a small share of rising wage inequality and that technology explained most of it. This lesson lays out both cases properly, because the dispute is genuine and unresolved, and because it is a good example of how two well-evidenced explanations can compete for the same phenomenon.

What has to be explained

Set the facts down before the theories.

The college wage premium in the United States, the ratio of average earnings of college graduates to high school graduates, grew substantially from around 1980 through the 2000s. Wage dispersion widened at the top and, in the 1980s, at the bottom. Manufacturing employment fell from its 1979 peak of about 19.5 million to about 12.8 million by 2019. And, as an earlier lesson established, employment growth was U-shaped rather than uniformly tilted toward skill.

Any successful explanation has to fit all four, including the last one, which is where several accounts fail.

The technology case

The canonical statement is Lawrence Katz and Kevin Murphy's, framed as a race. On the demand side, technological change raises the relative demand for skilled labour steadily over time. On the supply side, the share of the workforce with a college education grows. Whether the skill premium rises or falls depends on which runs faster.

The story then writes itself from American data. College attainment grew rapidly through the 1970s, and the premium fell. Growth in college completion slowed markedly after about 1980, while demand kept rising, and the premium rose. Claudia Goldin and Lawrence Katz developed this into a full account of the twentieth century in The Race Between Education and Technology, arguing that the mid-century compression of American wages was produced by an unusually fast expansion of high school and then college education, and that the late-century widening followed the slowdown in that expansion.

This account has real strengths. It is quantitative, it fits the long-run American series, and the mechanism is observable: firms did adopt computing on a very large scale, and the tasks that computing replaced were disproportionately performed by workers without degrees.

The objections, which are strong

David Card and John DiNardo pressed the objection that has never been fully answered. Computerisation proceeded steadily from the 1970s onward, but the growth in American wage inequality was episodic, concentrated heavily in the early 1980s and much slower afterward. A continuous cause does not naturally explain a discontinuous effect. They also noted that if computers rewarded skill in general, the gender wage gap should have moved with the skill premium, and it did not in the way the theory implies.

The second objection is polarisation. The canonical skill-biased account predicts that demand shifts monotonically toward more skill, so employment should grow at the top and shrink toward the bottom. What actually happened was growth at both ends and hollowing in the middle. The routine-biased refinement developed by David Autor, Frank Levy, and Richard Murnane fixes this, by making the relevant distinction routine against non-routine rather than skilled against unskilled. That is a genuine improvement, and it is worth noticing that it was forced by the failure of the earlier version.

The third objection is the one that does the most work in this lesson, and we will come back to it: the same technologies were adopted across all rich countries, and wage inequality rose by very different amounts in each.

The trade case, old and new

The theoretical prediction is old. The Stolper-Samuelson theorem implies that when a skill-abundant country trades more with labour-abundant countries, the relative wage of less-skilled workers in the skill-abundant country should fall.

The early 1990s assessments concluded the effect was modest, on the order of 10 to 20 percent of the rise in the skill premium. The reasoning was that imports from low-wage countries were a small share of American GDP, and that the price movements the theory requires were not clearly visible.

The China shock literature changed the terms of that assessment in two ways.

First, the magnitude. China's entry into world manufacturing after its 1990s reforms and its 2001 accession to the World Trade Organization was larger and faster than anything the earlier literature had contemplated.

Second, and more important, the level of analysis. Autor, Dorn, and Hanson compared local labour markets rather than the nation, using the fact that a commuting zone specialising in furniture or textiles was heavily exposed while one specialising in aerospace or medical devices was not. In heavily exposed places, manufacturing employment fell more, wages fell, unemployment and disability enrolment rose, and, crucially, the adjustment did not happen quickly. Standard trade models assume that displaced workers move to other sectors or other regions. In the data they largely did not, and the local effects persisted for more than a decade.

Key idea: The China shock literature's contribution is less about the aggregate size of trade effects than about the speed and geography of adjustment. The costs of trade were real, concentrated in particular places, and far more persistent than the models had assumed.

Where the trade case is weakest

A local labour market design compares places within one country. It therefore cannot directly measure the national effect, because whatever offsetting gains occurred elsewhere, in export industries, in lower consumer prices, in firms whose input costs fell, are partly absorbed into the comparison rather than measured by it. Economists who have tried to build the general equilibrium picture, including work emphasising employment supported by export growth, reach smaller net national numbers than the local estimates might suggest.

Autor and his co-authors have generally been careful about this distinction. Public discussion of their work has been much less careful.

The global picture, and the elephant

Christoph Lakner and Branko Milanovic plotted income growth by position in the global distribution between 1988 and 2008. The resulting curve, which resembles an elephant in profile, showed strong growth for people around the middle of the world distribution, largely industrialising Asia, strong growth at the global top, and weak growth around the 80th to 90th global percentile, a band that contains much of the working class of rich countries.

The chart became a Rorschach test. One reading is that globalisation raised hundreds of millions out of poverty at the cost of stagnation for rich-country workers. Another is that the flat segment reflects Japan and the post-communist transition countries rather than a general rich-country phenomenon, and that later revisions using different country coverage flatten the trunk considerably. Both readings are defensible, which is a reason to be careful with a chart that has been used to settle arguments it cannot settle.

The test that discriminates

Here is the comparison that does the most to separate the two accounts.

Germany imported the same industrial robots as the United States, adopted the same enterprise software, and faced the same Chinese competition, in some sectors more intensely, given its manufacturing specialisation. If technology or trade were sufficient explanations, German wage inequality should have risen comparably. It rose considerably less, and its manufacturing employment held up better.

What differed was institutional: sectoral collective bargaining covering a large share of employment, works councils with statutory rights over restructuring, apprenticeship systems that move young people into skilled positions, and short-time work schemes that hold employment relationships together through demand shocks.

The lesson is not that technology and trade do not matter. Both plainly do; Daron Acemoglu and Pascual Restrepo's study of American industrial robots estimated that each additional robot per thousand workers reduced the local employment-to-population ratio by around 0.2 percentage points and wages by a few tenths of a percent, which is a real effect measured with a credible design. The lesson is that a shock and its distributional consequences are two different things, and the second depends on the institutions the shock lands on.

The upshot: Technology and trade are shocks whose distributional consequences are mediated by bargaining institutions, education systems, and social policy. Countries facing the same shocks with different institutions got different distributions, which is the strongest single piece of evidence that neither shock is a sufficient explanation on its own.

Common misconceptions

  • Economists have settled that technology, not trade, caused rising inequality. That was the mainstream position in the 1990s and the China shock literature substantially revised it, particularly on the geography and persistence of adjustment.
  • The China shock estimate is a national jobs figure. It is identified from differences between local labour markets, and the aggregate national effect requires general equilibrium reasoning that yields smaller numbers.
  • Skill-biased technological change explains job polarisation. The canonical version predicts a monotonic shift toward skill; the U-shaped pattern required the routine-biased refinement.
  • Because rich-country workers stagnated, globalisation failed. The same period saw the largest reduction in global poverty ever recorded, and both facts are part of the same accounting.
  • If technology caused it, policy cannot change it. Countries facing identical technologies had very different distributional outcomes, which is the clearest evidence that institutions are doing real work.

Pulling it together

  • Autor, Dorn, and Hanson attributed roughly a quarter of the 1990 to 2007 decline in American manufacturing employment to Chinese import competition, with slow and persistent local adjustment.
  • The Katz and Murphy race between education and technology fits the long-run American skill premium series, with the premium rising once college completion growth slowed after 1980.
  • Card and DiNardo objected that computerisation was continuous while inequality growth was episodic, and the canonical account also fails to predict polarisation without the routine-biased refinement.
  • Local labour market designs identify concentrated and persistent costs but cannot by themselves measure aggregate national effects.
  • The elephant curve shows strong growth for the global middle and the global top with weak growth around the 80th to 90th global percentile, and its interpretation is sensitive to country coverage.
  • Germany faced comparable technology and trade shocks with much smaller increases in wage inequality, which points at bargaining institutions, apprenticeships, and short-time work as the mediating factor.

Sources

  1. Autor, D. H., Dorn, D., and Hanson, G. H. (2013). The China syndrome: Local labor market effects of import competition in the United States. American Economic Review, 103(6), 2121-2168. doi.org
  2. Card, D., and DiNardo, J. E. (2002). Skill-biased technological change and rising wage inequality: Some problems and puzzles. Journal of Labor Economics, 20(4), 733-783. doi.org
  3. Acemoglu, D., and Restrepo, P. (2020). Robots and jobs: Evidence from US labor markets. Journal of Political Economy, 128(6), 2188-2244. doi.org
  4. Our World in Data. (n.d.). Global economic inequality. ourworldindata.org
  5. Goldin, C., and Katz, L. F. (2008). The race between education and technology. Cambridge, MA: Belknap Press of Harvard University Press.
Key terms
Skill-biased technological change
The account in which technology raises relative demand for skilled labour, so the skill premium rises whenever demand growth outruns the growth of educated supply.
Race between education and technology
Goldin and Katz's framing in which the skill premium is determined by whether educational attainment or skill demand grows faster.
Routine-biased technical change
The refinement in which technology substitutes for codifiable tasks rather than for unskilled labour generally, which is what predicts polarisation.
Stolper-Samuelson theorem
The trade theory result implying that opening to trade with labour-abundant countries lowers the relative wage of less-skilled workers in skill-abundant countries.
China shock
The rapid expansion of Chinese manufacturing exports from the 1990s, whose effects were measured across differently exposed American local labour markets.
Local labour market design
Identification from differences in exposure across regions within a country, which measures concentrated effects well and aggregate national effects poorly.
Elephant curve
Lakner and Milanovic's plot of global income growth by position between 1988 and 2008, showing strong gains for the global middle and top and weak gains around the 80th to 90th percentile.
Short-time work
A policy that subsidises reduced hours instead of layoffs during demand shocks, used in Germany and credited with preserving employment relationships through downturns.

Equality of Opportunity Against Equality of Outcome

  • Distinguish formal from substantive equality of opportunity and state Roemer's circumstances-and-effort formalisation.
  • Reconstruct the strongest case for Rawls's difference principle and for Nozick's entitlement theory, including each one's best reply to the other.
  • Explain why equality of opportunity and equality of outcome cannot be pursued independently across generations.
  • Separate the parts of this dispute that evidence can settle from the parts it cannot.

Fifty-nine, eighty-four, thirty-two

Michael Norton and Dan Ariely surveyed 5,522 Americans in 2011. They asked respondents to estimate how the country's wealth was actually distributed across the five quintiles. The median respondent estimated that the top fifth held about 59 percent of it. The actual figure at the time was around 84 percent.

Then they asked a second question. Shown unlabelled distributions and asked which country they would rather join, not knowing their own position, respondents chose one in which the top fifth held about 32 percent. That preference held across income levels and across party identification, with differences of degree rather than of direction.

Three numbers, and each does different work. People underestimate concentration. People prefer less of it than they think exists. And people do not prefer equality: 32 percent for the top fifth is well above the 20 percent that an even distribution would give. Whatever most people want, it is neither the status quo nor equal shares.

This lesson is about the argument underneath that finding, and it is a real argument between positions held by serious people, not a puzzle with a solution at the back.

Two things called equality of opportunity

Before the dispute, a distinction, because the phrase is used for two quite different standards and a great deal of confusion follows from mixing them.

Formal equality of opportunity means careers open to talents: positions are advertised, selection is on relevant qualifications, and no one is excluded by law or by policy on grounds of birth, race, sex, or religion. This is the standard of anti-discrimination law, and in most rich countries it is substantially achieved as a matter of legal form.

Substantive equality of opportunity asks more. In John Rawls's formulation, those with the same native talent and the same willingness to use it should have the same prospects of success, regardless of the social class into which they were born. That is a far more demanding standard, and no society has come close to meeting it, because prospects are shaped by prenatal health, early language environment, school quality, neighbourhood, and the networks a family can supply.

John Roemer gave the second standard a formal shape that has been used in empirical work. Divide everything that determines an outcome into circumstances, for which a person cannot be held responsible, such as parental education, race, and place of birth, and effort, for which they can. Equality of opportunity obtains when outcomes do not depend on circumstances. Roemer's operational move was to define effort as a person's rank within their own circumstance group, which lets researchers estimate how much of the variance in outcomes is attributable to circumstances.

The move exposes the problem it was designed to manage. Effort is itself shaped by circumstance. The capacity for persistence, the belief that persistence pays, the health to sustain it, and the information about what to persist at are all distributed unequally by exactly the factors sitting in the circumstances column. Push the analysis and the effort column shrinks. How far you push it is a philosophical choice, not a statistical one.

Key idea: Formal equality of opportunity forbids exclusion. Substantive equality of opportunity requires equal prospects for equal talent and willingness. The second is far more demanding, and formalising it exposes that effort itself has causes.

The case for caring about outcomes

Rawls's argument in A Theory of Justice begins with a device. Imagine choosing the rules of a society from behind a veil of ignorance, not knowing your own class, talents, race, sex, or conception of the good life. What principles would a rational chooser accept?

Rawls argues they would accept equal basic liberties for all, and would permit inequalities only where two conditions hold: positions must be open to all under fair equality of opportunity, and the inequalities must work to the greatest benefit of the least advantaged. That second condition is the difference principle, and it is more permissive than it sounds. It does not require equal shares. It permits very large inequalities if they genuinely make the worst-off better off than any alternative arrangement would.

The deepest part of Rawls's case is his treatment of talent. He argues that the distribution of natural abilities is arbitrary from a moral point of view. Nobody deserves the genetic endowment they were born with, and nobody deserves the family that cultivated it. Rawls's own phrase is that the natural distribution is neither just nor unjust; what is just or unjust is how institutions deal with it. From that it follows that the rewards flowing to talent are not deserved in any deep sense, though they may be legitimately expected under rules that everyone benefits from.

Amartya Sen accepts much of this and asks a further question: equality of what? Equal incomes leave a disabled person worse off, because converting income into functioning costs more. Sen's answer is to focus on capabilities, the real freedoms people have to do and be things they have reason to value, which shifts attention from what people hold to what they can do with it.

The case against outcome-based standards

Robert Nozick's Anarchy, State, and Utopia is the strongest reply, and it deserves to be stated at full force rather than caricatured.

Nozick's entitlement theory says a distribution is just if it arose justly, full stop. Three principles govern this: justice in original acquisition, justice in transfer, and rectification of past violations of the first two. There is no pattern that a just distribution must match, because justice is historical rather than structural.

His argument for this is the Wilt Chamberlain example. Start from any distribution you consider just, including a perfectly equal one; call it D1. Now suppose a million people each voluntarily pay twenty-five cents extra to watch one basketball player, who ends up far richer than everyone else. Call the result D2. Nozick's question is sharp: if D1 was just, and every transfer was voluntary and involved nothing anyone was not entitled to give, on what ground is D2 unjust? And if you propose to restore D1, you must continuously prevent adults from doing what they wish with their own resources. His conclusion is that maintaining any patterned principle requires continuous interference with liberty.

Two things about this argument are usually missed. It is not an argument that the poor deserve their position; it is an argument about what makes a distribution just, and it applies equally to distributions that Nozick's political allies would dislike. And Nozick's third principle, rectification, is a substantial concession: if historical acquisitions were unjust, as in conquest, slavery, and expropriation, then the present distribution inherits that injustice, and correcting it may require large transfers. Nozick did not develop this principle, and a serious application of it to the American record would be radical.

Two further positions sharpen the field. Ronald Dworkin distinguished brute luck, which you did not choose, such as being born with a serious illness, from option luck, the outcome of a gamble you chose to take, and argued that a just society compensates for the first but not the second. Harry Frankfurt argued that equality is not what matters morally at all: what matters is that everyone has enough. On the sufficiency view, the gap between a comfortable person and a billionaire is not itself a moral problem, and treating it as one distracts from the people who lack enough.

Where the argument actually bites

Here is why these positions cannot simply be kept in separate boxes, and it is empirical rather than philosophical.

Today's outcomes are tomorrow's circumstances. The income distribution of one generation becomes the parental income distribution of the next, and everything in the mobility lesson shows that parental position transmits. The Great Gatsby Curve says the two are correlated across countries. The neighbourhood and education lessons supply the mechanisms.

So a policy of pure equality of opportunity, indifferent to outcomes within a generation, cannot be sustained across generations, because the outcome inequality it permits becomes the circumstance inequality it forbids. Anyone defending an opportunity standard has to say what they propose to do about that, and inheritance taxation, early childhood policy, and school finance are the usual answers.

The argument runs in the other direction too. A policy aimed purely at compressing outcomes has to answer Nozick's question about voluntary exchange, and has to reckon with incentive effects that the welfare state lesson quantified. Neither pure position survives contact with the other's strongest point.

Worth holding on to: Equality of opportunity and equality of outcome are not independent goals that a society can trade off freely, because the outcomes of one generation are the circumstances of the next. That is a finding about transmission, not a philosophical claim.

What people actually object to

A line of experimental work complicates the public opinion picture usefully. Christina Starmans, Mark Sheskin, and Paul Bloom reviewed evidence from laboratory studies and from developmental psychology and argued that people are not in fact bothered by inequality as such; they are bothered by unfairness. In experiments where unequal rewards track effort or contribution, both children and adults accept and often prefer the unequal distribution. Where the same inequality arises from an arbitrary or rigged process, they reject it.

If that reading is right, much public argument is misdescribed. The dispute is less about how large a gap should be and more about whether the process producing it was one anyone could accept, which is why procedural questions such as inherited advantage, admissions preferences, and lobbying generate more heat than the Gini coefficient does.

Michael Young saw this coming from a different direction. His 1958 book The Rise of the Meritocracy coined the word as a satire, not a compliment: he imagined a society that had genuinely sorted people by measured ability and found it crueller than the one it replaced, because those at the bottom could no longer tell themselves the sorting was unfair. A successful meritocracy, on Young's account, converts inequality into a verdict on the person.

What evidence can and cannot do

Close by separating the two kinds of question, because this is the discipline the whole course has been building.

Evidence can tell you how much of the variance in outcomes is attributable to circumstances, how large the mobility coefficients are, whether a given policy raised employment or lowered child poverty, and what a proposed reform would cost. Those are hard empirical questions with answers, and this course has spent fifteen lessons on them.

Evidence cannot tell you how much inequality is acceptable, how to weigh liberty against distribution, whether the untalented deserve less, or whether a gap that harms nobody is nonetheless objectionable. Those are questions about value, and no dataset settles them.

What evidence does do is constrain the arguments. Someone who defends existing inequality on the ground that it reflects effort has to reckon with the measured size of circumstance effects. Someone who proposes to eliminate inequality has to reckon with measured incentive responses and with Nozick's question. The facts do not decide, and they do rule things out.

Common misconceptions

  • Rawls demanded equal incomes. The difference principle permits large inequalities where they genuinely improve the position of the worst off.
  • Nozick argued that the rich deserve their wealth. His theory is historical rather than desert-based, and his rectification principle concedes that unjust acquisition taints present holdings.
  • Equality of opportunity is the moderate position that avoids the argument. Substantive equality of opportunity is extremely demanding, and it cannot be sustained across generations without addressing outcomes.
  • Public opinion favours equal distribution. Norton and Ariely's respondents chose a distribution giving the top fifth about 32 percent, well above equal shares and well below the actual figure.
  • People object to inequality itself. Experimental evidence suggests they object to unfair processes, and accept unequal rewards that track effort or contribution.

Where this leaves us

  • Norton and Ariely's respondents estimated the top fifth held 59 percent of American wealth against an actual 84 percent, and preferred a distribution giving it about 32 percent.
  • Formal equality of opportunity forbids exclusion; substantive equality of opportunity requires equal prospects for equal talent and willingness, and no society approaches it.
  • Roemer's split into circumstances and effort makes the standard measurable and exposes that effort itself is shaped by circumstance.
  • Rawls permits inequality only where it benefits the least advantaged, and denies that natural talent is deserved; Sen redirects the question to capabilities.
  • Nozick's entitlement theory makes justice historical rather than patterned, and the Wilt Chamberlain argument shows that maintaining any pattern requires continuous interference with voluntary exchange.
  • The two standards cannot be pursued independently, because one generation's outcomes are the next generation's circumstances, which is an empirical finding rather than a philosophical claim.

Sources

  1. Norton, M. I., and Ariely, D. (2011). Building a better America, one wealth quintile at a time. Perspectives on Psychological Science, 6(1), 9-12. doi.org
  2. Starmans, C., Sheskin, M., and Bloom, P. (2017). Why people prefer unequal societies. Nature Human Behaviour, 1, 0082. doi.org
  3. Stanford Encyclopedia of Philosophy. (n.d.). Equality of opportunity. Stanford University. plato.stanford.edu
  4. Stanford Encyclopedia of Philosophy. (n.d.). Equality. Stanford University. plato.stanford.edu
  5. Rawls, J. (1971). A theory of justice. Cambridge, MA: Harvard University Press.
  6. Nozick, R. (1974). Anarchy, state, and utopia. New York: Basic Books.
Key terms
Formal equality of opportunity
Careers open to talents: positions advertised, selection on relevant qualifications, and no exclusion by law or policy on grounds of birth.
Substantive equality of opportunity
The demanding standard on which people of equal talent and equal willingness should have equal prospects regardless of social origin.
Circumstances and effort
Roemer's division of the determinants of an outcome into factors a person cannot be held responsible for and factors they can, which makes opportunity measurable.
Veil of ignorance
Rawls's device for choosing principles of justice without knowing one's own class, talents, or conception of the good.
Difference principle
Rawls's condition that inequalities are permissible only where they work to the greatest benefit of the least advantaged members of society.
Entitlement theory
Nozick's account on which a distribution is just if it arose through just acquisition and just transfer, with rectification for past violations, and no pattern is required.
Brute and option luck
Dworkin's distinction between unchosen misfortune, which a just society compensates, and the outcome of a gamble deliberately taken, which it need not.
Sufficiency doctrine
Frankfurt's position that what matters morally is that everyone has enough, so gaps above the sufficiency threshold are not themselves objectionable.
Capability approach
Sen's redirection of the question from what people hold to what they are actually free to do and be, which accounts for unequal conversion of resources into functioning.

Open the interactive version with quizzes and progress →