Module 1: What Policy Is and Who Makes It
Public policy defined by what governments actually do and decline to do, then the machinery that produces it: a deliberately fragmented federal system, Congress and the presidency, the career bureaucracy, the courts, and the outside industry of interest groups, think tanks, and media. Closes each lesson by testing the tidy textbook model against how decisions really get made.
What Public Policy Is, and the Stages Heuristic
- Define public policy in a way that includes inaction and routine administration, not just landmark laws.
- Distinguish the substantive types of policy and explain why the type predicts the politics.
- Lay out the stages heuristic and state precisely what it clarifies and what it distorts.
The big picture
In 1962 a pharmacologist at the Food and Drug Administration named Frances Kelsey refused, six times, to approve an application for a sedative called thalidomide. The drug was already sold across Europe and the company was impatient. Kelsey kept asking for data on whether it crossed the placenta. She never got a satisfying answer, and while she stalled, reports arrived from West Germany of thousands of infants born with severe limb malformations. The United States was largely spared, and Congress responded by passing amendments that required drug makers to prove effectiveness, not merely safety, before marketing.
Notice how much policy is packed into that story, and how little of it looks like the thing you were probably taught to watch. There is a statute, yes, passed by Congress and signed by a president. But before the statute there was a career scientist exercising discretion inside an agency, and her decision was itself policy: a binding choice about what the government would and would not permit. There was a decision by the same agency in earlier decades to build a review process at all. There were companies, physicians, journalists, and a grieving public. And crucially, there was a period in which the United States government had no rule requiring proof of effectiveness, and that absence was also a policy, one with consequences.
This course is about how choices like that get made and how you analyze them. It is not about which choices are correct. That distinction will matter on every page, so let us fix it now. An analyst's job is to clarify what is at stake, lay out the real options, forecast what each would do, say honestly how confident anyone can be, and show who wins and who loses. The job is not to tell you what to want. Where this course covers contested ground, and most of it is contested, you will get the strongest version of each side and an honest account of where the evidence stops.
Defining public policy
The most durable definition in the field comes from Thomas Dye, who wrote that public policy is whatever governments choose to do or not to do. It sounds almost too simple. It earns its keep in two ways. First, it locates policy in action rather than intention. A legislature can announce a magnificent goal in a preamble, but if no money is appropriated, no agency is assigned, and no rule is written, the policy is the announcement and nothing more. Second, and more importantly, it counts inaction. When a government declines to regulate a chemical, declines to enforce a law already on the books, or declines to fund a program it authorized, it has made a policy choice with real distributional consequences.
Other definitions add useful nuance. James Anderson described policy as a purposive course of action followed by an actor or set of actors in dealing with a problem. The word purposive matters: policy is not a single decision but a pattern of them, a course. Charles Lindblom emphasized that policy is usually not chosen once and executed, but adjusted repeatedly by many hands, so that the operative policy at any moment is the accumulated residue of thousands of small decisions by legislators, budget officers, regulators, prosecutors, and clerks.
Put these together and a working definition emerges. Public policy is the pattern of authoritative action and inaction by which a government addresses, or declines to address, a public problem. Authoritative means backed by the state's power to bind. Pattern means it accumulates over time rather than arriving whole. And public problem is doing quiet, contested work in that sentence, because whether something counts as a public problem at all is one of the fiercest fights in politics. We will spend a full lesson on it in Module 3.
Key idea: Policy is what governments actually do and decline to do, accumulated across many decisions and many hands. Announcements without money, staff, and rules are not yet policy, and deliberate non-action is a policy choice with real consequences.
Instruments: the verbs of policy
Governments have a limited vocabulary of ways to act, and learning it early will make later lessons much easier. Scholars call these policy instruments or policy tools, and a common grouping distinguishes five.
| Instrument | How it works | Example |
|---|---|---|
| Regulation | Commands or forbids behavior, backed by penalties | Emissions limits on power plants; minimum wage law |
| Taxation and subsidy | Changes prices so the desired behavior becomes cheaper | Fuel taxes; the mortgage interest deduction; clean energy credits |
| Direct provision | Government produces the good or service itself | Public schools; the interstate highway system; military defense |
| Transfers and insurance | Moves money or risk to individuals | Social Security; unemployment insurance; food assistance |
| Information and persuasion | Changes what people know or notice | Nutrition labels; public health campaigns; disclosure rules |
Any given problem can usually be attacked with several of these, and a large part of practical analysis consists of noticing that fact. Suppose the goal is fewer traffic deaths. You could regulate (require seat belts and airbags, lower speed limits), tax (raise fuel taxes so people drive less), provide directly (build safer roads and better transit), transfer (subsidize vehicle safety upgrades), or inform (crash test ratings, public campaigns). These options differ enormously in cost, in who bears the burden, in how visible the government's hand is, and in how easily they can be undone by a later administration. A common beginner's error is to specify the goal and then leap to one instrument without noticing the others were available.
Types of policy, and why type predicts politics
In 1964 the political scientist Theodore Lowi made an observation that reversed the usual assumption. Everyone assumed politics produces policy. Lowi argued that policy also produces politics: the kind of policy at stake shapes the kind of political fight you get. His categories still organize the field.
Distributive policies hand out benefits to specific groups while spreading the cost thinly across everyone. A highway project, an agricultural subsidy, or a research grant program fits here. Because the benefits concentrate and the costs disperse, the politics tend to be quiet and cooperative, marked by logrolling, where legislators trade support for each other's projects rather than fighting.
Redistributive policies deliberately shift resources from one broad class to another. Progressive income taxation and means-tested benefits are the standard examples. Because both sides can see themselves clearly, the politics are ideological, visible, and bitter, and they tend to divide along stable party and class lines.
Regulatory policies impose rules on some identifiable group to benefit another. Environmental standards, workplace safety rules, and drug approval fall here. The politics are usually a contest between an organized industry and organized advocates, fought heavily in agencies and courts rather than on the floor of Congress.
Constituent policies concern the machinery of government itself: creating agencies, redrawing districts, setting budget procedures. They look technical and attract little public attention, which is exactly why they can be so consequential. A rule about how the budget is scored can matter more than a decade of speeches.
A related framework from James Q. Wilson sharpens the point by asking whether costs and benefits are concentrated or dispersed. When both are concentrated you get interest group politics, an open war between two organized sides. When benefits are concentrated and costs dispersed you get client politics, in which a small group quietly captures a benefit nobody else notices paying for. When costs are concentrated and benefits dispersed you get entrepreneurial politics, which usually requires a scandal, a disaster, or a determined policy entrepreneur to overcome the organized resistance. Kelsey's thalidomide case is a textbook example of the last kind: diffuse benefit to the public, concentrated cost to an industry, and a focusing event that broke the logjam.
Key idea: The structure of a policy's costs and benefits predicts the shape of the political fight over it. Before you predict who will fight, ask who bears the cost, who reaps the benefit, and whether each group is small enough to organize.
The stages heuristic
The single most common way to organize the policy process is the stages heuristic, sometimes called the policy cycle, associated with Harold Lasswell and later refined by Charles Jones and others. It breaks the process into a sequence.
- Agenda setting. Some conditions get defined as problems demanding government attention, while most do not.
- Policy formulation. Actors develop and refine possible responses.
- Adoption. An authoritative body chooses one, or some blend, and enacts it.
- Implementation. Agencies write rules, hire staff, spend money, and deliver.
- Evaluation. Someone asks whether it worked, and at what cost.
- Change or termination. The answer feeds back into the agenda, and the loop turns again.
The heuristic is genuinely useful. It gives you a checklist of questions, it reminds you that adoption is nowhere near the end of the story, and it tells you where to look for the actors who matter, since the cast changes completely from stage to stage. In Module 5 you will see that implementation failures kill more good ideas than legislative defeats do, and the stages model is what makes that observation sayable.
Why the stages heuristic is also wrong
Two scholars, Paul Sabatier and Hank Jenkins-Smith, mounted the most influential critique, and their objections are worth taking seriously because they will keep you from mistaking a filing system for a theory.
First, the stages do not actually occur in order. Agencies frequently implement in ways that redefine the problem, which sends the issue back to the agenda before evaluation has begun. Legislators formulate solutions and then hunt for problems those solutions can be attached to, a pattern you will meet again as Kingdon's garbage can. Evaluation often happens during implementation and changes it midstream.
Second, the model has no causal engine. It tells you that agenda setting happens; it does not tell you why one issue rises and another does not. A framework that describes a sequence without explaining any transition cannot generate testable predictions, which is what Sabatier meant when he called it a heuristic rather than a theory.
Third, it implies a single cycle for a single policy, when in reality dozens of overlapping cycles run at once across federal, state, and local levels, on interlocking issues, with different actors on different clocks. American climate policy at any moment consists of federal rulemaking, state programs, municipal ordinances, and litigation, all at different stages simultaneously.
The honest position is that the stages heuristic is a good map and a bad engine. Use it to organize what you are looking at and to remember stages you would otherwise skip. Do not use it to explain why anything happened. For explanation you need the frameworks in Module 3.
Key idea: The stages heuristic is a filing system, not a theory. It organizes attention well and explains causation not at all, so keep using it and stop expecting it to predict.
Working the definition on a real case
Try the whole apparatus on something concrete. In 2011 the federal government began requiring chain restaurants with twenty or more locations to post calorie counts on menus, a provision of the Affordable Care Act that took years to reach full enforcement in 2018.
Ask the instrument question: this is an information tool, not a regulation of what may be sold. Ask the Lowi question: costs concentrate on a definable industry, benefits disperse across everyone who eats out, which predicts entrepreneurial politics and a long fight in the agency rather than on the House floor, which is what happened. Ask the inaction question: for decades the policy was that no disclosure was required, and that was a choice. Ask the stages question and you can see agenda setting driven by rising obesity indicators, formulation over which chains and which items would be covered, adoption inside a much larger statute, an implementation phase lasting seven years, and an evaluation literature that is, to be blunt, mixed. Studies generally find small average reductions in calories ordered, on the order of a few dozen calories per transaction in some settings and no detectable effect in others. Whether that is a policy success depends on what you compare it to, and on how much you weight a small effect spread over an enormous number of meals against the compliance cost. Reasonable analysts land in different places, and this course will not pretend otherwise.
Common misconceptions
- Policy means legislation. Statutes are one source. Agency rules, court decisions, budget line items, prosecutorial priorities, and deliberate non-enforcement are all policy, and in volume they dwarf what Congress passes.
- Doing nothing is neutral. Non-action allocates benefits and burdens just as decisively as action does. The absence of a rule is a rule about who bears a risk.
- The stages heuristic explains the policy process. It describes a sequence. It identifies no causes and predicts no transitions, and its own defenders now call it a heuristic for that reason.
- Policy analysis tells you the right answer. Analysis narrows uncertainty and exposes trade-offs. Choosing among trade-offs requires values, which analysis cannot supply.
- Good policy design is mostly about picking the right goal. Goals are usually widely shared. The hard, contested work is choosing among instruments that reach the goal at different costs and load those costs onto different people.
Recap
- Public policy is the pattern of authoritative action and inaction through which government addresses a public problem, accumulated across many decisions.
- Governments act through five broad instruments: regulation, taxes and subsidies, direct provision, transfers, and information. Most problems admit several.
- Lowi and Wilson showed that the distribution of a policy's costs and benefits predicts the kind of political fight it produces.
- The stages heuristic organizes the process into agenda setting, formulation, adoption, implementation, evaluation, and change, and it is invaluable as a checklist.
- It is not a causal theory: stages overlap, run out of order, and multiply across levels of government, so explanation requires the frameworks in Module 3.
Sources
- Krutz, G., & Waskiewicz, S. (2021). 16.1 What is public policy? In American government 3e. OpenStax. openstax.org
- Krutz, G., & Waskiewicz, S. (2021). 16.2 Categorizing public policy. In American government 3e. OpenStax. openstax.org
- Encyclopaedia Britannica. (n.d.). Public policy. britannica.com
- U.S. Food and Drug Administration. (n.d.). Frances Oldham Kelsey: Medical reviewer famous for averting a public health tragedy. fda.gov
- Wikipedia contributors. (n.d.). Public policy. Wikipedia. en.wikipedia.org
- Key terms
- Public policy
- The pattern of authoritative action and inaction by which a government addresses, or declines to address, a public problem.
- Policy instrument
- A tool government uses to act: regulation, taxation and subsidy, direct provision, transfers and insurance, or information and persuasion.
- Distributive policy
- Policy that concentrates benefits on identifiable groups while spreading costs thinly, producing quiet logrolling politics.
- Redistributive policy
- Policy that deliberately shifts resources between broad classes, producing visible, ideologically charged conflict.
- Client politics
- Wilson's pattern in which concentrated benefits and dispersed costs let a small organized group win quietly.
- Entrepreneurial politics
- Wilson's pattern in which dispersed benefits and concentrated costs require a scandal, crisis, or determined advocate to overcome organized resistance.
- Stages heuristic
- The model dividing the policy process into agenda setting, formulation, adoption, implementation, evaluation, and change; useful for organizing attention, not for explaining causes.
- Logrolling
- Legislators trading votes so each secures a benefit for their own constituency.
The Fragmented System: Congress, the Presidency, and Federalism
- Explain why the American policy system was deliberately designed to be difficult to move, and what that design costs and buys.
- Trace how a policy idea survives or dies across committees, chambers, the filibuster, appropriations, and the veto.
- Analyze federalism as a policy variable, including preemption, grants, and the laboratory argument with its rebuttals.
The big picture
Here is a puzzle worth sitting with. Public opinion polling has shown majority support for background checks on gun sales, for some form of paid family leave, and for allowing Medicare to negotiate drug prices, in each case for years, and often across party lines. Yet for long stretches none of those became federal law. If you assume that policy tracks majority preference, this looks like corruption or conspiracy. It is mostly neither. It is architecture.
The people who designed the American system were not trying to build a machine that converts majorities into laws. They were trying to build one that makes it hard for any faction, including a majority faction, to impose its will quickly. James Madison said so plainly in Federalist 51: ambition must be made to counteract ambition. The system has many veto points, places where a determined minority can stop something, and comparatively few accelerators. Once you internalize that, American policy outcomes stop looking random and start looking like the predictable output of a specific institutional design.
This lesson maps that architecture. Your goal is not to memorize a civics diagram but to learn to ask, of any proposal, a practical question: where are the choke points, who controls each one, and what does the proposal have to give up at each to survive? That is how working analysts think about feasibility, and feasibility is not separate from good analysis. A brilliant option that cannot clear a single veto point is worth less to a decision maker than a mediocre one that can.
Counting the veto points
Follow an idea through the federal gauntlet. Suppose a member of the House wants to create a national program and has drafted a bill.
The bill is referred to a committee, and the committee chair decides whether it gets a hearing at all. Most bills die here silently: in a typical Congress, several thousand bills are introduced in the House and a small fraction receive committee action. If the bill survives markup, House leadership and the Rules Committee decide whether it reaches the floor and under what terms, including whether amendments are permitted. A floor majority is then required.
The Senate is where American policy most often dies. Under current practice, ending debate on most legislation requires sixty votes to invoke cloture, which means a minority of forty-one senators can prevent a vote. The filibuster is not in the Constitution. It emerged from a procedural accident in the 1800s and hardened into routine practice only in recent decades. Its defenders argue it forces the broad coalitions that make policy durable and protects against whipsaw reversals every time control changes. Its critics argue it entrenches minority rule, rewards obstruction with no cost, and pushes policymaking into executive action and the courts, which are less accountable still. Both arguments have real force, and where you land depends partly on how much you value stability against responsiveness. This course will not settle it for you.
There is an important exception. The budget reconciliation process allows certain fiscal legislation to pass the Senate by simple majority, subject to the Byrd rule, which strips provisions whose budgetary effects are merely incidental to their policy purpose. This is why so much major American legislation of the past twenty-five years has been shaped like tax and spending policy even when its purpose was social: reconciliation is the one reliable path around the sixty-vote threshold, and it bends the substance of policy to fit its procedural keyhole. That is a striking case of process determining content.
Assume the bill clears both chambers. The versions must be reconciled. The president may veto, and overriding requires two thirds of both chambers, which almost never happens. And even then the program does not exist until it is funded: authorization creates the legal permission, appropriation supplies the money, and they are separate acts by separate committees. Many authorized programs have never been meaningfully funded, which returns us to Lesson 1's point that inaction is policy.
Key idea: A proposal must clear a committee chair, chamber leadership, two floor majorities, usually a sixty-vote Senate threshold, a conference, a presidential signature, and a separate appropriation. Any single blockage is fatal, which is why the default outcome of the American system is that nothing happens.
What the design buys and what it costs
It would be easy to present fragmentation as pure dysfunction, and much popular commentary does. The honest account is a trade-off.
The case for the design: durable policy requires broad support, and policy that flips with every election is worse than no policy at all, because households and firms cannot plan against it. Multiple veto points give affected parties a chance to be heard, which surfaces problems before they are locked into law. And a system that is hard to move in one direction is equally hard to move in another, which is protection you may value most when your side is out of power.
The case against: veto points do not protect everyone equally. They systematically advantage those who are already organized, well funded, and satisfied with the status quo, because stopping something requires far less capacity than passing something. The design also produces accountability confusion, since when nothing happens voters struggle to identify whom to blame. And it pushes action toward the least deliberative channels: executive orders that the next president can reverse, agency rules subject to judicial reversal, and litigation.
Comparative context sharpens the trade-off. In a Westminster parliamentary system such as the United Kingdom's, the government commands a legislative majority by definition, so major legislation can pass in months. Britain established the National Health Service within three years of the 1945 election. The corresponding cost is reversibility and less protection for minorities. Germany sits between: a federal system with a second chamber representing the states and a strong constitutional court. Neither arrangement is simply better. They price stability and responsiveness differently.
The presidency as a policy actor
Presidents shape policy through several channels, and legislation is only one. Richard Neustadt's classic formulation is that presidential power is the power to persuade, meaning that a president's formal authorities are modest and their real leverage comes from bargaining, reputation, and public standing.
But the unilateral toolkit is not modest. Executive orders direct the executive branch and carry the force of law within its scope. Agencies issue rules under authority delegated by statute, and those rules fill many more pages annually than statutes do. Presidents set enforcement priorities, deciding which of many laws get resources. They negotiate executive agreements with other nations without Senate ratification. They propose the budget, which frames the fiscal debate even though Congress disposes. And through the Office of Information and Regulatory Affairs, the White House reviews significant agency rules and their cost-benefit analyses, a centralizing power institutionalized under Reagan and retained by every president since, of both parties.
The obvious limitation is durability. What one president does by pen, the next can undo by pen. Policies built on executive action are fast and fragile, and the past two decades of American immigration and environmental policy illustrate both halves of that sentence. An analyst asked to compare a statutory route with an executive route should say so explicitly: the executive route is faster, narrower, and reversible, and those are not small differences.
Key idea: Presidential policymaking trades durability for speed. Executive action moves quickly and can be reversed just as quickly, so the choice of route is itself a substantive policy decision, not a technicality.
Federalism as a policy variable
The United States is not one policy system but fifty-one interacting ones, plus tribal governments, territories, and roughly ninety thousand local governments including counties, municipalities, school districts, and special districts. For most policy questions, the first analytic move is to ask which level actually holds the lever.
| Mechanism | What it does | Example |
|---|---|---|
| Enumerated and implied federal powers | Ground national action, especially through the commerce and spending powers | Federal environmental and civil rights statutes |
| Reserved state powers | Leave education, most criminal law, family law, land use, and professional licensing to states | School funding formulas; occupational licensing rules |
| Preemption | Federal law displaces conflicting state law | Federal aviation and much of drug regulation |
| Categorical grants | Federal money for a narrow purpose with detailed conditions | Title I education funding |
| Block grants | Federal money for a broad purpose with state discretion | Temporary Assistance for Needy Families |
| Conditional funding | Money offered contingent on adopting a policy | Highway funds tied to a minimum drinking age of 21 |
Justice Louis Brandeis gave federalism its most famous defense in 1932, writing that a single courageous state may serve as a laboratory and try novel social and economic experiments without risk to the rest of the country. The argument is genuinely powerful, and it has substance behind it. Wisconsin's welfare reforms preceded the 1996 federal law. Massachusetts adopted an individual mandate and exchanges in 2006, four years before the Affordable Care Act borrowed the structure. California vehicle emission standards have repeatedly pulled national practice along. State minimum wage variation gave economists the natural experiments that drive the entire modern literature on employment effects, which you will meet in Module 5.
The rebuttals are equally real, and a good analyst holds both. States may compete downward rather than upward when mobile taxpayers or firms can leave, a dynamic sometimes called a race to the bottom, though economists disagree about how strong it actually is. Spillovers cross borders, so upstream pollution and downstream drinking water do not respect state lines. Fifty different rule sets impose genuine compliance costs on national firms. Fiscal capacity varies enormously, so equal effort produces unequal services and identical children receive very different public educations depending on the state line they were born behind. And the laboratory argument assumes experiments get evaluated and copied, which requires research infrastructure and political willingness that are not guaranteed.
Notice that these are not merely empirical disputes. Someone who prioritizes local self-government and policy variety will weigh the same evidence differently from someone who prioritizes uniform minimum guarantees for every citizen. That is a values difference, not a factual error on either side, and recognizing which disagreements are of which kind is one of the central skills this course is trying to build.
Common misconceptions
- The filibuster is constitutional. It is a Senate rule arising from procedural practice, changed several times in living memory, and it does not appear in the Constitution.
- A president can enact an agenda by executive order. Executive action is bounded by existing statutory authority, subject to judicial review, and reversible by the next administration.
- Authorization means a program exists. Authorization grants legal permission; appropriation supplies money. Programs authorized but never funded are common.
- Federalism is a fixed constitutional division. The boundary moves constantly through preemption, grant conditions, and litigation, and it is itself a live policy battleground.
- Gridlock means nothing is happening. When legislation stalls, policymaking migrates to agencies, courts, states, and cities. The activity moves; it does not stop.
Recap
- The American system was designed with many veto points, so the default outcome is inaction and the burden always falls on those seeking change.
- The design trades responsiveness for durability and minority protection; comparative cases such as the United Kingdom and Germany price that trade-off differently.
- Reconciliation is the main route around the Senate's sixty-vote threshold, and its procedural constraints visibly shape the substance of major legislation.
- Presidents act through persuasion, executive orders, agency rulemaking, enforcement priorities, and regulatory review, gaining speed at the price of reversibility.
- Federalism distributes policy levers across levels; the laboratory argument and the race-to-the-bottom and equity rebuttals are all serious, and choosing between them involves values as well as evidence.
Sources
- Krutz, G., & Waskiewicz, S. (2021). 3.1 The division of powers. In American government 3e. OpenStax. openstax.org
- Krutz, G., & Waskiewicz, S. (2021). 3.5 Advantages and disadvantages of federalism. In American government 3e. OpenStax. openstax.org
- Congressional Research Service. (n.d.). Reports on Congress and the legislative process. crsreports.congress.gov
- Encyclopaedia Britannica. (n.d.). Federalism. britannica.com
- Wikipedia contributors. (n.d.). Filibuster in the United States Senate. Wikipedia. en.wikipedia.org
- Key terms
- Veto point
- A place in the process where a single actor or small group can block a proposal, of which the American system has unusually many.
- Cloture
- The Senate procedure for ending debate, requiring sixty votes for most legislation and thus giving forty-one senators a blocking power.
- Reconciliation
- A budget process allowing certain fiscal legislation to pass the Senate by simple majority, constrained by the Byrd rule.
- Authorization versus appropriation
- Authorization creates legal permission for a program; appropriation supplies its money. Both are required for a program to operate.
- Preemption
- The displacement of state law by federal law where the two conflict or where federal law occupies the field.
- Categorical grant
- Federal funding for a narrowly defined purpose with detailed conditions attached, contrasted with a broader block grant.
- Laboratories of democracy
- Brandeis's argument that states can test policies at limited risk, producing evidence the nation can learn from.
- Regulatory review
- White House scrutiny of significant agency rules and their cost-benefit analyses, centralized through the Office of Information and Regulatory Affairs.
The Bureaucracy, the Courts, and the Influence Industry
- Explain how agencies convert vague statutes into operative policy through notice-and-comment rulemaking and enforcement discretion.
- Describe how courts shape policy through review of agency action, and why doctrinal shifts change the balance between agencies and judges.
- Assess the roles of interest groups, think tanks, and media, distinguishing evidence-based claims about influence from folk theories.
The big picture
Every year Congress enacts a few hundred public laws. Every year federal agencies issue roughly three thousand final rules. The Federal Register, the daily journal in which those rules appear, runs to tens of thousands of pages annually. If you measure policy by volume of binding text that changes what people and firms may do, Congress is a minor producer and the executive agencies are the main event.
This is not a scandal, and it is not an accident. Congress writes statutes at a level of generality it cannot escape. The Clean Air Act instructs the Environmental Protection Agency to set standards that protect public health with an adequate margin of safety. That sentence does not tell you the permissible concentration of fine particulate matter. Somebody with a laboratory, an epidemiological literature, and a staff must decide, and Congress lacks the technical capacity and the time to do it for every substance in every industry. So it delegates.
Delegation creates the central tension of the administrative state. Expertise and continuity live in agencies; democratic accountability lives in elected institutions. Every mechanism in this lesson, notice-and-comment procedure, judicial review, congressional oversight, White House regulatory review, is an attempt to get the benefits of expert administration without severing the connection to consent. None of them works perfectly. People who look at the same arrangement and weigh expertise against accountability differently reach genuinely different conclusions about whether the administrative state is a triumph or a problem, and this lesson will present both without adjudicating.
How a rule is actually made
The Administrative Procedure Act of 1946 established the process most federal rules follow, and it is worth knowing in detail because it is where an enormous amount of real policy gets decided.
- Statutory authority. An agency may only act where a statute authorizes it. Locating and interpreting that authority is the first fight, and often the last one.
- Proposed rule. The agency publishes a notice of proposed rulemaking in the Federal Register, including the text, the reasoning, and for significant rules a regulatory impact analysis with costs and benefits.
- Comment period. Anyone may comment, typically over thirty to ninety days. Major rules can draw hundreds of thousands of comments.
- Response and final rule. The agency must consider significant comments and respond to them in the preamble to the final rule. Failing to do so is a common ground for a court to strike the rule down.
- Review. Economically significant rules pass through the Office of Information and Regulatory Affairs at the White House. Congress may disapprove a rule under the Congressional Review Act within a limited window. Litigation may follow at any point.
Two things about this process surprise students. First, it is slow: a major rule commonly takes three to five years from proposal to effect, and longer if litigated. Second, the comment process is formally open to everyone and practically dominated by the organized. Writing a comment that will actually move an agency requires knowing the docket, the technical literature, and the legal standard. Regulated firms employ people who do this full time. Diffuse beneficiaries, by definition, do not. This is Wilson's collective action problem from Lesson 1 reappearing in procedural form.
Key idea: The bulk of binding American policy is made by agencies filling in statutory generalities, through a procedure that is open in form and unequal in practice because participation rewards concentrated, professionalized interests.
Enforcement discretion is policy
Rulemaking is only half of what agencies do. The other half is deciding which rules to enforce, against whom, and how hard. No agency has resources to enforce every provision it administers, so priorities must be set, and those priorities are policy in the fullest sense of Lesson 1's definition.
Consider what changes when an agency shifts inspection resources from small firms to large ones, or announces it will not pursue a category of violation, or reduces penalty amounts. Nothing in the statute or the regulation has changed. The operative policy has changed a great deal. Analysts who read only the law and not the enforcement data will systematically misdescribe what the government is doing. The Government Accountability Office, Congress's own audit and evaluation arm, exists in large part to close that gap, and its reports are among the most useful and least read documents in American government.
Courts as policymakers
American courts shape policy in three ways that are worth distinguishing.
The first is constitutional review: striking down statutes or executive actions as beyond constitutional limits. This is the most dramatic and the least frequent.
The second, and by far the most consequential in day-to-day policy, is review of agency action. Under the Administrative Procedure Act, a court may set aside agency action found to be arbitrary and capricious, or in excess of statutory authority. The practical question is how much deference judges owe an agency's reading of an ambiguous statute. For roughly forty years after the 1984 Chevron decision, the answer was substantial: if a statute was ambiguous and the agency's interpretation reasonable, courts generally deferred. In 2024, in Loper Bright Enterprises v. Raimondo, the Supreme Court overruled Chevron, holding that courts must exercise independent judgment in interpreting statutes. The consequences are still unfolding.
The competing arguments are both serious. Supporters of the change argue that interpreting law is the judicial function, that Chevron let agencies expand their own authority by finding ambiguity, and that policy shifting with each administration's reading of the same words is bad governance. Critics argue that generalist judges lack the technical competence to second-guess agency scientists, that the change transfers power from politically accountable agencies to life-tenured judges, and that it invites forum shopping and inconsistent nationwide rules. Which concern dominates depends on whether you think the greater risk is unaccountable agencies or unaccountable judges. That is not a question evidence alone will answer.
The third channel is remedial: courts supervising institutions over long periods through consent decrees and injunctions, as in school desegregation, prison conditions, and child welfare cases. This is judicial policymaking at its most direct, and it raises the sharpest questions about institutional competence and democratic legitimacy on both sides.
Key idea: Most judicial influence on policy runs through review of agency action rather than constitutional rulings, so doctrines about deference are among the highest-stakes and least visible policy questions in the system.
Interest groups: what the evidence actually shows
The folk theory is simple: money buys policy. The research is more interesting and more equivocal, and this is a place where careful analysts part company with confident commentators.
Start with what is well established. Interest groups are numerous, unequally distributed, and heavily concentrated in business and professional representation relative to the general public. Lobbying expenditures reported under federal disclosure law run into billions of dollars annually. Groups do gain access, and access is not nothing.
Now the contested part. A prominent 2014 study by Martin Gilens and Benjamin Page found that when the preferences of affluent Americans and organized business groups diverged from those of average citizens, policy outcomes tracked the former far more closely. The finding was widely reported as showing that ordinary preferences have near-zero influence. Subsequent work by Peter Enns, Omar Bashir, J. Alexander Branham and colleagues, and others challenged the interpretation, pointing out that rich and average Americans agree on the large majority of issues, that the cases of divergence are relatively few and often narrow, and that measurement choices strongly affect the estimated gap. The critics generally do not claim influence is equal; they claim the inequality is smaller and more conditional than the headline suggested.
Meanwhile a different line of research, associated with Richard Hall and Frank Baumgartner among others, finds that lobbying works less by changing minds than by subsidizing allies. Lobbyists give friendly legislators information, draft language, and analysis, allowing those legislators to be more effective on issues they already cared about. On this account the main effect of lobbying is on legislative attention and capacity rather than on vote switching. Baumgartner's large study of lobbying campaigns found that the side with more resources won only somewhat more often than not, and that the status quo won most of the time regardless.
The defensible summary is this: organized interests clearly shape which issues get attention, how proposals are drafted, and what dies quietly in committee. Whether they routinely override majority opinion on salient issues is genuinely disputed among competent researchers, and anyone who tells you the question is settled is overstating. Note also the normative dimension, since the First Amendment protects petitioning government, and one person's corrupting special interest is another's legitimate association defending its members.
Think tanks and the media
Think tanks occupy the space between universities and advocacy. Some, such as the Congressional Budget Office and the Congressional Research Service, are official nonpartisan bodies serving Congress directly, and their credibility rests on scrupulous neutrality. Others are independent research organizations with varying degrees of ideological commitment, ranging from those that resemble universities to those that function primarily as advocacy shops producing research-shaped material.
The practical skill is triage. When you encounter a policy study, ask: who funded it, was it peer reviewed or self-published, are the data and code available, does the organization publish findings that cut against its apparent priors, and does the paper report confidence intervals and limitations or only headline numbers? These questions are ideologically neutral and you should apply them symmetrically. An analyst who scrutinizes only research from organizations they dislike is not doing analysis.
Media shape policy chiefly through agenda setting, a well-supported finding usually traced to Maxwell McCombs and Donald Shaw: news coverage is much better at telling audiences what to think about than what to think. Coverage volume correlates with public salience, and salience determines what elected officials feel compelled to address. Framing effects, covered fully in Module 3, operate alongside this. The contemporary complication is fragmentation. When audiences select into different information environments, agenda setting becomes segmented, and the shared national agenda that made certain kinds of broad policy coalitions possible is harder to assemble.
Common misconceptions
- Bureaucrats make policy illegitimately. Agencies act under authority delegated by statute and constrained by procedure, review, and oversight. Whether the constraints are adequate is a genuine debate; the delegation itself is ordinary law.
- Public comment gives everyone equal voice. The process is formally open but rewards technical sophistication and sustained attention, which favors organized and professionalized interests.
- Money reliably buys policy outcomes. The best evidence supports influence over agenda, drafting, and legislative attention. Direct vote-buying on salient issues is much harder to demonstrate, and researchers disagree about magnitude.
- Overruling Chevron reduced judicial policymaking. It shifted interpretive authority from agencies to courts, which increases the policy role of judges rather than reducing it.
- Think tank equals biased, university equals objective. Quality varies within both categories. Apply the same funding, transparency, and peer-review questions to every source, including ones you agree with.
Recap
- Agencies produce most binding federal policy by filling in statutory generalities through notice-and-comment rulemaking under the Administrative Procedure Act.
- Enforcement discretion is policy: shifting priorities changes what government actually does without changing a word of law, which is why oversight bodies like GAO matter.
- Courts influence policy mainly by reviewing agency action, and the post-Chevron shift toward independent judicial interpretation reallocates power between judges and agencies with serious arguments on both sides.
- Interest group influence is real but contested in magnitude; the strongest evidence concerns agenda, drafting, and legislative capacity rather than vote switching.
- Media set the agenda more than they change minds, and source triage questions about funding, review, and transparency should be applied symmetrically to all research.
Sources
- U.S. Government Accountability Office. (n.d.). What GAO does. gao.gov
- Krutz, G., & Waskiewicz, S. (2021). 15.4 Controlling the bureaucracy. In American government 3e. OpenStax. openstax.org
- Encyclopaedia Britannica. (n.d.). Bureaucracy. britannica.com
- Congressional Research Service. (n.d.). Reports on administrative law and rulemaking. crsreports.congress.gov
- Wikipedia contributors. (n.d.). Administrative Procedure Act (United States). Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Loper Bright Enterprises v. Raimondo. Wikipedia. en.wikipedia.org
- Key terms
- Delegation
- Congress granting agencies authority to fill in statutory details, trading direct accountability for technical capacity.
- Notice-and-comment rulemaking
- The Administrative Procedure Act process requiring an agency to publish a proposed rule, accept public comment, and respond to significant comments in the final rule.
- Enforcement discretion
- An agency's choice of which violations to pursue and how hard, which changes operative policy without changing any legal text.
- Arbitrary and capricious review
- The standard under which courts set aside agency action that lacks reasoned explanation or ignores significant comments and evidence.
- Chevron deference
- The 1984 doctrine, overruled in 2024, under which courts generally accepted an agency's reasonable interpretation of an ambiguous statute.
- Legislative subsidy
- Hall's account of lobbying as supplying information, drafting, and analysis to already-sympathetic legislators rather than changing minds.
- Agenda setting (media)
- McCombs and Shaw's finding that news coverage strongly influences what audiences consider important, more than what they conclude about it.
- Congressional Review Act
- A statute allowing Congress to disapprove a recently issued agency rule by joint resolution within a limited window.
Module 2: Why Government Acts, and Why It Fails
The standard economic case for intervention through the four classic market failures, the separate and genuinely different equity rationale, and then the honest counterweight: public choice theory, regulatory capture, and unintended consequences. The point is to hold both critiques at once rather than deploying whichever is convenient.
Market Failure: The Standard Case for Government
- Identify the four classic market failures and diagnose which one is present in a real case.
- Explain externalities and public goods precisely enough to distinguish them from ordinary complaints about markets.
- Match plausible policy instruments to each failure type and state the assumptions each remedy requires.
The big picture
Suppose someone proposes that government should do something about a problem. Before asking whether the proposal is good, an analyst asks a prior question: why is a market not solving this on its own? The question is not hostile. It is diagnostic. Different reasons produce different remedies, and applying the wrong remedy to the right problem is one of the most reliable ways to waste public money.
Economics supplies a benchmark for that diagnosis. Under a demanding set of conditions, competitive markets allocate resources efficiently, meaning no reallocation could make someone better off without making someone else worse off. The conditions are strong: many buyers and sellers, no meaningful market power, complete and symmetric information, well defined and enforceable property rights, and no effects on third parties. When they hold, the market outcome is efficient and there is no efficiency case for intervention. When they fail, there may be.
Two cautions before we begin, and they matter. First, efficiency is not the only thing worth wanting. A perfectly efficient outcome can be one you find morally intolerable, and Lesson 5 takes that up in full. Second, showing a market failure does not establish that government intervention would improve matters. It establishes only that improvement is theoretically possible. Whether an actual government, with actual information and actual incentives, would do better is a separate empirical question, and answering it too quickly is a mistake made routinely on all sides.
Failure one: public goods
A public good has two properties. It is non-rival, meaning one person's consumption does not reduce what is available to others, and non-excludable, meaning you cannot practically prevent non-payers from benefiting. National defense is the standard example: protecting the country protects everyone in it, and your protection does not use up anyone else's.
Why does this break markets? Because if you cannot exclude non-payers, rational people have every reason to enjoy the benefit without paying, a behavior called free riding. Ask a hundred neighbors to contribute voluntarily to a levee that protects the whole neighborhood and most will reason that their individual contribution barely matters and the levee will be built or not regardless. Every individual reasons correctly and the levee does not get built. This is not selfishness in any moral sense; it is a structural feature of the situation. The standard remedy is compulsory finance, which is to say taxation, precisely because it removes the option to free ride.
Be precise about the definition, because it is widely misused. Public good does not mean good for the public, and it does not mean provided by government. Public education is largely government provided but it is both rival and excludable, so it is not a public good in the technical sense; the case for it rests on externalities and equity instead. Meanwhile basic scientific research is close to a genuine public good, since knowledge once published is non-rival and hard to exclude, which is the standard economic argument for public funding of basic research. Some goods are non-rival but excludable, like a streamed film, and are called club goods. Some are rival but non-excludable, like an ocean fishery, and are called common-pool resources. That last category produces the tragedy of the commons, and Elinor Ostrom won a Nobel Prize for demonstrating that real communities frequently solve it through local institutions rather than either markets or state ownership, which is an important corrective to the textbook story.
Key idea: Public goods fail in markets because non-excludability makes free riding individually rational. The technical definition is narrow, so test any candidate against both non-rivalry and non-excludability before applying the label.
Failure two: externalities
An externality is a cost or benefit imposed on a third party who is not part of the transaction. A factory that discharges pollution downstream imposes a negative externality on people who neither bought nor sold anything. A homeowner who vaccinates their child confers a positive externality on everyone that child would otherwise have infected.
The efficiency logic is straightforward once you see it. The factory weighs its private costs against its private benefits and produces where those balance. But the true social cost includes the damage downstream, which the factory does not pay. So it produces more than the socially efficient amount. Symmetrically, someone deciding whether to vaccinate weighs private benefit against private cost and ignores the benefit to others, so society gets less vaccination than would be efficient. Negative externalities produce overproduction; positive externalities produce underproduction.
The classic remedy, proposed by A. C. Pigou, is to make the private cost equal the social cost by taxing the harm or subsidizing the benefit, at a rate equal to the marginal external effect. A carbon tax is a Pigouvian tax. A vaccination subsidy is a Pigouvian subsidy.
Ronald Coase then complicated this beautifully. In a 1960 paper he argued that if property rights are clearly assigned and bargaining costs are low, the parties will negotiate to the efficient outcome regardless of who holds the right, though who holds it determines who pays whom. If the factory has the right to pollute, downstream residents can pay it to stop; if residents have the right to clean water, the factory can pay for permission. Either way, if the deal is worth doing it gets done. The Coase theorem is often cited as an argument against regulation, but read carefully it is an argument about transaction costs. With one factory and one downstream farm, bargaining is plausible. With a million drivers and eight billion people affected by climate change, transaction costs are prohibitive and Coasean bargaining is a fantasy. The theorem is best used as a diagnostic: ask how many parties there are and how cheaply they can contract, and let that answer tell you whether private bargaining is even on the table.
A third approach, developed by John Dales and others, is to create a property right where none existed and let it trade. Cap the total quantity of emissions, issue permits, and allow firms to buy and sell them. The market then finds the cheapest abatement wherever it lies. Module 6 examines how this worked in the American sulfur dioxide program and what happened when Europe tried it for carbon.
Key idea: Externalities drive a wedge between private and social cost, producing too much of harmful activity and too little of beneficial activity. Remedies include taxes, subsidies, tradable rights, and direct bargaining, and transaction costs largely determine which is feasible.
Failure three: information problems
Markets rely on buyers knowing what they are buying. When information is badly asymmetric, markets can shrink or collapse.
George Akerlof's 1970 paper on the market for lemons made the mechanism vivid. Sellers of used cars know their quality; buyers do not. Buyers, unable to distinguish, will only pay an average price. Owners of good cars find that price too low and withdraw, which lowers average quality, which lowers the price buyers will offer, which drives out the next tier. This unraveling is called adverse selection, and Akerlof shared a Nobel Prize for it.
The same logic runs through insurance, which is why it dominates health policy arguments. If insurers must charge everyone the same price and cannot screen, the people most likely to buy are those who expect high costs, which raises the average claim, which raises the premium, which drives out the healthiest remaining buyers. Every serious health financing system in the world contains some device to counter this: mandatory participation, automatic enrollment, heavy subsidies, or risk adjustment among insurers. They differ enormously in which device and in how much they rely on markets, but none of them ignores adverse selection.
The companion problem is moral hazard: once insured, people bear less of the cost of their choices and may take more risk or consume more care. Both problems are real, and there is no arrangement that eliminates both. Deductibles and copayments reduce moral hazard and also deter some genuinely valuable care, and the RAND Health Insurance Experiment of the 1970s and 1980s found exactly that, with cost sharing reducing both wasteful and useful utilization. Designing around this trade-off is a permanent feature of health policy, not a solvable problem.
Information remedies are often lighter touch than the alternatives: mandatory disclosure, licensing and certification, product testing, warranty rules, and standardized labels. They preserve choice, which many people value independently. They also fail when the information is too complex for the decision at hand, which is why disclosure-based consumer protection has a mixed empirical record.
Failure four: market power and natural monopoly
A firm with market power can raise price above marginal cost, producing less than the efficient quantity and capturing surplus from consumers. The general remedy is antitrust: prohibiting collusion, blocking anticompetitive mergers, and policing exclusionary conduct.
A special case is natural monopoly, which arises when a single firm can serve the whole market more cheaply than several can, because fixed costs are huge and marginal costs are small. Water distribution, electricity transmission, and sewer systems fit: nobody wants four competing sets of pipes under the street. Competition here is not merely difficult, it is wasteful.
Governments handle natural monopoly in three broad ways. Rate-of-return or price-cap regulation lets a private firm operate under a regulated price. Public ownership puts the asset in government hands, which is common for water in the United States and was common for electricity and rail in Europe before privatization. Franchise bidding, proposed by Harold Demsetz, auctions the right to serve for a period. Each has a characteristic weakness: rate-of-return regulation weakens cost discipline and invites the capture problems of Lesson 5, public ownership can be politically insulated from efficiency pressure, and franchise contracts are hard to write completely for long-lived assets. The comparative record across countries is mixed enough that reasonable economists disagree about which is preferable in which sector.
Diagnosing in practice
| Symptom | Likely failure | Common instruments |
|---|---|---|
| Everyone benefits, nobody will pay | Public good | Tax finance and direct provision |
| Third parties harmed or helped by others' transactions | Externality | Pigouvian tax or subsidy, tradable permits, standards |
| Buyers cannot judge quality; good sellers exit | Adverse selection | Disclosure, licensing, mandates, risk adjustment |
| One firm serves the market most cheaply | Natural monopoly | Price regulation, public ownership, franchise bidding |
| Few sellers, high prices, high margins | Market power | Antitrust enforcement, merger review |
Run one case. Antibiotic overuse: each patient and prescriber gains a small private benefit while resistance accumulates as a cost borne by future patients everywhere. That is a negative externality with an intergenerational and international dimension, and it also has a public good on the other side, since new antibiotic research produces widely diffused benefits that a developer cannot fully capture. Diagnose both and you can see why the policy conversation involves both stewardship rules restricting use and prize or subsidy schemes to pull new drugs into existence. One diagnosis would have produced half a policy.
Common misconceptions
- Public good means anything beneficial to the public. It is a technical term requiring both non-rivalry and non-excludability. Most government services are not public goods in this sense.
- Externality means any effect on others. Bidding up a price by entering a market affects others but transmits through prices, which is how markets are supposed to work; the term covers effects outside the price system.
- Demonstrating market failure justifies intervention. It shows improvement is theoretically possible. Whether real government would improve on the real market requires separate evidence.
- The Coase theorem shows regulation is unnecessary. It shows bargaining reaches efficiency when transaction costs are low, which is precisely the condition that fails in most environmental and public health problems.
- Moral hazard and adverse selection are the same thing. Adverse selection is about who buys, driven by information before the contract; moral hazard is about behavior changing after coverage begins.
Recap
- Competitive markets are efficient only under strong conditions, and the four classic failures name the main ways those conditions break.
- Public goods fail through free riding on non-excludable, non-rival benefits, which is why compulsory finance is the standard remedy.
- Externalities separate private from social cost; Pigouvian pricing, tradable permits, standards, and Coasean bargaining are all available, with transaction costs deciding feasibility.
- Information asymmetry produces adverse selection and moral hazard, which is why every functioning health system contains devices against both and none eliminates both.
- Natural monopoly makes competition wasteful, and price regulation, public ownership, and franchise bidding each carry characteristic weaknesses that the comparative record does not clearly rank.
Sources
- Greenlaw, S. A., & Shapiro, D. (2022). 13.4 Public goods. In Principles of microeconomics 3e. OpenStax. openstax.org
- Greenlaw, S. A., & Shapiro, D. (2022). 12.1 The economics of pollution. In Principles of microeconomics 3e. OpenStax. openstax.org
- U.S. Environmental Protection Agency. (n.d.). Environmental economics. epa.gov
- Wikipedia contributors. (n.d.). Market failure. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Coase theorem. Wikipedia. en.wikipedia.org
- Key terms
- Public good
- A good that is both non-rival in consumption and non-excludable, so free riding prevents voluntary market provision.
- Free rider problem
- The tendency to enjoy a non-excludable benefit without contributing, which leads to underprovision even when everyone values the good.
- Externality
- A cost or benefit falling on a third party outside a transaction, causing private and social costs to diverge.
- Pigouvian tax
- A tax set equal to the marginal external damage of an activity, aligning private cost with social cost.
- Coase theorem
- The proposition that with clear property rights and low transaction costs, private bargaining reaches an efficient outcome regardless of who holds the right.
- Adverse selection
- Market unraveling caused by buyers or sellers holding private information before contracting, as in Akerlof's market for lemons.
- Moral hazard
- Behavior change after being insured or protected, because the actor no longer bears the full cost of their choices.
- Natural monopoly
- A market in which one firm can serve total demand more cheaply than several, because fixed costs are large relative to marginal costs.
- Common-pool resource
- A good that is rival but non-excludable, such as a fishery, prone to overuse unless governed by rules.
Equity Rationales and the Honest Counterweight of Government Failure
- Distinguish efficiency from equity arguments and explain why efficiency alone cannot settle distributional questions.
- Explain public choice theory, regulatory capture, and unintended consequences as serious analytic tools rather than slogans.
- Apply both market failure and government failure reasoning to the same case without collapsing into either.
The big picture
Lesson 4 gave you the efficiency case for government action. This lesson gives you the two things that must sit beside it if your analysis is going to be honest.
The first is that efficiency is not the only value, and a great deal of what governments do has nothing to do with market failure. The second is that government is not a benevolent problem-solving machine that automatically corrects whatever markets get wrong. It is an institution staffed by people with their own interests, subject to its own systematic distortions. A framework that only names market failures will recommend intervention too readily; a framework that only names government failures will recommend it too rarely. You need both, applied symmetrically, and the discipline of applying them symmetrically to cases you feel strongly about is most of what separates an analyst from an advocate.
The equity rationale is genuinely separate
Economists sometimes describe the efficiency benchmark using the concept of Pareto efficiency: an allocation is Pareto efficient if no one can be made better off without making someone worse off. Notice what this permits. An allocation in which one person holds everything and everyone else holds nothing is Pareto efficient, because you cannot improve anyone's position without reducing that one person's. Efficiency is silent about distribution, entirely and by construction.
Because pure Pareto improvements are rare in policy, analysts often use the Kaldor-Hicks criterion instead: a change is an improvement if the winners gain enough that they could in principle compensate the losers and still come out ahead. This is what cost-benefit analysis operationalizes, and it is worth being clear-eyed about what it assumes. The compensation need not actually occur. A policy that adds a billion dollars of gains for a wealthy group while imposing nine hundred million in losses on a poor one passes the test. Whether that is a good policy is a question the test does not address.
So equity arguments enter on their own terms. Several distinct ones recur.
Insurance against the lottery of birth. John Rawls asked what principles you would choose for a society if you did not know what position you would occupy in it. He argued that from behind this veil of ignorance you would secure basic liberties and then accept inequalities only where they improve the position of the least advantaged. Whether or not you find the argument decisive, it captures why many people support redistribution without believing markets have failed at anything.
Equality of opportunity. Even people deeply skeptical of redistributing outcomes often accept that children should not have their life prospects determined by their parents' circumstances. This is the most widely shared equity premise in American politics and it underwrites public education, child nutrition, and early childhood programs across the political spectrum, though agreement on the premise coexists with sharp disagreement about which policies actually deliver it.
Specific egalitarianism. Some goods, such as emergency medical care, legal defense, and votes, are widely felt to warrant distribution by need or equality rather than ability to pay, even by people comfortable with market allocation of most goods. Notice that this is an intuition about particular goods rather than a general theory.
Against these sit serious counterarguments that deserve equal statement. Robert Nozick argued that if holdings arise from just acquisition and voluntary transfer, patterned redistribution requires continuous interference with liberty. Friedrich Hayek argued that market outcomes reflect dispersed knowledge no planner can assemble, and that treating them as distributions to be corrected misunderstands what prices do. Many economists across the spectrum note the efficiency cost of redistribution through the tax and transfer system, since taxes distort effort and investment and benefit phase-outs create high effective marginal rates on the poor. Arthur Okun called this the leaky bucket: some water is lost in transfer, and how much leakage is acceptable is a judgment, not a calculation.
Key idea: Efficiency criteria are silent about distribution by construction, so equity arguments are not an inferior kind of reasoning smuggled in afterward. They are a separate and legitimate basis for policy, and the serious objections to them are also about values, not just about facts.
Government failure: public choice
The public choice school, associated above all with James Buchanan and Gordon Tullock, applies the same behavioral assumptions to political actors that economics applies to market actors. Voters, legislators, and bureaucrats are assumed to pursue their own interests, not an abstract public interest. Buchanan received the Nobel Prize in 1986 for this work, sometimes summarized as politics without romance.
Several results follow, and each is useful even if you reject the framework's broader claims.
Rational ignorance. Anthony Downs observed that the cost of becoming informed about most policy questions exceeds the expected benefit to any individual voter, because one vote almost never decides anything. So most voters remain uninformed on most issues, rationally. This is not a criticism of voters; it is a structural prediction, and it explains why policies with diffuse costs survive scrutiny.
Concentrated benefits and diffuse costs. Mancur Olson's logic of collective action shows that small groups with much at stake organize more easily than large groups with little each at stake. A sugar tariff worth hundreds of millions to a few dozen producers and a few dollars to each consumer will be defended fiercely and opposed by nobody. This is Wilson's client politics from Lesson 1 with a formal engine underneath.
Rent seeking. Tullock coined this for resources spent obtaining favorable treatment rather than creating value: lobbying for a license restriction, a tariff, or a subsidy. The social cost is not merely the transfer but the real resources burned in the contest for it.
Budget maximizing. William Niskanen argued that agencies seek larger budgets because size brings prestige, salary, and discretion, producing systematic oversupply. This claim has fared less well empirically than the others; later work by Patrick Dunleavy and others suggests senior officials often prefer smaller, higher-status units to large service-delivery ones, and the empirical record is mixed enough that budget maximizing should be treated as a hypothesis rather than a finding.
Public choice has its own critics, and they should be heard. The self-interest assumption is a simplification; substantial evidence indicates that many public servants are motivated by mission and professional norms, a phenomenon studied as public service motivation. And the framework can slide from a model into a conclusion, treating any government action as presumptively rent seeking. Used well, it is a set of questions: who gains, who organizes, who pays attention, and what does each actor's incentive predict?
Key idea: Public choice supplies a symmetric standard. Apply to political actors the same assumption of self-interest that you apply to firms, and the burden is on anyone who wants to exempt either side.
Regulatory capture
George Stigler's 1971 economic theory of regulation argued that regulation is often acquired by the industry it governs and operated for its benefit. The mechanism is straightforward: the industry has intense, concentrated interest and technical expertise; the public has diffuse interest and little information; regulators depend on the industry for data and often for future employment.
The evidence is real but more nuanced than the slogan. Historical cases are strong. The Interstate Commerce Commission, created to police railroads, ended up restricting entry and stabilizing rates in ways railroads and later trucking firms welcomed, which is why deregulation in the 1970s and 1980s drew support from economists across the political spectrum, including Alfred Kahn under President Carter. The Minerals Management Service before the Deepwater Horizon disaster is a modern case, as is aviation certification practice examined after the Boeing 737 MAX crashes.
But capture is not universal, and Daniel Carpenter and David Moss led a body of work documenting when it does and does not occur. Agencies with strong professional cultures, transparent processes, organized opposing constituencies, active media attention, and independent scientific capacity resist capture better. Carpenter's own study of the FDA argues the agency built a reputation that gave it power over industry rather than the reverse. So the analytic question is not whether capture happens but under what conditions, and that turns it into an empirical claim you can actually test in a particular case.
Unintended consequences
Robert Merton's 1936 essay on the unanticipated consequences of purposive social action remains the best statement of why interventions surprise their designers: incomplete knowledge, error, the immediate interest overriding the long view, and the way basic values sometimes forbid considering certain consequences at all.
Concrete cases teach the pattern better than the abstraction.
| Policy | Intent | Unintended effect |
|---|---|---|
| Rent control | Keep existing tenants affordable and stable | Most economists find reduced maintenance and reduced long-run supply, though studies find real benefits to incumbent tenants; the distributional split is the whole debate |
| Cobra bounty in colonial Delhi | Reduce the cobra population | People bred cobras for the bounty, and released them when it ended |
| Benefit phase-outs | Target aid to those who need it | High effective marginal tax rates on earnings in the phase-out range |
| Occupational licensing | Assure competence and protect consumers | Restricted entry and higher prices, with contested effects on measured quality |
| Endangered species land restrictions | Protect habitat on private land | Some evidence of preemptive habitat clearing by landowners to avoid designation |
Two things must be said about this list. First, the existence of an unintended consequence does not by itself condemn a policy. Almost every policy has some. The question is whether the whole effect, intended and not, is better than the alternative, which is what Module 4 teaches you to assess. Second, the pattern is predictable enough to design against. Ask before enacting: how will the people affected change their behavior to their own advantage, what would I do if I were subject to this rule, and what does the rule reward that I did not intend to reward?
Holding both critiques at once
Here is the intellectual discipline this module is really teaching. Take a policy you support. Apply the government failure toolkit to it seriously: who is capturing this, what rent is being sought, whose organized interest explains its shape, how will people game it? Now take a policy you oppose. Apply the market failure toolkit seriously: what is the actual failure it addresses, and what would the counterfactual world without it look like?
Most people do the reverse instinctively, deploying market failure arguments for policies they like and government failure arguments for policies they dislike. Charles Wolf, whose work on nonmarket failure formalized the comparison, insisted the only defensible method is to compare imperfect markets with imperfect governments case by case, since both are imperfect and neither is uniformly worse. That comparison rarely yields a clean answer, and the honest output of good analysis is often a well-specified trade-off rather than a recommendation.
Common misconceptions
- Efficiency arguments are objective and equity arguments are political. Both rest on value premises. Kaldor-Hicks in particular embeds a specific and contestable judgment about uncompensated losses.
- Public choice proves government is bad. It supplies a symmetric assumption of self-interest. Used as a conclusion rather than a set of questions, it stops being analysis.
- Regulatory capture is inevitable. Research identifies conditions that make it more or less likely, which makes it a testable claim about a particular agency rather than a general law.
- Unintended consequences prove a policy failed. Nearly all policies have them; the relevant test is the net effect compared with the realistic alternative.
- Redistribution is costless, or redistribution is ruinous. Okun's leaky bucket names a real trade-off whose size is empirically contested, and how much leakage is tolerable is a value judgment.
Recap
- Efficiency criteria are silent on distribution, so equity rationales such as insurance against the lottery of birth, equality of opportunity, and specific egalitarianism stand on their own footing.
- Serious objections to redistribution from Nozick, Hayek, and the efficiency-cost literature are equally value-laden and deserve equal statement.
- Public choice predicts rational ignorance, concentrated-benefit politics, and rent seeking, and its budget-maximizing claim has weaker empirical support than the rest.
- Capture is well documented in some agencies and absent in others, and the useful question is which conditions produce it.
- The analytic discipline is symmetry: compare imperfect markets with imperfect governments case by case, and apply both toolkits hardest to the policies you already favor.
Sources
- Encyclopaedia Britannica. (n.d.). Public choice theory. britannica.com
- Wikipedia contributors. (n.d.). Regulatory capture. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Government failure. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Kaldor-Hicks efficiency. Wikipedia. en.wikipedia.org
- U.S. Government Accountability Office. (n.d.). Reports and testimonies. gao.gov
- Key terms
- Pareto efficiency
- An allocation in which no one can be made better off without making someone worse off; compatible with any distribution, however unequal.
- Kaldor-Hicks criterion
- The test that a change counts as an improvement if winners could in principle compensate losers, whether or not compensation actually occurs.
- Veil of ignorance
- Rawls's device of choosing principles of justice without knowing which position in society you will occupy.
- Leaky bucket
- Okun's image for the efficiency loss incurred when transferring resources from richer to poorer households.
- Rational ignorance
- Downs's observation that individual voters gain too little from becoming informed to justify the cost, given that one vote rarely decides anything.
- Rent seeking
- Spending real resources to obtain favorable treatment rather than to create value, as in lobbying for a protected license.
- Regulatory capture
- Stigler's account of regulation being acquired and operated for the benefit of the regulated industry; empirically variable rather than universal.
- Nonmarket failure
- Wolf's term for systematic shortcomings of government provision, meant to be compared symmetrically with market failure.
Module 3: Agenda Setting and Policy Formulation
How a condition becomes a problem, how a problem reaches the agenda, and how solutions attach themselves to problems. Covers problem definition as political struggle, causal stories and framing, Kingdon's multiple streams, punctuated equilibrium, and what focusing events actually do.
Problem Definition: The Most Political Act in Policy
- Explain the difference between a condition and a problem and why that difference is the central political struggle.
- Analyze causal stories, framing, and indicator choice as strategic moves with predictable consequences.
- Rewrite a problem definition several ways and show how each definition implies a different solution set.
The big picture
Every year in the United States, roughly forty thousand people die in motor vehicle crashes. That number is astonishing. It is roughly the population of a small city, gone annually, and it has been in that range for decades. If forty thousand Americans died each year from a novel infectious disease, it would be the dominant political story of the era. Instead, most years, traffic deaths are a condition rather than a problem: something regrettable that happens, like weather.
Now compare. In the 1960s, traffic deaths were framed as a matter of driver error, which is to say bad drivers. The policy implications were driver education, licensing rules, and campaigns urging care. Then Ralph Nader and a group of engineers and physicians reframed the same deaths as a product design failure: cars had rigid steering columns, dashboards studded with hard protrusions, and no restraints. Under the new definition, the relevant actor was the manufacturer, not the driver. Within a few years Congress created a federal vehicle safety agency and began mandating design standards. Deaths per mile traveled have fallen by roughly ninety percent since the mid twentieth century, driven substantially by exactly these design changes.
Nothing about the deaths changed. What changed was the definition. That is the subject of this lesson, and it is, in Deborah Stone's phrase, the essential political act.
Conditions versus problems
John Kingdon drew the distinction that organizes the field. A condition is a state of the world that people put up with. A problem is a condition that people have come to believe should be acted upon and could be acted upon. The gap between the two is not filled by the severity of the condition. It is filled by argument.
Three things typically move a condition across the line.
The condition is shown to be caused rather than fated. As long as something is understood as natural, accidental, or an act of God, there is nobody to hold responsible and nothing to demand. Once a human cause is identified, a demand becomes possible. Famine understood as crop failure is a tragedy; famine understood, following Amartya Sen's work, as a failure of entitlement and distribution is a policy failure.
The condition violates an important value. Facts alone do not motivate. The claim that something is unfair, unsafe, wasteful, or un-American converts a number into a grievance.
The condition is compared unfavorably. Comparison creates the sense that things could be otherwise. American life expectancy compared with its own past is one story; compared with peer countries at similar income, it is a much more uncomfortable one. Choosing the comparison is itself an argumentative move.
Key idea: Nothing is inherently a policy problem. A condition becomes a problem when it is successfully portrayed as caused, wrong, and fixable, and that portrayal is contested by parties who understand exactly what is at stake in it.
Causal stories
Deborah Stone's analysis of causal stories is the sharpest tool in this lesson. She argues that political actors compete to establish a causal account of harm, because the account determines who is responsible and therefore what should be done. She sorts causal claims along two dimensions: whether the actions producing the harm were intended or unintended, and whether the consequences were intended or unintended.
| Type | Story | Implied response |
|---|---|---|
| Mechanical or accidental cause | Nobody did it; it just happens | Nothing, or private adaptation |
| Inadvertent cause | Harm results from ignorance or carelessness | Education, warnings, better information |
| Intentional cause | Someone knowingly produced the harm | Punishment, prohibition, liability |
| Complex or systemic cause | The harm emerges from an institution or system | Structural reform; but diffuse blame weakens urgency |
Watch this play out in an argument you have certainly heard. Obesity as inadvertent cause implies nutrition labeling and public education. Obesity as intentional cause, produced by firms engineering products for maximum consumption, implies taxes, marketing restrictions, and litigation. Obesity as complex cause, produced by agricultural subsidies, urban design that requires driving, food deserts, work hours, and sleep, implies structural reform across many domains and, awkwardly, no obvious defendant. Each story is partly supported by evidence. Each implies a different politics. And each has organized constituencies who prefer it, which is why the fight over the story is fought as hard as the fight over the policy.
The systemic story has a particular weakness worth naming. It is often the most accurate and the least politically potent, because a diffuse cause gives voters and legislators nobody to be angry at. Analysts who insist on the complicated truth sometimes lose to advocates offering a simpler villain, and that tension is a permanent feature of the job rather than a failure of nerve.
Framing, and what the evidence says about it
Framing means selecting some aspects of a perceived reality and making them more salient, in Robert Entman's formulation, so as to promote a particular problem definition, causal interpretation, or evaluation. It is not the same as lying. A frame can be entirely truthful and still be doing enormous work by directing attention.
The experimental evidence is robust. Amos Tversky and Daniel Kahneman's classic Asian disease experiment presented identical outcomes described as lives saved or lives lost and produced systematically different choices, demonstrating that logically equivalent descriptions can reverse preferences. Political science work by Thomas Nelson and colleagues showed that framing a Ku Klux Klan rally as a free speech issue versus a public order issue moved tolerance for the rally substantially.
Some concrete policy frames and their consequences.
| Same policy | Frame A | Frame B |
|---|---|---|
| Levy on carbon emissions | Carbon tax | Carbon fee and dividend, or pollution charge |
| Levy on large estates | Estate tax | Death tax |
| Public health insurance expansion | Coverage for the uninsured | Government takeover of medicine |
| Work requirement for benefits | Promoting self-sufficiency | Taking food from poor families |
| Immigration enforcement | Rule of law and border security | Family separation |
Read that table carefully and notice something uncomfortable: in each row, both frames describe something real. The estate tax is levied at death. It is also a tax on large inherited wealth. A work requirement does aim at self-sufficiency, and it does remove benefits from people who fail to document compliance, including people who are working. The reason framing is powerful is precisely that the competing frames are usually not lies. They are selections.
An important caution against overreading. Framing effects in laboratory experiments are typically larger than in the world, where people encounter competing frames rather than one. James Druckman's work shows that when a strong opposing frame is present, or when people can consult trusted sources, framing effects shrink substantially. Frames are powerful but not hypnotic, and analysts who assume the public is infinitely manipulable make bad predictions.
Key idea: Frames work by selection rather than falsehood, which is why they survive fact-checking. Competing frames blunt each other, so the practical question is usually which frame reaches an audience unopposed.
Indicators are arguments
Numbers look like the antidote to framing. They are frequently its vehicle, because choosing what to count is itself a definitional act.
Consider unemployment. The Bureau of Labor Statistics publishes several measures. The headline U-3 rate counts people without work who actively searched in the past four weeks. U-6 adds people working part time who want full time work and people marginally attached to the labor force. In a weak economy the two can diverge by many percentage points. Neither is wrong. They answer different questions, and an advocate can pick the one that supports the story.
Consider poverty. The United States has an official poverty measure, derived from a 1960s calculation based on food budgets and adjusted for inflation, which counts pre-tax cash income. It does not count SNAP benefits, refundable tax credits, or housing assistance, and it does not adjust for local cost of living. The Census Bureau also publishes a Supplemental Poverty Measure that does all of those things. These two measures can tell substantially different stories about whether antipoverty policy is working, because the official measure by construction cannot detect the effect of the largest antipoverty programs, which operate through in-kind benefits and the tax code. When you read a claim that poverty has or has not fallen, the first question is which measure.
Consider crime. Reported crime from law enforcement agencies and victimization from household surveys measure different things, and the gap between them is itself informative, since it partly reflects willingness to report. A change in reporting practice can produce a crime wave in the statistics with no change in the world.
This is not a counsel of despair about numbers. It is a counsel of specificity. Always ask: what exactly is counted, who does the counting, what is excluded by the definition, and would a different reasonable definition tell a different story? If the answer to the last question is yes, say so in your analysis.
Redefining a problem: a worked exercise
Take a single condition and write it five ways. Suppose in a mid-sized city, emergency departments see a high volume of visits from a small number of people, many of them unhoused, most for conditions that could be treated elsewhere.
- As a hospital cost problem: uncompensated care strains hospital finances. Solutions point at reimbursement policy and care diversion.
- As a primary care access problem: people use the emergency department because nothing else is open or will take them. Solutions point at clinic hours, capacity, and Medicaid acceptance rates.
- As a housing problem: people without stable housing cannot manage chronic conditions. Solutions point at supportive housing, and there is real evidence that permanent supportive housing reduces emergency use for high utilizers, though evidence on total cost savings is more mixed than advocates often claim.
- As a behavioral health problem: untreated mental illness and substance use disorder drive much of the volume. Solutions point at treatment capacity and crisis response teams.
- As a public safety problem: this framing casts the same people as a disorder issue, and points at policing and enforcement.
Five definitions, five budgets, five lead agencies, five sets of winners and losers, one set of facts. Now notice that a good analyst does not simply pick a favorite. A good analyst tells the decision maker that the definition is contested, lays out what each definition implies, and where possible shows which definitions are better supported by the data on who these patients actually are. Definitions can be more or less accurate. They are just never determined by the facts alone.
Common misconceptions
- The biggest problems get the most attention. Attention tracks framing, salience, and organization far more than magnitude, which is why forty thousand annual traffic deaths draw less notice than far rarer hazards.
- Framing means spin or dishonesty. Effective frames are usually true selections rather than falsehoods, which is exactly why fact-checking rarely defeats them.
- Statistics settle problem definition. Choosing an indicator is a definitional act, as the gap between the official and supplemental poverty measures shows.
- The most accurate causal story wins. Systemic explanations are often the most accurate and the least mobilizing, because diffuse blame produces no defendant.
- Framing effects make the public infinitely manipulable. Competing frames and trusted sources substantially shrink framing effects outside the laboratory.
Recap
- Conditions become problems through argument, not severity; the claim must establish that the harm is caused, wrong, and fixable.
- Stone's causal stories determine who is blamed and therefore what remedies appear appropriate, which is why the causal fight is fought so hard.
- Frames select rather than falsify, so competing frames blunt each other and unopposed frames dominate.
- Indicators embed definitions, as with U-3 versus U-6 unemployment and official versus supplemental poverty; always ask what is counted and what is excluded.
- An analyst's job is to make the contest over definition visible, not to smuggle in a preferred definition as though it were the facts.
Sources
- U.S. Census Bureau. (n.d.). How the Census Bureau measures poverty. census.gov
- U.S. Bureau of Labor Statistics. (n.d.). Labor force statistics from the Current Population Survey: Concepts and definitions. bls.gov
- National Highway Traffic Safety Administration. (n.d.). Traffic safety research and data. nhtsa.gov
- Wikipedia contributors. (n.d.). Framing (social sciences). Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Agenda-setting theory. Wikipedia. en.wikipedia.org
- Key terms
- Condition versus problem
- Kingdon's distinction between a state of the world people tolerate and one they believe should and can be acted upon.
- Causal story
- Stone's term for a contested account of how a harm was produced, which determines who is blamed and what remedies seem appropriate.
- Framing
- Selecting and making salient certain aspects of a situation so as to promote a particular definition, interpretation, or evaluation.
- Indicator
- A measure chosen to represent a condition; because measurement choices embed definitions, indicators are arguments as well as data.
- U-3 and U-6
- Two Bureau of Labor Statistics unemployment measures, the narrower counting active searchers and the broader adding involuntary part-time and marginally attached workers.
- Supplemental Poverty Measure
- A Census measure counting in-kind benefits, tax credits, and local costs, which the official measure by construction excludes.
- Salience
- How prominent an issue is in public and elite attention, which drives what officials feel compelled to address.
How Issues Rise: Multiple Streams, Punctuated Equilibrium, and Focusing Events
- Explain Kingdon's three streams and the conditions under which a policy window opens.
- Describe punctuated equilibrium and the role of policy monopolies, venue shopping, and image change.
- Evaluate what focusing events do and do not accomplish, using contrasting real cases.
The big picture
Two questions about agendas are easy to confuse. The first is why some problems get attention. Lesson 6 answered that. The second is why attention translates into action at one moment and not another, when the problem was equally well understood for years beforehand.
Consider seat belts. The physics of restraint were understood in the 1950s. Volvo introduced the three-point belt in 1959 and made the patent freely available. Yet universal American state laws requiring belt use took until the 1980s and 1990s. Or consider the Affordable Care Act. Proposals for near-universal health coverage in the United States date to the Truman administration and earlier, and every serious attempt for sixty years failed. Then in 2010 one passed. The problem had not changed much. Something else had.
The frameworks in this lesson are attempts to explain that timing, and they share a common insight: policy change is not proportional to problem severity. Systems ignore problems for long stretches and then move suddenly, and the mechanism of that sudden movement is what we are after.
Kingdon's multiple streams
John Kingdon's 1984 study of health and transportation policy produced the most used framework in the field. He observed that federal officials could not usually tell him why a given idea rose when it did, which pointed him away from rational sequence models and toward a looser process adapted from the garbage can model of organizational choice.
Kingdon proposed three largely independent streams flowing through the system at all times.
The problem stream contains conditions that have been defined as problems. Kingdon identified three ways conditions get noticed: indicators that shift, such as a rising cost or infection curve; focusing events such as disasters and crises; and feedback from existing programs that reveals failure.
The policy stream contains proposals circulating among specialists: academics, agency staff, think tank researchers, congressional staff, and interest group analysts. Kingdon's memorable image is a policy primeval soup in which ideas float, combine, and recombine. Ideas survive by meeting criteria the community applies: technical feasibility, value acceptability among specialists, tolerable cost, and anticipated public and political acquiescence. Crucially, this stream evolves on its own clock. Proposals get worked out long before any problem is attached to them, a process Kingdon called softening up.
The political stream contains the national mood, election results, changes in administration and congressional control, and organized interest campaigns. It moves independently of whether any problem has worsened or any solution has improved.
A policy window opens when the streams couple. A problem is prominent, a worked-out solution is available, and the political conditions permit. Windows open predictably sometimes, as with scheduled reauthorizations and budget cycles, and unpredictably other times, after a disaster or an election. They close fast, often within months, because attention moves on, the political configuration shifts, or a crisis is deemed handled.
The person who joins the streams is the policy entrepreneur: someone who invests time, reputation, and money in advancing a proposal, and who has a solution ready and waiting for a problem to attach it to. This is the framework's most counterintuitive claim and its most useful one. Solutions frequently precede problems. Entrepreneurs do not respond to problems by inventing solutions; they carry solutions around looking for problems that can carry them.
Key idea: Policy change requires the coupling of an acknowledged problem, an available worked-out solution, and favorable politics. Because the three run on separate clocks, timing dominates merit, and a good idea with no window waits.
Reading the ACA through multiple streams
Apply it to 2009 and 2010 and the framework earns its keep.
The problem stream: the uninsured population had risen past forty-five million, health spending was approaching eighteen percent of GDP, and the 2008 financial crisis made job-linked coverage visibly fragile. Indicators and feedback both pointed the same way.
The policy stream: the specific architecture, regulated exchanges, an individual mandate, subsidies, and guaranteed issue with community rating, had been developed over two decades in policy circles, versions of it had appeared in proposals from across the spectrum, and it had been implemented in Massachusetts in 2006 with observable results. It was, in Kingdon's terms, thoroughly softened up.
The political stream: the 2008 election produced a president who had campaigned on it and, briefly, sixty Senate votes.
The window opened and closed quickly. When a special election in Massachusetts removed the sixtieth vote in January 2010, the House was forced to accept the Senate bill unchanged and use reconciliation for amendments, which is why the statute has the structure it does. The procedural constraint from Lesson 2 is visible in the final text. Note also what this analysis does not do: it says nothing about whether the ACA is good policy. Module 6 takes that up with the evidence on both sides. Multiple streams explains timing, not merit.
Punctuated equilibrium
Frank Baumgartner and Bryan Jones borrowed a term from evolutionary biology to describe a pattern they found across decades of American policy data: long periods of stability with incremental adjustment, interrupted by rare and dramatic bursts of change. Their empirical work, drawing on the Policy Agendas Project's coding of hearings, budgets, and coverage, found this pattern is not occasional but general. Budget change distributions have far more tiny changes and far more enormous changes than a normal distribution would predict.
Their explanation has two parts.
The first is the policy monopoly. Established policy areas develop a stable set of participants, usually an agency, the relevant congressional committees, and organized interests, who share an understanding of the issue and control the venue in which it is discussed. Combined with a supporting policy image, meaning a widely accepted way of understanding what the issue is about, the monopoly keeps outsiders out and change incremental.
The second is that monopolies break when the image changes or the venue changes. Baumgartner and Jones's central case is civilian nuclear power. For decades the image was clean, cheap, modern energy, and the venue was a supportive congressional committee and a promotional agency. Beginning in the late 1960s critics attacked the image, emphasizing safety, waste, and cost overruns, and moved the fight to new venues: state utility commissions, courts, and local siting hearings. The monopoly collapsed, and no new American reactor was ordered and completed for decades.
Venue shopping is the deliberate strategy behind this. If you lose in Congress, try the agencies. If you lose there, try the courts. If you lose federally, try the states or a city council. American federalism and separated powers make this unusually easy, which is one reason American policy conflicts are so persistent: losing in one arena is rarely final.
A related engine is attention scarcity. Jones emphasized that legislatures can attend to only a few issues at once, so agenda access is a genuinely rival good. Issues do not fail to advance only because of opposition; they fail because attention is finite and something else is occupying it.
Key idea: Stability is maintained by policy monopolies with a supporting image and a friendly venue. Change comes when challengers successfully redefine the image or move the fight to a venue where the monopoly does not control the outcome.
What focusing events actually do
Thomas Birkland studied disasters and crises systematically and found the popular intuition is only partly right. Focusing events do reliably increase attention. They do not reliably produce policy change.
The comparison that teaches this best is a pair of cases.
| Event | Attention | Policy result | Why |
|---|---|---|---|
| Deepwater Horizon spill, 2010 | Enormous | Significant regulatory reorganization and new drilling safety rules | A prepared community of experts had proposals ready; a single identifiable industry and agency were implicated |
| Repeated mass shootings | Enormous | Little federal legislative change for long stretches, though state-level change has been substantial in both directions | Organized groups on both sides, contested causal stories, no consensus solution in the policy stream |
Birkland's conclusion is that a focusing event opens a window; it does not fill one. If the policy community has already softened up a solution, the event lets it through. If the community is divided about the causal story or has no agreed remedy, the window opens and closes with nothing passing through it, and the same event recurs later with the same result. This is a genuinely important finding for anyone who expects that a sufficiently terrible event will eventually force action. Kingdon and Birkland both predict it will not, absent a prepared solution.
There is a further asymmetry. Events that fit an existing image reinforce it; events that contradict it are often absorbed as anomalies. Whether a given event becomes focusing is itself partly a product of framing, which returns us to Lesson 6. Nothing in this course lets you skip the definitional fight.
Other frameworks in brief
Two more deserve mention because you will encounter them.
The advocacy coalition framework, developed by Paul Sabatier and Hank Jenkins-Smith, organizes a policy subsystem into coalitions bound by shared beliefs rather than shared material interests. It distinguishes deep core beliefs, which almost never change, from policy core beliefs and secondary aspects, which can change through policy-oriented learning. It predicts that learning happens mostly on secondary matters, and that major change requires external shocks. It is especially useful for long-running technical conflicts such as water rights or air quality.
Incrementalism, from Charles Lindblom's 1959 essay on muddling through, is less a theory of change than a description of normal practice: decision makers compare a few alternatives marginally different from the status quo, because comprehensive rationality exceeds anyone's information and time. Lindblom argued this is not merely a limitation but often sensible, since small steps are reversible and errors are cheaper. Punctuated equilibrium can be read as incrementalism plus an account of when it breaks.
Common misconceptions
- Serious problems eventually force action. Without an available solution and favorable politics, severity alone does not move an issue, which is why some large problems persist for decades.
- Solutions are designed in response to problems. Kingdon found solutions typically circulate first and get attached to problems when a window opens.
- Disasters produce reform. Birkland found they open windows; whether anything passes through depends on whether the policy community has a prepared, agreed remedy.
- Policy change is gradual. Baumgartner and Jones documented distributions with far more tiny and far more enormous changes than gradualism predicts.
- Losing in one venue settles a question. Venue shopping across chambers, agencies, courts, and levels of government makes most American policy defeats provisional.
Recap
- Kingdon's problem, policy, and political streams run independently, and change requires a policy entrepreneur to couple them when a window opens.
- Windows are brief, sometimes scheduled and sometimes triggered, and the ACA's final structure shows how a closing window shapes the text of a law.
- Punctuated equilibrium explains long stability through policy monopolies with a supporting image and a friendly venue.
- Monopolies break through image change and venue shopping, both of which the fragmented American system makes easy.
- Focusing events raise attention reliably but produce change only when a prepared solution and a settled causal story already exist.
Sources
- Wikipedia contributors. (n.d.). Multiple streams framework. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Punctuated equilibrium in social theory. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Advocacy coalition framework. Wikipedia. en.wikipedia.org
- Congressional Research Service. (n.d.). Reports on health policy and the Affordable Care Act. crsreports.congress.gov
- Encyclopaedia Britannica. (n.d.). Patient Protection and Affordable Care Act. britannica.com
- Key terms
- Problem stream
- Kingdon's flow of conditions defined as problems through shifting indicators, focusing events, and program feedback.
- Policy stream
- The circulating supply of proposals among specialists, described by Kingdon as a primeval soup in which ideas combine and get softened up.
- Political stream
- National mood, election results, changes in control, and organized campaigns, moving independently of problems and solutions.
- Policy window
- A brief opportunity when problem, policy, and political streams can be coupled and change becomes possible.
- Policy entrepreneur
- An actor who invests resources to couple the streams, typically carrying a prepared solution in search of a problem.
- Policy monopoly
- A stable set of participants controlling an issue's venue and supported by a shared policy image, which keeps change incremental.
- Venue shopping
- Moving a policy conflict to a different institutional arena where the existing monopoly does not control the outcome.
- Focusing event
- A sudden, harmful, widely known event that raises attention; per Birkland it opens a window but does not itself produce change.
- Incrementalism
- Lindblom's account of decision makers comparing a few alternatives marginally different from the status quo, which he defended as often sensible.
Module 4: The Analyst's Toolkit
The craft itself, worked concretely. Constructing genuine alternatives and criteria, running a cost-benefit analysis with real discounting arithmetic, reasoning carefully about the value of a statistical life, using cost-effectiveness where valuation fails, exposing distribution, and knowing exactly where quantification stops helping.
From Problem to Alternatives: Defining the Question and Building the Options
- Write a problem statement that is specific, measurable, and does not smuggle a solution inside it.
- Construct a genuine set of alternatives including a properly specified baseline, avoiding the standard failure modes.
- Select and define evaluative criteria, and build an outcomes matrix that exposes trade-offs rather than hiding them.
The big picture
Suppose a mayor calls you in and says: we need more police officers downtown. Notice what has happened. You have been handed a solution wearing the costume of a problem. If you accept it, your analysis will be about how many officers and at what cost, and every alternative that does not involve officers has been eliminated before you started.
The single most valuable thing a policy analyst does happens in the first hour, before any spreadsheet opens: converting a demanded solution back into the problem it is meant to solve, and then generating the full set of things that could solve it. Eugene Bardach, whose eightfold path is the most widely taught practical method in the field, puts the warning bluntly: do not let the client's preferred solution define the problem, and do not define the problem in terms of a solution.
So you ask the mayor what she is worried about. It turns out that downtown retail vacancy has risen, merchants report shoplifting, and two assaults in the past year drew heavy coverage. Now you have something to work with. Notice that the problem might be crime, or it might be perception of crime, or it might be retail economics that has nothing to do with either, and those are distinguishable with data. That is Lesson 6's problem definition, now with a client waiting.
Writing the problem statement
A usable problem statement has four properties.
It is specific about magnitude and trend. Not downtown crime is a problem, but reported thefts from downtown retail establishments rose from 180 in 2021 to 310 in 2024, while citywide theft rose eleven percent over the same period. Numbers with a comparison give you something to test against later.
It names whose problem it is. Merchants losing inventory, employees who feel unsafe, residents whose evening use of downtown has dropped, and city finances dependent on the sales tax base are four different constituencies with four different stakes. They may want different things.
It contains no solution. Insufficient police presence is not a problem statement; it is a hypothesis about a cause and a solution in disguise.
Bardach adds a fourth discipline that students find surprisingly hard: quantify wherever possible even when the number is rough, because a rough magnitude tells you what scale of response is plausible. If total annual retail theft losses downtown are estimated at three hundred thousand dollars, a two-million-dollar program needs to justify itself on grounds other than the theft losses.
Key idea: The problem statement determines the entire analysis. If a solution is embedded in it, every alternative that solution excludes has been eliminated before the work begins.
Assembling evidence, not data
Bardach distinguishes data from information from evidence. Data are facts about the world. Information is data with meaning. Evidence is information that affects someone's belief about something that matters to the decision. Most analytic time is wasted collecting data that would not change any recommendation regardless of how it came out.
The practical test is to ask, before collecting anything: if this number came back high, what would I recommend, and if it came back low, what would I recommend? If the answer is the same, do not collect it. If a stakeholder cannot be moved by any value of the number, note that too, because it tells you which disagreements are empirical and which are about values, and only one of those kinds is addressable by more research.
Time is always the binding constraint. Real analysis happens under deadline, and the professional skill is knowing when the marginal hour of research stops improving the decision. Bardach's advice is to think about what you would do with the information before you go get it, and to start writing early because writing reveals which gaps actually matter.
Constructing alternatives
Here is where most student analyses fail, and the failures are predictable enough to name.
Failure one: the false pair. Presenting only do this or do nothing. Almost no real policy question has two options, and a two-option analysis is usually an advocacy document with extra steps.
Failure two: alternatives that differ only in size. Hire ten officers, twenty officers, or thirty officers is one alternative at three levels, not three alternatives. Scale variation belongs inside an option, not across the set.
Failure three: the straw alternative. Including an obviously terrible option so the preferred one looks good. Reviewers notice, and it destroys your credibility on everything else in the document.
Failure four: omitting the baseline. The status quo is always an alternative and must be specified as carefully as the others. Note that the baseline is not the world as it is today; it is the world as it will be if you do nothing, which usually means projecting current trends forward. If downtown theft is rising, the do-nothing baseline is more theft, not today's theft.
Techniques for generating better options: look at what other jurisdictions do, since somebody has almost certainly tried something; go back to your causal story from Module 3 and generate an option per link in the causal chain; run through the five instruments from Lesson 1 and force yourself to write an option for each; and separate the ends from the means by asking what else would produce the same outcome.
Applied to downtown: increased patrol presence is one option. Others include environmental design changes such as lighting and sightlines, a merchant-funded business improvement district with private security and cleaning, targeted social services for a small number of high-contact individuals, retail-specific loss prevention support and training, changes to prosecution thresholds, and a downtown activation strategy that raises foot traffic through events and residential conversion, on the theory that occupied streets are safer streets. Notice that these come from entirely different causal stories, and several could be combined.
Key idea: A real alternative set contains options that differ in kind, not in dose, includes a properly projected baseline, and excludes straw options. Generate options from causal links and from the instrument list, not from what the client already said.
Selecting criteria
Criteria are the standards by which you will judge options, and they must be selected before you evaluate anything, for the same reason a scientist preregisters a hypothesis. Choosing criteria after seeing results is how motivated reasoning enters analysis.
| Criterion | Question it asks | Typical measure |
|---|---|---|
| Effectiveness | Does it achieve the outcome, and by how much? | Reduction in the target indicator |
| Efficiency | Does the benefit exceed the cost? | Net present value, benefit-cost ratio, cost per unit of outcome |
| Equity | Who gains and who bears the cost? | Distribution of effects by income, place, race, or generation |
| Administrative feasibility | Can the implementing body actually do this? | Staffing, systems, timeline, prior experience |
| Political feasibility | Can it be adopted and survive? | Veto point analysis from Lesson 2 |
| Robustness and reversibility | What happens if key assumptions are wrong? | Sensitivity analysis; cost of unwinding |
| Legality | Is there authority to do it? | Statutory basis; litigation exposure |
Two cautions. First, do not use so many criteria that the matrix becomes unreadable; four to six well defined ones usually beat ten vague ones. Second, and more important, resist the temptation to collapse criteria into a single weighted score. Weighted scoring feels rigorous and hides exactly the judgment the decision maker is supposed to make. If effectiveness and equity point in opposite directions, saying so is the finding. Averaging them with weights you invented buries it.
Projecting outcomes and confronting trade-offs
Projection is the hardest step and the one where honesty matters most, because you are forecasting the future and you will often be wrong. Three practices help.
State magnitudes with ranges, not points. A projection of a fifteen to twenty-five percent reduction, based on a named source, is more useful and more honest than a single number with implied precision.
Name the source of each estimate. Evaluated program elsewhere, expert judgment, extrapolation from a related program, and pure assumption are four very different epistemic statuses, and the reader deserves to know which they are looking at.
Apply a break-even test when you cannot estimate. Instead of guessing how much a program reduces theft, ask how much reduction would be required for the program to pay for itself, then ask whether that magnitude is plausible given what similar programs achieve. Break-even framing converts an unanswerable estimation problem into an answerable plausibility judgment, and it is one of the most useful moves in the entire craft.
The outcomes matrix puts options in rows and criteria in columns. Its purpose is not to produce a winner; it is to make the trade-off legible. When one option dominates, meaning it is at least as good on every criterion and better on at least one, say so, and note that this is rare. When no option dominates, the matrix shows the decision maker precisely what they are choosing between, which is the real deliverable.
Telling the story
Bardach's final step, tell your story, is not a communications afterthought. A recommendation the decision maker cannot follow is worthless, and busy officials read the first page.
The house conventions of the field: lead with the recommendation and the reason, not the methodology; put the trade-off in the first paragraph, since a memo that only lists benefits will be distrusted; write for someone with no background in your analysis and no time; make the tables self-explanatory; and state your assumptions and your uncertainty explicitly, including what would change your recommendation. That last item, naming what evidence would change your mind, is the clearest signal that you did analysis rather than advocacy.
One more professional norm, taken up fully in Lesson 15. If you have a personal stake or a strong value commitment in the outcome, disclose it. The analyst's authority rests entirely on the reader's belief that the analysis would have come out differently had the evidence differed.
Common misconceptions
- The client's request is the problem. Clients usually hand you a preferred solution; converting it back into a problem statement is the first analytic act.
- Three funding levels are three alternatives. Dose variation belongs within an option; genuine alternatives differ in mechanism.
- The baseline is the world as it is now. It is the projected world without intervention, which for a worsening trend is materially different from today.
- A weighted score makes analysis objective. Weights are values in numerical costume, and collapsing criteria hides the judgment the decision maker should make.
- More data always improves analysis. Data that cannot change any recommendation is a cost with no benefit; ask first what you would do with each possible answer.
Recap
- A good problem statement is specific about magnitude and trend, names affected parties, and contains no embedded solution.
- Evidence is information that could change a recommendation; test each proposed data collection against that standard before spending time on it.
- Real alternative sets differ in mechanism, include a projected baseline, and avoid straw options; generate them from causal links and the instrument list.
- Criteria must be chosen before evaluation, kept few and well defined, and never collapsed into a single weighted score that hides trade-offs.
- Projections should carry ranges and named sources, break-even framing rescues unanswerable estimates, and the memo must state what would change your mind.
Sources
- U.S. Government Accountability Office. (n.d.). Reports and testimonies. gao.gov
- Congressional Budget Office. (n.d.). How CBO prepares cost estimates. cbo.gov
- Office of Management and Budget. (2003). Circular A-4: Regulatory analysis. The White House. whitehouse.gov archive
- Wikipedia contributors. (n.d.). Policy analysis. Wikipedia. en.wikipedia.org
- Key terms
- Problem statement
- A specific, quantified description of a condition and its trend that names affected parties and contains no embedded solution.
- Baseline
- The projected state of the world absent intervention, which for a worsening condition differs materially from present conditions.
- Evidence
- Information capable of changing someone's belief about something that matters to the decision, as distinct from data generally.
- Alternative
- A distinct mechanism for addressing the problem; options that differ only in funding level are one alternative at several doses.
- Criteria
- Standards for judging options, selected before evaluation, typically including effectiveness, efficiency, equity, feasibility, and robustness.
- Outcomes matrix
- A table of options against criteria whose purpose is to make trade-offs legible rather than to compute a winner.
- Dominance
- The rare situation in which one option is at least as good on every criterion and strictly better on at least one.
- Break-even analysis
- Asking how large an effect would have to be for a policy to justify its cost, converting an estimation problem into a plausibility judgment.
Cost-Benefit Analysis: Discounting and a Worked Example
- Carry out a complete cost-benefit calculation including present value discounting, net present value, and the benefit-cost ratio.
- Explain why the discount rate is a contested value choice and show numerically how much it can change a conclusion.
- Run sensitivity analysis and interpret a result whose sign flips across defensible assumptions.
The big picture
Cost-benefit analysis has a bad reputation among people who have never done one and an ambivalent reputation among people who have. Both reactions are informative. The technique does something genuinely valuable: it forces every consequence of a policy onto a common scale so they can be compared, and it forces the analyst to state assumptions explicitly enough that someone else can disagree with them precisely. It also does something genuinely dangerous: it produces a single number that looks authoritative and can conceal a chain of contestable judgments.
This lesson teaches you to do one and, just as importantly, to take one apart. By the end you should be able to look at any benefit-cost ratio and ask the four questions that usually determine the answer: what was counted, what was left out, what discount rate was used, and over what time horizon.
Cost-benefit analysis is not an academic curiosity in the United States. Since Executive Order 12866 in 1993, and under predecessors going back to Reagan, significant federal regulations have required a regulatory impact analysis reviewed by the Office of Information and Regulatory Affairs, following the methods set out in OMB Circular A-4. Every administration of both parties has retained this requirement. If you work anywhere near federal regulation, you will read these documents.
The logic, stated plainly
The underlying criterion is Kaldor-Hicks from Lesson 5: a policy is worth doing if total benefits exceed total costs, meaning winners could in principle compensate losers. Everything is converted to money, not because money is the only thing that matters, but because comparison requires a common unit and money is the unit for which we have the most information about what people are willing to trade.
The steps are mechanical once the judgments are made.
- Specify the policy and the baseline precisely, since all effects are differences between them.
- Identify all consequences to whomever they accrue, including effects on people outside the jurisdiction and on future generations.
- Quantify each in physical units: tons of emissions, hours of travel time, cases of illness.
- Monetize each using prices, willingness to pay, or established unit values.
- Discount future amounts to present value.
- Compute net present value and the benefit-cost ratio.
- Test the result against alternative assumptions.
- Report distribution separately, because the ratio deliberately ignores it.
Discounting: the arithmetic
A dollar next year is worth less than a dollar today, and not only because of inflation. Cost-benefit analysis is normally done in constant dollars, so inflation is already removed. The remaining reasons are that resources invested today earn a real return, so a dollar today can become more than a dollar later, and that people genuinely prefer benefits sooner, a preference economists call pure time preference.
The present value of an amount received in year t at real discount rate r is the amount divided by one plus r raised to the power t.
Work an example. A benefit of one million dollars arriving in year 20.
- At three percent: 1.03 raised to the 20th power is about 1.806, so the present value is about 553,700 dollars.
- At seven percent: 1.07 raised to the 20th power is about 3.870, so the present value is about 258,400 dollars.
The same future million is worth more than twice as much under one defensible rate as under another. Nothing about the world changed. Sit with that, because it is the whole reason the discount rate is contested.
For a stream of equal annual amounts over n years, you can use the annuity factor, which is one minus one over the quantity one plus r raised to the nth power, all divided by r. Multiply the annual amount by that factor to get the present value of the whole stream.
A complete worked example
A state agency proposes an energy efficiency retrofit of its public buildings. Assume the following, all in constant dollars.
- Capital cost in year zero: 10,000,000 dollars.
- Gross energy savings: 1,200,000 dollars per year in years 1 through 15.
- Added maintenance cost: 150,000 dollars per year in years 1 through 15.
- Net annual benefit: 1,050,000 dollars per year for 15 years.
- No residual value at year 15.
At a three percent discount rate. One plus r raised to the 15th power is about 1.5580. The annuity factor is one minus one divided by 1.5580, which is 0.3581, divided by 0.03, giving about 11.938. Present value of benefits is 1,050,000 times 11.938, or about 12,535,000 dollars. Net present value is 12,535,000 minus 10,000,000, or about positive 2,535,000 dollars. The benefit-cost ratio is about 1.25. The project passes.
At a seven percent discount rate. One plus r raised to the 15th power is about 2.7590. The annuity factor is one minus one divided by 2.7590, which is 0.6376, divided by 0.07, giving about 9.108. Present value of benefits is 1,050,000 times 9.108, or about 9,563,000 dollars. Net present value is about negative 437,000 dollars. The benefit-cost ratio is about 0.96. The project fails.
At a two percent discount rate. The annuity factor is about 12.849, present value of benefits is about 13,492,000 dollars, net present value is about positive 3,492,000 dollars, and the ratio is about 1.35. The project passes comfortably.
| Discount rate | PV of net benefits | Net present value | Benefit-cost ratio | Verdict |
|---|---|---|---|---|
| 2 percent | 13,492,000 | +3,492,000 | 1.35 | Passes |
| 3 percent | 12,535,000 | +2,535,000 | 1.25 | Passes |
| 7 percent | 9,563,000 | -437,000 | 0.96 | Fails |
This is not a contrived illustration. Front-loaded costs with long streams of modest benefits is the characteristic shape of infrastructure, prevention, education, and climate policy, and it is exactly the shape most sensitive to the discount rate. Any analyst who reports only one rate for a project like this is, wittingly or not, making the decision.
Key idea: When costs come first and benefits accrue over a long horizon, the discount rate can determine the sign of the answer. Report a range of rates, always, and show where the result flips.
Why the rate is contested
OMB Circular A-4 in its 2003 form directed agencies to present results at both three percent and seven percent. The seven percent figure was meant to approximate the pre-tax real return to private capital, on the theory that regulation displaces private investment. The three percent figure approximated the real rate of return on long-term government debt, taken as the social rate of time preference for consumption. In 2023 OMB revised the guidance toward a lower base rate of about two percent, reflecting the sustained decline in real returns on long-term Treasury securities over the intervening two decades. The revision was itself politically contested, because lower rates make long-horizon regulations look more attractive.
Beneath the technical dispute sits a genuine ethical one, sharpest in climate policy. The Stern Review of 2006 used a very low pure time preference rate on the argument that there is no defensible ethical reason to count the welfare of a person born in 2100 for less than that of a person alive now, and concluded that aggressive early action was justified. William Nordhaus, who later received the Nobel Prize, argued that the discount rate should be inferred from observed market behavior rather than asserted from ethics, and with market-based rates found a more gradual optimal path. Both are serious economists. Their disagreement is not primarily about climate science or about arithmetic. It is about whether the discount rate is a descriptive parameter or a moral commitment, and no amount of additional data resolves it.
A related refinement: many economists now favor declining discount rates over very long horizons, applying a lower rate to effects centuries away, partly because uncertainty about future interest rates mathematically implies a lower effective rate at long horizons. Several governments, including the United Kingdom in its Green Book guidance, use schedules of this kind.
Sensitivity analysis
Because a cost-benefit analysis is a chain of estimates, the responsible practice is to show how the conclusion responds when each link is varied.
One-way sensitivity varies a single parameter across its plausible range while holding others fixed. It answers which assumptions matter, and often reveals that only two or three drive the entire result.
Break-even or switching-value analysis, from Lesson 8, finds the value at which the conclusion changes. In the retrofit example, you can report that the project breaks even at a discount rate of roughly six and a half percent, and that at seven percent annual net savings would need to be about 1,098,000 dollars rather than 1,050,000 to break even. Those statements are far more useful to a decision maker than a single net present value, because they say exactly how much room for error there is.
Monte Carlo simulation assigns probability distributions to uncertain inputs and simulates many draws, producing a distribution of outcomes rather than a point. It is more informative and also more easily misread, since it can convey false precision if the input distributions were themselves guesses.
Key idea: The output of a cost-benefit analysis is not a number but a conditional statement: under these assumptions, the net present value is this, and it flips sign when these specific parameters cross these specific values.
What cost-benefit analysis reliably misses
Even done well, the technique has structural blind spots, and an honest analyst states them in the document rather than leaving them for critics to find.
It ignores distribution by construction. A positive net present value is compatible with severe losses concentrated on people least able to bear them. Lesson 10 takes up how to add distributional analysis alongside it.
It systematically undercounts what is hard to monetize. Effects with no market analogue, such as dignity, community cohesion, cultural loss, or procedural fairness, get a qualitative footnote while readily priced effects get a number, and numbers dominate footnotes in practice. This is a bias with a direction, not random noise.
It treats a dollar as a dollar regardless of whose it is, though a dollar means far more to a poor household than a rich one. Some frameworks address this with distributional weights, which are technically straightforward and politically fraught, since the weights are explicit value judgments.
It struggles with irreversibility and catastrophic risk. Standard expected-value reasoning handles a small chance of enormous, permanent loss poorly, which is why Martin Weitzman argued that fat-tailed catastrophic risks can dominate a climate calculation in ways ordinary discounting does not capture.
None of this means the technique should be abandoned. The alternative to explicit, checkable assumptions is implicit, uncheckable ones. It means the analysis is an input to judgment rather than a substitute for it, which is precisely what Circular A-4 itself says.
Common misconceptions
- Discounting is just adjusting for inflation. Cost-benefit analysis uses constant dollars, so inflation is already removed; discounting reflects opportunity cost of capital and time preference.
- The discount rate is a technical parameter with a correct value. Its choice embeds a judgment about how much to weight future people, which is why Stern and Nordhaus disagreed without either making an arithmetic error.
- A positive net present value means a policy should be adopted. It means aggregate benefits exceed aggregate costs under stated assumptions; distribution, feasibility, and unmonetized effects are separate questions.
- The benefit-cost ratio is the main result. Net present value is generally the better decision criterion, since ratios can be manipulated by classifying an item as a negative benefit rather than a cost.
- Unquantified effects are treated neutrally. They are systematically underweighted, because a qualitative caveat rarely outweighs a bolded number.
Recap
- Cost-benefit analysis implements Kaldor-Hicks by converting all consequences to a common monetary scale and comparing against a specified baseline.
- Present value discounting can more than halve the weight of a benefit twenty years out, and in the worked retrofit example it flips the verdict between two and seven percent.
- OMB guidance has used three and seven percent and more recently a two percent base, and the Stern-Nordhaus dispute shows the choice is partly ethical rather than purely technical.
- Sensitivity and switching-value analysis convert a fragile point estimate into a useful conditional statement about how much room for error exists.
- The technique structurally ignores distribution, undercounts unmonetized effects, and handles catastrophic irreversible risk poorly, all of which belong in the document.
Sources
- Office of Management and Budget. (2003). Circular A-4: Regulatory analysis. The White House. whitehouse.gov archive
- Office of Management and Budget. (n.d.). Office of Management and Budget. The White House. whitehouse.gov
- Congressional Budget Office. (n.d.). How CBO prepares cost estimates. cbo.gov
- U.S. Environmental Protection Agency. (n.d.). Guidelines for preparing economic analyses. epa.gov
- Wikipedia contributors. (n.d.). Cost-benefit analysis. Wikipedia. en.wikipedia.org
- Key terms
- Present value
- The value today of an amount arriving in the future, computed by dividing by one plus the discount rate raised to the number of years.
- Net present value
- Discounted benefits minus discounted costs; generally the preferred decision criterion because it is not manipulable by reclassifying items.
- Benefit-cost ratio
- Discounted benefits divided by discounted costs; intuitive but sensitive to whether an effect is booked as a cost or a negative benefit.
- Discount rate
- The annual rate at which future amounts are converted to present value, reflecting capital opportunity cost and time preference.
- Annuity factor
- The multiplier converting a constant annual stream into present value, equal to one minus one over the growth factor, divided by the rate.
- Sensitivity analysis
- Varying uncertain inputs across plausible ranges to determine which assumptions drive the conclusion.
- Switching value
- The parameter value at which a conclusion reverses, such as the discount rate at which net present value hits zero.
- Regulatory impact analysis
- The formal cost-benefit document required for significant federal rules and reviewed by the Office of Information and Regulatory Affairs.
The Value of a Statistical Life, Cost-Effectiveness, and the Limits of Numbers
- Explain what the value of a statistical life measures and what it does not, and where the estimates come from.
- Apply cost-effectiveness and cost-utility analysis where monetization is unavailable or objectionable.
- Conduct a distributional analysis alongside an efficiency analysis and state where quantification stops helping.
The big picture
In 2003 the Environmental Protection Agency published an analysis of a clean air rule that applied a lower value to preventing the death of a person over seventy than to preventing the death of a younger person. The reduction was about thirty-seven percent. The technical reasoning was defensible on its own terms: an older person has fewer remaining years, and some willingness-to-pay studies suggested older people would pay less for a given risk reduction.
The public reaction was ferocious. Critics named it the senior death discount, advocacy groups mobilized, and within months the agency's administrator announced it would not be used. Whatever you think of the underlying economics, the episode teaches something important about this lesson's subject. Monetizing mortality risk is unavoidable if you want to compare a rule that costs two billion dollars and prevents two hundred deaths with one that costs four billion and prevents six hundred. It is also an operation the public scrutinizes with unusual intensity, and analysts who cannot explain clearly what they are doing will be understood to be doing something they are not.
So let us be precise about what the value of a statistical life actually is, because almost every popular description of it is wrong.
What the value of a statistical life measures
It is not the value of a person's life. It is not what a court awards for a wrongful death. It is not what anyone would accept to be killed, which for essentially everyone is unbounded.
It is the aggregate amount a population is willing to pay for small reductions in the risk of death, scaled up to one expected death avoided. The arithmetic makes this concrete. Suppose a community of one hundred thousand people each would pay one hundred dollars per year for a safety improvement that reduces each person's annual risk of death by one in one hundred thousand. Across the community, that improvement prevents one expected death per year, and the community collectively paid ten million dollars for it. The value of a statistical life here is ten million dollars.
Notice what is being valued: a small change in probability, aggregated. No identified person's life is priced. Nobody is being asked what they would take to die. The unit exists so that risk reductions spread across large populations can be compared against costs, which is exactly the situation regulation is always in.
Key idea: The value of a statistical life is the sum of many people's willingness to pay for small risk reductions, divided by the expected deaths avoided. It prices probability changes across a population, never an identified individual.
Where the numbers come from
Two research traditions produce the estimates, and both have real weaknesses that a careful analyst names.
Revealed preference studies infer values from actual behavior. The dominant approach uses labor markets: comparing wages across jobs with different fatality rates, controlling for skill, education, industry, and other characteristics, to estimate the compensating wage differential for risk. W. Kip Viscusi has produced the most extensive body of this work. The strength is that people are making real choices with real money. The weaknesses are that workers may not accurately perceive small risks, that the estimates come from working-age people in risky occupations who may not represent the population, and that the controls are doing heavy lifting, since dangerous jobs differ from safe ones in many ways at once.
Stated preference studies ask people directly through survey experiments about hypothetical risk reductions. The strength is that you can ask about exactly the risk and population you care about. The weaknesses are hypothetical bias, since what people say they would pay exceeds what they do pay, and insensitivity to scope, the well-documented finding that respondents often report similar willingness to pay for risk reductions that differ by a factor of ten, which is logically incoherent and suggests they are responding to the idea rather than the magnitude.
Federal agencies converge on broadly similar figures while differing in detail. Recent Department of Transportation guidance has used a value in the range of roughly eleven to thirteen million dollars, updated annually for income and inflation; the Environmental Protection Agency has long used a central estimate derived from a set of studies, similarly adjusted. Agencies publish their methods, which is what makes disagreement productive rather than merely suspicious.
The genuine controversies
Four disputes recur, and none is resolved by better data alone.
Should the value vary by age? The senior discount episode is the American landmark. Arguments for varying: an intervention that adds forty expected years differs from one that adds three, and ignoring that treats unequal things equally. Arguments against: it makes the state assign different values to citizens' lives on the basis of age, which many people regard as a category error regardless of the arithmetic, and the empirical evidence that older people value risk reduction less is contested. Some analysts sidestep by using the value of a statistical life-year, though that has its own problems.
Should it vary by income or country? Willingness to pay rises with income, so revealed-preference methods produce lower values in poorer countries. Applying that consistently would justify weaker safety standards where people are poorer, which strikes many people as a mechanism for exporting hazard. Applying a rich-country value everywhere, however, can imply spending on risk reduction that a poor country would rather spend on other urgent needs, overriding its own priorities. Both positions have serious defenders.
Are all deaths equivalent? Evidence suggests people are willing to pay more to avoid deaths perceived as involuntary, dreaded, or catastrophic than to avoid statistically identical ordinary risks. Whether analysis should honor those preferences, since they are real preferences, or correct them as inconsistent, is contested.
Does using the number corrupt the deliberation? The strongest philosophical critique, associated with Elizabeth Anderson and with Frank Ackerman and Lisa Heinzerling, holds that some goods are wrongly valued by monetary comparison and that pricing them changes how we reason about them. The strongest reply, from Cass Sunstein among others, is that agencies with finite budgets implicitly assign values whether or not they state them, and the historical record of implicit valuation is far more erratic than the explicit one, with regulations that have ranged from well under a million dollars to hundreds of millions per life saved. Both sides here are making serious arguments and this course will not adjudicate between them.
Cost-effectiveness analysis
When monetizing the outcome is impossible or objectionable, you can still compare options that share an outcome. Cost-effectiveness analysis reports cost per unit of result: cost per life saved, per case of illness averted, per ton of carbon dioxide abated, per additional student reaching proficiency.
The key limitation is that it cannot tell you whether to do anything at all. It ranks options within a goal. Comparing eight thousand dollars per case averted against six thousand dollars per ton of carbon abated is meaningless, because the units differ. Cost-benefit analysis can make that comparison precisely because it forces both onto a monetary scale, which is the trade-off between the two techniques in a sentence.
Cost-utility analysis extends this in health by using a common outcome unit. The quality-adjusted life year, or QALY, weights each year of life by a quality factor from zero to one, so that ten years at a quality weight of 0.8 counts as eight QALYs. The disability-adjusted life year, or DALY, used widely in global health, counts years lost to premature death plus years lived with disability.
The incremental cost-effectiveness ratio compares two options: the difference in cost divided by the difference in QALYs. If a new treatment costs thirty thousand dollars more than the standard and yields 1.5 additional QALYs, the ratio is twenty thousand dollars per QALY.
Countries differ sharply in whether they use this to decide. The United Kingdom's National Institute for Health and Care Excellence applies a threshold in the region of twenty to thirty thousand pounds per QALY when advising on National Health Service adoption, with flexibility for end-of-life and highly specialized treatments. The United States does the opposite: federal law restricts the use of QALY-based thresholds in Medicare determinations, reflecting a durable political position that such measures systematically undervalue life extension for people with disabilities and the seriously ill. Disability rights advocates have made that argument forcefully and it is a substantive objection, not merely a political one, since quality weights derived from general-population surveys may not reflect how people actually living with a condition rate their own lives. The counterargument is that refusing an explicit standard does not avoid rationing; it moves rationing to price, insurance design, and administrative friction, where it is less visible and arguably less fair. Reasonable people land on both sides.
Key idea: Cost-effectiveness compares options within a single goal and cannot say whether the goal deserves resources at all. Cost-utility extends the comparison across health interventions, which is why it is both more powerful and more contested.
Distributional analysis
Because cost-benefit analysis is silent on distribution by construction, a complete analysis reports distribution separately. Three questions structure it.
Who bears the cost in the end? Statutory incidence is who legally pays; economic incidence is who ends up worse off after prices and wages adjust. A tax on a firm may be borne by consumers through prices, workers through wages, or owners through returns, in proportions that depend on elasticities. Analysts who report statutory incidence and call it distribution have not done the analysis.
Who receives the benefit, and how is it measured? Benefits measured by willingness to pay automatically skew toward the wealthy, because willingness to pay is bounded by ability to pay. A park improvement valued at one hundred dollars by a rich household and thirty by a poor one may matter more to the second household, and the metric cannot see that.
Across which dimensions? Income is the default, but place, race, age, and generation often matter as much or more, and effects can be progressive on one dimension and regressive on another.
Work an example. A carbon tax raises the price of gasoline, heating fuel, and electricity. Lower-income households spend a larger share of income on energy, so the direct burden is regressive. But that is only the first half. If the revenue is returned as an equal per-household dividend, the combined effect is typically progressive, because everyone receives the same rebate while wealthier households paid more in absolute terms. If instead the revenue funds a cut in corporate income tax rates, the combined effect is generally regressive. The lesson generalizes: for any revenue-raising instrument, the incidence of the tax and the incidence of the revenue use must be analyzed together, and reporting only the first is a standard and misleading practice on all sides of these debates.
Some frameworks apply distributional weights, formally counting a dollar to a poor household as worth more than a dollar to a rich one. This follows directly from diminishing marginal utility of income and is technically simple. It is also politically explosive, because the weights must be chosen and the choice is nakedly a value judgment. The 2023 revision to OMB Circular A-4 permitted agencies to present weighted results as a supplement, which was among its most debated features.
Where quantification stops helping
Four honest limits deserve statement in any analysis.
The streetlight effect. Analysts measure what can be measured. Program effects that are easy to observe get counted and effects that are diffuse or slow do not, so the measured effect is a biased subset of the real one, usually in a known direction.
Goodhart's law. When a measure becomes a target, it ceases to be a good measure. Once schools are judged on a test, effort shifts toward the test; once hospitals are judged on a wait-time metric, the metric improves in ways that may or may not track care. This is not cheating in most cases; it is rational response to incentives, and it is predictable enough to design against.
Aggregation hides structure. A net benefit of one hundred million dollars is compatible with severe concentrated harm. The aggregate is a real fact and an incomplete one.
Some values resist the operation. Procedural fairness, dignity, community continuity, and the sense of being treated as a citizen rather than a case do not reduce to willingness to pay without loss. The professional response is not to invent numbers for them and not to pretend they do not exist, but to name them explicitly in the analysis so the decision maker weighs them consciously rather than by default.
Common misconceptions
- The value of a statistical life prices a person. It aggregates many people's willingness to pay for small probability reductions and applies to no identified individual.
- Agencies invent the figure. Estimates come from wage-risk and survey studies with published methods, which is what allows the disagreements to be specific.
- Cost-effectiveness avoids valuation. It avoids monetizing the outcome but still requires someone to decide what ratio is acceptable, which is a valuation made elsewhere.
- QALYs are objective. Quality weights come from surveys whose respondents often do not live with the conditions being weighted, which is the core of the disability rights objection.
- A regressive tax is regressive overall. Incidence must be assessed jointly with the use of revenue; the same carbon tax can be progressive or regressive depending on where the money goes.
Recap
- The value of a statistical life aggregates willingness to pay for small risk reductions and is derived from wage-risk and stated-preference studies, each with named weaknesses.
- Whether the value should vary by age, income, or country, and whether monetizing mortality corrupts deliberation, are genuine and unresolved disputes.
- Cost-effectiveness ranks options within a goal; cost-utility uses QALYs or DALYs to compare across health interventions, which the United Kingdom does explicitly and United States law restricts.
- Distributional analysis must distinguish statutory from economic incidence and must analyze revenue use jointly with revenue raising.
- The streetlight effect, Goodhart's law, aggregation, and non-monetizable values are limits to state openly rather than to solve by inventing numbers.
Sources
- U.S. Department of Transportation. (n.d.). Departmental guidance on valuation of a statistical life in economic analysis. transportation.gov
- U.S. Environmental Protection Agency. (n.d.). Guidelines for preparing economic analyses. epa.gov
- Office of Management and Budget. (2003). Circular A-4: Regulatory analysis. The White House. whitehouse.gov archive
- Wikipedia contributors. (n.d.). Value of life. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Quality-adjusted life year. Wikipedia. en.wikipedia.org
- Key terms
- Value of a statistical life
- Aggregate willingness to pay across a population for small reductions in mortality risk, scaled to one expected death avoided.
- Revealed preference
- Inferring values from actual choices, most often the wage premium workers receive for accepting higher occupational fatality risk.
- Stated preference
- Estimating values from survey responses about hypothetical risk changes; flexible but prone to hypothetical bias and scope insensitivity.
- Cost-effectiveness analysis
- Comparing options by cost per unit of a shared outcome; it ranks within a goal but cannot say whether the goal merits resources.
- Quality-adjusted life year
- A year of life weighted by a quality factor between zero and one, used to compare health interventions on a common scale.
- Incremental cost-effectiveness ratio
- The difference in cost between two options divided by the difference in health outcome, such as dollars per additional QALY.
- Economic incidence
- Who actually bears a cost after prices and wages adjust, as opposed to statutory incidence, which is who legally pays.
- Distributional weights
- Explicit multipliers counting a dollar to a lower-income household as worth more, following from diminishing marginal utility of income.
- Goodhart's law
- The principle that a measure ceases to be a good measure once it becomes a target.
Module 5: Evidence, Evaluation, and Implementation
How to know whether a policy worked. The counterfactual problem stated plainly, randomized trials read through three real studies with their real limitations, the quasi-experimental designs that carry most of the policy literature, and the implementation research explaining why sound policies fail in delivery.
The Counterfactual Problem and What Randomized Trials Can Show
- State the fundamental problem of causal inference and identify selection bias in a real evaluation design.
- Explain how randomization solves the comparability problem and what it still cannot solve.
- Read the Oregon, Moving to Opportunity, and Perry Preschool studies accurately, including their genuine limitations.
The big picture
A job training program reports that seventy percent of its graduates found work within six months. Is the program effective?
You cannot tell, and the reason is not that the number is too small or the sample too narrow. The reason is that you have no idea what those same people would have done without the program. Perhaps sixty-five percent would have found work anyway. Perhaps only twenty percent would have. The seventy percent figure is compatible with a program that is transformative, useless, or actively harmful.
Worse, the people who enroll in a voluntary training program are systematically different from those who do not. They sought it out, which suggests motivation, information, and time. If motivated people find jobs faster with or without training, then comparing graduates to non-participants will make the program look effective even if it does nothing. This is selection bias, and it is the single most common way policy evaluation goes wrong. Notice that it can also run the other direction: if a program targets people with the worst job prospects, naive comparison will make an effective program look harmful.
This lesson is about the machinery that gets you from correlation to a defensible causal claim, and about how to read that machinery honestly when the results are inconvenient.
The fundamental problem of causal inference
Paul Holland gave the problem its standard name. The causal effect of a treatment on a unit is the difference between the outcome with treatment and the outcome without it, for that same unit, at that same time. That difference is never observable, because a person either received the program or did not. One of the two outcomes is always missing.
Every evaluation method is therefore a strategy for constructing a credible estimate of the missing outcome, the counterfactual. There is no method that avoids this. When you read a policy study, the first question is always: what is standing in for the counterfactual here, and why should I believe it resembles what would have happened?
Two tempting substitutes fail routinely.
Before and after comparison uses the same people at an earlier time. It fails whenever anything else changed, which is always. If a city adds streetlights and crime falls, crime may have been falling nationally. It also fails through regression to the mean: programs often start after an unusually bad period, and unusually bad periods tend to be followed by better ones regardless of intervention. A city that installs cameras after a spike in crime will usually see crime fall, because spikes subside.
Comparison with non-participants uses people who did not take the program. It fails through selection bias whenever participation is not random, which is nearly always.
Key idea: Every causal claim rests on a counterfactual that cannot be observed. Methods differ only in how credibly they construct a stand-in, so the reader's job is to interrogate the stand-in rather than the headline number.
What randomization does
If you assign people to treatment and control by lottery, the two groups are, in expectation, alike in everything: motivation, health, prior earnings, family support, and every characteristic you never thought to measure. That last clause is what makes randomization powerful. Statistical adjustment can only control for variables you observed. Randomization balances the unobserved ones too.
Some practical points that matter when reading trials.
Balance is probabilistic, not guaranteed. Any single randomization can produce unlucky imbalance, which is why studies publish baseline comparison tables and why small samples are more fragile.
Not everyone assigned takes the treatment. The effect of being offered a program, called the intent-to-treat effect, differs from the effect of actually receiving it. Intent-to-treat is usually the policy-relevant quantity, since real programs can only offer.
Randomization does not fix external validity. A trial tells you what happened to these people, in this place, under this implementation, at this time. Whether it generalizes is a separate argument requiring theory and replication.
Ethics constrain design. Randomization is defensible when there is genuine uncertainty about which arm is better, a condition called equipoise, and it is especially defensible when a program is oversubscribed, since a lottery is arguably fairer than first-come-first-served. Several of the studies below exist precisely because the program could not serve everyone.
The Oregon Health Insurance Experiment
In 2008 Oregon had funds to expand Medicaid to a limited number of low-income adults and far more applicants than slots. The state ran a lottery. Roughly ninety thousand people signed up for about ten thousand openings. This is one of the only randomized studies of health insurance coverage in a developed country, and it happened because a state needed a fair rationing device.
The findings, published by Amy Finkelstein, Katherine Baicker, and colleagues, are worth stating precisely because they are widely misdescribed by people on both sides.
Coverage increased use of care substantially: more office visits, more prescriptions, more preventive screening, more emergency department use rather than less. It sharply reduced financial strain, cutting catastrophic out-of-pocket expenditures and unpaid medical bills. It produced a large reduction in the rate of depression, roughly a third. It increased the diagnosis and treatment of diabetes. And at two years, it did not produce statistically significant improvements in measured blood pressure, cholesterol, or glycated hemoglobin.
How should an honest analyst read that last finding? Three things must be said together. First, it genuinely failed to confirm the large near-term improvements in physical health markers that some advocates had predicted, and that is a real result, not a technicality. Second, the confidence intervals were wide enough to include clinically meaningful improvements, so the study cannot establish that the effect is zero; absence of a significant estimate is not evidence of absence. Third, two years is a short window for chronic disease markers, and the mortality and long-run effects the debate ultimately cares about were beyond the study's reach.
What the study establishes solidly is that coverage improves financial security and mental health substantially. What it leaves open is the magnitude of physical health effects. Both supporters and critics of Medicaid expansion cite Oregon, and each is quoting a real part of the result while omitting another. That is the normal condition of policy evidence.
Moving to Opportunity
Between 1994 and 1998 the Department of Housing and Urban Development randomized roughly forty-six hundred families in five cities into three groups: a voucher usable only in a low-poverty neighborhood plus counseling, an unrestricted voucher, and a control group receiving neither.
The interim results published around 2008 were sobering for advocates. Adults who moved showed substantial improvements in mental health and, among women, reductions in obesity and diabetes. But there were no detectable effects on adult employment or earnings, and no overall effects on children's test scores. Many concluded that neighborhoods mattered less than believed.
Then in 2016 Raj Chetty, Nathaniel Hendren, and Lawrence Katz linked the participants to federal tax records and followed the children into adulthood. The pattern was strikingly age-dependent. Children who moved to a low-poverty neighborhood before roughly age thirteen had substantially higher earnings as adults, on the order of thirty percent, along with higher college attendance rates and lower rates of single parenthood. Children who moved as teenagers showed no gain and possibly slight harm, consistent with disruption costs outweighing shorter exposure.
The methodological lesson is as important as the substantive one. The same experiment produced a null result and a large positive result depending on the outcome measured and the horizon observed. If evaluation had stopped in 2008, the accepted conclusion would have been the opposite of the current one. Whenever you read that a program does not work, ask what was measured, on whom, and for how long.
Key idea: Null results are conditional on outcome, subgroup, and time horizon. Moving to Opportunity looked like a failure at ten years and a substantial success at twenty for the children young enough to benefit.
Perry Preschool, stated honestly
The Perry Preschool Project is probably the most cited social program evaluation in existence, and it is routinely oversold. Here is the accurate version.
Between 1962 and 1967 in Ypsilanti, Michigan, one hundred twenty-three low-income African American children assessed as being at high risk of school failure were assigned, largely at random, to either a high-quality preschool program with weekly home visits or to no program. They have been followed for decades.
The results are genuinely striking. Participants had higher rates of high school graduation, higher adult earnings, lower rates of arrest, and lower reliance on public assistance. James Heckman and colleagues estimated an annual rate of return in the range of seven to ten percent, which if accurate makes it an extraordinary public investment.
Now the limitations, all of which are acknowledged in the serious literature.
The sample is one hundred twenty-three children. That is very small for detecting anything, and it means estimates carry wide uncertainty and are sensitive to a handful of individual life trajectories. Heckman's own reanalysis addressed this with methods designed for small samples, which is itself an acknowledgment of the problem.
Randomization was imperfect. Some children were reassigned between groups for practical reasons, including maternal employment, which introduces a possible violation of the clean random assignment the design promised.
The counterfactual has changed. In 1962 the control group received essentially no preschool. Today most American four-year-olds have access to some form of preschool, so the relevant comparison is no longer program versus nothing but program versus whatever else exists.
Program intensity was very high. Perry used trained teachers with bachelor's degrees and master's-level training, low child-to-staff ratios, and weekly ninety-minute home visits. Scaled public programs rarely match this.
And the broader early childhood literature is genuinely mixed. The Abecedarian Project found comparable long-run benefits. But the randomized evaluation of Tennessee's Voluntary Pre-K program found initial achievement gains that faded and, by later elementary grades, reversed on some measures, a result the researchers themselves found unwelcome and published anyway. The Head Start Impact Study found initial gains that largely converged with the control group by third grade, while other work using different methods finds substantial long-run benefits from Head Start.
The defensible summary is that high-quality, intensive early childhood programs have shown large long-run benefits in small studies, that scaled programs have produced far more variable results, and that researchers disagree about how much of the gap is quality, counterfactual, or measurement. Anyone who tells you the early childhood evidence is settled in either direction is not describing the literature.
Common misconceptions
- A high success rate shows a program works. Without a counterfactual, an outcome rate says nothing about causal effect.
- Before-and-after comparison is a reasonable substitute. Secular trends and regression to the mean routinely generate apparent effects where none exist.
- Randomization guarantees balanced groups. It balances in expectation; any single draw can be unlucky, which is why baseline tables exist and small samples are fragile.
- A non-significant result proves no effect. Wide confidence intervals mean the study could not detect an effect, which is different from establishing that there is none.
- Perry Preschool proves universal pre-kindergarten works. It was 123 children in a uniquely intensive 1960s program against a no-preschool counterfactual, and the scaled evidence is genuinely mixed.
Recap
- The fundamental problem of causal inference is that a unit's untreated outcome is never observed, so every method is a strategy for constructing a credible counterfactual.
- Selection bias can make ineffective programs look effective or effective programs look harmful, depending on who selects in.
- Randomization balances unobserved as well as observed characteristics but cannot deliver external validity, and intent-to-treat is usually the policy-relevant estimate.
- Oregon showed large financial and mental health benefits from coverage and no significant two-year change in three physical markers, with intervals too wide to establish a zero.
- Moving to Opportunity reversed its own apparent verdict once children were followed into adulthood, and Perry Preschool is powerful, tiny, imperfectly randomized, and set against a counterfactual that no longer exists.
Sources
- Wikipedia contributors. (n.d.). Oregon Medicaid health experiment. Wikipedia. en.wikipedia.org
- U.S. Department of Housing and Urban Development, Office of Policy Development and Research. (n.d.). Moving to Opportunity for Fair Housing. huduser.gov
- Wikipedia contributors. (n.d.). Moving to Opportunity. Wikipedia. en.wikipedia.org
- HighScope Educational Research Foundation. (n.d.). Perry Preschool Project. highscope.org
- U.S. Department of Health and Human Services, Administration for Children and Families. (n.d.). Head Start research and evaluation. acf.hhs.gov
- Key terms
- Counterfactual
- What would have happened to the same units absent the policy; never observed, and therefore estimated by every evaluation design.
- Fundamental problem of causal inference
- Holland's observation that a unit's treated and untreated outcomes can never both be observed.
- Selection bias
- Distortion arising when participants differ systematically from non-participants in ways related to the outcome.
- Regression to the mean
- The tendency for extreme observations to be followed by less extreme ones, which can masquerade as a program effect.
- Intent-to-treat
- The effect of being offered a program regardless of whether the offer was taken up; usually the policy-relevant quantity.
- External validity
- Whether a finding generalizes beyond the specific people, place, implementation, and period studied.
- Equipoise
- Genuine uncertainty about which arm of a trial is better, which is the standard ethical justification for randomizing.
- Statistical power
- A study's ability to detect an effect of a given size; low power produces wide intervals and non-significant results that cannot rule out real effects.
Quasi-Experiments: Difference-in-Differences, Regression Discontinuity, and Reading a Literature
- Explain difference-in-differences and its parallel trends assumption using a real, contested policy study.
- Explain regression discontinuity, identify the running variable and cutoff, and state what population the estimate applies to.
- Assess a body of evidence rather than a single study, accounting for publication bias and genuine expert disagreement.
The big picture
Most policy questions cannot be randomized. You cannot randomly assign states to have a minimum wage, randomly assign people to turn sixty-five, or randomly assign a country to adopt a carbon price. Yet decisions must be made about all of these.
Quasi-experimental methods exploit situations where something other than a researcher's lottery creates variation that is plausibly as good as random. A law changes in one state and not its neighbor. A program cuts off eligibility at an income threshold, so a household one dollar above and one dollar below are nearly identical but treated differently. An election is decided by twenty votes.
These designs carry most of the empirical policy literature you will ever read. Learning to evaluate them is the difference between being persuaded by whichever study you saw last and being able to say what a body of evidence actually supports.
Difference-in-differences
The design compares the change over time in a treated group with the change over time in an untreated comparison group. Subtracting the second change from the first removes anything that affected both groups equally, which is why it is called a difference in differences.
The canonical policy application is David Card and Alan Krueger's 1994 study of the New Jersey minimum wage. New Jersey raised its minimum wage in April 1992; neighboring eastern Pennsylvania did not. The researchers surveyed fast food restaurants in both states before and after, and compared the change in employment.
Standard competitive theory predicted employment would fall in New Jersey relative to Pennsylvania. Card and Krueger found no such decline, and their point estimate was slightly positive. The finding was one of the most consequential in modern applied economics, and it prompted an unusually sharp response. David Neumark and William Wascher re-examined the question using payroll records rather than telephone survey data and reported employment declines. Card and Krueger responded using another data source. The exchange sharpened everyone's attention to data quality and continues to be taught as a model of how empirical disputes should proceed.
The critical assumption is parallel trends: absent the policy, the two groups would have moved in parallel. This is untestable, since it concerns a counterfactual, but it can be probed by checking whether the groups moved in parallel before the treatment. A study that does not show pre-trends is asking you to accept its central assumption on faith.
Other threats include compositional change, if the groups' membership shifts differently over time; spillovers, if the comparison group is affected by the treatment, as when workers commute across a state line; and anticipation, if behavior changes before implementation because the law was announced in advance. A newer methodological literature has also shown that when different units adopt a policy at different times, common estimation approaches can produce badly biased results, and better estimators are now standard.
Key idea: Difference-in-differences removes anything common to both groups but rests entirely on parallel trends. Always look for the pre-period plot, and treat its absence as a warning.
Where the minimum wage literature actually stands
This is an ideal case for practicing intellectual honesty, because the disagreement is real and runs through competent researchers rather than between honest and dishonest ones.
Points of broad agreement: modest increases in the minimum wage raise earnings for most affected workers who keep their jobs; effects on poverty are smaller than either side's rhetoric suggests, because many minimum wage workers are not in poor households and many poor households have no earner; and very large increases relative to local median wages carry greater employment risk than modest ones.
Points of genuine dispute: the size of employment effects at typical increases, whether adjustment occurs through hours and scheduling rather than headcount, whether effects appear only after several years, and how much the answer depends on the choice of comparison group. Researchers associated with Arindrajit Dube and colleagues, using neighboring-county comparisons, generally find small employment effects. Researchers associated with Neumark and Wascher, and separately work by Jeffrey Clemens and others, find larger disemployment effects and argue that neighboring-county designs discard useful variation. The Congressional Budget Office, in assessing proposals to raise the federal minimum wage, has published central estimates of job losses accompanied by wide uncertainty ranges that explicitly include outcomes near zero, which is an honest representation of the state of the field.
Layered on top is a values disagreement that no study can resolve. Suppose the truth is that a given increase raises pay substantially for a large majority of low-wage workers and eliminates jobs for a small minority. Whether that is a good trade depends on how you weigh gains to many against concentrated losses to a few, which is Lesson 5's territory, not an empirical question at all.
Regression discontinuity
Many programs assign treatment by a threshold. Scholarships go to students above a test score. Benefits phase out above an income level. Class-size rules trigger at an enrollment count. Medicare eligibility begins at sixty-five.
The insight is that units just above and just below the cutoff are nearly identical in everything except treatment status. A student scoring 1199 and one scoring 1201 do not differ meaningfully in ability, motivation, or family background, but one gets the scholarship. Comparing outcomes for people narrowly on either side of the line approximates a randomized experiment in that neighborhood.
The vocabulary: the running variable is what determines assignment (the test score); the cutoff is the threshold; a sharp design means treatment changes deterministically at the cutoff, while a fuzzy design means the probability of treatment jumps but not to certainty.
Real applications are everywhere. Joshua Angrist and Victor Lavy used an Israeli rule capping classes at forty students, which mechanically creates a sharp drop in class size when enrollment crosses a multiple of forty, to estimate class-size effects. David Card, Carlos Dobkin, and Nicole Maestas used the age-sixty-five Medicare threshold to estimate the effect of coverage on care use and health. David Lee used narrowly decided elections to estimate incumbency advantage.
Two limitations deserve emphasis. First, the estimate is local. It applies to units near the cutoff and need not generalize to those far from it. The effect of a scholarship on marginal students says little about its effect on the strongest applicants. Second, the design collapses if people can manipulate the running variable. If schools can nudge scores across a threshold, or families can report income just under a limit, then units on either side differ in exactly the way the design assumes they do not. The standard check, associated with Justin McCrary, is to plot the density of the running variable and look for bunching just on the favorable side of the cutoff.
Key idea: Regression discontinuity buys strong internal validity near a threshold at the cost of a local estimate, and its credibility depends on people being unable to place themselves on their preferred side of the line.
Two more designs worth recognizing
Instrumental variables use a factor that affects treatment but has no other path to the outcome. Angrist's use of the Vietnam draft lottery to estimate the effect of military service on later earnings is the classic case: the lottery number affected the probability of serving and nothing else about a person. The design's weakness is that the exclusion restriction, meaning the instrument affects the outcome only through the treatment, is an assumption rather than a finding, and it is often the weakest link.
Synthetic control, developed by Alberto Abadie and colleagues, builds a weighted combination of untreated units that tracks the treated unit's pre-treatment path, then uses that composite as the counterfactual. Their study of California's 1988 tobacco tax constructed a synthetic California from other states and attributed the post-1988 divergence in cigarette consumption to the policy. It is especially useful when a single large unit adopts a policy and no single comparison is convincing.
Reading a literature rather than a study
Any single study can be wrong. The competent move is to assess a body of work, and doing that requires knowing how bodies of work get distorted.
Publication bias. Journals and press releases favor statistically significant, surprising results. Null findings disproportionately stay in file drawers, so the published literature overstates average effects. Funnel plots and related tools can detect the asymmetry this produces.
Specification searching. With many defensible modeling choices, a researcher can arrive at significance without any conscious dishonesty. Preregistration, pre-analysis plans, and multiverse analyses reporting results across many specifications are the field's responses.
Replication. Social science has been through a decade of reckoning on this, with well-known findings failing to replicate. The healthy consequence is that a single striking study should now be treated as a hypothesis rather than a fact.
Effect sizes shrink at scale. Programs evaluated by their developers under ideal conditions typically perform worse when scaled, a pattern John List calls the voltage effect. Efficacy under ideal conditions and effectiveness under real ones are different quantities, and confusing them is a standard route to disappointment.
Practical habits: prefer systematic reviews and meta-analyses over single studies; check whether the review preregistered its inclusion criteria; look for whether independent teams with different priors reach similar conclusions; and give special weight to results that a research team clearly did not want, such as the Tennessee pre-kindergarten team publishing fade-out and reversal.
Common misconceptions
- Quasi-experiments are just correlational studies with better names. Well-designed ones exploit variation that is plausibly as good as random and state their identifying assumptions explicitly.
- Parallel trends can be verified. It is an assumption about a counterfactual; pre-trend evidence makes it plausible but never proves it.
- A regression discontinuity estimate applies to everyone. It is local to the cutoff and may not describe units far from it.
- Expert disagreement means the evidence is worthless. On the minimum wage, competent researchers agree on much and disagree about specific magnitudes, which is a normal and informative state.
- A single well-publicized study settles a question. Publication bias, specification searching, and scale effects all mean a striking result should be treated as a hypothesis until replicated.
Recap
- Difference-in-differences subtracts the comparison group's change from the treated group's, and stands or falls on parallel trends, which pre-period evidence can support but not prove.
- Card and Krueger's New Jersey study and the Neumark and Wascher response illustrate both the method and how a productive empirical dispute proceeds.
- The minimum wage literature has real agreement on some points, real disagreement on magnitudes, and an irreducible values question about trading broad gains against concentrated losses.
- Regression discontinuity exploits thresholds to obtain strong local estimates, and its main threats are limited generalizability and manipulation of the running variable.
- Judging a literature requires accounting for publication bias, specification searching, replication failures, and the shrinkage of effects at scale.
Sources
- Congressional Budget Office. (n.d.). Reports on the minimum wage. cbo.gov
- Wikipedia contributors. (n.d.). Difference in differences. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Regression discontinuity design. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Synthetic control method. Wikipedia. en.wikipedia.org
- U.S. Bureau of Labor Statistics. (n.d.). Characteristics of minimum wage workers. bls.gov
- Key terms
- Difference-in-differences
- A design comparing the treated group's change over time with a comparison group's change, removing influences common to both.
- Parallel trends
- The untestable assumption that treated and comparison groups would have moved together absent the policy; supported by pre-period evidence.
- Regression discontinuity
- A design exploiting a threshold rule, comparing units narrowly above and below a cutoff that determines treatment.
- Running variable
- The measure that determines assignment in a discontinuity design, such as a test score or income level.
- Manipulation
- Units placing themselves on the favorable side of a cutoff, which invalidates a discontinuity design and is checked by density tests.
- Instrumental variable
- A factor affecting treatment with no other path to the outcome, such as a draft lottery number affecting military service.
- Synthetic control
- A weighted combination of untreated units constructed to match a treated unit's pre-treatment path and serve as its counterfactual.
- Publication bias
- Systematic overrepresentation of significant and surprising findings in published literature, inflating apparent average effects.
- Voltage effect
- List's term for the tendency of program effects to shrink when moved from ideal evaluation conditions to full-scale delivery.
Implementation: Why Sound Policies Fail in Delivery
- Explain why clearance points and joint probabilities make multi-actor implementation fail more often than intuition suggests.
- Analyze street-level discretion and administrative burden as sources of the gap between policy on paper and policy as experienced.
- Apply implementation thinking at the design stage, using take-up rates and delivery constraints as analytic inputs.
The big picture
In 1965 the federal Economic Development Administration announced a program to create jobs for unemployed minority workers in Oakland, California. It was well funded, uncontroversial, and enjoyed enthusiastic support from every level of government involved. Roughly twenty-three million dollars was committed with great fanfare, and the projects included an airport hangar and a marine terminal.
Several years later almost nothing had been built and very few jobs had been created. Jeffrey Pressman and Aaron Wildavsky went to Oakland to find out why, and the book they produced in 1973 essentially founded the study of implementation. Its subtitle is the best summary of the field ever written: how great expectations in Washington are dashed in Oakland, or why it is amazing that federal programs work at all.
Here is what they found, and it is not what anyone expected. There was no villain. Nobody sabotaged the program. There was no scandal, no captured agency, no ideological opposition. What there was, instead, was a long chain of separate decisions, each requiring agreement from a different set of actors, each individually reasonable, each taking time.
The arithmetic of clearance points
Pressman and Wildavsky counted what they called clearance points: separate decisions requiring the assent of an actor whose agreement was not guaranteed. In the Oakland program they identified around seventy. Environmental review, engineering approvals, leasing negotiations, minority hiring plans, funding disbursements, and agency sign-offs each constituted a link.
Now do the arithmetic, because it is genuinely startling. Suppose each clearance point has a ninety-five percent chance of being resolved favorably and promptly, which is an optimistic assumption for any government process. With seventy independent clearance points, the probability that all of them clear is 0.95 raised to the seventieth power, which is roughly three percent. Even at a ninety-nine percent success rate per point, the joint probability is about half.
This is the central insight of implementation studies. Failure does not require anyone to fail. It requires only that many people succeed independently, which is unlikely at scale. Delay compounds the problem, because the longer implementation takes, the more likely it is that officials rotate out, priorities shift, budgets change, and the original agreement dissolves.
Key idea: Multi-actor implementation fails through the multiplication of independent probabilities, not through opposition. A design requiring many separate agreements is fragile even when every participant is competent and willing.
The practical design implication is direct: reduce the number of clearance points, or accept a long timeline and build political durability to survive it. Analysts who compare two options should count veto points in delivery just as Lesson 2 counted them in enactment.
Top-down and bottom-up perspectives
Two schools organized the field, and both have a piece of the truth.
The top-down perspective, developed by Paul Sabatier and Daniel Mazmanian among others, starts from the statute and asks what conditions produce faithful execution. Their conditions include clear and consistent objectives, a valid causal theory linking the intervention to the outcome, adequate legal and financial resources, a committed implementing agency, supportive interest groups and legislators over time, and stable socioeconomic conditions. It is a useful checklist, and its diagnosis of many failures is simply that the statute contained contradictory objectives because contradiction was the price of passage.
The bottom-up perspective, associated with Michael Lipsky and Benny Hjern, starts at the point of delivery and works upward. Its claim is that what a policy is, in practice, is what the caseworker, teacher, nurse, inspector, or officer does with it at the moment of contact.
Lipsky's concept of street-level bureaucrats is the most useful single idea here. These are public employees who interact directly with citizens and who exercise substantial discretion under conditions of chronic resource scarcity. They have more cases than time. They face ambiguous goals, since a school is asked to educate, socialize, sort, and supervise simultaneously. They cannot satisfy every demand, so they develop coping routines, and those routines become the operative policy.
Lipsky identified the recurring patterns. Rationing introduces friction, such as long waits and complex forms, that reduces demand without any formal denial. Creaming selects clients most likely to succeed, which makes performance metrics look good and pushes the hardest cases out. Routinizing reduces varied situations to a small number of categories so decisions can be made quickly, at the cost of fit.
None of this is corruption. It is what a competent person does with an impossible workload. Which means it cannot be fixed by exhortation, only by changing caseloads, simplifying the task, or changing what the front line is asked to do.
Administrative burden and take-up
Pamela Herd and Donald Moynihan gave the field its most policy-relevant recent concept. Every interaction with government imposes costs on the citizen, and they sort into three kinds.
| Cost | What it is | Example |
|---|---|---|
| Learning costs | Discovering that a program exists and whether you qualify | Not knowing you are eligible for a tax credit |
| Compliance costs | Time, documents, fees, appointments, and repeated recertification | Producing pay stubs, proof of residence, and identity documents in person |
| Psychological costs | Stigma, stress, and loss of autonomy in the encounter | Feeling judged applying for food assistance |
These costs determine take-up, the share of eligible people who actually receive a benefit, and take-up rates are far below one hundred percent for most American programs. The Earned Income Tax Credit reaches roughly four in five eligible households, which is comparatively high because it runs through tax filing. SNAP reaches a large majority of eligible people overall, but participation among eligible older adults has historically been much lower, often around half, driven by learning and psychological costs rather than by any rule.
The analytic consequence is sharp. A program's effect on the population equals its effect on participants multiplied by take-up. A policy with an excellent per-participant effect and forty percent take-up may accomplish less than a mediocre one at ninety percent. Analysts who model only the per-participant effect systematically overstate what a program will deliver.
An important observation from Herd and Moynihan: burden is often not accidental. It is sometimes a deliberate policy instrument, chosen precisely because it reduces enrollment and spending without requiring anyone to vote to narrow eligibility. That is a genuine policy choice with defenders, who argue that some friction is necessary to verify eligibility and deter fraud, and critics, who argue that documented improper payment rates in most programs are far smaller than the eligible non-participation created by the friction. Both sides are pointing at real numbers.
Key idea: A program's population-level effect is its per-participant effect times take-up, and take-up is driven by learning, compliance, and psychological costs that policy design controls.
When delivery meets technology and scale
Two modern failure modes deserve naming because they recur.
The healthcare.gov launch in October 2013 is the standard American case study in delivery failure. The policy had survived enactment and Supreme Court review, but the exchange website could not handle load, integration across federal systems failed, and the administration's signature domestic achievement nearly collapsed at the point of contact with users. Subsequent analysis pointed to fragmented contracting with no single accountable integrator, requirements that continued changing late, procurement rules that favored incumbent contractors over technical capability, and a testing schedule compressed to nothing. The eventual rescue, and the creation of digital service units in government afterward, treated delivery capacity as a policy variable rather than an afterthought.
Unemployment insurance in early 2020 is the other. State systems built on decades-old software, staffed for normal claim volumes, met an unprecedented surge along with brand-new federal programs to administer. The result was months of delay for many claimants in some states and, simultaneously, substantial improper payments. Both failures had the same root: administrative capacity that had been treated as a cost to be minimized during ordinary times.
The general lesson is that state capacity, meaning the staffing, systems, data, and expertise to actually execute, is itself a policy choice made incrementally over years, usually invisibly, and it constrains everything.
Designing for implementation
Implementation research is not merely a source of pessimism. It generates specific, testable design moves.
Reduce clearance points. Fewer required agreements means a higher joint probability of completion.
Use existing delivery channels. Programs that run through an established system, such as the tax code or an existing benefit, achieve higher take-up than programs requiring a new application.
Default and auto-enroll where legally possible. Automatic enrollment in retirement savings raised participation dramatically compared with opt-in, one of the more robust findings in behavioral policy. Data matching that enrolls people already verified through another program achieves the same thing in benefits administration.
Simplify the encounter. Shortening forms, prefilling known information, and extending recertification intervals measurably raise take-up.
Fund the front line. Caseload ratios are a policy variable, and Lipsky's coping behaviors respond to them.
Pilot, measure, and expect voltage drop. Assume effects will shrink at scale, and build measurement in from the start rather than commissioning evaluation after the fact.
One honest caution. The behavioral design literature has itself been through a correction. Some widely publicized nudge effects proved smaller on replication than the original reports suggested, and a meta-analytic debate has run over whether the published effect sizes were inflated by publication bias, exactly as Lesson 12 would predict. Automatic enrollment and simplification have held up well. Lighter-touch interventions such as single mailings or reminder messages generally produce small effects. That distinction matters: simplification changes the structure of the encounter, while a reminder only changes attention.
Common misconceptions
- Implementation failure indicates opposition or incompetence. Pressman and Wildavsky found neither in Oakland; the failure came from the joint probability of many independent agreements.
- Policy is what the statute says. What citizens experience is what front-line staff do under real caseloads, which is why bottom-up analysis is necessary.
- Eligible people receive benefits. Take-up is often far below full, and the gap is produced by learning, compliance, and psychological costs that design controls.
- Administrative burden is an accident of bureaucracy. It is frequently a deliberate instrument for limiting enrollment without changing eligibility rules, which is a defensible choice that should be stated openly and analyzed.
- Behavioral nudges reliably produce large effects. Structural changes such as auto-enrollment and simplification hold up; light-touch reminders generally produce small effects, and early estimates were inflated by publication bias.
Recap
- Multi-actor implementation fails through multiplied probabilities across clearance points, so fragility rises with the number of required agreements even among willing participants.
- Top-down analysis supplies conditions for faithful execution; bottom-up analysis shows that operative policy is what street-level staff do under scarcity.
- Lipsky's coping routines of rationing, creaming, and routinizing are rational responses to impossible workloads and respond to caseloads rather than exhortation.
- Administrative burden determines take-up, and population effect equals per-participant effect times take-up, so ignoring take-up overstates program impact.
- State capacity is a policy choice built over years, and design moves such as auto-enrollment, existing channels, and simplification improve delivery more reliably than light-touch nudges.
Sources
- U.S. Government Accountability Office. (n.d.). Reports and testimonies. gao.gov
- U.S. Department of Agriculture, Food and Nutrition Service. (n.d.). Supplemental Nutrition Assistance Program (SNAP). fns.usda.gov
- Internal Revenue Service. (n.d.). Earned Income Tax Credit (EITC). irs.gov
- Wikipedia contributors. (n.d.). Street-level bureaucracy. Wikipedia. en.wikipedia.org
- Congressional Research Service. (n.d.). Reports on program administration and oversight. crsreports.congress.gov
- Key terms
- Clearance point
- A separate decision requiring the assent of an actor whose agreement is not guaranteed; many such points multiply into a low joint probability of success.
- Top-down implementation
- Analysis beginning from the statute, asking what conditions produce faithful execution of legislative intent.
- Bottom-up implementation
- Analysis beginning at the point of delivery, treating operative policy as what front-line staff actually do.
- Street-level bureaucrat
- A public employee who interacts directly with citizens and exercises substantial discretion under chronic resource scarcity.
- Creaming
- Selecting clients most likely to succeed, which improves measured performance while excluding the hardest cases.
- Administrative burden
- The learning, compliance, and psychological costs a citizen bears when interacting with government.
- Take-up rate
- The share of eligible people who actually receive a benefit; population effect equals per-participant effect times take-up.
- State capacity
- The staffing, systems, data, and expertise available to execute policy, built incrementally over years and constraining all delivery.
Module 6: Domains as Case Studies, and the Analyst's Craft
Six contested policy domains worked as case studies with the strongest case stated on each side and the evidence given as it actually stands: health coverage, education, poverty and the safety net, climate and energy instruments, criminal justice, and immigration. Closes with the ethics of the job and where this work is actually done.
Health Coverage, Education, and the Safety Net
- Compare health coverage architectures across countries and identify which disagreements are empirical and which are about values.
- Summarize what the evidence does and does not establish about class size, teacher quality, and school choice.
- Analyze safety net programs using the poverty measurement, take-up, and evaluation tools from earlier modules.
The big picture
Everything in the previous five modules was preparation for this. Now you apply it. Three domains, each with an enormous literature, each with real disagreement, and each treated here the same way: the structure of the problem, what the evidence actually supports, where competent researchers disagree, and which parts of the argument are not about evidence at all.
You will notice a pattern across all three. The empirical disputes are usually narrower than the political ones, and the political ones usually persist because people weigh the same findings differently, not because one side is ignoring data. That pattern is the most useful thing this course can leave you with.
Health coverage: the structure of the problem
Start with the facts nobody disputes. The United States spends roughly seventeen to eighteen percent of gross domestic product on health care, far more than any peer country, most of which spend somewhere between nine and twelve percent. American life expectancy is below the peer average. The uninsured share of the population fell substantially after 2010, from roughly sixteen percent to figures around eight percent in recent years, with coverage gains concentrated in Medicaid expansion states and the subsidized marketplaces.
The most important empirical finding about American health spending is often misunderstood. Comparative work associated with Gerard Anderson and Uwe Reinhardt found that Americans do not consume dramatically more health care than people in peer countries; they pay dramatically more per unit. Hospital stays, procedures, imaging, and pharmaceuticals cost multiples of what they cost elsewhere. Administrative costs, spread across many payers each with its own rules, are also substantially higher. This matters for policy design because if the problem were overuse, the remedy would be demand-side cost sharing, and if the problem is prices, the remedy is somewhere in payment policy.
Now the architecture. A widespread misconception treats the choice as American markets versus socialized medicine. The comparative record does not support that binary. Every wealthy democracy has achieved universal or near-universal coverage, and they have done it through strikingly different mixes.
| Model | How it works | Examples |
|---|---|---|
| Single payer | One public insurer; providers may be private | Canada; Taiwan |
| National health service | Public financing and largely public provision | United Kingdom |
| Regulated multi-payer | Private or nonprofit insurers with mandates, subsidies, and price regulation | Germany; Netherlands; Switzerland; Japan |
| Mixed categorical | Different programs for different groups, with gaps | United States |
Switzerland and the Netherlands are especially instructive because they achieve universal coverage through competing private insurers, individual mandates, heavy subsidies, and tight regulation of what plans must cover and what they may charge. Their existence undercuts a claim made on both American flanks: that universal coverage requires a single government insurer, and that private insurance markets cannot deliver universality.
Key idea: Universal coverage has been achieved through single-payer, national service, and regulated private-insurer models alike, so the American debate is less about whether markets can work than about which regulatory bargain to strike and what transition costs to accept.
Health coverage: what the evidence supports, and where it stops
From Lesson 11 you already know the Oregon findings: coverage substantially improves financial protection and mental health, increases use of care, and did not produce significant two-year changes in three physical markers with intervals too wide to establish a zero.
Quasi-experimental work extends the picture. Studies by Benjamin Sommers and colleagues comparing states that expanded Medicaid with those that did not have found reductions in mortality, and Card, Dobkin, and Maestas used the age-sixty-five Medicare discontinuity to find increases in care use and some evidence of health benefits for the severely ill. These are quasi-experimental, so they carry the parallel-trends and local-estimate caveats of Lesson 12, and researchers differ over magnitudes.
Where the honest disagreements sit:
How much does coverage improve physical health, and over what horizon? Financial protection is well established. Mortality effects are supported by several quasi-experimental studies and contested in magnitude by others.
Would price regulation reduce innovation? The United States pays the highest pharmaceutical prices and hosts a large share of global biomedical research. Critics of price regulation argue those facts are causally linked and that lower revenues would reduce the flow of new drugs. Supporters argue that much foundational research is publicly funded, that a substantial share of industry spending goes to marketing and to drugs offering little therapeutic advance, and that other wealthy countries also innovate. The empirical work here is genuinely unsettled, and the size of the innovation elasticity is one of the most consequential unknown numbers in health policy.
How large would single-payer administrative savings be? Estimates vary widely depending on assumptions about provider payment rates, utilization increases from removing cost sharing, and transition costs. Analyses produced by advocates and by critics have differed by trillions of dollars over a decade, and most of the gap traces to a handful of stated assumptions rather than to hidden tricks. This is a case where reading the assumptions section is the entire exercise.
And the values question that no study resolves: how much should a society redistribute to finance care, and how much should individual choice of plan and provider be preserved when it conflicts with cost control? People who agree on every number can disagree here.
Education: three well-studied questions
Class size. Tennessee's Project STAR randomized roughly eleven thousand six hundred students and their teachers into small classes, regular classes, and regular classes with an aide in kindergarten through third grade. Small classes produced measurable achievement gains, larger for disadvantaged students, and later work linked participation to better long-run outcomes. This is one of the strongest randomized findings in education.
Then California reduced class sizes statewide in 1996, and the results were far weaker. The reason is a textbook voltage effect from Lesson 12: hiring tens of thousands of teachers rapidly reduced average teacher qualifications, and the schools most in need lost experienced staff to districts that could recruit more easily. The policy lesson is not that class size does not matter. It is that a reform requiring a large increase in a scarce input may destroy its own effect through the labor market.
Teacher quality. Work by Raj Chetty, John Friedman, and Jonah Rockoff using large administrative datasets found that students assigned to teachers with high value-added scores had better long-run outcomes including college attendance and adult earnings. Jesse Rothstein and others raised serious objections about whether students are sorted to teachers in ways the models cannot fully absorb. The broad conclusion that teachers vary substantially in effectiveness is widely accepted. Whether value-added scores are precise enough to be used for individual personnel decisions is much more contested, since year-to-year stability for a single teacher is modest.
School choice. This is where careful reading matters most, because the evidence is heterogeneous in a specific, informative way.
Lottery-based studies of oversubscribed urban charter schools, particularly in Boston and New York City, have found large achievement gains, especially for low-income and minority students. Because these use admission lotteries, they are randomized and internally strong.
Broader studies including charters nationally find results closer to a wash on average, with wide variation: some sectors clearly outperform, others clearly underperform, and the average conceals both.
Voucher evidence has surprised nearly everyone. Earlier evaluations in Milwaukee, Washington, and New York generally found small positive or null achievement effects, sometimes with graduation gains. Then statewide programs in Louisiana and Indiana produced significant negative achievement effects in early years, results that advocates did not expect and did not suppress. Some of those effects attenuated over time in follow-up work.
The defensible summary is that specific school models in specific settings produce large gains, that expanding choice as a general mechanism does not automatically reproduce them, and that program design, including which private schools participate and under what accountability, appears to matter enormously. Underneath sits a values dispute that evidence cannot settle: how much weight to give parental authority and religious liberty relative to measured achievement and to the effects on students who remain in district schools.
Key idea: In education, the same intervention can be strongly effective in a controlled or oversubscribed setting and weak or negative at scale, so the correct question is never whether an idea works but under what conditions, at what scale, and against what alternative.
Poverty and the safety net
Begin with measurement, because it drives everything. From Lesson 6 you know the official poverty measure counts pre-tax cash income and cannot see SNAP, refundable tax credits, or housing assistance, while the Census Bureau's Supplemental Poverty Measure counts all of them and adjusts for local costs.
Using the supplemental measure, the safety net has a large measured effect. The clearest natural experiment of recent years is the temporary expansion of the Child Tax Credit in 2021, which made the credit larger and fully refundable. Child poverty on the supplemental measure fell to a record low of roughly five percent in 2021, then rose to roughly twelve percent in 2022 when the expansion lapsed. The direction and rough magnitude of that swing are not seriously disputed.
What is disputed is what to conclude. Supporters read it as proof that child poverty is a policy choice with a known, affordable remedy. Critics raise three objections that deserve statement: the long-run effects of an unconditional child benefit on parental employment are contested, with different modeling teams producing very different projected work responses; the cost is substantial and competes with other uses; and a measure defined by counted income will mechanically improve when counted income is provided, which is true and does not by itself mean well-being improved by the same amount, though evidence on food insufficiency during the expansion pointed the same direction.
The Earned Income Tax Credit is the most studied American antipoverty program and the one with the broadest coalition. Research by Nada Eissa, Jeffrey Liebman, Bruce Meyer, and Dan Rosenbaum found it substantially increased employment among single mothers, and later work has associated it with improved infant health and children's academic outcomes. Its recognized weaknesses are the high effective marginal tax rates created by the phase-out range, its small size for workers without children, its dependence on tax filing, and an improper payment rate driven largely by the complexity of the qualifying-child rules rather than by deliberate fraud.
Work requirements are the sharpest live dispute. Supporters argue that conditioning aid on work promotes self-sufficiency and public legitimacy, and point to the sharp decline in cash assistance caseloads and the rise in single-mother employment after the 1996 welfare reform. Critics note that the late 1990s had an exceptionally strong labor market, that the employment gains are hard to disentangle from it and from the EITC expansion, and that research by Kathryn Edin and Luke Shaefer documented growth in extreme poverty among families with essentially no cash income.
One case is unusually clean. Arkansas imposed work reporting requirements on Medicaid in 2018. Research by Sommers and colleagues found roughly eighteen thousand people lost coverage with no measurable increase in employment, and survey evidence indicated many of those disenrolled were already working or exempt but did not navigate the reporting system. This is Lesson 13's administrative burden operating exactly as predicted, and it is a strong result about that specific design. It does not settle the broader question of whether work requirements can be designed to function, which supporters correctly note.
Cash versus in-kind. Economists generally prefer cash on efficiency grounds, since recipients know their own needs. In-kind provision is defended on paternalistic grounds, on grounds that it commands more political support, and because some goods have externalities. A widely cited review by David Evans and Anna Popova found little evidence that cash transfers increase spending on alcohol and tobacco, contrary to a common assumption. Guaranteed income pilots have proliferated: Finland's national experiment found improved well-being with little employment effect, and a large recent American unconditional cash study found modest reductions in work hours alongside other changes. Sample sizes, durations, and populations differ enough that a general conclusion about guaranteed income would be premature, and honest advocates on both sides acknowledge this.
Common misconceptions
- Universal coverage requires a single government insurer. Switzerland, the Netherlands, and Germany reach universality through regulated private or nonprofit insurers with mandates and subsidies.
- Americans overuse health care. Comparative evidence points primarily to higher prices per unit and higher administrative costs rather than greater utilization.
- Project STAR proves class size reduction works as policy. It proves small classes helped in a controlled setting; California's statewide reduction was diluted by the teacher labor market it created.
- School choice evidence points one way. Urban charter lotteries show large gains, national averages show a wash, and some statewide voucher programs showed negative effects, which together indicate that design and setting dominate.
- The 2021 child poverty drop settles the child allowance debate. The measured drop is not disputed; projected long-run employment effects and cost trade-offs are, and those are where the argument actually lives.
Recap
- American health spending is driven mainly by prices and administrative complexity, and universal coverage has been achieved abroad through several distinct architectures.
- Coverage clearly improves financial protection and mental health; mortality effects are supported and contested in magnitude, and innovation and administrative-savings estimates hinge on stated assumptions.
- Class size and teacher quality findings are strong in controlled settings and degrade at scale through labor market effects, which is the voltage effect in practice.
- School choice evidence is heterogeneous by design and setting, with urban charter lotteries positive and some statewide voucher programs negative.
- The safety net's measured antipoverty effect is large under the supplemental measure, and the live disputes concern employment responses, cost, and whether specific conditioning designs can work without excluding eligible people.
Sources
- U.S. Census Bureau. (n.d.). Income and poverty. census.gov
- Congressional Budget Office. (n.d.). Health care. cbo.gov
- National Center for Education Statistics. (n.d.). Fast facts and reports. nces.ed.gov
- Internal Revenue Service. (n.d.). Earned Income Tax Credit (EITC). irs.gov
- Congressional Research Service. (n.d.). Reports on health, education, and income security policy. crsreports.congress.gov
- Key terms
- Regulated multi-payer system
- Universal coverage achieved through competing private or nonprofit insurers under mandates, subsidies, and price regulation, as in Switzerland and the Netherlands.
- All-payer rate setting
- A payment approach in which all insurers pay regulated prices for the same service, used to address price-driven cost growth.
- Value-added measure
- An estimate of a teacher's contribution to student achievement gains; widely accepted as showing variation, contested for individual personnel use.
- Lottery-based study
- An evaluation using an oversubscribed program's admission lottery as random assignment, giving strong internal validity for that setting.
- Supplemental Poverty Measure
- The Census measure counting tax credits, in-kind benefits, and local costs, which the official measure by construction omits.
- Effective marginal tax rate
- The share of an additional dollar of earnings lost to taxes plus benefit phase-outs, which can be high for low-income households.
- Work requirement
- A condition tying benefit receipt to employment or reporting; the Arkansas Medicaid case illustrates how reporting burden can disenroll compliant people.
- Innovation elasticity
- How much pharmaceutical research responds to changes in expected revenue; a contested and consequential unknown in drug pricing policy.
Climate, Criminal Justice, Immigration, and the Ethics of the Job
- Compare carbon taxes, cap-and-trade, standards, and subsidies on efficiency, certainty, and political durability.
- Summarize what the evidence supports about policing, sentencing, and incarceration, and identify the genuinely open questions.
- State the professional roles and ethical obligations of a policy analyst and identify where this work is done.
The big picture
Three more domains, then the job itself. These three are chosen because each contains a different kind of difficulty. Climate policy is a case where economists largely agree on the instrument and political scientists largely doubt it can pass, which makes the efficiency and feasibility criteria from Lesson 8 collide head-on. Criminal justice is a case where the evidence is stronger than the debate suggests on some questions and much weaker on others, in a pattern that surprises people. Immigration is a case where a single famous empirical dispute has run for thirty years without resolution, and where the largest disagreements are frankly about values.
Climate and energy: four instruments
Greenhouse gas emissions are the textbook negative externality from Lesson 4, with two aggravating features: the harm is global, so no single jurisdiction captures the benefit of its own abatement, and it is intergenerational, which throws you straight into the discount rate fight from Lesson 9. The social cost of carbon, the estimated damage from an additional ton of carbon dioxide, is the parameter that ties it all together, and published federal estimates have ranged from roughly fifty dollars per ton to figures around one hundred ninety dollars per ton, with the spread driven mostly by the discount rate and the damage function rather than by disputes about physical science.
| Instrument | Certainty it provides | Main strength | Main weakness |
|---|---|---|---|
| Carbon tax | Price certain, quantity uncertain | Simple, raises revenue, equalizes marginal abatement cost | Emissions outcome not guaranteed; politically difficult to enact |
| Cap-and-trade | Quantity certain, price uncertain | Guarantees an emissions ceiling; permits can be given away to buy support | Price volatility; over-allocation destroys the signal |
| Performance standards | Neither, directly | Politically durable; costs hidden in product prices | Higher cost per ton; rebound effects |
| Subsidies and technology push | Neither | Creates constituencies; drives learning-by-doing cost declines | Costs the budget; can pay for actions that would have happened anyway |
Martin Weitzman's classic analysis of prices versus quantities explains why the choice between the first two is not arbitrary. If the marginal damage curve is steep, meaning a bit more emissions is much worse, quantity control is safer. If the marginal cost curve is steep, meaning a small overshoot in required abatement is very expensive, price control is safer. Hybrid designs with price floors and ceilings inside a cap try to get both, and most functioning systems now use them.
The record gives real evidence on each. The American acid rain program, which capped sulfur dioxide from power plants beginning in the 1990s, achieved its emissions reductions ahead of schedule at costs far below both industry and agency projections, and it is the strongest existing case for tradable permits. The European Union Emissions Trading System had a rough start: over-allocation of permits plus the 2008 recession collapsed the price to near nothing for years, which taught the field that a cap set too loosely is a cap in name only. Reforms including a market stability reserve later raised prices substantially. British Columbia's revenue-neutral carbon tax, introduced in 2008, has been studied extensively; most analyses find modest emissions reductions without evident harm to provincial economic growth, though estimates of the magnitude vary and the province is small and trade-exposed.
Here is the collision. Surveys of economists find broad support for carbon pricing as the most cost-effective instrument. Yet carbon pricing has repeatedly failed at the ballot box and in legislatures, including two defeated Washington State initiatives and the French fuel tax increase that triggered mass protests in 2018. Meanwhile standards and subsidies, which economists rate as less efficient per ton, have proven far more politically durable, partly because their costs are diffuse and invisible while a tax is visible and attributable. Political scientists including those who study policy feedback argue that subsidies build constituencies that defend the policy over time, which is a form of durability that a static efficiency calculation cannot see. Whether an efficient instrument that cannot pass is better than an inefficient one that can is not an empirical question. It is a judgment about how to weight the criteria from Lesson 8, and thoughtful people weight them differently.
Key idea: Carbon pricing dominates on cost-effectiveness per ton and has repeatedly lost politically, while standards and subsidies cost more per ton and stick. Ranking them requires weighing efficiency against feasibility and durability, which is a value judgment, not a calculation.
Criminal justice: what is known and what is not
The scale facts first. The United States incarcerates a far larger share of its population than any peer democracy, by a multiple rather than a margin. The state and federal prison population rose steeply from the 1970s, peaked around 2009, and has declined since, with substantial variation across states. Violent crime fell dramatically from the early 1990s through the mid 2010s, a decline whose causes remain one of the most argued questions in social science.
Findings with relatively strong support:
Police numbers reduce crime. Reviews by Aaron Chalfin and Justin McCrary, drawing on quasi-experimental studies that exploit sudden changes in police staffing, find that additional officers reduce serious crime, with effects on homicide that appear substantial. The identification strategies vary in quality and estimates differ, but the direction is consistent across many designs.
Certainty of apprehension matters more than severity of punishment. Daniel Nagin's synthesis of the deterrence literature reaches this conclusion repeatedly. Increasing the probability of being caught deters; lengthening sentences at the margin does much less, partly because offenders heavily discount distant consequences.
Hot-spots policing produces modest crime reductions. Meta-analyses of randomized and quasi-experimental trials find that concentrating police attention on small high-crime locations reduces crime there, and generally find diffusion of benefit to nearby areas rather than pure displacement.
Findings that are genuinely contested:
How much did incarceration contribute to the crime decline? Published estimates attribute somewhere between roughly ten and twenty-five percent of the 1990s decline to rising incarceration, with most analysts also concluding that returns diminish sharply at high incarceration levels, since additional imprisonment increasingly captures lower-risk people. Both the size and the shape of that curve are disputed.
What causes long-run crime trends? Candidate explanations include policing changes, demographics, drug market dynamics, economic conditions, and environmental lead exposure. No consensus exists on the weights.
Bail and pretrial reform. Evidence is accumulating and mixed across jurisdictions, partly because reforms differ substantially in design and partly because evaluation periods overlapped with pandemic disruption.
One case deserves special attention as a methodological lesson. Hawaii's HOPE probation program, applying swift and certain but modest sanctions for violations, produced dramatic results in an early evaluation and generated national enthusiasm. A multi-site randomized replication in four other jurisdictions largely failed to reproduce those effects. This is the voltage effect and the replication problem from Module 5 in a single case, and it is a good reason to be cautious about any program whose evidence base is one striking evaluation.
The values questions here are unusually explicit. Punishment serves deterrence, incapacitation, rehabilitation, and retribution, and these can conflict. A sentence that is efficient for incapacitation may be disproportionate as retribution, or the reverse. No amount of data tells you how to weigh a victim's claim to justice against a defendant's prospects, or how much collateral harm to children and communities to accept for a given reduction in offending.
Immigration: a thirty-year empirical dispute
The fiscal question has the most careful synthesis. The National Academies of Sciences, Engineering, and Medicine's 2016 report found that first-generation immigrants generate net fiscal costs at the state and local level, driven largely by the cost of educating children, while contributing positively at the federal level; that the second generation is among the strongest net fiscal contributors in the population; and that long-run aggregate effects are generally positive, though the estimate depends heavily on how public goods such as defense are allocated and on the time horizon chosen. That last clause is not a hedge. Different defensible accounting conventions move the answer substantially, and analysts who quote a single figure without stating the convention are omitting the key assumption.
The labor market question contains the field's most famous unresolved dispute. David Card studied the 1980 Mariel boatlift, in which roughly one hundred twenty-five thousand Cubans arrived in Miami over months, expanding the local labor force by several percent almost overnight. He found little effect on wages or unemployment of existing workers, including low-skilled ones. George Borjas later reanalyzed the episode focusing on a narrower subgroup of low-skilled men without a high school diploma and reported large wage declines. Giovanni Peri, Michael Clemens, and others responded that the subgroup sample was very small and that a contemporaneous change in survey composition could account for the result. Borjas replied in turn. The exchange has been running for years among serious economists with real methodological arguments on each side.
What most economists would sign: aggregate effects of immigration on native wages are small; immigration expands the economy and the tax base; effects are unevenly distributed, with the strongest competition falling on workers whose skills are closest substitutes, which often means earlier immigrants; and high-skilled immigration is associated with increased patenting and business formation, an area where the evidence is comparatively consistent.
What remains contested: the magnitude of wage effects on the least-educated native workers, how quickly capital and firm entry adjust to absorb new labor, and how much of measured complementarity depends on modeling choices about skill categories.
And the largest part of the immigration debate is not empirical at all. How much weight a country should give to humanitarian obligation, to the interests of prospective migrants who are not its citizens, to national self-determination over membership, to the rule of law in enforcement, and to cultural continuity are value questions. An analyst can clarify the fiscal and labor market consequences of a given policy with real precision. An analyst cannot tell a society how many people it ought to admit, and should say so plainly rather than dressing a value judgment in a fiscal estimate.
Key idea: Across all three domains, the empirical disagreements are narrower and more technical than the public argument suggests, and the largest divides are about how to weigh competing values that evidence cannot rank.
The analyst's roles and obligations
David Weimer and Aidan Vining describe three roles an analyst can occupy, and being clear about which one you are in is the beginning of professional ethics.
The objective technician sees analysis as the source of authority and keeps distance from the client's preferences, letting the analysis speak. The risk is irrelevance, since analysis nobody uses changes nothing.
The client's advocate sees the client relationship as the source of influence and works to advance the client's goals within the bounds of honesty. The risk is becoming a technician of whatever the client already wanted.
The issue advocate uses analysis to advance a substantive conception of the good. The risk is that the analytic work becomes a vehicle and stops being a check.
Most working analysts move among these. The obligation is not to occupy one permanently but to know which you are in and to disclose it.
A workable set of professional norms, drawn from the field's practice:
- State every material assumption where a reader can find it, and make the analysis reproducible enough that a skeptic could rerun it.
- Report uncertainty honestly, using ranges and switching values rather than false precision.
- Represent the literature as it stands, including findings that cut against your conclusion, and do not describe a contested question as settled.
- Disclose your own stake or strong value commitment in the outcome.
- Separate analytic findings from recommendations, so a reader who shares your facts but not your values can still use your work.
- Decline to produce a predetermined number. If a client wants a specific answer, the professional response is to say what the analysis shows and let them decide what to do with it.
- Say when you do not know. Aaron Wildavsky titled a book Speaking Truth to Power, and its most useful lesson is that the analyst's credibility is a capital asset spent one memo at a time.
Where this work is done
Policy analysis is a real occupation with identifiable employers. At the federal level, the Congressional Budget Office produces cost estimates and baseline projections, the Government Accountability Office audits and evaluates programs, and the Congressional Research Service writes nonpartisan issue briefs for members of Congress; all three are institutionally committed to nonpartisanship and are unusually good places to learn the craft. Executive agencies have policy and evaluation offices, and the Office of Information and Regulatory Affairs reviews regulatory analyses.
Beyond Washington, state legislative fiscal offices and legislative analyst offices do the same work with smaller staffs and shorter deadlines, city budget and performance offices increasingly employ analysts, and the analytic units of large school districts, transit agencies, and health systems do it inside single domains. Outside government there are research organizations, foundations, advocacy groups, consulting firms, and international institutions such as the World Bank and the OECD.
The skills that transfer across all of these are consistent: write clearly and briefly for busy readers; handle data and basic causal inference competently enough to read a study and to know when to call for help; understand budgeting and how money actually moves; and know one substantive domain deeply enough that people trust you in it. Nothing in the list is exotic. Most of it is what this course has been practicing.
Common misconceptions
- Economists disagree about whether carbon pricing is efficient. Broad agreement exists on cost-effectiveness per ton; the live dispute is about political feasibility and durability against standards and subsidies.
- Longer sentences deter crime proportionally. The deterrence literature consistently finds certainty of apprehension matters far more than severity at the margin.
- The Mariel dispute shows that immigration research is worthless. It shows that a specific subgroup estimate is fragile, while broader findings about aggregate effects and high-skilled immigration are comparatively consistent.
- A fiscal estimate can settle immigration policy. The fiscal answer depends on accounting conventions, and the largest questions in the debate are about values rather than accounting.
- Objectivity requires having no view. It requires disclosing your view, stating assumptions, representing contrary evidence, and keeping findings separable from recommendations.
Recap
- Carbon taxes fix price and leave quantity uncertain while caps do the reverse, and Weitzman's analysis explains when each is safer; hybrids with floors and ceilings are now standard.
- The acid rain program, the early EU trading system, and British Columbia's tax provide real evidence, and the efficiency-versus-durability trade-off against standards and subsidies is a value judgment.
- Police numbers, certainty over severity, and hot-spots policing have relatively strong support; incarceration's contribution to the crime decline and the causes of long-run trends remain contested, and HOPE's failed replication is a caution.
- Immigration's fiscal effects depend on accounting conventions, labor market effects on the least-educated remain disputed after thirty years, and the central questions are about values.
- Analysts occupy technician, client advocate, and issue advocate roles; the obligations are transparency about assumptions, honest uncertainty, faithful representation of the literature, disclosure of stakes, and refusal to produce a predetermined number.
Sources
- U.S. Environmental Protection Agency. (n.d.). Acid Rain Program. epa.gov
- Congressional Budget Office. (n.d.). Climate change and energy. cbo.gov
- Bureau of Justice Statistics. (n.d.). Corrections and correctional populations data. bjs.ojp.gov
- U.S. Department of Homeland Security. (n.d.). Immigration statistics. dhs.gov
- U.S. Government Accountability Office. (n.d.). What GAO does. gao.gov
- Key terms
- Social cost of carbon
- Estimated damage from an additional ton of carbon dioxide; published federal estimates have ranged widely, driven mainly by discount rate and damage function choices.
- Prices versus quantities
- Weitzman's result that a tax is safer when marginal abatement costs are steep and a cap is safer when marginal damages are steep.
- Over-allocation
- Issuing more emission permits than the market needs, which collapses the permit price and neutralizes a cap, as in the early EU trading system.
- Policy feedback
- The tendency of an enacted policy to create constituencies and expectations that make it politically durable over time.
- Certainty versus severity
- Nagin's finding that raising the probability of apprehension deters far more than lengthening sentences at the margin.
- Hot-spots policing
- Concentrating police attention on small high-crime locations; meta-analyses find modest reductions with diffusion of benefit rather than displacement.
- Objective technician
- Weimer and Vining's analyst role that keeps distance from client preferences and treats the analysis as the source of authority.
- Switching value disclosure
- The practice of reporting the parameter values at which a recommendation would reverse, so a reader can apply their own judgment.