Module 1: Foundations - What Epidemiology and Public Health Are
The definition and logic of epidemiology, the population perspective of public health, and the core measures used to count disease: incidence and prevalence.
What Epidemiology and Public Health Are
- Define epidemiology and describe its role as the basic science of public health.
- Explain the population perspective and how it differs from clinical medicine.
- Describe the epidemiologic triad and the aims of descriptive versus analytic epidemiology.
Epidemiology is the study of the distribution and determinants of health-related states and events in defined populations, and the application of that study to control health problems. Every phrase in that sentence is load-bearing. Distribution is the counting and mapping of health: who is affected, where, and when. Determinants are the causes, risk factors, and protective factors that explain the pattern. Defined populations anchors the work in groups with a countable denominator, not isolated individuals. And application to control reminds us that epidemiology is not knowledge for its own sake; it exists to prevent disease.
This makes epidemiology the basic science of public health, playing a role like the one physiology plays for clinical medicine. When a new illness appears, epidemiologists ask how many people have it, who is at greatest risk, how it spreads, and what would reduce it. The answers become the evidence on which prevention, treatment guidelines, and policy are built. A useful shorthand describes the core tasks as counting cases, relating those cases to a population at risk, and comparing groups to uncover causes. The chapters ahead take up each task in turn.
The population perspective
Clinical medicine asks, "What is wrong with this patient, and how do I treat them?" Public health asks, "Why do some populations have more disease than others, and how do we prevent it?" The unit of concern shifts from the individual to the population. A clinician improves one life at a time; an epidemiologist tries to shift the health of thousands at once. Both roles are essential and they inform each other, but the population lens changes which questions seem important and which solutions seem worthwhile.
This shift has a profound consequence known as the prevention paradox: a preventive measure that brings large benefit to a whole population may offer little to each participating individual. A small downward shift in the entire distribution of blood pressure prevents more strokes across a society than intensively treating only the few people with the very highest pressures, even though any one person feels almost nothing. Most cases often arise from the large number of people at modest risk rather than the small number at extreme risk, which is why the paradox matters so much.
Geoffrey Rose framed this as a contrast between two strategies. The high-risk strategy finds and treats the individuals whose risk is greatest; it is efficient per person helped and appeals to clinicians, but it reaches only the tail of the distribution and leaves the bulk of cases untouched. The population strategy nudges the whole distribution toward health, for example by lowering salt across the food supply; it can prevent far more disease overall, though each person gains only a little and may never know they were helped. Public health usually needs both.
Public health and its core functions
Public health is what a society does collectively to assure the conditions in which people can be healthy. Its three recognized core functions are assessment, policy development, and assurance. Assessment means measuring the health of populations, largely through epidemiology and surveillance. Policy development means using that evidence to craft laws, programs, and guidelines. Assurance means making sure needed services are actually delivered to the people who need them. Epidemiology supplies the facts that make all three possible; without sound measurement, policy is guesswork.
The triumphs of public health are largely invisible because they are absences: outbreaks that never happened, children who did not die. Clean water and sanitation, vaccination, safer workplaces, tobacco control, motor-vehicle safety, and the control of infectious disease account for far more of the twentieth-century gain in life expectancy in wealthy countries than clinical care did. A common misconception credits modern medicine with most of the historical fall in deaths from infection. In fact much of that decline preceded effective drugs and followed improvements in living conditions, nutrition, and sanitation.
A founding example: John Snow and the Broad Street pump
The methods of the field are often traced to London in 1854, when a severe cholera outbreak struck the Soho district. The physician John Snow, skeptical of the prevailing theory that disease spread through bad air, plotted each death on a street map. The cases clustered tightly around a single public water pump on Broad Street. Snow persuaded local officials to remove the pump handle, and the outbreak, already waning, subsided. He had reasoned from a population pattern to a cause and then to an action, decades before the cholera bacterium was ever seen under a microscope.
Snow went further, comparing districts supplied by different water companies, one drawing from a sewage-tainted stretch of the Thames and one from cleaner water upstream. Households served by the contaminated supply died of cholera at many times the rate of their neighbors. This was analytic epidemiology in all but name: two comparable groups differing mainly in one exposure, with a large difference in outcome. The example shows the whole arc of the discipline, from counting and mapping, to comparing groups, to guiding a concrete public health decision that saved lives.
The natural history of disease
Disease unfolds along a timeline called its natural history. It begins in a stage of susceptibility, when risk factors are present but disease has not started. Next comes a subclinical stage, in which biological changes have begun but the person feels nothing; the moment disease could first be caught by a test marks the start of the detectable preclinical phase. Then comes clinical disease, with signs and symptoms, and finally recovery, disability, or death. Understanding this timeline matters because each level of prevention is defined by where along it the action is aimed.
The epidemiologic triad and the levels of prevention
A classic model of disease causation is the epidemiologic triad: an agent (a microbe, a toxin, a nutrient deficiency), a host (the person, with their susceptibility), and an environment that brings the two together, often with a vector carrying the agent. Break any leg of the triad and transmission stops. The model was built for infectious disease but generalizes well: for a road-traffic injury the agent is kinetic energy, the host is the traveler, and the environment includes the road and the vehicle that concentrate and deliver that energy.
Prevention is layered to match the natural history. Primary prevention stops disease before it starts, acting in the stage of susceptibility, as with vaccination, seatbelts, or clean water. Secondary prevention detects disease early, in its subclinical phase when it is still treatable, as with screening tests. Tertiary prevention limits disability once disease is established, through rehabilitation and careful management of complications. Some authors add primordial prevention, which acts earlier still by keeping the risk factors themselves from arising, for instance through a food environment that never fosters widespread obesity.
Descriptive and analytic epidemiology
Epidemiology proceeds in two broad modes. Descriptive epidemiology characterizes health by person, place, and time: who is affected, where, and when. It examines age, sex, occupation, geography, and season, and it is the natural first step in any investigation because it generates hypotheses cheaply. When Snow mapped cholera deaths around the Broad Street pump, he was doing descriptive epidemiology in space, and the map itself suggested the waterborne hypothesis he then went on to test with comparison groups.
Analytic epidemiology takes those hypotheses and tests them with comparison groups, asking whether an exposure is truly associated with an outcome and, eventually, whether it causes that outcome. The presence of a comparison group is the hinge between the two modes: a description of cases with no group to compare against can raise a question but cannot answer it. The rest of this course builds the analytic toolkit, measure by measure and design by design, yet all of it rests on the counting that descriptive epidemiology does first.
Determinants, from proximal to upstream
Determinants are not only biological. Epidemiologists distinguish proximal causes near the individual, such as a virus or a cholesterol level, from distal or upstream causes in the social and physical environment, such as income, education, housing, and access to care. These social determinants of health shape who is exposed to proximal risks in the first place. A full account of why one population suffers more disease than another usually runs through both levels, which is why public health looks beyond the clinic to the conditions in which people live, learn, and work.
The many uses of epidemiology
Beyond finding causes, epidemiology serves several practical ends that recur throughout this course. It measures the burden of disease so that services and budgets can be planned. It evaluates whether interventions and programs actually work, using the comparison-group logic of analytic studies. It informs the diagnosis and prognosis that clinicians give, since a test result or a survival estimate is meaningful only against population data. And it underpins surveillance, the ongoing monitoring that detects outbreaks and tracks trends. Each of these uses reappears in later modules with its own methods and cautions.
The discipline is unified by a single habit of mind: always ask, compared with what? A number of cases means little until set against a population and a time frame; a risk means little until compared with the risk in an unexposed group; a treatment result means little until compared with a control. Holding that question steady is what separates epidemiologic reasoning from anecdote, and it is the thread that connects counting, comparison, and causal judgment across every lesson that follows.
Common misconceptions
Three errors recur among newcomers. First, that epidemiology is only about infectious outbreaks; in fact its methods apply equally to heart disease, cancer, injuries, and mental health. Second, that a larger number of cases always means a worse problem; without a denominator, a raw count says little, as the next lesson shows in detail. Third, that association proves causation; separating the two is so central that four later lessons are devoted to bias, confounding, chance, and the disciplined reasoning that carries evidence from mere correlation toward genuine cause.
Try it
Classify each action by level of prevention, then answer the population question. (a) A measles vaccine. (b) A mammogram in a woman with no symptoms. (c) Cardiac rehabilitation after a heart attack. (d) Fluoridating drinking water. Also: a town of 10,000 has 100 people at very high stroke risk and 9,900 at modest risk; if the modest-risk group produces about 70 percent of all strokes, does treating only the high-risk group address most of the cases?
Worked answer: (a) primary, because it stops disease before it starts; (b) secondary, early detection of silent disease; (c) tertiary, limiting disability from established disease; (d) primary. For the population question, no: because roughly 70 percent of strokes arise in the large modest-risk majority, a purely high-risk strategy misses most cases. That is the prevention paradox in numbers, and the argument for pairing a population strategy with targeted treatment.
Sources
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 1, Section 1: Definition of epidemiology. CDC Self-Study Course SS1978. archive.cdc.gov
- Rose, G. (1985). Sick individuals and sick populations. International Journal of Epidemiology, 14(1), 32-38. pubmed.ncbi.nlm.nih.gov
- Brody, H., Rip, M. R., Vinten-Johansen, P., Paneth, N., & Rachman, S. (2000). Map-making and myth-making in Broad Street: The London cholera epidemic, 1854. The Lancet, 356(9223), 64-68. pubmed.ncbi.nlm.nih.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 1, Section 9: Natural history and spectrum of disease. CDC Self-Study Course SS1978. archive.cdc.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 1, Section 6: Descriptive epidemiology. CDC Self-Study Course SS1978. archive.cdc.gov
- Munnangi, S., & Boktor, S. W. (2023). Epidemiology of study design. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Institute of Medicine, Committee for the Study of the Future of Public Health. (1988). The future of public health. National Academy Press. find source ↗
- Key terms
- Epidemiology
- The study of the distribution and determinants of health states in populations, applied to control health problems.
- Population perspective
- The public health focus on the health of whole populations rather than individual patients.
- Prevention paradox
- The observation that a measure benefiting a population greatly may offer little to each individual in it.
- Epidemiologic triad
- The model of disease as the interaction of agent, host, and environment (often with a vector).
- Levels of prevention
- Primary (before disease), secondary (early detection), and tertiary (limiting disability).
- Descriptive vs analytic epidemiology
- Characterizing disease by person, place, and time versus testing exposure-outcome hypotheses with comparison groups.
Measuring Disease Frequency: Incidence and Prevalence
- Distinguish incidence proportion, incidence rate, and prevalence and compute each.
- Explain the relationship among incidence, duration, and prevalence.
- Choose the appropriate frequency measure for a given epidemiologic question.
Before you can study causes, you must count cases. But a raw count, such as "500 cases of disease," is nearly meaningless without a denominator: the population at risk and the time over which those cases arose. Five hundred cases in a city of ten million over a decade is a rarity; five hundred in a village of a thousand over a month is a catastrophe. Epidemiology therefore expresses disease as a rate or proportion, always cases divided by an appropriate population, so that numbers from different places and periods can be compared fairly.
The two great families of frequency measures are incidence, which counts new cases, and prevalence, which counts existing cases. They answer different questions and are constantly confused, so it is worth learning each precisely before combining them. Incidence is about the risk or speed of becoming ill; prevalence is about how much illness is present right now. Getting the distinction wrong leads to misreading almost every health statistic you will encounter.
Rates, ratios, and proportions
Three words are used loosely in everyday speech but carefully here. A proportion is a part divided by the whole it belongs to, so it lies between 0 and 1; incidence proportion and prevalence are proportions. A rate in the strict sense has time in its denominator, as with an incidence rate of so many cases per person-year. A ratio divides two quantities that need not share a whole, such as a sex ratio of male to female cases. Keeping these straight prevents the common error of calling every fraction a rate.
Incidence: the rate of new disease
Incidence measures new cases arising in a population that was initially disease-free. It comes in two forms. The incidence proportion, also called cumulative incidence or the risk, is the number of new cases divided by the number of people at risk at the start, over a stated period. If 900 disease-free people are followed for one year and 30 develop the disease, the incidence proportion is 30 divided by 900, which equals 0.033, or 33 per 1,000 per year. It is a probability, ranging from 0 to 1, and it always needs a time frame attached.
The incidence proportion assumes you can follow everyone for the full period, a so-called closed cohort. When people are followed for different lengths of time, or enter and leave, we use the incidence rate, also called incidence density: new cases divided by the total person-time at risk. Person-time sums each individual's own time under observation, so someone watched for five years contributes five person-years and someone watched for six months contributes half a person-year. This lets a study handle staggered entry, dropout, and death without discarding partial information.
Work an example step by step. Five people enter a study. Person 1 stays well through 5 years. Person 2 develops disease after 2 years and stops contributing time. Person 3 is lost to follow-up after 3 years. Person 4 completes the full 5 years well. Person 5 develops disease after 4 years. Total person-time is 5 + 2 + 3 + 5 + 4 = 19 person-years, with 2 new cases. The incidence rate is 2 divided by 19, which equals 0.105 per person-year, or about 105 per 1,000 person-years.
Because its denominator is time, an incidence rate has no upper bound of 1 and cannot be read directly as a simple probability. A rate of 105 per 1,000 person-years does not mean 10.5 percent of people fall ill; it means that, on average, disease occurs about that often per unit of person-time observed. Rates are ideal for comparing the speed of new disease across groups followed for unequal durations, which is exactly the situation in most long cohort studies.
When risk and rate roughly agree
Over a short period in which little disease occurs and few people are lost, the incidence proportion and the incidence rate give nearly the same answer, because almost everyone contributes the full follow-up time. As the period lengthens, or as disease and dropout mount, the two diverge: the proportion is capped at 1, while the rate keeps climbing because its denominator is shrinking person-time. A practical rule is that for a rare outcome over a short interval you may treat them as interchangeable, but for a common outcome tracked over years you must respect the difference and report the rate.
Prevalence: the burden of existing disease
Prevalence is the proportion of a population that has the disease at a point in time, called point prevalence, or over an interval, called period prevalence. It counts old and new cases alike. If 240 of 1,200 people surveyed have the condition today, the point prevalence is 240 divided by 1,200, which equals 0.20, or 20 percent. Prevalence is a snapshot of burden, well suited to planning services, staffing clinics, and allocating resources, because it tells you how many people need care at a given moment.
The distinction between point and period prevalence matters in practice. Point prevalence freezes a single instant, like a census day, and is what most surveys estimate. Period prevalence counts everyone who had the disease at any moment during an interval, including those who developed it partway through and those who recovered before it ended, so it is always at least as large as the point prevalence within that window. Lifetime prevalence, a common variant, asks whether a person has ever had the condition, and suits disorders that recur or remit.
What prevalence cannot do is measure the risk of getting a disease, because it mixes new cases with survival. A condition can be common simply because those who have it live a long time with it, not because many new cases occur. Cross-sectional surveys, which measure prevalence, also tend to over-represent long-lasting cases and miss those who recovered quickly or died, a distortion revisited when we discuss study designs. Use prevalence to describe how much disease is present, never to estimate how fast it is appearing.
How incidence and prevalence relate
A helpful image is a bathtub. The water level is prevalence, the pool of existing cases. The tap filling the tub is incidence, the flow of new cases. The drain emptying it is recovery and death. When inflow and outflow balance, the level holds steady; raise the tap or narrow the drain and the level rises. This is why a treatment that keeps patients alive longer, narrowing the drain, can raise prevalence even though incidence, the tap, has not changed at all.
In a steady state, the relationship is approximately:
Prevalence ≈ Incidence × Duration
A disease with low incidence but long duration, such as a chronic non-fatal condition, can have high prevalence, while a common but brief illness like a cold has high incidence yet low prevalence at any moment. Worked values: if incidence is 0.02 per year and average duration is 5 years, prevalence is about 0.02 times 5, or 0.10 (10 percent); double the duration to 10 years and prevalence rises to about 0.20; halve incidence to 0.01 with duration 5 years and prevalence falls to about 0.05.
This single relationship explains a common trap. A rising prevalence can mean more people are getting sick, a higher incidence, or that patients are living longer with the disease, a longer duration that often signals better treatment. These are opposite stories told by the same number, which is why a prevalence trend alone can never tell you whether a population is getting healthier or sicker.
| Measure | Numerator | Denominator | Best for |
|---|---|---|---|
| Incidence proportion (risk) | New cases | People at risk at start | Probability of developing disease |
| Incidence rate | New cases | Person-time at risk | Speed of new disease with varying follow-up |
| Prevalence | All existing cases | Total population | Burden of disease for planning |
Mortality and case fatality
Death has its own frequency measures, and two of them are routinely mixed up. The mortality rate is essentially an incidence rate of death: deaths divided by the population (or person-time) over a period. If a city of 100,000 records 50 deaths from a disease in a year, the mortality rate is 50 divided by 100,000, or 50 per 100,000 per year. It measures the disease's toll across the whole population, healthy and ill together, and so reflects both how common the disease is and how deadly it is.
The case-fatality ratio is different: it is the proportion of people with the disease who die of it, a measure of how lethal the illness is once you have it. If an outbreak infects 500 people and 25 die, the case-fatality ratio is 25 divided by 500, or 5 percent. A disease can have a high case-fatality ratio yet a low mortality rate if it is rare, or a modest case-fatality ratio yet a high mortality rate if it is widespread. The two numbers answer distinct questions and should never be swapped.
Choosing the right measure
Match the measure to the question. To ask how likely a healthy person is to develop disease over a defined period, use the incidence proportion. To compare the speed of new disease across groups followed for unequal times, use the incidence rate with person-time. To describe how many people currently live with a condition, for planning care, use prevalence. To judge how deadly a disease is once contracted, use case fatality; to judge its toll on a whole community, use the mortality rate. The wrong measure can invert the story entirely.
Crude, specific, and adjusted measures
A single number for a whole population is a crude rate, mixing everyone together. Because disease frequency often varies sharply with age and sex, epidemiologists also compute specific rates within strata, such as the incidence in women aged 60 to 69. Comparing crude rates between two populations can mislead when they differ in age structure, since an older population will show more of most chronic diseases regardless of any real difference in risk. Adjusting or standardizing rates to a common age structure removes that distortion and lets fair comparisons be made across places and eras.
Common misconceptions
The frequent error is treating prevalence as if it were risk. Reading that a condition is "twice as prevalent" in one group does not mean members of that group are twice as likely to develop it, since longer survival or slower recovery could produce the same gap with identical incidence. A second error is quoting an incidence rate as a percentage, forgetting that its denominator is person-time, not people. A third is comparing crude counts across populations of different sizes without forming a proper rate, which makes a large population look sick merely for being large.
Try it
A clinic serves 4,000 adults. Today, 320 are living with a chronic condition, and over the past year 80 previously healthy adults were newly diagnosed. First compute the point prevalence and the one-year incidence proportion. Then, given an average disease duration of 8 years, use the steady-state relationship to estimate what long-run prevalence this incidence and duration would sustain. Show each step rather than jumping to the answer.
Worked answer: point prevalence = 320 / 4,000 = 0.08, or 8 percent, the burden now. Incidence proportion = 80 / (4,000 - 320 at risk) = 80 / 3,680 = 0.0217, about 21.7 per 1,000 per year, the new risk. Steady-state prevalence ≈ incidence × duration = 0.0217 × 8 ≈ 0.174, about 17 percent. The estimate exceeds today's 8 percent, suggesting prevalence would climb over time if this incidence and long duration persisted.
Sources
- Tenny, S., & Boktor, S. W. (2023). Incidence. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Tenny, S., & Hoffman, M. R. (2023). Prevalence. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Noordzij, M., Dekker, F. W., Zoccali, C., & Jager, K. J. (2010). Measures of disease frequency: Prevalence and incidence. Nephron Clinical Practice, 115(1), c17-c20. pubmed.ncbi.nlm.nih.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 3, Section 2: Morbidity frequency measures. CDC Self-Study Course SS1978. archive.cdc.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 3, Section 3: Mortality frequency measures. CDC Self-Study Course SS1978. archive.cdc.gov
- Hernandez, J. B. R., & Kim, P. Y. (2022). Epidemiology morbidity and mortality. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Ritchie, H., Spooner, F., & Roser, M. (2024). Causes of death. Our World in Data. ourworldindata.org
- Key terms
- Denominator
- The population at risk (and time) against which cases are counted to form a rate or proportion.
- Incidence proportion (risk)
- New cases divided by the number at risk at the start over a period; a probability from 0 to 1.
- Incidence rate (density)
- New cases divided by total person-time at risk; its denominator is time, not people.
- Person-time
- The sum of the time each individual is observed and at risk, used as the denominator of an incidence rate.
- Prevalence
- The proportion of a population that has the disease at a point or over a period; counts existing cases.
- Prevalence approximately equals incidence times duration
- In a steady state, prevalence rises with both how fast disease occurs and how long it lasts.
Module 2: Measuring Association - The Two-by-Two Table
The two-by-two table as the workhorse of epidemiology, and the measures of association it yields: relative risk, odds ratio, and attributable risk.
The Two-by-Two Table and Relative Risk
- Set up a standard two-by-two table of exposure against disease.
- Compute and interpret risk in exposed and unexposed groups and the relative risk.
- State what a relative risk of 1, above 1, and below 1 mean.
The central question of analytic epidemiology is comparative: do exposed people develop disease more often than unexposed people? Almost every measure of association is read off a single tool, the two-by-two table, which cross-classifies each person by exposure (yes or no) and disease (yes or no). Its power is that it turns a messy dataset into four numbers from which risks, ratios, and differences all follow. By long convention the cells are labeled a, b, c, and d:
| Disease + | Disease - | Total | |
|---|---|---|---|
| Exposed + | a | b | a + b |
| Exposed - | c | d | c + d |
Here a is exposed people with disease, b is exposed without disease, c is unexposed with disease, and d is unexposed without disease. Learning to read this table fluently is the single most useful skill in the course, because the design you used determines which measure the table can legitimately give you. The same four cells support several calculations, and a large part of skill is knowing which one your study earns the right to compute.
Reading the margins
Two habits make the table trustworthy. First, fill in the marginal totals: the row totals a + b and c + d give the number exposed and unexposed, while the column totals a + c and b + d give the number diseased and disease-free. Second, decide which denominator your design permits. In a cohort you divide along the rows, using a + b and c + d, because you followed those groups forward and watched disease appear. Reading down the columns instead answers a different question, and in a case-control study it is the only question you can answer.
Direction, and the danger of flipping the table
Because the table looks symmetric, it is easy to compute the wrong ratio. Always be explicit about which group is the reference. If you accidentally divide the unexposed risk by the exposed risk, you get the reciprocal, turning a relative risk of 3 into 0.33 and reversing the story. A useful check is to state the claim in words before trusting the number: the exposed have how many times the risk of the unexposed? If the arithmetic and the sentence disagree, the table was read the wrong way around, and the fix is simply to swap the order of division.
Risk in each group
When you have followed a defined group forward (a cohort), you can compute the risk (incidence proportion) in each row. The risk in the exposed is a divided by (a + b). The risk in the unexposed is c divided by (c + d). These two numbers are the raw material of comparison, and every ratio and difference that follows is built from them.
Suppose a = 40, b = 160, c = 20, and d = 180. The risk in the exposed is 40 divided by 200, which equals 0.20, and the risk in the unexposed is 20 divided by 200, which equals 0.10. Already you can see the exposed group carrying twice the risk of the unexposed. The next step simply expresses that comparison as one number instead of two.
Relative risk (the risk ratio)
The relative risk (RR), or risk ratio, is the risk in the exposed divided by the risk in the unexposed:
RR = [a / (a + b)] ÷ [c / (c + d)]
Its interpretation is direct. RR = 1 means the exposed and unexposed have equal risk, so there is no association. RR > 1 means exposure is associated with more disease, a possible risk factor. RR < 1 means exposure is associated with less disease, a possible protective factor, as we hope for a vaccine. An RR of 3 means the exposed have three times the risk; an RR of 0.5 means the exposed have half the risk.
The size of the ratio conveys the strength of the association. A relative risk of 1.1 or 1.2 is weak and easily produced by undetected bias or confounding, so it demands caution. A relative risk of 5, 10, or more is strong and much harder to explain away, which is one reason the early smoking studies, with their large ratios, carried such weight. Strength alone never proves causation, but a large, repeatable ratio is a serious signal that something real is happening.
Worked example
Suppose we follow 3,000 smokers and 3,000 non-smokers for ten years for a heart-disease outcome. Among smokers, 90 develop disease; among non-smokers, 30 do. The table is:
| Disease + | Disease - | Total | |
|---|---|---|---|
| Smokers | 90 | 2,910 | 3,000 |
| Non-smokers | 30 | 2,970 | 3,000 |
Risk in the exposed (smokers) = 90 / 3,000 = 0.030, or 30 per 1,000. Risk in the unexposed = 30 / 3,000 = 0.010, or 10 per 1,000. Therefore:
RR = 0.030 / 0.010 = 3.0
Smokers in this study had three times the risk of heart disease compared with non-smokers. The relative risk answers a proportional question, how many times the risk, and it does so cleanly. What it does not tell you, by itself, is how much disease that threefold difference represents in absolute terms, a gap filled two lessons from now by the attributable-risk measures that use subtraction rather than division.
A protective exposure: relative risk below 1
Not every exposure is harmful, and the relative risk handles protection just as naturally. Imagine a cohort in which 2,000 vaccinated people yield 20 cases of an infection while 2,000 unvaccinated people yield 100 cases. The risk in the vaccinated is 20 / 2,000 = 0.010, and in the unvaccinated 100 / 2,000 = 0.050. The relative risk is 0.010 / 0.050 = 0.20, well below 1, meaning the vaccinated carried one-fifth the risk. Expressed as protection, the vaccine cut risk by 80 percent, a quantity called the relative risk reduction, which is simply 1 minus the relative risk.
The historical anchor: cohorts that built the field
The relative risk is not an abstraction; it is the number that reshaped medicine. In the British Doctors Study begun in 1951, Richard Doll and Austin Bradford Hill followed tens of thousands of physicians for decades and found that heavy smokers had many times the lung-cancer death rate of non-smokers, a large and durable relative risk. The Framingham Heart Study, launched in 1948, followed a whole town across generations and gave the world the very idea of a cardiovascular risk factor by computing relative risks for blood pressure, cholesterol, and smoking. Both were cohorts, and both earned the right to report relative risks.
Precision: the confidence interval
A relative risk computed from a sample is an estimate of the true value, and like any estimate it carries uncertainty. Researchers report a 95 percent confidence interval around it, a range that would contain the true relative risk in most repeated studies of the same kind. The key habit is to check whether that interval includes 1, the value of no association. An RR of 3.0 with an interval from 2.1 to 4.3 excludes 1 and is convincing; an RR of 3.0 with an interval from 0.7 to 12 includes 1 and could easily be chance.
Precision improves with sample size. A large study yields a narrow interval and a stable estimate; a small one yields a wide interval that may span from clearly protective to clearly harmful. This is why a single small study with a striking relative risk should be read cautiously, and why replication and pooling, taken up in the final module, matter so much. The point estimate is where to look first, but the interval tells you how much to trust it.
Risk ratio versus rate ratio
The term relative risk is used loosely for two closely related ratios. The risk ratio divides two incidence proportions, as we have done here. The rate ratio divides two incidence rates built on person-time. When the outcome is uncommon and follow-up is similar between groups, the two are nearly equal, and both are reported as relative risks. When follow-up differs greatly, the rate ratio is the sounder choice because it accounts for the person-time each group actually contributed, echoing the incidence-rate distinction from the previous lesson.
When you can and cannot compute a relative risk
You can compute a relative risk only when the design gives you true risks in each group, which means a cohort that follows exposed and unexposed people forward and counts new cases. A case-control study cannot, because the investigator fixes how many cases and controls to enroll, so any row proportion is an artifact of that sampling rather than a real risk. There the odds ratio takes over, the subject of the next lesson. A cross-sectional study yields a prevalence ratio, useful for description but subject to the survival distortions of prevalence.
From two numbers to a decision
The reason epidemiology bothers to reduce two risks to one ratio is that a single, comparable number supports decisions. A relative risk lets you rank exposures by strength, compare findings across studies that used different baseline populations, and communicate a result in a phrase a clinician or policymaker can act on. But the ratio is a beginning, not an end. A careful analysis pairs it with an absolute measure to show how much disease is at stake, a confidence interval to show how firm the estimate is, and a judgment about bias and confounding to show whether it can be believed at all.
Common misconceptions
The most important caution is that a relative risk describes strength, not amount. An RR of 10 sounds alarming, but if the unexposed risk is 1 in a million, tenfold is still only 10 in a million, a tiny absolute burden. Conversely, an RR of 1.2 on a common disease can mean many extra cases across a population. A relative risk also says nothing about the baseline risk it multiplies, so the identical ratio can be trivial in one setting and grave in another. The absolute measures two lessons ahead close this gap.
A related error is assuming a relative risk transfers unchanged to every population. Because it multiplies a baseline risk, the same ratio produces very different absolute effects in a high-risk and a low-risk group, so applying one study's relative risk to a very different population requires care. Another frequent slip is confusing the relative risk with the odds ratio when reading case-control results. The two coincide only when disease is rare, as the next lesson explains, and treating a large odds ratio on a common disease as if it were a relative risk overstates the effect.
Try it
In a cohort study, 800 exposed workers are followed and 96 develop the disease, while 1,200 unexposed workers are followed and 72 develop it. Lay out the two-by-two table with its margins, compute the risk in each group, then compute and interpret the relative risk. Work each step rather than guessing the final ratio.
Worked answer: risk in the exposed = 96 / 800 = 0.12; risk in the unexposed = 72 / 1,200 = 0.06. Relative risk = 0.12 / 0.06 = 2.0. The exposed workers had twice the risk of disease compared with the unexposed. Because this is a cohort with real risks in each row, the relative risk is a valid measure here, unlike in the case-control design coming next.
Sources
- Tenny, S., & Hoffman, M. R. (2023). Relative risk. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 3, Section 5: Measures of association. CDC Self-Study Course SS1978. archive.cdc.gov
- Doll, R., & Hill, A. B. (1954). The mortality of doctors in relation to their smoking habits: A preliminary report. British Medical Journal, 1(4877), 1451-1455. pubmed.ncbi.nlm.nih.gov
- Doll, R., Peto, R., Boreham, J., & Sutherland, I. (2004). Mortality in relation to smoking: 50 years' observations on male British doctors. BMJ, 328(7455), 1519. pubmed.ncbi.nlm.nih.gov
- Dawber, T. R., Meadors, G. F., & Moore, F. E. (1951). Epidemiological approaches to heart disease: The Framingham Study. American Journal of Public Health, 41(3), 279-286. pubmed.ncbi.nlm.nih.gov
- Mahmood, S. S., Levy, D., Vasan, R. S., & Wang, T. J. (2014). The Framingham Heart Study and the epidemiology of cardiovascular disease: A historical perspective. The Lancet, 383(9921), 999-1008. pmc.ncbi.nlm.nih.gov
- Zhang, J., & Yu, K. F. (1998). What's the relative risk? A method of correcting the odds ratio in cohort studies of common outcomes. JAMA, 280(19), 1690-1691. pubmed.ncbi.nlm.nih.gov
- Key terms
- Two-by-two table
- A table cross-classifying people by exposure (yes/no) and disease (yes/no) into cells a, b, c, and d.
- Risk in the exposed
- The incidence proportion among exposed people, a / (a + b).
- Risk in the unexposed
- The incidence proportion among unexposed people, c / (c + d).
- Relative risk (risk ratio)
- The risk in the exposed divided by the risk in the unexposed; how many times the risk exposure carries.
- RR = 1
- No association: exposed and unexposed have equal risk.
- RR below 1
- Exposure is associated with lower risk, suggesting a protective effect.
The Odds Ratio
- Define odds and compute the odds ratio from a two-by-two table.
- Explain why case-control studies use the odds ratio instead of relative risk.
- State when the odds ratio approximates the relative risk.
The relative risk is intuitive, but it requires knowing the true risk in each group, which means following a cohort forward. Many important questions, especially about rare diseases, are studied instead with the case-control design, which starts from people who already have the disease (cases) and compares them with people who do not (controls), looking backward at exposure. In that design you cannot compute a risk, because you, the investigator, decided how many cases and controls to enroll. The measure that works is the odds ratio.
This is not a minor technicality. The number of cases and controls is set by the study's budget and design, not by nature, so any proportion computed across those fixed totals is arbitrary. What the data can still reveal is whether exposure is more common among the diseased than among the healthy, and the odds ratio captures exactly that, in a form that behaves well mathematically and connects to the methods used to adjust for confounding.
Odds versus probability
The odds of an event is the probability it happens divided by the probability it does not, equivalently the ratio of the number with the event to the number without it. If 3 of 4 people are exposed, the probability of exposure is 3/4 = 0.75, while the odds of exposure is 3 to 1, or 3. Odds and probability carry the same information in different form, and odds have a mathematical convenience that makes them natural for case-control data and for regression.
Converting between them is worth practicing. From a probability p, the odds are p / (1 - p): a probability of 0.20 becomes odds of 0.20 / 0.80 = 0.25, or 1 to 4. From odds o, the probability is o / (1 + o): odds of 3 become a probability of 3 / 4 = 0.75. For small probabilities the two are close, which will matter shortly; for large probabilities they diverge sharply, which is the root of the odds ratio's later distortion.
The word odds comes from gambling, and the betting sense is exactly right. Odds of 3 to 1 on exposure mean three exposed people for every one unexposed, the same information as a 75 percent probability of exposure. Epidemiologists borrow the concept not for its glamour but because, when a study fixes the number of cases and controls, odds of exposure remain interpretable while probabilities of disease do not. Holding that fact in mind keeps the odds ratio from feeling arbitrary and explains why it, rather than risk, is the measure a case-control study can honestly report.
The odds ratio
The odds ratio (OR) compares the odds of exposure in cases with the odds of exposure in controls. Remarkably, using the a, b, c, d cells this reduces to the simple cross-product:
OR = (a × d) ÷ (b × c)
Its interpretation mirrors the relative risk: OR = 1 means no association, OR > 1 means exposure is more common in cases (a risk factor), and OR < 1 means exposure is less common in cases (protective). The direction and rough magnitude read just like a risk ratio, which is why the measures are so often discussed together.
Where does the cross-product come from? In a case-control study you sample cases and controls, then ask about exposure. The odds of exposure among cases is exposed cases over unexposed cases, a / c. The odds of exposure among controls is b / d. The odds ratio is the first divided by the second, (a/c) divided by (b/d), which rearranges algebraically to (a times d) divided by (b times c). The same four cells that would give a relative risk in a cohort give an odds ratio here, read down the columns rather than across the rows.
The symmetry that makes case-control work
A remarkable algebraic fact underlies the whole method: the odds ratio of exposure given disease equals the odds ratio of disease given exposure. Both equal ad divided by bc. This symmetry is why a case-control study, which samples on disease status and looks back at exposure, still recovers the exposure-disease association we ultimately care about. It is also why the odds ratio, alone among the common measures, can be estimated from a design in which the investigator sets the number of cases and controls and true risks are unknowable.
Why bother with odds at all?
Newcomers reasonably ask why epidemiology does not simply use probabilities everywhere. The answer is partly practical and partly mathematical. Practically, the odds ratio is the only valid ratio measure a case-control study can produce, and case-control studies are indispensable for rare diseases. Mathematically, odds range from zero to infinity and their logarithm is symmetric around no effect, which makes odds the natural currency of logistic regression and of pooling results across studies. Those properties are why the odds ratio, despite being less intuitive than risk, appears throughout the literature.
Worked example
Investigators enroll 200 lung-cancer cases and 200 healthy controls and ask about past exposure to a workplace chemical. Among the 200 cases, 120 were exposed and 80 were not. Among the 200 controls, 60 were exposed and 140 were not. The table is:
| Cases (Disease +) | Controls (Disease -) | |
|---|---|---|
| Exposed + | a = 120 | b = 60 |
| Exposed - | c = 80 | d = 140 |
Applying the cross-product:
OR = (120 × 140) ÷ (60 × 80) = 16,800 ÷ 4,800 = 3.5
The odds of the chemical exposure were 3.5 times higher in cases than in controls, evidence that the exposure is associated with the cancer. Notice we never computed a risk, only odds of exposure, which is exactly why the odds ratio is the correct measure when the sampling is done by disease status. Had we tried to read a risk across the rows, the numbers would have reflected only how many controls the investigators chose to recruit.
A protective odds ratio
The odds ratio describes protection as readily as harm. Suppose a case-control study of an infection enrolls 150 cases and 150 controls. Among cases, 30 had been vaccinated and 120 had not; among controls, 90 had been vaccinated and 60 had not. The odds ratio is (30 × 60) divided by (120 × 90) = 1,800 divided by 10,800 = 0.17. An odds ratio well below 1 marks the vaccine as protective, and if the infection is uncommon it also approximates a relative risk near 0.17, implying roughly an 83 percent reduction in risk among the vaccinated.
When the odds ratio approximates the relative risk
A crucial fact makes case-control studies so valuable: when the disease is rare, roughly under 10 percent in the population, the odds ratio closely approximates the relative risk. This is the rare disease assumption. Consider a population table with a = 10, b = 990, c = 5, d = 995. The relative risk is [10/1000] divided by [5/1000] = 2.0, and the odds ratio is (10 × 995) divided by (990 × 5) = 9,950 divided by 4,950 = 2.01, essentially identical.
Now make the disease common. Take a = 200, b = 800, c = 100, d = 900. The relative risk is (200/1000) divided by (100/1000) = 2.0, but the odds ratio is (200 × 900) divided by (800 × 100) = 180,000 divided by 80,000 = 2.25. The odds ratio has drifted above the relative risk. Push the disease commoner still and the gap widens, always away from 1, so a common-disease odds ratio exaggerates the apparent strength if it is mistaken for a risk ratio.
Choosing controls well
The validity of any odds ratio rests on the controls. Controls should come from the same source population that produced the cases, so that their exposure history reflects what the cases' exposure would have looked like had they stayed disease-free. Drawing controls from a convenient but different group, such as other hospital patients whose own illnesses relate to the exposure, distorts b and d and therefore the cross-product. Much of the craft of case-control research lies in this single choice, and a beautifully computed odds ratio built on the wrong controls is simply wrong.
Matched designs and the matched odds ratio
Investigators often match each case to one or more controls on variables like age and sex to remove those as confounders by design. Matching changes the analysis: instead of the simple cross-product, a matched study focuses on discordant pairs, those in which the case and control differ in exposure, and estimates the odds ratio from the ratio of the two kinds of discordant pair. The lesson to carry forward is that the design and the analysis must agree, a theme that returns when confounding is treated formally in a later lesson.
The odds ratio in regression
The odds ratio has a second home that guarantees its ubiquity: logistic regression, the workhorse method for adjusting an association for many confounders at once, reports its results as adjusted odds ratios. So even in cohort and cross-sectional analyses, where a risk ratio could in principle be computed, findings are often presented as odds ratios because of the model used. Reading the literature therefore demands comfort with the odds ratio regardless of design, and the same rare-disease caution applies whenever it is interpreted as though it were a relative risk.
Precision and significance
Like the relative risk, an odds ratio from a sample is an estimate reported with a 95 percent confidence interval, and the same test applies: an interval that includes 1 is compatible with no association. Odds ratios from small case-control studies can have very wide intervals, because a handful of exposed controls can swing the cross-product substantially. A large point estimate built on thin data therefore warrants caution until it is replicated in a larger or independent study.
Reading an odds ratio in a paper
When a study reports an adjusted odds ratio of, say, 1.8 with a 95 percent confidence interval of 1.3 to 2.5, read it in three moves. The point estimate, 1.8, gives strength and direction: exposure is associated with higher odds. The interval, excluding 1, says the finding is unlikely to be pure chance. The word adjusted signals that the authors used regression to hold measured confounders constant. What the number still cannot settle on its own is causation, which requires the weighing of evidence taken up in the module on validity and causal inference.
Common misconceptions
The cardinal error is announcing an odds ratio as though it were a relative risk. Only when the disease is rare do the two nearly coincide; for a common outcome the odds ratio is larger, and saying the exposed have "so many times the risk" overstates the effect. A second error is forgetting that the odds ratio, like every measure in this module, reflects association only. A valid, precise odds ratio can still arise from bias or confounding, which the later modules examine in detail before any causal claim is warranted.
Try it
A case-control study of a rare disease enrolls 100 cases and 100 controls. Among the cases, 80 were exposed and 20 were not; among the controls, 40 were exposed and 60 were not. Lay out the two-by-two table, compute the odds ratio by the cross-product, and state whether it also approximates the relative risk in this instance.
Worked answer: a = 80, b = 40, c = 20, d = 60. OR = (a × d) divided by (b × c) = (80 × 60) divided by (40 × 20) = 4,800 divided by 800 = 6.0. The odds of exposure were six times higher in cases than controls. Because the disease is rare, this odds ratio also approximates the relative risk, so exposed people carried roughly six times the risk of the unexposed.
Sources
- Tenny, S., & Hoffman, M. R. (2026). Odds ratio. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Tenny, S., Kerndt, C. C., & Hoffman, M. R. (2023). Case control studies. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Bland, J. M., & Altman, D. G. (2000). Statistics notes: The odds ratio. BMJ, 320(7247), 1468. pmc.ncbi.nlm.nih.gov
- Szumilas, M. (2010). Explaining odds ratios. Journal of the Canadian Academy of Child and Adolescent Psychiatry, 19(3), 227-229. pmc.ncbi.nlm.nih.gov
- Zhang, J., & Yu, K. F. (1998). What's the relative risk? A method of correcting the odds ratio in cohort studies of common outcomes. JAMA, 280(19), 1690-1691. pubmed.ncbi.nlm.nih.gov
- Doll, R., & Hill, A. B. (1950). Smoking and carcinoma of the lung: Preliminary report. British Medical Journal, 2(4682), 739-748. pubmed.ncbi.nlm.nih.gov
- Berkson, J. (1946). Limitations of the application of fourfold table analysis to hospital data. Biometrics Bulletin, 2(3), 47-53. pubmed.ncbi.nlm.nih.gov
- Key terms
- Case-control study
- A design that starts from cases and controls and looks backward at exposure; yields an odds ratio.
- Odds
- The probability an event occurs divided by the probability it does not; the ratio with-event to without-event.
- Odds ratio
- The ratio of the odds of exposure in cases to that in controls, computed as the cross-product ad / bc.
- Cross-product
- The shortcut (a times d) divided by (b times c) that yields the odds ratio from a two-by-two table.
- Rare disease assumption
- When disease is uncommon, the odds ratio closely approximates the relative risk.
- OR = 1
- No association between exposure and disease; the odds of exposure are equal in cases and controls.
Attributable Risk and the Impact of Exposure
- Compute the attributable risk (risk difference) and interpret it.
- Calculate the attributable risk percent and distinguish it from relative risk.
- Explain the population attributable risk and its use in public health priorities.
Relative risk tells you how strongly an exposure and a disease are linked. But public health also needs to know how much disease an exposure actually causes, the absolute impact. That is the job of the attributable risk family of measures, which use subtraction rather than division. Where a ratio asks "how many times the risk," a difference asks "how many extra cases," and prevention is planned in cases, not in ratios.
The distinction is not academic. A ministry of health deciding where to spend a fixed budget cares about the number of illnesses it can avert, which is an absolute quantity. A ratio can be enormous while the number of cases behind it is tiny, or modest while the number is large. The attributable-risk measures translate association into the currency of prevention, and learning to compute them is what turns a strength of association into a case for action.
Attributable risk (the risk difference)
The attributable risk (AR), also called the risk difference or excess risk, is the risk in the exposed minus the risk in the unexposed:
AR = [a / (a + b)] - [c / (c + d)]
It estimates the amount of disease among the exposed that is attributable to the exposure, assuming the association is causal. Using our smoking cohort, with risk in smokers 0.030 and in non-smokers 0.010, the attributable risk is 0.030 - 0.010 = 0.020, or 20 excess cases per 1,000 smokers. In words: of every 1,000 smokers, about 20 heart-disease cases would not have occurred without smoking. This absolute number is what shapes prevention, because it states how many cases could actually be prevented by removing the exposure.
The attributable risk is measured in the same units as risk itself, a proportion or a rate, and it can be scaled to any population size. Twenty per 1,000 is 200 per 10,000 and 2,000 per 100,000, the same fact expressed at different scales. Whenever an attributable risk is quoted, check that a time frame and a denominator travel with it, exactly as they must for any incidence measure.
Number needed to harm and number needed to treat
The reciprocal of the attributable risk has a vivid interpretation. The number needed to harm (NNH) is 1 divided by the attributable risk: for our smokers, 1 / 0.020 = 50. On average, 50 people must carry the exposure for one extra case of disease to appear. The mirror image for a beneficial intervention is the number needed to treat (NNT), 1 divided by the absolute risk reduction, the number of patients who must be treated for one to benefit.
Clinicians favor these numbers because they are concrete. A drug that cuts risk from 0.050 to 0.030 has an absolute risk reduction of 0.020 and an NNT of 50, meaning 50 patients treated to prevent one event. A rival drug boasting a large relative risk reduction on a very rare outcome may have an NNT in the thousands, which reframes an impressive-sounding ratio as a modest practical gain. Absolute measures, once again, keep the story honest.
Relative and absolute measures can tell different stories
A large relative risk on a rare outcome can mean few actual cases, while a modest relative risk on a common outcome can mean many. Suppose exposure A triples a rare risk from 1 to 3 per 100,000 (RR = 3, but AR = only 2 per 100,000), while exposure B raises a common risk from 100 to 150 per 100,000 (RR = 1.5, but AR = 50 per 100,000). Exposure B, with the smaller relative risk, causes twenty-five times more excess disease.
This is why reporting a relative risk alone can badly mislead, and why absolute measures belong beside it. A headline that an exposure "doubles your risk" is meaningless until you know the baseline it doubles. Doubling a one-in-a-million risk is trivial; doubling a one-in-ten risk is grave. The mature reading of any study holds the ratio and the difference in view together, letting strength and impact inform each other.
Attributable risk percent
The attributable risk percent (AR%) expresses the excess as a fraction of the total risk in the exposed:
AR% = (Risk in exposed - Risk in unexposed) ÷ Risk in exposed × 100
For the smokers: (0.030 - 0.010) / 0.030 × 100 = 66.7 percent. Among exposed people who got the disease, about two-thirds of their risk is attributable to the exposure. Equivalently, using the relative risk, AR% = (RR - 1) / RR × 100 = (3 - 1) / 3 × 100 = 66.7 percent, the same answer by a different route. The two formulas are algebraically identical, so you can use whichever inputs a study happens to report.
Read carefully, the attributable risk percent answers a specific question: of the disease that occurs among the exposed, what share would vanish if the exposure were removed? It is sometimes called the attributable fraction in the exposed. It says nothing about the unexposed, and nothing yet about the whole population, which is where the exposure's prevalence finally enters.
Population attributable risk
Public health cares not only about the exposed but about the whole population, where impact depends on how common the exposure is. The population attributable risk (PAR) is the excess disease rate in the total population attributable to the exposure, and the population attributable risk percent (PAR%) is the fraction of all disease in the population that would be eliminated if the exposure were removed. A moderately harmful exposure that is very widespread can carry a larger population attributable risk than a highly dangerous but rare one.
Work it through. Keep risk in the exposed at 0.030 and in the unexposed at 0.010, and suppose 40 percent of the population is exposed. The total population risk is 0.40 × 0.030 + 0.60 × 0.010 = 0.012 + 0.006 = 0.018. The population attributable risk is 0.018 - 0.010 = 0.008, meaning 8 excess cases per 1,000 in the whole population. The PAR% is 0.008 / 0.018 × 100 = 44.4 percent, so removing the exposure would prevent about 44 percent of all cases.
A compact formula gives the same result from the relative risk and the exposure prevalence Pe: PAR% = Pe(RR - 1) / [Pe(RR - 1) + 1] × 100. Here that is 0.40 × 2 / [0.40 × 2 + 1] × 100 = 0.8 / 1.8 × 100 = 44.4 percent, matching the direct calculation. Notice how the answer depends jointly on strength (RR) and reach (Pe): weaken either and the population impact shrinks.
A worked population example
Consider a stylized illustration with round numbers. In a population of 1,000,000, suppose 300,000 are exposed with a disease risk of 0.020 and 700,000 are unexposed with a risk of 0.005. The exposed produce 6,000 cases and the unexposed 3,500, for 9,500 cases in all. Had the exposed instead faced only the unexposed risk of 0.005, they would have produced 1,500 cases, so 4,500 of the 9,500 total cases, about 47 percent, are attributable to the exposure across the population. That is the population attributable risk percent computed straight from case counts.
From individual to population attribution
It helps to keep three related quantities distinct. The attributable risk percent describes the exposed only. The population attributable risk percent describes everyone, and it can never exceed the attributable risk percent, because diluting the exposure with unexposed people can only lower the share of total disease it explains. The gap between the two is governed entirely by how common the exposure is: as exposure prevalence approaches 100 percent the population figure rises toward the exposed figure, and as it approaches zero the population figure vanishes.
Setting public health priorities
This is precisely the calculation that guides where public health should spend its effort. It weighs not only the most dangerous exposures but those whose combination of risk and prevalence produces the most preventable disease. A rare occupational toxin with a tenfold relative risk may harm few people if almost no one is exposed, while a common exposure like physical inactivity, with a modest relative risk, may account for a large slice of chronic disease simply because so many people share it.
The practical upshot is that a population strategy aimed at a widespread modest risk can prevent more disease than an intense effort aimed at a rare extreme risk, echoing the prevention paradox from the opening lesson. Population attributable risk turns that intuition into a number, letting competing programs be compared on a single scale: how many cases would each actually prevent across the whole community it serves.
Attribution in historical perspective
The attributable-risk framework did real historical work. Once cohort studies had established that smoking causes lung cancer, epidemiologists could estimate what fraction of lung-cancer cases in a population was attributable to smoking, a number that proved very large and made the public health case unavoidable. The same logic later quantified the share of cardiovascular disease attributable to high blood pressure and to elevated cholesterol, drawing on decades of follow-up from studies such as Framingham. Attribution translated causal knowledge into the scale of preventable loss, which is what moved policy.
Attributable and prevented fractions
The same subtraction logic runs in reverse for protective exposures. When an exposure lowers risk, as a vaccine does, the prevented fraction measures the share of potential disease that the exposure actually averts among those who received it, and the population prevented fraction extends that to the whole population given how many are protected. These mirror the attributable-risk and population-attributable-risk measures for harmful exposures, and they are the natural language for reporting the population benefit of a preventive program.
The causal assumption behind attribution
Every attributable-risk measure carries a warning label: the word attributable assumes the association is causal. If the link between exposure and disease is really due to confounding or bias, then removing the exposure will not prevent the cases the calculation promises. The arithmetic is always valid as a description of the data, but its prevention interpretation stands or falls on whether the exposure truly causes the disease. That is why these measures are most trustworthy for exposures whose causal status is already well established.
Common misconceptions
Three errors recur. The first is confusing the attributable risk (an absolute difference) with the attributable risk percent (a fraction); they answer different questions and carry different units. The second is inferring impact from a ratio alone, forgetting that a large relative risk can sit atop a tiny baseline. The third is adding population attributable percentages across several correlated risk factors and expecting them to sum to 100; because factors overlap and interact, attributable fractions can total well over 100 percent and are not simply additive.
Try it
In a cohort, the risk of disease is 0.08 in the exposed and 0.02 in the unexposed, and 25 percent of the population is exposed. Compute the relative risk, the attributable risk, the attributable risk percent, the number needed to harm, and the population attributable risk percent. Show each step.
Worked answer: RR = 0.08 / 0.02 = 4.0. AR = 0.08 - 0.02 = 0.06, or 60 excess cases per 1,000 exposed. AR% = (0.08 - 0.02) / 0.08 × 100 = 75 percent. NNH = 1 / 0.06 ≈ 17, so about 17 exposed people per extra case. PAR% = Pe(RR - 1) / [Pe(RR - 1) + 1] × 100 = 0.25 × 3 / [0.25 × 3 + 1] × 100 = 0.75 / 1.75 × 100 ≈ 42.9 percent.
Sources
- Rockhill, B., Newman, B., & Weinberg, C. (1998). Use and misuse of population attributable fractions. American Journal of Public Health, 88(1), 15-19. pmc.ncbi.nlm.nih.gov
- Cook, R. J., & Sackett, D. L. (1995). The number needed to treat: A clinically useful measure of treatment effect. BMJ, 310(6977), 452-454. pubmed.ncbi.nlm.nih.gov
- Laupacis, A., Sackett, D. L., & Roberts, R. S. (1988). An assessment of clinically useful measures of the consequences of treatment. New England Journal of Medicine, 318(26), 1728-1733. pubmed.ncbi.nlm.nih.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 3, Section 5: Measures of association. CDC Self-Study Course SS1978. archive.cdc.gov
- Tenny, S., & Hoffman, M. R. (2023). Relative risk. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Rose, G. (1985). Sick individuals and sick populations. International Journal of Epidemiology, 14(1), 32-38. pubmed.ncbi.nlm.nih.gov
- Doll, R., Peto, R., Boreham, J., & Sutherland, I. (2004). Mortality in relation to smoking: 50 years' observations on male British doctors. BMJ, 328(7455), 1519. pubmed.ncbi.nlm.nih.gov
- Key terms
- Attributable risk (risk difference)
- The risk in the exposed minus the risk in the unexposed; the excess risk from exposure.
- Excess risk
- Another name for the attributable risk; the absolute additional risk carried by the exposed.
- Attributable risk percent
- The excess risk as a proportion of the total risk in the exposed, equal to (RR - 1) / RR times 100.
- Population attributable risk percent
- The fraction of disease in the whole population removable by eliminating the exposure; depends on exposure prevalence.
- Relative vs absolute measures
- Ratios show strength of association; differences show the actual amount of disease caused.
- Preventable fraction
- The share of cases that could be avoided by removing a causal exposure, captured by attributable-risk measures.
Module 3: Study Designs
The main observational and experimental designs of epidemiology - cross-sectional, case-control, cohort, and the randomized controlled trial - and how to match a design to a question.
Descriptive and Cross-Sectional Studies
- Describe case reports, case series, and ecological studies and their limits.
- Explain the cross-sectional design and what it can and cannot show.
- Define the ecological fallacy.
Study designs form a ladder from simply describing what is seen to rigorously testing causes. The lower rungs generate hypotheses cheaply; the higher rungs test them convincingly. Choosing wisely means matching the design to the question, the disease, and the resources at hand. This lesson covers the lower rungs, the descriptive and cross-sectional designs, which are indispensable for raising questions even though they rarely settle them.
None of these designs is a poor relation to be dismissed. Every serious investigation begins descriptively, and much of routine public health, from planning services to spotting the first hint of a new problem, lives on these rungs. Their limits become dangerous only when their results are read as if they had come from a stronger design, so the aim here is to learn both what each can do and what it cannot.
The hierarchy of study designs
It helps to place the designs on a map before examining each. At the base are descriptive studies without a comparison group: case reports and case series. Next come studies that compare, whether groups (ecological) or individuals at one time (cross-sectional). Above them sit the analytic observational designs of the next lesson, case-control and cohort, which compare exposed and unexposed experience over time. At the top is the experiment, the randomized trial. Reliability rises as you climb, and so, usually, does cost and effort.
Two axes organize the map. One is whether the investigator intervenes: observational designs watch what people do, while experimental designs assign the exposure. The other is the direction and timing of observation: descriptive studies simply portray, cross-sectional studies look at one instant, case-control studies look backward from disease, and cohort studies look forward from exposure. Almost every design you will meet is a combination of positions on these two axes, and naming those positions is often enough to predict a design's strengths and weaknesses.
Descriptive epidemiology: person, place, and time
Underneath these designs lies the descriptive core of the field: characterizing health by person (age, sex, occupation, and other traits), place (where cases cluster geographically), and time (trends, seasons, and sudden changes). Organizing observations along these three axes is what turns scattered reports into a pattern a hypothesis can grab. Snow's map was descriptive epidemiology in the dimension of place; a seasonal peak in influenza is descriptive epidemiology in time. Every analytic study begins by describing its cases this way before it attempts to explain them.
Descriptive designs
A case report describes a single striking patient; a case series describes several. They cannot establish cause because there is no comparison group, yet they are often the first alarm. In 1981, clusters of Pneumocystis pneumonia and Kaposi sarcoma in previously healthy young men, published as case series, were the first signal of what became the AIDS epidemic. A single alert clinician noticing something out of the ordinary has repeatedly launched major discoveries, which is why case reports retain a place in the literature.
The strength of a case series is sensitivity to the unexpected; its weakness is that it cannot tell whether the pattern is real or coincidental. With no denominator and no comparison, a series of patients who share an exposure might reflect a true hazard or merely the popularity of that exposure. The series raises the question; a designed study must answer it. Treating a vivid case series as proof of cause is one of the most common and costly errors in reading medical evidence.
An ecological study climbs one rung by comparing whole groups rather than individuals, correlating, for example, average fat intake against heart-disease rates across countries. Such data are easy to obtain because they use figures already collected for populations, and they can reveal broad patterns worth investigating. They also carry a specific and famous danger that gives the next section its name.
The ecological fallacy
The ecological fallacy is the error of drawing conclusions about individuals from group-level data. If countries with higher average salt intake have higher stroke rates, it does not follow that the high-salt individuals within each country are the ones having strokes. The association at the group level may not hold at the individual level, because a group average hides all the variation inside the group. Ecological studies suggest hypotheses; they do not confirm individual-level causes.
Why does the fallacy arise? A national average blends people with very different exposures and risks, and a correlation between two averages can be produced by a third factor that varies across the groups. Countries differ in wealth, diet, health care, and much else at once, so a cross-country correlation is heavily confounded by everything that travels with national development. The group-level pattern is real as a pattern; the error is only in transferring it to individuals.
The mirror-image error, sometimes called the atomistic fallacy, is assuming that a relationship found in individuals must hold at the group level. Cross-level inference is hazardous in both directions, which is why epidemiologists prefer to study exposure and disease in the same units, ideally the individual, whenever they can. Ecological data still earn their place for generating hypotheses cheaply and for exposures that are genuinely collective, such as a law, a tax, or an air-pollution level that everyone in a region shares.
More first alarms from case reports
History is full of discoveries that began as description. An astute pharmacologist linking a cluster of newborn limb malformations to a sedative taken in pregnancy, reports connecting a rare vaginal cancer in young women to a hormone their mothers had taken, and the recognition of Legionnaires disease after an unexplained pneumonia outbreak among conventioneers in 1976 all started from careful description of unusual cases. In each, the case series or outbreak report raised the alarm, and analytic studies then confirmed the cause. Description opens the door; it does not close the case.
Cross-sectional studies
A cross-sectional study measures exposure and disease at the same point in time in a sample of individuals, a snapshot, like a single survey. It is the natural design for measuring prevalence, as in "what fraction of adults have hypertension right now?" and for describing the health of a population. Because it captures everyone at once, it is fast and relatively cheap, and it can examine many exposures and many outcomes together in a single pass.
Large national health surveys are cross-sectional in spirit: they interview and examine a representative sample and report how common conditions and behaviors are. Because they measure so much at once, they are efficient engines of description and hypothesis generation. Repeating such a survey at intervals, a design called serial or repeated cross-sectional, tracks how prevalence changes over time even though each round samples different individuals, which is how populations monitor trends in conditions such as obesity or hypertension.
What a cross-sectional study can measure
Beyond a simple prevalence, a cross-sectional study can compare prevalence between exposed and unexposed people, yielding a prevalence ratio or a prevalence odds ratio. These describe association at a moment and can be genuinely useful, but they inherit the weaknesses of prevalence. Because they count existing cases, they blend the rate of new disease with how long disease lasts, so an apparent association can reflect differences in survival rather than in the risk of becoming ill.
A worked cross-sectional calculation
Put numbers to it. A survey of 1,000 adults classifies each by a behavior and by whether they currently have a condition. Among the 400 with the behavior, 80 have the condition; among the 600 without it, 60 do. The prevalence in the exposed is 80 / 400 = 0.20, and in the unexposed 60 / 600 = 0.10. The prevalence ratio is 0.20 / 0.10 = 2.0, so the condition is twice as common among those with the behavior. The overall prevalence is 140 / 1,000 = 0.14, or 14 percent.
The prevalence ratio of 2.0 looks like a relative risk, and it is computed the same way, but it must be read differently. It compares existing cases at a moment, not new cases over time, so it cannot distinguish whether the behavior raises the chance of developing the condition or merely lengthens how long people live with it. That single caution separates a prevalence ratio from a true risk ratio, and it is why the identical arithmetic supports a weaker conclusion here than it would in a cohort.
Temporal ambiguity and its cousins
The great limitation of the cross-sectional design is temporal ambiguity: because exposure and disease are measured simultaneously, you usually cannot tell which came first. If a survey finds that people who exercise less are more depressed, does inactivity cause depression, or does depression reduce activity? The snapshot cannot say. This threat of reverse causation is why a cross-sectional association, however strong, cannot on its own support a causal claim.
Cross-sectional studies also tend to over-represent long-lasting cases and miss those who already recovered or died, a distortion called prevalence-incidence bias or Neyman bias. A disease that kills quickly or resolves fast is under-counted in any snapshot, so the survivors who remain to be surveyed may differ systematically from all who ever had the disease. For these reasons cross-sectional studies describe burden well but test causes poorly.
| Design | Comparison group? | Establishes time order? | Main use |
|---|---|---|---|
| Case report / series | No | No | First description of rare events |
| Ecological | Groups | No | Hypothesis generation (risk of ecological fallacy) |
| Cross-sectional | Individuals, one time point | No | Prevalence and population description |
When a cross-sectional design is the right tool
Chosen for the right job, the cross-sectional study is powerful. It is the correct design to estimate how common a condition is, to plan services and allocate resources, to monitor a population's burden over repeated rounds, and to generate hypotheses that stronger designs will test. For exposures that do not change, such as blood type or genotype, temporal ambiguity even eases, since the exposure clearly precedes the outcome. The design's reputation for weakness comes almost entirely from using it for the wrong question, causation.
Common misconceptions
The recurring error is reading a cross-sectional association as evidence of cause, forgetting that the snapshot cannot fix time order. A second is treating an ecological correlation as if it described individuals, the ecological fallacy. A third is mistaking a high prevalence for a high risk of getting the disease, when long duration alone can inflate prevalence. Each error has the same root: applying a descriptive design to a question that only an analytic or experimental design can answer.
Try it
A single survey of 2,000 office workers finds that those who report high job stress have a higher prevalence of back pain than those reporting low stress. A colleague concludes that job stress causes back pain. Identify the design, state the two-part reason this conclusion is premature, and name one design that could test the hypothesis properly.
Worked answer: this is a cross-sectional study. The conclusion is premature first because of temporal ambiguity, since stress and back pain were measured together and either could have come first, and second because prevalence-incidence bias may over-represent people whose back pain is chronic. A prospective cohort, measuring stress in pain-free workers and following them for new back pain, would establish time order and test the hypothesis properly.
Sources
- Wang, X., & Cheng, Z. (2020). Cross-sectional studies: Strengths, weaknesses, and recommendations. Chest, 158(1S), S65-S71. pubmed.ncbi.nlm.nih.gov
- Sedgwick, P. (2015). Understanding the ecological fallacy. BMJ, 351, h4773. pubmed.ncbi.nlm.nih.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 1, Section 6: Descriptive epidemiology. CDC Self-Study Course SS1978. archive.cdc.gov
- Centers for Disease Control and Prevention. (1981). Pneumocystis pneumonia - Los Angeles. Morbidity and Mortality Weekly Report, 30(21), 250-252. pubmed.ncbi.nlm.nih.gov
- Fraser, D. W., Tsai, T. R., Orenstein, W., Parkin, W. E., Beecham, H. J., Sharrar, R. G., ... Broome, C. V. (1977). Legionnaires' disease: Description of an epidemic of pneumonia. New England Journal of Medicine, 297(22), 1189-1197. pubmed.ncbi.nlm.nih.gov
- Herbst, A. L., Ulfelder, H., & Poskanzer, D. C. (1971). Adenocarcinoma of the vagina: Association of maternal stilbestrol therapy with tumor appearance in young women. New England Journal of Medicine, 284(15), 878-881. pubmed.ncbi.nlm.nih.gov
- Perez-Guerrero, E. E., Guillen-Medina, M. R., Marquez-Sandoval, F., Vera-Cruz, J. M., Gallegos-Arreola, M. P., Rico-Mendez, M. A., ... Gutierrez-Hurtado, I. A. (2024). Methodological and statistical considerations for cross-sectional, case-control, and cohort studies. Journal of Clinical Medicine, 13(14), 4005. pmc.ncbi.nlm.nih.gov
- Key terms
- Case report / series
- A description of one or several patients with no comparison group; useful for first alarms, not causation.
- Ecological study
- A study comparing groups rather than individuals, correlating group-level exposure and disease.
- Ecological fallacy
- The error of inferring individual-level relationships from group-level data.
- Cross-sectional study
- A snapshot measuring exposure and disease at the same time; ideal for prevalence.
- Temporal ambiguity
- The inability of a snapshot to show whether exposure preceded disease or the reverse.
- Prevalence-incidence bias
- The tendency of cross-sectional data to over-represent long-lasting cases and miss brief or fatal ones.
Cohort and Case-Control Studies
- Contrast the direction of inquiry in cohort and case-control studies.
- Match each design to the diseases and exposures it suits, and to its measure of association.
- Explain prospective versus retrospective cohorts and the key biases of each design.
The two workhorse observational analytic designs are the cohort study and the case-control study. Both compare exposed and unexposed experience, but they start from opposite ends and suit opposite situations. Understanding their logic is central to reading almost any epidemiologic paper, because most of what we know about the causes of chronic disease was learned from one or the other.
The key to telling them apart is to ask what the investigator knew first. In a cohort, exposure is known and disease is awaited. In a case-control study, disease is known and exposure is sought. Everything else, the measure of association, the strengths, the biases, follows from that single difference in starting point.
Cohort studies: exposure to outcome
A cohort study classifies people by exposure first and then follows them forward in time to see who develops the disease. Because it observes incidence directly, it yields risks and therefore the relative risk and the attributable risk introduced in earlier lessons. This is the design that measures how often disease actually appears, which is why it, and not the case-control study, can produce a true risk.
A cohort can be prospective, assembling the group now and following it into the future, or retrospective, using existing records to reconstruct a cohort defined by past exposure and following it forward from then to now. A retrospective cohort still moves from exposure to outcome; it simply uses history to compress the waiting. The distinction is only about when the follow-up happens relative to the study, not about the logical direction, which is always forward from exposure.
Cohorts excel for a rare exposure, because you can deliberately enroll exposed people, such as workers at a particular plant or survivors of a specific event, and compare them with an unexposed group. They also shine for studying multiple outcomes of a single exposure, since once a cohort is assembled you can watch for every disease that follows. A study of an unusual occupational exposure can, in one design, examine cancers, lung disease, and mortality all at once.
Their weaknesses are cost, long duration, and loss to follow-up. Following thousands of people for years is expensive and slow, and participants drop out, move, or die of other causes. If that dropout is related to both the exposure and the outcome, the remaining sample is skewed and the results distort. A cohort is also inefficient for a rare disease, because even a large group followed for years may yield only a handful of cases, wasting most of the effort on people who never fall ill.
Cohorts and person-time
A cohort's forward design makes it the natural home of person-time. When members enter at different dates, drop out, or die, each contributes only the time actually observed, and the incidence rate divides new cases by that summed person-time, as the frequency lesson showed. Dynamic cohorts, whose membership changes as people join and leave a population, rely entirely on person-time denominators. This is why the rate ratio, not just the risk ratio, is a natural product of cohort analysis, and why cohorts handle staggered, incomplete follow-up gracefully rather than discarding it.
Prospective versus retrospective in practice
The two flavors of cohort trade speed against data quality. A prospective cohort lets investigators decide in advance exactly what to measure and how, yielding clean, purpose-built exposure data, but it must wait years for outcomes to accrue. A retrospective cohort reads exposure from records that already exist, delivering results quickly, but it is at the mercy of whatever those records happened to capture, which may be incomplete or measured inconsistently. Choosing between them balances the value of tailored data against the cost of waiting.
Case-control studies: outcome to exposure
A case-control study starts from the disease first, assembling cases (people with the disease) and controls (comparable people without it), then looks backward to compare past exposure. Because sampling is by disease status, it yields the odds ratio, not the relative risk, as the odds-ratio lesson explained in detail. You choose how many cases and controls to enroll, so no true risk can be read from the table.
Case-control studies shine for a rare disease, because you start with the scarce cases instead of waiting years for them to appear, and for diseases with a long latency between exposure and onset. They are fast and inexpensive relative to cohorts, and a single study can examine many exposures at once, asking cases and controls about a whole history of possible causes. For a newly recognized rare cancer with a dozen suspected triggers, this efficiency is decisive.
Their central challenge is bias. Recall bias arises when cases, motivated to explain their illness, remember past exposures more keenly than controls. Selection bias creeps in when controls are chosen poorly. Choosing controls that come from the same population that produced the cases is the hardest part of a good case-control study, and getting it wrong can manufacture an association or hide a real one. The design also cannot, by itself, prove that exposure preceded disease if the exposure history is uncertain.
The source population and valid comparison
Both designs live or die by comparability. In a cohort, the unexposed must resemble the exposed in everything except the exposure, so that a difference in outcome can be pinned on the exposure rather than on who chose to be exposed. In a case-control study, the controls must be drawn from the same source population that produced the cases, representing the exposure distribution the cases would have shown had they stayed well. When either comparison group is chosen carelessly, the resulting measure is precise but wrong, a theme the module on bias develops.
The direction of inquiry
Picture the two designs as arrows on a timeline. The cohort arrow points forward: start with exposure status, wait, and count disease. The case-control arrow points backward: start with disease status, look back, and reconstruct exposure. This mirror image explains almost every practical difference between them, from which measure they yield to which biases threaten them. Fixing the arrow in your mind is the fastest way to reason about an unfamiliar study.
Hybrid designs: nested case-control and case-cohort
The nested case-control study combines the strengths of both by drawing cases and controls from within an existing cohort. It keeps the efficiency of case-control sampling, since you analyze only the cases and a sample of controls, while retaining the cohort's high-quality exposure data, often collected and stored before any disease appeared. Because that exposure information predates the outcome, it sidesteps recall bias entirely, one of the case-control design's worst hazards.
A close relative, the case-cohort study, compares all cases with a random subcohort sampled from the full cohort at the outset, which allows one control group to serve studies of several different outcomes. Both hybrids are ways of spending a cohort's expensive, carefully collected data more efficiently, and both are common in modern research built on large stored biobanks and registries.
Which measure each design yields
Tie the designs back to the measures. A cohort observes incidence, so it produces risks, the relative risk, the attributable risk, and, with person-time, incidence rates and rate ratios. A case-control study samples on disease, so it produces the odds ratio, which approximates the relative risk when the disease is rare. Knowing the design tells you immediately which measure a paper can legitimately report, and a mismatch between design and measure is a red flag worth pausing over.
| Feature | Cohort study | Case-control study |
|---|---|---|
| Starts from | Exposure | Disease |
| Direction | Forward to outcome | Backward to exposure |
| Measure of association | Relative risk (and attributable risk) | Odds ratio |
| Best for | Rare exposures; multiple outcomes | Rare diseases; long latency |
| Key weakness | Cost, time, loss to follow-up | Recall and selection bias |
Choosing the design
The choice follows the question. To study many possible effects of a single unusual exposure, say the health of astronauts or of a chemical plant's workforce, a cohort is natural, because it starts from that exposure and can watch for every outcome. To study a rare cancer with dozens of suspected causes, a case-control study is far more efficient, because it starts from the scarce cases and looks back at all the exposures at once. A memorable heuristic: cohorts are good for rare exposures, case-control studies are good for rare diseases.
When resources and an existing cohort allow, the nested case-control design captures the best of both, delivering case-control efficiency with cohort-quality, pre-collected exposure data. In practice the decision also weighs how common the disease is, how long its latency runs, how many exposures and outcomes are of interest, and how much time and money are available. No single design is best in the abstract; the best design is the one matched to the specific question.
Efficiency, in a sentence
The deepest reason the two designs suit opposite problems is arithmetic. A cohort spends its effort in proportion to how many people it must follow to see enough cases, so a rare disease makes it wasteful. A case-control study spends its effort in proportion to how many cases and controls it must interview, so a rare exposure makes it wasteful, since few in either group will have been exposed. Match the scarce quantity, disease or exposure, to the design that starts from it, and the study stays efficient.
A historical pairing: smoking studies
The case for smoking and lung cancer was built on both designs in tandem. In 1950 Richard Doll and Austin Bradford Hill published a case-control study comparing lung-cancer patients with other hospital patients and found far heavier smoking among the cases. The following year they launched the British Doctors Study, a prospective cohort that followed physicians for decades and confirmed the same association forward in time, with a clear dose-response. Two designs, opposite in direction, pointing to one conclusion, illustrate how converging evidence is assembled.
Common misconceptions
A frequent error is conflating a retrospective cohort with a case-control study. Both use existing records, but a retrospective cohort still starts from exposure and moves forward to outcome, while a case-control study starts from disease and looks back, so they yield different measures. Another error is expecting a case-control study to report a relative risk; it cannot, and the odds ratio it does report only approximates the relative risk when the disease is rare. A third is assuming a cohort is always superior, when for a rare disease it can be hopelessly inefficient.
Try it
You must study a very rare liver cancer thought to be linked to several past chemical and dietary exposures, and you need results within two years. State which design fits, give two reasons, and name the measure of association it will produce and whether that measure approximates a relative risk here.
Worked answer: a case-control study fits. First, it starts from the scarce cases, which is efficient for a rare disease that a cohort would take many years and huge numbers to accumulate. Second, it can examine several suspected exposures at once by asking about a whole history, and it is fast and inexpensive. It yields an odds ratio, which, because this cancer is rare, will closely approximate the relative risk.
Sources
- Song, J. W., & Chung, K. C. (2010). Observational studies: Cohort and case-control studies. Plastic and Reconstructive Surgery, 126(6), 2234-2242. pmc.ncbi.nlm.nih.gov
- Tenny, S., Kerndt, C. C., & Hoffman, M. R. (2023). Case control studies. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Munnangi, S., & Boktor, S. W. (2023). Epidemiology of study design. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Doll, R., & Hill, A. B. (1950). Smoking and carcinoma of the lung: Preliminary report. British Medical Journal, 2(4682), 739-748. pubmed.ncbi.nlm.nih.gov
- Doll, R., Peto, R., Boreham, J., & Sutherland, I. (2004). Mortality in relation to smoking: 50 years' observations on male British doctors. BMJ, 328(7455), 1519. pubmed.ncbi.nlm.nih.gov
- Bao, Y., Bertoia, M. L., Lenart, E. B., Stampfer, M. J., Willett, W. C., Speizer, F. E., & Chavarro, J. E. (2016). Origin, methods, and evolution of the three Nurses' Health Studies. American Journal of Public Health, 106(9), 1573-1581. pubmed.ncbi.nlm.nih.gov
- Perez-Guerrero, E. E., Guillen-Medina, M. R., Marquez-Sandoval, F., Vera-Cruz, J. M., Gallegos-Arreola, M. P., Rico-Mendez, M. A., ... Gutierrez-Hurtado, I. A. (2024). Methodological and statistical considerations for cross-sectional, case-control, and cohort studies. Journal of Clinical Medicine, 13(14), 4005. pmc.ncbi.nlm.nih.gov
- Key terms
- Cohort study
- A design that classifies people by exposure and follows them forward to observe disease; yields relative risk.
- Case-control study
- A design that starts from cases and controls and looks backward at exposure; yields the odds ratio.
- Prospective vs retrospective cohort
- Following a cohort into the future versus reconstructing past exposure from records and following forward.
- Loss to follow-up
- Dropout of cohort participants; damaging if related to both exposure and outcome.
- Recall bias
- Differential accuracy of remembered exposure between cases and controls, a hazard of case-control studies.
- Nested case-control study
- A case-control study drawn from within a cohort, combining efficiency with pre-collected exposure data.
The Randomized Controlled Trial
- Explain how randomization controls confounding and defines the RCT.
- Describe blinding, placebo control, and intention-to-treat analysis.
- State the strengths, limits, and ethical constraints of experimental designs.
All the designs so far are observational: the investigator watches exposures that people chose or encountered on their own. The randomized controlled trial (RCT) is experimental, meaning the investigator assigns the exposure, usually a treatment or preventive intervention, by chance. This one feature makes the RCT the strongest single design for establishing that an intervention causes an effect, and the foundation of evidence-based medicine.
The power comes at a price in cost, time, and ethical constraint, so trials are reserved for questions where they are both feasible and justified. Understanding how a good trial is built, and where even a trial can go wrong, is essential both for reading treatment evidence and for appreciating why so much of epidemiology must rely on observational designs instead.
Why randomization is powerful
In observational studies, people who take a treatment often differ systematically from those who do not. They may be healthier, wealthier, or more health-conscious, and these differences, the confounders of a later lesson, can masquerade as effects. Randomization, assigning each participant to a group by chance, tends to distribute all characteristics, both those we know and those we have never thought of, evenly between the groups. That is randomization's unique gift: it balances even unknown confounders, something no amount of statistical adjustment in an observational study can guarantee. If the groups then differ in outcome, the difference can be attributed to the intervention.
What randomization buys, concretely
Imagine a confounder, say baseline severity, that strongly affects outcome. In an observational comparison, sicker patients might gravitate toward a new treatment, making it look worse than it is. Randomization assigns treatment by a coin flip that ignores severity entirely, so on average the two arms contain similar mixes of mild and severe cases, and this holds for every other characteristic at once, including ones no one measured. With a large enough sample, chance makes the arms nearly identical at baseline, which is why a trial report's baseline table should show the groups looking alike before treatment begins.
How people are randomized
Randomization has its own craft. Simple randomization is a straight coin flip per participant. Block randomization assigns within small blocks to keep the arms equal in size as enrollment proceeds. Stratified randomization first divides participants by an important factor, such as disease stage, then randomizes within each stratum, guaranteeing balance on that factor rather than trusting chance alone. Each method still uses chance; they differ only in how tightly they control the arithmetic of group sizes and key characteristics.
Equally important is allocation concealment: the person enrolling a participant must not know or predict the next assignment, or they might steer sicker patients toward one arm and undo the whole point. Concealment protects the integrity of randomization at the moment of assignment, and it is distinct from blinding, which protects the study afterward. A trial can be well randomized on paper yet ruined by predictable, unconcealed allocation.
The machinery of a good trial
- Control group - a comparison arm receiving a placebo (an inert look-alike) or the current standard of care, so the intervention's effect can be isolated from the natural course of disease and the placebo effect.
- Blinding (masking) - keeping participants, and ideally the investigators and outcome assessors, unaware of who received which treatment. A double-blind trial blinds both patient and researcher, preventing expectations from coloring behavior or assessment.
- Intention-to-treat analysis - analyzing every participant in the group to which they were randomized, regardless of whether they completed the assigned treatment. This preserves the balance randomization created and gives a realistic estimate of the intervention's effect in practice; analyzing only those who complied can reintroduce bias.
These three pieces work together. The control group isolates the intervention from the disease's natural course, blinding keeps expectation from tainting behavior and measurement, and intention-to-treat protects the balance that randomization created. Weaken any one and the trial's causal claim weakens with it, which is why reports are scrutinized for all three.
The placebo effect and why controls matter
The placebo effect is the real improvement people report simply from being treated, through expectation and attention, independent of any active ingredient. Without a control arm, this effect and the natural tendency of many conditions to improve on their own would be wrongly credited to the intervention. A placebo or standard-care comparison subtracts both, isolating what the treatment itself adds. This is why an uncontrolled before-and-after comparison, however dramatic it looks, cannot substitute for a controlled trial.
Intention-to-treat and per-protocol
The choice of analysis population deserves its own attention. An intention-to-treat analysis keeps every participant in the arm they were randomized to, whatever they actually did, and it answers the practical question of how a policy of offering the treatment performs. A per-protocol analysis restricts to those who followed the assigned treatment faithfully, and while tempting, it quietly reintroduces confounding, because adherers differ systematically from non-adherers. Trials report intention-to-treat as the primary analysis for exactly this reason.
Efficacy versus effectiveness
Trials come in two temperaments. An explanatory or efficacy trial asks whether an intervention works under ideal conditions, with selected patients and tight control, maximizing the chance of detecting a true effect. A pragmatic or effectiveness trial asks whether it works in ordinary practice, with typical patients and routine care. Efficacy answers "can it work," effectiveness answers "does it work here," and a treatment can look strong in the first while disappointing in the second, which is one root of the generalizability problem below.
Variations: crossover and cluster trials
Not every trial randomizes individuals to parallel arms. In a crossover trial each participant receives both the treatment and the control in sequence, serving as their own comparison, which suits stable chronic conditions but fails when one treatment leaves a lasting carryover effect. In a cluster-randomized trial whole groups, such as clinics or villages, are randomized together, which is necessary when an intervention operates at the group level, like a community education program. Each variant keeps randomization at its heart while fitting a different kind of question.
Phases of clinical trials
New treatments move through phases. Phase I tests safety and dosing in a small group. Phase II looks for preliminary evidence of benefit and further safety. Phase III is the large, definitive randomized controlled trial that compares the new treatment against a control to establish efficacy. Phase IV is post-marketing surveillance after approval, watching for rare or delayed harms that a trial of a few thousand people could never detect. Each phase asks a narrower, more demanding question than the last.
Precision, power, and stopping
A trial must be large enough to detect a meaningful effect if one exists, a property called statistical power, fixed in advance by a sample-size calculation. Results are reported with confidence intervals, and an interval that includes no difference is inconclusive rather than proof of no effect. Long trials often include planned interim analyses and pre-set stopping rules, overseen by an independent data monitoring committee that can halt a trial early for clear benefit, clear harm, or futility, protecting participants from a comparison that has already been answered.
A worked trial result
Numbers make it concrete. Suppose 1,000 patients are randomized, 500 to a new drug and 500 to placebo. Over a year, 40 events occur in the drug arm and 80 in the placebo arm. The risk is 40 / 500 = 0.08 with the drug and 80 / 500 = 0.16 with placebo, so the relative risk is 0.08 / 0.16 = 0.50, a halving of risk.
The absolute measures complete the picture. The absolute risk reduction is 0.16 - 0.08 = 0.08, and the number needed to treat is 1 / 0.08 = 12.5, so about 13 patients must be treated for one to avoid an event. Because assignment was random and the analysis was intention-to-treat, these differences can be read as caused by the drug, not by who chose it.
Strengths, limits, and ethics
The RCT's strength is unmatched control of confounding and clear temporal order, so it sits atop the hierarchy of evidence for treatment questions. But it has real limits. Trials are expensive and often short. Their tightly selected participants may not resemble ordinary patients, limiting generalizability (external validity). And crucially, RCTs are bound by ethics: you cannot randomize people to a suspected harm.
No one may assign volunteers to smoke or to be exposed to a toxin, which is exactly why the causal case against such exposures rests on observational evidence weighed by judgment, the subject of the next module. The principle of equipoise, genuine uncertainty about which arm is better, is what makes randomizing patients ethical in the first place.
Ethical conduct has further requirements. Participants must give informed consent, understanding the risks and their freedom to withdraw. An independent ethics board, often called an institutional review board, must approve the protocol before enrollment. And the stopping rules mentioned above exist so that no one is left in an inferior arm once the evidence is decisive. These safeguards are not obstacles to good science but part of what makes a trial legitimate.
The RCT and the hierarchy of evidence
For a question about whether a treatment works, the randomized controlled trial sits near the top of the evidence hierarchy, above cohort and case-control studies, and is itself surpassed only by a systematic review that pools several sound trials. This ranking is specific to intervention questions. For questions about harm from an exposure people cannot ethically be assigned, well-conducted cohort studies are often the best available evidence, and no trial will ever outrank them because no trial can be run. Placement in the hierarchy always depends on the question being asked.
When randomization is impossible
Because you cannot ethically assign people to a suspected harm, some causal questions can never be answered by a trial, and epidemiology falls back on observational evidence weighed by judgment. Occasionally nature or policy performs a rough randomization on its own, a so-called natural experiment, as when a law changes for some regions but not others. Such quasi-experiments borrow part of the trial's logic without deliberate assignment, and they are valuable precisely where a real trial would be unethical or impossible.
Common misconceptions
One error is believing randomization guarantees balanced arms in any single trial; it only makes balance likely, and a small trial can still land lopsided by chance, which is why baseline tables are checked. Another is confusing randomization with blinding; the first balances the groups, the second keeps expectations from distorting the results, and a trial needs both. A third is treating the RCT as always the ideal, when for harmful exposures, rare outcomes, or lifelong behaviors it is simply impossible, and a well-designed cohort is the strongest evidence available.
Try it
In a drug trial, many patients in the treatment arm stop taking the drug because it makes them feel unwell, and they fare poorly. An analyst proposes excluding these dropouts and comparing only patients who took the full course. Name the analysis the analyst is proposing, explain why it is a mistake, and state which analysis should be primary and why.
Worked answer: the analyst is proposing a per-protocol analysis. It is a mistake because dropping non-adherers breaks the balance randomization created, since people who tolerate a drug differ systematically from those who do not, reintroducing confounding. The primary analysis should be intention-to-treat, keeping everyone in their assigned arm, which preserves that balance and gives an honest estimate of the drug's real-world effect.
Sources
- Medical Research Council Streptomycin in Tuberculosis Trials Committee. (1948). Streptomycin treatment of pulmonary tuberculosis: A Medical Research Council investigation. British Medical Journal, 2(4582), 769-782. pmc.ncbi.nlm.nih.gov
- Schulz, K. F., Altman, D. G., & Moher, D. (2010). CONSORT 2010 statement: Updated guidelines for reporting parallel group randomised trials. BMJ, 340, c332. pmc.ncbi.nlm.nih.gov
- Schulz, K. F., & Grimes, D. A. (2002). Allocation concealment in randomised trials: Defending against deciphering. The Lancet, 359(9306), 614-618. pubmed.ncbi.nlm.nih.gov
- Gupta, S. K. (2011). Intention-to-treat concept: A review. Perspectives in Clinical Research, 2(3), 109-112. pmc.ncbi.nlm.nih.gov
- Freedman, B. (1987). Equipoise and the ethics of clinical research. New England Journal of Medicine, 317(3), 141-145. pubmed.ncbi.nlm.nih.gov
- Bayot, M. L., Moore, N., Brannan, J. M., & Brannan, G. D. (2026). Human subjects research design. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Keyes, D., & Varacallo, M. A. (2026). Evidence-based medicine. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Key terms
- Randomized controlled trial
- An experiment in which investigators assign the intervention by chance, the strongest design for causal inference.
- Randomization
- Chance assignment to groups, which balances known and unknown confounders across arms.
- Placebo and blinding
- An inert comparison plus concealment of assignment (double-blind blinds patient and researcher) to prevent expectation bias.
- Intention-to-treat analysis
- Analyzing participants by their assigned group regardless of adherence, preserving randomization's balance.
- Generalizability (external validity)
- The extent to which trial results apply to patients beyond the study sample.
- Equipoise
- Genuine uncertainty about which trial arm is better, the ethical precondition for randomizing patients.
Module 4: Threats to Validity and Causation
The three alternative explanations for any association - bias, confounding, and chance - and how epidemiologists reason from association toward causation using the Bradford Hill viewpoints.
Bias: Selection and Information
- Define bias and distinguish selection bias from information bias.
- Recognize common named biases and how each distorts results.
- Explain why bias, unlike chance, is not reduced by a larger sample.
Whenever a study reports an association, three rival explanations must be excluded before believing it reflects a true effect: bias, confounding, and chance. This lesson takes the first. Bias is any systematic error in the design, conduct, or analysis of a study that produces a wrong estimate of the association. The word systematic is essential: unlike random error (chance), bias pushes results consistently in one direction, and, a point students often miss, a larger sample does not fix bias. It only makes a biased estimate more precisely wrong. Bias comes in two broad families.
It helps to separate internal from external validity. Internal validity asks whether the study correctly measured the association in the people it actually studied; bias is fundamentally a threat to internal validity. External validity, or generalizability, asks whether that result extends to other populations. A study can be internally valid but not generalizable, yet a study crippled by bias fails at the first hurdle, so its results cannot be trusted even for its own participants.
Bias, confounding, and chance
It is worth naming the three rivals precisely, since the rest of the module separates them. Chance is random error, the luck of sampling, and it shrinks as the study grows. Confounding is a real association carried by a third variable, addressed in the next lesson, and it can sometimes be removed in analysis. Bias is systematic error built into how the study was done, and it is neither reduced by a larger sample nor reliably removed afterward. Deciding which of the three could explain a result is the core skill of critical appraisal.
Selection bias
Selection bias arises when the people included in the study, or retained in it, differ systematically from the target population in ways related to both exposure and outcome, so the study groups are not comparable. Examples include:
- Healthy worker effect - employed people are healthier than the general population, so a workforce may appear to have lower disease rates than it truly causes.
- Berkson's bias - using hospitalized controls, who are themselves sick, can distort exposure comparisons.
- Loss to follow-up (attrition) bias - if dropout in a cohort is related to both exposure and outcome, the remaining sample is skewed.
- Self-selection (volunteer) bias - volunteers differ from non-volunteers in health and behavior.
Two further selection problems deserve names. Non-response bias arises when those who decline to participate differ systematically from those who agree, so the sample misrepresents the population. Prevalence-incidence (Neyman) bias, met earlier, arises when a study of existing cases misses those who died quickly or recovered before they could be enrolled. Both share the signature of selection bias: the people analyzed are not a fair representation of the people the study meant to describe.
Selection bias in who stays, not just who enters
Selection bias is often imagined as a fault in recruitment, but it can arise just as easily from who remains. If participants leave a study for reasons tied to both their exposure and their outcome, the analyzed sample no longer represents the group that started, and the association is distorted even though enrollment was sound. This is why loss to follow-up is treated as a selection problem, and why trials and cohorts report how many participants were lost and whether that loss differed between the groups being compared.
Berkson's bias, examined
Berkson's bias illustrates how control selection can mislead. If a case-control study draws its controls from hospital patients, and those patients are admitted for conditions themselves linked to the exposure under study, the controls' exposure rate no longer reflects the source population. The comparison is then between two selected groups rather than between cases and the community that produced them, and the odds ratio is distorted before any interview begins. Community controls, when feasible, avoid this trap by representing the population the cases came from.
Information bias
Information bias arises from systematic error in measuring exposure or outcome, so the data themselves are mismeasured. Key forms include:
- Recall bias - cases remember past exposures more completely than controls (a classic hazard of case-control studies).
- Interviewer bias - an interviewer who knows a subject's disease status probes exposures differently.
- Misclassification - errors in sorting people as exposed or diseased. Non-differential misclassification (errors unrelated to the other variable) usually blurs an association toward the null, weakening it; differential misclassification (errors that differ by group) can push an association in either direction, sometimes creating one that does not exist.
Recall bias, examined
Recall bias earns its fame in case-control studies of pregnancy and birth outcomes. A mother whose infant was born with a defect will often scour her memory for anything unusual during pregnancy, while a mother of a healthy infant recalls little, so an ordinary exposure looks falsely elevated among cases. The distortion is not dishonesty; it is the ordinary human tendency to seek explanations for a bad outcome. Designs that capture exposure from records made before the outcome, rather than from later memory, are the surest defense.
A worked example of dilution
Non-differential misclassification is worth seeing in numbers. Suppose the truth is 100 cases among 1,000 truly exposed people, a risk of 0.10, and 50 cases among 1,000 truly unexposed, a risk of 0.05, for a true relative risk of 2.0. Now suppose the exposure is measured with error that is unrelated to disease: 10 percent of the truly exposed are mislabeled unexposed, and 10 percent of the truly unexposed are mislabeled exposed. Watch what happens to the two groups.
The observed exposed group loses 100 truly exposed people carrying 10 cases but gains 100 truly unexposed carrying 5 cases, giving 95 cases in 1,000, a risk of 0.095. The observed unexposed group gains those 100 exposed with 10 cases and loses 100 unexposed with 5 cases, giving 55 cases in 1,000, a risk of 0.055. The observed relative risk is 0.095 / 0.055 = 1.73, dragged from the true 2.0 toward the null. The error did not invent an effect; it concealed a real one.
When measurement invents an effect
Differential misclassification is the more dangerous cousin, because it can create an association from nothing. Suppose exposure is truly unrelated to disease, with 60 cases among 600 exposed and 60 among 600 unexposed, a relative risk of 1.0. Now suppose diseased people, prompted to search their memories, over-report exposure, so that 20 truly unexposed cases are recorded as exposed. The case counts shift while the healthy counts do not, and the exposed group now shows more cases than the unexposed. A relative risk that was truly 1.0 is pushed above it, a spurious association built entirely from unequal measurement.
Detection and other biases
Some biases hide in how outcomes are found. Detection or surveillance bias occurs when exposed people are watched more closely and so have more disease discovered, inflating the association even when true risk is unchanged. In screening research, lead-time and length-time biases, treated fully in the screening lesson, are information biases that make early detection appear to prolong life. The lesson to carry is that measurement, not just selection, can manufacture or erase an association.
Susceptibility and immortal-time pitfalls
Two further biases round out the catalogue. Susceptibility bias arises when the exposed and unexposed differ in baseline risk for reasons tied to why they were exposed, as when sicker patients are the ones prescribed a drug, a problem sometimes called confounding by indication that behaves like selection bias. Immortal-time bias arises when a stretch of follow-up during which the outcome could not occur is misassigned to the exposed group, artificially lowering their apparent risk. Both are reminders that how time and eligibility are defined can bias a study as surely as how people are sampled.
Which way does the bias point?
A useful discipline is to ask not only whether a bias exists but which direction it pushes. Non-differential misclassification typically pulls toward the null, weakening a real effect. Differential recall may inflate an association in a case-control study. The healthy worker effect can make a hazardous workplace look protective. Reasoning explicitly about direction turns a vague worry into a concrete judgment about whether the reported effect is likely overstated, understated, or invented outright.
Guarding against bias
Because bias is built into the study's structure, it must be prevented by design. It cannot be removed afterward the way confounding sometimes can. Sound sampling, objective and identically applied measurements, blinding of assessors, choosing controls from the same source population as cases, and minimizing dropout are the defenses. Where an exposure or outcome can be measured by a record made before the study, rather than by memory, recall bias is avoided entirely, which is one reason nested case-control designs are prized.
The critical reader therefore asks, for every reported association: could the way people were selected, or the way data were collected, have manufactured this result? Answering it requires knowing how participants entered and stayed, how exposure and outcome were measured, and who knew what during measurement. These questions do more to separate credible findings from artifacts than any statistical test, because a test can quantify chance but not bias.
Reducing information bias in practice
Practical defenses against information bias are concrete. Use objective measures, such as a laboratory value or a pre-existing record, instead of self-report where possible. Apply the same instruments and definitions to every group, so that any error is at least non-differential rather than differential. Blind those who measure the outcome to the exposure status, and those who ascertain exposure to the outcome, so that knowledge cannot steer the measurement. Standardized, blinded, records-based measurement is the antidote to most of the biases in this lesson.
Common misconceptions
The first misconception is that bias, like chance, shrinks with a larger sample. It does not; more data only sharpen a wrong answer. The second is that bias can be adjusted away in analysis the way a measured confounder can; in general it cannot, because the flaw is in how the data were generated. The third is that bias always exaggerates an effect. Non-differential misclassification usually hides one, so the useful question is how large a bias plausibly is and which way it points, not merely whether it exists.
Try it
A study compares cancer rates among current employees at a factory with rates in the surrounding general population, finds the employees have lower cancer rates, and concludes the factory is safe. Name the bias, explain its mechanism, and propose a comparison group that would avoid it.
Worked answer: this is the healthy worker effect, a form of selection bias. Employed people are healthier than the general population because illness tends to remove people from the workforce, so current workers can look healthier regardless of any factory hazard. A sounder comparison uses a different employed population, or follows the factory's own workers over time including those who left, rather than benchmarking against a general population that includes the too-sick-to-work.
Sources
- Sackett, D. L. (1979). Bias in analytic research. Journal of Chronic Diseases, 32(1-2), 51-63. pubmed.ncbi.nlm.nih.gov
- Delgado-Rodriguez, M., & Llorca, J. (2004). Bias. Journal of Epidemiology and Community Health, 58(8), 635-641. pubmed.ncbi.nlm.nih.gov
- Popovic, A., & Huecker, M. R. (2023). Study bias. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Berkson, J. (1946). Limitations of the application of fourfold table analysis to hospital data. Biometrics Bulletin, 2(3), 47-53. pubmed.ncbi.nlm.nih.gov
- Suissa, S. (2008). Immortal time bias in pharmaco-epidemiology. American Journal of Epidemiology, 167(4), 492-499. pubmed.ncbi.nlm.nih.gov
- Rothman, K. J., & Greenland, S. (2005). Causation and causal inference in epidemiology. American Journal of Public Health, 95(S1), S144-S150. pubmed.ncbi.nlm.nih.gov
- Kramer, B. S., & Croswell, J. M. (2009). Cancer screening: The clash of science and intuition. Annual Review of Medicine, 60, 125-137. pubmed.ncbi.nlm.nih.gov
- Key terms
- Bias
- A systematic error in a study that yields an incorrect estimate of the association; not reduced by larger samples.
- Selection bias
- Error from how participants are chosen or retained, making study groups non-comparable.
- Information bias
- Error from inaccurate measurement of exposure or outcome.
- Recall bias
- Differential accuracy of remembered exposure between diseased and non-diseased participants.
- Non-differential misclassification
- Measurement error unrelated to the other variable, usually biasing the association toward the null.
- Healthy worker effect
- A selection bias in which employed populations appear healthier than the general population.
Confounding and How to Control It
- Define a confounder by its three defining conditions.
- Distinguish confounding from effect modification.
- List methods to control confounding in design and in analysis.
Confounding is the mixing of the effect of the exposure of interest with the effect of another variable, producing a distorted, sometimes entirely spurious, association. It is the reason correlation is not causation, and recognizing it is a defining skill of the epidemiologist. Unlike bias, which comes from how a study is built, confounding is a real feature of the world: a third variable genuinely related to both exposure and outcome creates a true association that is nonetheless not the causal one we sought.
Because it is real rather than an artifact of measurement, confounding can often be dealt with, either by design or by analysis, and much of the technical machinery of epidemiology exists for exactly that purpose. The first task, though, is recognition, so we begin with the precise definition that separates a genuine confounder from an innocent bystander.
What makes a variable a confounder
A variable is a confounder only if it meets three conditions:
- It is associated with the exposure.
- It is an independent risk factor for the outcome (associated with the disease apart from the exposure).
- It is not on the causal pathway between exposure and outcome (it is not a step by which the exposure causes the disease).
The classic illustration: a study finds that coffee drinkers have more lung cancer. But coffee drinkers are also more likely to smoke, and smoking causes lung cancer. Smoking is associated with the exposure (coffee), is an independent cause of the outcome (lung cancer), and is not a step by which coffee might cause cancer. Smoking confounds the coffee-cancer association, creating an apparent link that largely vanishes once smoking is accounted for.
Age and sex confound so many associations that they are adjusted for almost by reflex. Older people differ from younger in exposure to countless things and in the risk of nearly every disease, so any crude comparison that ignores age quietly mixes an age effect into whatever it reports. Checking the three conditions for a candidate variable is the disciplined alternative to guessing, and it is worth doing explicitly for each suspected confounder.
A worked example: confounding by age
See it in numbers. Suppose an exposure looks harmful in a crude comparison. Among 200 exposed people, 50 develop disease, a crude risk of 0.25; among 200 unexposed, 30 develop disease, a crude risk of 0.15. The crude relative risk is 0.25 / 0.15 = 1.67, apparently a real effect. But suppose the exposed group is mostly older and the unexposed group mostly younger, and age itself strongly drives the disease.
Now split by age. Among the young, the risk is 0.10 whether exposed or not, so the stratum relative risk is 1.0. Among the old, the risk is 0.30 whether exposed or not, again a stratum relative risk of 1.0. Within each age group the exposure does nothing; the entire crude association of 1.67 came from the exposed being older. Stratifying by the confounder has revealed the truth that the crude number hid, which is precisely what adjustment is for.
Detecting confounding in the data
How do you know confounding is present? The operational sign is a meaningful gap between the crude estimate and the adjusted one. If the relative risk moves from 1.67 to 1.0 after stratifying by age, age was confounding the association. A common rule of thumb treats a change of about ten percent or more in the estimate as evidence that the variable is worth adjusting for, though prior knowledge and judgment matter more than any fixed cutoff. Comparing the crude and adjusted results is the everyday diagnostic for confounding.
A second worked example: a hidden effect
Confounding can also mask or reverse a real effect. Suppose an exposure is genuinely protective, cutting risk in half within every age group, yet the exposed happen to be much older. The age penalty can cancel the protection, so the crude comparison shows no difference or even apparent harm, while stratifying by age uncovers the true relative risk of about 0.5 in each stratum. The direction of confounding depends on how the confounder relates to both exposure and outcome, so it can inflate, dilute, or entirely flip a crude estimate.
Confounding is not effect modification
These two are often confused. Confounding is a nuisance to be removed, a distortion of the true effect. Effect modification (interaction) is a real biological finding to be reported: the exposure's effect genuinely differs across levels of a third variable. If a drug lowers risk in younger patients but not older ones, age is an effect modifier, and reporting a single overall effect would hide the truth. The practical test: if adjusting for the third variable changes the estimate, suspect confounding; if the effect is truly different in each subgroup, that is effect modification, and the subgroup results should be presented separately.
Mediators are not confounders
A third relationship is easy to mishandle. A mediator lies on the causal pathway: the exposure causes the mediator, which causes the outcome. If a diet raises blood pressure, which then causes strokes, blood pressure is a mediator, not a confounder, because it is a step by which the diet acts. Adjusting for a mediator is a mistake, because it removes part of the very effect you are trying to measure, making a real cause look innocent. The third condition of confounding exists precisely to keep mediators out.
Directed acyclic graphs, briefly
Epidemiologists increasingly draw a directed acyclic graph, a simple diagram of arrows showing which variables cause which, to decide what to adjust for. A confounder sits with arrows pointing to both the exposure and the outcome, so it should be controlled. A mediator sits on the arrow from exposure to outcome, so it should not. A third structure, a collider, receives arrows from both and must not be adjusted, because conditioning on it can create a false association. The graph turns a verbal argument about the three conditions into a picture you can reason about.
Controlling confounding
Unlike bias, confounding can be addressed at two stages, in the design of the study or in its analysis.
| Stage | Method | How it works |
|---|---|---|
| Design | Randomization | Balances all confounders, known and unknown (RCTs only) |
| Design | Restriction | Admit only one level of the confounder (e.g. only non-smokers) |
| Design | Matching | Pair exposed and unexposed (or cases and controls) on the confounder |
| Analysis | Stratification | Analyze within strata of the confounder, then pool |
| Analysis | Multivariable adjustment | Use regression to statistically hold confounders constant |
The great limitation of every analytic method is that you can only adjust for confounders you have measured. Residual confounding from unmeasured or imperfectly measured variables always lurks in observational studies, which is precisely why randomization, the one method that handles the unknown, is so prized.
Design versus analysis
Each approach has trade-offs. Restriction cleanly removes a confounder by admitting only one of its levels, but it narrows the study and prevents you from examining that variable's effect. Matching forces balance on chosen variables but complicates the analysis and can waste data if overdone. Stratification is transparent and shows the effect within each level, but it strains when several confounders must be handled at once. Multivariable regression adjusts for many confounders simultaneously, at the cost of relying on modeling assumptions the reader cannot fully see.
Propensity scores and modern adjustment
Modern studies often summarize many confounders into a single propensity score, the estimated probability that a person received the exposure given their measured characteristics. Comparing exposed and unexposed people with similar propensity scores mimics some of the balance that randomization provides, and it is useful when the exposure is common but the outcome is rare. Like every analytic method, it can only balance the confounders that were actually measured, so it does not escape the fundamental limit that unmeasured confounding survives.
Stratified analysis and pooling
The worked age example is exactly a stratified analysis: compute the effect within each stratum of the confounder, then combine. When the stratum-specific estimates are similar, they can be pooled into a single adjusted estimate, classically by the Mantel-Haenszel method, which weights the strata and yields one summary free of the confounder's influence. When the stratum-specific estimates differ markedly, that is a signal of effect modification, and pooling them into one number would be misleading rather than clarifying.
Residual and unmeasured confounding
Even a careful adjustment leaves gaps. If a confounder is measured crudely, some of its effect leaks through as residual confounding. If a confounder was never measured, or never suspected, no analysis can touch it. A special case is confounding by indication, where the reason a treatment was prescribed, the patient's underlying severity, both drives the choice of treatment and predicts the outcome, so treated patients differ from untreated in ways records may not capture. These limits are why a single observational study, however well adjusted, rarely settles a causal question by itself.
Why randomization stands apart
Every method in the table except randomization shares one weakness: it can only handle confounders that were identified and measured. Randomization is unique because it balances variables the investigator never thought of, distributing unknown confounders evenly by the same chance that balances the known ones. That is why the randomized trial, for questions where it is ethical and feasible, occupies the top of the hierarchy for causal inference, and why observational findings are always read with the question of what unmeasured factor might still be at work.
Confounding across a body of evidence
No single study fully escapes confounding, so causal judgment rests on the pattern across many. When different studies, using different populations and adjusting for different sets of confounders, keep finding the same association, it becomes harder to blame any one lurking variable, since a confounder would have to operate consistently everywhere. This convergence is part of how the Bradford Hill reasoning of the next lesson builds a causal case from imperfect observational pieces.
Common misconceptions
The first misconception is confusing confounding with bias; confounding is a real third-variable effect that can often be removed, while bias is a structural flaw that usually cannot. The second is over-adjustment, controlling for a mediator or a collider and thereby distorting the very effect under study. The third is believing that a long list of adjusted variables guarantees a clean estimate; it does not, because unmeasured and residual confounding survive every regression, and only randomization addresses the confounders no one thought to measure.
Try it
A study finds that people who drink diet soda have higher rates of obesity and concludes that diet soda causes obesity. Propose a plausible confounder, verify it against the three conditions, and name one design method and one analytic method that could control it.
Worked answer: a plausible confounder is a person's baseline weight or tendency to gain weight. It is associated with the exposure, since heavier people and dieters more often choose diet soda; it is an independent risk factor for obesity; and it is not a step by which diet soda causes obesity, so all three conditions hold. It also raises reverse causation. Control it by restriction, studying people of similar baseline weight, or by multivariable adjustment for baseline weight; randomization would balance it but is impractical here.
Sources
- Jager, K. J., Zoccali, C., Macleod, A., & Dekker, F. W. (2008). Confounding: What it is and how to deal with it. Kidney International, 73(3), 256-260. pubmed.ncbi.nlm.nih.gov
- Skelly, A. C., Dettori, J. R., & Brodt, E. D. (2012). Assessing bias: The importance of considering confounding. Evidence-Based Spine-Care Journal, 3(1), 9-12. pubmed.ncbi.nlm.nih.gov
- VanderWeele, T. J., & Shpitser, I. (2013). On the definition of a confounder. Annals of Statistics, 41(1), 196-220. pubmed.ncbi.nlm.nih.gov
- Greenland, S., Pearl, J., & Robins, J. M. (1999). Causal diagrams for epidemiologic research. Epidemiology, 10(1), 37-48. pubmed.ncbi.nlm.nih.gov
- Austin, P. C. (2011). An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate Behavioral Research, 46(3), 399-424. pmc.ncbi.nlm.nih.gov
- Popovic, A., & Huecker, M. R. (2023). Study bias. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Rothman, K. J., & Greenland, S. (2005). Causation and causal inference in epidemiology. American Journal of Public Health, 95(S1), S144-S150. pubmed.ncbi.nlm.nih.gov
- Key terms
- Confounding
- Distortion of an exposure-disease association by a third variable linked to both.
- Confounder
- A variable associated with the exposure, an independent risk factor for the outcome, and not on the causal pathway.
- Effect modification (interaction)
- A real difference in the exposure's effect across levels of a third variable, to be reported not removed.
- Restriction
- Limiting the study to one level of a potential confounder to remove its effect.
- Stratification
- Analyzing the association separately within strata of a confounder and then pooling.
- Residual confounding
- Confounding that remains because a confounder was unmeasured or imperfectly measured.
From Association to Causation: The Bradford Hill Viewpoints
- Explain why statistical association alone does not prove causation.
- Summarize the Bradford Hill viewpoints for weighing causal evidence.
- Apply the viewpoints to a real body of evidence such as smoking and lung cancer.
Suppose a well-conducted study reports an association, and bias, confounding, and chance have been reasonably excluded. Is the exposure a cause of the disease? Epidemiology rarely proves causation from a single study the way a laboratory experiment might. Instead, causal judgment is reached by weighing an entire body of evidence against a set of considerations proposed in 1965 by the British statistician Sir Austin Bradford Hill. These are not a checklist to be scored, and no single one is necessary or sufficient. They are viewpoints that, taken together, strengthen or weaken the case for causation.
Hill offered them in a now-famous address, drawing on his own work establishing that smoking causes lung cancer. He was explicit that they were aids to judgment, not rules, and that demanding all of them would be as mistaken as ignoring them. A century of practice has treated them as the field's standard framework for the hardest question it faces, moving from a measured association to a claim about cause.
Association is not causation
Before applying the viewpoints, recall why an association might not be causal at all. It could be chance, a fluke of sampling that a confidence interval helps gauge. It could be bias, a systematic flaw in how the study was built. It could be confounding, a real third-variable effect masquerading as the one of interest. Or it could be reverse causation, where the disease changes the exposure rather than the other way around. The viewpoints are applied only after these rivals have been weighed, not instead of weighing them.
The nine viewpoints
- Strength of association - a large relative risk is harder to explain away by undetected bias or confounding than a small one.
- Consistency - the association is found repeatedly, by different investigators, in different populations, with different methods.
- Specificity - one exposure leads to one outcome. This is the weakest viewpoint, since many exposures have multiple effects, and its absence does not argue against cause.
- Temporality - the exposure precedes the disease. This is the one viewpoint that is absolutely required: a cause must come before its effect.
- Biological gradient (dose-response) - more exposure produces more disease, as smoking more cigarettes raises lung-cancer risk further.
- Plausibility - a credible biological mechanism exists (though this is limited by the science of the day).
- Coherence - the causal interpretation does not conflict with what is known about the disease's natural history and biology.
- Experiment - removing the exposure reduces the disease, as when quitting smoking lowers risk or removing a contaminated water source ends an outbreak.
- Analogy - established causal relationships for similar exposures make a new one more believable.
No one viewpoint carries the argument. The case for causation is built like a legal case, from converging lines of evidence that are individually inconclusive but together compelling. The more of these that point the same way, and the harder each is to explain by bias or confounding, the stronger the causal verdict.
Temporality is non-negotiable
Of all nine, only temporality is a strict requirement. Strength, consistency, and dose-response are powerful when present, but the exposure must come first, which is why cross-sectional studies, unable to fix time order, cannot establish cause on their own, and why longitudinal designs are so valued. If the disease could have preceded the exposure, no other viewpoint can rescue the causal claim, because effect cannot come before cause.
Strength, consistency, and dose-response
Three viewpoints do much of the heavy lifting. Strength matters because a large relative risk would require a correspondingly large hidden bias or confounder to explain away, and such a large lurking factor is usually implausible or would already be known. A weak association, by contrast, could be produced by modest, easily overlooked flaws, so it demands more caution before any causal reading.
Consistency and dose-response reinforce strength. When many studies of different designs in different populations find the association, a single shared flaw becomes hard to blame. When more exposure yields more disease in a graded way, the pattern is difficult to fake, since few biases produce a neat gradient. Together these three convert a striking correlation into a serious causal candidate.
A worked example of dose-response
Dose-response is worth making concrete. Imagine a study reporting lung-cancer relative risks by smoking level: about 1 for non-smokers by definition, roughly 5 for light smokers, 10 for moderate smokers, and 20 or more for heavy smokers. Each step up in exposure brings a step up in risk, a monotone gradient. Such a pattern is powerful evidence because few biases or confounders would arrange themselves to rise smoothly with dose. When quitting then moves a person's risk back down over years, the gradient runs in reverse, reinforcing the causal reading.
Strength has limits as evidence
Strength is persuasive but not decisive on its own. A large association can still be produced by a large bias, as the healthy worker effect or severe confounding by indication can generate impressive ratios that are not causal. Conversely, a genuinely causal effect can be modest, as many real risk factors for common diseases are, so a small relative risk is not evidence against causation. Strength raises or lowers the prior, but it is weighed alongside the other viewpoints rather than trusted alone.
The viewpoints in action: smoking and lung cancer
The mid-twentieth-century case that smoking causes lung cancer, built entirely from observational data because randomizing people to smoke was unthinkable, illustrates the viewpoints well. The association was strong, with heavy smokers carrying many times the risk. It was consistent across dozens of studies and countries. It showed a clear dose-response gradient, more cigarettes bringing more cancer. It respected temporality, smoking preceding cancer by years. It was biologically plausible and coherent with carcinogens identified in smoke.
It even satisfied the experiment viewpoint, because those who quit saw their risk fall over time. No single study proved it; the convergence did, drawing on both the 1950 case-control study and the British Doctors cohort begun in 1951. A landmark 1964 government report weighed exactly this body of evidence, by essentially the reasoning Hill would formalize, and concluded that smoking causes lung cancer. This is how epidemiology reasons its way from association to cause, and it is the intellectual heart of the discipline.
The experiment viewpoint and natural experiments
The experiment viewpoint carries special weight because it comes closest to intervention. When removing an exposure lowers disease, the causal case strengthens sharply. Snow's removal of the Broad Street pump handle in 1854 and the subsequent fall in cholera is an early instance; the decline in lung-cancer risk after people quit smoking is another. These are not randomized trials, but they show that acting on the suspected cause changes the outcome, which is hard to explain by confounding. Whenever a deliberate or natural change in exposure tracks a change in disease, this viewpoint is in play.
Consistency and the weight of replication
No viewpoint does more quiet work than consistency. A single study can be wrong for many reasons, but an association found repeatedly, by different teams, in different countries, with cohort and case-control designs alike, is hard to attribute to one shared flaw, because each study's biases and confounders differ. Replication does not guarantee truth, since a shared misconception can propagate through a literature, but broad, methodologically varied consistency is among the strongest signals epidemiology can offer short of an experiment.
What causation means
Underlying the viewpoints is a counterfactual idea of cause: an exposure causes an outcome if the outcome would not have happened, in that person or population, had the exposure been absent. Because we can never observe the same people both exposed and unexposed, epidemiology approximates the missing counterfactual with a comparison group, which is why comparability and confounding dominate the earlier lessons. A companion idea, developed in the chronic-disease lesson, is that most outcomes arise from a set of component causes acting together rather than from one cause alone.
Coherence and plausibility, together
Coherence asks that the causal story fit what is already known about the disease, and plausibility asks that a mechanism be credible. The two overlap but differ: coherence concerns consistency with the broader picture, such as trends in disease tracking trends in exposure, while plausibility concerns a specific biological pathway. Both are constrained by current knowledge, so their absence is a weak argument against cause. When present, however, they make a causal interpretation more comfortable and help rule out alternative explanations.
Using and misusing the viewpoints
The viewpoints are aids to judgment, not a scorecard, and misusing them is common. Specificity is the weakest, since real causes often have many effects, so its absence argues nothing. Plausibility and coherence are bounded by the science of the day, and a mechanism unknown now may be discovered later, so a missing mechanism is not a refutation. Analogy is suggestive at best. Treating the list as a tally to be summed, rather than as considerations to be weighed, is the classic error.
Reasoning under uncertainty
The viewpoints matter because epidemiology usually cannot run the decisive experiment. For a suspected harm, no ethical trial is possible, so the field must reach a judgment from observational evidence that is individually fallible. The Bradford Hill viewpoints are a disciplined way to weigh that evidence, converting a scattered literature into a reasoned verdict about cause. They are the intellectual bridge from the measures and designs of earlier modules to the decisions of public health, which often cannot wait for a certainty that may never arrive.
From verdict to action
A causal verdict is not the end but the beginning of policy. Once the weight of the viewpoints supports cause, the attributable-risk measures estimate how much disease is at stake, and the tools of the final module turn that estimate into surveillance, guidelines, and prevention. The smoking case shows the whole arc: decades of observational evidence, weighed by these viewpoints, justified tobacco control that has prevented an enormous number of deaths. Causal reasoning is what licenses action in the absence of a trial.
Common misconceptions
The first misconception is treating the nine viewpoints as a checklist that must all be met; only temporality is required, and the rest accumulate weight. The second is expecting a single study to prove causation, when the framework is designed for a body of evidence. A further error is demanding every viewpoint before accepting cause: a case can be compelling on strength, consistency, temporality, dose-response, and experiment while scoring poorly on the softer viewpoints, and insisting on a perfect card would paralyze judgment where action is needed.
Try it
A cross-sectional survey finds that people with a certain diet also have more of a disease, but it cannot tell whether the diet or the disease came first. Name the Bradford Hill viewpoint that is unmet, explain why its absence blocks a causal claim, and state what kind of study would supply it.
Worked answer: the unmet viewpoint is temporality. Because a cause must precede its effect, and the snapshot cannot establish which came first, no causal conclusion can be drawn no matter how strong the association looks. Temporality is the one viewpoint that is truly required. A prospective cohort, measuring the diet in disease-free people and following them forward for new cases, would establish that the exposure preceded the outcome and supply the missing viewpoint.
Sources
- Hill, A. B. (1965). The environment and disease: Association or causation? Proceedings of the Royal Society of Medicine, 58(5), 295-300. pmc.ncbi.nlm.nih.gov
- Fedak, K. M., Bernal, A., Capshaw, Z. A., & Gross, S. (2015). Applying the Bradford Hill criteria in the 21st century: How data integration has changed causal inference in molecular epidemiology. Emerging Themes in Epidemiology, 12, 14. pmc.ncbi.nlm.nih.gov
- Rothman, K. J., & Greenland, S. (2005). Causation and causal inference in epidemiology. American Journal of Public Health, 95(S1), S144-S150. pubmed.ncbi.nlm.nih.gov
- Dhawan, R., & Shay, D. (2024). Principles of causation. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- U.S. Department of Health and Human Services. (2014). The health consequences of smoking - 50 years of progress: A report of the Surgeon General. Centers for Disease Control and Prevention. ncbi.nlm.nih.gov
- Doll, R., Peto, R., Boreham, J., & Sutherland, I. (2004). Mortality in relation to smoking: 50 years' observations on male British doctors. BMJ, 328(7455), 1519. pubmed.ncbi.nlm.nih.gov
- U.S. Surgeon General's Advisory Committee on Smoking and Health. (1964). Smoking and health: Report of the Advisory Committee to the Surgeon General of the Public Health Service (PHS Publication No. 1103). U.S. Public Health Service. find source ↗
- Key terms
- Bradford Hill viewpoints
- Nine considerations for weighing whether an association is causal; guidelines, not a rigid checklist.
- Temporality
- The requirement that the exposure precede the disease; the one strictly necessary causal criterion.
- Strength of association
- A large effect size is harder to attribute to undetected bias or confounding.
- Consistency
- Repeated observation of the association across studies, populations, and methods.
- Biological gradient (dose-response)
- Greater exposure producing greater risk, strong evidence for causation.
- Specificity
- One exposure leading to one outcome; the weakest viewpoint, since exposures often have many effects.
Module 5: Screening and Diagnostic Tests
How epidemiology evaluates tests that detect disease - sensitivity, specificity, and predictive values - and why the same test performs differently in different populations.
Sensitivity, Specificity, and Predictive Values
- Build a test-by-disease two-by-two table and compute sensitivity and specificity.
- Compute positive and negative predictive values and explain their dependence on prevalence.
- Explain the trade-off between sensitivity and specificity.
Secondary prevention depends on screening: applying a test to apparently healthy people to detect disease early. To evaluate any test, screening or diagnostic, epidemiologists compare its result against the truth, established by a gold standard, in a familiar two-by-two table. This time the rows are the test result and the columns are the true disease status:
| Disease + (truth) | Disease - (truth) | |
|---|---|---|
| Test + | True Positive (TP) | False Positive (FP) |
| Test - | False Negative (FN) | True Negative (TN) |
The four cells name every possible outcome. A true positive and a true negative are correct calls. The two errors are a false positive, alarming a healthy person, and a false negative, missing real disease. Everything in this lesson is a ratio built from these four numbers, read either down the columns to describe the test or across the rows to describe what a result means to a patient.
Sensitivity and specificity: properties of the test
Sensitivity is the proportion of people with the disease whom the test correctly calls positive: TP / (TP + FN). A highly sensitive test rarely misses disease, so a negative result from it is reassuring, a good rule-out test, captured by the mnemonic SnNout: a Sensitive test that is Negative rules disease out. Specificity is the proportion of people without the disease whom the test correctly calls negative: TN / (TN + FP). A highly specific test rarely falsely alarms, so a positive result is convincing, a good rule-in test, by the mnemonic SpPin.
Crucially, sensitivity and specificity are relatively stable properties of the test itself and do not change with disease prevalence. They are measured by reading down the columns, within the diseased and within the healthy separately, so the mix of diseased to healthy in the population does not affect them. This stability is what lets a test be characterized once and then applied in many settings, though, as the next lesson shows, what a result means still shifts with prevalence.
Worked example
A new screening test is applied to 1,000 people, of whom 100 truly have the disease (10 percent prevalence). The test correctly flags 90 of the diseased and misses 10; among the 900 healthy people it correctly clears 720 and falsely alarms 180. The table:
| Disease + | Disease - | Total | |
|---|---|---|---|
| Test + | TP = 90 | FP = 180 | 270 |
| Test - | FN = 10 | TN = 720 | 730 |
| Total | 100 | 900 | 1,000 |
Sensitivity = 90 / (90 + 10) = 90 / 100 = 0.90 = 90 percent.
Specificity = 720 / (720 + 180) = 720 / 900 = 0.80 = 80 percent.
The complements are worth naming. The false-negative rate is 1 minus sensitivity, here 10 percent, the share of diseased people the test misses. The false-positive rate is 1 minus specificity, here 20 percent, the share of healthy people it wrongly flags. That 20 percent applied to a large healthy majority will matter enormously for what a positive result means, as the predictive values now reveal.
Two kinds of error and their costs
The two errors are not equally costly. A false negative misses real disease, delaying treatment and offering false reassurance, which is grave when the disease is dangerous and treatable. A false positive alarms a healthy person, triggering anxiety, further tests, and sometimes harm from those tests. Which error matters more depends on the disease and the follow-up available, and that judgment drives whether a program is tuned to favor sensitivity or specificity.
Predictive values: what a result means for a patient
A patient does not ask "how sensitive is the test"; they ask "I tested positive, do I have the disease?" That is the positive predictive value (PPV): the proportion of test-positives who truly have the disease, TP / (TP + FP). The negative predictive value (NPV) is the proportion of test-negatives who are truly disease-free, TN / (TN + FN). These are read across the rows, so they mix the diseased and the healthy in whatever proportion the population supplies.
PPV = 90 / (90 + 180) = 90 / 270 = 0.333 = 33.3 percent.
NPV = 720 / (720 + 10) = 720 / 730 = 0.986 = 98.6 percent.
Strikingly, even though this test is 90 percent sensitive, only one in three positives actually has the disease. That gap, between how good the test is and what a positive means, is governed by prevalence, the subject of the next lesson. The negative predictive value, by contrast, is reassuringly high here, so a negative result is trustworthy in this setting.
Predictive value and Bayes' theorem
The predictive values are an application of Bayes' theorem, which updates the probability of disease in light of a test result. Written with prevalence, sensitivity, and specificity, the positive predictive value is (sensitivity × prevalence) divided by [(sensitivity × prevalence) + ((1 - specificity) × (1 - prevalence))]. Substituting our numbers: (0.90 × 0.10) divided by [(0.90 × 0.10) + (0.20 × 0.90)] = 0.09 / (0.09 + 0.18) = 0.09 / 0.27 = 0.333, the same 33.3 percent as before.
The formula makes the dependence on prevalence explicit. The denominator's second term, the false positives, grows as (1 - prevalence) grows, so as disease becomes rarer the false positives swell and the positive predictive value falls. Sensitivity and specificity stay fixed while the predictive value slides, which is the single most important and counterintuitive fact about testing, developed fully in the screening lesson.
A worked prevalence comparison
To feel the prevalence effect before the next lesson, apply this same test, sensitivity 0.90 and specificity 0.80, to a lower-prevalence group. Suppose only 1 percent of 10,000 people have the disease, so 100 are diseased and 9,900 healthy. True positives are 0.90 × 100 = 90, but false positives are 0.20 × 9,900 = 1,980. The positive predictive value is 90 / (90 + 1,980) = 90 / 2,070 = 0.043, about 4 percent. The identical test that gave a 33 percent PPV at 10 percent prevalence gives only 4 percent here, because the healthy majority floods the positives with false alarms.
Likelihood ratios
A compact way to summarize a test uses likelihood ratios. The positive likelihood ratio is sensitivity divided by (1 - specificity), here 0.90 / 0.20 = 4.5, meaning a positive result is 4.5 times as likely in a diseased person as in a healthy one. The negative likelihood ratio is (1 - sensitivity) divided by specificity, here 0.10 / 0.80 = 0.125. Likelihood ratios do not depend on prevalence and can be combined with a patient's pre-test odds to yield post-test odds, a clean Bayesian update.
The sensitivity-specificity trade-off
For any test with a numeric cutoff, moving the threshold trades one property for the other. Loosening the cutoff to catch more true cases raises sensitivity but lowers specificity, producing more false alarms; tightening it does the reverse. There is no free lunch. Which way to lean depends on the stakes: a screening test for a dangerous, treatable disease favors high sensitivity, do not miss cases, accepting more false positives that a confirmatory test will later sort out.
ROC curves and choosing a cutoff
Plotting sensitivity against 1 minus specificity as the cutoff varies traces a receiver operating characteristic (ROC) curve. A test that is all guesswork lies on the diagonal; a good test bows toward the top-left corner, and the area under the curve summarizes overall discrimination, from 0.5 for useless to 1.0 for perfect. The curve does not choose the cutoff for you; that choice weighs the costs of the two errors and the prevalence of disease, but it shows the full menu of sensitivity-specificity pairs available.
Cutoffs in practice
The abstract trade-off becomes concrete once a test yields a number, like a blood concentration, rather than a simple positive or negative. Setting the threshold low labels more people positive, raising sensitivity and lowering specificity; setting it high does the reverse. A program for a serious, treatable disease with a safe confirmatory step will set the cutoff to prioritize sensitivity, while a test whose positives trigger a risky procedure will set it to protect specificity. The single number chosen determines every cell of the two-by-two table.
Screen, then confirm
A common strategy uses the two properties in sequence. A first, highly sensitive test screens the population, rarely missing true cases but generating many false positives. Everyone who screens positive then receives a second, highly specific confirmatory test that clears most of the false alarms. The sensitive test rules disease out for the many who are negative; the specific test rules it in for the few who remain positive. Sequential testing captures the strengths of both while limiting the harms of each.
Why accuracy alone misleads
It is tempting to judge a test by accuracy, the share of all results that are correct, (TP + TN) divided by the total. But accuracy is treacherous when disease is rare. A test that simply calls everyone negative will be right about a rare disease almost every time, scoring high accuracy while catching not a single case. That is why sensitivity, specificity, and the predictive values, which examine the diseased and the healthy separately, tell you far more than a single accuracy figure ever can.
Reference standards and their limits
Every one of these measures is defined against a gold standard taken as truth, but real reference tests are imperfect. If the standard itself misses some disease, a new test that correctly flags those cases will be counted as producing false positives, unfairly lowering its apparent specificity. When no perfect standard exists, methods that combine several imperfect tests are used, but the reader should always ask what the reported sensitivity and specificity were measured against, since that answer sets the ceiling on how good any test can appear.
Common misconceptions
The commonest error is confusing sensitivity with positive predictive value. Sensitivity asks how many of the diseased test positive; predictive value asks how many of the positives are diseased, and the two can be wildly different, as the 90 percent sensitive test with a 33 percent PPV shows. A related error treats a positive result as a diagnosis. A screening positive is a signal to investigate, not a verdict, especially when the disease is uncommon and most positives are false, so honest communication of that fact is part of running a program ethically.
Try it
A test with sensitivity 0.85 and specificity 0.90 is applied to 2,000 people, of whom 200 truly have the disease. Build the two-by-two table from these figures, then compute the positive and negative predictive values, showing each step.
Worked answer: among 200 diseased, TP = 0.85 × 200 = 170 and FN = 30. Among 1,800 healthy, TN = 0.90 × 1,800 = 1,620 and FP = 180. PPV = TP / (TP + FP) = 170 / (170 + 180) = 170 / 350 = 0.486, about 49 percent. NPV = TN / (TN + FN) = 1,620 / (1,620 + 30) = 1,620 / 1,650 = 0.982, about 98 percent. At 10 percent prevalence, fewer than half of positives are real, even with a strong test.
Sources
- Shreffler, J., & Huecker, M. R. (2023). Diagnostic testing accuracy: Sensitivity, specificity, predictive values and likelihood ratios. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Simon, D., & Boring, J. R. (1990). Sensitivity, specificity, and predictive value. In H. K. Walker, W. D. Hall, & J. W. Hurst (Eds.), Clinical methods: The history, physical, and laboratory examinations (3rd ed.). Butterworths. ncbi.nlm.nih.gov
- Altman, D. G., & Bland, J. M. (1994). Statistics notes: Diagnostic tests 1 - Sensitivity and specificity. BMJ, 308(6943), 1552. pmc.ncbi.nlm.nih.gov
- Altman, D. G., & Bland, J. M. (1994). Statistics notes: Diagnostic tests 2 - Predictive values. BMJ, 309(6947), 102. pmc.ncbi.nlm.nih.gov
- Trevethan, R. (2017). Sensitivity, specificity, and predictive values: Foundations, pliabilities, and pitfalls in research and practice. Frontiers in Public Health, 5, 307. pmc.ncbi.nlm.nih.gov
- Hajian-Tilaki, K. (2013). Receiver operating characteristic (ROC) curve analysis for medical diagnostic test evaluation. Caspian Journal of Internal Medicine, 4(2), 627-635. pmc.ncbi.nlm.nih.gov
- Gigerenzer, G., & Edwards, A. (2003). Simple tools for understanding risks: From innumeracy to insight. BMJ, 327(7417), 741-744. pubmed.ncbi.nlm.nih.gov
- Key terms
- Sensitivity
- The proportion of truly diseased people the test correctly calls positive, TP / (TP + FN); a sensitive test rules disease out when negative.
- Specificity
- The proportion of truly non-diseased people the test correctly calls negative, TN / (TN + FP); a specific test rules disease in when positive.
- Positive predictive value
- The proportion of test-positives who truly have the disease, TP / (TP + FP); depends on prevalence.
- Negative predictive value
- The proportion of test-negatives who are truly disease-free, TN / (TN + FN).
- Gold standard
- The reference test taken as truth against which a new test's accuracy is judged.
- Sensitivity-specificity trade-off
- Shifting a test's cutoff raises one of sensitivity or specificity while lowering the other.
Prevalence, Predictive Value, and Screening Programs
- Explain why positive predictive value falls as prevalence falls.
- Apply the same test to high- and low-prevalence settings and compare results.
- List the criteria for a worthwhile screening program and the biases that flatter it.
The most counterintuitive fact in test evaluation is this: the same test, with fixed sensitivity and specificity, gives very different answers to the patient depending on how common the disease is. Sensitivity and specificity belong to the test; predictive values belong to the population. Understanding why is essential to interpreting any positive result honestly, and to deciding whom it is worth screening in the first place.
This single idea explains why a positive on a rare-disease screen is usually a false alarm, why screening is targeted to high-risk groups, and why a test that performs beautifully in a specialty clinic can mislead when applied to the general public. It follows directly from Bayes' theorem, introduced in the previous lesson, made concrete with whole-population counts.
Why prevalence drives predictive value
When a disease is rare, the vast majority of people tested are healthy, so even a small false-positive rate generates a large number of false positives, which can swamp the true positives and drive the positive predictive value down. When a disease is common, positives are much more likely to be real. Consider a test with sensitivity 95 percent and specificity 90 percent applied to 100,000 people at three different prevalences:
| Prevalence | True positives | False positives | PPV | NPV |
|---|---|---|---|---|
| 1 percent (rare) | 950 | 9,900 | 8.8 percent | 99.9 percent |
| 10 percent | 9,500 | 9,000 | 51.4 percent | 99.4 percent |
| 50 percent (common) | 47,500 | 5,000 | 90.5 percent | 94.7 percent |
At 1 percent prevalence, a positive result is right less than 1 time in 10, because the 9,900 false positives (10 percent of the 99,000 healthy people) dwarf the 950 true positives. Yet the identical test at 50 percent prevalence gives a positive that is right 9 times in 10. This is why screening the general low-risk population for a rare disease produces mostly false alarms, and why we target screening to higher-risk groups, where prevalence, and thus PPV, is higher.
It is also the mathematics behind the anxiety of a positive result on a rare-disease screen: most such positives are false, a fact that ought to shape how results are communicated. A person told only that they screened positive, without being told the predictive value in their group, may badly overestimate the chance they are truly ill.
A worked look at the rare-disease case
Return to the 1 percent row and trace the counts. Of 100,000 people, 1,000 are diseased and 99,000 are healthy. The test finds 950 of the diseased at 95 percent sensitivity, but wrongly flags 9,900 of the healthy, since a 90 percent specificity leaves a 10 percent false-positive rate. Among the 10,850 who screen positive, only 950 are truly ill, so the positive predictive value is 950 / 10,850 = 8.8 percent. More than nine in ten positives are false, not because the test is poor but because the disease is rare, which is the arithmetic behind cautious reading of any low-prevalence screen.
Bayes made visual
The clearest way to see this is to reason in natural frequencies, whole counts rather than percentages. Instead of juggling a 95 percent sensitivity and a 1 percent prevalence in the abstract, imagine 100,000 real people, watch 950 true positives emerge alongside 9,900 false positives, and read the predictive value straight off the counts. Studies of how clinicians and patients reason show that this framing prevents the errors that percentages invite, which is why the table above uses counts rather than probabilities alone.
Targeted screening and pre-test probability
The lever that makes screening useful is pre-test probability, essentially the prevalence in the group actually tested. Restricting a screen to people whose age, family history, symptoms, or exposures raise their risk lifts the prevalence in the tested group, which lifts the positive predictive value without touching the test. This is why screening guidelines specify who should be screened and when, rather than urging everyone to be tested for everything. Aim a test at a low-risk crowd and it drowns in false positives; aim it at a high-risk group and it earns its keep.
Screening changes what a positive means, not the test
It bears repeating in a different form. When a program moves from a high-risk clinic to the general population, the test's sensitivity and specificity do not budge, yet its positive predictive value can collapse from most positives being real to most being false. Nothing about the instrument changed; only the pre-test probability of the people fed into it did. This is why the same blood test or swab can be a sound tool in one setting and a generator of false alarms in another, a distinction that belongs at the center of any screening decision.
What makes screening worthwhile
Screening is not automatically good; a program must satisfy conditions long associated with Wilson and Jungner. The disease should be serious and reasonably common, with a recognizable early or latent stage, and its natural history understood. The test should be safe, acceptable to the public, and sufficiently sensitive and specific for the setting. And, most important, there must be an effective treatment that works better when begun early, since screening for an untreatable disease only lengthens the time a person knows they are ill without changing the outcome.
Beyond these classic criteria, the program itself matters. An organized screening program with defined intervals, quality control, and reliable follow-up of positives generally outperforms scattered opportunistic testing, because a screen is only as good as the workup and treatment that follow a positive result. A test that identifies disease no one then investigates or treats delivers no benefit and all of the harm.
The harms of screening
Screening is often imagined as pure benefit, but it carries real harms that must be weighed. False positives subject healthy people to anxiety, further testing, and sometimes invasive and risky procedures. Overdiagnosis detects disease that would never have caused symptoms in the person's lifetime, and because doctors cannot always tell harmless from dangerous disease, overdiagnosis leads to overtreatment, with its own side effects. These harms fall on many people to benefit a few, so a screen must clear a high bar before it is offered to a population.
Interval cases and program sensitivity
A program's real-world sensitivity differs from a test's laboratory sensitivity. Cases that arise between scheduled screens, called interval cases, are often the fast-growing, aggressive ones that appeared and advanced within a single screening interval. A program with few interval cases is catching disease effectively; one with many is missing exactly the cases that matter most. Monitoring interval cases is one way organized programs check whether their interval and test are actually protecting the population, rather than merely finding indolent disease.
Biases that make screening look better than it is
Evaluating screening is treacherous because two biases can make it appear to prolong life when it does not. Lead-time bias is the illusion of longer survival created simply by diagnosing disease earlier. The clock starts sooner, so measured survival looks longer even if death comes at the same moment it would have anyway.
Length-time bias is the tendency of screening to preferentially catch slow-growing, indolent cases, which spend longer in a detectable phase, making screen-detected disease look more survivable than it truly is. Fast, aggressive cases are more likely to surface between screens, so they are under-represented among screen-detected cases, flattering the program's apparent results.
Overdiagnosis is the extreme of length-time bias: cases so slow that they never would have surfaced at all. Counted as cured because they never progress, they inflate survival statistics while representing disease that needed no treatment. Because of these traps, screening's real value must be judged by whether it lowers mortality, not by comparing the survival times of screened and unscreened patients.
Why mortality, not survival, is the right measure
The safeguard against all three biases is to compare disease-specific mortality between people randomized to be screened or not, following whole groups rather than only the diagnosed. Survival time from diagnosis is fooled by lead time, length time, and overdiagnosis, but the death rate in a randomized population is not, because randomization balances the fast and slow cases and the clock starts at randomization, not at diagnosis. This is why a screening program's evidence base rests on randomized trials with mortality endpoints, not on how long diagnosed patients appear to live.
Number needed to screen
A useful summary for a program is the number needed to screen, how many people must be screened over a period to prevent one death from the disease. It is often large, because most screened people never had the disease and some who did would have done well anyway. Setting that number against the harms, the false positives, procedures, and overdiagnosis generated along the way, is how the net value of a screening program is honestly judged, rather than by counting cases found.
Informed choice
Because screening carries harms as well as benefits, sound programs support informed choice rather than blanket exhortation. A person deciding whether to be screened is weighing a chance of earlier, more treatable disease against the chances of a false alarm, an unnecessary procedure, and overdiagnosis. Presenting those trade-offs honestly, ideally in natural frequencies, respects the person and reflects that for several screens the balance of benefit and harm is genuinely close rather than obvious.
The screening paradox
Screening produces a paradox that mirrors the prevention paradox from the opening lesson. The individuals most likely to benefit, those with fast, deadly, early-stage disease, are a small minority, while the majority screened either never had the disease, had harmless disease, or would have fared the same without screening. A program can therefore be worthwhile for a population yet offer little expected benefit to most individuals in it, which is exactly why the choice to screen is best framed as a careful weighing of a modest average benefit against a real average harm.
Common misconceptions
The first misconception is that more screening is always better; because harms scale with the number screened, a poorly targeted program can do net harm. The second is reading a positive as a diagnosis, when at low prevalence most positives are false. The third is trusting improved survival among the screened as proof of benefit, when lead-time and length-time biases can manufacture that improvement with no lives saved. Each error is corrected by the same discipline: judge screening by randomized mortality evidence and by predictive values in the group actually tested.
Try it
A test with sensitivity 0.90 and specificity 0.95 screens 100,000 people. In population A the disease prevalence is 2 percent; in population B it is 20 percent. Compute the positive predictive value in each population and explain, in one sentence, why they differ.
Worked answer: in population A, diseased = 2,000 so TP = 0.90 × 2,000 = 1,800, and healthy = 98,000 so FP = 0.05 × 98,000 = 4,900, giving PPV = 1,800 / (1,800 + 4,900) = 1,800 / 6,700 = 0.269, about 27 percent. In population B, diseased = 20,000 so TP = 18,000, and healthy = 80,000 so FP = 4,000, giving PPV = 18,000 / (18,000 + 4,000) = 18,000 / 22,000 = 0.818, about 82 percent. The identical test yields 27 versus 82 percent because lower prevalence produces relatively more false positives, dragging the predictive value down.
Sources
- Wilson, J. M. G., & Jungner, G. (1968). Principles and practice of screening for disease (Public Health Papers No. 34). World Health Organization. find source ↗
- Andermann, A., Blancquaert, I., Beauchamp, S., & Dery, V. (2008). Revisiting Wilson and Jungner in the genomic age: A review of screening criteria over the past 40 years. Bulletin of the World Health Organization, 86(4), 317-319. pmc.ncbi.nlm.nih.gov
- Hall, K. (2020). Max Wilson and the principles and practice of screening for disease. International Journal of Neonatal Screening, 6(1), 15. pmc.ncbi.nlm.nih.gov
- Givler, D. N., & Givler, A. (2023). Health screening. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Kramer, B. S., & Croswell, J. M. (2009). Cancer screening: The clash of science and intuition. Annual Review of Medicine, 60, 125-137. pubmed.ncbi.nlm.nih.gov
- Welch, H. G., & Black, W. C. (2010). Overdiagnosis in cancer. Journal of the National Cancer Institute, 102(9), 605-613. pubmed.ncbi.nlm.nih.gov
- Gigerenzer, G., & Edwards, A. (2003). Simple tools for understanding risks: From innumeracy to insight. BMJ, 327(7417), 741-744. pubmed.ncbi.nlm.nih.gov
- Key terms
- Predictive values depend on prevalence
- PPV rises and NPV falls as disease becomes more common, though sensitivity and specificity stay fixed.
- False positives swamp true positives
- In rare disease, the many healthy people generate enough false alarms to lower PPV sharply.
- Targeted screening
- Applying a screening test to higher-risk (higher-prevalence) groups to raise its positive predictive value.
- Wilson and Jungner criteria
- Conditions a worthwhile screening program should meet regarding the disease, the test, and available treatment.
- Lead-time bias
- Apparent survival gain from diagnosing disease earlier without changing the time of death.
- Length-time bias
- Screening's tendency to catch slow-growing cases, overstating how survivable screen-detected disease is.
Module 6: Infectious Disease Epidemiology
The transmission dynamics of infectious disease, the reproduction number and herd immunity, and the systematic steps of an outbreak investigation.
Transmission, R-naught, and Herd Immunity
- Describe chains of transmission and the components needed for spread.
- Define the basic reproduction number and interpret its value.
- Compute the herd immunity threshold from the reproduction number.
Infectious disease epidemiology adds a feature no chronic disease has: one case can create another. Because cases are not independent, the mathematics of spread is distinctive, and controlling it means breaking chains of transmission rather than only reducing individual risk. A single infection can seed an epidemic, so the field studies not just who is at risk but how the pathogen moves from person to person.
That dependence between cases is what makes prevention so powerful in infectious disease: protecting one person can protect others down the chain, and stopping transmission at any link can halt an outbreak entirely. The quantities in this lesson, above all the reproduction number, capture that chain behavior in numbers.
The chain of infection
Transmission requires a linked chain: an agent, a reservoir where it lives (humans, animals, or the environment), a portal of exit, a mode of transmission, a portal of entry, and a susceptible host. Every control measure works by breaking one link: killing the agent through disinfection, removing the reservoir, blocking transmission with masks, mosquito nets, or hand hygiene, or protecting the susceptible host through vaccination. A person who harbors and spreads a pathogen without symptoms is an asymptomatic carrier, a link that makes some diseases especially hard to stop.
Timing along the chain matters too. The incubation period is the interval from infection to symptoms, while the latent period is the interval from infection to becoming infectious, and the two need not coincide. When people become infectious before they feel ill, or without ever feeling ill, isolating only the visibly sick cannot break the chain, which is why pre-symptomatic and asymptomatic transmission so complicates control.
Modes of transmission
Pathogens travel by characteristic routes, and each suggests its own defense. Direct contact passes an agent skin to skin or through bodily fluids. Droplet spread carries larger respiratory particles over short distances, while airborne spread suspends smaller particles that travel farther and linger. Vehicle transmission uses food, water, or blood, and vector-borne transmission relies on mosquitoes, ticks, or other carriers. Snow's cholera was vehicle-borne through water; malaria is vector-borne. Naming the mode points directly at the link to break, from safe water to bed nets to ventilation.
The basic reproduction number
The single most important quantity in infectious disease epidemiology is the basic reproduction number, R-naught (R₀): the average number of new infections produced by one infectious person in a completely susceptible population. Its value decides whether an epidemic grows or dies out:
- R₀ > 1: each case produces more than one new case, so the disease spreads and an epidemic can occur.
- R₀ = 1: each case replaces itself; the disease is endemic and stable.
- R₀ < 1: cases fail to replace themselves and the outbreak fades.
Measles is among the most contagious diseases known, with R₀ often cited around 12 to 18 in textbooks; seasonal influenza is much lower, commonly given as roughly 1 to 2. These canonical ranges illustrate how widely transmissibility varies. As a population gains immunity, the effective reproduction number falls below R₀, and control is achieved once it drops below 1.
The effective reproduction number
R₀ describes a fully susceptible population, a condition that rarely lasts. The effective reproduction number, sometimes written R or Rt, is the average number of new infections per case given the immunity and behavior present at a moment. Vaccination, prior infection, distancing, and masking all push it below R₀. The goal of every control effort is to hold the effective number under 1, at which point each case fails to replace itself and the epidemic recedes. Watching this number rise or fall is how the trajectory of an outbreak is read.
The SIR model in words
A simple way to picture an epidemic divides the population into three compartments: susceptible, infectious, and recovered, the classic SIR model. People flow from susceptible to infectious as they catch the disease, and from infectious to recovered as they clear it and gain immunity. Early on, with almost everyone susceptible, each case produces close to R₀ new infections and cases rise. As the susceptible pool shrinks, the effective reproduction number falls, the epidemic peaks when it passes 1, and then declines. Some people remain untouched at the end, because transmission stops before every susceptible is reached.
Generation time and epidemic growth
How fast an epidemic grows depends not only on R₀ but on the generation time, the typical interval between one case and the next it infects, often estimated by the observed serial interval between symptom onsets. A disease with a high reproduction number and a short generation time explodes, while the same reproduction number with a long generation time spreads slowly. This is why two pathogens with similar R₀ values can produce very different epidemic curves, and why speed of response is judged against the generation time.
Superspreading and averages
The reproduction number is an average, and averages can hide important structure. For many pathogens, most infected people transmit to no one, while a few, in the right settings, infect many, a pattern called superspreading. Two diseases with the same R₀ can behave differently if one spreads evenly and the other in rare large bursts, since bursty transmission is more prone both to fizzling out by chance and to exploding in the wrong venue. Control that targets the settings where superspreading occurs can be unusually efficient.
Contagiousness versus deadliness
Contagiousness and lethality are different axes, and confusing them is common. The basic reproduction number measures how readily a pathogen spreads, not how dangerous it is to the infected. A highly contagious disease can be mild, and a poorly transmitted one can be deadly. Case fatality, from the frequency lesson, measures deadliness once infected, while R₀ measures spread. A full picture of an infectious threat needs both numbers, since the effort to contain spread and the clinical severity of illness are driven by different quantities.
Herd immunity
Herd immunity is the indirect protection that unimmunized people receive when enough of the surrounding population is immune that sustained transmission cannot occur. The pathogen simply cannot find enough susceptible hosts to keep chains going. The fraction that must be immune, the herd immunity threshold (Hc), follows directly from R₀:
Hc = 1 - (1 / R₀)
The more contagious the disease, the greater the share that must be immune. Worked values make the pattern clear: at R₀ = 2, Hc = 1 - 1/2 = 50 percent; at R₀ = 3, Hc = 1 - 1/3 = 66.7 percent; at R₀ = 5, Hc = 1 - 1/5 = 80 percent; at R₀ = 10, Hc = 1 - 1/10 = 90 percent. The threshold rises steeply as transmissibility climbs, and then flattens toward, but never reaches, 100 percent.
| Disease example | R₀ | Hc = 1 - 1/R₀ |
|---|---|---|
| Lower-contagion illness | 2 | 1 - 1/2 = 50 percent |
| Moderate-contagion illness | 4 | 1 - 1/4 = 75 percent |
| Measles (highly contagious) | 16 | 1 - 1/16 = 93.75 percent |
This is why measles requires roughly 95 percent immunity to prevent outbreaks, while a less contagious disease needs far less. It also explains why coverage that slips even a few points can let a highly contagious disease return: once the immune fraction falls below Hc, the effective reproduction number rises above 1 and transmission resumes among the susceptible.
From threshold to vaccine coverage
The threshold is a fraction that must be immune, but vaccines are not perfectly protective, so the fraction that must be vaccinated is higher. If a vaccine has efficacy VE, the required coverage is roughly the threshold divided by the efficacy, Hc / VE. For measles, with Hc near 0.94 and a vaccine efficacy around 0.97, the coverage needed is about 0.94 / 0.97 = 0.97, close to 95 percent. This calculation explains why highly contagious diseases demand both an excellent vaccine and very high uptake to be held in check.
Herd immunity through infection carries a toll
Immunity can be reached in two ways, through vaccination or through infection, and the two are not equivalent in cost. Reaching a high herd immunity threshold by letting a pathogen spread means most of the population must first be infected, paying whatever case fatality and long-term harm the disease inflicts along the way. Vaccination reaches the same threshold without that toll, which is the central public health argument for immunization. The threshold formula says how many must be immune; it does not say the route to immunity is free.
Who herd immunity protects
Herd immunity matters most for those who cannot be vaccinated or do not respond: newborns too young for a vaccine, people with immune-suppressing conditions or treatments, and the small fraction in whom a vaccine fails to take. These people depend on the immunity of everyone around them, which gives vaccination a communal dimension beyond individual protection. A decision to vaccinate is partly a decision about the safety of one's neighbors, a point at the heart of the ethics of immunization policy.
Reservoirs and why some diseases can be eradicated
Whether a disease can be eliminated depends heavily on its reservoir. A pathogen that lives only in humans, with no animal or environmental refuge, can in principle be driven to extinction by interrupting human-to-human transmission, as smallpox was, declared eradicated in 1980 after a global vaccination campaign. A pathogen with an animal reservoir, by contrast, can always spill back into people, so eradication is far harder. The chain of infection is not only a control framework but a map of a disease's vulnerabilities.
Breaking the chain in practice
Every classic control measure maps onto a link in the chain. Isolating cases and quarantining exposed contacts cut the portal of exit and the route to new hosts. Safe water, food safety, and sanitation close vehicle transmission. Vector control, from bed nets to reducing breeding sites, targets the carriers. Vaccination protects the susceptible host and, at scale, delivers herd immunity to the whole community. Reading an outbreak as a chain turns a vague sense of threat into a checklist of specific, breakable links.
Common misconceptions
One misconception treats R₀ as a fixed property of a pathogen, when it also depends on contact patterns and setting, so the same microbe can have different values in different populations. Another equates herd immunity with everyone being immune, when the threshold is usually well below 100 percent. A third assumes a low reproduction number means a mild disease, confusing spread with severity. Each error dissolves once R₀, the effective reproduction number, and case fatality are kept as distinct quantities measuring distinct things.
Try it
A pathogen has a basic reproduction number of 8. Compute the herd immunity threshold. Then, if the available vaccine is 90 percent effective, estimate the fraction of the population that must be vaccinated to reach that threshold, and state in plain words what R₀ = 8 means.
Worked answer: Hc = 1 - 1/8 = 1 - 0.125 = 0.875, so 87.5 percent must be immune. With a vaccine efficacy of 0.90, the required coverage is about Hc / VE = 0.875 / 0.90 = 0.972, roughly 97 percent vaccinated. An R₀ of 8 means that one infectious person would, on average, infect eight others in a fully susceptible population, so the outbreak would grow quickly until immunity built up.
Sources
- Delamater, P. L., Street, E. J., Leslie, T. F., Yang, Y. T., & Jacobsen, K. H. (2019). Complexity of the basic reproduction number (R0). Emerging Infectious Diseases, 25(1), 1-4. pmc.ncbi.nlm.nih.gov
- Fine, P., Eames, K., & Heymann, D. L. (2011). Herd immunity: A rough guide. Clinical Infectious Diseases, 52(7), 911-916. pubmed.ncbi.nlm.nih.gov
- Fine, P. E. M. (1993). Herd immunity: History, theory, practice. Epidemiologic Reviews, 15(2), 265-302. pubmed.ncbi.nlm.nih.gov
- Anderson, R. M., & May, R. M. (1985). Vaccination and herd immunity to infectious diseases. Nature, 318(6044), 323-329. pubmed.ncbi.nlm.nih.gov
- Guerra, F. M., Bolotin, S., Lim, G., Heffernan, J., Deeks, S. L., Li, Y., & Crowcroft, N. S. (2017). The basic reproduction number (R0) of measles: A systematic review. The Lancet Infectious Diseases, 17(12), e420-e428. pubmed.ncbi.nlm.nih.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 1, Section 10: Chain of infection. CDC Self-Study Course SS1978. archive.cdc.gov
- World Health Organization. (n.d.). Smallpox. WHO Health Topics. who.int
- Key terms
- Chain of infection
- The linked sequence (agent, reservoir, exit, transmission, entry, susceptible host) required for spread.
- Mode of transmission
- How a pathogen passes between hosts: direct contact, droplet, airborne, vehicle, or vector-borne.
- Basic reproduction number (R-naught)
- The average new infections from one case in a fully susceptible population; above 1, disease spreads.
- Effective reproduction number
- The average new infections per case given current immunity; control is reached when it falls below 1.
- Herd immunity
- Indirect protection of the susceptible when enough of the population is immune to stop sustained transmission.
- Herd immunity threshold
- The immune fraction needed to halt spread, equal to 1 minus 1 over R-naught.
Outbreak Investigation
- Define an outbreak and epidemic and interpret an epidemic curve.
- List the standard steps of an outbreak investigation.
- Use attack rates to identify a likely source.
When disease appears in excess of what is expected, public health mounts an outbreak investigation, epidemiology at its most urgent and practical. An outbreak or epidemic is the occurrence of cases clearly in excess of normal expectancy in an area; a pandemic is an epidemic spread over many countries or continents; and an endemic disease is one at its usual, baseline level in a population. The whole toolkit of earlier lessons, counting, comparison, and causal reasoning, is deployed here against the clock.
Because every one of these terms is defined against what is expected, an investigation cannot even begin without a sense of the baseline. A dozen cases may be an ordinary week for a common illness and a genuine emergency for a rare one, so the first question is always whether the observed number truly exceeds the usual.
Endemic, epidemic, pandemic
The vocabulary of levels is worth fixing precisely. A disease at its usual baseline in a population is endemic. A rise clearly above that baseline in an area is an outbreak or epidemic, the two words differing mainly in scale and tone. An epidemic crossing many countries or continents is a pandemic. Isolated, unlinked cases are sporadic. Because every term is relative to what is expected, none can be applied without first knowing the baseline, which is why surveillance data underpin the very declaration of an outbreak.
Confirming the outbreak
The first task is to make sure there is an outbreak at all. An apparent rise can be an artifact of increased testing, a new laboratory method, better reporting, or a coding change rather than a true increase in disease. Investigators verify the diagnosis, often with laboratory confirmation, and compare the current count against the expected baseline from surveillance. Only once excess disease is real and the diagnosis is sound does the fuller investigation proceed, though in a fast-moving, dangerous situation these early steps and control may begin together.
The steps of an investigation
Investigations follow a well-worn sequence, though steps often overlap and iterate:
- Confirm the outbreak and verify the diagnosis - is there truly excess disease, and is the diagnosis correct?
- Define a case - write an explicit case definition (clinical features plus limits of person, place, and time) so cases are counted consistently.
- Find cases and count them - conduct systematic case-finding and record each case.
- Describe by person, place, and time - the descriptive epidemiology, including the epidemic curve.
- Develop hypotheses about the source and mode of transmission.
- Test hypotheses analytically, typically with a case-control or retrospective cohort study using attack rates.
- Implement control measures - often begun early, as soon as a plausible source is suspected.
- Communicate findings to stakeholders and the public.
The numbering suggests an order, but real investigations loop back constantly. New cases sharpen the case definition, a fresh hypothesis sends investigators back to find more cases, and control measures start the moment a plausible source is suspected rather than waiting for a finished analysis. The list is a scaffold for thinking, not a rigid recipe.
The case definition
The case definition is the backbone of consistent counting. It combines clinical criteria, such as specific symptoms or a laboratory result, with limits of person, place, and time that fence the outbreak. Definitions are often tiered into confirmed, probable, and suspected cases by how much evidence each has. Early in an investigation a broad, sensitive definition casts a wide net so no cases are missed; later a narrower, more specific definition sharpens the analysis. Changing the definition changes the counts, so the definition must always travel with the numbers.
Case-finding and the outbreak's true size
Finding cases is more than waiting for reports. Investigators actively search, alerting clinicians and laboratories, reviewing records, and sometimes surveying the exposed group directly, because the cases that come forward on their own are only part of the picture. A line list records each case with its key facts, one row per person, becoming the working dataset from which the epidemic curve and attack rates are built. Undercounting mild or resolved cases can distort every later step, so systematic case-finding is worth the effort it takes.
Descriptive epidemiology, first
Before any hypothesis is tested, the cases are described by person, place, and time, exactly the descriptive core from the study-design module. Who is affected by age, sex, and other traits; where they cluster; and when their illness began. This portrait often suggests the source on its own, as a cluster around one venue or a sharp spike after one event does, and it always shapes the hypotheses the analytic study will then test. Jumping straight to a favorite suspect, before describing the outbreak, is a common and costly shortcut.
Reading the epidemic curve
The epidemic curve, or epi curve, is a histogram of case counts by time of onset, and its shape is deeply informative. A point-source outbreak, in which everyone is exposed at one time, such as a contaminated meal at an event, produces a single sharp peak whose width reflects the range of incubation periods. The rise is steep and the fall more gradual, and the whole curve is compressed into roughly one incubation period's span.
A continuous common-source outbreak, such as an ongoing contaminated water supply, shows a plateau that persists as long as the source operates. A propagated outbreak, spread person to person, shows a series of progressively taller peaks, each about one incubation period apart, as each generation infects the next. Reading which of these shapes a curve resembles immediately narrows the hypotheses about how the disease is spreading.
What the curve reveals
Beyond its shape, the curve carries quantitative clues. In a point-source outbreak, the spread of onset times reflects the disease's incubation period, and the interval from a known exposure to the peak estimates the median incubation. Working backward, if the incubation period is known, the position of the peak can pinpoint when exposure likely occurred, helping identify the responsible event or meal. The epidemic curve is therefore both a picture of the outbreak and a clock for reconstructing it.
Attack rates point to the source
The key analytic tool in a foodborne outbreak is the attack rate: the proportion of an exposed group that becomes ill, essentially an incidence proportion for the outbreak. By computing the attack rate among those who did and did not eat each food, investigators find the item with both a high attack rate in the exposed and a large difference from the unexposed.
Worked example: of 150 guests who ate the potato salad, 45 fell ill (attack rate 45 / 150 = 30 percent); of 100 who did not eat it, 5 fell ill (attack rate 5 / 100 = 5 percent). The relative risk is 0.30 / 0.05 = 6.0, strongly implicating the potato salad. The food with the largest attack-rate difference, especially when most cases ate it and few non-eaters fell ill, is the prime suspect, and removing it should end the outbreak, satisfying the Bradford Hill experiment viewpoint in real time.
A worked attack-rate comparison
Investigations compute an attack rate for each suspect food among eaters and non-eaters, then look for the largest, most consistent difference. Suppose chicken shows an attack rate of 40 percent among eaters and 38 percent among non-eaters, almost no difference, while custard shows 65 percent among eaters and 6 percent among non-eaters. The custard's relative risk is about 65 / 6, near 11, and nearly all cases ate it while few non-eaters fell ill. The food with a high attack rate in eaters, a low rate in non-eaters, and most cases exposed is the prime suspect.
Cohort or case-control at an outbreak
The design used to test food hypotheses depends on the setting. When the exposed population is defined and reachable, as at a wedding with a known guest list, a retrospective cohort is natural: interview everyone, compute attack rates for each food, and compare. When there is no roster, as in an outbreak scattered across a city, a case-control study is used instead, comparing the food histories of cases and community controls with odds ratios. The choice mirrors the design lessons from earlier, applied under time pressure.
Confirming and controlling
Epidemiologic suspicion is strengthened by laboratory and environmental work: culturing the pathogen from patients and, ideally, from the suspected food or its source, and tracing that source back through the supply chain. Control acts on whichever link is reachable, recalling a product, closing a kitchen, correcting a water system, or isolating cases. When the outbreak curve falls after the suspected source is removed, that decline is itself powerful confirmatory evidence that the right source was found.
Why speed matters
Outbreak work is unusual in epidemiology for its urgency. Cases are still occurring while the study is underway, so control measures are often begun on a plausible hypothesis before the analysis is complete, then refined as evidence firms up. Removing a suspected source and watching the epidemic curve fall doubles as confirmation. The investigator constantly weighs the cost of acting on an incomplete picture against the cost of letting an outbreak run while certainty is pursued, and usually errs toward protecting the public early.
Communicating findings
An investigation ends by communicating what was learned, to health authorities, to affected institutions, and often to the public. Good communication is timely, clear, and actionable, telling people what happened, what is being done, and what they should do. Because outbreaks generate fear and rumor, a trusted, transparent voice is itself a control measure, guiding behavior and preventing the panic that misinformation breeds. The final written report also feeds surveillance and prevents the next outbreak by recording what went wrong.
Common misconceptions
The first misconception is that a rise in cases always means an outbreak, when better testing or reporting can mimic one without any true increase. The second is implicating a food merely because many cases ate it, when the telling sign is a large gap between the attack rates in eaters and non-eaters, not a high count alone. The third is treating the steps as strictly sequential, when investigations iterate and control need not wait for a completed analysis. Each error is corrected by grounding judgment in baselines and comparisons.
Try it
At a catered lunch attended by 200 people, investigators find that of the 120 who ate the shrimp, 84 became ill, while of the 80 who did not eat shrimp, 8 became ill. Compute the attack rate in each group and the relative risk, and state whether the shrimp is implicated and why.
Worked answer: attack rate among shrimp eaters = 84 / 120 = 0.70, or 70 percent; among non-eaters = 8 / 80 = 0.10, or 10 percent. The relative risk is 0.70 / 0.10 = 7.0. The shrimp is strongly implicated: it shows a high attack rate among eaters, a low rate among non-eaters, and a large difference between them, exactly the pattern that points to an outbreak's source.
Sources
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 6, Section 2: Steps of an outbreak investigation. CDC Self-Study Course SS1978. archive.cdc.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 6, Section 1: Introduction to investigating an outbreak. CDC Self-Study Course SS1978. archive.cdc.gov
- Reingold, A. L. (1998). Outbreak investigations - A perspective. Emerging Infectious Diseases, 4(1), 21-27. pmc.ncbi.nlm.nih.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 1, Section 11: Epidemic disease occurrence. CDC Self-Study Course SS1978. archive.cdc.gov
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 3, Section 2: Morbidity frequency measures. CDC Self-Study Course SS1978. archive.cdc.gov
- Fraser, D. W., Tsai, T. R., Orenstein, W., Parkin, W. E., Beecham, H. J., Sharrar, R. G., ... Broome, C. V. (1977). Legionnaires' disease: Description of an epidemic of pneumonia. New England Journal of Medicine, 297(22), 1189-1197. pubmed.ncbi.nlm.nih.gov
- Tenny, S., Kerndt, C. C., & Hoffman, M. R. (2023). Case control studies. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Key terms
- Outbreak / Epidemic
- Occurrence of cases clearly in excess of what is normally expected in an area.
- Endemic vs pandemic
- The usual baseline level of a disease versus an epidemic spanning many countries or continents.
- Case definition
- Explicit clinical and person-place-time criteria used to count cases consistently in an investigation.
- Epidemic curve
- A histogram of cases by time of onset whose shape suggests point-source, continuous, or propagated spread.
- Attack rate
- The proportion of an exposed group that becomes ill; an incidence proportion used to find an outbreak's source.
- Point-source outbreak
- An outbreak from a single common exposure at one time, producing a single sharp epidemic-curve peak.
Module 7: Chronic Disease Epidemiology and Health Policy
How epidemiology adapts to slow, multifactorial chronic diseases, and how its evidence is turned into surveillance, guidelines, and public health policy.
Chronic Disease Epidemiology
- Contrast chronic and infectious disease epidemiology.
- Explain multifactorial causation and the web of causation.
- Describe the role of risk factors and long-term cohort studies.
As infectious diseases were tamed in wealthy nations, the leading causes of death shifted to chronic diseases such as heart disease, cancer, stroke, diabetes, and chronic lung disease. This long shift is called the epidemiologic transition. It rewards a different epidemiologic style, because the causes of chronic disease are slow, multiple, and entangled with behavior, biology, and society. The measures and designs from earlier modules still apply, but they must be stretched across decades rather than the days or weeks of an outbreak.
This lesson adapts the whole toolkit to that longer timescale. We examine how chronic disease differs from infection, how many causes combine in a web, how the smoking and lung cancer story became the paradigm for the field, how relative and absolute risk send very different prevention messages, and how population and high-risk strategies each prevent disease. Throughout, the central habit stays the same: compare groups, and ask how much disease an exposure truly explains.
How chronic disease differs
Infectious disease often has a single necessary agent and a short incubation of days or weeks. Chronic disease typically has no single necessary cause, a long latency of years or decades between exposure and illness, and a course measured in years rather than days. Koch's postulates, which tie one microbe to one disease, simply do not apply. There is no organism to culture, no single exposure whose removal reliably ends the disease, and no sharp onset to mark on an epidemic curve.
These features reshape method. Long latency means a study must follow people for many years, or reach far back into their histories, before outcomes appear. The absence of a necessary agent means causation is a matter of probability and degree, not presence or absence. And because the same disease can arise many ways, prevention has many possible entry points rather than one chokepoint. Chronic disease epidemiology therefore trades the urgency of the outbreak for patience, large samples, and careful control of competing explanations.
Risk factors are not the same as causes
Chronic disease epidemiology speaks of risk factors: characteristics or exposures that raise the probability of disease without being strictly necessary or sufficient. Smoking, high blood pressure, elevated cholesterol, obesity, physical inactivity, and diet are risk factors, not agents. Most people with a given risk factor never develop the disease, and some people without it do. A risk factor shifts the odds; it does not seal a fate. The term itself entered medicine through the Framingham Heart Study, the long cohort that began in 1948 and gave cardiovascular epidemiology much of its vocabulary.
It helps to separate three ideas that beginners often merge. A marker merely predicts disease without causing it, like a gray hair predicting a heart attack because both track age. A cause actually produces disease, so that removing it lowers risk. A modifiable risk factor is a cause we can change. Prevention depends on modifiable causes, not on markers, which is why distinguishing them is not academic hair-splitting but the difference between a policy that works and one that wastes effort.
Multifactorial causation and the web
Because many factors act together, epidemiologists picture a web of causation: an interconnected network in which distal causes, such as poverty, education, and the built environment, shape proximal ones, such as diet, smoking, and blood pressure, which in turn combine to produce disease. Pulling any single strand rarely collapses the web, but it does loosen it. The image captures why chronic disease has no magic bullet, and why interventions at many different points, from farm policy to statins, can each move the same outcome.
A sharper companion idea is the sufficient-component cause model, often drawn as a pie. A disease occurs when a full set of component causes, one complete pie, is assembled. Different people may complete different pies and reach the same disease by different routes. A component present in every pie is a necessary cause; any single component, once removed, prevents the cases whose pie contained it. The model explains both why one exposure raises risk without guaranteeing disease and why removing it still prevents some, though not all, cases.
The smoking and lung cancer paradigm
The defining achievement of chronic disease epidemiology is the demonstration that smoking causes lung cancer, and it remains the template for how such claims are built. In the early 1950s Richard Doll and Austin Bradford Hill studied the question first with a case-control comparison of lung cancer patients and controls, then with the British Doctors Study, a prospective cohort begun in 1951 that followed tens of thousands of physicians for decades. The cohort watched exposure precede disease, the design feature a case-control study cannot supply directly.
The evidence pointed one way on several fronts at once. The association was strong, with heavy smokers dying of lung cancer at many times the rate of non-smokers, on the order of ten to twenty times in the heaviest users. It showed a clear biological gradient: the more one smoked, the higher the risk. It was consistent across studies and populations, temporally correct because the cohort measured smoking before cancer appeared, and coherent with laboratory and pathological findings. No single study proved it; the weight of many did.
This is exactly the reasoning Module 4 named after Bradford Hill, and it is no coincidence: Hill articulated his viewpoints in 1965 partly to formalize how the smoking case had been made. Chronic disease causation could not be settled by a trial that deliberately assigned people to smoke, which would be unethical and impossible. It had to be assembled from observational cohorts and case-control studies, weighed against the Hill considerations. The smoking story shows that observational evidence, when strong, graded, consistent, and coherent, can support a firm causal claim.
The tools: risk factors and long cohorts
The signature instrument of the field is the long-term prospective cohort study, which enrolls large groups of healthy people, measures many exposures at baseline, and follows them for years or decades while outcomes accumulate. Because it records exposure before disease, it establishes temporality and can study many outcomes at once. The Framingham Heart Study did this for cardiovascular disease from 1948 onward; the EPIC study, the European Prospective Investigation into Cancer and Nutrition, is an established example built on the same logic for diet and cancer.
Such cohorts are expensive and slow, and they cannot randomize the exposures that matter most, so confounding is a constant worry addressed by careful measurement and statistical adjustment. Their strength is that they can capture the long, quiet latency of chronic disease that no short study could see. When a harmful exposure cannot ethically be assigned, the prospective cohort is the strongest design available, and the accumulated weight of many cohorts is what carries a chronic disease claim from association toward cause.
Relative risk and absolute risk in prevention messaging
A recurring confusion in chronic disease prevention is the gap between relative risk and absolute risk. Relative risk says how many times more likely disease is in the exposed; the absolute risk difference says how many extra cases per person actually occur. The two can point in very different practical directions, because a large relative risk applied to a tiny baseline still yields few extra cases, while a modest relative risk applied to a common disease can yield many.
Work an example. Suppose an exposure doubles risk, so the relative risk is 2.0 in both scenarios. For a common disease with a baseline 10-year risk of 20 per 1,000, the exposed risk is 40 per 1,000, an absolute difference of 20 per 1,000; roughly 50 people would need the exposure removed to prevent one case. For a rare disease with baseline 1 per 100,000, the exposed risk is 2 per 100,000, an absolute difference of just 1 per 100,000, and preventing a single case would require changing exposure in about 100,000 people.
Same relative risk, very different public health payoff. A headline that an exposure "doubles your risk" means little without the baseline, yet baselines are routinely omitted. This is why prevention is planned in absolute terms and in cases prevented, not in ratios. Relative risk is the better measure of how strongly an exposure and disease are linked; absolute risk is the better measure of how much a prevention effort will actually accomplish. Good messaging reports both.
Levels of prevention for chronic disease
The three levels of prevention from Module 1 map cleanly onto the long natural history of chronic disease. Primary prevention acts before disease begins, during susceptibility: reducing smoking, improving diet, or lowering blood pressure across a population. Secondary prevention detects disease early in its silent phase, as mammography or colonoscopy aim to do, so that treatment starts sooner. Tertiary prevention limits disability once disease is established, through cardiac rehabilitation, careful diabetes management, and the prevention of complications.
Some authors add primordial prevention, which acts earliest of all by keeping risk factors from ever arising, for instance by building neighborhoods that make walking easy or a food supply low in harmful fats. Because chronic disease is largely driven by modifiable behaviors and conditions, it is in principle highly preventable, and the earlier levels tend to yield the largest population gains. The long latency that makes chronic disease hard to study is also the long window in which prevention can act.
Two prevention strategies: high-risk and population
Geoffrey Rose distinguished two ways to lower a population's disease burden. The high-risk strategy finds the individuals at greatest risk and intervenes on them; it is efficient per person and appeals to clinicians. The population strategy shifts the entire risk distribution slightly toward health, for example by lowering salt across the food supply; each person gains little, but the whole curve moves. Rose argued that because most cases arise from the many people at modest risk, the population strategy often prevents more disease overall.
Put numbers on it. Imagine a town of 100,000 adults and a 10-year stroke risk that rises with blood pressure. Say 5,000 people have very high pressure with a 10-year stroke risk of 0.04, expecting 200 strokes, while 95,000 have mild-to-moderate elevation with a risk of 0.008, expecting 760 strokes. The total is 960 strokes, and the large moderate-risk majority produces 760 of them, about 79 percent, even though each such person is individually at low risk.
Now compare strategies. A high-risk strategy that halves risk in the 5,000 highest-risk people prevents 5,000 times 0.02, or 100 strokes, leaving the 760 in the majority untouched. A population strategy that lowers everyone's risk by 20 percent prevents 0.20 times 960, or 192 strokes, nearly double the high-risk yield, and it reaches the majority where most cases actually are. Yet each moderate-risk person's 10-year risk falls only from 0.008 to 0.0064, a change of 1.6 per 1,000 they will never feel.
That last fact is the prevention paradox: a measure bringing large benefit to a population may offer little to each individual within it. It explains why population-wide measures can be at once the most effective public health tools and the hardest to motivate, since no single person experiences a dramatic personal gain. In practice the two strategies are complements, not rivals: treat the high-risk tail clinically while shifting the whole distribution through policy. Each covers what the other misses.
Life-course epidemiology
Life-course epidemiology studies how exposures across an entire lifetime, and even before birth, shape the risk of chronic disease decades later. It distinguishes a critical period, when an exposure has a lasting effect that later change cannot undo, from the accumulation of risk, in which insults add up over years, and from chains of risk, in which one exposure raises the odds of the next. Early disadvantage can thus echo into late-life disease through many linked steps.
A well-known illustration is the developmental-origins idea, the observation that low birth weight and poor early growth are associated with higher later rates of heart disease and diabetes, plausibly through early programming of metabolism. Whatever the mechanism in any single case, the framework reminds us that a chronic disease diagnosed at sixty may have roots reaching back sixty years. It widens the web of causation along the time axis and pushes prevention earlier, toward childhood and even the prenatal period.
Common misconceptions
Several errors recur. The first is treating a risk factor as a guarantee, when most exposed people never develop the disease and a risk factor only shifts probability. The second is confusing relative and absolute risk, so that "doubles the risk" alarms without the baseline that gives it meaning. The third is expecting a single cause or single cure for diseases that arise from a web of many contributing factors. Each mistake dissolves once causation is seen as multifactorial, probabilistic, and best measured in absolute cases prevented.
A fourth misconception is that because chronic disease cannot be studied by trials of harmful exposures, its causal claims must be weak. The smoking and lung cancer story refutes this directly: strong, graded, consistent, temporally correct, and coherent observational evidence, weighed by the Bradford Hill viewpoints, established causation without a single trial assigning anyone to smoke. Observational evidence is not second-class; it is simply the evidence that the questions of chronic disease allow, and it can be entirely conclusive.
Try it
A news headline reports that a dietary habit "doubles the risk" of a certain cancer. In the unexposed population, the 10-year risk of this cancer is 2 per 1,000. (a) What is the exposed group's risk and the relative risk? (b) What is the absolute risk difference? (c) Roughly how many people would need to change the habit to prevent one case over 10 years? (d) Explain why "doubles" can mislead the public.
Worked answer: (a) exposed risk = 2 times 2 per 1,000 = 4 per 1,000, and the relative risk = 4 / 2 = 2.0. (b) absolute risk difference = 4 - 2 = 2 per 1,000, or 0.002. (c) about 1 / 0.002 = 500 people must change the habit to prevent one case. (d) "doubles" is purely relative; against a low baseline of 2 per 1,000 the true extra burden is only 2 cases per 1,000, so the relative figure sounds far more alarming than the small absolute impact it represents.
Sources
- Omran, A. R. (2005). The epidemiologic transition: A theory of the epidemiology of population change. The Milbank Quarterly, 83(4), 731-757. (Original work published 1971). pmc.ncbi.nlm.nih.gov
- Rothman, K. J. (1976). Causes. American Journal of Epidemiology, 104(6), 587-592. pubmed.ncbi.nlm.nih.gov
- Kannel, W. B., Dawber, T. R., Kagan, A., Revotskie, N., & Stokes, J. (1961). Factors of risk in the development of coronary heart disease - Six-year follow-up experience: The Framingham Study. Annals of Internal Medicine, 55(1), 33-50. pubmed.ncbi.nlm.nih.gov
- Doll, R., Peto, R., Boreham, J., & Sutherland, I. (2004). Mortality in relation to smoking: 50 years' observations on male British doctors. BMJ, 328(7455), 1519. pubmed.ncbi.nlm.nih.gov
- Rose, G. (1985). Sick individuals and sick populations. International Journal of Epidemiology, 14(1), 32-38. pubmed.ncbi.nlm.nih.gov
- Ben-Shlomo, Y., & Kuh, D. (2002). A life course approach to chronic disease epidemiology: Conceptual models, empirical challenges and interdisciplinary perspectives. International Journal of Epidemiology, 31(2), 285-293. pubmed.ncbi.nlm.nih.gov
- Riboli, E., Hunt, K. J., Slimani, N., Ferrari, P., Norat, T., Fahey, M., ... Saracci, R. (2002). European Prospective Investigation into Cancer and Nutrition (EPIC): Study populations and data collection. Public Health Nutrition, 5(6B), 1113-1124. pubmed.ncbi.nlm.nih.gov
- Key terms
- Epidemiologic transition
- The shift in a population's leading causes of death from infectious to chronic diseases as it develops.
- Chronic disease
- A long-lasting condition, typically with no single cause and long latency, such as heart disease or diabetes.
- Risk factor
- A characteristic or exposure that raises disease probability without being necessary or sufficient.
- Web of causation
- A network model in which many interconnected distal and proximal factors combine to cause disease.
- Sufficient-component cause model
- The idea that disease occurs when a sufficient set of component causes is completed, reachable by different combinations.
- Long-term prospective cohort
- A study following large groups for decades, the signature tool for identifying chronic disease risk factors.
Surveillance, Evidence, and Health Policy
- Explain the purpose and types of public health surveillance.
- Trace how epidemiologic evidence becomes guidelines and policy.
- Describe the ethical principles and trade-offs in public health decisions.
Epidemiology exists to be used. Its final purpose is action: monitoring the health of populations and turning evidence into policies that prevent disease. This closing lesson connects the science of the previous modules to the system that applies it, following the path from raw surveillance data, through the weighing of evidence, to the laws, programs, and guidelines a society actually adopts.
Three questions organize the lesson. How do we keep watch on the health of a population? How does a mass of studies become a trustworthy recommendation? And how should a society balance the good of the many against the freedom of the individual when it acts? Each question has its own methods and its own hard cases, and each draws on the measures and designs built earlier in the course.
Public health surveillance
Surveillance is the ongoing, systematic collection, analysis, and interpretation of health data, tied to the timely dissemination of that information to those who can act. The classic phrase is "information for action," and the last word matters: data gathered but never used to guide a decision is not surveillance. Its jobs are to detect outbreaks early, track trends over time, identify who is most affected, evaluate whether programs work, and point limited resources where they will do the most good.
Surveillance comes in two broad modes. Passive surveillance relies on routine reporting: clinicians and laboratories notify authorities of specified conditions as they encounter them. It is inexpensive and continuous but incomplete, because it depends on others remembering to report. Active surveillance reverses the flow: agencies reach out to hospitals, laboratories, and providers to seek cases directly. It is more complete and accurate but costly and labor-intensive, so it is usually mounted for a limited time, often during an outbreak or for a specific high-priority disease.
The measures from Module 1 are the currency of every surveillance report. Incidence tracks new cases, prevalence tracks existing burden, and derived rates such as mortality, deaths in a population, and case fatality, deaths among cases, summarize severity. A surveillance system that reports counts without denominators repeats the first error of the whole course, so rates, not raw counts, are what let officials compare places and periods and recognize when something has truly changed.
Notifiable diseases and how reports flow
Passive surveillance rests on the concept of notifiable, or reportable, diseases: conditions that law requires clinicians and laboratories to report to public health authorities. The list typically covers diseases that are dangerous, controllable, or of broad public concern, such as tuberculosis, measles, and certain foodborne and sexually transmitted infections. Making a disease notifiable creates the legal and routine channel through which cases flow upward from clinics and labs to local, regional, and national agencies.
The system's great weakness is under-reporting. Busy clinicians forget, mild cases never reach care, and laboratories vary in what they flag, so the reported count is usually a fraction of the true total. This does not make the data useless: as long as the reported fraction stays roughly stable, trends over time remain informative even when absolute counts are too low. Trouble arises when reporting itself changes, since a new law or a publicized outbreak can raise reports without any true rise in disease.
Sentinel and syndromic surveillance
Because universal reporting is incomplete and slow, agencies supplement it with focused systems. Sentinel surveillance recruits a selected network of reporting sites, such as a sample of clinics or laboratories, that report chosen conditions carefully and consistently. It sacrifices complete coverage for timeliness and data quality, and for many purposes a well-chosen sample tracks trends as faithfully as a costly attempt to capture every case. Influenza is often watched this way, through a set of volunteer practices reporting influenza-like illness.
Syndromic surveillance goes earlier still, monitoring pre-diagnostic signals such as emergency-department visits for particular symptoms, school absences, or over-the-counter medication sales. These signals are less specific than confirmed diagnoses, but they arrive days sooner, giving an early warning that something may be starting before laboratory confirmation catches up. The price of that speed is more false alarms, which is the recurring tension of all surveillance: the faster and more sensitive the signal, the noisier it tends to be.
What makes a surveillance system good
A surveillance system is judged on several attributes that trade off against one another. Sensitivity is the proportion of true cases the system detects. Timeliness is how quickly a case moves from occurrence to a report that can drive action. Positive predictive value is the proportion of reported cases that are genuinely cases. Others include simplicity, flexibility, representativeness, and acceptability, since a system too burdensome to use will not be used at all.
The central tension is between sensitivity and predictive value, governed by how broad the case definition is. Work an example. Suppose a city truly has 200 cases of a condition in a month. A passive system detects 120 of them and also generates 30 reports that turn out to be false. Its sensitivity is 120 / 200 = 60 percent. The total reports number 120 + 30 = 150, so its positive predictive value is 120 / 150 = 80 percent. One reported case in five is a false alarm, and two of every five true cases are missed.
Now broaden the case definition to catch more disease. Say sensitivity rises to 90 percent, detecting 180 of the 200 true cases, but the looser net also raises false reports to 60. Positive predictive value falls to 180 / (180 + 60) = 180 / 240 = 75 percent. Sensitivity climbed while predictive value dropped, and each false report costs investigation time, which can hurt timeliness. Designers choose where to sit on this curve by the stakes: early in a dangerous outbreak, missing cases is worse than chasing false alarms, so a sensitive definition is preferred.
From evidence to policy
Sound policy rests on a hierarchy of evidence that ranks study designs by how well they control bias and confounding. At the base sit expert opinion and case reports; above them the observational designs from Module 3, cross-sectional, then case-control, then cohort; higher still the randomized controlled trial; and at the very top the systematic review and meta-analysis. A single study, however good, can mislead; the summit of the hierarchy is not one study but a disciplined synthesis of them all.
A systematic review gathers every relevant study through an explicit, reproducible search, appraises each for quality, and summarizes them, reducing the cherry-picking that plagues informal reviews. A meta-analysis goes further and statistically pools the results into a single, more precise estimate, giving greater weight to larger and better studies. Because it combines many populations and settings, a well-conducted meta-analysis offers the most reliable answer epidemiology can provide to the question of whether an intervention works.
Evidence alone does not set policy; it is weighed against cost, feasibility, and values. Cost-effectiveness analysis makes the cost trade explicit by comparing the health gained, often measured in quality-adjusted life years (QALYs), against the money spent. Suppose program A costs 500,000 to gain 100 QALYs, or 5,000 per QALY, while program B costs 900,000 for the same 100 QALYs, or 9,000 per QALY. With a fixed budget, A buys more health per dollar, so it is preferred; the analysis turns a values question into a comparison limited resources can actually decide.
The precautionary principle
Sometimes a decision cannot wait for the top of the evidence hierarchy. The precautionary principle holds that when an activity threatens serious or irreversible harm, a lack of full scientific certainty should not by itself justify postponing protective action. It shifts the default toward caution when the downside is grave and hard to undo, as with a novel contaminant in drinking water or an emerging pathogen whose properties are not yet fully known.
The principle is powerful but double-edged. Acting on incomplete evidence risks costly or unnecessary intervention, while waiting for certainty risks preventable harm, and both errors are real. The mature position treats it not as a license to act on any fear, but as a call to weigh the seriousness and reversibility of the potential harm against the costs of acting early. It complements the evidence hierarchy rather than replacing it, governing what to do while the evidence is still being built.
When should a society screen? The Wilson and Jungner criteria
Screening entered this course in Module 5 as a test problem; as a policy it is a decision about whole populations, and the classic guide is the set of principles Wilson and Jungner set out for the World Health Organization in 1968. They ask, in effect, whether screening for a given condition will do more good than harm before a program is launched, and most modern screening guidelines are elaborations of their logic.
Their key conditions can be grouped. The disease should be an important health problem with a recognizable early or latent stage and a reasonably understood natural history. The test should be suitable, acceptable to the population, and accurate enough to keep false positives and negatives tolerable. The treatment should exist, work better when started early, and be backed by agreed policy on whom to treat. And the program as a whole should be cost-effective and continuous, its total cost weighed against its benefit.
The criteria explain why we do not simply screen for everything. A test for a disease with no effective early treatment offers detection without benefit while still causing real harms: false positives, anxiety, and the overdiagnosis of conditions that would never have caused trouble. Recalling the predictive-value lesson, screening a low-prevalence population with an imperfect test yields mostly false positives, so a program sound in one setting can fail in another. Screening policy is thus a population judgment, not a property of the test alone.
Policy instruments and the levels of prevention
Public health acts through many levers, not just clinical care. It uses laws and regulation (seatbelt and clean-air laws, food-safety rules), taxation (tobacco and alcohol taxes that raise price and cut consumption), the design of the built environment, health education, and the direct provision of services such as vaccination and screening. Each instrument maps onto the levels of prevention from Module 1, and each suits some problems better than others.
The most powerful instruments often operate upstream, changing the conditions in which people live rather than relying on individual choice one person at a time. A tax that lowers sugary-drink consumption across a whole population can, through the prevention paradox of Module 1, avert more disease than counseling high-risk individuals, because it shifts the entire distribution. Upstream measures also tend to reduce inequities, since they reach people regardless of their access to care or their capacity to act on advice.
Health in all policies
Because the deepest determinants of health lie outside the health sector, in housing, transport, education, agriculture, and employment, an approach called health in all policies asks every branch of government to weigh the health consequences of its decisions. A transport ministry that designs streets for walking, or a housing authority that removes damp and cold, may do more for population health than a clinic ever could. The idea follows directly from the social determinants of Module 1: if upstream conditions drive disease, then health cannot be the concern of the health ministry alone.
Ethics and trade-offs
Because public health acts on whole populations, and sometimes constrains individuals for the common good, it is inseparable from ethics. Its central tension pits individual liberty against collective benefit: quarantine, isolation, and vaccination mandates protect the many by restricting the few. A guiding rule is the least restrictive means, the principle that when a coercive measure is justified, officials should choose the option that achieves the public health goal with the smallest intrusion on liberty.
Two further principles constrain action. Justice and health equity demand that the burdens and benefits of policy fall fairly, and that measures do not deepen the disadvantage of groups already worse off; a well-meant program that helps the advantaged fastest can widen the very gaps it meant to close. And any coercive measure carries a duty to rest on solid evidence and proportionality, since restricting liberty on weak grounds forfeits the public trust on which all public health finally depends.
Common misconceptions
A first misconception treats a rise in reported cases as proof of a real rise in disease, when improved reporting, a broadened case definition, or a publicized scare can lift the count with no true change. A second imagines that a single striking study should drive policy, when the reliable signal comes from systematic reviews that pool many studies. A third assumes that any available screening test should be offered, when screening is warranted only when early treatment helps and the program does more good than harm.
A fourth misconception frames public health ethics as a simple contest in which either liberty or safety must always win. In practice the task is proportionality: matching the intrusiveness of a measure to the seriousness of the threat and the strength of the evidence, and choosing the least restrictive option that works. Recognizing that both unchecked coercion and paralyzed inaction can cause harm is what turns the numbers of this course into responsible judgment.
Try it
During one month, special studies indicate that about 1,000 true cases of a reportable condition occurred in a city. The passive surveillance system received 350 reports, of which 300 were confirmed as true cases and 50 were ruled out. (a) Estimate the system's sensitivity. (b) Compute its positive predictive value. (c) If officials broaden the case definition so that sensitivity rises but predictive value falls, what have they traded, and why might that be worthwhile early in an outbreak?
Worked answer: (a) sensitivity = true cases detected / all true cases = 300 / 1,000 = 0.30, or 30 percent, a reminder that passive systems capture only a fraction of cases. (b) positive predictive value = true cases among reports / all reports = 300 / 350 = 0.857, about 86 percent. (c) they have traded predictive value for sensitivity, catching more real cases at the cost of more false alarms and more investigation. Early in an outbreak, missing true cases is more dangerous than chasing a few false ones, so a sensitive definition is usually preferred first and narrowed later.
Epidemiology supplies the facts; a democratic society weighs them against its values. Understanding both the numbers and their limits, which is the work of this whole course, is what lets you take part in that judgment responsibly, whether as a scientist, a clinician, or a citizen. The discipline that began by counting cases around a London water pump ends here, in the reasoned use of evidence to protect the health of populations.
Sources
- Centers for Disease Control and Prevention. (2012). Principles of epidemiology in public health practice (3rd ed.), Lesson 5, Section 2: Purpose and characteristics of public health surveillance. CDC Self-Study Course SS1978. archive.cdc.gov
- Thacker, S. B., & Berkelman, R. L. (1988). Public health surveillance in the United States. Epidemiologic Reviews, 10, 164-190. pubmed.ncbi.nlm.nih.gov
- Buehler, J. W., Hopkins, R. S., Overhage, J. M., Sosin, D. M., & Tong, V. (2004). Framework for evaluating public health surveillance systems for early detection of outbreaks: Recommendations from the CDC Working Group. MMWR Recommendations and Reports, 53(RR-5), 1-11. pubmed.ncbi.nlm.nih.gov
- Sackett, D. L., Rosenberg, W. M. C., Gray, J. A. M., Haynes, R. B., & Richardson, W. S. (1996). Evidence based medicine: What it is and what it isn't. BMJ, 312(7023), 71-72. pmc.ncbi.nlm.nih.gov
- Higgins, J. P. T., Thomas, J., Chandler, J., Cumpston, M., Li, T., Page, M. J., & Welch, V. A. (Eds.). (2024). Cochrane handbook for systematic reviews of interventions. Cochrane. training.cochrane.org
- Childress, J. F., Faden, R. R., Gaare, R. D., Gostin, L. O., Kahn, J., Bonnie, R. J., ... Nieburg, P. (2002). Public health ethics: Mapping the terrain. Journal of Law, Medicine & Ethics, 30(2), 170-178. pubmed.ncbi.nlm.nih.gov
- Wilson, J. M. G., & Jungner, G. (1968). Principles and practice of screening for disease (Public Health Papers No. 34). World Health Organization. find source ↗
- Key terms
- Surveillance
- Ongoing systematic collection and analysis of health data with timely dissemination for action.
- Passive vs active surveillance
- Relying on routine reporting versus agencies actively seeking cases; cheaper but less complete versus costlier but more complete.
- Hierarchy of evidence
- Ranking of study designs by reliability, from expert opinion up to systematic reviews and meta-analyses.
- Systematic review and meta-analysis
- A rigorous synthesis pooling many studies to give the most reliable overall estimate, atop the evidence hierarchy.
- Cost-effectiveness analysis
- Comparing health gained against cost across options so limited resources buy the most health.
- Health equity
- Fairness in the distribution of health and of the burdens and benefits of health policy across groups.