Module 1: The Philosophy of Science
The epistemological foundations that determine what counts as knowledge and how methods acquire their warrant.
Ontology, Epistemology, and Paradigms
- Distinguish ontology, epistemology, and methodology as nested levels of a research paradigm.
- Explain why methodological choices are downstream of philosophical commitments, not independent of them.
Before you choose a survey over an interview, or a regression over a thematic analysis, you have already made philosophical commitments, usually without noticing. The purpose of this lesson is to make those commitments explicit, because a defensible research design is one whose methods follow coherently from its assumptions about reality and knowledge.
Doctoral examiners rarely fail a proposal because a coefficient was miscomputed. They fail it because the candidate cannot say what kind of knowledge the study was built to produce, or why the chosen technique was the right instrument for that knowledge. This lesson equips you to answer that question at the level of first principles, so that every later decision about sampling, measurement, and analysis has a foundation to rest on. Get the foundation wrong and the errors compound downstream, because a mis-specified paradigm quietly licenses the wrong sampling logic, the wrong instrument, and the wrong standard of proof.
Keep two researchers in mind for the whole lesson. Maria and Devon both study burnout among hospital nurses. Maria administers the 22-item Maslach Burnout Inventory to 400 nurses across six hospitals, then tests whether burnout scores rise with mandatory overtime. Devon interviews twelve nurses for ninety minutes each. He asks what being burned out means to them, how they name it, and when they first used the word. Maria can end up saying that more monthly overtime predicts higher emotional exhaustion in those hospitals. She cannot say what exhaustion feels like from the inside. Devon can describe how nurses tell honest tiredness apart from the deadened feeling they fear most. He cannot say how common either state is. Neither researcher is doing it wrong. They have made different commitments about what burnout is, and those commitments decide what each may claim.
Key idea: You choose your beliefs about reality and knowledge before you choose your methods, and those beliefs set the limit on what your study can honestly claim.
In plain terms
Three questions sit under every study. What is real? How could I know it? What should I do to find out? Answer them in that order. The first answer limits the second. The second limits the third. Say burnout is a real thing people have more or less of. Then a scale can measure it, and a survey makes sense. Say burnout is a word people use to make sense of their week. Then a scale misses the point, and you need to ask them. The mistake is starting at the bottom. Somebody likes running regressions, so they design a study around a regression. That gets the order backward. A statistical test does not tell you what kind of thing you are studying. One more piece of vocabulary. Methodology is your reason for using a family of procedures. A method is one procedure. Naming the method tells a reader almost nothing.
Three nested levels
A research paradigm is a shared set of beliefs about how the world works and how it can be studied. It is usefully decomposed into three levels, each constraining the next.
- Ontology (what kinds of things are real) asks what exists and what reality is like. Is there one reality that sits there whether or not anyone observes it? That is a realist ontology. Or is reality built out of human meaning-making, and therefore multiple and local? That is a relativist ontology. Maria's design assumes burnout is the first kind of thing. Devon's assumes it is the second.
- Epistemology (how we decide something counts as knowledge) asks what can be known and how the knower relates to the known. Can the researcher stand outside the thing studied and observe it objectively? Or is knowledge always shaped by the researcher's position and values?
- Methodology (the strategy that justifies your procedures) asks how we should proceed, given the first two answers. Methodology is the logic of inquiry. Methods are the specific techniques that carry it out, such as a t-test or a coding scheme.
The direction of constraint matters. Ontology conditions epistemology, which conditions methodology, which in turn selects methods. Read the arrow downward and it works. A position on what exists narrows what can be known about it. That narrows how you may credibly study it. That in turn narrows the toolkit. Read the arrow upward and it fails. A t-test does not imply an ontology. You therefore cannot reverse-engineer a coherent design by starting from a favorite technique and working back to a worldview.
Consider a study of teacher burnout. A realist treats burnout as a real, stable attribute that people possess in degrees, measurable with a validated scale such as the Maslach Burnout Inventory. A relativist treats burnout as a category that teachers construct and contest, better approached by asking how they narrate exhaustion in their own words. Same word, two ontologies, two incompatible measurement commitments. Neither is naive; each simply obligates a different method.
Keep methodology and method distinct, because conflating them produces weak proposals. Methodology is the argument for why a class of procedures yields warranted answers to your question. Method is the procedure itself. Interviews are a method; grounded theory and phenomenology are methodologies that both use interviews yet analyze them under different logics and toward different ends. Naming the method tells a reader little. Naming the methodology tells them how your evidence becomes an inference.
Key idea: Ontology sets what is real, epistemology sets what can be known, methodology sets how to proceed, and only then do you pick a method - the arrow never runs backward.
Realism, relativism, and the space between
Ontological positions form a spectrum rather than a binary. At one pole, naive realism holds that reality is directly knowable exactly as it is. Few philosophers of science defend it today, because observation is theory-laden and instruments are fallible. At the other pole, radical relativism holds that there is no reality apart from the accounts we construct, so no account can be judged more accurate than another.
Between them sits critical realism, associated with Roy Bhaskar, which most post-positivist researchers find congenial. It grants that a mind-independent reality exists and has causal powers, while insisting our knowledge of it is always partial, fallible, and mediated by concepts. This is why a critical realist runs a study, expects to be wrong in part, and treats replication as the mechanism that gradually corrects error. Locating your ontology on this spectrum is the first commitment a methods section should make explicit.
Put the spectrum on Maria's study. A naive realist would say her inventory reads burnout off a nurse the way a thermometer reads a fever. A radical relativist would say her 400 scores are just marks on paper with no reality behind them. A critical realist says something more useful. Burnout is real, it has effects such as sick days and resignations, and the inventory is a fallible instrument that gets at part of it. That middle position is why Maria reports a confidence interval instead of a verdict, and why she wants another team to repeat the study in different hospitals.
Key idea: Most working researchers are critical realists - they hold that something real is out there while admitting every measurement of it is partial and correctable.
A map of the major paradigms
Four paradigm families recur across the social and health sciences. Learning their standard ontological and epistemological positions lets you locate any study on the map and predict its criteria of rigor before you read a single result.
- Positivism: a realist ontology of one law-governed reality and an objectivist epistemology in which a detached observer can measure it. Favors deduction, experiments, and quantification.
- Post-positivism: critical realism in Bhaskar's sense, one reality that exists but is only imperfectly and probabilistically knowable, so all findings are fallible. Objectivity becomes a community achievement through replication and peer critique.
- Constructivism and interpretivism: a relativist ontology of multiple constructed realities and a subjectivist epistemology in which knower and known co-create meaning. Favors induction, interviews, and thick description.
- Critical and transformative: a historical realism in which reality is shaped over time by power, so inquiry aims not merely to describe but to expose and change unjust structures. Values are explicit and emancipatory.
Pragmatism sits slightly apart. It brackets the ontological dispute and asks a blunter question: which combination of methods best answers this practical problem? That is why pragmatism is the usual philosophical home for the mixed-methods work you will study later. The point of the map is not orthodoxy. It is that once you declare a cell, your obligations follow, and a reader can hold you to them. Declaring a paradigm is a promise about two things. It says which questions you will and will not try to answer. It also names the yardstick by which your evidence should be judged.
Key idea: Each paradigm comes with its own test of good work, so naming yours tells the reader in advance which yardstick to use on you.
Why this is not merely academic
Reviewers and dissertation committees rarely reject work over a miscalculated statistic. They reject it because the design is internally incoherent. One version claims an interpretivist interest in lived meaning, then forces experience into pre-set Likert categories. Another claims positivist generalization from a purposively chosen sample of six. Naming your paradigm forces you to check that the pieces fit.
The incoherence is usually invisible to the author because each piece looks defensible in isolation. A well-validated scale is good; a rich interview is good; a large sample is good. The error is combinatorial. A study that wants to generalize a prevalence rate needs a probability sample, not six information-rich cases, however eloquent. A study that wants to theorize meaning needs open-ended data, not a closed instrument that presupposes the very categories under investigation. Coherence is the property that the goal, the ontology, and the technique all agree.
Take a concrete case. Suppose a candidate asks how first-generation students experience belonging at elite universities. The verb "experience" signals an interpretivist aim. If the methods section then reports a 40-item belonging scale analyzed by regression, the reviewer's objection writes itself: the instrument answers how much belonging, not how belonging is lived. The remedy is not more statistics. It is realigning the method to the question, or rewriting the question to match the method.
Key idea: Incoherence is a mismatch, not a mistake in arithmetic - the goal, the view of reality, and the technique all have to agree.
Axiology and reflexivity
A fourth level, axiology, concerns the role of values in inquiry. Positivist traditions aspire to value-free observation; interpretivist and critical traditions hold that values are inescapable and should instead be disclosed. This is why qualitative writing so often includes a reflexivity statement in which the researcher accounts for how their own background may have shaped data collection and interpretation. Reflexivity is not a confession of bias to be apologized for; it is a methodological control appropriate to a paradigm that denies the possibility of a view from nowhere.
Reflexivity has two useful forms. Personal reflexivity examines how the researcher's identity, history, and stake in the topic condition what participants disclose and what the analyst notices. Epistemic reflexivity examines the assumptions built into the concepts and instruments themselves. Consider an interviewer studying police trust who is visibly an outsider to the community. She will elicit different accounts than an insider would. A rigorous report treats that fact as evidence about the conditions of knowledge, not as noise to be hidden.
Devon owes a reflexivity statement for the same reason. He is a former ward nurse, so nurses tell him things they would not tell a stranger with a clipboard. That access is an asset. It is also a filter, because his own memory of burnout shapes which phrases he hears as important. Naming both effects is not an apology. It tells the reader the conditions under which this knowledge was produced. Maria's design controls researcher influence differently, by standardizing the instrument so every nurse answers the same 22 items in the same order.
Key idea: Values cannot be scrubbed out of interpretive work, so it discloses them - reflexivity is a control, not a confession.
Incommensurability and the paradigm wars
Thomas Kuhn argued that competing paradigms can be incommensurable: they define their central terms so differently that adherents partly talk past one another. In research methods this fueled decades of paradigm wars over whether quantitative and qualitative approaches rest on assumptions too opposed to combine. The strong incommensurability thesis implies you must pick one camp and stay inside it.
Most working researchers now hold a weaker view. Paradigms do differ. Even so, one project can use different lenses for different sub-questions, on two conditions. The researcher must be explicit about the shift. And the researcher must not claim, say, statistical generalization from interpretive data. This pragmatic settlement is what makes principled mixed-methods design possible. Paradigms still matter. They matter enough to be named and reconciled rather than quietly blurred.
Key idea: Paradigms can be combined in one project only if you say out loud which lens answers which sub-question and refuse to borrow the other lens's warrant.
Worked contrast: one topic, two paradigms
Take the topic of workplace surveillance. A post-positivist design operationalizes perceived monitoring and job strain as scale scores, samples employees across firms, and tests whether monitoring predicts strain while controlling for workload. Its warrant rests on measurement validity, sampling, and ruling out confounds. Its payoff is an estimated effect that, if it survives replication, generalizes to comparable workplaces.
An interpretivist design instead asks how employees make sense of being watched. It gathers observations and interviews and builds a grounded account of, say, the rituals of resistance and performance that workers enact under the camera. Its warrant rests on prolonged engagement and the fit between claims and testimony. Its payoff is conceptual: a vocabulary that later quantitative work can operationalize. Neither answers the other's question. A design that promised both from one dataset would owe the reader an account of how.
Key idea: One topic supports many studies, but each paradigm buys a different payoff - an effect estimate that travels, or a vocabulary that explains what is going on.
Where people get stuck
Three confusions account for most of the trouble in this material.
- Ontology confused with epistemology. Ontology is about the world; epistemology is about our access to it. Two people can agree that burnout is real (same ontology) and still disagree about whether a questionnaire can capture it (different epistemology). Test yourself with the question form: "what is there?" is ontology, "how would I know?" is epistemology.
- Paradigm confused with data type. Numbers are not positivism and words are not interpretivism. Devon could count how many of his twelve nurses used the word "hollow" without becoming a positivist. Maria could quote a nurse's comment without becoming an interpretivist. What fixes the paradigm is the logic of inference, not the shape of the data.
- Methodology confused with method. "I used interviews" names a method and settles nothing. "I used interpretive phenomenological analysis of interviews to describe the structure of the experience" names a methodology and tells the reader how testimony turns into a finding.
Key idea: Ask "what is real", "how would I know", and "by what logic does my evidence become a claim" as three separate questions, and most paradigm confusion dissolves.
Try it
Exercise. A doctoral student proposes to study misinformation belief. Their draft says: "I will conduct 12 in-depth interviews to measure the prevalence of false-belief endorsement in the national population and generalize the rate with a margin of error." Identify the ontological and inferential incoherence, then propose two internally coherent redesigns, one positivist and one interpretivist.
Model answer. The incoherence is a mismatch between an interpretivist method and a positivist goal. Twelve purposively chosen interviews cannot estimate a population prevalence or support a margin of error, both of which presuppose probability sampling and quantified measurement. The phrase "measure the prevalence and generalize the rate" belongs to a realist, objectivist frame that the chosen method cannot deliver.
A coherent positivist redesign defines false-belief endorsement operationally with a validated item battery. It draws a probability sample of the target population and reports prevalence with a confidence interval. A coherent interpretivist redesign keeps the 12 interviews but changes the question. It asks how people reason about and justify contested claims, then reports a grounded account of that reasoning and makes no prevalence claim. Each version aligns ontology, epistemology, and method. That alignment is the standard every later chapter will hold you to.
Recap
- Ontology (what is real) constrains epistemology (what counts as knowledge), which constrains methodology, which selects methods; the arrow never runs upward from a favorite technique.
- Methodology is the argument for a class of procedures, and a method is the procedure itself: interviews are a method, grounded theory is a methodology.
- Critical realism is the common middle position - a real world exists, and our fallible measures of it need replication to correct.
- Positivism, post-positivism, constructivism, and critical theory each pair an ontology with a standard of rigor; pragmatism brackets the dispute to serve the question.
- A design is incoherent when goal and technique come from different paradigms, as when twelve interviews are asked to yield a population prevalence rate.
- Axiology asks what role values play, and reflexivity is the control an interpretivist design uses in place of pretending to a view from nowhere.
- Maria's 400-nurse survey and Devon's twelve interviews are both good research, but each may claim only what its own paradigm licenses.
Sources
- Chakravartty, A. (2017). Scientific realism. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Reiss, J., & Sprenger, J. (2020). Scientific objectivity. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Bird, A. (2018). Thomas Kuhn. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Longino, H. (2019). The social dimensions of scientific knowledge. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Bhattacherjee, A. (2012). Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Understanding science. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Korstjens, I., & Moser, A. (2017). Series: Practical guidance to qualitative research. Part 2: Context, research questions and designs. European Journal of General Practice, 23(1), 274-279. pmc.ncbi.nlm.nih.gov
- Key terms
- Ontology
- The branch of philosophy concerned with the nature of reality and what exists.
- Epistemology
- The theory of knowledge; what can be known and the relation between knower and known.
- Methodology
- The logic and strategy of inquiry that justifies the choice of specific methods.
- Paradigm
- A shared framework of ontological, epistemological, and methodological assumptions guiding research.
- Axiology
- The study of the role of values in inquiry, including whether research can be value-free.
- Reflexivity
- The researcher's critical self-examination of how their position shapes the research process.
Positivism, Post-positivism, and Interpretivism
- Contrast the aims, logic, and criteria of positivist and interpretivist traditions.
- Explain the post-positivist revisions of naive positivism, including fallibilism and theory-ladenness.
The two families of assumption you meet most often are positivism and interpretivism, with post-positivism as the dominant modern refinement of the first. Knowing what each claims, and what it does not, lets you read a methods section and immediately see the criteria against which the work wants to be judged.
These are not merely labels for quantitative versus qualitative work. They are packages of commitments about what reality is, what counts as a good explanation, and what makes a claim trustworthy. A researcher can use numbers within an interpretivist frame or words within a post-positivist one. What fixes the tradition is the underlying logic of inference, not the data type, and that is what this lesson trains you to detect.
A brief caution about the word positivism. In ordinary academic speech it is often used loosely, and sometimes as an insult, to mean any use of numbers or any claim to objectivity. That casual usage obscures more than it reveals. In this course the term keeps its technical sense: a specific ontological and epistemological package with a history, not a synonym for quantitative method. Holding the term precise is part of writing a methods section a committee will respect.
Use one topic to hold the whole lesson together: why some parents delay childhood vaccines. Team A mails a validated hesitancy scale to 2,000 parents, then tests whether hesitancy scores drop after parents read one of three message versions. Team B sits in clinic waiting rooms for four months and interviews eighteen parents about how they decided. Team A ends with a number: message version two lowered mean hesitancy by 0.4 points, 95% CI 0.2 to 0.6. Team B ends with a description: parents were not weighing evidence so much as deciding whom to trust after a bad experience with a rushed appointment. Both findings are real. They are not the same kind of finding, and they answer to different judges.
Key idea: Positivism, post-positivism, and interpretivism are packages of assumptions, not code words for numbers and words.
In plain terms
Strip the labels away and the ideas are simple. A positivist thinks there is one world out there. It follows rules. A careful observer can measure those rules. So the researcher stands back, takes measurements, and tests a guess. A post-positivist agrees there is one world. But every measure of it can be wrong. So findings are always provisional. Other people have to repeat the study. An interpretivist starts somewhere else. People build the social world out of meanings. You cannot measure a meaning from the outside. You have to ask, listen, and work out what an action means to the person doing it. Two extra families matter. Critical theory asks who gains from the way things are described. Pragmatism asks only which tool answers the question. That is the whole map. The rest of this lesson fills in the detail and the history.
Positivism
Classical positivism makes two claims. There is a single objective reality governed by regularities. And the proper aim of science is to discover those regularities through systematic observation. Its hallmarks are a realist ontology and an objective epistemology, in which the researcher stays separate from the studied. It also prefers deductive reasoning: derive hypotheses from theory, then test them against data. Explanation takes a causal, often law-like form. Quantification is prized because it disciplines observation and supports replication.
The tradition traces to Auguste Comte. In the nineteenth century he proposed a positive science of society modeled on physics. The Vienna Circle sharpened it in the 1920s under the name logical positivism. That movement advanced a strict verification principle: a statement is meaningful only if it can be empirically verified or is true by definition. Metaphysical claims that no observation could confirm were dismissed as literally meaningless. Few hold the verification principle today, yet its spirit survives wherever researchers insist that a construct be tied to observable indicators.
A related commitment is operationalism, associated with the physicist Percy Bridgman, which defines a concept by the operations used to measure it. Intelligence, on a strict operationalist reading, just is the score a given test yields. The appeal is discipline, since it forbids vague talk of unmeasured essences. The cost is that it can trivialize a construct by equating it with a single fallible instrument. It would be like saying temperature is nothing beyond the reading on one particular thermometer. Post-positivists took that objection seriously.
A textbook positivist study looks like this. A team hypothesizes that sleep deprivation impairs working memory. It operationalizes deprivation as hours awake and memory as digit-span score. It randomly assigns volunteers to sleep conditions and tests the mean difference. The reasoning is deductive and the variables are quantified. The researcher stays outside the phenomenon. The payoff is a causal generalization expected to replicate in other laboratories. Every choice reflects one ontology: a stable, observer-independent regularity waiting to be measured.
Key idea: Positivism aims at law-like causal regularities, measured by a detached observer and tested by deduction.
Post-positivism
Twentieth-century philosophy of science made naive positivism untenable, and post-positivism absorbed the lessons rather than abandoning the goal of objective knowledge. Three revisions matter most:
- Fallibilism. All knowledge is provisional. We do not prove theories true; at best we fail to refute them. This reframes science as the elimination of error rather than the accumulation of certainty.
- Falsifiability. Karl Popper argued that what demarcates science is that its claims forbid something observable - they are falsifiable. A theory compatible with every conceivable outcome explains nothing. This is why a well-posed hypothesis specifies what result would count against it.
- Theory-ladenness of observation. There are no pure, uninterpreted facts; what we notice is shaped by the concepts and instruments we bring. Objectivity is therefore reconceived as an achievement of the community - through peer review, replication, and critical scrutiny - rather than a property of a lone observer.
These revisions do not restore certainty. They complicate falsification itself. The Duhem-Quine thesis notes that we never test a hypothesis in isolation. It always comes bundled with auxiliary assumptions about the instrument, the sample, and the setting. So a failed prediction never tells us for certain which element was wrong. Suppose Team A's message has no effect. Was the theory wrong, or was the scale insensitive, or did parents skim the leaflet? Imre Lakatos answered with research programmes that protect a hard core while adjusting a protective belt. Kuhn described science as long stretches of normal problem-solving punctuated by revolutions. The working lesson for a doctoral researcher is modesty: evidence constrains theory without ever fully determining it.
Post-positivism is the default operating system of much contemporary quantitative social and health science, even when researchers never name it. You can hear it in the phrasing. A report that speaks of failing to reject the null, offers a confidence interval instead of certainty, or calls for replication before belief is speaking post-positivist. Recognizing this lineage helps you see something useful. The statistical rituals of later chapters are not arbitrary conventions. They are fallibilism turned into procedure.
Key idea: Post-positivism keeps the goal of objective knowledge but treats every finding as fallible, so objectivity becomes the work of a critical community.
Interpretivism
Interpretivism has roots in hermeneutics (the study of how texts and actions are interpreted) and in Weber's notion of Verstehen, or interpretive understanding. It begins from a different ontology. The social world is not a set of external objects. It is a web of meanings that people actively construct. You cannot understand a strike, a diagnosis, or a ritual by measuring its surface behavior alone. You must grasp what it means to the actors. The knower and the known are entangled, so knowledge is co-constructed and bound to its context. The aim shifts from causal explanation and generalization toward rich, contextual understanding. Reasoning tends to be inductive, building concepts up from data rather than testing them from above.
Interpretivism draws on several streams. Wilhelm Dilthey distinguished the human sciences, which seek understanding, from the natural sciences, which seek causal explanation. Alfred Schutz brought phenomenology into sociology. He argued that everyday actors already interpret their world before any researcher arrives. Peter Berger and Thomas Luckmann showed how repeated interactions harden into institutions that later feel objective and external. So the interpretivist's task is second-order interpretation. You make sense of the sense that participants have already made of their own situation. Team B's parents had already built an account of who deserved their trust. Team B's job was to understand that account, not to correct it.
An interpretivist study of fatigue among hospital staff would look wholly different from the memory experiment above. Rather than scoring digit spans, the researcher shadows night-shift nurses and interviews them about how exhaustion reshapes their sense of competence and care. The output is not a mean difference but a set of themes, grounded in participants' own words, about what being depleted means for professional identity. The warrant is depth and fidelity to lived experience, not comparison across randomized conditions.
| Dimension | Positivism / Post-positivism | Interpretivism |
|---|---|---|
| Ontology | Single (fallibly knowable) reality | Multiple, socially constructed realities |
| Aim | Causal explanation, prediction | Understanding of meaning in context |
| Logic | Primarily deductive, hypothesis-testing | Primarily inductive, concept-building |
| Researcher role | Detached, minimizing influence | Instrument, reflexively engaged |
| Quality criteria | Internal/external validity, reliability | Credibility, transferability, dependability |
Neither family is superior in the abstract. The right question is which set of assumptions matches your object of study and your inferential goal. A pharmacologist testing a dose-response curve and an ethnographer studying how nurses improvise care are answering different kinds of questions. Each borrows its standards of rigor from a different tradition. The table is a compass, not a cage. Many strong programs of research move between its columns over a career.
Key idea: Interpretivism seeks the meaning an action has for the people doing it, so its warrant is depth and fidelity rather than control and comparison.
Critical and pragmatic alternatives
Two further families complete the landscape. Critical theory, rooted in the Frankfurt School and extended by feminist, critical-race, and participatory traditions, shares the interpretivist premise that reality is socially constructed but adds that those constructions serve power. Its aim is not neutral description but critique and change. Its quality standard asks a further question: does the research advance the interests of the studied, or only the researcher's career? A participatory action study of tenants organizing against eviction sits squarely here. A critical study of vaccine hesitancy would not treat hesitant parents as deficient in information. It would ask which communities were experimented on in the past, and who profits from the mistrust that followed.
Pragmatism is associated with Peirce, James, and Dewey, and it was revived by mixed-methods scholars. It refuses to let the ontological quarrel dictate method. Its criterion is workability: choose whatever combination of tools answers the question and yields useful, defensible knowledge. Pragmatism is why one dissertation can legitimately pair a survey with interviews. The researcher still owes an account of how the two kinds of evidence speak to one integrated question rather than sitting side by side unreconciled.
Key idea: Critical theory asks who benefits from the way reality is currently described; pragmatism asks only which tools will answer the question at hand.
Reading a methods section for its paradigm
Paradigms are rarely announced, so you infer them from tells. Verbs betray aims. Phrases such as "determine the effect of," "predict," and "control for" signal a post-positivist frame. Phrases such as "explore," "understand," and "interpret the meaning of" signal an interpretivist one. Sampling betrays goals too. Probability sampling points toward generalization, purposive sampling toward depth. Even the treatment of the author is a tell. A detached, absent narrator implies objectivism. A reflexive first-person voice implies a co-constructed epistemology.
The skill matters because it lets you apply the right yardstick. Faulting an ethnography for a small, non-random sample misunderstands its warrant, which never rested on statistical generalization. Faulting a randomized trial for thin description of participants' inner lives misunderstands its warrant, which rested on comparison and control. A competent reviewer first identifies the tradition a study is working in, then judges it by that tradition's standards rather than by an imported checklist.
A caution follows from this. Paradigms are packages, so mixing their vocabularies carelessly reproduces the incoherence flagged in the previous lesson. Imagine an author who promises to explore lived experience and then, one sentence later, tests whether the effect is significant at the .05 level. That author owes the reader an explicit account of how one project houses both aims. Often the honest fix is to split the study into phases with distinct questions. That is precisely the disciplined move mixed-methods design formalizes.
Key idea: Read the verbs, the sampling, and the narrator's voice to find a study's paradigm, then judge the study by that tradition's rules instead of an imported checklist.
Where people get stuck
Two mix-ups do most of the damage here.
- Positivism is not the same as quantitative. Team B could report that fourteen of eighteen parents named a single bad appointment as the turning point. That is a count inside an interpretivist study. Meanwhile a survey full of numbers can be interpretivist if it is used to map how respondents categorize their own world rather than to estimate an effect. The tradition lives in the inference, not the arithmetic.
- Post-positivism is not softened positivism about method. The difference is epistemological, not procedural. A positivist and a post-positivist may run the same randomized trial. The post-positivist reports a confidence interval, says the finding could be wrong, and asks for replication. The classical positivist expects the measurement to deliver the regularity itself.
One further trap concerns criteria. Internal validity and reliability belong to the post-positivist column. Credibility, transferability, dependability, and confirmability belong to the interpretivist column. Applying a criterion across the line is the most common reviewing error in this material, and the exercise below is built on it.
Key idea: Data type does not fix paradigm, and criteria do not travel across paradigms - a small non-random sample is a flaw only if the study claimed generalization.
Try it
Exercise. A journal abstract reads: "Drawing on 60 hours of participant observation in a hospice, we develop a grounded theory of how nurses negotiate hope with dying patients." A reviewer objects that the study "cannot be generalized because the sample is one hospice and no statistics are reported." Diagnose the reviewer's error in paradigm terms, and state the criterion by which the study should instead be judged.
Model answer. The reviewer has imported a post-positivist standard, statistical generalization, into an interpretivist study that never claimed it. A grounded theory built from participant observation seeks conceptual insight into meaning and process, not a prevalence estimate, so the absence of random sampling and inferential statistics is appropriate rather than a defect.
The proper criteria come from the qualitative tradition. Credibility is achieved through prolonged engagement and by grounding claims in observed instances. Transferability is supported by thick description that lets readers judge fit to other settings. Dependability and confirmability come from an auditable analytic trail. If the theory illuminates the negotiation of hope in ways a reader can trace to the data, the study succeeds on its own terms. The reviewer should ask whether the account is trustworthy and useful, not whether it generalizes to a population it was never designed to represent.
Recap
- Positivism assumes one law-governed reality that a detached observer can measure, and it reasons deductively from theory to test.
- Post-positivism keeps that goal but adds fallibilism, falsifiability, and theory-ladenness, so objectivity becomes a community achievement through replication and critique.
- The Duhem-Quine thesis explains why a failed prediction never says exactly which assumption broke, which is why evidence constrains theory without determining it.
- Interpretivism treats the social world as a web of constructed meaning and aims at understanding, using inductive concept-building and the researcher as instrument.
- Critical theory adds power and an emancipatory aim; pragmatism sets the ontological quarrel aside and judges methods by workability.
- You can infer a study's paradigm from its verbs, its sampling logic, and its narrative voice, then apply that tradition's criteria rather than an imported checklist.
- Team A's effect estimate and Team B's account of trust are both good findings, but neither may be judged by the other's yardstick.
Sources
- Creath, R. (2023). Logical empiricism. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Thornton, S. (2023). Karl Popper. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Bourdeau, M. (2022). Auguste Comte. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Mantzavinos, C. (2023). Hermeneutics. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Makkreel, R. (2021). Wilhelm Dilthey. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Legg, C., & Hookway, C. (2021). Pragmatism. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Bohman, J., Flynn, J., & Celikates, R. (2021). Critical theory. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Science and common sense. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Key terms
- Positivism
- A tradition holding that a single objective reality can be known through systematic, quantifiable observation.
- Post-positivism
- A refined positivism accepting fallibilism, falsifiability, and the theory-ladenness of observation.
- Interpretivism
- A tradition holding that social reality is constructed through meaning and must be understood in context.
- Falsifiability
- Popper's criterion that scientific claims must forbid some observable outcome and thus be capable of refutation.
- Theory-ladenness
- The idea that observation is always shaped by prior concepts and instruments, so no fact is purely neutral.
- Verstehen
- Weber's concept of interpretive understanding, grasping the meaning an action holds for the actor.
Induction, Deduction, and Abduction
- Differentiate deductive, inductive, and abductive inference by their logical form and risk.
- Map each mode of inference onto the phase of research where it typically operates.
Every empirical study moves between theory and evidence, and it does so using one or more of three forms of inference. Being precise about which you are using clarifies what your conclusion can legitimately claim, and it exposes a whole class of errors that arise when one mode is quietly swapped for another.
The three modes are not competitors from which you pick a favorite. They are complementary moves that a mature research program makes at different moments. This lesson defines each by its logical form and its characteristic risk, then shows how they chain together into the loop by which inquiry actually advances.
Getting the vocabulary right also disciplines how you write. A dissertation that labels its exploratory data-mining as hypothesis testing, or its single case study as proof of a general law, invites exactly the objections a committee is trained to raise. When you can name the inference behind each claim, you can calibrate the verb: suggests and is consistent with for ampliative moves, entails or is refuted by for deductive ones.
Here is a case to carry through the lesson. A hospital quality team notices that one surgical ward readmits 18% of patients within thirty days while a nearly identical ward readmits 9%. That gap is a surprise, and the surprise is where inference starts. Someone guesses that the high ward discharges patients before their pain is controlled, so they return. That guess is abduction. Someone else works out what must follow if the guess is right: patients on the high ward should show worse pain scores on the day of discharge. That step is deduction. The team then pulls 300 discharge records, finds the predicted pattern, and says it probably holds for next quarter's patients too. That last step is induction. One anomaly, three different kinds of thinking, three different levels of certainty.
Key idea: Deduction tests, induction generalizes, and abduction invents - and a claim is only as strong as the kind of inference behind it.
In plain terms
Three ways of reasoning, three plain descriptions. Deduction works downward. Start with a rule, apply it to a case, and the answer follows for certain. Nothing new is learned. Induction works upward. Look at many cases and guess the rule. You might be wrong, and the next case can prove it. Abduction works sideways. Something surprising happens, and you invent a reason for it. That reason might also be wrong, but at least now you have something to test. Real research uses all three, in a loop. Something odd shows up. You guess why. You work out what should follow if your guess is right. You go and look. Then you count how often the pattern holds. The main sin in this area is mislabeling. If the data gave you the idea, the same data cannot test it. Say so, and nobody can complain.
Deduction
Deductive inference moves from general premises to a specific conclusion that follows necessarily. If the premises are true and the argument is valid, the conclusion cannot be false. Classic form: All humans are mortal; Socrates is human; therefore Socrates is mortal. Deduction is truth-preserving but not ampliative - it never tells you more than was already contained in the premises. In research, deduction dominates the confirmatory phase: you derive a testable prediction from a theory and confront it with data.
Two properties of a deductive argument must be kept apart. Validity concerns form. The conclusion follows from the premises whether or not those premises are true. Soundness adds that the premises are in fact true. A perfectly valid argument built on a false premise delivers a false conclusion with complete logical propriety. That is why deductive rigor in a proposal never substitutes for the empirical truth of its starting assumptions. A committee probes both the logic and the premises.
Deductive logic is what makes hypothesis testing possible. Suppose a theory of cognitive load predicts something specific. If the theory is true, students given a split-attention worksheet will score lower than students given an integrated one. The argument then runs in three steps. If the theory holds, the integrated group outperforms. The integrated group did not outperform. Therefore the theory as specified is wrong. That pattern is modus tollens, or denying the consequent. It is why a single disconfirming result can challenge a theory. Confirmation is weaker, because many rival theories could predict the same success.
Key idea: Deduction can refute cleanly but never discovers anything new, and a valid argument from a false premise is still worthless.
Induction
Inductive inference generalizes from specific observations to a broader claim. Observe many white swans and you infer that swans are white. Induction is ampliative, meaning the conclusion says more than the premises do. For that very reason it is not truth-preserving. The conclusion can be false even when every premise is true, as the first black swan demonstrates. This is Hume's problem of induction: no finite run of observations guarantees the next case. Induction dominates the exploratory, pattern-finding phase. It is also central to grounded-theory work that builds concepts from data.
Not all induction is the crude enumerative kind that counts swans. Statistical induction generalizes from a sample to a population using probability, and the entire apparatus of sampling and confidence intervals you will meet later is a disciplined form of it. Eliminative induction, by contrast, strengthens a conclusion by ruling out rival explanations rather than piling up confirming instances. Mill's methods of agreement and difference are early formalizations. Strong empirical work rarely rests on enumeration alone. It designs comparisons that eliminate competitors. The readmission team does this when it checks whether the two wards differ in patient age, surgery type, and staffing before blaming pain control.
Hume's problem has never been dissolved, only managed. We cannot justify induction without circularly assuming that the future will resemble the past. The practical response, congenial to post-positivism, is to treat inductive conclusions as provisional and to expose them to further test rather than to seek a guarantee. This is also why replication carries such weight: it does not prove a generalization, but each successful repetition narrows the space in which a black swan might still be hiding.
Induction also has a qualitative face. In grounded theory a researcher reads transcripts, notices recurring ideas, and builds categories upward until a theoretical account emerges. The logic is inductive because the concepts are not imposed in advance but abstracted from instances. The same fallibility applies. A category induced from twenty interviews may not hold in the twenty-first. That is why qualitative analysts actively hunt for disconfirming cases instead of collecting only confirming ones. The trustworthiness chapter formalizes that discipline.
Key idea: Induction can always be overturned by the next case, so its conclusions are provisional and replication is how they earn trust.
Abduction
Abductive inference, articulated by C. S. Peirce, is inference to the best explanation. You face a surprising observation. You ask which hypothesis, if true, would make that observation unsurprising. Then you provisionally adopt the most plausible one. A physician reasoning from symptoms to the most likely diagnosis is reasoning abductively. Abduction is the logic of discovery, because it generates candidate explanations. Those candidates are then developed deductively into predictions and tested. Like induction, abduction is fallible. Unlike enumerative induction, it can introduce genuinely new theoretical terms.
Abduction raises an obvious question: best by what standard? Philosophers list several marks of a good explanation, and you can use them as a checklist. A strong candidate has explanatory scope, so it accounts for more of the data. It has simplicity, so it posits fewer unsupported entities. It has plausibility, so it fits established knowledge. And it has testability, so it yields new predictions. The danger is the just-so story. That is an explanation that fits the data at hand precisely because it was tailored to them, and that forbids nothing else. "The high ward has a difficult patient culture" is a just-so story. It explains the gap but predicts nothing you could go and check.
Consider a worked diagnostic case. A patient presents with fever, cough, and a specific chest-imaging pattern. Pneumonia would make all three unsurprising, so the clinician abduces pneumonia as the best explanation. She then reasons deductively: if this is bacterial pneumonia, a culture should grow a particular organism. She orders the test. The abduction generated the hypothesis, deduction turned it into a prediction, and the culture supplies the inductive check. Research programs run the same loop at a larger scale and a slower pace.
Key idea: Abduction is where hypotheses come from, and a good one earns its place by explaining more, assuming less, and predicting something you can still go and test.
| Mode | Direction | Truth-preserving? | Typical role |
|---|---|---|---|
| Deduction | General to specific | Yes | Testing predictions |
| Induction | Specific to general | No (ampliative) | Generalizing patterns |
| Abduction | Observation to best explanation | No (ampliative) | Generating hypotheses |
The hypothetico-deductive method
The dominant template of confirmatory science formalizes this interplay. The hypothetico-deductive method has four steps. State a hypothesis. Deduce an observable consequence that would follow if it were true. Arrange conditions under which that consequence could fail to appear. Then observe. Confirmation underdetermines theory, meaning several rival theories can fit the same result. So the method leans on the asymmetry Popper emphasized. A clean disconfirmation is logically decisive in a way that a lone confirmation is not. Strong designs therefore court risky predictions, the kind a false theory would most likely fail.
This explains a stylistic feature of good hypotheses. A prediction that "something will differ" is nearly unfalsifiable and therefore weak. A prediction of a specific direction and rough magnitude sticks its neck out. The more a hypothesis forbids, the more informative its survival. Precision is not decoration. It is what gives a test its evidential bite. That is why later chapters insist on stating directional predictions and effect sizes before the data are seen.
Key idea: A hypothesis is strong when it rules something out, so state the direction and rough size of the effect you expect before you look.
The cycle in practice
Mature research programs cycle through all three. A surprising anomaly prompts an abductive leap to a candidate explanation. The explanation is then elaborated deductively into specific, falsifiable predictions. Data collection and inductive generalization assess how far the pattern holds. Leftover anomalies restart the cycle. Knowing where you stand in this loop keeps you honest. Exploratory findings arrived at inductively should not be dressed up as confirmatory tests. And a hypothesis abduced from a dataset cannot then be confirmed on that same dataset without circularity.
A historical example makes the loop concrete. Ignaz Semmelweis observed in the 1840s that one maternity clinic had a far higher rate of fatal childbed fever than another, a surprising anomaly. He abduced that some contaminating matter carried on physicians' hands was responsible, a hypothesis suggested but not proven by the pattern. He then deduced that if so, handwashing in a chlorine solution should cut mortality, intervened, and observed a sharp inductive confirmation as deaths fell. The germ theory that would later explain why came only afterward.
Notice what the example does not show. Semmelweis could not deduce his hypothesis from the mortality data, because deduction elaborates and tests but never discovers. Nor did the falling death rate prove the contamination theory beyond revision, since other explanations remained logically possible. The case models good practice for one reason. Each inferential move does the job it is suited to, and none is asked to carry more certainty than its form allows.
Key idea: Real research programs loop through all three modes, and honesty means labeling which mode produced each claim.
The circularity trap
The most common inferential error in applied research has a simple shape. A hypothesis generated by induction or abduction is reported as though it had been confirmed on the very data that suggested it. Suppose you comb a dataset, notice that one subgroup responded, and then test that subgroup effect on the same data. You have shown almost nothing. The pattern was nearly guaranteed to appear, because you picked it out for that reason. Dressing exploratory discovery as confirmatory testing is known as HARKing, or hypothesizing after the results are known.
The remedy is to keep the phases honest and, ideally, separate in time. Exploratory analyses legitimately generate hypotheses; those hypotheses then earn their status only against fresh data, whether a held-out sample, a preregistered replication, or an independent study. A later chapter returns to this as a pillar of open science. For now, the discipline is simple. Label each claim by the inference that produced it. Then resist promoting a hunch to a finding without a genuinely new test.
Key idea: The data that suggested a hypothesis cannot also test it - a new sample, a held-out set, or a preregistered replication is required.
Where people get stuck
Three distinctions repay careful attention.
- Valid is not the same as true. Validity is about the shape of the argument; soundness adds true premises. "All wards with low staffing readmit more; this ward has low staffing; therefore it readmits more" is valid even if the first premise is false.
- Abduction is not induction. Induction says the pattern will keep holding. Abduction says here is why the pattern exists. Counting that 18% of the high ward's patients return is inductive. Proposing uncontrolled pain as the cause is abductive.
- Exploratory is not inferior, only different. There is nothing wrong with mining data for patterns. The error is labeling the result confirmatory. Say "this analysis was exploratory and generated the hypothesis we now plan to test", and the same work becomes respectable.
A last practical note on verbs. Match the verb to the inference: suggests, is consistent with, and raises the possibility for ampliative moves; entails, predicts, and is inconsistent with for deductive ones. Reviewers read verbs closely, because an overclaiming verb is the cheapest tell that an author has confused the modes.
Key idea: Most inferential errors are labeling errors - the work was fine, but the verb promised more than the logic delivered.
Try it
Exercise. A team notes that in an existing survey, respondents who drink coffee report better mood. They write: "We hypothesize that coffee improves mood, and our data confirm this hypothesis, since the correlation is positive and significant." Name the inferential mode actually at work, identify the error, and describe a design that would let the causal claim be tested rather than assumed.
Model answer. What actually occurred is an inductive or abductive move: a correlation in existing data suggested the explanation that coffee lifts mood. Reporting that the same data confirm the hypothesis is the circularity trap, since the pattern that generated the idea cannot also serve as its independent test. Correlation also fails to establish direction or to rule out confounds such as sleep, income, or sociability.
A legitimate test specifies the prediction in advance and gathers new data under controlled conditions. Randomly assign participants to caffeinated or decaffeinated coffee. Hold expectancy constant with blinding. Then measure mood afterward. Random assignment addresses confounding. The fresh, preregistered prediction restores the deductive, confirmatory logic that the original report only pretended to have. Only then does the word confirm become honest rather than decorative.
Recap
- Deduction runs from general premises to a necessary conclusion; it is truth-preserving but never tells you anything the premises did not already contain.
- Validity concerns the form of an argument and soundness adds true premises, so a valid argument from a false premise still yields a false conclusion.
- Induction runs from cases to a general claim, gains new content, and can always be overturned by the next observation - Hume's problem of induction.
- Abduction runs from a surprising observation to the explanation that would make it unsurprising, judged by scope, simplicity, plausibility, and testability.
- The hypothetico-deductive method chains them together, and it prizes risky predictions because disconfirmation is decisive in a way confirmation is not.
- HARKing is treating a hypothesis abduced from a dataset as confirmed by that same dataset; the remedy is new data, ideally with a preregistered prediction.
- Match your verbs to your inference, since an overclaiming verb is the clearest sign that the modes have been confused.
Sources
- Douven, I. (2021). Abduction. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Henderson, L. (2022). The problem of induction. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Andersen, H., & Hepburn, B. (2021). Scientific method. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Burch, R. (2024). Charles Sanders Peirce. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Stanford, K. (2023). Underdetermination of scientific theory. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review, 2(3), 196-217. pubmed.ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Thinking like a researcher. In Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Key terms
- Deduction
- Inference from general premises to a conclusion that follows necessarily; truth-preserving but not ampliative.
- Induction
- Inference generalizing from specific observations to a broader claim; ampliative but fallible.
- Abduction
- Inference to the best explanation; generates plausible hypotheses from surprising observations.
- Ampliative
- Describing an inference whose conclusion contains more information than its premises.
- Problem of induction
- Hume's observation that no finite set of observations can guarantee a universal generalization.
- Confirmatory research
- Research that tests pre-specified hypotheses, relying chiefly on deductive inference.
Module 2: Questions, Theory, and the Literature
Turning a broad interest into an answerable question, a testable hypothesis, and a defensible relationship to prior work.
From Problem to Answerable Research Question
- Convert a broad topic into a focused, feasible, and answerable research question.
- Distinguish descriptive, relational, and causal question types and their design implications.
A dissertation lives or dies on the quality of its question. A vague topic ("social media and mental health") is not a research question; it is a field. Your task is to carve out of that field one question that is narrow enough to answer with available resources and important enough to be worth answering.
Committees can often predict a project's fate from its central question alone. A sharp question implies its own design, sample, and analysis, so the rest of the proposal almost writes itself. A blurred question hides its difficulties until data collection is underway and the problems become expensive. This lesson gives you tools to interrogate a question before it commits you to anything, so that its weaknesses surface on paper rather than in the field.
Watch one question get built. A doctoral student, Priya, starts with "social media and teenage mental health." By the end of this lesson that topic has become: "Among 14- to 16-year-olds at three public high schools, does a two-week, app-enforced limit of one hour of evening social media, compared with unrestricted use, lower scores on the PHQ-A depression screen measured at four weeks?" Nothing important was lost. Everything vague was replaced with something checkable. Notice what the second version tells you that the first does not: who is studied, what is done to them, what they are compared against, what is measured, and when.
Key idea: A topic names a field; a research question names one observation that would settle something.
In plain terms
A good question passes a simple test. Can you say, in one sentence, what result would answer it? If not, keep working. Four things make a question answerable. Name who you are studying. Name what you are doing to them, or measuring about them. Name what you are comparing that against. Name what counts as the outcome, and when you will measure it. That is all PICO means. Then run five quick checks. Can I actually recruit these people? Would anyone care about the answer? Is the answer already known? Can I do this ethically? Would it change anything? Two traps to avoid. Do not ask two things in one question. And do not use a word like "improves" unless your design can support a causal claim. If you cannot draw the table that shows your answer, the question is not finished yet.
Where good questions come from
Researchable questions rarely arrive by inspiration. They are manufactured from identifiable sources, and knowing the sources helps you generate candidates on purpose. Three sources are the most productive. Theoretical tensions arise where two well-supported ideas make conflicting predictions. Empirical contradictions arise where studies disagree and someone must adjudicate. Boundary conditions arise where an established effect may or may not hold in a new population, setting, or time period.
Practical problems supply a second stream. Think of a clinic with an unexplained readmission rate, a district with a stubborn achievement gap, or a firm with rising attrition. Each poses a question that matters to stakeholders and often to theory as well. Methodological limitations supply a third stream. An important finding may rest on a weak measure or an unrepresentative sample, so a better-designed replication becomes a genuine contribution. Doctoral originality usually lies in one of these three, not in an unheard-of topic.
Key idea: Good questions are manufactured from tensions, contradictions, boundary conditions, practical problems, and weak prior methods - not from waiting for an original topic.
Three families of question
- Descriptive questions ask what is happening: how prevalent, how distributed, how it varies. "What proportion of first-year graduate students report clinical-range anxiety?" A description needs a defensible sample and sound measurement but no comparison group.
- Relational questions ask how two or more variables covary: "Is screen time associated with sleep quality?" These require variation on both variables and appropriate association statistics, but association alone does not license a causal claim.
- Causal questions ask whether changing one variable produces a change in another: "Does reducing evening screen time improve sleep quality?" Causal claims demand a design that can rule out alternative explanations - ideally random assignment, or a strong quasi-experimental substitute.
Priya's three versions make the difference concrete. "What share of teens use social media after 10 p.m.?" is descriptive and needs only a good sample and a sound measure. "Is late-night use associated with depression scores?" is relational and needs variation on both variables. "Does capping evening use lower depression scores?" is causal and needs a comparison group and, ideally, random assignment. Same topic, three questions, three completely different studies and price tags.
The type of question dictates the design. A common error is to ask a causal question and answer it with a cross-sectional correlational study, then hedge with the word "impact" as if the hedge conferred causal warrant. It does not. Verbs such as affect, influence, improve, and reduce smuggle in causal claims. Only a design that can rule out alternatives supports them.
A fourth family sits alongside these three. Interpretive or process questions ask how and why something unfolds as experienced by participants: "How do first-generation students make sense of academic setbacks?" These are not lesser versions of causal questions. They target meaning and mechanism rather than magnitude. They call for qualitative designs judged by trustworthiness rather than by statistical inference. Misreading such a question as descriptive-quantitative is a frequent source of design and reviewer conflict.
Key idea: Descriptive questions need a good sample, relational questions need variation on both variables, causal questions need a comparison, and interpretive questions need open-ended data.
The grammar of a question
The wording of a question quietly encodes its paradigm and design, so it repays close editing. Quantitative stems tend to begin "To what extent," "How much," "Is there a difference in," or "Does X predict Y," and they presuppose measured variables. Qualitative stems begin "How do," "What is the experience of," or "In what ways," and they presuppose open-ended data. Mixing stems within one question ("How do students experience anxiety, and does mindfulness reduce it?") usually signals two questions in disguise.
A useful editing test is to underline the key nouns and verbs, then ask whether each one is measurable or interpretable as written. "Success," "engagement," and "wellbeing" are placeholders until you operationalize them. "Improves" and "causes" are promises about design. Force every term to be either measurable or explicitly interpretive. That single habit is often what turns a topic into a question a reader can evaluate.
Key idea: The wording of a question already commits you to a design, so edit the verbs and nouns as carefully as you would edit a hypothesis.
The FINER criteria
A widely used checklist holds that a good question is Feasible, Interesting, Novel, Ethical, and Relevant. Feasibility covers access to participants, an adequate sample, time, and skills. It is where most ambitious projects fail. Novelty does not require a wholly new topic. Replication in a new population, or resolving a contradiction in the literature, is a legitimate contribution.
Take each in turn against a concrete case, a proposal to study burnout among emergency nurses. Feasible asks whether you can actually recruit enough nurses, secure their time during shifts, and obtain hospital cooperation. Interesting asks whether the answer would matter to someone beyond you. Novel asks what the study adds that a reader could not already find. Ethical asks whether the burdens on a stressed population are justified and whether consent can be genuinely voluntary. Relevant asks whether the answer would change practice, policy, or theory.
Feasibility deserves special emphasis, because it is the criterion candidates most often flatter. Do the arithmetic. A power analysis says Priya needs 180 students. Schools have let past researchers recruit about 12% of enrolled students. Three schools with 900 students each therefore yield roughly 320 consenting families, but only if consent forms come back at the historical rate. Halve that rate and the study is underpowered before it begins. Many first drafts silently assume a response rate two or three times what the setting will yield, and the whole timeline rests on that optimism. Naming the assumption turns a hope into a number a committee can assess.
Key idea: Feasibility is arithmetic, not attitude - required sample divided by a believable recruitment rate tells you how many months the design really costs.
Sharpening with PICO
In applied and health fields the PICO frame disciplines a comparative question by forcing four elements to the surface: Population, Intervention (or exposure), Comparison, and Outcome. "Among first-year doctoral students (P), does a four-week sleep-hygiene program (I) versus a waitlist (C) reduce self-reported insomnia severity (O)?" Notice how much that specifies. It fixes the sample frame, the manipulation, the counterfactual, and the measured endpoint. Each element is now a design decision you can defend or critique.
Variants extend the frame to other designs. Adding T for time yields PICOT, which fixes the follow-up window. That matters because an effect at two weeks may vanish at six months. For observational work, exposure replaces intervention, giving PECO. For qualitative synthesis the parallel SPIDER frame substitutes sample, phenomenon of interest, design, evaluation, and research type. A manipulated intervention and a comparison group simply do not fit an interpretive question. The point is not the acronym. It is the habit of naming every element a reader would need in order to judge the study.
The C in PICO deserves emphasis, because it names the counterfactual. The counterfactual is the comparison against which any effect is judged. "Does the program reduce insomnia?" is incomplete until you say reduce compared to what. A waitlist? An active alternative? The same people beforehand? Each comparison answers a different question and carries different threats, a theme the experimental-design chapters develop at length. A question without an explicit comparison is usually a causal question missing half of its structure.
Key idea: PICO forces the four things a reader needs - who, what was done, compared with what, and measured how - and the comparison is the part most often left out.
Scope and the funnel
Think of question formation as a funnel. A broad area narrows to a specific problem. The problem narrows to a gap identified in the literature. The gap narrows to a precise question, and the question narrows to hypotheses or aims. Here is the test: can you state in one sentence exactly what observation would answer your question? If not, it is not yet sharp enough. Write that sentence before you design anything else. Every later choice of sample, measure, and analysis exists to serve it.
The one-sentence test has a useful corollary: sketch the result table before you collect data. Priya's table has two rows, capped and uncapped, and two columns, mean PHQ-A at baseline and at four weeks. Once she can draw those rows and columns, her question is operational. If you cannot draw the table, some element is still undefined. This discipline catches vague outcomes and missing comparison groups earlier than any amount of prose about statistical significance ever will.
Key idea: If you cannot sketch the table that would display your answer, your question is not finished.
From question to aims and hypotheses
Keep three related objects distinct. A research question is interrogative and names what you seek to learn. An aim or objective is a declarative statement of what the study will do to answer it. A hypothesis is a specific, testable prediction. It suits confirmatory quantitative work but is usually out of place in an exploratory qualitative study, which states aims instead of predictions.
Alignment across these three is what examiners check first. Suppose your question asks whether a program reduces insomnia. Then your aim is to evaluate that program's effect. Your hypothesis predicts the direction of the difference. Your design supports a causal comparison. And your analysis tests that difference. When any link in the chain does not match the others, the proposal reads as incoherent, no matter how polished each individual piece looks on its own.
Key idea: Question, aim, hypothesis, design, and analysis form one chain, and examiners check the links before they check the details.
Common failure modes
Several recurring faults sink otherwise promising questions. A compound question asks two things at once, as in "Does the program improve outcomes and satisfaction?" It cannot be cleanly answered or analyzed, so split it. A leading question presupposes its own answer, as in "Why is remote work more productive?" It biases the whole design, so ask whether before you ask why. An unfalsifiable question admits no observation that could count against the expected answer, which strips the study of evidential value.
Two further faults concern scope. An overbroad question such as "What causes inequality?" cannot be settled by any single study. Narrow it to a specific mechanism and population. An under-motivated question is answerable but trivial, adding a decimal place to something already established. The remedy for both is the funnel. Keep narrowing until the question is answerable, then check upward that the narrowed version still connects to something a reader has reason to care about.
Key idea: Split compound questions, de-bias leading ones, narrow overbroad ones, and check that the narrowed version still matters to someone.
Where people get stuck
Two pairs cause most of the confusion in this material.
- Topic versus question. A topic can be researched forever; a question can be answered and then closed. "Burnout in nursing" is a topic. "Do twelve-hour shifts raise emotional-exhaustion scores relative to eight-hour shifts among medical-surgical nurses over six months?" is a question. If you cannot imagine writing the sentence "the answer turned out to be...", you still have a topic.
- Aim versus hypothesis. An aim says what you will do ("to compare exhaustion scores across shift lengths"). A hypothesis says what you expect ("twelve-hour shifts will show higher exhaustion"). Exploratory qualitative studies have aims and no hypotheses, and that is correct rather than sloppy. Writing a hypothesis for a phenomenological study is a category error, not extra rigor.
A third trap is subtler. Novelty is often confused with unfamiliarity of topic. A committee does not need a subject nobody has studied. It needs a question whose answer is not already known, which includes replications in new populations and better measurement of old claims. Priya's screen-time question has been asked before, but not with an enforced cap, a comparison group, and a validated adolescent screen.
Key idea: Novelty means the answer is not yet known, not that the topic is unheard of.
Try it
Exercise. A student proposes: "This study will explore the impact of remote work on employee productivity and wellbeing." Diagnose at least three weaknesses, then rewrite it as one sharp, answerable question using an appropriate frame, and state the matching aim and hypothesis.
Model answer. The draft has several defects. "Explore ... impact" mixes an interpretive verb with a causal claim. "Productivity" and "wellbeing" are two unoperationalized outcomes bundled into one question. There is no named population, comparison, or time frame. And "remote work" is undefined as to dose. As written, no single observation could answer it.
A sharper version, using PICOT: "Among full-time knowledge workers at one firm (P), does fully remote work (I) compared with hybrid work (C) change self-reported job satisfaction on a validated scale (O) measured after six months (T)?" The matching aim is to estimate the six-month difference in job satisfaction between remote and hybrid workers. The matching hypothesis predicts a direction, for instance that remote workers report higher satisfaction. Productivity is a distinct construct, so it becomes a separate question rather than a smuggled second outcome. Each element is now defensible, measurable, and testable. A reviewer can see at a glance exactly what result would answer it.
Recap
- A topic is a field and a question is answerable; the test is whether you can state in one sentence what observation would settle it.
- Questions are manufactured from theoretical tensions, empirical contradictions, boundary conditions, practical problems, and weak prior methods.
- Descriptive, relational, causal, and interpretive questions each demand a different design, and causal verbs require a design that rules out alternatives.
- FINER checks feasibility, interest, novelty, ethics, and relevance; feasibility fails most often and should be checked with arithmetic, not optimism.
- PICO and PICOT force population, intervention or exposure, comparison, outcome, and time to the surface, and the comparison is the element most often missing.
- Question, aim, hypothesis, design, and analysis must align; exploratory qualitative work has aims and legitimately has no hypotheses.
- Compound, leading, unfalsifiable, overbroad, and under-motivated questions are the standard failure modes, and each has a specific repair.
Sources
- Farrugia, P., Petrisor, B. A., Farrokhyar, F., & Bhandari, M. (2010). Practical tips for surgical research: Research questions, hypotheses and objectives. Canadian Journal of Surgery, 53(4), 278-281. pmc.ncbi.nlm.nih.gov
- Richardson, W. S., Wilson, M. C., Nishikawa, J., & Hayward, R. S. (1995). The well-built clinical question: A key to evidence-based decisions. ACP Journal Club, 123(3), A12-A13. pubmed.ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Generating good research questions. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Korstjens, I., & Moser, A. (2017). Series: Practical guidance to qualitative research. Part 2: Context, research questions and designs. European Journal of General Practice, 23(1), 274-279. pmc.ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Pautasso, M. (2013). Ten simple rules for writing a literature review. PLoS Computational Biology, 9(7), e1003149. pmc.ncbi.nlm.nih.gov
- Andersen, H., & Hepburn, B. (2021). Scientific method. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Key terms
- Research question
- A focused, answerable interrogative that a study is designed to resolve.
- Descriptive question
- A question about the state, prevalence, or distribution of a phenomenon, needing no comparison group.
- Causal question
- A question about whether changing one variable produces change in another, requiring a design that rules out alternatives.
- FINER criteria
- A checklist for good questions: Feasible, Interesting, Novel, Ethical, Relevant.
- PICO
- A framing device specifying Population, Intervention/exposure, Comparison, and Outcome.
- Feasibility
- The practical achievability of a study given access, time, sample size, and skills.
Theory, Hypotheses, and the Null
- Formulate directional and non-directional hypotheses derived from theory.
- Explain the logic of null-hypothesis significance testing and the meaning of a p-value.
A hypothesis is a specific, testable statement about the relationship between variables, ideally deduced from a theory so that testing the hypothesis also tests the theory. This lesson ties hypothesis writing to the inferential machinery of null-hypothesis significance testing (NHST), and corrects the misinterpretations that plague reported results.
The material rewards precision because almost every quantitative results section in your field runs on this logic, and a good deal of it runs on the logic incorrectly. Learning exactly what a test does and does not license separates two kinds of researcher. One can defend a claim. The other merely reports a number a software package produced.
One study will run through the lesson. A reading program is given to 60 third-graders while 60 others get the usual lessons. At the end of the year the program group reads 4 words per minute faster on average. The test returns p = 0.03, Cohen's d = 0.31, and a 95% confidence interval from 0.5 to 7.5 words per minute. Every idea below is a way of saying precisely what that set of numbers does and does not entitle you to claim. Notice already that the interval is wide. The honest summary is "probably some benefit, size unclear", not "the program works."
Key idea: A significance test answers one narrow question, and most reporting errors come from reading a broader answer into it.
In plain terms
Here is what a significance test actually does. You assume nothing is going on. Then you ask how odd your data would look under that assumption. If the data would be very odd, you stop believing the assumption. That is all. The p-value is the answer to that one question. It is not the chance your theory is wrong. It is not the chance the finding will repeat. And it says nothing about how big the effect is. For size you need the effect size. For how sure you are, you need the interval. Two errors are possible. You can cry wolf when nothing is there. Or you can miss a real wolf. The first rate is alpha, which you choose. The second depends on your sample size, and you can only fix it before you collect data. Afterward is too late.
From construct to prediction
A theory posits relationships among abstract constructs. To test it you derive a hypothesis stated in terms of measured variables. The research (alternative) hypothesis, denoted H1, asserts that an effect or relationship exists. It may be directional ("group A scores higher than group B") when theory predicts a direction, or non-directional ("the groups differ") when it does not. A directional hypothesis is a stronger, more falsifiable claim and, correspondingly, warrants a one-tailed test only when the direction is specified in advance for principled reasons.
Consider a worked chain. A theory of self-determination holds that autonomy support raises intrinsic motivation, an abstract construct. You operationalize autonomy support as a specific classroom manipulation and intrinsic motivation as a validated scale score. The hypothesis then reads, in measured terms, that students in the autonomy-supported condition will report higher scale scores than controls. Testing that measured prediction is how you put the abstract theory at risk, which is the entire purpose of stating it.
The choice between directional and non-directional forms is not cosmetic. A one-tailed test concentrates the whole rejection region in one direction. That raises power to detect an effect in that direction. But it leaves you unable to claim anything if the result comes out the other way, however strong. Researchers are tempted to switch tails after seeing the data. So most methodologists advise two-tailed tests, unless the direction was fixed in advance for genuine theoretical reasons and stated in a preregistration.
Key idea: A hypothesis is only testable once every construct in it has become a measured variable, and a directional prediction risks more than a vague one.
The null hypothesis
The null hypothesis H0 typically states that there is no effect - no difference between groups, or no association. NHST does not attempt to prove H1 directly. Instead it asks: if H0 were true, how surprising would data at least this extreme be? That probability is the p-value. If it is small (below a pre-set threshold alpha, conventionally 0.05), we reject H0 as an implausible account of the data. This indirect logic mirrors the falsificationist stance of the previous module: we do not confirm theories, we fail to reject nulls.
Two intellectual traditions are fused, somewhat uneasily, in modern practice. Ronald Fisher introduced the p-value as a continuous measure of evidence against the null, to be weighed by judgment. Jerzy Neyman and Egon Pearson introduced a decision framework with fixed error rates, alpha and beta, and the accept-or-reject dichotomy. The hybrid taught today borrows Fisher's p-value and the Neyman-Pearson threshold, which is why the logic can feel internally strained. Knowing the two sources clarifies many debates you will meet in the reform literature.
Why reject rather than accept? The asymmetry is deliberate and mirrors falsification. A single set of data consistent with the null does not establish that the effect is exactly zero, because low power or a small sample could hide a real effect. But data very unlikely under the null give grounds to abandon it. The reasoning resembles proof by contradiction: assume no effect, show the data would be surprising under that assumption, and conclude that the assumption is doubtful. We never earn the right to declare the null true.
A standard criticism deserves airing. In much research the null is a nil null: the claim that the effect is exactly zero. In the messy social world that is almost never literally true. If the null is false by default, then rejecting it with a large enough sample is close to guaranteed, and therefore uninformative. This is one reason the field increasingly asks two further questions. How large is the effect? And is it large enough to matter to theory or practice?
Key idea: The test never proves there is no effect; it only asks how surprising your data would be if there were none.
What a p-value is not
The p-value is among the most misreported quantities in science. It is not the probability that the null hypothesis is true. It is also not the probability that the results occurred by chance. Formally, it is the probability of obtaining a test statistic at least as extreme as the one observed, assuming H0 is true. Three consequences follow.
- A non-significant result (p greater than alpha) does not prove the null. Absence of evidence is not evidence of absence.
- Statistical significance is not the same as practical importance. With a large enough sample, a trivial effect can be significant.
- A significant result does not report the size of the effect. Always accompany p-values with an effect size and a confidence interval.
Return to the reading program, which reported p = 0.03. The correct reading is this. If the program truly had no effect, data at least this extreme would turn up about 3 times in 100 such studies. It is wrong to say there is a 3% chance the null is true. It is wrong to say there is a 97% chance the finding is real. And it is wrong to say there is a 97% chance it will replicate. The p-value is a statement about data under an assumption. It says nothing about the probability of the assumption itself.
In 2016 the American Statistical Association took the unusual step of issuing a formal statement on p-values. It warned that p-values do not measure the probability that a hypothesis is true, nor the importance of a result. It also warned that scientific conclusions should not rest on whether a p-value passes a threshold. The statement did not ban the tool. It insisted the tool be understood. The practical upshot for your writing is simple: report exact p-values, effect sizes, and intervals together, and interpret them jointly rather than ceremonially.
Key idea: A p-value tells you how odd your data would look if nothing were going on - never how likely your hypothesis is, and never how big the effect is.
Two kinds of error
| H0 actually true | H0 actually false | |
|---|---|---|
| Reject H0 | Type I error (alpha) | Correct decision (power) |
| Fail to reject H0 | Correct decision | Type II error (beta) |
A Type I error is a false positive: rejecting a true null. Its long-run rate is set by alpha. A Type II error is a false negative: failing to detect a real effect. Its rate is denoted beta. Statistical power equals 1 minus beta, the chance of detecting an effect that truly exists. Power rises with larger samples, larger true effects, and lower measurement error. Planning the sample size to reach adequate power, commonly 0.80, is a hallmark of rigorous confirmatory design. It must be done in advance, because power calculated after a null result is uninformative.
Put numbers on it. If the reading program's true effect is d = 0.3, then 60 children per group gives roughly 45% power. In other words, a real effect of that size would be missed more often than found. Doubling to 120 per group lifts power to about 78%. That single calculation, done before recruitment, is the difference between a study that can answer its question and a study that cannot.
Alpha and beta trade off against each other for a fixed sample. Lowering alpha to 0.01 to reduce false positives makes the test more conservative and, all else equal, raises beta, increasing the risk of missing a real effect. The only way to reduce both at once is to buy more information, chiefly through a larger sample or a more precise measure. This is why sample-size planning is not bureaucratic box-ticking but the mechanism that sets both error rates at acceptable levels before a single participant is enrolled.
The post-hoc power fallacy is worth stating plainly. After a non-significant result, computing power from the observed effect tells you nothing new. It is merely a re-expression of the p-value, and it will always look low. Power analysis is a design tool. Run it before data collection, using the smallest effect size that would interest you. Do not run it afterward to excuse a null. A committee that sees observed power offered as an explanation for a null result reads it as a misunderstanding of what power means.
Key idea: Alpha sets your false-positive rate and power sets your chance of finding a real effect, and only a bigger or better-measured sample improves both at once.
Effect size and estimation
An effect size expresses the magnitude of a relationship on a scale independent of sample size. For a difference between two means, Cohen's d reports the gap in standard-deviation units. A d of 0.5 means the groups differ by half a standard deviation. Correlations, odds ratios, and variance-explained measures play the same role for other designs. Significance mixes size with sample size, so it cannot answer "how much." The effect size can, and that is usually the question a decision-maker cares about most.
Conventional labels such as small, medium, and large are useful starting points. Interpret them against the effects typical in your own research area rather than applying them mechanically. An effect that counts as small in laboratory cognition may be substantial for a public-health intervention delivered to millions. The reading program's d of 0.31 is modest in the lab. Across a district of 20,000 children, at a cost of a few dollars per pupil, it may be excellent value.
Many methodologists now favor an estimation approach. It leads with effect sizes and confidence intervals rather than with a reject-or-not verdict. A 95% confidence interval conveys two things at once: a best estimate, and the precision around it. A very wide interval warns that the study has learned little, significant or not. Our reading study is a case in point. Its interval runs from 0.5 to 7.5 words per minute, so the data are compatible with a benefit too small to notice and with one a teacher would welcome. Reporting the interval also supports cumulative science, because later reviews can pool intervals across studies in a meta-analysis. A bare significant-or-not verdict can never be pooled that way.
The reform conversation has produced several alternatives you should recognize. Some methodologists propose lowering the default threshold to 0.005 for claims of a new discovery, reserving 0.05 for suggestive evidence. Others advocate a Bayesian framework, which updates the probability of a hypothesis given the data and prior information. That framework answers the question people wrongly think a p-value answers. None of this is required for your dissertation. A defensible results section does need to show that you chose your framework deliberately rather than by inertia.
Key idea: Lead with the effect size and its interval, because they say how much and how sure, which is what a p-value can never tell you.
Where people get stuck
Four confusions account for nearly all misreported results.
- "Not significant" versus "no effect." A p of 0.21 says the data are unsurprising under the null. It does not say the effect is zero. Only a narrow interval sitting close to zero, or a formal equivalence test, supports a claim of no meaningful effect.
- "Significant" versus "important." With 20,000 children, a gain of 0.3 words per minute could be significant and useless. With 40 children, a gain of 8 words per minute could be non-significant and worth acting on. Size and certainty are separate questions.
- Alpha versus p. Alpha is a threshold you choose before seeing data; p is a quantity the data produce. Adjusting alpha after the fact, or reporting p less than 0.05 when p was 0.049 and the plan said 0.01, quietly breaks the logic.
- Power before versus after. Power computed in advance from the smallest effect worth detecting is a design decision. Power computed afterward from the observed effect is just the p-value wearing a disguise.
Key idea: Keep four pairs apart - not significant is not no effect, significant is not important, alpha is not p, and planned power is not observed power.
Try it
Exercise. A colleague writes: "We found no significant difference between the training and control groups (p equal to 0.21), which proves the training has no effect. Post-hoc power was 0.34." Identify the two inferential errors, then describe what the study should report and what it should do next.
Model answer. The first error is treating a non-significant result as proof of no effect. A p of 0.21 means the data are not surprising under the null. It does not mean the null is true. Absence of evidence is not evidence of absence, and a small or imprecise study can easily miss a real effect. The second error is offering post-hoc power as reassurance. Observed power is a transform of the p-value and is uninformative by construction.
The study should report the observed effect size with its 95% confidence interval, which honestly displays how much remains uncertain. Suppose that interval runs from a trivial value to a meaningful one. Then the correct conclusion is that the study was inconclusive, not that the effect is absent. Two next steps are defensible. One is a properly powered replication, with the sample size set in advance from the smallest effect worth detecting. The other is a formal equivalence test, which asks whether any effect is small enough to be treated as negligible.
Recap
- A hypothesis restates a theory's claim in measured variables, and a directional prediction forbids more than a vague one.
- NHST asks how surprising the data would be if the null were true, so it can reject a null but never establish one.
- Modern practice fuses Fisher's continuous p-value with the Neyman-Pearson accept-reject framework, which is why the logic feels internally strained.
- A p-value is not the probability that the null is true, not the probability of chance, and not a measure of effect size or replicability.
- Type I error is a false positive at rate alpha, Type II error is a false negative at rate beta, and power equals 1 minus beta.
- Power must be planned in advance from the smallest effect worth detecting; observed power after a null result is just the p-value restated.
- Report effect size and confidence interval alongside p, because a wide interval says the study is inconclusive whatever the p-value shows.
Sources
- Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. pmc.ncbi.nlm.nih.gov
- Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133. amstat.org
- Amrhein, V., Greenland, S., & McShane, B. (2019). Scientists rise up against statistical significance. Nature, 567(7748), 305-307. nature.com
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Understanding null hypothesis testing. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Using theories in psychological research. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Sprenger, J., & Hartmann, S. (2020). Philosophy of statistics. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. pubmed.ncbi.nlm.nih.gov
- Key terms
- Hypothesis
- A specific, testable statement about the relationship between measured variables.
- Null hypothesis (H0)
- The statement of no effect or no difference that significance testing attempts to reject.
- p-value
- The probability of data at least as extreme as observed, assuming the null hypothesis is true.
- Type I error
- A false positive: rejecting a null hypothesis that is actually true; its rate is alpha.
- Type II error
- A false negative: failing to reject a null hypothesis that is actually false; its rate is beta.
- Statistical power
- The probability (1 minus beta) of detecting an effect that genuinely exists.
The Literature Review as Argument
- Distinguish narrative, systematic, and scoping reviews by purpose and method.
- Structure a literature review as a synthesized argument that motivates a specific gap.
A literature review is not an annotated bibliography and not a chronological march through everything ever published. It is an argument: a synthesis of prior work organized to demonstrate that your question is unanswered, answerable, and worth answering. This lesson distinguishes review types and lays out how to build the argument rigorously.
Think of the review as the case for your study's necessity. Every source you cite is evidence marshaled toward a verdict: that a specific gap exists and that your design is the right instrument to close it. Read this way, the review is not a preliminary chore to clear before the real work. It is the foundation on which the whole project's claim to originality rests.
One example runs through this lesson. A candidate, Tomas, wants to study whether peer mentoring keeps first-generation students enrolled. He finds 31 studies. Three report large gains, four report nothing, and one reports a small loss. The weak version of his review lists all 31 in turn. The strong version notices something: every study reporting a gain lasted at least two semesters, while every null study lasted one. That observation is worth more than the other 30 summaries combined, because it turns a pile of citations into a claim about why the literature disagrees.
Key idea: A review is an argument that your question is unanswered, answerable, and worth answering - not a catalogue of what others have written.
In plain terms
A literature review has one job. Convince the reader your study needs doing. Everything else follows from that. First pick your type. A narrow question in a settled field gets a systematic review. A messy new field gets a scoping review. Second, search properly. List your search terms, name your databases, note the dates, and follow the reference lists both backward and forward. Write it all down so someone else could repeat it. Third, and this is the part people skip, write about the literature rather than about each paper. If your paragraph starts with an author's name, you are summarizing. If it starts with a claim and cites authors as proof, you are synthesizing. Fourth, end on a gap with teeth. "Nobody has done this" is weak. "People are making a decision on evidence that points the other way" is strong.
Types of review
- A narrative (traditional) review surveys a body of work to characterize its themes and debates. It is flexible and interpretive but vulnerable to selection bias, because the author chooses what to include without a pre-registered protocol.
- A systematic review answers a specific question using an explicit, reproducible protocol. That means pre-specified search strings across named databases, transparent inclusion and exclusion criteria, and often a formal quality appraisal of each study. Reporting standards such as PRISMA require you to document how many records were identified, screened, and excluded, and why. When studies are statistically pooled, the systematic review becomes a meta-analysis.
- A scoping review maps the extent, range, and nature of evidence on a broad topic - useful when a field is too heterogeneous or immature for a narrow systematic question. It charts what exists rather than adjudicating an effect.
Several other forms fill particular niches. An integrative review combines empirical and theoretical literature to generate a new conceptual framework. A rapid review streamlines the systematic method under time pressure, trading some comprehensiveness for speed. An umbrella review synthesizes existing systematic reviews when a field already has many. For qualitative evidence, a meta-synthesis integrates findings across studies while respecting their interpretive character rather than pooling numbers.
Choosing among them follows from your question and the field's maturity. A narrow, well-defined effect in a mature literature calls for a systematic review or meta-analysis. A broad, emerging area with mixed methods calls for a scoping review that maps the terrain first. Tomas has 31 studies with wildly different designs and outcome measures, so a scoping review is defensible and a meta-analysis would pool apples with oranges. Matching the review type to the state of the evidence is a methodological decision your committee will expect you to justify, not a habit to fall into.
Key idea: Pick the review type from the state of the evidence - systematic for a settled question, scoping for a messy field, meta-analysis only when studies are comparable enough to pool.
Searching systematically
Even a narrative review benefits from a disciplined search. Identify the concepts in your question. List synonyms and controlled vocabulary for each. Then combine them with Boolean operators, using OR within a concept and AND across concepts. Tomas searches ("peer mentoring" OR "peer coaching") AND ("first-generation" OR "first generation college") AND (retention OR persistence OR "drop out"). Run the search in more than one database, because coverage differs, and record the exact strings and dates. This turns an invisible, idiosyncratic process into one a reader could reproduce, which is the minimum standard of transparency.
Two complementary techniques catch what keyword searches miss. Backward snowballing mines the reference lists of key papers for earlier work. Forward snowballing uses citation indexes to find later work that cited them. Systematic reviews then document the funnel from records identified to studies included, using a PRISMA flow diagram that states how many were screened and why each exclusion occurred. The discipline is not bureaucratic. It is what stops a search from quietly becoming a hunt for confirming citations.
Key idea: Write down your search strings, databases, and dates, because a search a reader cannot reproduce is an opinion rather than evidence.
Synthesis, not summary
The defining move of a strong review is synthesis. You group studies by construct, method, or finding, then say what the body of work collectively shows. Where does it agree? Where does it conflict? And why? A telltale sign of a weak review is the paragraph that runs "study X found... study Y found... study Z found..." It summarizes serially without integrating anything. Organize thematically or conceptually, never one source per paragraph.
A practical device is the synthesis matrix, a grid whose rows are studies and whose columns are the constructs, methods, or findings you care about. Reading down a column rather than across a row forces integration, because you are comparing how many studies bear on one theme rather than narrating one study at a time. The prose that results speaks about the theme first and cites studies as evidence for claims about it, which is the grammatical signature of genuine synthesis.
Organization can be thematic, methodological, or theoretical, depending on your argument. A thematic structure groups findings by substantive topic. A methodological structure contrasts what different designs have found and why they might disagree. A theoretical structure arrays studies by the framework they assume. Tomas needs the methodological structure, because his real finding is that program duration tracks the result. The right choice is the one that makes your eventual gap visible, so reverse-engineer the structure from the argument it must deliver.
Key idea: Synthesis means writing about themes and citing studies as evidence, not writing about studies one at a time.
Theoretical framing
Doctoral reviews are expected to do more than catalogue findings. They position the study within a theoretical or conceptual framework that names the constructs and their presumed relationships. The framework tells the reader which lens you will look through and why. It also disciplines the review by giving it an organizing spine. A study of technology adoption framed by a theory of acceptance reads very differently from one framed by practice theory. The review should make that commitment explicit rather than leaving it implicit in the choice of citations.
The framework also constrains what counts as a relevant finding. Suppose your lens foregrounds social influence. Then studies of individual cognition are context rather than core, and the review can say so instead of treating all adjacent work as equally central. Naming the framework early does double duty. It justifies your selection of sources, and it previews the variables and relationships your own design will examine.
Key idea: Name your framework early; it is the spine that decides which studies are central and which are merely nearby.
Identifying a defensible gap
The review must land on a gap. But "no one has studied exactly this" is the weakest kind of gap, and it often signals that the question is unimportant. Four kinds are stronger. A contradiction in the evidence that your study can resolve. A population or context to which an established finding has never been extended. A methodological limitation shared across prior studies that a better design would overcome. Or a theoretical tension that competing frameworks predict differently. State the gap explicitly, and show that answering it advances knowledge rather than merely filling a blank.
Apply a blunt "so what" test to any candidate gap. If a skeptical reader could reply that the blank is empty because nobody needed it filled, the gap is not yet defensible. A strong gap statement pairs the absence with a consequence. Because this contradiction is unresolved, practitioners cannot choose between two treatments. Because this effect has never been tested in this population, a policy is being applied on untested assumptions. Tomas can say something sharper than "no one has studied rural campuses." He can say that colleges are buying one-semester mentoring programs, and the evidence suggests one semester is exactly the dose that does not work. The consequence is what converts a gap into a rationale.
Key idea: An empty cell is not a gap - a gap is an absence with a consequence attached.
Appraising quality, not just counting
Counting studies is not the same as weighing them. A rigorous review appraises each source for internal validity, using structured tools suited to the design, and lets stronger studies carry more argumentative weight. Evidence hierarchies place systematic reviews and randomized trials above observational and case studies. That ranking is a useful heuristic for questions of effectiveness. It misleads when the question is about meaning or mechanism, where a well-conducted qualitative study is the higher-quality evidence.
Frameworks such as GRADE formalize this. GRADE rates a body of evidence for risk of bias, consistency, directness, and precision, then reports a confidence level in the overall conclusion. You need not adopt GRADE wholesale, but its logic should inform your prose. State not only what studies found but how much trust their designs warrant. A review that treats a small unblinded pilot and a large preregistered trial as equal voices has given up its critical function.
Key idea: Weigh studies, do not count them, and remember that the top of the evidence hierarchy depends on whether your question is about effect or about meaning.
Managing bias and the corpus
Two hazards deserve attention. The first is publication bias. Studies with significant, positive results are more likely to be published, so a review of the published literature alone can overstate an effect. Systematic reviewers counter this by searching grey literature and trial registries. The second hazard is the temptation to cite only work that supports your thesis. A credible review engages seriously with contrary findings and explains them. Keep a transparent record of search terms, dates, and decisions, so a reader could in principle reconstruct how you arrived at your corpus. That transparency separates a scholarly review from an opinion essay with footnotes.
Publication bias can be probed as well as guarded against. In a meta-analysis a funnel plot displays each study's effect against its precision. A symmetric funnel is reassuring. A gap where small null studies should sit hints that they were never published. Related diagnostics examine the distribution of p-values just below the significance threshold for signs of selective reporting. You are not expected to run these for a narrative review. Understanding them keeps you properly skeptical of a suspiciously tidy literature.
The subtler hazard is your own confirmation bias. Once you have invested in a thesis, contrary studies become easy to overlook or explain away. A credible review does the opposite. It seeks out the strongest disconfirming work and engages it on the merits. Readers trust an argument more when it has clearly survived contact with its best objections. An auditable log of search terms, dates, and inclusion decisions is the procedural safeguard that makes this discipline visible to others.
Key idea: Published findings skew positive, and so will you - the defenses are grey literature, registries, an audit trail, and deliberate engagement with contrary work.
Positioning your study
The review's final paragraphs should hand off to your study as the natural next move. You have established what is known, where it conflicts, and what remains unaddressed. Now name the specific question that follows and preview the design that answers it. Done well, a reader reaches your methods section already convinced the study is necessary. The review and the study are one argument in two acts, not a survey with an unrelated project bolted on afterward.
Key idea: The last paragraph of the review should make your methods section feel inevitable.
Where people get stuck
Three difficulties recur.
- Summary versus synthesis. Check the grammar of your sentences. If most begin with an author's name, you are summarizing. If most begin with a claim about the literature and cite authors as support, you are synthesizing. Tomas's sentence "programs lasting two or more semesters show gains, while single-semester programs do not (Lee, 2019; Ortiz, 2021; Chen, 2022)" is synthesis.
- Systematic versus scoping. A systematic review answers one narrow question and can be pooled; a scoping review maps a whole territory and cannot. Choosing systematic for an immature literature produces a review with four includable studies, which impresses nobody.
- Gap versus absence. Every possible study is absent until someone does it. A gap needs a stake: a decision someone is making badly, a contradiction blocking practice, or a theory that has never been tested where it would most likely fail.
Key idea: Ask of every paragraph whether it makes a claim about the literature; if it only reports what one author did, it belongs in a table instead.
Try it
Exercise. A draft review ends: "In summary, Smith (2019) found a positive effect, Jones (2020) found none, and Lee (2021) found a negative effect. No study has examined this in rural clinics. The present study will do so." Diagnose the weaknesses in both the synthesis and the gap, then rewrite the closing as a defensible argument.
Model answer. The synthesis is serial summary, not integration: it lists three discrepant findings without asking why they differ. The gap is the weakest kind, a mere empty cell ("no study has examined rural clinics"), with no consequence attached. Nothing tells the reader why the disagreement matters or why the rural setting would change anything about it.
A stronger closing integrates and motivates. It might read like this. Across these studies the effect ranges from positive to negative, and the pattern tracks a methodological difference. The two positive studies used clinician-rated outcomes, while the null and negative studies used self-report. That suggests the discrepancy may be an artifact of measurement rather than a true absence of effect.
The rewrite then attaches a consequence and a design. Rural clinics rely disproportionately on self-report, so the ambiguity has direct practical stakes there. The present study therefore tests the effect in rural clinics using both outcome types, so that measurement is no longer confounded with setting. This version explains the conflict, ties the gap to a consequence, and shows how the design resolves it. Serial summary and an empty-cell gap could do none of those things.
Recap
- A literature review is an argument for the necessity of your study, and every citation is evidence marshaled toward that verdict.
- Narrative, systematic, scoping, integrative, rapid, umbrella, and meta-synthesis reviews serve different purposes, and the choice must be justified by the state of the evidence.
- A disciplined search uses Boolean logic within and across concepts, more than one database, backward and forward snowballing, and a recorded audit trail.
- Synthesis groups studies by theme, method, or theory and explains disagreement; serial summary is the signature of a weak review.
- A theoretical framework gives the review its spine and decides which findings are central rather than merely adjacent.
- The strongest gaps are contradictions, untested populations, shared methodological weaknesses, and theoretical tensions - each paired with a consequence.
- Appraise quality rather than counting studies, and guard against publication bias and your own confirmation bias with registries, grey literature, and engagement with contrary work.
Sources
- Pautasso, M. (2013). Ten simple rules for writing a literature review. PLoS Computational Biology, 9(7), e1003149. pmc.ncbi.nlm.nih.gov
- Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., ... Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. pmc.ncbi.nlm.nih.gov
- Baethge, C., Goldbeck-Wood, S., & Mertens, S. (2019). SANRA: A scale for the quality assessment of narrative review articles. Research Integrity and Peer Review, 4, 5. pmc.ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Reviewing the research literature. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- EQUATOR Network. (2025). Reporting guidelines for main study types. EQUATOR Network, University of Oxford. equator-network.org
- Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. pmc.ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Key terms
- Literature review
- A synthesized argument from prior work that motivates and situates a study's question.
- Systematic review
- A review answering a specific question via an explicit, reproducible protocol and appraisal.
- Meta-analysis
- A quantitative pooling of results from multiple studies to estimate an overall effect.
- Scoping review
- A review mapping the extent, range, and nature of evidence on a broad topic.
- Synthesis
- Integrating studies to show collective patterns, agreements, and conflicts rather than summarizing serially.
- Publication bias
- The tendency for statistically significant, positive results to be preferentially published.
Module 3: Research Ethics and Human Subjects
The moral and regulatory framework governing research with people, from foundational principles to IRB review.
Foundations: From Nuremberg to the Belmont Report
- Trace the historical abuses that produced modern research-ethics regulation.
- State the three Belmont principles and connect each to a concrete application.
The rules that govern research with human participants were written in response to real harm. Understanding that history is not ceremonial; it explains why the requirements exist and guards against treating them as bureaucratic friction to be minimized.
Grasp the reasoning behind a safeguard and you can apply it to a situation nobody anticipated. Memorize only the forms and you are lost the moment a study departs from the template. This lesson traces the landmark cases and the principles they produced, so that the ethics of your own design rest on understanding rather than rote compliance.
Here is the test case to keep in mind. Suppose you want to know whether a cash transfer reduces school dropout. The cheapest way to run it is to offer money to families in a shelter, who will almost certainly say yes, and to withhold the program from a comparison group who would also benefit. Nothing in that plan involves a needle or a drug. Every principle in this lesson still applies to it. Ethics is not a medical specialty. It is a set of questions about who bears the burden of finding something out.
Key idea: Every rule in research ethics exists because someone was harmed, so knowing the history tells you what each rule is trying to prevent.
In plain terms
Here is the whole lesson in short sentences. Doctors in Nazi camps ran deadly experiments on prisoners. The trial that followed produced the Nuremberg Code in 1947. Its first rule was that people must agree freely to take part. Doctors then wrote their own version, the Declaration of Helsinki, in 1964. It said the patient comes before the study. Meanwhile the United States ran its own scandal. For forty years, doctors watched syphilis go untreated in Black men and lied to them about it. That was Tuskegee. When the public found out, Congress acted. A commission wrote the Belmont Report in 1979. It gave three rules. Respect people and let them choose. Weigh the harm against the good. Share the burden fairly. Those three rules became federal law. Everything else in this lesson is detail about how they work.
The historical record
The Nuremberg Code (1947) emerged from the trials of physicians who conducted lethal experiments on concentration-camp prisoners. Its first and central principle is that voluntary consent of the human subject is absolutely essential. The Declaration of Helsinki was first adopted by the World Medical Association in 1964 and has been revised repeatedly since. It extended ethical guidance for medical research and insisted on a priority: the well-being of the individual takes precedence over the interests of science and society.
The Nuremberg Code set out ten points. Voluntary consent is the first. The Code also required that an experiment yield fruitful results unobtainable by other means, that it avoid unnecessary suffering, and that the subject be free to withdraw at any time. The Code arose from a criminal tribunal rather than a professional body. It therefore carried great moral authority but almost no enforcement machinery, which is part of why further instruments followed over the next decades.
In the United States, the pivotal domestic scandal was the Tuskegee syphilis study. From 1932 to 1972 the U.S. Public Health Service observed the untreated progression of syphilis in Black men. Penicillin became the standard cure in the 1940s, and it was withheld. Participants were deceived about their own condition for four decades. Public revelation of Tuskegee led to the National Research Act of 1974, and to the commission that produced the Belmont Report (1979), the ethical cornerstone of U.S. human-subjects regulation.
Tuskegee was not an isolated aberration. In 1963 physicians at the Jewish Chronic Disease Hospital injected live cancer cells into debilitated patients without consent. From the 1950s into the 1970s the Willowbrook studies deliberately infected institutionalized children with hepatitis to study the disease. In 1966 the anesthesiologist Henry Beecher catalogued twenty-two such ethically troubling studies in a leading medical journal, showing that violations were widespread rather than exceptional. This accumulating record, not any single event, forced systemic reform.
A related case concerns tissue rather than treatment. In 1951 cells were taken from Henrietta Lacks, a Black woman treated for cervical cancer. They were cultured and distributed worldwide without her knowledge or her family's, becoming the HeLa cell line used in laboratories everywhere. The case shows that consent and justice reach beyond bodily risk to the use of biological materials and data. Current regulation on identifiable specimens is still working out that problem.
These cases are not safely in the past. In 2014 a large social-media experiment manipulated the emotional content of users' feeds without their knowledge. It provoked a lasting debate about consent for online behavioral research. Members of the Havasupai Tribe learned that blood samples given for a diabetes study had been used for unrelated genetic research they found objectionable, and the dispute was settled in 2010. Each case shows the same principles reasserting themselves in new technological form. That is why understanding the reasoning matters more than memorizing any single regulation.
Key idea: Nuremberg answered coerced experiments, Helsinki put the individual ahead of science, and Tuskegee produced the Belmont Report and U.S. regulation.
The three Belmont principles
- Respect for persons. Individuals are autonomous agents entitled to make their own decisions, and persons with diminished autonomy are entitled to protection. In practice this principle grounds informed consent and special safeguards for vulnerable populations.
- Beneficence. Researchers must maximize possible benefits and minimize possible harms. This principle grounds the risk-benefit analysis that sits at the center of ethical review; a study is justifiable only when its risks are reasonable in relation to its anticipated benefits.
- Justice. The benefits and burdens of research must be distributed fairly. No group should bear the risks while another reaps the rewards. This principle grounds equitable selection of participants and speaks directly to why Tuskegee, which concentrated harm on an already-marginalized population, was so grievous a violation.
Each principle translates into a concrete obligation. That is what makes Belmont operational rather than merely aspirational. Respect for persons yields the consent process and the protection of those who cannot fully consent. Beneficence yields the systematic weighing of risks against benefits. Justice yields scrutiny of who is recruited and who ultimately gains. It forbids designs that load risk onto the powerless and deliver benefit to the comfortable.
Run the cash-transfer study past all three. Respect for persons asks whether families in a shelter can decline freely when the researcher also controls something they need. Beneficence asks whether the study risks stigma or loss of benefits, and whether those risks have been cut to the minimum. Justice asks whether the families bearing the burden are the ones who would get the program if it works. Notice that each principle asks a different question, and that a design can pass one while failing another.
The three principles are deliberately left unranked. Belmont does not say that respect always outranks beneficence, or that justice always outranks respect. Real cases require judgment about which consideration weighs most in the situation at hand. That refusal to supply a mechanical priority is a feature rather than a gap. It obliges researchers and boards to reason openly about tradeoffs instead of hiding behind a formula. It is also why two thoughtful reviewers can reach different but defensible conclusions about the same protocol.
Key idea: Respect for persons means consent, beneficence means weighing harm against benefit, and justice means fair distribution of who is put at risk.
From principles to regulation
Belmont is a statement of principle; its force in the United States comes through federal regulation. The relevant rules are codified at 45 CFR 46, and the core subpart adopted across many federal agencies is known as the Common Rule, most recently revised in 2018. Oversight rests with the Office for Human Research Protections, and local enforcement rests with the institutional review boards you will meet in the next lesson. The Food and Drug Administration maintains parallel rules for research on regulated products.
The Belmont Report itself named three applications that map onto the three principles, and the regulations elaborate each. Informed consent operationalizes respect for persons. The assessment of risks and benefits operationalizes beneficence. The equitable selection of subjects operationalizes justice. When you complete a protocol form, its sections are not arbitrary. Each one exists to document that one of these principles has been honored in the specifics of your design.
Key idea: Belmont supplies the principles, 45 CFR 46 and the Common Rule supply the enforceable rules, and your protocol form is where the two meet.
Weighing risks and benefits
Beneficence sits at the center of review, so it is worth seeing how a board actually reasons about it. Risk is assessed along two dimensions: how likely a harm is, and how bad it would be. It is also assessed across several kinds of harm, including physical, psychological, social, economic, and legal. Consider a breach of confidentiality that exposes a participant's immigration status or drug use. That can be far more damaging than a passing physical discomfort. Informational risk is therefore taken seriously even in studies that never touch the body.
Benefits are assessed just as carefully, and a common error is to inflate them. Benefits to the individual participant are kept separate from benefits to society. Payment is explicitly not counted as a benefit that offsets risk, because that would let money purchase the acceptability of danger. The board then asks one further question. Has the design cut risk as far as it can while still answering the question? An avoidable risk is never justified by a large benefit.
Key idea: Risk includes social, legal, and informational harm, benefits to society are not benefits to the participant, and payment never buys the acceptability of a risk.
International and disciplinary scope
Ethical governance is not uniquely American, and your obligations depend on where and with whom you work. The Declaration of Helsinki governs much biomedical research internationally, the CIOMS guidelines address research in resource-limited settings, and many countries maintain their own national frameworks and review bodies. Multi-site and cross-national studies must satisfy the standards of every jurisdiction involved, which occasionally conflict and must then be reconciled explicitly in the protocol.
Disciplines differ too. The foundational cases are biomedical, so some requirements fit awkwardly onto social, behavioral, and educational research, where risks are often informational rather than physical. Regulations therefore distinguish levels of risk and provide for exemptions and expedited handling, discussed next. The principles are constant even so. An ethnographer and a surgeon owe their participants the same respect, beneficence, and justice, even though the concrete risks and safeguards differ in kind.
Key idea: The principles travel across countries and disciplines even when the specific rules and risk categories do not.
A second domain: research integrity
Human-subjects protection is one branch of research ethics. Research integrity is the other, and the two are often confused. Integrity concerns honesty in the conduct and reporting of research. Its cardinal violations are three. Fabrication is inventing data. Falsification is altering data or results. Plagiarism is appropriating others' work. These are governed not by the Belmont principles but by institutional and funder misconduct policies, and a finding of misconduct can end a career.
The domains intersect in practice. Fabricating outcomes in a clinical trial harms future patients as well as the scientific record. Concealing a conflict of interest can distort a risk-benefit judgment. A complete research ethics therefore covers two things: how you treat participants, and how honestly you handle what you learn from them. A strong dissertation attends to both rather than assuming that board approval exhausts its obligations.
Key idea: Human-subjects protection and research integrity are separate systems, and approval from one says nothing about compliance with the other.
Principles to applications
| Principle | Core commitment | Primary application |
|---|---|---|
| Respect for persons | Autonomy and protection of the vulnerable | Informed consent |
| Beneficence | Maximize benefit, minimize harm | Risk-benefit assessment |
| Justice | Fair distribution of benefits and burdens | Equitable participant selection |
Notice that the principles can pull against one another. A study promising large social benefit satisfies beneficence. The easiest way to run it may be to recruit a captive or desperate population, which violates justice. It may also work better scientifically without full disclosure, which strains respect for persons. Ethical review is the structured negotiation of these tensions, not the mechanical application of a single rule. Keep all three principles in view as you design, because they are the criteria a board will apply to your protocol.
Key idea: The principles conflict on purpose, and your job in a protocol is to show which tradeoff you made and why.
Where people get stuck
Three points cause repeated trouble.
- Consent is not a form. A signature documents a process; it does not constitute one. If a participant cannot explain what will happen to them, no signature repairs that. Tuskegee's participants signed nothing meaningful, but the deeper wrong was that they were told they were being treated.
- Voluntary is not the same as willing. Families in a shelter may enthusiastically agree to a study because they need what the researcher controls. Eagerness can be evidence of undue inducement rather than of free choice. The test is whether a reasonable person in that position could comfortably decline.
- Minimal risk is not zero risk. The regulatory category means risk no greater than that of ordinary daily life. It does not mean nothing could go wrong, and it does not exempt a study from confidentiality safeguards.
One further distinction is worth fixing. Beneficence weighs risk against benefit; justice asks who receives each. A study can be perfectly balanced overall and still be unjust, if the risk sits with one group and the benefit with another. That was the precise failure at Tuskegee, and it is the failure the cash-transfer design above would repeat.
Key idea: Beneficence asks whether the tradeoff is worth it; justice asks who is on each side of the trade.
Try it
Exercise. A pharmaceutical company proposes to test a new but risky drug by recruiting participants from a low-income community that lacks other access to health care, offering free medical examinations as the inducement. The drug, if approved, will be priced beyond that community's reach. Identify which Belmont principles are in tension and how the design might be revised.
Model answer. The design most sharply violates justice: the burdens of risky research fall on a marginalized group that will not share in the benefit, since the approved drug will be unaffordable to them. The same feature strains respect for persons, because free care offered to people with no alternative can function as undue inducement that compromises the voluntariness of consent. Beneficence is engaged as well, since the risk-benefit balance for these particular participants is poor.
A revision would address each principle in turn. Justice improves if the study recruits from populations who could plausibly benefit from the eventual product, and if provisions are made for post-trial access. Voluntariness improves in two ways. Set compensation at reasonable reimbursement rather than an offer that overwhelms judgment. And do not make essential care contingent on participation. The risk-benefit balance improves with additional safety monitoring and clear withdrawal rights. The point is not that the study is impossible. Its ethics turn on the distribution of risk and benefit, which is exactly the lesson Tuskegee wrote into the regulations.
Recap
- The Nuremberg Code (1947) made voluntary consent absolutely essential but had moral authority without enforcement machinery.
- The Declaration of Helsinki (1964 onward) established that the individual's well-being takes precedence over the interests of science and society.
- Tuskegee, Willowbrook, the Jewish Chronic Disease Hospital case, and Beecher's 1966 catalogue showed abuse was systemic, which produced the National Research Act and the Belmont Report.
- Belmont's three principles are respect for persons, beneficence, and justice, applied as informed consent, risk-benefit assessment, and equitable selection.
- The principles are deliberately unranked, so ethical review is reasoned negotiation of tradeoffs rather than a formula.
- Risk includes psychological, social, economic, legal, and informational harm, and payment is never counted as a benefit that offsets risk.
- Research integrity, covering fabrication, falsification, and plagiarism, is a separate system from human-subjects protection, and both bind you.
Sources
- National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (1979). The Belmont Report: Ethical principles and guidelines for the protection of human subjects of research. U.S. Department of Health and Human Services. pubmed.ncbi.nlm.nih.gov
- United States Holocaust Memorial Museum. (2025). The doctors trial: The medical case of the subsequent Nuremberg proceedings. Holocaust Encyclopedia. ushmm.org
- World Medical Association. (2024). WMA Declaration of Helsinki: Ethical principles for medical research involving human participants. World Medical Association. wma.net
- Brandt, A. M. (1978). Racism and research: The case of the Tuskegee Syphilis Study. Hastings Center Report, 8(6), 21-29. pubmed.ncbi.nlm.nih.gov
- Katz, R. V., Kegeles, S. S., Kressin, N. R., Green, B. L., James, S. A., Wang, M. Q., ... Claudio, C. (2008). Awareness of the Tuskegee Syphilis Study and the US presidential apology and their influence on minority participation in biomedical research. American Journal of Public Health, 98(6), 1137-1142. pmc.ncbi.nlm.nih.gov
- National Institute of Environmental Health Sciences. (2025). Research ethics timeline. NIEHS, National Institutes of Health. niehs.nih.gov
- Office of the Federal Register. (2025). Protection of human subjects (45 C.F.R. Part 46). Electronic Code of Federal Regulations. ecfr.gov
- Key terms
- Nuremberg Code
- The 1947 code establishing that voluntary consent of the human subject is absolutely essential.
- Declaration of Helsinki
- The World Medical Association's ethical statement for medical research, first adopted in 1964.
- Tuskegee syphilis study
- The 1932 to 1972 U.S. study that withheld syphilis treatment from Black men, prompting modern U.S. reform.
- Belmont Report
- The 1979 report establishing respect for persons, beneficence, and justice as core U.S. research principles.
- Beneficence
- The principle of maximizing benefits and minimizing harms, grounding risk-benefit analysis.
- Justice (research ethics)
- The principle that benefits and burdens of research be distributed fairly across groups.
Informed Consent, IRB Review, and Vulnerable Populations
- Enumerate the required elements of valid informed consent.
- Distinguish exempt, expedited, and full-board IRB review and identify vulnerable-population safeguards.
This lesson turns the Belmont principles into the operational machinery every researcher must navigate: the elements of valid consent, the tiers of Institutional Review Board (IRB) review, and the extra protections owed to vulnerable groups. The U.S. framework is codified in federal regulation known as the Common Rule (45 CFR 46), which many countries mirror in substance.
The machinery can feel like paperwork, but each requirement is a principle in operational dress. Read a protocol form as a translation of Belmont rather than as a hurdle. Then you can anticipate what a board will ask and design a study that clears review without repeated revision. It also protects you, because a documented, well-reasoned ethics process is your best defense if a participant later raises a concern.
One protocol will illustrate the machinery. A doctoral student, Amara, plans to interview 25 nursing-home residents about loneliness. Her study looks harmless. It is not simple. Some residents have mild dementia, so capacity to consent varies from person to person and even from morning to afternoon. The activities director who introduces her to residents also decides who gets seats at the weekly outing, which puts pressure behind an invitation. And a resident who says the staff ignore her could face consequences if that remark is traced back. No needles, no drugs, three real ethical problems.
Key idea: Consent, review tier, and extra protections are the three machines that turn ethical principles into a study a board can approve.
In plain terms
Consent needs three things at once. Tell people what will happen. Make sure they can understand what you told them. And make sure they can say no without it costing them anything. A signature proves none of that. It only records that a conversation happened. If someone cannot explain your study back to you, you do not have consent yet. Next, review. Low-risk studies get a light touch. Higher-risk studies go to the full committee. You do not get to pick which category you are in. The board does, because you have an interest in the answer. Finally, some people need extra care. Children agree, and a parent gives permission. Prisoners cannot refuse as freely as other people. Anyone who depends on you is in a weak position to say no. The question is always the same. Could this person comfortably decline?
Valid informed consent
Consent is a process, not merely a signed form. To be valid it requires three components working together:
- Disclosure of the information a reasonable person would need. That covers the study's purpose, its procedures, foreseeable risks and discomforts, anticipated benefits, alternatives, confidentiality protections, and whom to contact with questions.
- Comprehension. Information must be presented in language and at a reading level the participant can understand. Disclosure that cannot be understood is not disclosure.
- Voluntariness. Agreement must be free of coercion (a threat) and of undue influence (an excessive inducement). Critically, consent must be revocable: participants may withdraw at any time without penalty.
The difference between the consent document and the consent process matters in practice. A signature captures a moment, but voluntariness must persist. That is why a participant may withdraw at any time, and why long studies re-confirm willingness as they go. The document should be written at roughly an eighth-grade reading level. It should avoid jargon and separate essential information from fine print. A form only a specialist can parse fails the comprehension requirement no matter how complete it is.
Amara's consent process is therefore not one event. She explains the study, waits, and returns the next day to ask again. At each interview she checks that the resident still wants to talk. When one resident says "you can ask me anything, I have nothing else going on," Amara treats that as a signal to slow down rather than as clean permission.
Two subtler threats to valid consent recur. The therapeutic misconception arises when participants in a clinical study believe that procedures designed to answer a research question are individualized treatment for them. It blurs the line between care and research. The second problem is capacity. Consent requires the competence to understand and reason about the choice, and competence can fluctuate with illness, medication, or age. A useful safeguard is the teach-back. Ask the participant to restate the study in their own words. That tests comprehension instead of assuming it.
The 2018 rule also reshaped the document itself. Consent must now begin with a concise key-information summary of the study's purpose, risks, benefits, and alternatives. It sits up front so a prospective participant can grasp the essentials before wading into detail. The reform responds to years of consent forms that had grown into unreadable legal documents, defeating the very comprehension they were meant to secure. Brevity at the top is now a regulatory expectation, not merely good style.
Key idea: Valid consent needs disclosure, comprehension, and voluntariness together, and a signature proves none of the three on its own.
When consent may be waived or altered
Consent is the default, not an absolute. The Common Rule permits an IRB to waive or alter consent only when four conditions are met. The research must be no more than minimal risk. The waiver must not adversely affect participants' rights and welfare. The research could not practicably be carried out without the waiver. And where appropriate, participants are given the pertinent information afterward. Secondary analysis of existing de-identified records is the common case. Contacting everyone would be impossible, and the informational risk is low.
For biospecimens and large data resources, the 2018 revisions introduced broad consent. This is a one-time agreement to unspecified future research within stated limits. It is optional and administratively demanding. It recognizes that no biobank can name every future use in advance. The through-line is simple. Any departure from full, specific consent must be justified to the board against explicit criteria. It is never something a researcher may assume for convenience.
Key idea: A waiver of consent must clear four written conditions, and only the board can grant it.
Tiers of IRB review
Not all research carries equal risk, and review is calibrated accordingly. Under the Common Rule the categories are, in ascending order of scrutiny:
- Exempt. Certain low-risk categories may qualify as exempt from ongoing review. Examples include some anonymous survey research and research in established educational settings. Here is the crucial point. The researcher does not decide exemption. The IRB or its designee makes that determination.
- Expedited. Research posing no more than minimal risk and fitting defined categories may be reviewed by the chair or an experienced member rather than the full board. Minimal risk means the probability and magnitude of harm are no greater than those met in ordinary daily life or in routine examinations.
- Full board. Research exceeding minimal risk requires review at a convened meeting of the full committee, with a majority present.
An IRB is a committee, not an individual, and its composition is regulated. Under the Common Rule a board must have at least five members of varying backgrounds. It must include at least one scientist and one non-scientist. It must also include at least one member unaffiliated with the institution, often a community representative. This diversity is deliberate. A purely scientific panel might discount the informational or dignitary harms that a community member or ethicist would flag. Members with a conflict of interest must recuse themselves from that protocol.
The 2018 revisions streamlined parts of the system. Annual continuing review is no longer required for most minimal-risk studies reviewed by the expedited route. Single-IRB review is encouraged for multi-site studies, to avoid duplicative approvals. And the broad-consent pathway was added. The core logic did not change. Scrutiny scales with risk, and the decision about which tier applies belongs to the board rather than to the investigator, who has an interest in the answer.
Key idea: Review runs from exempt to expedited to full board as risk rises, and the investigator never chooses the tier.
Vulnerable populations
Some groups have compromised autonomy or heightened exposure to coercion, so they receive additional safeguards. Children give assent while a parent or guardian gives permission. Prisoners are protected because confinement raises coercion concerns. Pregnant women and fetuses, and persons with cognitive impairments, are covered too. For participants who cannot consent for themselves, a legally authorized representative may consent on their behalf, subject to stricter risk limits.
The regulations codify these protections in named subparts. Subpart B governs research involving pregnant women, fetuses, and neonates. Subpart C governs prisoners, whose environment makes voluntary consent especially fragile. Subpart D governs children. For children, permissible research is tiered by risk. The tiers run from minimal risk, through a minor increase over minimal risk that holds out the prospect of direct benefit, up to research requiring special federal review. Assent from the child is sought alongside permission from a parent or guardian.
Vulnerability is not only a property of formal categories. Economic or educational disadvantage can compromise voluntariness. So can undocumented status, or dependence on the researcher, as when a professor studies their own students. Amara's residents are a case in point. No subpart is titled "nursing-home residents," yet dependence on staff and fluctuating capacity make them vulnerable in exactly the sense the regulations care about. The ethical question is always the same. Can this particular person, in this particular situation, give a free and informed yes? Design out the pressures that would make a no costly.
Key idea: Vulnerability is about the situation, not only the named category - dependence on the researcher can undermine consent in any group.
Confidentiality, deception, and debriefing
Beyond consent, researchers must protect data. Confidentiality limits who can link data to identities. Anonymity, where feasible, means collecting no identifiers at all. Some designs require deception, which means withholding or misstating the true purpose to prevent reactive behavior. Deception is permissible only under three conditions. The study could not be done otherwise. The risks are minimal. And participants are debriefed afterward: told the true purpose, given the chance to withdraw their data, and helped to a proper understanding. The bar is high because deception pulls against respect for persons. It must be justified to the IRB, never assumed.
Confidentiality is protected by both procedure and law. In practice, researchers separate identifiers from data, store the linking key securely, and report results in aggregate. In law, a federal Certificate of Confidentiality can shield sensitive data from compelled disclosure. That matters when studying illegal behavior or stigmatized conditions. Confidentiality has limits, though. Mandatory-reporting laws for child abuse or imminent harm override research promises. Consent forms must disclose those limits honestly instead of overpromising secrecy. Amara's protocol therefore quotes no resident by name, room, or job title, and it tells residents up front that a report of abuse would have to be passed on.
Debriefing does more than reveal a study's true purpose. A thorough debriefing has two further parts. Dehoaxing corrects any false beliefs the manipulation created. Desensitizing helps participants who were led to act in ways they might find troubling, so that they leave in no worse a state than they arrived. Debriefing also offers the chance to withdraw data collected under deception. Deception trades against respect for persons, so before permitting it the board asks whether any honest design could produce the same knowledge.
Key idea: Deception is a last resort that must be repaid with a full debriefing, and confidentiality promises must be honest about their legal limits.
Consent and data in the digital era
Modern studies raise consent and confidentiality questions the original regulations did not anticipate. Online recruitment, mobile sensors, and social-media data blur two things at once. Are participants truly informed? And are "public" data fair to use without consent? The European General Data Protection Regulation adds obligations for any research touching EU residents' personal data. Those include a lawful basis for processing and rights of access and erasure. A protocol that ignores data-protection law because it satisfies the Common Rule may still be non-compliant.
Re-identification is the quiet threat. Data stripped of names can often be re-linked to individuals by combining quasi-identifiers such as ZIP code, birthdate, and sex. Genuine de-identification therefore requires more than deleting a name column. Sound practice pairs technical measures, such as aggregation and access controls, with a data-management plan. The plan states who may see the data, where they are stored, and when they will be destroyed. Boards increasingly expect that plan as part of the protocol.
Key idea: Removing names does not de-identify data, and satisfying the Common Rule does not satisfy data-protection law.
Where people get stuck
Four distinctions are worth memorizing precisely.
- Assent versus permission versus consent. A child gives assent, a parent gives permission, and an adult gives consent. Using "consent" for a nine-year-old signals to a board that you have not read Subpart D.
- Anonymity versus confidentiality. Anonymous means you never knew who they were. Confidential means you know and will not tell. Amara's interviews cannot be anonymous, so everything rests on confidentiality.
- Exempt versus minimal risk. These are different categories, and neither is self-certified. Exempt refers to specific listed activities; minimal risk is a risk threshold used for expedited review.
- Waiver of consent versus waiver of documentation. A board may waive the signature requirement while still requiring a full consent conversation, which is common when a signature itself would create risk.
Key idea: Precision in these four pairs is the difference between a protocol a board approves and one it returns.
Try it
Exercise. A doctoral student plans an online experiment in which participants are told they are testing a "memory app," when the real aim is to measure how they respond to fabricated social-media comments criticizing their answers. The student proposes no debriefing, arguing that participants are anonymous and will never know. Identify the ethical problems and specify what the protocol must include to be approvable.
Model answer. Several requirements are unmet. The design uses deception, which is permissible only if the study cannot be done otherwise, the risks are minimal, and participants are debriefed; the student has simply omitted the debriefing. Exposing people to fabricated criticism poses a real psychological risk, so the minimal-risk claim needs support, and anonymity does not remove the obligation to debrief, since the purpose of debriefing is to correct false beliefs and offer withdrawal, not to protect the researcher.
To be approvable, the protocol needs several changes. It should justify the deception as necessary to prevent reactive behavior. It should minimize risk by limiting the intensity of the fabricated comments. It may disclose in advance that some information is being withheld, which is permitted. It must include a full debriefing that reveals the true purpose, dehoaxes participants by explaining that the comments were fabricated, offers them the chance to delete their data, and provides resources if any distress was caused. Finally, the tier of review is not the student's call. The risk may exceed minimal, so the board may require full-board review rather than exemption.
Recap
- Consent is a process requiring disclosure, comprehension, and voluntariness; the signed form documents it but does not constitute it.
- The 2018 Common Rule requires a concise key-information summary at the top of the consent document.
- Watch for therapeutic misconception and for fluctuating capacity, and test comprehension with a teach-back rather than assuming it.
- An IRB may waive or alter consent only when four stated conditions are met, and broad consent covers unspecified future use of biospecimens within limits.
- Review is tiered as exempt, expedited, and full board according to risk, and the board rather than the investigator decides the tier.
- Subparts B, C, and D protect pregnant women and fetuses, prisoners, and children, where children assent and guardians give permission.
- Deception requires necessity, minimal risk, and full debriefing including dehoaxing and desensitizing, and confidentiality promises must disclose their legal limits.
- Removing names is not de-identification, and a data-management plan plus attention to laws such as the GDPR is now expected.
Sources
- Office of the Federal Register. (2025). General requirements for informed consent (45 C.F.R. 46.116). Electronic Code of Federal Regulations. ecfr.gov
- Office of the Federal Register. (2025). Criteria for IRB approval of research (45 C.F.R. 46.111). Electronic Code of Federal Regulations. ecfr.gov
- Office of the Federal Register. (2025). Additional protections for children involved as subjects in research (45 C.F.R. Part 46, Subpart D). Electronic Code of Federal Regulations. ecfr.gov
- Grady, C. (2015). Institutional review boards: Purpose and challenges. Chest, 148(5), 1148-1155. pmc.ncbi.nlm.nih.gov
- Nijhawan, L. P., Janodia, M. D., Muddukrishna, B. S., Bhat, K. M., Bairy, K. L., Udupa, N., & Musmade, P. B. (2013). Informed consent: Issues and challenges. Journal of Advanced Pharmaceutical Technology & Research, 4(3), 134-140. pmc.ncbi.nlm.nih.gov
- National Institutes of Health. (2025). Guiding principles for ethical research. National Institutes of Health. nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Putting ethics into practice. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Key terms
- Common Rule
- The U.S. federal policy (45 CFR 46) governing human-subjects research across federal agencies.
- Informed consent
- A process combining disclosure, comprehension, and voluntariness that must remain revocable.
- Minimal risk
- Risk no greater in probability or magnitude than that of daily life or routine examinations.
- Expedited review
- IRB review by the chair or a member for minimal-risk studies in defined categories.
- Assent
- A minor's affirmative agreement to participate, alongside a guardian's permission.
- Debriefing
- Post-study disclosure of true purpose after deception, with a chance to withdraw one's data.
Module 4: Measurement, Validity, and Reliability
Turning abstract constructs into measurable variables and evaluating whether the measures are accurate and consistent.
Constructs, Variables, and Operationalization
- Distinguish constructs from variables and classify variables by role and level of measurement.
- Operationalize a construct and articulate the resulting definitional trade-offs.
Research tests relationships among abstractions - motivation, poverty, trust - but data can only be collected on concrete observables. Operationalization is the bridge: the specification of exactly how an abstract construct will be measured in a given study. The quality of that bridge determines whether your findings mean anything.
The stakes are easy to underestimate. Two studies can use the same construct name and reach opposite conclusions, purely because they operationalized it differently. A reader who ignores measurement is reading a mirage. This lesson builds the vocabulary that lets you specify, classify, and defend every variable before any statistic is chosen. That is the order in which competent design actually proceeds.
Start with a case that shows how much rides on this. Two teams both study whether poverty harms children's reading. Team 1 defines poverty as household income below the federal poverty line, a yes-or-no flag. Team 2 defines it as an index built from income, food insecurity, housing instability, and parental education. Team 1 finds no effect. Team 2 finds a large one. Nobody miscalculated anything. Team 1's flag treats a family at 99% of the line and a family at 20% as identical, and it misses hungry children in households just above the cutoff. The construct was the same word. The measure was not the same thing.
Key idea: The words in your theory cannot be measured; only the procedures you substitute for them can, and that substitution is a claim you have to defend.
In plain terms
The core idea is short. Theories talk about big words. Poverty. Trust. Motivation. You cannot measure a big word. You can only measure something you can count or record. So you swap the big word for a recipe. "Poverty" becomes "income below this line." That swap is called operationalization. It is a choice, and you have to defend it. Different recipes give different answers, as our two teams found. Next you sort your variables. Which one do you think is the cause? Which is the effect? What else might explain the link? Then you check what kind of numbers you have. Names with no order. Ranks with uneven gaps. Real units with no true zero. Real units with a true zero. That last question decides which statistics are allowed. Do all of this before you pick a test, never after.
Constructs and their operational definitions
A construct is an abstract concept that a theory invokes and that nobody can observe directly. "Socioeconomic status" is one. An operational definition states the procedure that will stand in for it, perhaps a composite of income, education, and occupational prestige. Any single operationalization is a partial, contestable rendering of the construct. Different reasonable operationalizations can yield different results, as our two poverty teams found. So you must justify yours and say what it leaves out. The gap between construct and measure is where construct validity, treated in the next lesson, will live.
Turning a construct into a measure is called construct explication, and it runs in two directions. Downward, you specify indicators concrete enough to observe. Upward, you check that those indicators, taken together, actually capture the concept rather than something adjacent. A construct like "resilience" might be indicated by items on recovery from setbacks, but if the items mostly tap current mood, the measure has quietly drifted from the construct it claims to name.
Indicators relate to constructs in two contrasting ways. In a reflective model the construct causes its indicators. The items are interchangeable symptoms of the same latent trait and should correlate highly. Depression, for instance, causes low mood, poor sleep, and fatigue. In a formative model the indicators compose the construct instead. Income, education, and occupation together constitute socioeconomic status without being caused by it. The distinction changes how you build and evaluate a scale. Mislabeling one as the other is a frequent measurement error. Notice that Team 2's poverty index is formative, so there is no reason its four components should correlate highly, and a low alpha would not be a defect.
Two naming fallacies deserve caution. The jingle fallacy assumes that two measures sharing a label measure the same thing, when "engagement" on one scale may differ sharply from "engagement" on another. The jangle fallacy assumes that two differently named measures capture different things, when "grit" and "conscientiousness" may largely overlap. Both are reasons to inspect the actual items rather than trusting the construct name printed at the top of an instrument.
A construct earns its meaning from its place in a nomological network. That is the web of theoretical relationships it is expected to have with other constructs. Cronbach and Meehl introduced the idea to explain how we validate a measure of something we cannot see. We check whether it relates to other measures as theory predicts. A good measure of test anxiety should correlate with physiological arousal before exams. It should not correlate strongly with unrelated traits. Assembling that pattern of expected relationships is what a validation study does.
Key idea: A construct is a word in a theory; an operational definition is a recipe, and two honest recipes for the same word can give different answers.
Variables by role
- An independent variable (IV) is the presumed cause or predictor; in an experiment it is the manipulated factor.
- A dependent variable (DV) is the presumed effect or outcome, the thing measured for change.
- A confounding variable is a third variable associated with both IV and DV that can produce a spurious association. Uncontrolled confounds are the chief enemy of causal inference.
- A mediator lies on the causal path between IV and DV and explains how the effect occurs. A moderator changes the strength or direction of the IV-DV relationship, specifying when or for whom it holds.
A worked example separates the roles that students most often confuse. Suppose a study finds that a mentoring program (IV) raises graduation rates (DV). If the program works by increasing students' sense of belonging, which in turn raises graduation, then belonging is a mediator, the mechanism on the causal path. If the program helps first-generation students more than others, then first-generation status is a moderator, altering the effect's size across groups. Mediators answer how; moderators answer for whom.
Two further terms round out the vocabulary. A control variable is one you hold constant or statistically adjust for, so as to remove its influence. A covariate is a measured variable included in a model to increase precision or to adjust for baseline differences. A confound is dangerous precisely when it is neither controlled nor adjusted. It then offers a rival explanation for any association you observe, a threat the design chapters treat in depth.
Key idea: Mediators explain how an effect happens, moderators say for whom it happens, and confounds offer a rival story for why you saw anything at all.
Levels of measurement
Stevens' taxonomy classifies variables by the information their numbers carry, and this classification constrains which statistics are legitimate.
| Level | Property | Example | Permissible center |
|---|---|---|---|
| Nominal | Categories, no order | Blood type | Mode |
| Ordinal | Ordered, unequal intervals | Likert agreement | Median |
| Interval | Equal intervals, no true zero | Celsius temperature | Mean |
| Ratio | Equal intervals, true zero | Reaction time | Mean, geometric mean |
The distinction between interval and ratio hinges on a true (absolute) zero, a zero that means the quantity is absent. Celsius has no true zero, because 0 degrees is not the absence of temperature. Ratios are therefore meaningless: 20 degrees C is not twice as hot as 10 degrees C. Reaction time has a true zero, so ratio statements about it are valid. Treating an ordinal variable as interval is a common shortcut with Likert data. It is a modeling assumption to be defended, not a fact.
Stevens' taxonomy is a starting point rather than the final word, and you should know the live debate. A strict reading forbids computing a mean of ordinal data. That would rule out the routine practice of averaging Likert items. A pragmatic school answers that an ordinal scale with several ordered categories often behaves close enough to equal intervals, so treating it as interval yields usable and often robust results. The defensible position is to state the assumption openly. Where it matters, check whether your conclusions survive an ordinal-appropriate analysis.
Some variables sit at special points on the scale. A dichotomous variable such as pass or fail is nominal with two categories, yet it behaves conveniently in many models. Count data, such as number of hospital visits, are ratio but often skewed and bounded at zero. That is why counts call for count-specific models rather than ordinary regression. Recognizing these cases stops you from applying techniques built for continuous, symmetric data to numbers that plainly violate their assumptions.
Key idea: The level of measurement decides which statistics are even meaningful, so classify every variable before you choose an analysis.
Building or borrowing a measure
Faced with a construct, your first question should be whether a validated instrument already exists. Borrowing one imports its accumulated evidence of reliability and validity. It also makes your results comparable to prior work. Reinventing a scale is rarely worth the cost, unless no existing measure fits your construct or population. When you do adapt an instrument, be careful. Even small changes to wording, response options, or context can alter its properties, so the adaptation needs justification and, ideally, a pilot.
Building a new measure is a project in itself. Start by defining the construct's dimensions. Write a pool of items that samples each one. Choose a response format, such as a Likert scale or a frequency count. Items should be clear, single-barreled, and free of leading language. Refine the pool through expert review and pilot testing before any confirmatory use. The survey chapter revisits item wording in detail. The point here is that a measure is designed, not improvised.
Key idea: Borrow a validated instrument when one fits, because it brings evidence and comparability that a homemade scale cannot.
Latent variables and factor structure
Reflective constructs are latent variables. They are inferred from patterns among observed items rather than measured directly. Factor analysis is the statistical tool that examines those patterns. Exploratory factor analysis discovers how many underlying dimensions the items share. Confirmatory factor analysis tests whether the items load on the dimensions a theory specifies. Suppose a scale meant to measure one construct splits into three factors. That result is telling you the construct, as operationalized, is not unidimensional, and it changes how scores should be computed and interpreted.
Key idea: Factor analysis checks whether your items really hang together as the one thing you claimed to be measuring.
A measurement workflow
Assembling these ideas yields a repeatable sequence. First, define the construct and its boundaries, saying what it includes and excludes. Second, choose or build an operational definition, preferring an existing validated instrument when one fits. Reinventing a scale forfeits accumulated evidence. Third, classify each variable by role and by level of measurement. Only then, fourth, select analyses whose assumptions match those classifications. Reversing the order means choosing a favorite analysis and forcing the variables to fit it. That reversal is the source of much published nonsense.
Documenting these decisions pays off at the writing stage. Name each construct, its operational definition, the instrument and its source, the response format, and the level of measurement. That gives reviewers exactly what they need in order to judge whether your later statistics are appropriate. Vague phrasing such as "we measured motivation" invites the very question the whole workflow was meant to answer. A precise specification closes it in advance.
Key idea: Construct, then operational definition, then variable roles and levels, then analysis - and never the reverse.
Why this discipline matters
Measurement level is not pedantry. Computing a mean of nominal categories is nonsense. Applying a technique that assumes interval data to genuinely ordinal data can distort conclusions. Fix the construct. Choose an operational definition you can defend. Classify each variable's role and level. Only then select an analysis. Doing these in the wrong order is how researchers end up with sophisticated statistics applied to meaningless numbers.
Where people get stuck
Four pairs cause most of the trouble.
- Construct versus variable. "Poverty" is a construct; "household income below the federal poverty line" is a variable. Reviewers object when a paper slips between them, claiming results about poverty from a study that measured only one income threshold.
- Mediator versus moderator. If mentoring works by raising belonging, belonging is a mediator on the causal path. If mentoring works better for first-generation students, that status is a moderator. Mediators are about mechanism, moderators about conditions.
- Confound versus covariate. Both are third variables. A covariate is one you measured and adjusted for; a confound is one you did not. The same variable can be either, and which it is depends entirely on your design.
- Reflective versus formative. Reflective items should correlate, because one latent cause drives them. Formative components need not correlate, because they add up to the construct. Reporting Cronbach's alpha for a formative index is a category error.
One further point about the jingle and jangle fallacies. Never trust the name printed at the top of an instrument. Read the items. Two scales called "burnout" may measure exhaustion and cynicism respectively, and a study that swaps one for the other has changed its research question without saying so.
Key idea: Read the items, not the label - the name of a scale is a marketing claim until you check what it actually asks.
Try it
Exercise. A researcher studies whether "school climate" affects "student achievement." They measure climate with a single yes-or-no item ("Is this a good school?") answered by principals, measure achievement as a pass or fail flag, and plan to compute the mean climate score and correlate it with mean achievement. Diagnose the operationalization and measurement problems, then propose a defensible alternative.
Model answer. The operationalization of school climate is inadequate: a single yes-or-no item answered by one interested informant is a thin, possibly invalid indicator of a rich multidimensional construct, and it risks the jingle fallacy of assuming the label guarantees the content. The measurement levels are mishandled too. A yes-or-no item is dichotomous nominal, so computing its "mean" as if it were interval is not meaningful, and achievement reduced to pass or fail discards information that a continuous score would retain.
A defensible alternative has three parts. Use a validated multi-item school-climate scale completed by multiple respondents, treated as at least an approximate interval measure with the assumption stated. Keep the achievement outcome on its original continuous scale, such as a standardized test score. Then model climate as a predictor of achievement, controlling for confounds such as prior achievement and school poverty. The redesign fixes both halves of the bridge: a construct-valid operationalization, and measurement levels matched to the intended statistics.
Recap
- Constructs are theoretical and unobservable; operational definitions are the procedures that stand in for them, and every such substitution is contestable.
- Two defensible operationalizations of one construct can produce opposite findings, so measurement decisions belong in the methods section in full.
- Reflective indicators are caused by the construct and should correlate; formative indicators compose the construct and need not.
- The jingle and jangle fallacies warn that scale names are unreliable guides to scale content, so inspect the items.
- Variables are classified by role as independent, dependent, confounding, mediating, moderating, control, or covariate.
- Stevens' levels - nominal, ordinal, interval, ratio - constrain which statistics are meaningful, and treating ordinal data as interval is an assumption to defend.
- The workflow is construct, operational definition, variable roles and levels, then analysis; reversing it produces sophisticated statistics on meaningless numbers.
Sources
- Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. psychclassics.yorku.ca
- Tal, E. (2020). Measurement in science. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Understanding psychological measurement. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Basic concepts. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Strauss, M. E., & Smith, G. T. (2009). Construct validity: Advances in theory and methodology. Annual Review of Clinical Psychology, 5, 1-25. pmc.ncbi.nlm.nih.gov
- Sullivan, G. M., & Artino, A. R. (2013). Analyzing and interpreting data from Likert-type scales. Journal of Graduate Medical Education, 5(4), 541-542. pmc.ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Measurement of constructs. In Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Key terms
- Construct
- An abstract, unobservable concept invoked by a theory, such as motivation or status.
- Operationalization
- Specifying the concrete procedure by which a construct will be measured in a study.
- Confounding variable
- A third variable related to both IV and DV that can create a spurious association.
- Mediator
- A variable on the causal path between IV and DV that explains how the effect occurs.
- Moderator
- A variable that changes the strength or direction of the relationship between IV and DV.
- Level of measurement
- Stevens' classification (nominal, ordinal, interval, ratio) governing permissible statistics.
Reliability: The Consistency of Measurement
- Define reliability and distinguish its major forms.
- Interpret reliability coefficients and relate reliability to measurement error.
Reliability is the consistency or reproducibility of a measurement. A reliable bathroom scale gives the same reading when you step on it twice in a row. Reliability is a precondition for validity but not a guarantee of it: a scale miscalibrated by five kilograms is perfectly reliable and perfectly wrong. This lesson defines the forms of reliability and shows how to read the coefficients that quantify them.
Keep the scale analogy in mind throughout, because it isolates what reliability is and is not. Consistency concerns whether repeated measurement lands in the same place. It says nothing about whether that place is correct. A measure can be reliable and biased. And it cannot be valid without first being reliable. Reliability is the floor on which validity is built, never a substitute for it.
Picture four bathroom scales to fix the idea. Scale A reads 71, 71, 71 kg for the same person: reliable and, let us say, accurate. Scale B reads 76, 76, 76: reliable and wrong by five kilograms every time. Scale C reads 64, 78, 71: unreliable, though its average happens to be right. Scale D reads 61, 82, 55: unreliable and wrong. Only A is usable. Notice that B looks better than C on every reliability index, and that is the point. Consistency is necessary and nowhere near sufficient.
Key idea: Reliability asks whether a measure repeats itself, not whether it is right.
In plain terms
Reliability asks one question. If I measure this again, do I get the same answer? That is it. It does not ask whether the answer is correct. A scale that is five kilograms out gives the same wrong number every time, and every reliability test will call it excellent. Four kinds of repeat are worth checking. Measure the same people twice and correlate the scores. Check whether the items on a scale agree with each other. Check whether two raters watching the same thing write down the same thing. Or check whether two versions of a test give the same result. Which one you need depends on where your noise comes from. Two practical points. Every score deserves a band around it, not a single number. And a noisy measure hides real effects, so better measurement buys you the same thing a bigger sample would.
Classical test theory
The organizing idea is simple: any observed score is the sum of a true score and error, written X = T + E. Reliability is the proportion of observed-score variance that reflects true-score variance rather than random error. Perfectly reliable measurement (all true score, no error) has a reliability of 1; pure noise has a reliability of 0. The aim is to shrink the error term relative to the signal.
The model rests on assumptions worth stating. Error is taken to be random. It has a mean of zero over many measurements and no correlation with the true score. That is why averaging repeated measurements converges on the truth. A practical quantity follows: the standard error of measurement (SEM). The SEM expresses, in the score's own units, how much an individual's observed score would wobble on re-measurement. It equals the standard deviation times the square root of one minus the reliability.
A worked SEM example makes it concrete. Suppose a test has a standard deviation of 10 points and a reliability of 0.84. Then SEM = 10 x sqrt(1 - 0.84) = 10 x sqrt(0.16) = 10 x 0.4 = 4 points. A person who scored 70 has an approximate 95% band of 70 +/- 2 x 4, or roughly 62 to 78. The SEM turns a single number into an interval. It is also why high-stakes decisions should never rest on small score differences that fall inside the measurement error. Two applicants scoring 70 and 73 on this test are, for practical purposes, tied.
Key idea: Every observed score is a true score plus error, and the standard error of measurement tells you how wide a band to put around any one person's number.
Sources of measurement error
Error is not monolithic, and naming its sources suggests how to reduce it. Transient error reflects momentary states of the person, such as fatigue or mood on the day of testing, and it is what test-retest reliability captures. Item-specific error reflects quirks in how particular items are interpreted, and it is what internal consistency captures. Rater error reflects differences between observers, captured by inter-rater indices. Each form of reliability targets a different slice of error, which is why a single coefficient never tells the whole story.
A crucial distinction cuts across these. Classical test theory concerns random error, which varies unpredictably and averages out over repeated measurement. It says nothing about systematic error, a constant bias that pushes every score the same way. Scale B above is the model case. Systematic error leaves reliability untouched while destroying validity. That is the precise sense in which a measure can be perfectly consistent and perfectly wrong at once.
Key idea: Each form of reliability catches a different kind of random error, and none of them catches systematic bias.
Forms of reliability
- Test-retest reliability assesses stability over time: administer the same measure to the same people on two occasions and correlate the scores. It presumes the construct itself is stable across the interval - inappropriate for genuinely changeable states such as mood.
- Internal consistency assesses whether items intended to tap one construct hang together. The most reported index is Cronbach's alpha. It rises with the average inter-item correlation and with the number of items. Values around 0.70 to 0.80 are commonly treated as acceptable for research scales. A very high alpha, above roughly 0.95, usually signals redundant near-duplicate items rather than a virtue.
- Inter-rater reliability assesses agreement between independent observers coding the same events. Because raw percent agreement is inflated by chance, use a chance-corrected index such as Cohen's kappa for two raters, which subtracts the agreement expected by chance.
- Parallel-forms reliability correlates two equivalent versions of an instrument built from the same content domain. It is useful when repeated testing risks memory effects.
Match the form to the threat you actually face. A depression scale given twice needs test-retest evidence, but only if you believe depression is stable over the interval. A new eight-item scale needs internal consistency. A study where two coders classify 400 classroom videos needs inter-rater agreement, and internal consistency is irrelevant to it. Reporting the wrong coefficient is a common way to look rigorous while proving nothing.
Key idea: Test-retest covers stability, internal consistency covers items, inter-rater covers observers, and parallel forms cover versions - pick the one that matches your design.
A worked inter-rater example
Because raw percent agreement is inflated by chance, inter-rater reliability is usually reported with a chance-corrected index. Consider two raters independently coding 100 interview transcripts for whether distress is present. They agree that distress is present in 40 transcripts and absent in 35, so they agree on 75 of 100, a raw agreement of 0.75. Across all transcripts, Rater A marked distress present in 50 and Rater B in 55.
Cohen's kappa corrects that figure for chance. The expected chance agreement is the sum, across categories, of the product of the two raters' marginal proportions: for "present," 0.50 x 0.55 = 0.275, and for "absent," 0.50 x 0.45 = 0.225, giving expected agreement of 0.50. Kappa is then observed minus expected, divided by one minus expected: kappa = (0.75 - 0.50) / (1 - 0.50) = 0.25 / 0.50 = 0.50. The raw 75% agreement shrinks to a moderate kappa of 0.50 once chance is removed.
The lesson generalizes. When categories are few, or one category is common, chance agreement is high and raw percentages badly overstate reliability, which is exactly the situation kappa was designed for. For more than two raters you would use Fleiss' kappa, for ordinal categories a weighted kappa that penalizes larger disagreements more heavily, and for continuous ratings an intraclass correlation. Reporting only raw agreement, without any chance correction, is a common and avoidable error in coding-based research.
Kappa has its own quirks that a careful analyst should flag. When one category is very rare or very common, kappa can be low even though raw agreement is high. That pattern is sometimes called the prevalence paradox. It does not mean the raters disagree wildly. It means there was little room to agree beyond chance. Report the underlying agreement table alongside kappa rather than the coefficient alone. Readers can then see the marginal distribution that produced it and judge the number for themselves.
Key idea: Raw percent agreement flatters you, because two raters guessing at random would already agree much of the time.
Understanding Cronbach's alpha
Internal consistency asks whether items meant to measure one construct behave as if they do. Cronbach's alpha summarizes this, and conceptually it can be understood as the average of all possible split-half reliabilities, corrected for the number of items. It rises with two things: the average correlation among items and the number of items. That second dependency is why a long scale of mediocre items can post a respectable alpha while a short scale of excellent items posts a lower one.
Two cautions follow. First, alpha is not a measure of unidimensionality. A scale tapping two distinct factors can still return a high alpha. Pair alpha with evidence about factor structure rather than treating it as proof that the items measure one thing. Second, an alpha above roughly 0.95 usually signals redundancy: several items asking essentially the same question. That wastes respondents' effort without adding information. Values around 0.70 to 0.80 are conventional for research use, and higher is demanded for consequential decisions about individuals.
One practical implication follows from alpha's dependence on length. The Spearman-Brown prophecy formula predicts how reliability changes as you lengthen or shorten a test. It shows that adding good items raises reliability, with diminishing returns. The same logic applies to raters. Averaging the judgments of several trained coders is more reliable than relying on one. That is why observational studies often pool multiple raters instead of trusting a lone observer's codes.
Key idea: Alpha goes up when items correlate and when there are more of them, so a high alpha may mean a good scale or merely a long, repetitive one.
Reading a coefficient
Reliability coefficients generally range from 0 to 1, and higher is better. The acceptable threshold depends on stakes. A screening instrument used for group-level research tolerates lower reliability than a test used to make decisions about individuals, where 0.90 or above is often demanded. Kappa has its own conventions. Values above about 0.60 are frequently described as substantial agreement, though such labels are heuristics rather than laws.
Interpretation should resist rigid cutoffs. Textbook benchmarks such as kappa above 0.60 or alpha above 0.70 are heuristics. Their appropriateness depends on the construct, the stakes, and the field. A reliability that is fine for a group-level research correlation may be unacceptable for classifying an individual patient. Report the coefficient alongside its practical implication rather than a bare verbal label.
Key idea: There is no universal threshold - the reliability you need depends on whether you are describing a group or deciding about a person.
Reliability and error
Reliability caps validity, so unreliable measurement attenuates observed relationships. Random error in the variables biases correlations toward zero, and a real effect can be masked by a noisy instrument. Improving reliability therefore increases your ability to detect true effects. You do it by adding good items, training raters, standardizing administration, or clarifying wording. The practical lesson is that measurement quality is not a formality to report and forget. It governs statistical power and the credibility of every relationship you estimate. It is often cheaper to buy precision through better measurement than through a larger sample.
The attenuation effect can be quantified, which sharpens the point. If two variables are each measured with error, the observed correlation is roughly the true correlation multiplied by the square root of the product of the two reliabilities. A true correlation of 0.50 between constructs measured with reliabilities of 0.64 each would appear as about 0.50 x sqrt(0.64 x 0.64) = 0.50 x 0.64 = 0.32. The real relationship is stronger than the noisy data reveal. That is why unreliable measurement hides effects that genuinely exist and inflates the apparent number of null results.
Key idea: Noisy measures shrink your observed correlations toward zero, so poor reliability costs you power just as surely as a small sample does.
Where people get stuck
The reliability-versus-validity confusion is the single most common error in this material, so make the two questions explicit.
- Reliability asks: does it repeat? Step on the scale three times and see whether the numbers match. This is answered by correlating repeated measurements, items, raters, or forms.
- Validity asks: is it measuring the right thing? Scale B, wrong by five kilograms every time, has perfect reliability and no validity. No reliability coefficient will ever detect that error, because reliability cannot see a constant bias.
Three further traps follow from that asymmetry. High reliability never demonstrates validity, though low reliability does limit it, since a measure too noisy to repeat cannot track anything consistently. A high alpha does not show that a scale measures one thing, so pair it with factor analysis. And raw percent agreement is not inter-rater reliability, because chance agreement has not been removed. When you write a methods section, report which coefficient, computed on which data, and what the number implies for your specific use.
Key idea: Reliability is necessary but never sufficient - a consistently wrong instrument passes every reliability test you can run on it.
Try it
Exercise. Two coders rate 80 social-media posts as "misinformation" or "not." They agree on 68 posts. Coder 1 labeled 20 posts as misinformation; Coder 2 labeled 24. A student reports "85% agreement, so reliability is excellent." Compute Cohen's kappa, and explain why the student's conclusion is premature.
Model answer. Observed agreement is 68 / 80 = 0.85. To find expected agreement, use the marginals. Coder 1 called 20/80 = 0.25 misinformation and 0.75 not; Coder 2 called 24/80 = 0.30 misinformation and 0.70 not. Expected agreement for "misinformation" is 0.25 x 0.30 = 0.075, and for "not" is 0.75 x 0.70 = 0.525, summing to 0.60.
Kappa is then (0.85 - 0.60) / (1 - 0.60) = 0.25 / 0.40 = 0.625. The chance-corrected agreement is therefore about 0.63. That falls in the substantial range but not the excellent one, well below what the raw 85% implied. The student's error was reading raw agreement as if it were reliability. Because misinformation is the minority category here, a large share of the raw agreement was expected by chance. Only kappa reveals how much genuine agreement remains once that chance component is subtracted.
Recap
- Reliability is consistency, and it is necessary but not sufficient for validity: a measure can repeat perfectly and still be biased.
- Classical test theory writes observed score as true score plus random error, and the standard error of measurement converts a reliability into a score band.
- Random error averages out over repeated measurement, while systematic error leaves reliability untouched and destroys validity.
- Test-retest, internal consistency, inter-rater, and parallel-forms reliability each target a different source of error, so choose the one your design needs.
- Cohen's kappa corrects raw agreement for chance, and raw percent agreement systematically overstates reliability when one category is common.
- Cronbach's alpha rises with item correlation and item count; it does not establish unidimensionality, and above 0.95 it signals redundancy.
- Unreliable measurement attenuates correlations toward zero, so better measurement buys statistical power that would otherwise cost a larger sample.
Sources
- Tavakol, M., & Dennick, R. (2011). Making sense of Cronbach's alpha. International Journal of Medical Education, 2, 53-55. pmc.ncbi.nlm.nih.gov
- Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155-163. pmc.ncbi.nlm.nih.gov
- McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276-282. pmc.ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Reliability and validity of measurement. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Sullivan, G. M., & Artino, A. R. (2013). Analyzing and interpreting data from Likert-type scales. Journal of Graduate Medical Education, 5(4), 541-542. pmc.ncbi.nlm.nih.gov
- Substance Abuse and Mental Health Services Administration. (2019). Reliability of key measures in the National Survey on Drug Use and Health. NCBI Bookshelf. ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Scale reliability and validity. In Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Key terms
- Reliability
- The consistency or reproducibility of a measurement across time, items, or raters.
- Classical test theory
- The model in which an observed score equals a true score plus error (X = T + E).
- Test-retest reliability
- Stability of scores when the same measure is re-administered to the same people.
- Cronbach's alpha
- An index of internal consistency reflecting average inter-item correlation and item count.
- Inter-rater reliability
- Agreement between independent observers coding the same events.
- Cohen's kappa
- A chance-corrected index of agreement between two raters.
Validity of Measurement
- Differentiate content, criterion, and construct validity.
- Diagnose construct-validity threats including convergent and discriminant evidence.
Reliability asks whether a measure is consistent. Validity asks whether it measures what it claims to measure. A measure can be reliable without being valid. It cannot be valid without being reliable. Validity is therefore the more demanding and ultimately more important criterion. One warning about the word. This lesson concerns the validity of measurement. The validity of a study's causal and generalizing inferences, called internal and external validity, is a separate topic taken up in Module 6.
A modern refinement frames the whole subject precisely. Validity is not a property of a test. It is a property of the inferences drawn from its scores for a particular purpose. The same reading test may support valid inferences about decoding skill and invalid inferences about general intelligence. So whenever you read that an instrument "is valid," translate that into "supports valid inferences of a specific kind." That translation is what a rigorous validation argument has to establish.
Here is a case to hold onto. A district adopts a 15-minute tablet test and calls it a measure of "reading ability." Third-graders take it in a noisy library, tapping answers under time pressure. The test correlates 0.91 with itself two weeks later, so it is highly reliable. It also correlates 0.72 with typing speed and only 0.31 with teachers' judgments of which children can read a story and explain it. Something is being measured very consistently. It is not reading. Every idea below is a tool for detecting that gap before a district spends a million dollars on the wrong instrument.
Key idea: Validity is a claim about what your scores let you conclude, and it has to be argued rather than asserted.
In plain terms
Four plain questions cover this whole lesson. First: does the test cover the whole subject, or only a corner of it? That is content validity. Second: does the score line up with something we already trust, or predict something that happens later? That is criterion validity. Third: does the score go up with things it should, and stay flat with things it should not? That is convergent and discriminant evidence, and together they make construct validity. Fourth: what happens to people because of the score? Messick said that counts too. One more thing. Do not say "this test is valid." Say "this test supports this conclusion, for these people, for this purpose." A test can be right for one job and useless for another. Our tablet test measures something. It just is not reading.
Content validity
Content validity asks whether the items adequately sample the full domain of the construct. Consider a statistics exam that tests only descriptive methods yet claims to assess "statistical competence." It under-represents the domain, so it lacks content validity. Content validity is judged largely by expert review against a clear specification of the domain. A closely related but more superficial notion is face validity: whether the measure merely looks relevant to lay observers.
Content validity is built, not discovered. The standard procedure begins with a table of specifications. That table lays out the construct's domains and the weight each should receive. You then write items to sample each domain in proportion. A panel of subject-matter experts rates whether each item is essential, and indices such as Lawshe's content validity ratio summarize the degree of expert consensus. The descriptive-only statistics exam fails for a structural reason. Its items may each be fine. Its blueprint omitted whole regions of the domain it claims to cover. The reading tablet test has the same defect: it samples word recognition and skips comprehension entirely.
Face validity is the weakest cousin, and it should not be confused with content validity. Face validity is only whether a measure looks relevant to a lay respondent. That affects cooperation and motivation. It says nothing about whether the items sample the domain. Sometimes low face validity is even deliberate, as in measures that disguise their purpose to reduce socially desirable responding. Face validity is a usability consideration, not evidence of measurement quality.
Key idea: Content validity is about coverage of the domain, and it is designed in with a blueprint rather than checked afterward.
Criterion validity
Criterion validity asks whether the measure relates as expected to an external criterion. It comes in two temporal flavors:
- Predictive validity: the measure predicts a future criterion (an admissions test predicting later grades).
- Concurrent validity: the measure agrees with a criterion assessed at the same time (a new depression screener correlating with an established diagnostic interview).
Criterion validity is only as good as the criterion, which creates the criterion problem. Suppose you validate a new hiring test against supervisor ratings, and those ratings are themselves biased or unreliable. A strong correlation then validates your test against a flawed standard. A weak correlation may indict the criterion rather than the test. Choosing a defensible criterion is part of the validation argument. It is not a given you may assume.
Two technical issues shape criterion coefficients. The first is restriction of range, which attenuates them. If you can only observe the later grades of admitted students, the narrowed spread of admission-test scores understates the test's true predictive validity in the full applicant pool. The second concerns screening instruments. There the relevant evidence is often sensitivity and specificity: the rates of correctly identifying true cases and true non-cases. Those matter more than a single correlation when a measure will be used to classify individuals.
Put numbers on why sensitivity and specificity matter. Screen 1,000 people for depression where 100 are truly depressed. A screener with 90% sensitivity finds 90 of those 100. With 90% specificity it also mislabels 90 of the 900 healthy people as positive. So 180 people screen positive and only half of them are cases. A coin flip on a positive result. A single validity correlation hides that completely. Reporting sensitivity, specificity, and the resulting predictive values shows exactly how the instrument will behave when it is used on real people.
A further useful concept is incremental validity. It asks whether a measure predicts the criterion beyond what cheaper or existing measures already predict. A new, expensive personality inventory may correlate handsomely with the criterion and still add nothing to what a short questionnaire already achieved. Its incremental validity is then near zero. Reviewers increasingly ask this question, because a measure earns its place by what it adds rather than by what it repeats.
Key idea: Criterion validity is only as trustworthy as the criterion, and for classification decisions sensitivity and specificity say far more than a correlation.
Construct validity
Construct validity is the overarching question. Does the measure truly capture the theoretical construct? Modern measurement theory treats content and criterion evidence as feeding into it. Two complementary lines of evidence are central.
- Convergent validity: the measure correlates strongly with other measures of the same or theoretically related constructs.
- Discriminant (divergent) validity: the measure correlates weakly with measures of different, theoretically unrelated constructs. A self-esteem scale should relate to related well-being measures (convergent) but not to, say, verbal IQ (discriminant).
The classic tool for examining both at once is the multitrait-multimethod matrix. It arranges correlations among several traits, each measured by several methods. That layout lets you compare convergent evidence, meaning the same trait measured by different methods, against discriminant evidence and against the confound of a shared method.
Convergent and discriminant evidence work as a pair, and a worked pattern shows why. Imagine validating a new empathy scale. Convergent evidence would be a strong correlation with an established empathy measure, plus a moderate one with prosocial behavior. Both are theoretically expected. Discriminant evidence would be a weak correlation with an unrelated construct such as spatial reasoning. Now suppose the new scale correlated as highly with spatial reasoning as with the established empathy measure. Its scores would be capturing something other than, or in addition to, empathy. That is exactly what the tablet reading test showed with its 0.72 correlation with typing speed.
A further strand is known-groups validity. If a measure of clinical anxiety is valid, patients diagnosed with an anxiety disorder should score higher than a matched community sample. Theory predicts the groups differ on the construct. Confirming that predicted difference adds evidence. Failing to find it, when the groups genuinely do differ, warns that the measure is not tracking the construct. Known-groups checks are especially useful early in validation, before a full nomological network of relationships has been assembled.
Key idea: Construct validity is built from a pattern - strong links to what should be related, weak links to what should not, and predicted differences between groups.
Validity as a unified argument
The historical categories of content, criterion, and construct validity are now widely regarded as facets of one unified concept. Samuel Messick developed that argument, and professional testing standards reflect it. On this view, construct validity is the whole of validity. Content and criterion evidence are simply among the sources that feed it. Validation is therefore a cumulative case. It marshals several kinds of evidence toward one conclusion: that a given interpretation and use of scores is warranted.
Messick also pressed a further, sometimes uncomfortable, point. The consequences of test use are part of validity. Suppose a supposedly neutral assessment systematically disadvantages a group for reasons unrelated to the construct. That consequence bears on whether the intended interpretation holds. Our tablet test is a live example. Because it rewards fast tapping, it will underrate children with limited screen access, and that pattern is not a fact about their reading. You need not resolve every debate about consequential validity. You should recognize that a defensible validation attends to what is done with a score, and to whom, and not only to what it correlates with.
The current professional Standards for Educational and Psychological Testing organize this argument into five sources of validity evidence. They are evidence based on test content, on response processes, on internal structure, on relations to other variables, and on the consequences of testing. The list is worth memorizing, because it doubles as a checklist. Read any validation section against it to see which sources have been supplied and which have merely been assumed. The gaps are usually where a reviewer will press.
Key idea: Validity is one cumulative argument built from five kinds of evidence, and what happens to people because of the scores is part of that argument.
Building a validation program
Validation is a program of studies rather than a single analysis, and it usually unfolds in a sequence. Early work establishes the domain and content, refines items, and checks internal structure through factor analysis. Middle work gathers convergent, discriminant, and known-groups evidence to show the scores behave as the construct predicts. Later work establishes criterion validity and, where relevant, predictive and incremental validity against consequential outcomes. Each stage can send you back to revise items. That is why validation is described as iterative rather than linear.
This is also why borrowing a well-validated instrument is so valuable. Adopt a measure with an established validation record in a comparable population and you inherit the accumulated argument. You then need only confirm that it holds in your context. Building a new measure obliges you to supply the whole case yourself. Many dissertations underestimate that undertaking until a committee asks for the evidence.
Key idea: Validation is a program, not a coefficient, which is why borrowing a validated instrument saves years of work.
Threats to construct validity
Two systematic threats deserve special vigilance. Construct under-representation occurs when the measure taps too narrow a slice of the construct. The descriptive-only statistics exam is the model case. Construct-irrelevant variance occurs when the measure is contaminated by something extraneous. A "math ability" test whose word problems are so linguistically dense that it partly measures reading ability is one example. The tablet reading test is another, since it partly measures manual dexterity. Both threats distort the meaning of scores. Establishing construct validity is therefore not a single coefficient. It is a cumulative, theory-guided argument assembled from several sources of evidence. That is exactly why a serious methods section devotes real space to defending its measures instead of assuming them.
Several further threats recur in practice. Method variance arises when scores reflect the measurement method rather than the construct. It happens when every variable is self-reported and correlations are inflated by a shared response style. Reactivity arises when being measured changes the behavior being measured. And a mismatch between the population a measure was validated on and the population it is applied to can silently void its validity. An instrument validated with adults may misbehave with children. Each threat is a reason that a validity claim does not transfer automatically to a new use.
Key idea: A measure can fail by covering too little of the construct or by quietly measuring something else as well.
Where people get stuck
Two pairs of terms cause most of the difficulty, and the first is the pair students confuse most often in the entire course.
- Reliability versus validity. Reliability asks whether the measure repeats itself; validity asks whether it measures the intended thing. The tablet test scored 0.91 on retest and still failed, because consistency cannot detect that you are measuring the wrong construct. Reliability is necessary, never sufficient.
- Measurement validity versus study validity. This lesson is about whether a score means what you say it means. Internal validity, covered in Module 6, is about whether your causal claim holds. External validity is about whether your finding travels. A study can have flawless internal validity built on an invalid measure, and then it has proved something precise about the wrong variable.
- Content versus face validity. Content validity is expert judgment against a domain blueprint. Face validity is whether the thing looks plausible to a respondent. Only the first is evidence.
- Convergent versus discriminant evidence. Convergent evidence is a correlation you want to be high. Discriminant evidence is a correlation you want to be low. A validation reporting only the first half has told you nothing about what else the scores might be picking up.
Key idea: Reliability, measurement validity, and study validity are three different questions, and answering one never answers another.
Try it
Exercise. A team develops a 10-item "leadership potential" questionnaire. To validate it, they administer it twice two weeks apart, obtain a test-retest correlation of 0.88, and conclude the measure is valid. Critique the validation, and outline what evidence a credible construct-validity argument would require.
Model answer. The team has demonstrated reliability, not validity. A test-retest correlation of 0.88 shows the questionnaire is stable over two weeks, but stability says nothing about whether the items capture leadership potential rather than, say, general confidence or verbal fluency. A measure can be highly reliable and still measure the wrong construct, which is the exact gap between the two criteria.
A credible construct-validity argument would assemble several strands. Content evidence would show, through a specification and expert review, that the ten items sample the domain of leadership rather than a narrow slice of it. Convergent evidence would show the scale correlates with established leadership measures and with relevant behavior such as supervisor ratings. Discriminant evidence would show weaker correlations with unrelated constructs. Criterion evidence, ideally predictive, would show the scale forecasts later leadership outcomes. An incremental-validity check would show it adds prediction beyond existing measures. Only this cumulative case licenses the claim that the instrument is valid for its intended use. A single reliability coefficient never does.
Recap
- Validity is a property of inferences from scores for a stated purpose, not a badge a test carries around with it.
- Content validity concerns coverage of the domain and is built with a table of specifications and expert review; face validity is appearance only.
- Criterion validity comes in predictive and concurrent forms, is limited by the quality of the criterion, and is attenuated by restriction of range.
- For classification uses, sensitivity, specificity, and predictive values reveal behavior that a single correlation conceals.
- Construct validity rests on convergent, discriminant, and known-groups evidence, with the multitrait-multimethod matrix as the classic tool.
- Messick unified the categories: construct validity is the whole of validity, and the consequences of test use form part of the argument.
- The two systematic threats are construct under-representation and construct-irrelevant variance, joined in practice by method variance and reactivity.
- Reliability, measurement validity, and internal or external validity are three separate questions, and high reliability never demonstrates validity.
Sources
- Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. psychclassics.yorku.ca
- Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. find source ↗
- Strauss, M. E., & Smith, G. T. (2009). Construct validity: Advances in theory and methodology. Annual Review of Clinical Psychology, 5, 1-25. pmc.ncbi.nlm.nih.gov
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association. find source ↗
- Tal, E. (2020). Measurement in science. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Reliability and validity of measurement. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Practical strategies for psychological measurement. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Key terms
- Validity (measurement)
- The degree to which a measure assesses the construct it claims to assess.
- Content validity
- Whether items adequately sample the full domain of the intended construct.
- Criterion validity
- Whether a measure relates as expected to an external criterion, predictively or concurrently.
- Construct validity
- The overarching evidence that a measure truly captures its theoretical construct.
- Convergent validity
- Evidence that a measure correlates with measures of the same or related constructs.
- Discriminant validity
- Evidence that a measure correlates weakly with measures of unrelated constructs.
Module 5: Sampling
Selecting cases so that inferences to a target population are warranted, and recognizing when they are not.
Populations, Frames, and Probability Sampling
- Distinguish target population, sampling frame, and sample, and identify coverage error.
- Compare simple random, systematic, stratified, and cluster sampling.
Sampling is the logic by which a manageable subset licenses claims about a larger whole. Whether that license is valid depends first on clean definitions and then on how cases are selected. Probability sampling - in which every element has a known, non-zero chance of selection - is what makes formal statistical inference to a population defensible.
The word "known" is doing heavy work here. It is not enough that selection be haphazard, or that the researcher intend to be fair. The selection probability must be specifiable. The mathematics of inference uses that probability to move from what you saw in the sample to a statement about the population, with a quantified margin of error. This lesson defines the populations involved and the standard probability designs, so that you can match a design to a study's inferential goal and its budget.
One case runs through the lesson. A ministry of health wants to know what percentage of the country's 90,000 registered nurses plan to leave the profession within two years. It cannot ask all 90,000. It has a professional registry with 84,000 names, of whom about 3,000 have retired and 1,200 appear twice. It gets cooperation from hospitals in four of twelve regions. Notice how many different groups have already appeared. All nurses, the registry, the four regions, and finally the 900 nurses who answer. Keeping those groups straight is most of what this lesson teaches.
Key idea: Probability sampling means every person has a known, non-zero chance of being picked, and that known chance is what licenses a margin of error.
In plain terms
Keep four groups straight and most of this lesson is easy. There are the people you want to describe. There is the list you actually have. There are the people you can reach. And there are the people who end up in your data. Every gap between those four is a place bias can hide. The one rule that makes statistics work is this. Every person must have a known chance of being picked, and that chance must not be zero. Once you have that, you can put a margin of error on your answer. Four ways to pick people. Pull names at random. Take every tenth name. Split the list into groups and take some from every group, which buys precision. Or pick a few whole groups and study everyone in them, which saves money. That last one costs you precision, and your analysis has to admit it.
Three populations, not one
- The target population is the entire group you wish to describe (all registered nurses in a country).
- The sampling frame is the operational list from which you actually draw (a professional registry). The frame is almost never a perfect image of the target.
- The sample is the set of elements selected from the frame.
The mismatch between target and frame is coverage error. If the registry omits recently licensed or unregistered nurses, no sampling technique applied to it can recover them; the flaw is baked in before selection begins. Identifying frame limitations is a required part of any honest sampling account.
In practice a fourth term is useful. The accessible (study) population is the portion of the target you could realistically reach, and it sits between the target and the frame. Our ministry's target is all 90,000 nurses, but its accessible population is nurses in four cooperating regions. Being explicit about this narrowing matters for a hard reason. Your statistical inference runs strictly to the accessible population. Any further generalization to the full target is an argument you must make on substantive grounds. The statistics do not supply it.
Frames fail in several specific ways beyond simple omission. A frame may contain ineligible units, as a business registry does when it lists closed firms, or the nursing registry does with its 3,000 retirees. It may contain duplicates, as with the 1,200 nurses listed twice, who thereby get double the chance of selection. Or it may show clustering, where several elements sit under one listing, as when a household contains several eligible adults. Each problem needs a fix: screen out ineligibles, de-duplicate, or set a rule for picking one element per listing. Documenting these problems and their remedies is part of a transparent sampling account, not an admission of sloppiness.
Key idea: Target population, sampling frame, accessible population, and sample are four different things, and the gaps between them are where bias enters before you select anyone.
Probability designs
- Simple random sampling (SRS): every element and every combination of elements is equally likely to be chosen. It is the theoretical baseline for inference but requires an enumerable frame.
- Systematic sampling: choose a random start and take every k-th element. Efficient, but dangerous if the frame has a hidden periodicity that aligns with the interval k.
- Stratified sampling: divide the frame into homogeneous strata, say by region, and sample within each one. This guarantees representation of every stratum. When strata are internally homogeneous, it also yields more precise estimates than SRS at the same sample size. Proportionate stratification samples each stratum in proportion to its size. Disproportionate stratification deliberately over-samples small but important strata, and corrects for it later with weights.
- Cluster sampling: divide the population into naturally occurring clusters such as schools or city blocks, randomly select whole clusters, and study all or a random subset of elements within them. It is economical when a full frame is unavailable or fieldwork is spread out geographically. But members of a cluster resemble one another, so cluster samples are less precise than an SRS of the same size.
Two additions complete the toolkit. Multistage sampling chains several methods together. A survey might first sample districts, then schools within districts, then students within schools. That is how national surveys operate when no single frame of individuals exists. Probability-proportional-to-size sampling gives larger clusters a greater chance of selection, so that each ultimate element still has an equal overall probability. It balances efficiency against representativeness. Real large-scale surveys almost always combine these designs rather than using any one in pure form.
The warning about systematic sampling deserves a concrete case. Suppose a frame lists apartments, and every fourth unit is a corner apartment that is larger and more expensive. A systematic sample with an interval of four would select either all the corner units or none of them. Any estimate of apartment size or rent would be badly distorted. That periodicity is invisible unless you inspect how the frame is ordered. It is exactly the hidden structure that turns an efficient design into a biased one.
Key idea: Simple random sampling is the benchmark, systematic sampling is cheap but vulnerable to periodicity, stratification buys precision, and clustering buys affordability.
Design effect and effective sample size
Cluster sampling's precision penalty can be quantified. Members of a cluster tend to resemble one another, an effect measured by the intraclass correlation. So each additional person inside an already-sampled cluster adds less new information than an independent person would. The design effect summarizes how much the variance of an estimate is inflated relative to simple random sampling of the same size. A design effect of 2 means your 1,000-person cluster sample carries the precision of a 500-person simple random sample. Ask 30 nurses on the same ward about leaving and you learn less than you would from 30 nurses at 30 different hospitals, because ward culture makes their answers alike.
Two practical implications follow. Clustered designs must enroll more participants to reach a target precision. And analyses must account for the clustering, or they will understate uncertainty and produce false confidence. Ignoring the design effect is a frequent error. It happens whenever researchers analyze cluster or multistage data with formulas that assume independent observations, and it inflates statistical significance in a way peer reviewers increasingly catch.
Key idea: People inside a cluster resemble each other, so a clustered sample of 1,000 may be worth only 500 independent people - and your analysis has to say so.
Stratified versus cluster
These two are easily confused, yet their structure is opposite. In stratified sampling you divide into groups and sample from every group. You want homogeneity within each stratum, because that is what buys precision. In cluster sampling you divide into groups and sample only some groups, accepting a precision penalty in exchange for logistical feasibility. Here is the heuristic worth remembering. Stratify to reduce sampling error. Cluster to reduce cost.
| Design | All groups sampled? | Effect on precision vs SRS | Main motive |
|---|---|---|---|
| Stratified | Yes, every stratum | Can improve precision | Representation, precision |
| Cluster | No, selected clusters | Typically reduces precision | Cost, feasibility |
Because probability designs support quantified uncertainty, they let you attach a margin of error and confidence interval to estimates. That formal bridge from sample to population is the payoff for the discipline of a known selection probability - and it is precisely what the non-probability methods in the next lesson cannot provide.
A brief numerical intuition shows why stratification helps. Imagine estimating mean income in a city split evenly between a low-income and a high-income district. Within each district incomes vary little, but the gap between districts is large. A simple random sample might by chance draw too many from one district and miss the mark. Stratify and sample equally from both, and each district is guaranteed representation. That removes a whole source of error, so the stratified estimate is more precise at the same total sample size.
Key idea: Stratifying samples from every group; clustering samples from only some groups - the first improves precision, the second saves money.
What drives the required sample size
The size of a probability sample depends on three quantities, and not on the size of the population. That surprises many students. Precision comes first. A narrower desired margin of error requires a larger sample, and halving the margin roughly quadruples the required n. Variability comes second. A more heterogeneous population needs a larger sample to pin down an estimate. Confidence comes third. Demanding 99% rather than 95% confidence widens the interval or requires more cases. Once the population is large relative to the sample, its total barely matters. The ministry needs about the same 1,100 nurses whether the country has 90,000 of them or 900,000.
Key idea: Sample size is driven by the precision you want, the variability you face, and the confidence you demand - not by how big the population is.
Weighting and inference
Sometimes selection probabilities differ across elements, as in disproportionate stratification or probability-proportional-to-size sampling. Unweighted estimates are then biased toward the over-sampled groups. The remedy is design weights. Each case is weighted by the inverse of its selection probability, so over-sampled units count proportionally less. Suppose the ministry over-samples rural nurses at three times the urban rate to get a stable rural estimate. Each rural nurse must then carry one third the weight in the national figure. Public-use survey datasets ship with such weights, and using them is not optional. Analyzing a complex sample as if it were a simple random sample throws away the very structure that made the estimates representative.
Key idea: If some people were more likely to be chosen, weight them down - otherwise your national estimate is really an estimate of whoever you over-sampled.
Selection versus assignment
Two uses of randomness are routinely confused and must be kept apart. Random selection is a sampling procedure. It supports generalizing from a sample to a population, and it addresses external validity. Random assignment is an experimental procedure. It distributes participants across conditions to support causal inference, and it addresses internal validity. A study can have one without the other. A lab experiment often uses random assignment on a convenience sample. A national survey uses random selection with no assignment at all. Confusing the two leads researchers to claim generalizability from an experiment, or causation from a survey.
Key idea: Random selection gets you from sample to population; random assignment gets you from association to cause - and neither substitutes for the other.
A cautionary case
The value of probability sampling shows up most clearly when it is abandoned. Opt-in online polls, in which anyone who encounters a link may respond, can gather enormous numbers yet systematically misrepresent a population, because the people who choose to answer differ from those who do not, and no known probability links the respondents to the whole. Their apparent authority comes from raw counts, but the selection mechanism is unknown, so there is no valid bridge from the sample back to the population it claims to describe.
The moral is one this course returns to repeatedly. Size does not fix bias. A sample of ten million drawn from a self-selected pool estimates that pool very precisely. A few thousand drawn with known probabilities estimates the population. The discipline of a specified selection probability earns the right to generalize; sheer volume does not. The next lesson dissects the seductive but treacherous large non-probability sample and its most famous historical failure. The point is more urgent than ever, because huge datasets scraped from the internet constantly tempt researchers to mistake volume for representativeness.
Key idea: A huge self-selected sample measures the self-selected group extremely well and the population not at all.
Where people get stuck
Three distinctions repay careful attention here.
- Population versus frame versus sample. The population is who you want to describe. The frame is the list you actually draw from. The sample is who you got. The registry with 3,000 retirees and 1,200 duplicates is a frame, not a population, and no amount of clever selection repairs a frame that omits the people you care about.
- Stratified versus cluster. Both split the population into groups. Stratified sampling then takes people from every group; cluster sampling takes only some groups and everyone, or many people, inside them. If you find yourself saying "we stratified by hospital and then sampled four hospitals," you clustered.
- Random selection versus random assignment. One is about who is in your study. The other is about what happens to them once they are in. A randomized trial on volunteers has excellent internal validity and unknown external validity, and a national survey has the reverse.
One more habit is worth building. When you read a survey estimate, ask what the frame was and who could not appear in it. A telephone survey excludes people without phones. A patient-portal survey excludes patients who do not use the portal. Those exclusions are usually correlated with the thing being measured, which is precisely what makes coverage error dangerous rather than merely untidy.
Key idea: Ask who could not possibly have been selected, because that group is where the bias lives.
Try it
Exercise. A ministry wants a national estimate of average reading achievement among fourth-graders, with a stated margin of error, but there is no list of individual pupils, only a list of schools. Recommend a sampling design, name its main tradeoff, and state what the analysis must account for.
Model answer. With no frame of individuals, a pure simple random sample of pupils is impossible, so a multistage cluster design is the natural choice: randomly select schools, ideally with probability proportional to enrollment, then sample classrooms or pupils within the selected schools. Because the ministry wants a national estimate, the design should also stratify schools by region and by urban or rural status before selection, guaranteeing representation of each and improving precision.
The main tradeoff is precision. Clustering pupils within schools reduces precision relative to a simple random sample of the same size, because pupils in a school resemble one another. The design therefore needs more pupils than a naive calculation suggests, inflated by the design effect implied by the intraclass correlation. The analysis must also incorporate the sampling design. That means design weights for the unequal selection probabilities, plus cluster-robust variance estimation. Then the reported margin of error honestly reflects the clustered, stratified structure instead of pretending the pupils were drawn independently.
Recap
- Probability sampling requires a known, non-zero selection probability for every element, and that is what licenses a margin of error.
- Target population, sampling frame, accessible population, and sample are distinct; the gap between target and frame is coverage error.
- Frames also fail through ineligible units, duplicates, and multiple elements under one listing, and each needs a documented remedy.
- Simple random, systematic, stratified, and cluster designs trade precision against cost, with systematic sampling vulnerable to hidden periodicity.
- Clustering inflates variance by the design effect, so clustered studies need more participants and cluster-aware analysis.
- Required sample size depends on desired precision, population variability, and confidence level, not on the size of the population.
- Unequal selection probabilities require design weights, and complex samples must not be analyzed as if simple random.
- Random selection supports generalization while random assignment supports causal inference; a study may have either without the other.
Sources
- Elfil, M., & Negida, A. (2017). Sampling methods in clinical research: An educational review. Emergency, 5(1), e52. pmc.ncbi.nlm.nih.gov
- Martinez-Mesa, J., Gonzalez-Chica, D. A., Duquia, R. P., Bonamigo, R. R., & Bastos, J. L. (2016). Sampling: How to select participants in my research study? Anais Brasileiros de Dermatologia, 91(3), 326-330. pmc.ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Overview of survey research. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Bhattacherjee, A. (2012). Sampling. In Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Krosnick, J. A. (1999). Survey research. Annual Review of Psychology, 50, 537-567. pubmed.ncbi.nlm.nih.gov
- Substance Abuse and Mental Health Services Administration. (2019). National Survey on Drug Use and Health: Methodological summary and definitions. NCBI Bookshelf. ncbi.nlm.nih.gov
- Boynton, P. M., & Greenhalgh, T. (2004). Selecting, designing, and developing your questionnaire. BMJ, 328(7451), 1312-1315. pmc.ncbi.nlm.nih.gov
- Key terms
- Target population
- The entire group about which conclusions are ultimately desired.
- Sampling frame
- The operational list of elements from which a sample is actually drawn.
- Coverage error
- Error arising when the sampling frame does not match the target population.
- Simple random sampling
- A design in which every element and combination has an equal chance of selection.
- Stratified sampling
- Dividing the frame into homogeneous strata and sampling within every stratum.
- Cluster sampling
- Randomly selecting whole naturally occurring groups and sampling elements within them.
Non-Probability Sampling and Sampling Error
- Describe common non-probability methods and their appropriate uses.
- Distinguish sampling error from systematic bias and explain why large biased samples do not help.
Not all research can or should use probability sampling. Qualitative studies, hard-to-reach populations, and early exploratory work routinely rely on non-probability sampling, in which selection probabilities are unknown. The methods are legitimate for their purposes, but they do not license formal statistical generalization to a population, and pretending otherwise is a serious error.
The task is not to avoid non-probability sampling. It is to use it honestly. Choose it when it fits the question, report how cases were selected, and frame conclusions in terms the design can support. This lesson surveys the common methods, then draws the pivotal distinction between sampling error and bias. Confusing those two is behind many overconfident claims built on large but skewed datasets.
Two studies will make the difference concrete. Study X posts a link on a fitness app and collects 50,000 responses about sleep quality. Study Y draws 900 people by random digit dialing and gets 40% of them to answer. X has 55 times more data. Y is the study you would bet on. X measures the opinions of people motivated enough to click a survey link inside a wellness app, extremely precisely. Nothing about having 50,000 of them makes that group resemble the country. Y measures a known population imperfectly, and its imperfection is one you can estimate and report.
Key idea: Non-probability sampling is a legitimate tool for the right question, but it never earns a margin of error.
In plain terms
Sometimes you cannot pick people at random. You take whoever is available, or you choose people on purpose because they know something, or you ask participants to refer their friends. All of that is fine. What is not fine is then reporting a margin of error, because that number is computed from selection probabilities you do not have. Now the single most important idea in this lesson. There are two kinds of error, and they behave in opposite ways. One shrinks as you collect more data. That is ordinary sampling noise. The other does not shrink at all, ever. That is bias. A survey of fifty thousand self-selected people has almost no noise and possibly enormous bias, so it gives you a very precise answer about the wrong group. Ask yourself one question: would more of the same people fix this? If not, it is bias.
Common non-probability methods
- Convenience sampling: selecting whoever is easiest to reach (students in your class, shoppers at one mall). Cheap and fast, but its representativeness is unknown and usually poor.
- Purposive (judgment) sampling: deliberately selecting cases that fit criteria relevant to the research question. That means information-rich cases in qualitative work, or extreme, typical, and deviant cases chosen on principled grounds. Here selectivity is a feature. It aims at insight rather than generalization.
- Quota sampling: setting targets for subgroup counts, for example equal numbers in each age band, and filling them non-randomly. It mimics the structure of stratification without random selection inside the cells, so bias can enter through how the quotas get filled.
- Snowball sampling: existing participants refer others, useful for hidden or stigmatized populations reachable only through social networks. Because referrals follow network ties, the sample over-represents the well-connected.
Two further variants appear especially in qualitative work. Theoretical sampling is central to grounded theory. It selects each new case based on what the emerging analysis needs next, so sampling and analysis proceed together rather than in sequence. Maximum-variation sampling deliberately spans a wide range of cases, to see whether a pattern holds across diversity. In these designs the absence of random selection is a tool rather than a flaw. The goal is conceptual insight into how and why, not a population estimate.
Whatever the method, honesty about it is non-negotiable. A study using a convenience sample should say so plainly. It should describe who was included and excluded. And it should confine its claims to the group actually studied, offering only cautious, reasoned suggestions about wider relevance. The failure mode to avoid is a convenience sample dressed up in the language of generalization, reporting margins of error or population percentages that its design cannot justify.
Key idea: Purposive and theoretical sampling choose cases on purpose because insight, not representativeness, is the goal - and they should never report a margin of error.
Two fundamentally different problems
It is essential to separate two sources of error that behave in opposite ways as sample size grows.
- Sampling error is the random gap between a sample estimate and the population value. It arises purely because you observed a subset. It is not a mistake. It is inherent to sampling, it is quantifiable in probability samples through the standard error, and it shrinks as sample size increases.
- Bias is a systematic tendency for estimates to be off in a particular direction. It comes from flawed selection, coverage, nonresponse, or measurement. Bias does not shrink with sample size. A larger biased sample simply estimates the wrong value more precisely.
The two errors combine into total error. A compact way to see their relationship is that the mean squared error of an estimate equals its variance plus the square of its bias. Sampling error is the variance term, which more data reduces. Bias is the other term, which more data does not touch. A design can be low in one and high in the other. So a precise, tightly clustered set of estimates can still sit on the wrong value. It is the statistical equivalent of a rifle that groups its shots tightly, well away from the bullseye. Study X above has almost no sampling error and unknown, probably large, bias.
Key idea: More data shrinks sampling error and does nothing at all to bias.
Confidence intervals, properly understood
In a probability sample, sampling error is quantified by the standard error, and a confidence interval turns that into a range. A 95% confidence interval is constructed by a procedure that, across many repeated samples, would contain the true population value about 95% of the time. That is a statement about the long-run behavior of the method, not a probability that this one particular interval contains the truth, which either does or does not.
The common misreadings are worth naming. A 95% interval does not mean there is a 95% probability the parameter lies inside this one interval, nor that 95% of future sample estimates will fall in it, nor that 95% of the population lies within it. Interpreted correctly, a wide interval is a candid confession of imprecision, and reporting it alongside the point estimate is far more informative than a bare number that hides how much the estimate might move in another sample.
A worked reading fixes the idea. A survey estimates that 62% of a population approves of a policy, with a 95% confidence interval from 58% to 66%. Two things can be said defensibly. The data are consistent with true approval anywhere in that band. And the estimating procedure captures the truth 95% of the time in the long run. A claim that approval is exactly 62% ignores the interval. A claim that there is a 95% chance approval lies between 58 and 66 misstates what the interval formally means, even though the practical upshot is similar.
Key idea: A confidence interval describes how the procedure behaves over many samples, and it quantifies sampling error only - it is silent about bias.
The large-sample fallacy
The most famous illustration is the 1936 Literary Digest poll, which mailed millions of ballots and collected about 2.4 million responses, yet wrongly predicted the U.S. presidential election. Its frame - drawn heavily from telephone and automobile owners during the Depression - was skewed toward the affluent, and those who returned ballots were self-selected. Enormous size could not rescue a biased design; a much smaller but better-sampled contemporaneous poll got it right. The lesson is permanent: quantity does not cure bias. When you read a study boasting a huge sample, ask first how cases were selected, not how many there were.
Selection bias wears many disguises, and one classic is survivorship bias. During the Second World War, analysts studying returning bombers wanted to armor the areas most riddled with bullet holes. The statistician Abraham Wald pointed out the opposite conclusion. The holes marked places where a plane could be hit and still come home. So the armor belonged where the survivors showed no damage, because those hits were on planes that never returned. The sample of survivors systematically excluded the cases most relevant to the question. That is the essence of a biased frame.
Key idea: Millions of self-selected responses got the 1936 election wrong while a few thousand well-sampled ones got it right, and quantity has not cured bias since.
When bigger can be worse
A striking modern result sharpens the warning against large biased samples. Analyses of recent election surveys have shown that a self-selected sample of millions can carry the effective precision of only a few hundred randomly drawn respondents, because a small correlation between whether someone responds and how they would answer contaminates every additional case. As the sample grows, that correlation does not average away; it scales, so the confidence interval shrinks around an estimate that is confidently wrong.
The practical lesson updates Literary Digest for the era of massive online data. Precision and representativeness are different virtues. Modern datasets often deliver the first while quietly lacking the second. Hand a researcher a dataset of a million records and the right question is not how much the size reduces sampling error. The right question is whether the process that generated the records is tied to the outcome under study. That tie, not the count, governs whether the data can be trusted.
Key idea: When responding is even slightly related to the answer, extra cases sharpen a wrong number instead of correcting it.
Nonresponse
Even a well-drawn probability sample degrades when those who respond differ systematically from those who do not. That is nonresponse bias. A 15% response rate from a perfect frame can mislead more than a modest but complete one. The missing 85% may share a relevant characteristic, say disengagement, that is correlated with the outcome. Reporting response rates and probing who is missing is therefore part of a defensible sampling account, not an optional courtesy.
Two nuances refine the nonresponse story. First, distinguish unit nonresponse from item nonresponse. Unit nonresponse means a sampled person does not participate at all. Item nonresponse means a respondent skips particular questions. They call for different remedies, broadly weighting for the former and imputation for the latter. Second, a low response rate does not by itself prove bias. Bias arises only when respondents differ from non-respondents on the variable of interest. So a study should investigate that possibility rather than assuming a high rate is automatically safe and a low one automatically fatal.
Where nonresponse threatens representativeness, researchers can compare respondents to known population benchmarks. They can then apply post-stratification weights to bring the sample into line on observed characteristics. This helps, but it cannot fully cure the problem. It adjusts only for differences you can measure, leaving any unmeasured difference between responders and non-responders uncorrected. Transparency about the response rate, and about who is likely missing, lets readers judge how far to trust the estimates.
Key idea: A low response rate is a warning, not a verdict - what matters is whether the people who answered differ from those who did not on the thing you measured.
Defending a non-probability sample
Non-probability samples can still support careful inference when their limits are respected. Qualitative work substitutes transferability for statistical generalization. The researcher describes the sample and setting richly, and readers judge whether the findings might apply to other contexts. Modern opt-in surveys use model-based approaches instead. They match or weight the sample to a reference population on many covariates. That reduces bias without eliminating it, and it shifts the burden onto the correctness of the adjustment model rather than removing the burden. The honest posture is to name which population the adjusted estimate is meant to represent, and to treat that claim as a modeling assumption open to challenge.
Key idea: A non-probability sample earns credibility by describing itself honestly, not by borrowing the vocabulary of probability sampling.
Where people get stuck
Three confusions dominate this material.
- Sampling error versus bias. Sampling error is random and shrinks with n. Bias is directional and does not. The tell is simple: if adding respondents would fix it, it is sampling error. If adding more of the same kind of respondent would not fix it, it is bias.
- Probability versus non-probability sampling. The dividing line is not size, quality, or effort. It is whether every element's chance of selection is known and non-zero. Quota sampling looks like stratification and is not, because selection within cells is not random.
- Sample size versus sample quality. Study X's 50,000 responses give it a tiny confidence interval around a number that describes app users who like surveys. Study Y's 900 calls give a wider interval around a number that describes the country. Only one of those intervals means what it says.
Finally, keep saturation out of this discussion. Saturation is a stopping rule in qualitative sampling, judged by whether new data are still generating new concepts. It is not a sample-size calculation and it is not a substitute for one. A quantitative study cannot claim saturation, and a qualitative study should not claim a margin of error.
Key idea: Ask whether more data would fix the problem; if not, you are looking at bias, and no sample size will save you.
Try it
Exercise. A wellness company reports that "in a survey of 50,000 app users, 80% said our program improved their sleep, with a margin of error of +/-0.4 percentage points, so the program clearly works for the general population." Identify every inferential problem, and state what the data can and cannot support.
Model answer. The claim commits several errors at once. The sample is a convenience sample of self-selected app users, so it cannot be generalized to the general population, and the margin of error is meaningless here because it quantifies only sampling error, which does not apply when there is no known selection probability. The huge size makes the estimate precise about the wrong group, the large-sample fallacy in miniature. Nonresponse and self-selection are likely severe, since satisfied users are more inclined both to keep using the app and to answer favorably.
There is also no comparison and no causal design. "Improved their sleep" rests on self-report without a counterfactual, so the program is confounded with expectation and with regression to the mean. What the data can support is one descriptive statement: among these responding users, 80% reported improved sleep. What they cannot support is generalization to the population, or a causal claim that the program works. A defensible study would need a probability or well-modeled sample for generalization, and a controlled comparison for causation. A satisfaction survey provides neither.
Recap
- Non-probability sampling covers convenience, purposive, quota, snowball, theoretical, and maximum-variation designs, each legitimate for particular purposes.
- None of them licenses a margin of error, because a margin of error is computed from known selection probabilities.
- Sampling error is random and shrinks as n grows; bias is systematic and does not shrink at all.
- A confidence interval describes the long-run behavior of the estimating procedure and quantifies only sampling error.
- The Literary Digest poll and Wald's bombers show that a huge sample from a skewed frame estimates the wrong quantity precisely.
- Nonresponse bias depends on whether respondents differ from non-respondents on the outcome, not on the response rate alone.
- Non-probability designs defend themselves through rich description and transferability, or through explicit model-based adjustment whose assumptions are open to challenge.
Sources
- Elfil, M., & Negida, A. (2017). Sampling methods in clinical research: An educational review. Emergency, 5(1), e52. pmc.ncbi.nlm.nih.gov
- Palinkas, L. A., Horwitz, S. M., Green, C. A., Wisdom, J. P., Duan, N., & Hoagwood, K. (2015). Purposeful sampling for qualitative data collection and analysis in mixed method implementation research. Administration and Policy in Mental Health, 42(5), 533-544. pmc.ncbi.nlm.nih.gov
- Martinez-Mesa, J., Gonzalez-Chica, D. A., Duquia, R. P., Bonamigo, R. R., & Bastos, J. L. (2016). Sampling: How to select participants in my research study? Anais Brasileiros de Dermatologia, 91(3), 326-330. pmc.ncbi.nlm.nih.gov
- Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. pmc.ncbi.nlm.nih.gov
- Moser, A., & Korstjens, I. (2018). Series: Practical guidance to qualitative research. Part 3: Sampling, data collection and analysis. European Journal of General Practice, 24(1), 9-18. pmc.ncbi.nlm.nih.gov
- Krosnick, J. A. (1999). Survey research. Annual Review of Psychology, 50, 537-567. pubmed.ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Sampling. In Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Key terms
- Non-probability sampling
- Selection in which elements' probabilities of inclusion are unknown.
- Convenience sampling
- Selecting the most easily accessible cases, with unknown representativeness.
- Purposive sampling
- Deliberately choosing information-rich or criterion-fitting cases for insight.
- Snowball sampling
- Recruiting further participants through referrals from existing ones.
- Sampling error
- Random discrepancy between sample and population that shrinks as sample size grows.
- Bias (sampling)
- A systematic directional error that does not diminish with larger samples.
Module 6: Experimental and Quasi-Experimental Design
Designing studies that can support causal claims, and the validity threats that determine whether they succeed.
The Logic of Experiments and Randomization
- Explain the counterfactual logic of causation and how random assignment addresses it.
- Distinguish random assignment from random sampling and identify their distinct roles.
The experiment is the strongest single design for causal inference, because it is built to satisfy the logical requirements of a causal claim. John Stuart Mill stated those requirements and later work refined them. There are three. The cause must covary with the effect. The cause must precede the effect in time. And alternative explanations must be ruled out. Correlational designs can establish the first. Only a strong design secures all three.
The third condition is the hard one, and it is where experiments earn their status. Ruling out every alternative explanation for an association is nearly impossible in observational data. Some unmeasured difference between the compared groups can always be lurking. The experiment's contribution is a procedure. Random assignment addresses even the alternatives you never thought to measure. That is why the logic of this lesson underlies every causal claim you will make or judge.
A single case will carry the lesson. A company offers a 20-hour management training course and lets employees volunteer. A year later, the 140 volunteers have performance ratings 15% higher than the 300 who did not sign up. The training looks excellent. Now ask a harder question. Who volunteers for 20 hours of extra work in the evenings? Ambitious people. People already being groomed for promotion. People whose managers encouraged them. Every one of those traits raises performance ratings without any training at all. The 15% gap is real. What produced it is completely unknown, and no larger sample will tell you.
Key idea: A causal claim needs covariation, time order, and the elimination of rival explanations, and the third is what almost every weak design fails.
In plain terms
To say A caused B you need three things. A and B have to move together. A has to come first. And nothing else can explain it. The first two are usually easy. The third is where studies fail. The problem is simple to state. You would like to know what would have happened to this person without the treatment, and you can never see that. So you use another group as a stand-in. That only works if the two groups were alike to begin with. Let people choose their own group and they will not be alike, because the kind of person who volunteers is different. Flip a coin instead and the groups end up alike on everything, including the things you never thought to measure. That is the whole reason random assignment is special. Nothing else does that.
The counterfactual and its problem
A causal effect is defined by a counterfactual. It is the difference between what happened to a unit under treatment and what would have happened to that same unit without treatment. Here is the difficulty, known as the fundamental problem of causal inference. We never observe both outcomes for the same unit, because a person is either treated or not. Experiments solve this at the group level. If two groups are equivalent in expectation before treatment, then the untreated group's outcome estimates the counterfactual for the treated group.
Statisticians formalize this with the potential outcomes framework, associated with Donald Rubin. Each unit has two potential outcomes, one under treatment and one under control. The individual causal effect is their difference. Only one is ever observed, so the individual effect is unknowable. But the average treatment effect across a group is estimable if the groups are exchangeable. This reframing makes precise what "equivalent in expectation" means, and why it licenses using one group's outcome in place of the other's missing counterfactual.
In observational data the two groups are usually not exchangeable, and that is the source of selection bias. People who choose a treatment differ from those who do not, in ways that also affect the outcome. A raw comparison then confounds the treatment with those differences. One famous pattern makes the danger vivid. Patients who adhere to any regimen, even a placebo, tend to have better outcomes, because adherence stands in for conscientiousness and health behavior. Our training volunteers are the same phenomenon in a workplace. The experiment breaks the link by taking the choice away and letting chance assign it.
Key idea: You can never see what would have happened to the same person under the other condition, so a comparison group has to stand in for it.
Random assignment as the engine
Random assignment allocates participants to conditions by a chance mechanism. Its power lies in one property. In expectation, it balances the groups on all characteristics before treatment: measured and unmeasured, known and unknown. That is what neutralizes confounding. Any pre-existing difference is distributed by chance rather than tied to condition. Ambition still exists among the trainees, but now roughly as much of it sits in each group. No other technique controls unmeasured confounders, which is why the randomized controlled trial (RCT) is the benchmark for causal claims. Randomization does not guarantee balance in any single small study. It guarantees the absence of systematic bias and makes the probability of imbalance calculable.
The mechanics matter for credibility. Assignment should use a genuine chance device, such as a computer random-number generator. It also needs allocation concealment, so whoever enrolls participants cannot foresee or steer the next assignment. Simple randomization is a coin flip for each unit. Block randomization assigns within small blocks to keep group sizes balanced over time. Stratified randomization randomizes within strata to ensure balance on a few important variables. These refinements reduce the chance imbalance that pure simple randomization can leave in a small trial.
A common misunderstanding is that randomization must produce balanced groups. It promises no such thing in any one study. It promises the absence of systematic bias and a calculable probability of imbalance. That is why trials report a baseline table comparing the groups on measured characteristics. The table is not there to prove balance was achieved. It lets readers see the allocation that actually occurred. A chance imbalance on a measured covariate can be adjusted for in the analysis. An unmeasured confound cannot, and only randomization addresses it.
An experiment also needs enough participants to detect the effect it seeks, which links assignment back to the power analysis of an earlier lesson. A trial can randomize cleanly and still enroll too few participants to find a real effect. Its null result is then inconclusive rather than evidence of no effect. Setting the sample size in advance, from the smallest effect worth detecting, is as much a part of experimental design as the randomization procedure itself.
Key idea: Random assignment is the only tool that balances the confounders you never thought to measure, and that is the whole reason it is the gold standard.
A critical distinction
Two uses of "random" are routinely conflated but do entirely different jobs:
| Random sampling | Random assignment | |
|---|---|---|
| Question addressed | To whom can results generalize? | Did the treatment cause the effect? |
| Supports | External validity | Internal validity |
| Mechanism | Selecting cases from a population | Allocating selected cases to conditions |
A study can have one without the other. A tightly controlled lab experiment on undergraduate volunteers uses random assignment, which gives strong internal validity, but not random sampling, which leaves external validity limited. A survey may use random sampling, giving generalizable description, with no assignment and therefore no causal warrant. Rigorous design requires knowing which one you have.
Key idea: Random sampling tells you who the finding applies to; random assignment tells you whether the treatment caused it.
Core experimental elements
Beyond assignment, experiments deploy several controls. A control group provides the counterfactual baseline. Blinding keeps participants unaware of condition, which is single-blind, and often keeps experimenters unaware too, which is double-blind. It prevents expectancy effects from contaminating outcomes. A placebo control isolates the specific effect of a treatment from the effect of merely receiving something. And a manipulation check verifies that the independent variable was actually delivered as intended. None of these is decoration. Each closes a specific route by which an apparent effect could turn out to be an artifact.
Control conditions come in several forms, and each answers a different question. A no-treatment control shows what happens with nothing added. A placebo control separates the specific ingredient from the effect of receiving attention. An active or standard-care control compares the new treatment against the current best, and it is often the ethically required comparison when an effective treatment already exists. A waitlist control eventually delivers the treatment, which addresses the ethics of withholding while still providing a baseline. Choosing the control defines exactly which causal question the experiment answers. If the training trial uses a waitlist control, it answers "does training beat nothing." If it uses a control group given 20 hours of unrelated seminars, it answers the sharper question "does the content matter, or just the time away from the desk."
Experiments also run into an ethical boundary. Randomly assigning people to conditions is justified only under equipoise: genuine uncertainty in the expert community about which condition is better. Once strong evidence establishes that a treatment helps, withholding it from a control group becomes hard to defend. That is why active-control and waitlist designs exist, and why data-monitoring boards can stop a trial early when one arm proves clearly superior. The strongest design is not always the permissible one.
Key idea: Your choice of control group decides which question the experiment actually answers.
Between, within, and factorial designs
Experiments differ in how conditions map onto participants. A between-subjects design assigns different people to each condition. It avoids carryover but needs more participants. A within-subjects design exposes the same people to every condition. It gains power, because each person serves as their own control, but it pays for that with order and carryover effects. Counterbalancing is designed to neutralize those. A factorial design crosses two or more independent variables at once. It is efficient, and it reveals interactions, in which the effect of one factor depends on the level of another.
An interaction is often the most interesting result. Suppose a factorial study crosses caffeine, yes or no, with sleep, rested or deprived. It finds that caffeine helps the deprived and does nothing for the rested. Neither main effect tells the whole story. The payoff is the interaction, which says the effect of caffeine depends on sleep state. Averaging over the sleep factor would hide it completely. That is why factorial designs, by testing factors jointly, reveal conditional effects that separate single-factor experiments would miss.
Key idea: Between-subjects designs avoid carryover, within-subjects designs buy power, and factorial designs are the only way to see that an effect depends on something else.
From assignment to analysis
Randomization can be undone at the analysis stage if you are careless. The intention-to-treat principle analyzes participants in the group to which they were assigned, whether or not they complied. Moving dropouts or non-compliers between groups reintroduces exactly the self-selection that randomization removed. Suppose the 12 trainees who skipped most sessions get reclassified into the control group. Those are the least motivated people, and shifting them makes the training group look better for reasons that have nothing to do with training. A per-protocol analysis, restricted to compliers, answers a different and more fragile question. It should supplement the intention-to-treat result, never replace it.
Attrition is the practical threat to this discipline. When participants drop out differentially across conditions, the groups that remain are no longer the groups that were randomized, and comparability erodes. Reporting standards such as CONSORT require a flow diagram tracking every participant from enrollment through analysis. Its purpose is precisely to let readers see whether attrition threatens the causal claim. A trial that loses a third of one arm and a tenth of the other owes an explicit account of why.
Key idea: Analyze people in the group they were assigned to, because moving non-compliers around quietly restores the self-selection randomization removed.
What randomization does not fix
Random assignment secures the internal validity of the comparison at the moment of allocation. Several threats can still creep in afterward. Differential attrition, already noted, unbalances the groups after randomization. Diffusion occurs when the control group learns and adopts elements of the treatment, which shrinks the apparent effect. In our company, trained employees will naturally tell untrained colleagues what they learned over lunch. Demand characteristics and experimenter effects arise when participants or researchers act on their guesses about the hypothesis. Blinding is what prevents those.
The Hawthorne effect names a related worry. People may change their behavior simply because they know they are being observed, so an apparent treatment effect may partly reflect the attention itself. A well-designed experiment anticipates these problems. It blinds where it can, keeps conditions separated, and includes controls that receive equivalent attention. The next lesson catalogues the full set of threats to internal validity. The point here is narrower: randomization is necessary but not sufficient for a clean causal claim.
Key idea: Randomization protects the comparison at the moment of assignment and nothing after it - attrition, diffusion, and expectancy can still ruin the study.
Where people get stuck
Three points cause most of the confusion.
- Random selection versus random assignment. Selection is about who is studied and governs generalization. Assignment is about who gets what and governs causation. A trial on volunteers is strong on the second and weak on the first, and saying so is not a criticism but a description.
- Balance is not promised. Randomization guarantees no systematic bias, not equal groups in your particular study. A baseline imbalance in a small trial is expected, is not evidence of a broken randomization, and can be adjusted for if it was measured.
- Control group does not mean untreated. No-treatment, placebo, active, and waitlist controls answer four different questions. "We had a control group" tells a reader almost nothing until you say which kind.
One further habit: whenever you see a comparison between people who chose something and people who did not, name the trait that drives the choice. Volunteers for training are ambitious. Patients who adhere to a placebo are conscientious. Families who enroll in a preschool program are engaged. That named trait is your rival explanation, and only assignment by chance removes it.
Key idea: When people choose their own condition, ask what kind of person chooses it - that answer is the confound.
Try it
Exercise. A company tests a new training program by letting employees volunteer for it, then comparing the volunteers' later performance to that of non-volunteers and concluding the program raised performance by 15%. Redesign the study as a true experiment, and explain which specific threat your redesign removes that the original could not.
Model answer. The original is observational, not experimental, because employees selected themselves into the program. Volunteers likely differ from non-volunteers in motivation, ambition, and prior performance, all of which affect the outcome, so the 15% gap confounds the program with those pre-existing differences. This is textbook selection bias, and no amount of data collected this way can separate the program's effect from the kind of person who volunteers.
A true experiment fixes this. Randomly assign willing employees to receive the training now or to a waitlist control, using a random-number generator with allocation concealment, and analyze everyone by intention to treat. Random assignment makes the two groups equivalent in expectation on all characteristics, measured and unmeasured. A later difference in performance can then be attributed to the program rather than to who chose it. The specific threat removed is confounding by self-selection, which the original design guaranteed and randomization dissolves. Blinding the performance raters to condition would further guard against expectancy effects.
Recap
- A causal claim requires covariation, temporal precedence, and elimination of alternative explanations, and only the third is hard.
- Causal effects are counterfactual, and because both potential outcomes are never observed for one unit, a comparison group must stand in.
- Random assignment balances measured and unmeasured characteristics in expectation, which is why it and nothing else controls unknown confounders.
- Randomization promises no systematic bias rather than balance in any single study, which is why trials publish a baseline table.
- Random sampling supports generalization and random assignment supports causal inference, and a study may have either alone.
- Control groups, blinding, placebos, and manipulation checks each close a specific alternative explanation, and the choice of control fixes the question asked.
- Between-subjects, within-subjects, and factorial designs trade participants against carryover, and only factorial designs reveal interactions.
- Intention-to-treat analysis preserves randomization, while attrition, diffusion, demand characteristics, and Hawthorne effects can still undermine it.
Sources
- Sibbald, B., & Roland, M. (1998). Understanding controlled trials: Why are randomised controlled trials important? BMJ, 316(7126), 201. pmc.ncbi.nlm.nih.gov
- Hariton, E., & Locascio, J. J. (2018). Randomised controlled trials: The gold standard for effectiveness research. BJOG, 125(13), 1716. pmc.ncbi.nlm.nih.gov
- Kang, M., Ragan, B. G., & Park, J. H. (2008). Issues in outcomes research: An overview of randomization techniques for clinical trials. Journal of Athletic Training, 43(2), 215-221. pmc.ncbi.nlm.nih.gov
- Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science & Medicine, 210, 2-21. pmc.ncbi.nlm.nih.gov
- Schulz, K. F., Altman, D. G., & Moher, D. (2010). CONSORT 2010 statement: Updated guidelines for reporting parallel group randomised trials. EQUATOR Network. equator-network.org
- Woodward, J. (2016). Causation and manipulability. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Experiment basics. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Key terms
- Counterfactual
- The outcome that would have occurred for a unit under the condition it did not receive.
- Fundamental problem of causal inference
- That both potential outcomes can never be observed for the same unit.
- Random assignment
- Allocating participants to conditions by chance, balancing all confounders in expectation.
- Randomized controlled trial
- An experiment using random assignment to a treatment and control, the benchmark for causal claims.
- Blinding
- Concealing condition from participants and/or experimenters to prevent expectancy effects.
- Manipulation check
- A measure verifying that the independent variable was delivered as intended.
Threats to Internal Validity
- Define internal validity and identify the classic single-group and multi-group threats.
- Match a design flaw to the specific validity threat it introduces.
Internal validity is the degree to which a study establishes that the independent variable, and not something else, caused the change in the dependent variable. Campbell and Stanley catalogued the recurring threats, and knowing them by name lets you diagnose a weak design instantly. Many threaten single-group or pre-post designs; randomization neutralizes several but not all.
Treat the catalogue as a checklist to run against any causal claim, your own or someone else's. For each threat, ask one question. Could this produce the observed result even if the treatment did nothing? A design is strong to the exact degree that it has an answer for each. Naming threats is what turns a vague unease about a study into a specific, defensible critique a committee will respect.
One study will get diagnosed over and over in this lesson. A middle school gives a reading test in September, puts the lowest-scoring 40 students into a daily tutoring program, retests them in May, and reports a 14-point average gain. The principal is delighted. By the end of this lesson you will be able to name five separate reasons those 40 students would have scored higher in May even if the tutor had read them the phone book, and you will be able to name the single design change that rules out all five at once.
Key idea: Internal validity is whether your treatment, rather than something else, produced the change you observed.
In plain terms
Someone shows you a before-and-after gain and says the programme worked. Run five checks before you agree. Did something else happen in the world during that time? That is history. Would people have changed anyway, just by getting older or more practised? That is maturation. Did taking the test the first time help them do better the second time? That is testing. Did the measuring instrument or the person scoring it drift? That is instrumentation. And were people picked because they scored at an extreme? If so, they were always going to drift back toward average. That is regression to the mean. All five look exactly like a treatment effect. One design change kills all five at once: add a comparison group that goes through the same year and the same tests without the treatment. Then subtract.
Four kinds of validity
Internal validity is one of four types that Cook and Campbell distinguished, and keeping them separate prevents muddled critiques. Statistical conclusion validity concerns whether the inference about covariation is sound, threatened by low power, violated assumptions, or fishing. Internal validity concerns whether the covariation is causal. Construct validity concerns whether the operations match the intended constructs, the measurement question of earlier lessons. External validity concerns whether the effect generalizes, the subject of the next lesson.
The four often trade against one another. A tightly controlled laboratory study maximizes internal validity by stripping away real-world noise, but that same artificiality can weaken external validity. Recognizing which validity a design privileges lets you read its limitations charitably rather than faulting a lab experiment for not resembling the world or a field study for imperfect control. A research program advances by combining designs whose strengths cover one another's weaknesses.
Campbell called internal validity the sine qua non, the indispensable condition, of a causal study. A result that cannot be trusted as causal in the sample cannot meaningfully generalize anywhere. This does not make external validity unimportant. It sets a priority. First secure that the effect is real here. Then ask how far it travels. A beautifully generalizable finding that is actually an artifact of regression or selection generalizes nothing worth having.
Key idea: Statistical conclusion validity, internal validity, construct validity, and external validity are four different questions, and internal validity comes first.
Threats especially dangerous to single-group pre-post designs
- History: an external event between pretest and posttest that affects the outcome, such as a policy change during a study of an economics intervention. The observed change may be due to the event rather than the treatment. In our school, the district also bought new library books in January.
- Maturation: natural changes within participants over time, such as growing older, more tired, or more practiced, that mimic a treatment effect. Nine-year-olds read better in May than in September whatever anyone does.
- Testing: the effect of taking a pretest on posttest scores. Participants improve simply from prior exposure to the test format.
- Instrumentation: a change in the measuring instrument or the observers over time. A rater grows more lenient, so the change reflects the instrument rather than the participants.
- Regression to the mean: when participants are selected for extreme scores, their scores tend to move toward the average on retest for purely statistical reasons, with no treatment involved.
These threats are most acute in the one-group pretest-posttest design, which measures one group before and after treatment with no comparison. Any gain could be the treatment, but it could equally be history, maturation, testing, instrumentation, or regression, and the design cannot tell them apart. This is why a single before-and-after comparison, however intuitive it feels, is among the weakest bases for a causal claim, and why reviewers treat it with immediate suspicion.
Two of these threats are easy to confuse. History is an external event, something that happens in the world during the study, such as a recession striking midway through a jobs program. Maturation is an internal process within participants, such as children becoming more articulate simply by growing older. The test is whether the change would have happened without any outside event: if yes, suspect maturation; if it took a specific occurrence, suspect history. Both masquerade as treatment effects in a one-group design.
Instrumentation is subtle, because the participants may not change at all. Imagine a study where trained observers rate classroom disruption before and after an intervention. Through fatigue or shifting standards, the observers become stricter over the term. Disruption scores then rise even when behavior is constant. A lenient drift in the other direction can manufacture an apparent improvement. Human raters are instruments, and instruments drift. That is why you blind raters to time and condition, and check their agreement periodically.
Key idea: In a one-group before-and-after study, history, maturation, testing, instrumentation, and regression all look exactly like a treatment effect.
Regression to the mean, the subtlest threat
Regression to the mean deserves its own treatment, because it fools even experienced researchers. Whenever a measure is imperfectly reliable, individuals selected for extreme scores will on average score less extremely the second time. The reason is mechanical. Chance contributed to their first extreme value, and chance does not repeat itself. The more extreme the selection, and the less reliable the measure, the larger the regression.
A concrete case shows the trap. A clinic enrolls patients with the highest blood-pressure readings, treats them, and finds their average pressure lower a month later. Some of that drop was guaranteed by regression alone. A high reading is high partly by momentary chance, and the next reading tends to fall back. Attributing the whole drop to treatment overstates the effect. Only a control group, selected the same way and measured the same way, shows how much of the change is regression and how much is real. Our 40 tutored students were chosen precisely because they scored lowest, so some of their 14-point gain was arithmetically certain before the tutor arrived.
Key idea: Pick people for their extreme scores and they will drift back toward average on retest, no matter what you do to them in between.
Threats especially dangerous to multi-group designs
- Selection: pre-existing differences between groups when assignment is not random, so a posttest difference may reflect who was in each group rather than the treatment. This is the threat random assignment is designed to remove.
- Attrition (mortality): differential dropout between groups. If the treatment group loses its least-motivated members while the control retains everyone, the groups are no longer comparable even if assignment began random.
- Selection interactions: selection can combine with maturation or history so that non-equivalent groups also change at different natural rates.
Cook and Campbell extended the catalogue with threats that arise when groups notice the comparison and react to it. Compensatory rivalry, sometimes called the John Henry effect, occurs when a control group resents its status and works unusually hard, closing the gap. Resentful demoralization is the opposite. A control group gives up and performs worse, which exaggerates the treatment's apparent benefit. Compensatory equalization occurs when administrators, uneasy about withholding a desirable treatment, quietly give the control group compensating resources. A principal who cannot bear to leave 40 struggling readers untutored will find them a volunteer, and the study loses its comparison.
A further subtlety is local history, sometimes called selection-by-history: an event affects one group but not the other during the study, such as a fire drill interrupting only the treatment classroom's posttest. Because it strikes the groups unequally, randomization does not fully neutralize it, which is why researchers try to administer conditions under matched circumstances and to keep groups from contaminating one another.
A tempting but flawed remedy deserves warning. When randomization is impossible, researchers often match or statistically adjust non-equivalent groups on measured variables, then treat the comparison as if it were clean. Matching helps with the confounds you can see. It cannot balance the ones you cannot see, and those are exactly what selection threatens. This is the core limitation of observational causal claims. It is also why the next lesson develops quasi-experimental designs that try, with varying success, to approximate the counterfactual randomization would have supplied.
Key idea: Selection is the threat that random assignment removes, and attrition is the threat that can undo random assignment after the fact.
A worked diagnosis
Suppose a reading clinic enrolls the lowest-scoring 5 percent of students, delivers tutoring, and reports large gains at retest. Before crediting the tutoring, a careful reader flags at least three threats. Regression to the mean predicts that an extreme-low group will rise on retest regardless. Maturation predicts reading improves with time in school anyway.
Testing predicts that familiarity with the assessment inflates the second score. Without a comparable control group that experiences the same passage of time, the same testing, and the same statistical regression, the design cannot separate the tutoring's effect from these artifacts. The remedy is structural. Add a randomized control group. Both groups then experience history, maturation, testing, and regression equally, so subtracting one from the other removes all four at once.
Key idea: A control group that goes through everything except the treatment subtracts every threat both groups share.
Reading designs in Campbell and Stanley notation
Campbell and Stanley introduced a compact notation that makes a design's threats visible at a glance. O denotes an observation or measurement, X a treatment, and R random assignment, with each group written on its own row in time order. The weak one-group design is O X O. The much stronger pretest-posttest control-group design is written as two rows, R O X O and R O O, showing that both groups are randomized, both are measured before and after, and only one receives the treatment.
Reading the notation, you can locate each threat. The second row is the control group. It experiences the same history, maturation, testing, and regression as the first, so subtracting the two rows removes those threats. The R at the front of each row removes selection. Two things the notation cannot show are attrition, which happens after assignment, and local history, which strikes one group only. That is why even a strong design still requires vigilant reporting of who dropped out and what happened during the study.
Key idea: Write a design as rows of R, O, and X and its threats become visible at a glance.
Why the catalogue matters
Each named threat corresponds to a design feature that removes it. A pretest-posttest control-group design handles most single-group threats through the control group. Random assignment handles selection. Tracking and reporting dropout addresses attrition. Stable, blinded instruments handle instrumentation. Design is largely the systematic anticipation and closure of these specific threats before data collection. It is never a hopeful patch applied afterward.
Key idea: Every named threat has a matching design feature, so learning the list is learning how to build studies.
Where people get stuck
Four pairs are worth pinning down.
- History versus maturation. History is something that happens in the world; maturation is something that happens inside the participants. New library books are history. Children simply getting older is maturation.
- Testing versus instrumentation. In testing, the participants change because they took the test before. In instrumentation, the measuring device or the rater changes while the participants stay the same.
- Selection versus attrition. Selection is a difference between groups at the start. Attrition is a difference created by who leaves. Randomization fixes the first and does nothing about the second.
- Internal versus external validity. Internal validity asks whether the effect is real here. External validity asks whether it holds elsewhere. Our tutoring study has poor internal validity, so asking whether it generalizes is premature: there is nothing yet to generalize.
Here is the practical drill. Take any before-and-after claim and ask, item by item, whether history, maturation, testing, instrumentation, or regression could produce it on its own. If you can answer yes to any of them and the design has no control group, the causal claim is unsupported no matter how large the change or how small the p-value.
Key idea: Run the five single-group threats against any before-and-after claim; if even one survives, the causal reading fails.
Try it
Exercise. A district notes that students scoring in the bottom 10 percent on a fall test are placed in a remedial program, and that by spring their average score has risen sharply. The district credits the program. Name the most serious threats, and describe the design that would let the program's true effect be estimated.
Model answer. The strongest objection is regression to the mean. Because students were selected for extreme-low fall scores on an imperfectly reliable test, their spring average would rise substantially even with no program, as chance-driven lows partly rebound. Maturation compounds this, since students generally improve over a school year, and testing may add familiarity gains. The one-group before-and-after structure cannot separate the program from any of these artifacts.
A design that isolates the program would randomly assign eligible bottom-scoring students either to the remedial program or to a control condition. Both groups then experience the same year, the same tests, and the same regression, and you compare spring scores across the two randomized groups. If exact randomization is impossible for policy reasons, a regression-discontinuity design can approximate it. It compares students just above and just below the cutoff, who are similar in everything except program placement. Either way, the fix is structural. Supply a counterfactual that undergoes the same artifacts, so that only the program itself differs systematically between the groups.
Recap
- Internal validity is whether the independent variable, rather than something else, produced the observed change.
- Cook and Campbell distinguished statistical conclusion, internal, construct, and external validity, and internal validity is the indispensable one for causal claims.
- History, maturation, testing, instrumentation, and regression to the mean all mimic treatment effects in one-group pre-post designs.
- Regression to the mean is guaranteed whenever participants are chosen for extreme scores on an imperfectly reliable measure.
- Selection threatens non-randomized group comparisons, and random assignment is the only remedy for unmeasured selection differences.
- Attrition, compensatory rivalry, resentful demoralization, compensatory equalization, and local history can undermine even a randomized design after assignment.
- Campbell and Stanley notation using R, O, and X displays a design's threats, and a randomized pretest-posttest control-group design closes most of them.
Sources
- Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Houghton Mifflin. find source ↗
- Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin. find source ↗
- Barnett, A. G., van der Pols, J. C., & Dobson, A. J. (2005). Regression to the mean: What it is and how to deal with it. International Journal of Epidemiology, 34(1), 215-220. pubmed.ncbi.nlm.nih.gov
- McCambridge, J., Witton, J., & Elbourne, D. R. (2014). Systematic review of the Hawthorne effect: New concepts are needed to study research participation effects. Journal of Clinical Epidemiology, 67(3), 267-277. pmc.ncbi.nlm.nih.gov
- Sedgwick, P., & Greenwood, N. (2015). Understanding the Hawthorne effect. BMJ, 351, h4672. pubmed.ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Quasi-experimental research. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Rothman, K. J., & Greenland, S. (2005). Causation and causal inference in epidemiology. American Journal of Public Health, 95(S1), S144-S150. pubmed.ncbi.nlm.nih.gov
- Key terms
- Internal validity
- The degree to which a study shows the treatment, not something else, caused the outcome.
- History (threat)
- An external event between pre- and posttest that could produce the observed change.
- Maturation
- Natural change within participants over time that can mimic a treatment effect.
- Regression to the mean
- Statistical tendency of extreme scores to move toward the average on retest.
- Selection
- Pre-existing differences between non-randomly formed groups that confound comparisons.
- Attrition
- Differential dropout across conditions that undermines initial group comparability.
External Validity and Quasi-Experimental Designs
- Define external validity and its principal threats to generalization.
- Describe key quasi-experimental designs and how each approximates a counterfactual without randomization.
A study can be internally airtight and still tell us little about the world, if its result does not travel. External validity is the extent to which a causal finding generalizes beyond the specific participants, settings, treatments, and times of the study. Randomization is also often impossible for ethical or practical reasons. So this lesson introduces quasi-experimental designs as well. They pursue causal inference without random assignment, by approximating the counterfactual in other ways.
The two topics belong together, because they mark the frontier where clean experiments give way to real-world constraints. In the field you rarely control who receives a treatment or where. You therefore face two problems at once. Does a lab effect generalize? And how do you estimate an effect when you cannot randomize? This lesson equips you to reason about both without overclaiming what a design can support.
Two concrete cases will run through the lesson. First, a 20-minute writing exercise raises exam grades in a randomized study of 200 psychology undergraduates at one selective university. Does it work for adults returning to community college at night? That is the external-validity question, and the trial cannot answer it. Second, one state raises its minimum wage and its neighbor does not. Nobody randomized anything, yet the comparison can still support a causal estimate. That is the quasi-experimental question. Keep both in view: the first study has strong internal and weak external validity, and the second has to earn its internal validity by argument.
Key idea: Internal validity asks whether the effect is real here; external validity asks who else it applies to, and a perfect experiment can be strong on the first and silent on the second.
In plain terms
This lesson answers two plain questions. Question one: your study worked here, but will it work anywhere else? That is external validity. It can fail four ways. The people were unusual. The place was unusual. The treatment was delivered differently. Or the moment in history was unusual. You test it by running the study again with different people in different places. Question two: what do you do when you cannot flip a coin? Some things cannot be randomized. You cannot randomly assign a state to raise its wages. So you build a stand-in for the comparison you are missing. Watch the same unit for many months before and after. Or compare people just above and just below a cutoff. Or subtract the change in a group that got nothing. Each trick rests on one assumption. Say what it is, then show evidence for it.
Threats to external validity
- Interaction of selection and treatment: the treatment works for the studied sample but not for other populations. Such samples are often WEIRD, meaning Western, educated, industrialized, rich, and democratic. A finding on undergraduates may not hold for older, less-educated, or non-Western groups.
- Interaction of setting and treatment: an effect obtained in a controlled laboratory may vanish in a messy field setting, or vice versa.
- Interaction of history and treatment: an effect tied to a particular historical moment may not replicate at another time.
There is often a real tension between internal and external validity. Tightening experimental control, with a sterile lab and an artificial task, strengthens causal inference. The same tightening makes the situation less like the world it is meant to inform. Neither should be sacrificed reflexively. The balance depends on the study's purpose. Generalization is ultimately established through replication across samples and settings, not asserted from a single study.
The selection-treatment threat has a name worth knowing. Much of psychology's evidence rests on WEIRD samples. Henrich and colleagues coined the acronym for Western, Educated, Industrialized, Rich, and Democratic participants. Such people turn out to be unusual on many measures rather than representative of humanity. A finding robust among North American undergraduates may not describe people in other cultures. That is why claims of universality require testing across genuinely different populations, rather than assuming a convenient sample stands in for everyone. Our writing exercise was tested on 200 students at a selective university, which is about as WEIRD as a sample gets.
A related concept is ecological validity. It asks whether the task and setting resemble the real situation the finding is meant to inform. A memory experiment using nonsense syllables under bright laboratory lights has strong control. It may not reflect how people remember a conversation. Ecological validity is not identical to external validity, because an artificial task can still yield a generalizable principle. Still, a large gap between the study setting and the target setting is a reason to test the effect closer to where it will be applied.
Key idea: External validity fails when the sample, the setting, the treatment, or the moment turns out to change how the effect works.
Generalization is a claim about moderators
It helps to see generalization not as a vague hope but as a set of claims about moderators. To say an effect generalizes from sample A to population B is to claim something specific. It claims that the variables on which A and B differ do not moderate the effect. If age, culture, or setting changes the effect, then the effect does not transfer unchanged. The honest move is to identify the boundary conditions rather than assert universality. For our writing exercise, the suspected moderators are obvious once named: age, prior academic confidence, and whether students already have a private place to write.
This reframing turns generalization into an empirical question with a research strategy. Instead of debating whether a single study generalizes, you design or await studies that vary the suspected moderators. Then you map where the effect holds, weakens, or reverses. Generalization is therefore built by replication across deliberately varied samples and settings. That is why a program of studies, not one decisive experiment, establishes how far a finding actually reaches.
Campbell offered a useful reframing of how generalization is justified in practice: the proximal similarity model. Rather than claiming a result applies universally, you argue that it applies to persons, settings, and times closely resembling those studied. Confidence fades as the target grows more dissimilar. Generalization is thus a matter of degree and of argued resemblance, not an all-or-nothing property. Stating the zone of similarity you are willing to defend is more honest than a blanket claim.
Generalization can also fail at the level of the treatment itself. "The same" intervention delivered by different staff, at a different intensity, or in a different context may not be the same treatment in any meaningful sense. An effect can therefore evaporate not because the population differs but because the treatment was not faithfully reproduced. This is why implementation fidelity is reported alongside outcomes. It is also why a failed replication sometimes reflects a different treatment rather than a fragile effect.
Key idea: To say a finding generalizes is to bet that the differences between your sample and the new one do not change the effect - so name those differences and test them.
Quasi-experimental designs
When you cannot assign at random, you can still design to rule out alternatives.
- Nonequivalent groups design: compare a treatment group and a comparison group that were not randomly formed, say two schools. Use pretests where possible, so you can gauge and adjust for initial differences. The central worry is selection, and the pretest plus careful matching are the defenses.
- Interrupted time-series design: take many measurements before and after an intervention on the same unit. A change in the level or the slope of the series at the intervention point, against the established trend, supports a causal reading. It also helps rule out simple maturation.
- Regression-discontinuity design: use it when treatment is assigned by a cutoff on a continuous measure, such as a scholarship for everyone scoring above a threshold. Compare units just above and just below the cutoff, which are nearly equivalent by chance. Under its assumptions this design approaches the credibility of a randomized experiment near the threshold.
- Difference-in-differences: compare the before-after change in a treated group to the before-after change in an untreated comparison group. This differences out stable group differences and common time trends, under a parallel-trends assumption.
Two further tools extend the quasi-experimental kit. Propensity-score methods estimate each unit's probability of receiving the treatment from measured covariates. They then match or weight to build comparison groups balanced on those covariates, but only on the ones actually measured. Instrumental-variable methods exploit a variable that affects treatment and influences the outcome only through treatment, using it to isolate causal variation. Both can strengthen a causal claim. Both also live or die by assumptions that, unlike randomization, cannot be guaranteed by design and must be argued.
The quasi-experimental designs are not equal in credibility. A well-implemented regression-discontinuity design can rival a randomized trial for units near the threshold, because assignment there is nearly random. Difference-in-differences and interrupted time-series are strong when their trend assumptions hold. A simple nonequivalent-groups comparison is the weakest and the most exposed to selection. A skilled designer strengthens the weaker forms by adding comparison groups, more pre-intervention observations, or a second untreated series. Each addition closes another alternative explanation.
Key idea: Every quasi-experimental design replaces randomization with an assumption, so the design is only as strong as the argument for that assumption.
A worked quasi-experiment
Return to the minimum wage. One state raises it and a neighboring state does not. A naive before-after comparison in the treated state alone would confound the policy with any economy-wide trend. A difference-in-differences design compares the change in employment in the treated state to the change in the untreated neighbor over the same period. Put numbers on it. Employment falls 2 points in the treated state and 1.5 points in the neighbor. The estimated policy effect is 0.5 points, not 2. Subtracting the neighbor's change removes whatever common trend hit both, and it isolates the policy's incremental effect - on one condition. Absent the policy, the two states must have been on course to move in parallel.
The design's credibility rests entirely on that parallel-trends assumption. That is why analysts probe it, plotting the two series for several periods beforehand to see whether they tracked each other. If the treated state was already diverging before the policy, the assumption fails and the estimate is suspect. This is the general pattern of quasi-experimental work. Name the identifying assumption, then marshal evidence for and against it. The assumption is doing the causal work that randomization would otherwise have done for you.
Key idea: Difference-in-differences subtracts the comparison group's change to remove the trend they share, and the whole estimate depends on the two groups having been headed the same way.
The unifying idea
Every quasi-experimental design tries to construct a credible counterfactual without the randomizer's guarantee. Each rests on assumptions that must be argued, probed, and where possible tested. Those assumptions include no differential selection trend, parallel trends, and continuity at the cutoff. Quasi-experiments are not second-rate experiments to apologize for. They are the responsible way to pursue causal questions where randomization is unethical or infeasible, provided their identifying assumptions are made explicit and defended with the rigor an experimenter brings to randomization.
Key idea: A good quasi-experiment states its identifying assumption out loud and then supplies evidence for it.
Replication and its varieties
No single study settles generalizability, so replication is the mechanism by which findings earn their reach. A direct (close) replication repeats a study's procedures as faithfully as possible, testing whether the original result recurs. It isolates whether the effect is real. A conceptual replication tests the same hypothesis with different operations, populations, or settings. It probes whether the underlying relationship holds, not just the specific procedure.
The two play complementary roles. Direct replications guard against false positives and questionable practices. Conceptual replications, when they succeed across varied conditions, map the breadth of an effect and identify its moderators. Repeating the writing exercise with the same students and materials is a direct replication. Running it with night-school adults using a different prompt is a conceptual one, and only the second speaks to whether the finding travels. A result supported by both is far more trustworthy than one resting on a single elegant study. That is the disciplinary logic behind the open-science reforms taken up in the final lesson. On this view a single striking result is a hypothesis about the world, and it becomes knowledge only after others reproduce it.
Key idea: Direct replication tests whether an effect is real; conceptual replication tests how far it reaches.
Where people get stuck
Two confusions do most of the damage here, and the first is among the most common in the whole course.
- Internal versus external validity. Internal validity is about the inside of the study: did the treatment cause the change in these participants? External validity is about the outside: does the finding hold for other people, places, or times? The writing-exercise trial has excellent internal validity and unknown external validity. A large national survey of study habits has the reverse. Neither shortcoming is a mistake; each is a limit to state plainly.
- Randomization is not representativeness. Random assignment makes groups comparable within your sample. It says nothing about whether your sample resembles anyone else. A trial can be perfectly randomized and perfectly unrepresentative at the same time, and 200 students at one selective university is exactly that case.
A third trap concerns the quasi-experimental designs. They are not ranked by complexity. A regression-discontinuity design with a sharp cutoff and dense data near the threshold is more credible than a propensity-score analysis with dozens of covariates, because the first rests on an assumption you can inspect and the second rests on having measured everything that matters. When you read a quasi-experimental paper, find the sentence that states the identifying assumption. If there is no such sentence, that absence is your critique.
Key idea: Randomizing who gets the treatment does not make your sample look like the population - internal and external validity are bought separately.
Try it
Exercise. A hospital introduces a new hand-hygiene protocol on one ward and, comparing infection rates the month before and after, reports a large drop. A skeptic notes that infection rates fall every winter-to-spring anyway. Propose a quasi-experimental design that addresses the skeptic, and state its key identifying assumption.
Model answer. The skeptic is describing a secular trend that a simple before-after comparison on one ward cannot rule out, a version of history or maturation at the unit level. A stronger design adds an untreated comparison ward in the same hospital and uses difference-in-differences: compare the before-after change in infections on the treated ward to the before-after change on a similar ward that kept the old protocol over the same months.
Subtracting the comparison ward's change removes the seasonal trend that affected both wards, isolating the protocol's incremental effect. Alternatively, an interrupted time-series design would collect many months of infection data before and after the change on the treated ward and test for a shift in level or slope at the intervention point, against the established trend, which also distinguishes a real effect from a smooth seasonal decline.
The key identifying assumption for difference-in-differences is parallel trends. Without the new protocol, the two wards' infection rates would have changed by the same amount. The analyst should support that assumption by showing the wards tracked each other in the months before the intervention. The analyst should also be candid that any pre-existing divergence would undermine the causal reading.
Recap
- External validity is whether a causal finding travels across people, settings, treatments, and times, and it is separate from internal validity.
- Threats arise from interactions of treatment with selection, setting, and history, and WEIRD samples are the standard example of the first.
- Generalization is a claim that the differences between your sample and the target do not moderate the effect, which makes it testable rather than rhetorical.
- Campbell's proximal similarity model treats generalization as argued resemblance that fades with distance from the studied case.
- Quasi-experiments approximate a counterfactual without randomization through nonequivalent groups, interrupted time series, regression discontinuity, and difference-in-differences.
- Propensity scores and instrumental variables extend the toolkit, and both depend on assumptions that must be argued rather than guaranteed by design.
- Difference-in-differences relies on parallel trends, regression discontinuity on continuity at the cutoff, and both assumptions can and should be probed.
- Direct replication tests whether an effect is real; conceptual replication maps how far it extends and which variables moderate it.
Sources
- Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin. find source ↗
- Bernal, J. L., Cummins, S., & Gasparrini, A. (2017). Interrupted time series regression for the evaluation of public health interventions: A tutorial. International Journal of Epidemiology, 46(1), 348-355. pmc.ncbi.nlm.nih.gov
- Austin, P. C. (2011). An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate Behavioral Research, 46(3), 399-424. pmc.ncbi.nlm.nih.gov
- Geldsetzer, P., & Fink, G. (2017). Quasi-experimental study designs series - paper 2: Complementary approaches to advancing global health knowledge. Journal of Clinical Epidemiology, 89, 12-16. pubmed.ncbi.nlm.nih.gov
- Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science & Medicine, 210, 2-21. pmc.ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Quasi-experimental research. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Camerer, C. F., Dreber, A., Holzmeister, F., Ho, T.-H., Huber, J., Johannesson, M., ... Wu, H. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2(9), 637-644. pubmed.ncbi.nlm.nih.gov
- Key terms
- External validity
- The extent to which a finding generalizes across people, settings, treatments, and times.
- WEIRD samples
- Western, educated, industrialized, rich, democratic samples that may not represent humanity broadly.
- Quasi-experiment
- A design pursuing causal inference without random assignment by approximating a counterfactual.
- Nonequivalent groups design
- Comparing non-randomly formed treatment and comparison groups, often with pretests.
- Interrupted time-series
- Repeated measures before and after an intervention to detect a change in level or trend.
- Regression-discontinuity
- Exploiting a cutoff on a continuous assignment variable to compare near-equivalent units.
Module 7: Survey and Qualitative Methods
Two major traditions of primary data collection: structured self-report and the interpretive study of meaning.
Survey Research and Questionnaire Design
- Identify sources of survey error within the total survey error framework.
- Recognize and repair common questionnaire flaws that bias responses.
Surveys are the workhorse of social measurement, but their apparent simplicity is deceptive: small wording choices can swing results by double-digit margins. The total survey error framework organizes everything that can go wrong into errors of representation (who answers) and errors of measurement (what their answers mean), and good design attacks both.
The framework is valuable because it prevents lopsided effort. Picture two researchers. One obsesses over sample size and ignores question wording. The other crafts elegant items and tolerates a 12% response rate. Each has optimized one half of the problem and left the other wide open. Every design decision, from the frame to the final item, can be located as an attempt to reduce one of these error sources. A defensible survey shows deliberate attention to each.
Here is a real demonstration of how much wording matters. Ask people whether the government should "help the poor" and support runs high. Ask whether it should spend more on "welfare" and support falls sharply. Same policy, same respondents, two very different numbers. Now consider an item a doctoral student actually wrote: "How often do you experience work-related stress?" with options Never, Rarely, Sometimes, Often. Two nurses working identical shifts can answer Sometimes and Often, because "often" has no fixed meaning and neither does "stress." Nothing about a larger sample repairs that. This lesson is about the half of survey error that lives inside the question.
Key idea: Total survey error has two halves - who answers, and what their answers mean - and a strong survey attacks both.
In plain terms
Two things can go wrong with a survey. The wrong people answer, or the right people misread your questions. Both matter, and researchers usually work hard on one and neglect the other. On the first half: your list may miss people, your sample adds noise, and the people who reply may differ from the people who do not. On the second half, here are the classic faults. Asking two things in one question. Hinting at the answer you want. Using loaded words. Using vague words like "often." And writing statements people will agree with out of politeness. There is a simple fix for the vague-word problem. Ask for something countable. "On how many of the last seven days" beats "often" every time. One last rule. Test your questions on a few real people first and ask them to think aloud. You will be surprised what they thought you meant.
Errors of representation
These concern the gap between the population and the people whose answers you obtain. There are three. Coverage error means the frame misses part of the population. Sampling error is random variation from observing a subset. Nonresponse error means those who answer differ from those who do not. Module 5 treated these in detail. The point here is narrower: a beautifully worded questionnaire cannot rescue a study whose respondents are unrepresentative.
Mode of administration interacts with representation in ways worth planning for. A telephone survey misses people without listed numbers and reaches the willing-to-answer. A web survey excludes those without internet access or skills. A mail survey depends on literacy and motivation. An in-person survey is costly but achieves the best coverage of hard-to-reach groups. Each mode reaches a different slice of the population and evokes different answering behavior, so the choice of mode is a representation decision as much as a budget one.
Key idea: Coverage, sampling, and nonresponse error all concern who answers, and no amount of careful wording repairs them.
Raising and defending response rates
Because nonresponse threatens representation, survey methodologists have developed systematic ways to reduce it. Dillman's tailored-design method treats the questionnaire as an exchange, using multiple contacts, a respectful and clear layout, and modest incentives to raise cooperation. Repeated, varied reminders typically help more than any single feature, and a small prepaid incentive often outperforms a larger promised one, because it signals trust rather than payment.
Yet a high response rate is not the whole story, as an earlier lesson stressed. What matters is whether respondents differ from non-respondents on the survey's variables. A careful report therefore compares the achieved sample to known population benchmarks. Where they differ, it applies weights and discusses the residual risk. Chasing the rate for its own sake, while ignoring who is systematically missing, optimizes a number rather than the validity that number is supposed to protect.
Key idea: Multiple respectful contacts and small prepaid incentives raise response, but the number that matters is whether non-respondents differ from respondents.
Errors of measurement
Even with the right people, the instrument can distort. Watch for these classic questionnaire defects:
- Double-barreled questions ask two things at once ("Do you support increased funding for schools and hospitals?"), so a single answer cannot be interpreted. Split them.
- Leading questions embed a preferred answer ("Don't you agree that the mayor has failed?"), pushing responses toward it.
- Loaded language uses emotionally charged terms that shift responses regardless of the substantive content.
- Ambiguous or vague terms such as "often" and "regularly" mean different things to different respondents, which injects noise. Replace them with countable quantities: "on how many of the last 7 days" beats "often" every time.
- Response-set biases. Acquiescence is the tendency to agree with statements regardless of content, countered by balancing positively and negatively worded items. Social desirability bias is the tendency to answer in a way that looks good. It is strongest for sensitive topics and is mitigated by anonymity, neutral wording, and self-administration.
Key idea: Split double barrels, strip loaded words, replace vague frequencies with countable ones, and balance item polarity so agreement is not a default.
The four steps of answering a question
Survey methodology models what a respondent does with each item as four cognitive steps, following Tourangeau and colleagues. Comprehension works out the question's meaning. Retrieval pulls relevant information from memory. Judgment integrates what was retrieved. Response maps that judgment onto the options offered. An error can enter at any step, and diagnosing which step failed tells you how to fix the item. Our stress item fails at step one and step four: "stress" is undefined, and "often" has no shared meaning.
Retrieval errors are especially common with behavioral questions. Asked how many times they exercised last month, respondents estimate rather than count, and telescoping pulls distant events into the reference period, inflating reports. Shorter reference periods, memory cues, and anchoring to landmark events reduce these errors. The lesson is that a factual question is not automatically accurate; its accuracy depends on whether an ordinary memory can actually supply the answer it demands.
A unifying concept for many measurement errors is satisficing, Krosnick's term for respondents who, instead of doing the full cognitive work, take a shortcut to a merely acceptable answer. Weak satisficing shows up as choosing the first plausible option or the midpoint; strong satisficing shows up as agreeing regardless of content, or answering questions the respondent knows nothing about. Satisficing worsens with respondent fatigue, low motivation, and difficult items, which is one more reason to keep surveys short and questions easy.
Related is the problem of non-attitudes. Classic studies found that sizeable fractions of respondents will offer opinions on entirely fictitious issues rather than admit they have none. The survey situation itself pressures them to answer. That is why an explicit "no opinion" option and neutral framing matter. Without them, a survey can manufacture attitudes that do not exist and then report them as public opinion. Telling a real distribution of views apart from an artifact of the question is part of reading survey data honestly.
Key idea: Answering a question takes real cognitive work, and when the work is too hard respondents take shortcuts that look like data.
Order and format effects
Question order matters. An earlier item can prime a frame that colors later answers. Ask about a specific worry before asking about overall happiness and the happiness rating drops. For closed items, three further choices shape the distribution of responses: whether to offer an explicit "don't know," whether to include a neutral midpoint, and how many scale points to use. None of these choices is neutral, so each should be deliberate.
Two order phenomena recur. In an assimilation effect a prior item pulls later answers toward it, while in a contrast effect it pushes them away, and which occurs depends on how respondents frame the relationship between the items. Primacy and recency effects mean that in long lists, options seen first on paper or heard last aloud are more likely to be chosen, which is why response-option order is sometimes randomized across respondents to average the bias out.
Scale design carries its own choices. The number of points shapes the distribution, and so does whether the scale is fully labeled or labeled only at the endpoints, and whether a midpoint is offered. A midpoint lets genuinely ambivalent respondents say so, and it also invites satisficing. An explicit "don't know" reduces guessing, and it can also discard real if uncertain opinions. None of these is universally correct. Choose each for the construct and defend it, rather than adopting it by habit.
Key idea: Order, scale points, midpoints, and "don't know" options all move the numbers, so treat them as design decisions rather than formatting.
Assembling the whole instrument
Item quality is necessary but not sufficient; the questionnaire as a whole has an architecture. A common structure funnels from easy, engaging warm-up items, through a substantive core grouped by topic to minimize jarring shifts, to sensitive and demographic items near the end, when rapport is highest and a reluctant respondent has already invested effort. Abrupt topic changes and early sensitive questions raise break-off rates, so the sequence is a design variable, not an afterthought.
Length is the quiet enemy. Every added item raises the burden, and burden provokes satisficing. A disciplined designer justifies each question by the analysis it will feed and cuts the ones that will never be reported. The test to apply to a draft instrument is simple. For each item, name the result it will produce and the decision or claim that result supports. Items that fail the test are padding, and they degrade the answers to the questions that actually matter.
Key idea: Order the instrument from easy to sensitive, and cut any item whose eventual result you cannot name.
Asking sensitive questions
Sensitive topics, from drug use to income to prejudice, invite social desirability bias, the pull toward answers that make the respondent look good. Several techniques reduce it. Self-administered modes, which remove a human interviewer, lower the pressure to impress. Neutral, non-judgmental wording that normalizes a range of behaviors makes an honest answer easier. And for the most sensitive items, specialized methods such as the randomized-response technique or list experiments let respondents convey information without directly revealing their own answer, trading some precision for candor.
The guiding principle is simple. Make an honest answer easier to give than a flattering one. That is a matter of design, not of urging respondents to be truthful. Every choice, from mode to wording to placement, either lowers or raises the cost of candor. On sensitive topics that cost often decides whether your data reflect behavior or self-presentation.
Key idea: On sensitive topics, design lowers the cost of telling the truth - asking people to be honest does not.
Open versus closed and the value of piloting
Closed-ended items are easy to analyze and compare but constrain answers to the researcher's categories, risking missing what matters to respondents. Open-ended items capture unanticipated content but demand coding. Because respondents interpret items in ways designers rarely foresee, the single most valuable safeguard is a pilot test - a small-scale trial run, ideally with cognitive interviewing in which respondents think aloud as they answer, revealing where wording is misread. A survey deployed without piloting is an instrument of unknown validity, and its polished appearance offers no assurance that respondents understood the questions as intended.
Cognitive interviewing repays a closer look, because it catches problems no amount of armchair editing will. Watch a handful of respondents narrate their thinking and you will find the term read two ways, the reference period that felt ambiguous, the option nobody saw themselves in. These are failures at the comprehension and response steps. They are invisible in the final data, because a respondent who misreads an item still supplies an answer that looks perfectly valid. A short round of piloting is therefore not a courtesy. It is the cheapest validity evidence a survey can buy.
Key idea: A misread question still produces a tidy answer, so piloting with think-aloud interviews is the only way to see the misreading.
Where people get stuck
Three distinctions repay attention.
- Representation versus measurement error. Nonresponse is a representation problem, about who is in the data. A double-barreled item is a measurement problem, about what the answers mean. Weighting fixes some of the first and none of the second.
- Acquiescence versus social desirability. Acquiescence is agreeing with whatever is put in front of you, and the cure is mixing item polarity. Social desirability is shading answers to look better, and the cure is anonymity, neutral wording, and self-administration. They are different biases with different repairs.
- Reliability of a scale versus validity of an item. A five-item stress scale can report alpha of 0.88 while every item uses the word "often." Consistency among vague items is consistency, not accuracy, which is the reliability-validity distinction from Module 4 reappearing in survey form.
A useful habit when critiquing any item: read it aloud and ask what two different reasonable people could take it to mean. If you can construct two readings, so can your respondents, and their answers will be a mixture of both.
Key idea: If an item can be read two ways, your data will contain both readings mixed together and you will never be able to separate them.
Try it
Exercise. Critique this item and rewrite it: "How often do you use our excellent new app to conveniently manage your busy finances, and would you recommend it to friends and family? (Always / Often / Sometimes / Never)." Name each defect and produce a clean version.
Model answer. The item has several defects. It is double-barreled, asking both how often the person uses the app and whether they would recommend it, so one answer cannot serve both. "Excellent" and "conveniently" are loaded, leading language that pushes toward a positive view. The response options mix a frequency scale with the recommendation question they do not fit, and "often" and "sometimes" are vague terms respondents will read differently.
A clean version splits the item and neutralizes the wording. First, a frequency question with a concrete reference period and defined options: "In the past 7 days, on how many days did you use the app? (0, 1 to 2, 3 to 4, 5 to 6, 7)." Second, a separate recommendation item on a labeled scale: "How likely are you to recommend the app to others? (Very unlikely / Unlikely / Neither / Likely / Very likely)."
The rewrite removes the double barrel, strips the persuasive adjectives, anchors frequency to a countable period, and gives each question a scale that fits it. A pilot with think-aloud interviews would then confirm that respondents read the reference period and the options as intended. That closes the loop between careful drafting and demonstrated comprehension.
Recap
- Total survey error splits into errors of representation, about who answers, and errors of measurement, about what their answers mean.
- Mode of administration is a representation decision, because telephone, web, mail, and in-person surveys reach different slices of a population.
- Response rates can be raised by multiple contacts and prepaid incentives, but the diagnostic question is whether respondents differ from non-respondents.
- The classic item defects are double-barreled, leading, loaded, and vague wording, plus acquiescence and social desirability response sets.
- Answering a question requires comprehension, retrieval, judgment, and response, and satisficing is the shortcut respondents take when the work is too hard.
- Order effects, scale length, midpoints, and don't-know options all shift the distribution of answers and must be chosen deliberately.
- Sensitive questions need designs that lower the cost of candor, including self-administration, neutral wording, and indirect techniques.
- Piloting with cognitive interviewing is the cheapest validity evidence available, because misread items still yield plausible-looking data.
Sources
- Boynton, P. M., & Greenhalgh, T. (2004). Selecting, designing, and developing your questionnaire. BMJ, 328(7451), 1312-1315. pmc.ncbi.nlm.nih.gov
- Krosnick, J. A. (1999). Survey research. Annual Review of Psychology, 50, 537-567. pubmed.ncbi.nlm.nih.gov
- Sullivan, G. M., & Artino, A. R. (2013). Analyzing and interpreting data from Likert-type scales. Journal of Graduate Medical Education, 5(4), 541-542. pmc.ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Constructing survey questionnaires. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Conducting surveys. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Substance Abuse and Mental Health Services Administration. (2019). Reliability of key measures in the National Survey on Drug Use and Health. NCBI Bookshelf. ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Survey research. In Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Key terms
- Total survey error
- A framework partitioning survey error into representation and measurement components.
- Double-barreled question
- A single item asking about two issues at once, so answers cannot be interpreted.
- Leading question
- An item worded to push respondents toward a particular answer.
- Acquiescence bias
- The tendency to agree with statements regardless of their content.
- Social desirability bias
- The tendency to answer in a way that presents the respondent favorably.
- Cognitive interviewing
- A pretest method in which respondents think aloud to reveal how they interpret items.
Qualitative Inquiry: Interviews, Ethnography, and Traditions
- Compare major qualitative traditions by their aim and product.
- Explain core qualitative data-collection methods and the researcher-as-instrument stance.
Qualitative research seeks to understand meaning, process, and context in depth rather than to quantify variation across many cases. Its logic, established in Module 1, is interpretivist: the researcher is the primary instrument, and analysis is typically inductive. This lesson surveys the main traditions and data-collection methods so you can match approach to question.
A frequent doctoral error is to call a study simply "qualitative" as though that named a method. It does not. Qualitative inquiry is a family of distinct traditions. Each has its own question, its own data conventions, and its own criteria for a good result. Choosing among them deliberately is what separates a designed study from a loose collection of interviews. The traditions below are the ones you are most likely to adopt or encounter.
Take one topic and watch it split five ways. Suppose you want to study homeless shelter residents. Ask "what is it like to sleep in a shelter?" and you are doing phenomenology. Ask "how do residents come to trust or distrust staff, and what theory explains that process?" and you are doing grounded theory. Ask "how does this shelter work as a community, with its own rules and hierarchy?" and you are doing ethnography. Ask "what can we learn from this one shelter that switched to a low-barrier policy?" and you have a case study. Ask "how do residents tell the story of how they arrived here?" and you have narrative inquiry. Same building, same people, five different studies, five different products. Deciding which one you are doing is the first design choice, not an afterthought.
Key idea: Qualitative is a family of traditions, not a method, and the tradition you choose fixes your data, your analysis, and what counts as a good finding.
In plain terms
Five traditions, five plain questions. What is it like? That is phenomenology. Why does this happen, and what theory explains it? That is grounded theory. How does this group live? That is ethnography. What can we learn from this one place? That is a case study. How do people tell their story? That is narrative inquiry. Pick one and stick to its rules. Then choose how you gather data. Talk to people one at a time. Sit and watch. Put people in a room together. Or read what they wrote. Most good studies use two or three of these. Choose whom to study on purpose, not at random. Keep going until new interviews stop teaching you anything new. And write down who you are and how that shaped what you heard. You are the instrument here.
Five traditions
| Tradition | Central question | Typical product |
|---|---|---|
| Phenomenology | What is the lived experience of a phenomenon? | The essence of an experience |
| Grounded theory | What theory is grounded in these data? | A theory built from data |
| Ethnography | How does this culture-sharing group work? | A cultural portrait |
| Case study | What can be learned from this bounded case? | An in-depth case analysis |
| Narrative inquiry | How do people story their experience? | A restoried account |
These are not interchangeable. Grounded theory aims to build theory from data through systematic coding and constant comparison. Phenomenology aims to distill the essence of a lived experience. Ethnography studies a culture-sharing group in its natural setting, classically through extended immersion. Choosing a tradition commits you to its data-collection and analytic conventions.
Phenomenology studies the essence of a lived experience. It asks what it is like to undergo something, such as receiving a serious diagnosis. It descends from two lines. Husserl's descriptive phenomenology asks the researcher to bracket, or set aside, prior assumptions through the epoche, so as to see the experience freshly. Heidegger's interpretive or hermeneutic phenomenology holds instead that the researcher's situated understanding is unavoidable and even productive. The product is a structured description of the phenomenon's essential features across participants.
Grounded theory, introduced by Glaser and Strauss in 1967, builds theory upward from data instead of testing a theory from above. It runs on two engines. Constant comparison compares each new datum with the codes already built. Theoretical sampling lets the emerging analysis dictate who or what to study next. Later strands diverge. The Straussian version offers structured coding procedures. Charmaz's constructivist version treats the resulting theory as co-constructed rather than simply discovered in the data.
Ethnography studies a culture-sharing group in its own setting through extended immersion, seeking to render its practices intelligible. Fieldworkers distinguish the emic view, the meanings insiders hold, from the etic view, the analyst's external framework, and aim for the thick description that Geertz argued separates a wink from a mere blink by supplying the layers of meaning behind an act. The product is a cultural portrait grounded in prolonged participation.
Case study investigates a bounded system, a program, event, or organization, in depth and within its real context. Methodologists differ in emphasis: Yin frames case study as a strategy for tracing how and why within a case using multiple sources of evidence, while Stake distinguishes an intrinsic case, studied for its own sake, from an instrumental case, studied to illuminate a broader issue. A multiple-case design compares several cases to strengthen a claim through replication logic.
Narrative inquiry takes the stories people tell as the unit of analysis. It examines how individuals make sense of their lives through the plots, characters, and turning points they construct. Following Clandinin and Connelly, the researcher often restories a participant's account into a coherent chronology while preserving its meaning. Attention goes to time, place, and the interaction between personal and social forces. The product is an interpreted account that treats the story itself, not only its content, as data.
Key idea: Phenomenology yields an essence, grounded theory a theory, ethnography a cultural portrait, case study an analysis of one bounded system, and narrative inquiry a restoried account.
Other approaches you will meet
Several further approaches appear across the social sciences. Action research and participatory designs pursue change alongside understanding, involving stakeholders as collaborators rather than subjects. Discourse and conversation analysis treat language use itself as the object, examining how talk and text construct social reality. Qualitative description offers a lower-inference option for studies that need a rich summary of participants' views without committing to a full theoretical tradition.
Two cautions accompany this variety. First, mixing traditions carelessly produces a muddle that reviewers call method slurring. Collecting phenomenological interviews and then analyzing them with grounded-theory coding is the classic instance, and it needs either avoiding or explicit justification. Second, the choice among approaches is a matter of fit rather than taste, because each carries assumptions about what counts as data and as a good finding. Naming the tradition and following its conventions is what makes a qualitative study legible to others in the field.
Key idea: Do not borrow one tradition's data collection and another's analysis without saying why - reviewers call that method slurring.
Data-collection methods
- Interviews range from structured (fixed questions, near-survey) through semi-structured (a guide of topics with freedom to probe, the qualitative workhorse) to unstructured (open, conversational). Good interviewing centers on open questions, active listening, and probing without leading.
- Participant observation, the backbone of ethnography, places the researcher in the setting along a continuum from complete observer to complete participant, recorded in detailed field notes.
- Focus groups use group interaction itself as data, surfacing shared understandings and points of contention that individual interviews might miss.
- Documents and artifacts - records, images, material objects - provide unobtrusive evidence not shaped by the researcher's presence.
Interview quality is a learned craft, not a matter of following a script. Strong questions are open and non-leading, and the interviewer's main skills are active listening and the well-timed probe, a neutral prompt such as "tell me more about that" that deepens an answer without steering it. Silence is itself a tool, since a pause often draws out more than a follow-up question would, and the practiced interviewer resists the urge to fill it too quickly.
Practical craft surrounds the interview itself. An interview guide lists topics and a few core questions while leaving room to follow the participant, and it is piloted like any instrument. Interviews are typically audio-recorded with consent and transcribed verbatim, because working from memory or notes loses the exact wording that later analysis depends on. Building rapport early, and starting with easier questions before sensitive ones, raises the quality and candor of what follows.
Key idea: Semi-structured interviews, participant observation, focus groups, and documents are the four workhorses, and each gives you a different kind of access.
Most strong qualitative studies combine several of these methods rather than relying on one. An ethnography pairs observation with interviews and documents; a case study deliberately draws on multiple sources of evidence about the same bounded system. Using more than one window onto the phenomenon lets the analyst check whether accounts converge, a form of triangulation the next lesson develops as a criterion of rigor.
Sampling and saturation
Qualitative sampling is almost always purposive. Cases are chosen because they are information-rich for the question, not to represent a population statistically. A common stopping rule is data saturation: the point at which additional data yield no new codes, themes, or insight. Saturation is a judgment about informational redundancy rather than a fixed number. It still needs to be reasoned about and reported rather than left implicit.
Saturation deserves scrutiny, because it is more often asserted than shown. Empirical work suggests that for a fairly homogeneous group and a focused question, most themes emerge within roughly the first dozen interviews. Heterogeneous samples and broad questions require more. Methodologists also distinguish two kinds. Code saturation arrives when no new codes appear. Meaning saturation arrives later, when the understanding of each code stops deepening. A credible report says how saturation was judged, not merely that it was reached. In our shelter study, that means something concrete: interviews 13 through 16 produced no new codes, and interviews 17 and 18 added only detail to the trust code that was already established.
The scale of a qualitative study reflects its logic. A survey seeks many shallow observations. A qualitative study seeks few deep ones. So a design with eight richly analyzed interviews can be more defensible than one with eighty thin ones. Depth is the currency: pages of transcript per participant, hours in the field, iterations of analysis. Head count is not the currency. Reviewers judge a qualitative sample by whether it is rich and appropriate for the question, not by whether it is large.
Key idea: Saturation is a claim about conceptual yield that you must demonstrate, not a number you assert - and it is not a power analysis.
The researcher as instrument
Because a human being, not a standardized scale, collects and interprets the data, subjectivity is not a contaminant to be eliminated but a resource to be managed through reflexivity. The qualitative researcher documents their assumptions, keeps an audit trail of decisions, and sometimes brackets (in phenomenology) prior beliefs to see the phenomenon afresh. This is the disciplined counterpart to the experimenter's controls: different machinery for a different kind of knowledge, held to standards appropriate to its own tradition rather than borrowed wholesale from the quantitative one.
Positionality is the specific form this takes. A researcher's insider or outsider status, identity, and relationship to participants shape both what is disclosed and what is noticed, so many qualitative reports include a positionality statement that makes these influences visible. A reflexive journal kept throughout the study records evolving interpretations and the reasons for analytic choices, which both disciplines the analysis and supplies the audit trail that later underwrites claims of trustworthiness.
Reflexivity is not license for the researcher's views to overwrite the participants'. The discipline works in the other direction. Making the researcher's lens explicit lets both the analyst and the reader watch for places where a prior belief may have shaped an interpretation. Some studies add member reflections, returning tentative findings to participants and asking whether the account rings true. That keeps the interpretation anchored to the experiences it claims to represent instead of floating free of them.
Key idea: In qualitative work the researcher is the instrument, so disclosing and examining your own position is the equivalent of calibrating a machine.
Where people get stuck
Three confusions are worth heading off.
- Saturation versus sample size. Saturation is a stopping rule about conceptual yield. A sample size calculation is an arithmetic requirement for a target precision. Writing "saturation was reached at n = 15" without saying what stopped appearing is asserting a conclusion, not reporting evidence.
- Method versus methodology, again. "I conducted semi-structured interviews" names a method, and three different traditions could say the same sentence. The tradition determines whether those transcripts become an essence, a theory, or a cultural portrait.
- Small sample versus weak sample. Eight participants is not a limitation to apologize for if each was chosen because they were information-rich and each interview ran ninety minutes. Eighty thin interviews with whoever answered a flyer is the weaker design, even though it looks bigger.
One further caution about the phrase "researcher bias." In a quantitative design, researcher influence is error to be minimized. In an interpretive design it is a condition of knowledge to be documented and examined. Both stances are rigorous. What is not rigorous is claiming the interpretive stance while providing no positionality statement, no audit trail, and no evidence that disconfirming cases were sought.
Key idea: A small purposive sample analyzed deeply is a design choice; a large careless one is not a bigger version of it.
Try it
Exercise. For each question, name the qualitative tradition that best fits and justify the match: (a) "What is it like to live with chronic pain?" (b) "How do the staff of one hospice negotiate end-of-life decisions over a year of fieldwork?" (c) "What process do new immigrants go through in forming a sense of belonging, and what theory explains it?"
Model answer. Question (a) fits phenomenology, because it asks about the essence of a lived experience, the felt quality of living with chronic pain, which phenomenological description is designed to capture. Question (b) fits ethnography or a case study: the year of immersion in a single culture-sharing setting, studying how staff actually work, is ethnographic, and because the hospice is a bounded system it could equally be framed as a case study depending on whether the emphasis is cultural understanding or the bounded case itself.
Question (c) fits grounded theory. It explicitly seeks a theory of a process built from the data, which is grounded theory's defining aim. The phrase "what theory explains it" signals constant-comparison, theory-building logic rather than a description of essence or culture. The general principle here is worth keeping. The verb and object of the question, whether experience, culture, process, or bounded case, point to the tradition. Naming that tradition then commits you to its data-collection and analytic conventions.
Recap
- Qualitative research is a family of traditions rather than a single method, and the tradition determines the question, the data, and the product.
- Phenomenology seeks an essence, grounded theory a theory of process, ethnography a cultural portrait, case study an in-depth bounded analysis, and narrative inquiry a restoried account.
- Mixing traditions without justification is method slurring, and reviewers treat it as a design error.
- The main data-collection methods are semi-structured interviews, participant observation, focus groups, and documents, and strong studies combine several.
- Sampling is purposive because information richness rather than representativeness is the goal.
- Saturation is a demonstrated judgment about conceptual yield, with code saturation arriving before meaning saturation, and it is not a sample-size calculation.
- The researcher is the instrument, so positionality statements, reflexive journals, and audit trails are the qualitative equivalent of calibration.
Sources
- Korstjens, I., & Moser, A. (2017). Series: Practical guidance to qualitative research. Part 2: Context, research questions and designs. European Journal of General Practice, 23(1), 274-279. pmc.ncbi.nlm.nih.gov
- Moser, A., & Korstjens, I. (2018). Series: Practical guidance to qualitative research. Part 3: Sampling, data collection and analysis. European Journal of General Practice, 24(1), 9-18. pmc.ncbi.nlm.nih.gov
- Saunders, B., Sim, J., Kingstone, T., Baker, S., Waterfield, J., Bartlam, B., ... Jinks, C. (2018). Saturation in qualitative research: Exploring its conceptualization and operationalization. Quality & Quantity, 52(4), 1893-1907. pmc.ncbi.nlm.nih.gov
- Vasileiou, K., Barnett, J., Thorpe, S., & Young, T. (2018). Characterising and justifying sample size sufficiency in interview-based studies. BMC Medical Research Methodology, 18, 148. pmc.ncbi.nlm.nih.gov
- Reeves, S., Kuper, A., & Hodges, B. D. (2008). Qualitative research methodologies: Ethnography. BMJ, 337, a1020. pubmed.ncbi.nlm.nih.gov
- Tenny, S., Brannan, J. M., & Brannan, G. D. (2022). Qualitative study. In StatPearls. NCBI Bookshelf. ncbi.nlm.nih.gov
- Price, P. C., Jhangiani, R., Chiang, I.-C. A., Leighton, D. C., & Cuttler, C. (2017). Qualitative research. In Research methods in psychology (2nd Canadian ed.). BCcampus. opentextbc.ca
- Key terms
- Grounded theory
- A tradition that builds theory inductively from data via systematic coding and constant comparison.
- Phenomenology
- A tradition seeking the essence of the lived experience of a phenomenon.
- Ethnography
- The study of a culture-sharing group in its natural setting, often through immersion.
- Semi-structured interview
- An interview using a topic guide while allowing flexible probing.
- Participant observation
- Data collection by taking part in and observing a setting, recorded in field notes.
- Data saturation
- The point at which new data cease to yield new codes, themes, or insight.
Analyzing Qualitative Data and Establishing Trustworthiness
- Describe the process of qualitative coding and thematic development.
- Apply Lincoln and Guba's trustworthiness criteria and corresponding techniques.
Qualitative analysis transforms unstructured text - transcripts, field notes, documents - into a defensible interpretive account. It is systematic, not impressionistic, and it has its own standards of rigor. This lesson covers the mechanics of coding and the trustworthiness framework that plays the role validity and reliability play in quantitative work.
One misconception needs dislodging first. Qualitative analysis is not subjective in the pejorative sense, a matter of a researcher reporting impressions. It is a disciplined, auditable process. Its claims must be traceable to specific data and defensible against rival readings. What differs from quantitative work is not the presence of rigor but the form rigor takes. That is the argument this lesson builds toward.
Here is the study we will work with. A researcher interviews 15 secondary teachers about burnout, records 22 hours of audio, and ends up with roughly 300 pages of transcript. The question is how that pile becomes a finding somebody should believe. Two versions are possible. In the weak version she reads the transcripts, notices that teachers talk about workload, support, and recognition, and reports those three as themes with a quotation each. In the strong version she codes every transcript line by line, tracks how codes cluster, actively hunts for teachers who contradict the emerging pattern, and ends with a claim: that unmanageable workload was experienced as a betrayal of a vocation, not merely as too much work. The second version is longer, slower, and defensible. This lesson is about the difference.
Key idea: Qualitative analysis is a traceable procedure, and its output is an argued interpretation rather than a list of topics people mentioned.
In plain terms
The work has two halves. Half one is coding. Read a transcript. Put a short label on each chunk that says what is going on there. Do it for every transcript. Then look at your labels and see which ones belong together. A group of labels that share one idea becomes a theme. A theme has to say something, not just name a topic. "Workload" names a topic. "Workload felt like a betrayal" says something. Half two is showing your work. Anyone reading you should be able to ask "where did that come from?" and get an answer. So keep four things. Quotes that back each theme. A record of how your labels changed. A note on who you are and what you brought. And at least one case that did not fit, plus what you did about it. That is what trustworthiness means in practice.
Coding
Coding is the assignment of short labels to segments of data that capture their content or meaning, enabling the researcher to retrieve, compare, and organize related material. Analysts often distinguish two logics of code creation:
- Inductive (open) coding derives codes from the data themselves, letting categories emerge - the default in grounded theory.
- Deductive (a priori) coding applies a predetermined codebook drawn from theory or prior research.
In grounded theory the sequence typically moves through three stages. Open coding fractures the data into concepts. Axial coding relates categories to subcategories. Selective coding integrates everything around a core category. Throughout, constant comparison keeps categories grounded by continually comparing new segments against existing codes. Memo-writing records the analyst's developing interpretations along the way. Codes then cluster into broader themes, which are patterns of meaning that answer the research question. A theme is not merely a frequent word. It is an interpretive construct the analyst argues for from the evidence.
Two finer distinctions sharpen the craft. An in vivo code uses the participant's own words as the label. It preserves their language rather than the analyst's, which guards against imposing outside categories too early. A codebook records each code's name, its definition, and an example, whether the codes were built inductively or from theory. That makes coding consistent and communicable. Software such as NVivo or Atlas.ti manages the mechanics of tagging and retrieval. It does not do the interpretation. The analyst still decides what each segment means.
Qualitative analysis rarely waits until all data are collected. In grounded theory especially, analysis begins with the first interview and shapes the next, through the theoretical sampling and constant comparison of earlier lessons. Even outside grounded theory, early coding surfaces gaps and ambiguities that improve later data collection. Treat analysis as a phase that starts only after fieldwork ends and you forfeit that feedback. You often end up with data that cannot answer the question you sharpened along the way.
Key idea: Coding is not filing - it is the work of deciding what each piece of data means, and the codebook makes those decisions visible to others.
Thematic analysis, step by step
Braun and Clarke's widely used approach to thematic analysis lays out six recursive phases that make the process visible. First, become familiar with the data by reading and re-reading. Second, generate initial codes across the dataset. Third, search for themes by grouping codes. Fourth, review the candidate themes against both the coded extracts and the whole dataset. Fifth, define and name each theme. Sixth, produce the report, weaving themes and evidence together.
The phases are recursive rather than linear. Reviewing themes often sends the analyst back to recode. Defining a theme can reveal that two candidates are really one. Braun and Clarke stress a further point. Themes are actively constructed by the analyst. They do not passively "emerge" from data as if waiting to be found. Owning that constructive role, and showing the work at each phase, is what makes a thematic analysis auditable instead of a set of unsupported assertions.
Key idea: The six phases loop back on each other, and themes are built by the analyst rather than discovered lying in the data.
A worked coding illustration
Take a fragment from one of our teacher interviews. "By November I stopped planning creative lessons. I just needed to survive each day, and I felt guilty about giving the kids less." An open coder might tag three segments here: "abandoning creative practice," "survival mode," and "guilt about students." These in vivo and descriptive codes stay close to the data. They capture what the teacher said before any theory is imposed on it.
Across many transcripts, such codes recur and cluster. "Survival mode," "just getting through," and "running on empty" might group under a candidate theme of depletion as a coping strategy. "Guilt about students" and "feeling like a fraud" might group under moral distress. The move from code to theme is interpretive. The analyst argues that these scattered segments share an underlying meaning, and supports the claim with representative extracts. That argued link is what makes a theme. The frequency of a word does not.
A frequent weakness is the theme that is really just a topic bucket. Reporting that "participants discussed workload, support, and recognition" summarizes domains rather than making an analytic point. It stops exactly where analysis should begin. A genuine theme says something. Not that teachers discussed workload, but that unmanageable workload was experienced as a betrayal of a vocation. The test is whether the theme has a central organizing idea that answers the research question, rather than merely naming an area participants talked about.
Key idea: A topic is what people talked about; a theme is what you claim it means, and only the second one is a finding.
Two stances toward coding
Qualitative analysts differ on whether coding should be checked for reliability. Codebook and coding-reliability approaches are common where teams work together, or where a positivist audience expects it. Multiple coders apply a shared codebook and their agreement is quantified, much as inter-rater reliability is handled in quantitative work. This suits deductive, structured analyses and larger teams.
Reflexive thematic analysis takes the opposite view. Braun and Clarke argue that coding is an interpretive act by a situated analyst, so hunting for a single correct coding to agree on misunderstands the method. Here rigor comes from depth of engagement, reflexivity, and a clear analytic trail rather than from an agreement coefficient. Neither stance is universally right. What matters is that the study's approach to coding matches its paradigm and is stated openly rather than left ambiguous.
Key idea: Some traditions want two coders to agree; others treat coding as interpretation and want a visible trail instead - just say which one you are doing.
Trustworthiness
Positivist validity and reliability presuppose a single objective reality. Lincoln and Guba therefore proposed four parallel criteria for judging qualitative rigor, each with its own associated techniques.
| Criterion | Quantitative parallel | Illustrative techniques |
|---|---|---|
| Credibility | Internal validity | Triangulation, member checking, prolonged engagement |
| Transferability | External validity | Thick description enabling readers to judge fit |
| Dependability | Reliability | Audit trail of decisions and changes |
| Confirmability | Objectivity | Reflexivity, documentation linking findings to data |
- Credibility asks whether the findings are believable representations of participants' realities. Triangulation - convergence across multiple data sources, methods, investigators, or theories - and member checking - returning interpretations to participants for their reaction - are its main supports.
- Transferability concerns whether findings can inform other contexts. The researcher's job is not to claim generalization but to provide thick description - richly detailed context - so readers can judge applicability to their own settings.
- Dependability concerns whether the process is consistent and traceable, evidenced by an audit trail.
- Confirmability concerns whether findings are grounded in data rather than in the researcher's bias. It is supported by reflexivity and by clear documentation linking claims to evidence.
Apply the four to our teacher study. Credibility asks whether the burnout account rings true against other evidence and against teachers who disagreed. Transferability asks whether the report describes the schools well enough that a reader in a different district can judge the fit. Dependability asks whether someone could follow how the codes changed over three rounds of analysis. Confirmability asks whether each claim can be traced to specific extracts rather than to the researcher's own memories of teaching. Four different questions, four different pieces of evidence.
Key idea: Credibility, transferability, dependability, and confirmability are the qualitative counterparts of internal validity, external validity, reliability, and objectivity - and they are not interchangeable.
Techniques that build credibility
Credibility rests on a toolkit worth knowing by name. Triangulation comes in several forms. Data triangulation compares sources or times. Investigator triangulation compares analysts. Theory triangulation compares frameworks. Methodological triangulation compares methods. Prolonged engagement and persistent observation build enough familiarity to tell the central from the incidental. Peer debriefing exposes the analysis to a knowledgeable colleague's challenge. Negative case analysis actively hunts for data that contradict the emerging account, then revises the account until it can accommodate them.
These techniques share a logic. Each is a deliberate attempt to disconfirm the researcher's preferred reading. They are the qualitative counterpart to the experimenter's controls. A study that reports only confirming quotations, never a disconfirming case, and never a challenge from a second analyst has not earned its credibility claim, however vivid its excerpts. Rigor is demonstrated by the effort to be wrong, not by the polish of the conclusion. One technique is rarely enough. Credible studies combine several, choosing those that fit the tradition and the specific threats the study faces.
Key idea: Every credibility technique is a way of trying to prove yourself wrong, which is why a study with no disconfirming case has not demonstrated rigor.
Rigor as argument
Notice that responsibility for transferability is deliberately shared. The researcher supplies thick description, and the reader decides fit. That is a very different stance from statistical generalization. Across all four criteria, qualitative rigor is an argument: that the interpretation is well grounded, transparently produced, and defensible against alternatives. It is assembled from triangulation, thick description, audit trails, and reflexivity rather than from a single coefficient.
Lincoln and Guba later added a set of authenticity criteria that go beyond method to the ethics and consequences of the inquiry, asking whether the research fairly represents different views and whether it deepens participants' and readers' understanding. You need not invoke all of these in every study, but they signal that qualitative quality is not only technical but also about doing justice to the people and meanings studied, which is consistent with the paradigm's underlying values.
Whatever the approach, the report must let a reader follow the chain of evidence from raw data to claim. This means presenting enough extracts to show that themes are grounded, describing how codes became themes, and being candid about disconfirming cases and how they were handled. A results section that offers polished conclusions with a scattering of confirming quotes, but no visible analytic path, asks the reader to trust rather than to verify, which is precisely what the trustworthiness framework exists to prevent.
The aim is not to eliminate the researcher's interpretive role, which is both impossible and undesirable. The aim is to make that role transparent enough that others can assess it. Rigor and interpretation are partners here rather than opposites, and a strong qualitative report shows both at work on every claim it makes.
Key idea: A reader should be able to walk backward from any claim to the extracts that produced it - that traceability is what rigor means here.
Where people get stuck
Four distinctions do the heavy lifting in this material.
- Credibility versus transferability. Credibility is about this study being a believable account of these participants. Transferability is about whether a reader elsewhere can judge the fit to their own setting. The first is your job through triangulation and negative cases; the second is your job only in the sense of supplying thick description, after which the reader decides.
- Dependability versus confirmability. Dependability asks whether the process was consistent and traceable, which an audit trail supplies. Confirmability asks whether the findings came from the data rather than from you, which reflexivity and a chain of evidence supply. One is about procedure, the other about grounding.
- Code versus theme. A code labels a segment. A theme makes an interpretive claim across many segments. "Survival mode" is a code; depletion as a coping strategy is a theme.
- Theme versus topic. "Workload" is a topic bucket. "Unmanageable workload experienced as a betrayal of a vocation" is a theme, because it has a central organizing idea and answers the question.
A useful memory aid for the four criteria: credibility is about truth, transferability about reach, dependability about process, confirmability about grounding. When a committee says a study "has not established trustworthiness," ask which of those four is missing, because each has a different remedy.
Key idea: Credibility is truth, transferability is reach, dependability is process, and confirmability is grounding - name which one a critique is aimed at before you try to fix it.
Try it
Exercise. A dissertation reports: "I interviewed 15 teachers, read the transcripts, and identified three themes about burnout, which I illustrate with quotations." A committee member writes that the analysis "has not established trustworthiness." Specify what is missing and describe concrete steps that would address each of Lincoln and Guba's four criteria.
Model answer. The description reports a result but not a defensible process, so each criterion is unmet. For credibility, the student should add techniques that test the reading: triangulate the interviews against another source such as observation or documents, seek disconfirming cases, debrief with a peer, and return themes to some teachers as member reflections. Illustrative quotations alone show the themes exist in the data but not that rival interpretations were ruled out.
For transferability, the student should provide thick description of the teachers, the schools, and the context, so readers can judge fit to their own settings. Claiming the findings generalize is not the move. For dependability, an audit trail documenting how codes were developed, merged, and revised makes the process traceable and consistent. For confirmability, two things are needed: a reflexivity statement about the student's own stance as a former teacher, and a clear chain of evidence linking each theme to specific coded extracts. Together these turn a bare report into the argued, auditable account that trustworthiness requires.
Recap
- Coding assigns labels to data segments, inductively or from a codebook, and grounded theory moves from open through axial to selective coding.
- Themes are interpretive constructs the analyst argues for, not frequent words, and a topic bucket is not a theme.
- Braun and Clarke's six phases of thematic analysis are recursive, and themes are constructed rather than passively emerging.
- Coding-reliability approaches quantify agreement between coders, while reflexive thematic analysis treats coding as situated interpretation; state which you are using.
- Lincoln and Guba's four trustworthiness criteria are credibility, transferability, dependability, and confirmability.
- Credibility techniques include triangulation, prolonged engagement, peer debriefing, member checking, and negative case analysis, all of which try to disconfirm the preferred reading.
- Transferability is a shared responsibility: the researcher supplies thick description and the reader judges fit.
- Rigor here is a traceable argument from raw data to claim, which is why an audit trail and reflexivity statement are part of the evidence.
Sources
- Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. Sage. find source ↗
- Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77-101. find source ↗
- Korstjens, I., & Moser, A. (2018). Series: Practical guidance to qualitative research. Part 4: Trustworthiness and publishing. European Journal of General Practice, 24(1), 120-124. pmc.ncbi.nlm.nih.gov
- Saunders, B., Sim, J., Kingstone, T., Baker, S., Waterfield, J., Bartlam, B., ... Jinks, C. (2018). Saturation in qualitative research: Exploring its conceptualization and operationalization. Quality & Quantity, 52(4), 1893-1907. pmc.ncbi.nlm.nih.gov
- Birt, L., Scott, S., Cavers, D., Campbell, C., & Walter, F. (2016). Member checking: A tool to enhance trustworthiness or merely a nod to validation? Qualitative Health Research, 26(13), 1802-1811. pubmed.ncbi.nlm.nih.gov
- Carter, N., Bryant-Lukosius, D., DiCenso, A., Blythe, J., & Neville, A. J. (2014). The use of triangulation in qualitative research. Oncology Nursing Forum, 41(5), 545-547. pubmed.ncbi.nlm.nih.gov
- Olmos-Vega, F. M., Stalmeijer, R. E., Varpio, L., & Kahlke, R. (2023). A practical guide to reflexivity in qualitative research: AMEE Guide No. 149. Medical Teacher, 45(3), 241-251. pubmed.ncbi.nlm.nih.gov
- O'Brien, B. C., Harris, I. B., Beckman, T. J., Reed, D. A., & Cook, D. A. (2014). Standards for reporting qualitative research (SRQR). EQUATOR Network. equator-network.org
- Key terms
- Coding (qualitative)
- Assigning short labels to data segments to capture and organize their meaning.
- Constant comparison
- Continually comparing new data against existing codes to keep categories grounded.
- Theme
- An interpretive pattern of meaning across the data that addresses the research question.
- Triangulation
- Seeking convergence across multiple sources, methods, investigators, or theories.
- Member checking
- Returning interpretations to participants to verify credibility.
- Thick description
- Richly detailed contextual reporting that lets readers judge transferability.
Module 8: Mixed Methods, Writing, and Reproducibility
Integrating paradigms, reporting research transparently, and meeting modern standards for reproducible science.
Mixed-Methods Research Designs
- Justify mixing methods and articulate the pragmatist rationale.
- Distinguish the core mixed-methods designs by timing and integration.
Mixed-methods research intentionally combines quantitative and qualitative data within a single study or program, on the premise that the two together can answer questions neither could answer alone. Numbers can establish that an effect exists and how large it is; words can explain why it occurs and what it means to those involved. The value lies not in doing both side by side but in integration - deliberately connecting the strands so the whole exceeds the sum.
Mixed methods is sometimes called the third methodological movement, after the quantitative and qualitative traditions. Its promise is real but easy to overstate. Attaching a few interviews to a survey does not make a study mixed-methods in any meaningful sense. What earns the label is a design in which the strands are planned to inform one another and are brought together at a defined point. That discipline is what this lesson develops.
One case will show what integration means in practice. A health system rolls out a text-message reminder program for missed appointments. A trial across twelve clinics finds that no-show rates drop from 22% to 17% overall. Good news, except that at three clinics nothing changed at all. The numbers say what happened and cannot say why. So the team returns to those three clinics, interviews twelve staff and eighteen patients, and learns something the trial could never have produced: at those sites the reminders arrived in English to households that spoke Spanish, and were read as debt-collection messages. Notice the structure. The quantitative result chose the cases. The qualitative work explained the anomaly. Neither strand alone would have got there.
Key idea: Mixed methods means the two strands are planned to inform each other and are actually joined at a stated point - not simply reported in the same document.
In plain terms
Numbers tell you what happened and how much. Words tell you why, and what it meant to the people involved. Mixed methods uses both. But putting a survey chapter next to an interview chapter is not mixed methods. The two have to meet somewhere, and you have to say where. Three basic shapes exist. Run both at once and lay the results side by side. Run the numbers first, then use interviews to explain something puzzling in them. Or talk to people first, then build a measure from what they said and test it on a crowd. Pick the shape that matches your reason for mixing. Then answer one question in your write-up: what did you conclude from putting the two together that neither could have told you alone? If you cannot answer it, you did two studies, not one.
The pragmatist rationale
Mixing methods raises a philosophical worry. Positivism and interpretivism rest on opposed assumptions, so can they be combined coherently? The most common answer is pragmatism. It treats the research question as paramount and methods as tools chosen for their usefulness in answering it. It declines to make the paradigm war a precondition for empirical work. On this view the paradigm-incompatibility objection is set aside in favor of what actually illuminates the problem. Pragmatism is not an evasion. It is an explicit stance with its own philosophical lineage.
Pragmatism is not the only philosophical home for mixing. A dialectical stance, associated with Greene, deliberately holds the two paradigms in tension instead of dissolving it. It treats the friction between quantitative and qualitative findings as a source of insight. A transformative stance, associated with Mertens, frames mixing around advancing justice for marginalized groups. An emancipatory aim then guides which methods are combined and how. Naming your stance clarifies why the mixing is coherent rather than merely opportunistic.
Key idea: Pragmatism licenses mixing by putting the question first, and the dialectical and transformative stances offer principled alternatives.
Why mix at all
Greene, Caracelli, and Graham identified five distinct purposes for combining methods, and naming yours disciplines the design. Triangulation seeks convergence of results from different methods on the same question. Complementarity uses each method to illuminate a different facet. Development uses one method to build the next, as when interviews shape a survey. Initiation deliberately seeks paradox and fresh perspectives by juxtaposing the strands. Expansion extends the study's scope by using different methods for different components. The text-message study is a case of complementarity shading into explanation: the trial measured the drop, the interviews explained the exceptions.
The purpose should precede the design, not follow it. A study whose aim is complementarity needs both strands to run in parallel on related but distinct questions, while a study aiming at development must run them in sequence so the first can inform the second. Choosing a design without first naming the purpose is how mixed-methods projects drift into two disconnected studies that never actually meet.
These purposes also carry ethical and interpretive weight. An initiation purpose courts contradiction, so it requires the researcher to resist smoothing conflicting results into a tidy story. There the paradox is the finding. A development purpose obliges honesty about how much the first strand really shaped the second, rather than claiming a grounding that a rushed qualitative phase never provided. Purpose therefore sets two things: the structure of the design, and the standard against which its integration will be judged.
Key idea: Name your purpose for mixing before you choose a design, because the purpose is what makes one structure right and another wrong.
Core designs
Mixed-methods designs are distinguished by two features. The first is the timing of the strands, sequential or concurrent. The second is the priority given to each. Three prototypes recur.
- Convergent (parallel) design: collect quantitative and qualitative data at roughly the same time, analyze each separately, then merge the results to compare or corroborate. Used when you want a fuller picture by triangulating findings from both strands.
- Explanatory sequential design: quantitative first, then qualitative. You obtain numerical results and then use qualitative follow-up to explain them - especially surprising, non-significant, or outlier findings. Numbers point to what needs explaining; words supply the explanation.
- Exploratory sequential design: qualitative first, then quantitative. You explore a poorly understood phenomenon qualitatively, then use those insights to build something testable - for instance, developing a survey instrument or hypotheses grounded in the qualitative phase, and then testing them on a larger sample.
| Design | Timing | Core purpose |
|---|---|---|
| Convergent | Concurrent | Merge to corroborate or complete |
| Explanatory sequential | Quant then qual | Explain quantitative results |
| Exploratory sequential | Qual then quant | Build and test from qualitative insight |
A simple notation records the priority of the strands. Uppercase marks the dominant method and lowercase the secondary one. So QUAN plus qual denotes a mainly quantitative study with a supporting qualitative component, and a sequence such as QUAL then quan shows order and priority at once. Our text-message study is QUAN then qual. Beyond the three prototypes, more complex designs exist. Embedded designs nest one strand inside a larger design of the other. Multiphase programs chain several studies over time, which is common in program evaluation.
Key idea: Convergent designs merge, explanatory sequential designs explain a number, and exploratory sequential designs build a measure from insight.
The three prototypes in practice
A convergent example: a study of patient satisfaction administers a validated satisfaction scale and, in the same period, conducts interviews about care experiences. The scores say how satisfied patients are and the interviews say what satisfaction means to them, and merging the two in a joint display shows whether high scorers describe the same things the scale assumes. Where scores and stories diverge, the design has found something the scale alone would have missed.
An explanatory-sequential example is our text-message study. A trial finds an effect on average with puzzling variation, and interviews with high and low responders explain it. The numbers set the agenda and the words answer it. An exploratory-sequential example runs the other way. Interviews with an under-studied group reveal dimensions of a construct no existing scale captures. Those dimensions become items on a new instrument, and the instrument is then tested on a large sample.
Notice how the purpose from earlier maps onto the design. The convergent case serves triangulation or complementarity. The explanatory case serves explanation of a prior result. The exploratory case serves development, building a measure grounded in local meaning. When purpose and design agree, the strands reinforce each other. When they do not, the strands sit awkwardly and the integration feels forced.
Key idea: Match the design to the purpose - and if the strands never needed each other, you had two studies rather than one mixed-methods study.
Integration and its challenges
The intellectual core, and the hardest part, is integration. You join the strands by merging them, by connecting them so one phase informs the sampling or instruments of the next, or by embedding one inside the other. A joint display is the standard integration device: a table or figure arraying quantitative and qualitative results side by side against common dimensions.
Integration also creates distinctive difficulties. Divergent findings must be reconciled when the strands disagree, which is often revealing rather than embarrassing. The two traditions have different sample-size logics that must be balanced. And competent mixed-methods work demands genuine fluency in both toolkits. Done well, integration is exactly what makes a mixed-methods study more than a quantitative and a qualitative study stapled together.
Integration can occur at three levels, and strong designs are explicit about which. At the design level, the strands are structurally linked through timing and priority. At the methods level, they connect through sampling, instruments, or data transformation, as when qualitative themes become variables to count. At the interpretation level, the analyst draws meta-inferences, conclusions that arise from considering both strands together and that neither alone could support.
Integration sometimes involves transforming one kind of data into the other. Quantitizing converts qualitative codes into counts or presence indicators that can enter a statistical analysis, while qualitizing turns quantitative profiles into narrative types. These moves can be illuminating but carry risk, since counting themes can strip the context that gave them meaning, and they should be reported transparently so readers can judge what was gained and lost in the conversion.
A useful vocabulary describes what happens when the strands meet. Confirmation occurs when qualitative and quantitative findings agree, which strengthens confidence. Expansion occurs when they address different aspects and together give a fuller picture. Discordance occurs when they conflict. Far from being a failure, discordance often points to a measurement problem, an unrecognized moderator, or a subtlety worth pursuing. Our three flat clinics are discordance turned into the study's most useful finding. A joint display arraying the two strands against common dimensions is where these relationships become visible.
Key idea: Integration is the study - state where the strands meet, what you concluded from meeting them, and what you did when they disagreed.
Judging a mixed-methods study
Mixing methods multiplies the ways a study can go wrong, so it has developed its own quality vocabulary. Beyond the validity of each strand, analysts speak of legitimation, the trustworthiness of the integrated inference, which includes whether the two samples are compatible enough to combine, whether the strengths of one strand offset the weaknesses of the other, and whether the meta-inferences genuinely follow from both strands rather than leaning on the one the researcher prefers.
A practical checklist helps. Is the reason for mixing stated, and does the design match it? Are both strands executed competently by their own standards, not just the strand the author is comfortable with? Is there a real integration point, visible in a joint display or a meta-inference, rather than two results reported in separate chapters? A study that fails these is mixed in name only, and a committee will say so.
Mixing is not always the right choice. It roughly doubles the demands of a project. It requires competence, time, and often funding for two full methodologies. A question that one strand answers well gains nothing from a token second strand. The honest test is whether the integrated design answers a question neither method could answer by itself. If a single well-designed study suffices, adding a perfunctory second component for appearance is effort spent on decoration rather than knowledge.
Key idea: Judge a mixed-methods study on whether both strands are competent and whether they actually meet - and be willing to conclude that mixing was unnecessary.
Where people get stuck
Three points cause repeated trouble.
- Mixed methods versus multi-method. Running a survey and an interview study in the same dissertation is multi-method. It becomes mixed-methods only when there is a stated point of contact and a conclusion drawn from the pair. The test: name the meta-inference. If you cannot, you have two studies.
- Explanatory versus exploratory sequential. The names describe what the second strand does. Explanatory sequential runs quantitative first and uses qualitative work to explain the number. Exploratory sequential runs qualitative first and uses quantitative work to test what the qualitative phase revealed. Order and purpose move together.
- Divergence versus failure. When the strands disagree, the instinct is to hide it. Discordance is often the most informative result you have, because it points at a bad measure, a missing moderator, or a local condition nobody anticipated. Report it, then explain it.
A last note on sample logic. The two strands answer to different standards, and forcing one standard on both is a common error. It is entirely coherent to run a trial with 2,400 patients and interview 30 of them. The trial needs numbers for precision; the interviews need depth for meaning. A reviewer who asks why the qualitative sample was not also 2,400 has misunderstood both traditions.
Key idea: The two strands keep their own sample-size logic, so a big trial and a small interview study inside one design is correct rather than lopsided.
Try it
Exercise. A researcher finds, in a large survey, that a new mentoring program raises retention overall but has no effect at three of twelve sites, and cannot tell from the numbers why those sites differ. Name the mixed-methods design that fits, describe its sequence and integration point, and state the meta-inference it could support.
Model answer. This calls for an explanatory sequential design, because a quantitative result already exists and the need is to explain a puzzling pattern within it. The quantitative phase is complete; the qualitative phase follows, purposively sampling the three null sites and a few successful ones for interviews or observation to understand what differs in how the program was implemented or received there.
The integration point is the connection between phases: the quantitative finding selects the cases and focuses the qualitative questions, an instance of connecting rather than merging. A joint display could array each site's quantitative effect beside the qualitative themes about its implementation. The meta-inference might be that the program works only where local mentors were trained and matched deliberately. Neither strand could reach that conclusion alone. The survey shows where the effect failed, and the interviews show why. That combined explanation is actionable for scaling the program, and it is precisely the payoff integration is meant to deliver.
Recap
- Mixed methods combines quantitative and qualitative strands with deliberate integration; reporting both without joining them is multi-method work.
- Pragmatism is the usual philosophical home for mixing, with dialectical and transformative stances as principled alternatives.
- Greene and colleagues' five purposes are triangulation, complementarity, development, initiation, and expansion, and the purpose must precede the design.
- The three prototypes are convergent, explanatory sequential, and exploratory sequential, distinguished by timing and priority.
- Integration happens at the design, methods, and interpretation levels, and joint displays and meta-inferences are where it becomes visible.
- Quantitizing and qualitizing convert data between forms, and both must be reported transparently because context can be lost.
- Confirmation, expansion, and discordance describe what happens when strands meet, and discordance is frequently the most informative result.
- Mixing doubles a project's demands, so the honest test is whether the integrated design answers a question neither strand could answer alone.
Sources
- Fetters, M. D., Curry, L. A., & Creswell, J. W. (2013). Achieving integration in mixed methods designs: Principles and practices. Health Services Research, 48(6, Pt 2), 2134-2156. pmc.ncbi.nlm.nih.gov
- Tariq, S., & Woodman, J. (2013). Using mixed methods in health research. JRSM Short Reports, 4(6). pmc.ncbi.nlm.nih.gov
- O'Cathain, A., Murphy, E., & Nicholl, J. (2010). Three techniques for integrating data in mixed methods studies. BMJ, 341, c4587. pubmed.ncbi.nlm.nih.gov
- Palinkas, L. A., Horwitz, S. M., Green, C. A., Wisdom, J. P., Duan, N., & Hoagwood, K. (2015). Purposeful sampling for qualitative data collection and analysis in mixed method implementation research. Administration and Policy in Mental Health, 42(5), 533-544. pmc.ncbi.nlm.nih.gov
- Creswell, J. W., & Plano Clark, V. L. (2018). Designing and conducting mixed methods research (3rd ed.). Sage. find source ↗
- Legg, C., & Hookway, C. (2021). Pragmatism. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Levitt, H. M., Bamberg, M., Creswell, J. W., Frost, D. M., Josselson, R., & Suarez-Orozco, C. (2018). Journal article reporting standards for qualitative primary, qualitative meta-analytic, and mixed methods research in psychology. American Psychologist, 73(1), 26-46. pubmed.ncbi.nlm.nih.gov
- Key terms
- Mixed-methods research
- Research intentionally combining quantitative and qualitative data with deliberate integration.
- Pragmatism
- A stance treating the research question as central and methods as tools chosen for usefulness.
- Convergent design
- Collecting quant and qual data concurrently, then merging results to compare or corroborate.
- Explanatory sequential design
- Quantitative followed by qualitative work that explains the numerical results.
- Exploratory sequential design
- Qualitative followed by quantitative work that builds and tests from qualitative insight.
- Joint display
- A table or figure arraying quantitative and qualitative results together for integration.
Scientific Writing, Transparency, and Reproducibility
- Structure a scholarly research report and describe the peer-review process.
- Explain the replication crisis and the open-science practices that respond to it.
Research that is not communicated does not exist as knowledge, and research that cannot be checked does not deserve trust. This final lesson covers the conventions of scholarly reporting and the transparency practices that have become central to credible science.
The two halves are connected. Clear reporting is not a cosmetic final step. It is the mechanism by which a private finding becomes public knowledge that others can scrutinize, build on, or overturn. A study written so that no one could reproduce it has failed at the last hurdle, however clever its design. So has a study reported so selectively that its result cannot be trusted. This lesson closes the course by showing how to communicate work in a way that earns belief.
Start with a concrete comparison. Two papers report the same experiment on a memory technique. Paper A says: "Participants completed a memory task. Outliers were removed and the analysis showed a significant improvement (p < .05)." Paper B says: "Sixty-four undergraduates completed 30 trials of a paired-associate recall task. Two participants with accuracy below 20%, a threshold set before data collection, were excluded. The technique group recalled 4.1 more words on average, 95% CI [1.2, 7.0], d = 0.52; the analysis plan and data are posted at the linked repository." Paper A cannot be checked by anyone. Paper B can be checked by anyone. That difference, not the eloquence of the prose, is what this lesson is about.
Key idea: A finding becomes knowledge only when it is reported clearly enough that someone else could check it.
In plain terms
A paper has four jobs. Say why the question matters. Say exactly what you did. Say what you found. Say what it means. That is the whole shape. The second job is the one people rush, and it is the one that matters most. Name the scale. State the rule you used to drop cases. Report the analysis you actually ran, not the tidy one. Then the second half of the lesson. Around 2011 researchers started checking whether famous findings held up. Many did not. The reasons were mostly not fraud. People ran twenty analyses and reported one. People found a pattern and then said they had predicted it. People used tiny samples. Journals printed the exciting results and binned the boring ones. The fix is simple to state. Decide what you will do before you look at the data. Write it down. Share your data and your code. Then anyone can check you.
The structure of a report
Most empirical papers follow the IMRaD structure - Introduction, Methods, Results, and Discussion - which maps onto the logic of inquiry:
- Introduction: the problem, the gap, and the research question or hypotheses (the "why").
- Methods: participants, materials, procedure, and analysis in enough detail that another competent researcher could reproduce the study (the "how"). This is the section on which reproducibility depends.
- Results: what was found, reported without interpretation, with appropriate statistics, effect sizes, and uncertainty.
- Discussion: what the findings mean, how they relate to prior work, their limitations, and their implications (the "so what").
Scholarly work is vetted through peer review, in which independent experts evaluate a manuscript before publication. Peer review is a quality filter, not a guarantee of truth: it can miss errors and fraud, and it operates on the work as described rather than as actually performed.
Several conventions surround the core structure. A precise title and an informative abstract are what most readers actually read, so they must convey the question, method, and finding honestly rather than for maximum drama. Field-specific reporting guidelines, such as CONSORT for trials, PRISMA for systematic reviews, and the APA Journal Article Reporting Standards, specify exactly what a methods and results section must contain, and following them is the simplest way to ensure a reader has what they need to evaluate the work.
The Discussion is where overclaiming most often creeps in, so it deserves discipline. Its job is to interpret the results without exceeding them. Say what the data support. Relate the finding to prior work. State limitations candidly rather than burying them. A correlational result must not acquire causal language in the Discussion that the Results could not license. A small or preliminary study should present itself as one. A frank limitations paragraph strengthens a paper, because it shows the author sees the study's boundaries as clearly as its contributions.
Key idea: IMRaD maps the logic of inquiry - why, how, what, and so what - and peer review filters quality without guaranteeing truth.
Writing that can be checked
Clarity is a reproducibility tool, not a stylistic luxury. A methods section that says a scale was "administered" without naming it cannot be reproduced. Nor can one that says outliers were "removed" without stating the rule. In both cases the reader cannot tell what was done. Precision means three things: naming instruments, stating exact procedures and decision rules, and reporting the analysis actually run, including the choices a skeptic would want to interrogate. Vagueness is not modesty. It is a barrier to the very checking that makes a claim trustworthy.
Good scholarly prose also reads well, which is a service to the reader rather than a concession. Prefer short sentences to sprawling ones. Break dense paragraphs at their natural seams. Choose the plain word over the ornate when both are exact. The goal is not to sound impressive but to be understood on the first reading, so a reader spends effort evaluating your reasoning instead of decoding your syntax.
Key idea: If a reader cannot tell exactly what you did, they cannot check it, and an unclear methods section is therefore a scientific failure rather than a stylistic one.
Authorship and credit
Who counts as an author is an ethical question with real stakes for careers. Widely used criteria, such as those of the International Committee of Medical Journal Editors, reserve authorship for those who made a substantial intellectual contribution, helped draft or revise the work, approved the final version, and agree to be accountable for it. Gift authorship, adding a senior name that did no such work, and ghost authorship, omitting someone who did, both corrupt the record of who is responsible for a study.
Related transparency norms round out honest reporting. Contributions are increasingly spelled out in author-contribution statements. Funding sources and conflicts of interest are disclosed so readers can weigh possible bias. Those who helped without meeting authorship criteria are named in acknowledgments. None of this is bureaucratic nicety. Knowing who did what, and who paid for it, is part of the evidence a reader uses to judge how much to trust a result.
Key idea: Authorship tracks intellectual contribution and accountability, so both adding a name that did nothing and omitting one that did are falsifications of the record.
The replication crisis
Large-scale replication efforts in the 2010s, most prominently in psychology, found that a substantial fraction of published findings did not replicate when studies were repeated. Several practices were implicated:
- P-hacking: trying many analyses, subsets, or exclusions and reporting only those that cross the significance threshold, which inflates false positives.
- HARKing (Hypothesizing After the Results are Known): presenting a hypothesis discovered in the data as though it had been predicted in advance, erasing the distinction between exploratory and confirmatory work.
- Underpowered studies: small samples that, when they do reach significance, tend to overestimate effect sizes and replicate poorly.
- Publication bias: journals' preference for positive results, which distorts the cumulative record.
Two related ideals are worth distinguishing carefully. Reproducibility usually means obtaining the same results from the same data and code. Replicability means obtaining consistent results from new data collected under the same procedure. Both are needed. Both were found wanting.
The crisis was documented rather than merely alleged. The Open Science Collaboration's 2015 project repeated 100 psychology studies. It reproduced a significant effect in only about a third to two-fifths of them, and the replicated effects averaged roughly half the original size. A landmark methodological paper by Simmons, Nelson, and Simonsohn showed how ordinary researcher degrees of freedom can push the false-positive rate far above the nominal 5% almost effortlessly. Those degrees of freedom are simply undisclosed flexibility in how data are collected and analyzed.
Key idea: Reproducibility means the same data give the same answer; replicability means new data give a consistent answer - and the 2010s showed both were weaker than assumed.
How false positives arise
A concrete illustration shows how easily significance can be manufactured. Suppose a researcher measures four outcomes, splits by two subgroups, and considers two possible covariate adjustments. Testing every combination gives sixteen tests, and with a few extra outlier rules it is easily twenty. Each test carries its own chance of a false positive, so at least one "significant" result becomes nearly inevitable even when nothing is going on. Report only the winner and hide the other nineteen, and you have p-hacked. It is often done with no intent to deceive at all.
Gelman's phrase, the garden of forking paths, captures a subtler version. A researcher may run a single analysis and still have chosen it after seeing the data, from among many that would have seemed reasonable had the data looked different. The false-positive rate inflates just the same, with no explicit multiple testing anywhere. This is why the timing of analytic decisions matters so much, and why the reforms below focus on fixing those choices before the data can influence them.
Key idea: Twenty analyses produce a significant result by chance almost every time, and reporting only the winner is the most common way honest researchers publish false positives.
Open-science responses
The reform movement offers concrete remedies. Pre-registration - publicly time-stamping hypotheses and the analysis plan before data collection - hardens the line between confirmatory and exploratory analysis and blocks HARKing and much p-hacking. The registered report goes further: peer review and in-principle acceptance occur before results are known, so publication no longer hinges on the outcome. Open data and open materials let others reproduce analyses and reuse instruments.
And a priori power analysis guards against underpowered designs. None of these makes fraud impossible, and none replaces good judgment. Together they shift the incentives toward transparency, which is the direction credible, cumulative science must move. The through-line of this entire course holds here too. Rigorous method is the disciplined anticipation of the ways an inference can go wrong, and honest reporting is what lets the community check that the anticipation succeeded.
The reforms extend into infrastructure as well. Open code, ideally in a form another analyst can actually run, addresses computational reproducibility: the ability to regenerate a paper's numbers from its inputs. Data shared under FAIR principles, meaning findable, accessible, interoperable, and reusable, serve the wider community beyond the original team. Preprints speed dissemination and open review. Journal policies such as the Transparency and Openness Promotion guidelines and open-science badges nudge a whole field toward these norms, rather than leaving them to individual virtue.
The reforms also change how the literature accumulates. When null results and full analyses are shared rather than hidden, later meta-analyses can combine studies without the distortion of publication bias, and the field's collective estimate of an effect grows more accurate over time. Large collaborative replication projects, running the same protocol across many labs, now test key findings directly. The shift is from prizing the single surprising result toward valuing the reliable, cumulative record that a mature science actually runs on.
These practices are moving from optional to expected. Funders and journals increasingly mandate data-sharing plans and preregistration for confirmatory work. Doctoral training now treats transparency as a core competence rather than an add-on. Adopting these habits early is both an ethical stance and a practical one. Work that is transparent and reproducible is harder to dismiss and easier to build a reputation on.
Key idea: Preregistration, registered reports, open data and code, and planned power all work by fixing decisions before the data can influence them.
Where people get stuck
Four distinctions matter here, and the first two are frequently swapped.
- Reproducibility versus replicability. Reproducibility is same data, same answer. Replicability is new data, consistent answer. A paper can be perfectly reproducible and fail to replicate, because the original result was a false positive that the code faithfully regenerates.
- Exploratory versus confirmatory. Exploratory analysis is legitimate and valuable. What is not legitimate is presenting it as confirmatory. Preregistration does not forbid exploration; it labels it, so a reader knows which claims were predicted and which were discovered.
- P-hacking versus fraud. P-hacking is usually done in good faith by researchers who believe each individual decision was reasonable. That is exactly why procedural safeguards work better than exhortations to integrity.
- Peer review versus verification. Reviewers evaluate the work as described, not as performed. They rarely see the data or rerun the code. Peer review filters; open data and code verify.
One habit will serve you for a career. Before you look at your data, write down the primary outcome, the analysis, the exclusion rule, and the sample size, and time-stamp it somewhere you cannot quietly edit. Everything else you find afterward is still worth reporting, labeled as exploratory. That single discipline eliminates most of the practices this lesson has catalogued.
Key idea: Write down your primary outcome, analysis, exclusion rule, and sample size before you look at the data, and report everything found afterward as exploratory.
Try it
Exercise. A colleague describes their study: "We collected data until the effect became significant, tried a few outlier rules until one worked, tested several outcomes and reported the one that reached p below 0.05, and framed the successful result as our original hypothesis." Identify each questionable practice and describe how a preregistered design would prevent it.
Model answer. Four practices are present. Collecting data until significance is optional stopping, which inflates false positives because the researcher keeps sampling until noise happens to cross the line. Trying outlier rules until one works, and testing several outcomes while reporting only the winner, are forms of p-hacking that exploit researcher degrees of freedom. Framing a data-driven result as the original prediction is HARKing, which erases the exploratory-confirmatory distinction the course has stressed throughout.
A preregistration would prevent each by fixing decisions before the data are seen. It would specify the sample size from an a priori power analysis, so stopping is not contingent on the result; it would state the outlier rule and the single primary outcome in advance, removing the flexibility that p-hacking exploits; and it would record the hypothesis with a time stamp, so a data-driven finding could not later be dressed as a prediction.
Anything discovered beyond the plan remains legitimate as clearly labeled exploratory analysis, to be confirmed in a fresh study. This is the through-line of the entire course. Rigorous method is the disciplined anticipation of the ways an inference can go wrong. Honest, transparent reporting is what lets the community verify that the anticipation succeeded.
Recap
- IMRaD maps the logic of inquiry, and the Methods section is where reproducibility is either secured or lost.
- Peer review evaluates work as described rather than as performed, so it filters quality without guaranteeing truth.
- Precision in reporting - named instruments, stated decision rules, the analysis actually run - is a reproducibility tool rather than a stylistic preference.
- Authorship requires substantial intellectual contribution and accountability, which is why gift and ghost authorship corrupt the record.
- Reproducibility means same data and same answer; replicability means new data and a consistent answer.
- P-hacking, HARKing, optional stopping, underpowered designs, and publication bias together explain much of the replication crisis.
- The garden of forking paths shows that a single analysis chosen after seeing the data inflates false positives just as multiple testing does.
- Preregistration, registered reports, open data and code, FAIR sharing, and a priori power analysis all work by fixing decisions before the data can influence them.
Sources
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. pubmed.ncbi.nlm.nih.gov
- Munafo, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., ... Ioannidis, J. P. A. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, 0021. pmc.ncbi.nlm.nih.gov
- Nosek, B. A., Alter, G., Banks, G. C., Borsboom, D., Bowman, S. D., Breckler, S. J., ... Yarkoni, T. (2015). Promoting an open research culture. Science, 348(6242), 1422-1425. pmc.ncbi.nlm.nih.gov
- Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. PNAS, 115(11), 2600-2606. pmc.ncbi.nlm.nih.gov
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. pubmed.ncbi.nlm.nih.gov
- National Academies of Sciences, Engineering, and Medicine. (2019). Reproducibility and replicability in science. NCBI Bookshelf. ncbi.nlm.nih.gov
- Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., ... Mons, B. (2016). The FAIR guiding principles for scientific data management and stewardship. Scientific Data, 3, 160018. pmc.ncbi.nlm.nih.gov
- Center for Open Science. (2025). Transparency and Openness Promotion (TOP) guidelines. Center for Open Science. cos.io
- Appelbaum, M., Cooper, H., Kline, R. B., Mayo-Wilson, E., Nezu, A. M., & Rao, S. M. (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1), 3-25. pubmed.ncbi.nlm.nih.gov
- Key terms
- IMRaD
- The Introduction, Methods, Results, and Discussion structure of an empirical report.
- Peer review
- Evaluation of a manuscript by independent experts before publication; a filter, not a guarantee.
- P-hacking
- Running many analyses and reporting only those that reach statistical significance.
- HARKing
- Presenting a post hoc hypothesis found in the data as if it had been predicted in advance.
- Reproducibility
- Obtaining the same results from the same data and analysis code.
- Pre-registration
- Publicly recording hypotheses and analysis plans before data collection.