Module 1: The Logic and Paradigms of Qualitative Inquiry
What makes inquiry qualitative, the interpretive paradigms that ground it, and how to turn an interest into an answerable qualitative question and design.
What Qualitative Research Is - and Is Not
- Define qualitative research by its purpose and logic rather than merely by the absence of numbers.
- Distinguish the questions qualitative inquiry answers well from those it answers poorly.
Qualitative research is often introduced by what it lacks - no numbers, no hypotheses, no statistics - which is exactly the wrong way to understand it. Defined negatively, it looks like a deficient version of quantitative work. Defined positively, it is a distinct and demanding way of producing knowledge, with its own logic, its own criteria of quality, and its own kinds of questions that no survey or experiment can answer.
Consider a concrete case. A hospital finds that nurse reports of near-miss medication errors on one ward have fallen by forty percent in a year. The quantitative record is unambiguous and, on its face, encouraging. A qualitative study of that same ward finds something the counts cannot show: nurses stopped reporting because a redesigned electronic form now requires the prescriber to be named, and naming a consultant is socially expensive on a small unit. The reporting rate fell; the error rate did not. No amount of additional counting would have produced that explanation, because the explanation lives in what the form means to the people filling it in and what filling it in costs them. That is the territory qualitative research occupies.
Key idea: Qualitative research is not quantitative research minus the statistics. It is a different inferential enterprise with a different unit of warrant. A quantitative study asks you to believe a claim because a defined procedure was applied to a sample that stands in for a population. A qualitative study asks you to believe a claim because an interpretation has been built from, and repeatedly tested against, a body of evidence that you as reader can inspect.
A working definition
Qualitative research is the systematic study of how people interpret and make meaning of their experiences, actions, and social worlds, conducted through close, sustained engagement with participants in their own settings and reported largely in words. Its data are texts and images - interview transcripts, fieldnotes, documents, recordings - rather than counts. Its aim is understanding (in Weber's sense of Verstehen) rather than prediction, and its findings are typically interpretations defended with evidence rather than estimates with margins of error.
What follows from that definition
- The research is naturalistic: phenomena are studied in situ, not stripped of context in a laboratory. Context is not noise to be controlled away; it is part of what is being explained.
- The design is emergent: questions, sampling, and even the focus can shift as understanding develops. This is a feature, not sloppiness, but it must be documented and defended.
- The researcher is the primary instrument of data collection and analysis. There is no questionnaire that stands between investigator and world; the investigator's own perception, rapport, and judgment do the gathering, which is precisely why reflexivity becomes a methodological requirement.
- Reasoning is largely inductive and abductive, building concepts up from data and inferring the best explanation, rather than deducing predictions from theory and testing them.
Key idea: Those four features hang together as a package. Because the setting is part of the explanation, you have to be in it; because you are in it, the design cannot be fully fixed in advance; because you are the instrument, your influence has to be examined rather than assumed away; and because you are building concepts rather than testing them, your reasoning runs from evidence upward. A proposal that claims one of these features while denying the others is usually incoherent.
Analytic generalization, not statistical generalization
The single most consequential thing to understand about qualitative inference is the kind of generalization it supports. Statistical generalization reasons from a probability sample to a population by frequency: because 14 percent of a random sample reported X, roughly 14 percent of the population does, within a stated margin. Qualitative sampling is not probabilistic, so this route is closed. What remains open is analytic generalization (sometimes called theoretical or conceptual generalization): reasoning from a case to a concept or a mechanism that may operate elsewhere, with the reader judging the fit.
Two derived ideas matter. Transferability, in Lincoln and Guba's vocabulary, places the burden on the reader: the researcher's job is to supply enough contextual detail - who the participants were, what the setting was like, what constraints they worked under - that a reader can judge whether the findings plausibly travel to their own setting. Naturalistic generalization, Robert Stake's term, describes what readers do when a rich case resonates with their own tacit knowledge and reshapes how they see a situation they know. Neither is a weaker version of statistical inference; both are different logics with their own demands, chief among them the demand for detail thick enough to support the reader's judgment.
A worked illustration: one topic, four questions
Take a single topic - burnout among emergency physicians - and watch how the question dictates the method.
- How prevalent is burnout among emergency physicians in this health system? This needs a validated instrument, a defined sampling frame, and a response rate. Qualitative methods cannot answer it, and should not try.
- Does a scheduling intervention reduce burnout scores relative to control? This needs comparison and, ideally, random assignment. Again not a qualitative question.
- What does burnout mean to emergency physicians who describe themselves as burned out? Now the concept itself is at stake. If clinicians use the word to name something other than what the instrument measures, the instrument's validity is in question, and only close listening will reveal it.
- How does a physician move from engaged to depleted over a career, and what turns that trajectory? This is a process question about sequence and contingency, best pursued through longitudinal interviews or fieldwork.
The third and fourth questions are qualitative not because they are softer but because they concern meaning and process. Notice also the productive relationship: if the interview study finds that physicians consistently describe burnout as a loss of moral agency rather than as emotional exhaustion, that finding is a direct challenge to the construct the survey measures - and a research program has been born.
Questions it answers well
Qualitative inquiry is the right tool when the question is about meaning, process, or context. How do first-generation students experience belonging at an elite university? What does recovery mean to people who have left addiction treatment? How does a decision actually get made inside a hospital ethics committee? These are questions of how and why that require understanding an insider's perspective, capturing a process as it unfolds, or discovering categories that no one has yet named. When the relevant concepts are not yet well enough understood to be measured, forcing them into a closed survey would fabricate precision.
Questions it answers poorly
Qualitative work does not estimate prevalence ("what percentage of students feel they belong?"), does not establish the average size of a causal effect, and does not support statistical generalization to a population. A study of twelve interviewees cannot tell you how common a view is. Claiming otherwise - reporting that "most participants" felt something as though the sample were representative - is a category error that discredits otherwise strong work. Knowing the boundary of your method is itself a mark of rigor.
The complementarity
None of this makes qualitative inferior or superior; the two families answer different questions. Quantitative methods excel at measuring how much and how often across many cases; qualitative methods excel at illuminating what something means and how it happens within cases. Much of the strongest research programs cycle between them: qualitative work discovers and defines the constructs that quantitative work later measures, and quantitative anomalies send investigators back to the field to understand what the numbers cannot say.
A first map of the traditions
"Qualitative research" names a family, not a method. Later modules treat each member in detail, but it helps to have the map early. Ethnography, descended from Malinowski and Boas and given its interpretive turn by Clifford Geertz, studies culture through sustained participation in a setting. Phenomenology, rooted in Edmund Husserl's descriptive project and redirected by Martin Heidegger and later Max van Manen toward interpretation, seeks the structure of lived experience. Grounded theory, launched by Barney Glaser and Anselm Strauss in 1967 and since split into Glaserian, Straussian, and Kathy Charmaz's constructivist variants, builds explanatory theory from data. Case study, in Robert Yin's more postpositivist form or Robert Stake's more interpretive one, examines a bounded system in depth. Narrative inquiry, associated with D. Jean Clandinin and F. Michael Connelly and with Catherine Kohler Riessman's analytic techniques, treats stories as the unit of analysis. Discourse analysis examines language as social action rather than as a window onto an inner state.
Alongside these sits qualitative description, which Margarete Sandelowski defended as a legitimate design in its own right rather than a failed attempt at one of the others. A study that stays close to participants' own terms and does not claim to have produced a theory, an essence, or a cultural account is not deficient; it is descriptive, and saying so plainly is more honest than borrowing a tradition's label without doing its work.
Common misconceptions
Qualitative research is exploratory groundwork for the real study. Sometimes it is, but treating that as its definition is a mistake. Questions about meaning are terminal questions, not preliminary ones. A study of how patients understand a diagnosis is not waiting for a survey to complete it.
Qualitative research is subjective, so anything goes. The opposite is closer to the truth. Precisely because the researcher is the instrument, qualitative work carries an unusually heavy burden of procedural transparency: how participants were chosen, how the interviews were conducted, how codes were developed, what disconfirming evidence was found and what was done about it. Nicholas Mays and Catherine Pope made this case early and directly - the relevant question is not whether the researcher influenced the study but whether that influence has been examined and reported.
Small samples are a weakness to apologize for. Sample size in qualitative work is governed by informational adequacy, not by power calculations. A single case studied intensively can carry an argument, as Sandelowski's defense of the "n of one" argues; the failure mode is not smallness but thinness.
Numbers are forbidden. They are not. Counting how many participants raised a concern can be a legitimate analytic aid, and reporting sample characteristics numerically is standard. What is forbidden is treating those counts as estimates of population prevalence.
Try it
Take a research interest you already have and write it three ways: as a prevalence question, as a causal question, and as a meaning-or-process question. Then ask which version you actually care about. Doctoral students frequently discover at this point that the question that motivated them was qualitative all along and that they had been translating it into a quantitative form because that form felt more legitimate.
Recap
- Qualitative research is defined positively by its purpose - understanding how people make meaning - not by the absence of numbers.
- Its four hallmarks are naturalistic setting, emergent design, researcher-as-instrument, and inductive or abductive reasoning; they stand or fall together.
- It supports analytic and naturalistic generalization and transferability, never statistical generalization to a population.
- It answers questions of meaning, process, and context well; it answers questions of prevalence and average causal effect badly.
- Rigor comes from transparency about procedure and from engaging disconfirming evidence, not from claiming detachment.
Sources
- Tenny, S., Brannan, J. M., & Brannan, G. D. (2022). Qualitative study. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
- Hammarberg, K., Kirkman, M., & de Lacey, S. (2016). Qualitative research methods: When to use them and how to judge them. Human Reproduction, 31(3), 498-501. pubmed.ncbi.nlm.nih.gov
- Moser, A., & Korstjens, I. (2017). Series: Practical guidance to qualitative research. Part 1: Introduction. European Journal of General Practice, 23(1), 271-273. pmc.ncbi.nlm.nih.gov
- Mays, N., & Pope, C. (1995). Rigour and qualitative research. BMJ, 311(6997), 109-112. pmc.ncbi.nlm.nih.gov
- Sandelowski, M. (2010). What's in a name? Qualitative description revisited. Research in Nursing & Health, 33(1), 77-84. pubmed.ncbi.nlm.nih.gov
- Sandelowski, M. (1996). One is the liveliest number: The case orientation of qualitative research. Research in Nursing & Health, 19(6), 525-529. pubmed.ncbi.nlm.nih.gov
- Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. Sage Publications. find source ↗
- Key terms
- Qualitative research
- Systematic study of how people interpret and make meaning of experience, reported largely in words.
- Verstehen
- Weber's notion of interpretive understanding: grasping the meaning an action holds for the actor.
- Naturalistic inquiry
- Studying phenomena in their real-world settings rather than stripping them of context.
- Emergent design
- A design in which questions, sampling, and focus may evolve as understanding develops.
- Researcher as instrument
- The principle that in qualitative work the investigator's own perception and judgment gather and analyze data.
- Statistical generalization
- Inference from a sample to a defined population by frequency, which qualitative sampling does not support.
Interpretive Paradigms: Constructivism, Critical, and Pragmatism
- Contrast the ontology, epistemology, and aims of the constructivist, critical, and pragmatic paradigms.
- Explain how a stated paradigm licenses particular qualitative methods and criteria of quality.
Every qualitative study rests on assumptions about what social reality is and how it can be known. Making those assumptions explicit - naming your paradigm - is not academic throat-clearing; it is what lets a reader judge whether your methods and your criteria of quality actually fit your claims. Three paradigms dominate contemporary qualitative work.
The vocabulary comes from a specific text. In 1994 Egon Guba and Yvonna Lincoln published "Competing Paradigms in Qualitative Research," which laid out paradigms along three axes that every doctoral proposal is now expected to address. Ontology asks what the nature of reality is. Epistemology asks what the relationship between knower and known can be. Axiology asks what role values play in inquiry. A fourth axis, methodology, follows from the first three: given those commitments, how can one proceed? The discipline of the framework is that the axes constrain each other. If you say reality is multiple and constructed, you cannot also say your job was to record it neutrally.
Key idea: A paradigm is not a decoration on the front of a proposal. It is a set of promises about what your claims will and will not assert, and every method you subsequently choose either honors those promises or breaks them.
Constructivism (interpretivism)
The constructivist or interpretivist paradigm holds a relativist ontology: there is no single social reality waiting to be discovered, but multiple realities constructed by people as they interpret their worlds. Its epistemology is subjectivist and transactional - knowledge is co-created in the interaction between researcher and participant, so the investigator cannot and should not pretend to stand outside. The aim is to understand meaning from the participant's frame of reference, and quality is judged by trustworthiness criteria such as credibility and by the depth and coherence of the interpretation. Most interview-based studies, phenomenology, and constructivist grounded theory sit here.
The practical test of a constructivist commitment is what you do with disagreement between participants. A researcher operating under a realist assumption treats contradiction as measurement error and tries to establish who is correct. A constructivist treats contradiction as data: two participants describe the same committee meeting incompatibly because they occupy different positions in it, and the analytic task is to account for both accounts, not to adjudicate between them. Kathy Charmaz's constructivist grounded theory extends this to the researcher, insisting that the resulting theory is a construction from a particular vantage point rather than an emergent discovery.
Post-positivism, the paradigm qualitative researchers most often occupy by accident
Post-positivism deserves naming because a large share of applied qualitative work sits here without saying so. Its ontology is critical realist: there is a real world, but our access to it is imperfect and theory-laden. Its epistemology is modified objectivist - objectivity is a regulative ideal approached through triangulation, peer review, and the search for disconfirming cases. Health services research, much implementation science, and Robert Yin's version of case study operate in this register, which is why they speak comfortably of validity, reliability, and even "bias" - terms a thoroughgoing constructivist would reject as category mistakes.
This matters for coherence. A proposal that promises to capture "the participants' subjective realities" and then proposes to establish inter-coder reliability at kappa above 0.80 and to "control for researcher bias" has mixed two paradigms. Neither commitment is wrong on its own; the combination is what a committee will notice. If your criteria are reliability and bias control, say you are working post-positivistically and defend it.
The critical paradigm
The critical paradigm (with roots in Marxist thought, feminism, critical race theory, and the Frankfurt School) shares the view that reality is shaped socially, but adds that it is shaped by power - historically sedimented structures of class, race, gender, and colonial relation that come to feel natural. Its epistemology is explicitly value-laden: the researcher takes a stance, and neutrality is regarded as complicity with the status quo.
The aim is not only to understand but to critique and transform - to expose how arrangements that seem inevitable are contingent and unjust, and often to work with participants toward change. Participatory action research, much feminist and decolonizing methodology, and critical ethnography belong here. Quality includes whether the work is catalytic - whether it prompts recognition and action.
Pragmatism
Pragmatism sidesteps the metaphysical dispute. Rather than beginning from a fixed ontology, it begins from a problem and asks what inquiry would usefully address it. Truth, for the pragmatist, is what works to resolve the problem at hand; methods are tools chosen for the job, not expressions of a worldview. Pragmatism is the usual philosophical home of mixed-methods research, because it authorizes combining qualitative and quantitative tools whenever doing so answers the question better than either alone. Quality is judged by whether the inquiry yields actionable, warranted understanding of the problem.
Pragmatism's founding figures - Charles Sanders Peirce, William James, and John Dewey - offer more than a permission slip for method mixing. Peirce's account of abduction, inference to the best available explanation, is arguably the reasoning form qualitative analysts actually use when they move from a puzzling excerpt to a candidate interpretation. Dewey's insistence that inquiry begins in an indeterminate situation and ends when the situation has been rendered determinate maps closely onto how applied qualitative projects are actually justified to stakeholders. Citing pragmatism therefore commits you to something specific: your account will be judged by whether it resolves the problem, and a beautiful interpretation that changes nothing is a weaker result than a plainer one that does.
| Paradigm | Ontology | Aim | Stance on values |
|---|---|---|---|
| Constructivism | Multiple constructed realities | Understand meaning in context | Values disclosed, knowledge co-created |
| Critical | Reality shaped by power | Critique and transform injustice | Explicitly value-laden and committed |
| Pragmatism | Bracketed; problem-driven | Solve a practical problem usefully | Values instrumental to the problem |
Why naming the paradigm disciplines the design
A paradigm supplies the standard against which your work will be judged. A constructivist who claims to have discovered the single true meaning of an event has drifted into realism and invited an objection they cannot answer. A critical researcher who reports findings with detached neutrality has abandoned the very commitment that justified the design. And a study that mixes methods without a pragmatic or otherwise reasoned rationale looks opportunistic rather than principled. The paradigm is the hinge on which coherence turns: state it, then make sure every downstream choice - tradition, sampling, analysis, and criteria of rigor - follows from it.
A worked illustration: one topic, three paradigms
Suppose the topic is the experience of asylum seekers using a city's emergency shelter system. Watch how the paradigm reshapes every element of the design.
Constructivist version. Question: How do residents make sense of shelter rules and of themselves as people subject to them? Method: repeated unstructured interviews plus observation of common areas. Sampling: maximum variation across length of stay, family status, and country of origin. Analysis: reflexive thematic analysis producing themes about dignity, waiting, and surveillance. Quality criteria: credibility through prolonged engagement, an audit trail, and thick description supporting transferability. Claim made: an interpretation of how residents construct meaning, offered as one reading among possible readings, defended by evidence.
Critical version. Question: How do shelter rules reproduce the precarity they claim to relieve, and what would residents change? Method: participatory action research with a resident advisory group co-designing the interview guide. Sampling: theoretically driven toward those most affected by the rules, plus staff who enforce them, in order to expose the structure rather than to represent a population. Analysis: attention to how official categories ("non-compliant," "voluntary departure") do political work. Quality criteria: credibility plus catalytic validity - did participants come to see their situation differently, and did anything change? Claim made: a critique with an explicit normative commitment, plus an account of the action taken.
Pragmatic version. Question: Why do 30 percent of families leave the shelter within a week, and what would reduce that? Method: administrative data on exits combined with exit interviews and staff focus groups. Sampling: purposive for the interviews, complete records for the exits. Analysis: qualitative findings used to interpret patterns in the exit data, and the exit data used to select whom to interview next. Quality criteria: whether the resulting account is actionable and holds up when the shelter tries a change. Claim made: a warranted, usable explanation of a defined problem.
None of these three is more rigorous than the others. Each is rigorous by its own standard, and each would be judged a failure by the others' standards - which is precisely why the paradigm has to be declared.
Common misconceptions
Paradigm talk is empty ritual. It becomes ritual when it is written as a detached opening paragraph and never referred to again. It is substantive when it is doing work: when it explains why you did not calculate inter-rater reliability, or why you gave participants editorial control over quotations, or why you count "the group decided to petition the housing office" as a finding rather than as contamination.
Constructivism means all accounts are equally good. Relativism about reality is not relativism about evidence. A constructivist study can still be badly done: interpretations unsupported by data, disconfirming cases ignored, participants' words wrenched from context. Relativist ontology sets what you can claim; it does not lower the evidentiary bar for claiming it.
Pragmatism means you need not think about philosophy. Pragmatism is itself a philosophical position, descended from Peirce, James, and Dewey, with a distinctive theory of truth as warranted assertibility. Invoking it as a license to skip the question is a misuse; invoking it because your inquiry is genuinely problem-driven is legitimate.
Mixed methods requires pragmatism. It is the usual home, not the only one. Critical realists mix methods routinely, and some critical researchers combine survey data with participatory work. What is required is a stated rationale for why combination serves the question.
Try it
Take a published qualitative article in your field and locate its paradigm without reading the methods section's label. Look at three things: what the findings claim to be (interpretations, structures, essences, or facts), what quality criteria the authors invoke, and how they describe their own role. Then check the stated paradigm. Mismatches are common, and finding them trains the eye you will need for your own proposal.
Recap
- Guba and Lincoln's axes - ontology, epistemology, axiology, methodology - constrain one another; incoherence across them is the most common paradigm failure.
- Constructivism treats contradictory participant accounts as data to be explained, not error to be adjudicated.
- Post-positivism is where much applied qualitative work actually sits; its vocabulary of validity, reliability, and bias control is coherent only within it.
- The critical paradigm adds power and transformation, and adds catalytic validity to its quality criteria.
- Pragmatism is problem-driven and is the usual home of mixed methods, but it is a philosophical position rather than an exemption from one.
Sources
- Guba, E. G., & Lincoln, Y. S. (1994). Competing paradigms in qualitative research. In N. K. Denzin & Y. S. Lincoln (Eds.), Handbook of qualitative research (pp. 105-117). Sage Publications. find source ↗
- Moser, A., & Korstjens, I. (2017). Series: Practical guidance to qualitative research. Part 1: Introduction. European Journal of General Practice, 23(1), 271-273. pmc.ncbi.nlm.nih.gov
- Korstjens, I., & Moser, A. (2017). Series: Practical guidance to qualitative research. Part 2: Context, research questions and designs. European Journal of General Practice, 23(1), 274-279. pmc.ncbi.nlm.nih.gov
- Charmaz, K., & Keller, R. (2016). A personal journey with grounded theory methodology: Kathy Charmaz in conversation with Reiner Keller. Forum: Qualitative Social Research, 17(1). qualitative-research.net
- Starks, H., & Trinidad, S. B. (2007). Choose your method: A comparison of phenomenology, discourse analysis, and grounded theory. Qualitative Health Research, 17(10), 1372-1380. pubmed.ncbi.nlm.nih.gov
- Bhattacherjee, A. (2012). Social science research: Principles, methods, and practices (2nd ed.). University of South Florida. digitalcommons.usf.edu
- Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. Sage Publications. find source ↗
- Key terms
- Paradigm
- A shared set of ontological, epistemological, and axiological assumptions guiding inquiry.
- Constructivism
- The paradigm holding that multiple social realities are constructed and knowledge is co-created with participants.
- Critical paradigm
- A value-laden paradigm aiming to expose and transform power-based injustice.
- Pragmatism
- A problem-driven paradigm that treats methods as tools chosen for what works, the usual home of mixed methods.
- Axiology
- The study of the role of values in inquiry, on which the three paradigms sharply differ.
- Catalytic validity
- A critical-paradigm criterion asking whether research prompts recognition and action toward change.
Qualitative Research Questions and Emergent Design
- Write a central qualitative research question and aligned sub-questions in the appropriate grammar.
- Explain the logic and safeguards of emergent design and align paradigm, tradition, question, and method.
A qualitative study is only as good as its question, and qualitative questions have a distinctive grammar. Where a quantitative question asks about relationships between variables ("what is the effect of X on Y?"), a qualitative question asks about meaning, experience, and process. Getting the wording right is the first act of design, because everything downstream - who you talk to, what you ask, how you analyze - is in service of answering it.
Key idea: A qualitative research question is a promise about what kind of answer you will deliver. If the question asks what something means, the answer must be an account of meaning; if it asks how a process unfolds, the answer must have sequence and contingency in it. Committees fail proposals not because the question was uninteresting but because the promised answer could not be produced by the design underneath it.
Where qualitative questions come from
Questions are manufactured, not received. Four sources reliably generate them. The first is a construct that resists measurement: a concept everyone uses - resilience, engagement, moral distress, trust - that turns out to mean different things to different people, so that the instruments purporting to measure it may be measuring different phenomena in different groups. The second is an unexplained quantitative pattern: an intervention works in one clinic and fails in another, adherence collapses in month four, a policy produces the opposite of its intended effect. The third is a process no one has described: how does a family actually arrive at a decision to move a parent into residential care? The fourth is a silenced or unheard perspective: the experience of a group that has been studied about but rarely studied with.
Each source implies a different tradition. A contested construct points toward phenomenology or constructivist grounded theory. An unexplained pattern points toward case study or ethnography in the setting where the pattern occurs. An undescribed process points squarely at grounded theory. A silenced perspective points toward narrative inquiry or a critical participatory design. Identifying your source is a fast way to narrow the tradition before you have written a word of methodology.
The grammar of a qualitative question
Strong central questions typically open with how or what rather than whether or how much, and they use process verbs - experience, understand, construct, negotiate, make sense of - rather than causal ones. Compare:
- Weak (smuggles in a variable relationship): "Does mentoring increase the retention of first-generation doctoral students?"
- Strong (opens onto meaning and process): "How do first-generation doctoral students experience mentoring relationships, and what meaning do they attach to them?"
The strong version does not presuppose the answer, invites the participant's frame of reference, and can surface unexpected categories. A good design usually has one central question and a small set of sub-questions that break it into researchable facets without fragmenting it into a survey.
Each tradition also imposes its own grammar, and using the wrong verb signals to a reviewer that you have not internalized the tradition you claim. A phenomenological question asks about the meaning or essence of the lived experience of a phenomenon for those who have lived it. A grounded-theory question asks what process or theory explains how something comes about, usually naming a core concern that participants are managing. An ethnographic question asks how members of this group share and enact patterns of belief and behavior. A case-study question asks how and why something happened within a bounded system. A narrative question asks how a person composes and tells a story of some stretch of life. Read those five verbs again and notice that each licenses a different kind of finding.
A worked illustration: revising one question five times
Start with a doctoral student's opening draft: "What is the impact of telehealth on rural patients?" This is a topic, not a question, and "impact" is a causal verb that no interview study can honor.
- Draft 2 (still weak): "How do rural patients feel about telehealth?" - "Feel about" invites an attitude survey and produces evaluations rather than understanding.
- Draft 3 (phenomenological): "What is the meaning of receiving care at a distance for rural patients living with a chronic illness?" - Promises an account of experiential structure; requires participants who share the phenomenon and interviews that pursue lived detail rather than opinion.
- Draft 4 (grounded theory): "How do rural patients with chronic illness manage the work of maintaining a therapeutic relationship across a screen, and what process explains variation in how they do it?" - Promises a process model; requires theoretical sampling toward variation and constant comparison.
- Draft 5 (case study): "How and why did a telehealth program in one rural health district achieve high uptake among older adults while an adjacent district's failed?" - Promises an explanation bounded to two cases; requires multiple data sources and a replication logic across the two.
- Draft 6 (critical): "How does the design of telehealth services distribute the burden of access - travel, broadband, digital literacy - and whose convenience does it serve?" - Promises a critique; requires attention to structure and, typically, participant involvement in framing.
All five later drafts are defensible. None is the "right" question in the abstract. What makes one right is that the rest of your design can deliver it.
Sub-questions that help rather than fragment
Good sub-questions are facets of the central question, not a survey in disguise. Under the phenomenological Draft 3, useful sub-questions might be: What is the texture of waiting and connecting in a remote consultation? How is the clinician's presence experienced when the body is not co-present? How does the home, as the site of care, alter the meaning of the encounter? Notice that each remains a meaning question and each is answerable from the same corpus of interviews. A failing set would read: How old are participants? How often do they use telehealth? What barriers do they report? - three items that a questionnaire answers better and that fragment the phenomenon into variables.
Alignment: the through-line of a defensible study
The single most important property of a qualitative proposal is alignment - a visible through-line connecting paradigm, tradition, question, data, and analysis so that each element implies the next.
| Element | Question it answers |
|---|---|
| Paradigm | What do I assume reality and knowledge to be? |
| Tradition | What kind of qualitative study will this be? |
| Research question | What exactly do I want to understand? |
| Sampling and data | Whom and what must I study to answer it? |
| Analysis | How will I move from data to defensible interpretation? |
A misalignment anywhere is a fatal flaw a committee will find. A phenomenological question ("what is the essence of the experience of...") paired with grounded-theory coding aimed at building a process model is incoherent, because the two traditions want different things from the data.
Emergent design and its safeguards
Qualitative designs are deliberately emergent: unlike a pre-registered experiment, the study is expected to evolve as early data reshape the researcher's understanding. You may add a new type of participant because interviews revealed a perspective you had not anticipated, or refocus the central question because the phenomenon turned out to differ from your assumptions. This flexibility is a strength - it lets inquiry follow the phenomenon rather than force it - but it is also where undisciplined work goes wrong.
The safeguard is documentation, not rigidity. Every substantive change to the design should be recorded in a decision log or set of analytic memos, with the reasoning that prompted it. This audit trail lets a reader see that the study evolved for principled reasons in response to data, not to chase a preferred conclusion. Emergent design without a record of its emergence is indistinguishable from making it up as you go; emergent design with a transparent trail is one of the method's great advantages.
A concrete decision-log entry shows what this looks like in practice. "Week 6. After interviews 7 and 8, both participants spontaneously described the discharge coordinator as the person who decided what home care they received, a role absent from my original sampling frame. Two options: treat this as background context, or extend sampling to coordinators. Extending, because the central question concerns how the decision is made, and the data now indicate that a key actor was excluded by my initial assumption that families decide. Consequence: revised interview guide to add a question about who else was involved; sought amendment to the ethics approval to include staff participants. Risk noted: staff accounts may reframe the study toward organizational process, so I will keep the family perspective as the analytic anchor." Two hundred words like these, written when the decision was made rather than reconstructed at write-up, are worth more to your defense than a page of abstract claims about rigor.
What emergence does not license
Emergence has limits, and they are the ones an ethics committee and a reader will police. It does not license changing the phenomenon of interest so far that the consent participants gave no longer covers what you are doing; substantive shifts require an amendment. It does not license dropping data that resist your developing interpretation - that is the difference between following the phenomenon and chasing a conclusion. It does not license abandoning the tradition's logic midstream, so that a study begun as phenomenology quietly becomes thematic analysis without acknowledgment. And it does not license an unwritten question: if you cannot state, at any moment in the study, what your current central question is, you have drifted rather than emerged.
Common misconceptions
A good qualitative question must avoid the word "why." Not quite. "Why" questions asked of participants often produce rationalizations rather than accounts, so as an interview probe "why" is weak. As a research question, "how and why" is standard in case study, where it signals explanatory rather than descriptive intent.
The question should be finalized before entering the field. The central question should be firm enough to justify the design and loose enough to be refined by data. What must be fixed in advance is the phenomenon, the tradition, and the criteria by which you will judge whether a change is warranted.
Emergent design is incompatible with preregistration. Qualitative preregistration and published protocols are increasingly common and are compatible with emergence, provided the protocol states in advance which elements are fixed and which are expected to evolve, and provided deviations are reported.
Try it
Write your central question, then hand it to a colleague with one instruction: describe the shape of the answer this question demands - is it an essence, a process model, a cultural account, an explanation of a case, or a story? If they cannot say, the question is not yet doing its job. If they can, check that your proposed sampling and analysis could actually produce that shape.
Recap
- Qualitative questions ask how and what about meaning and process; causal and prevalence verbs signal a design mismatch.
- Each tradition has its own question grammar, and the verb you choose licenses a particular kind of finding.
- Sub-questions should be facets of the central question answerable from the same corpus, not variables in disguise.
- Alignment among paradigm, tradition, question, sampling, and analysis is the property reviewers check first.
- Emergence is legitimate when documented in a contemporaneous decision log and bounded by consent, tradition, and evidence.
Sources
- Korstjens, I., & Moser, A. (2017). Series: Practical guidance to qualitative research. Part 2: Context, research questions and designs. European Journal of General Practice, 23(1), 274-279. pmc.ncbi.nlm.nih.gov
- Starks, H., & Trinidad, S. B. (2007). Choose your method: A comparison of phenomenology, discourse analysis, and grounded theory. Qualitative Health Research, 17(10), 1372-1380. pubmed.ncbi.nlm.nih.gov
- Hammarberg, K., Kirkman, M., & de Lacey, S. (2016). Qualitative research methods: When to use them and how to judge them. Human Reproduction, 31(3), 498-501. pubmed.ncbi.nlm.nih.gov
- Crowe, S., Cresswell, K., Robertson, A., Huby, G., Avery, A., & Sheikh, A. (2011). The case study approach. BMC Medical Research Methodology, 11, 100. pmc.ncbi.nlm.nih.gov
- Korstjens, I., & Moser, A. (2018). Series: Practical guidance to qualitative research. Part 4: Trustworthiness and publishing. European Journal of General Practice, 24(1), 120-124. pmc.ncbi.nlm.nih.gov
- Aguas, P. P. (2022). Fusing approaches in educational research: Data collection and data analysis in phenomenological research. The Qualitative Report, 27(1), 1-20. nsuworks.nova.edu
- Merriam, S. B., & Tisdell, E. J. (2016). Qualitative research: A guide to design and implementation (4th ed.). Jossey-Bass. find source ↗
- Key terms
- Central question
- The single overarching interrogative that a qualitative study is designed to answer.
- Sub-questions
- A small set of researchable facets that break the central question down without fragmenting it.
- Alignment
- The visible through-line by which paradigm, tradition, question, data, and analysis imply one another.
- Emergent design
- A design expected to evolve as early data reshape the researcher's understanding.
- Audit trail
- A documented record of design decisions and their rationale that makes an emergent study transparent.
- Decision log
- A running record of substantive changes to a study and the reasoning behind them.
Module 2: The Major Traditions
The five most influential qualitative traditions - ethnography, phenomenology, grounded theory, case study, and narrative - each with its own question, data, and analytic logic.
Ethnography: Studying Culture in the Field
- State the aim and defining commitments of ethnography, including fieldwork and the emic/etic distinction.
- Explain what 'thick description' contributes and how participant observation generates it.
Ethnography is the oldest of the qualitative traditions, born in anthropology and carried into sociology, education, and organizational studies. Its object is culture - the shared, learned, and largely taken-for-granted patterns of meaning, practice, and language through which a group makes sense of its world. Its central question is some version of: what is it like to be a member of this group, and what tacit knowledge organizes their way of life?
Key idea: Ethnography licenses claims about culture - about shared, patterned, learned meaning - and about how practice is accomplished in situ. It does not license claims about individual psychology, about prevalence, or about what would happen if conditions changed. When a dissertation says "an ethnographic study of nurses' burnout" but reports twelve interviews with no time in the ward, the label has been borrowed and the claim cannot be honored.
Fieldwork as the method
Ethnography's signature is prolonged fieldwork: the researcher enters a setting and remains, often for months or years, through participant observation - simultaneously taking part in the group's activities and systematically observing them. Bronislaw Malinowski's insistence that the ethnographer live among the people studied, learn the language, and grasp "the native's point of view" set the template still followed today. The extended stay is not incidental; it is what allows the tacit and the routine - the things members never think to mention because they are obvious to them - to become visible.
Participation is a continuum rather than a single stance, and naming your position on it is part of the method. Raymond Gold's classic typology runs from complete participant (a full member whose research role is concealed, now rarely defensible ethically) through participant as observer (a genuine participant whose research role is known) and observer as participant (present and known, participating minimally) to complete observer (present without interaction). Most contemporary fieldwork sits in the middle two. The choice is consequential: deeper participation buys tacit knowledge and rapport at the cost of analytic distance and of a heavier reflexive burden, while lighter participation preserves distance but leaves you unable to feel what members feel.
Alongside observation, ethnographers conduct ethnographic interviews, which James Spradley distinguished from ordinary research interviews by their embeddedness in the setting and their use of specific question types. Descriptive questions invite a member to walk you through a routine ("Could you take me through a whole shift, from arriving to leaving?"). Structural questions probe how members organize their knowledge ("What kinds of patients are there, from your point of view?"). Contrast questions sharpen the boundaries of a category ("What is the difference between a difficult patient and a demanding one?"). Those three types, used in sequence, do more to surface a cultural taxonomy than any amount of open-ended prompting.
Emic and etic
Two complementary vantage points structure ethnographic analysis. The emic perspective is the insider's: the categories, meanings, and distinctions that members themselves use. The etic perspective is the outsider's, analytic frame: the concepts the researcher brings to compare this culture with others and to build theory. Good ethnography holds both. A study that reports only emic accounts becomes journalism; one that imposes only etic categories misses the meanings that make the practice intelligible to those who live it. The art is to render the insider's world faithfully and then step back to interpret it.
Thick description
The interpretive turn in ethnography, associated with Clifford Geertz, reframed the goal as thick description. Geertz's famous illustration contrasts a wink with an identical twitch of the eyelid: a thin description records the physical movement; a thick description conveys what the wink means - conspiracy, parody, rehearsal of a parody - within a web of cultural convention. The same contraction of muscle carries entirely different social meanings, and only thick description captures them. Thick description is therefore not merely detailed; it is interpretive detail that situates action within the meanings that make it sensible.
Contemporary forms
- Critical ethnography studies culture with an explicit eye to power and inequality, aiming to expose and challenge domination rather than only to describe.
- Autoethnography turns the lens on the researcher's own experience, using systematic self-reflection connected to wider cultural patterns; done well it is analytic, not merely confessional.
- Institutional and organizational ethnography studies the cultures of hospitals, firms, schools, and agencies, where the field is a workplace rather than a village.
- Focused (or rapid) ethnography compresses fieldwork by narrowing to a specific practice, exploiting the researcher's prior insider knowledge, and using intensive short bursts with team-based analysis. It is a legitimate adaptation for applied settings, but it must be named as such rather than passed off as classical fieldwork.
- Multi-sited and digital ethnography follow a phenomenon across locations or into online settings, on the argument that many contemporary cultural formations are not contained in a single place.
Site selection deserves the same care as participant sampling in interview studies, and it follows a similar logic. You choose a site because it is typical of a class of settings, because it is extreme or deviant in a way that makes an ordinarily invisible process visible, because it is critical in the sense that if the phenomenon does not appear here it will appear nowhere, or because it is convenient - and if the last, you owe the reader an honest account of what convenience selected for. Within a site, sampling continues: which shifts, which meetings, which corridors, which members. Ethnographers sample time and events, not only people, and a defense that names those decisions is markedly stronger than one that says only "I spent six months in the ward."
Across all its forms, ethnography trades breadth for depth. It cannot tell you how a practice varies across a nation, but it can tell you, from the inside and in context, what that practice means and how it is accomplished - knowledge that no survey could reach.
A worked illustration: from fieldnote to cultural claim
Here is a raw jotting from a hospital ward, expanded the same evening into a full fieldnote:
"14:10. Handover in the corridor, not the office. Sr. K stands with her back to the med room door, blocking it. Junior nurse B starts to describe bed 6 and Sr. K cuts in: 'Just tell me if he's a worry.' B pauses, restarts: 'He's not a worry but he's not right.' Sr. K nods, writes nothing, moves on. Total time on bed 6: eleven seconds. Later I ask B what 'not right' means. She says: 'You can't put it in the notes, but you know.'"
A thin description here would read: handovers are brief and informal. The ethnographic work begins with the details that a stranger would not notice. Why the corridor? Why the physical blocking of the med room? Why does "a worry" function as a sufficient category while a full clinical account is treated as excessive? Why is "not right" sayable aloud but not writable? An emic reading recovers the members' own taxonomy of patients - worry, not a worry, not right - and notes that these categories do work the formal risk-assessment tool does not. An etic reading connects that taxonomy to the literature on tacit clinical knowledge and to the mismatch between documented and undocumented systems of alertness. The thick description then situates the eleven seconds inside a web of meaning: brevity is not carelessness but a display of competence, since needing detail marks a nurse as not yet able to read a ward.
Note the discipline required. The claim "nurses maintain an informal risk taxonomy that runs alongside the documented one" is a cultural claim, and it needs the same kind of evidence a court would want: repeated instances across shifts and staff, disconfirming instances (the handovers where full detail was given, and what distinguished them), and members' own commentary on the pattern.
How long is long enough, and what counts as adequate data
Ethnography has no saturation formula. What it has is a set of adequacy questions. Have you been present across the setting's full temporal cycle - nights, weekends, seasonal peaks, the annual audit - rather than only during convenient hours? Have you seen the setting under stress as well as in routine? Can you predict, and then check, what will happen next in a familiar sequence? Have you moved from being told the official version to being told the version members tell each other? Michael Agar's notion of the rich point - the moment when something happens that your existing frame cannot explain - is a useful marker: fieldwork is productive while rich points keep arriving, and is approaching adequacy when new observations mostly confirm the frame you have built.
The crisis of representation and its consequences
Ethnography underwent a self-examination in the 1980s, often dated to Writing Culture (Clifford and Marcus, 1986), which argued that ethnographic texts are constructed literary artifacts rather than transparent windows, and that the authority of the detached ethnographic voice concealed the author's position and power. Three practical consequences followed and are now expected in doctoral work: the researcher's positionality is written into the text rather than hidden behind the passive voice; participants' own interpretations are represented rather than only summarized; and the account acknowledges its partiality instead of claiming to be the description of the culture. Critical ethnography and autoethnography are in part responses to that crisis.
Common misconceptions
Ethnography just means qualitative fieldwork. It means the study of culture. Observation alone does not make a study ethnographic; the analytic object must be shared, patterned meaning.
"Going native" is the main risk. Over-identification is a real hazard, but the more common failure in doctoral ethnography is the opposite - too little time in the field to see past the official account, producing what critics call "blitzkrieg ethnography."
Autoethnography is self-indulgent by nature. Weak autoethnography is confessional; strong autoethnography uses the researcher's experience as a systematically analyzed case connected to wider cultural patterns, with the same demands for evidence and analytic distance as any other design.
Recap
- Ethnography's object is culture; its method is prolonged participant observation supplemented by ethnographic interviewing and documents.
- Gold's participation continuum forces you to state and defend your stance in the field.
- Spradley's descriptive, structural, and contrast questions surface members' own taxonomies.
- Thick description in Geertz's sense is interpretive, not merely detailed; the wink and the twitch are physically identical.
- Adequacy is judged by temporal coverage, exposure to stress and routine, and the drying up of rich points, not by a saturation count.
- Post-Writing Culture ethnography writes in the researcher's position and acknowledges the account's partiality.
Sources
- Reeves, S., Kuper, A., & Hodges, B. D. (2008). Qualitative research methodologies: Ethnography. BMJ, 337, a1020. pubmed.ncbi.nlm.nih.gov
- Savage, J. (2000). Ethnography and health care. BMJ, 321(7273), 1400-1402. pmc.ncbi.nlm.nih.gov
- Goodson, L., & Vassar, M. (2011). An overview of ethnography in healthcare and medical education research. Journal of Educational Evaluation for Health Professions, 8, 4. pmc.ncbi.nlm.nih.gov
- Kawulich, B. B. (2005). Participant observation as a data collection method. Forum: Qualitative Social Research, 6(2). qualitative-research.net
- Mulhall, A. (2003). In the field: Notes on observation in qualitative research. Journal of Advanced Nursing, 41(3), 306-313. pubmed.ncbi.nlm.nih.gov
- Geertz, C. (1973). Thick description: Toward an interpretive theory of culture. In The interpretation of cultures (pp. 3-30). Basic Books. find source ↗
- Hammersley, M., & Atkinson, P. (2019). Ethnography: Principles in practice (4th ed.). Routledge. find source ↗
- Key terms
- Ethnography
- The study of culture through prolonged fieldwork, aiming to render a group's shared way of life.
- Participant observation
- Simultaneously taking part in a group's activities and systematically observing them.
- Emic perspective
- The insider's categories and meanings as used by members of the culture themselves.
- Etic perspective
- The outsider's analytic frame that the researcher brings to compare cultures and build theory.
- Thick description
- Interpretive detail that situates action within the cultural meanings that make it sensible.
- Autoethnography
- Ethnography that analyzes the researcher's own experience in connection with wider cultural patterns.
Phenomenology: The Structure of Lived Experience
- State the aim of phenomenology and distinguish its descriptive and interpretive (hermeneutic) branches.
- Explain bracketing, the lifeworld, and the search for the essence of an experience.
Phenomenology asks a question no other tradition asks so directly: what is the essence of a particular lived experience? Not how common it is, not what causes it, but what it is like - the fundamental structure of the experience as it is lived, before we theorize about it. Rooted in the philosophy of Edmund Husserl and developed by Heidegger, Merleau-Ponty, and others, it has become a major method for studying experiences such as living with chronic pain, grieving a spouse, or being a caregiver.
Key idea: Phenomenology licenses claims about the structure of experience - what must be present for this to be the experience it is. It does not license claims about why the experience arises, how it varies across social groups, or what should be done about it. A phenomenological study that concludes with recommendations for service redesign has left its warrant behind, however welcome the recommendations may be.
Why the philosophy cannot be skipped
More than any other tradition, phenomenology is abused by borrowing its name for what is really thematic analysis of interview data. The tell is easy to spot: a study announces a phenomenological design, then reports six themes with illustrative quotations under each. Themes are categories of content; a phenomenological finding is a structure of experience. The two are not the same thing, and reviewers who know the tradition will say so. If your analysis produces themes, call it thematic analysis - a perfectly respectable choice - and drop the philosophical apparatus.
The phenomenological attitude
Phenomenology begins from the lifeworld - the world as we immediately and pre-reflectively experience it, prior to scientific abstraction. The aim is to describe phenomena as they present themselves to consciousness. Central to this is the concept of intentionality: consciousness is always consciousness of something; experience is always directed at objects, and phenomenology studies that relationship between the experiencing person and what is experienced.
Bracketing (epoche)
To reach the experience itself, Husserl argued the researcher must set aside - bracket - their preconceptions, prior theories, and even the assumption that the object exists in the ordinary sense, so that the phenomenon can appear on its own terms. This suspension is called the epoche. Practically, in descriptive (Husserlian) phenomenology, bracketing is a discipline: the researcher writes out their assumptions in advance and consciously holds them in abeyance during data collection and analysis, so that the description is faithful to participants' experience rather than a projection of the researcher's expectations.
Two branches
- Descriptive (Husserlian) phenomenology seeks the essence - the invariant structures without which the experience would not be that experience. The analyst reads accounts closely, identifies meaning units, and moves through imaginative variation (mentally removing features to test which are essential) toward a description of the essence common across participants.
- Interpretive (hermeneutic) phenomenology, following Heidegger, denies that bracketing is fully possible or even desirable. Because we are always already immersed in a world of meaning (Heidegger's Dasein, being-in-the-world), interpretation is unavoidable; the researcher's fore-understanding is a resource to be used reflectively rather than eliminated. Analysis proceeds through the hermeneutic circle, moving iteratively between parts (specific passages) and the whole (the emerging overall meaning), each revising the other.
Within these branches sit named procedures you should be able to distinguish. Amedeo Giorgi's descriptive phenomenological psychological method proceeds in explicit steps: read the whole for a sense of it, mark meaning units at points where the meaning shifts, transform each unit from the participant's everyday language into disciplinary language while remaining faithful to it, and synthesize a general structure. Clark Moustakas's transcendental approach adds horizontalization (treating every statement as initially of equal value), clustering into themes, and the construction of textural (what was experienced) and structural (how it was experienced, the conditions of its appearing) descriptions, which are then combined into a composite. Max van Manen's hermeneutic phenomenology of practice orients analysis around four existentials - lived body, lived space, lived time, and lived human relation - and treats the writing itself as the method rather than as reportage. Interpretative phenomenological analysis (IPA), developed by Jonathan Smith and colleagues, is idiographic: it analyzes each case in depth before moving across cases, and it is explicit that the researcher is making sense of the participant making sense of their experience - a "double hermeneutic."
Choosing among them is not a matter of taste. Giorgi's method suits a study seeking a single general structure from participants who share a well-defined phenomenon. Van Manen suits a study whose contribution is an evocative account that changes how practitioners see something familiar. IPA suits small, homogeneous samples where individual particularity is the point, and its typical sample of six to ten is a feature of the design, not a limitation to apologize for.
What phenomenology delivers and demands
Done well, phenomenology yields a description so faithful that a reader who has undergone the experience recognizes it, and one who has not begins to understand it from within. This demands a particular kind of data: rich, first-person accounts of concrete experience ("describe a specific time when..."), not opinions or generalizations. It also demands unusual analytic restraint, because the constant temptation is to explain or categorize the experience rather than to describe it. The discipline of staying with the experience - rendering its texture rather than its causes - is what makes the tradition distinctive and difficult.
Sampling and interviewing for a phenomenological study
Sampling is criterion sampling in its purest form: every participant must have lived the phenomenon, and homogeneity on that criterion matters more than variation on demographics. A study of the experience of being told a cancer diagnosis by telephone should not dilute itself with participants told in person, however tempting the comparison, because the essence being sought belongs to the phenomenon as specified. Recruit for the experience, not for the category of person.
The interview itself is unusual. It is closer to guided remembering than to questioning. Useful moves include anchoring in a single episode ("Take me to the moment the phone rang"), asking for sensory and temporal detail ("Where were you standing? What did you do with your hands?"), following silence rather than filling it, and resisting the interviewer's own interpretive offers. Two habits ruin phenomenological data: asking "why," which converts experience into explanation, and asking "generally, how do you find it," which converts experience into summary. Some phenomenologists supplement interviews with written accounts, diaries, or elicited descriptions from literature and art, on van Manen's argument that experiential material need not come only from participants.
A worked illustration: from excerpt to meaning unit to structure
A participant in a study of living with early-onset Parkinson disease says:
"The worst part isn't the shaking. It's that I plan my whole day around a door. Whether there'll be a door I have to open while holding something. I used to be someone who just went places. Now I go places I have already been."
A thematic analyst would code this as "loss of spontaneity" and move on. A phenomenological analyst does something different. First, mark the shifts in meaning: (1) a rejection of the obvious symptom as the phenomenon; (2) the door as the organizing object of a day; (3) a contrast between two selves, one who "just went places" and one who returns to known ground.
Then transform each into disciplinary language without adding explanation. Unit 2 becomes: the lived world contracts around anticipated points of bodily failure, so that the environment is experienced in advance as a series of obstacles rather than encountered as it arrives. Unit 3 becomes: the future is experienced as repetition of the already-known, so that lived time loses its openness. Notice that both transformations name structures of experience - the pre-lived environment, the closing of futurity - rather than topics.
Now test with imaginative variation. Could this be the same experience if the environment were fully accessible? If the answer is yes for some participants, then "obstacle anticipation" is not invariant and belongs to a variant rather than the essence. Could it be the same experience without the contrast between past and present selves? Across accounts, if every participant renders the phenomenon through a before-and-after self, the structure is a candidate for the essence. In van Manen's vocabulary you would say the phenomenon shows up primarily through lived space (the door) and lived time (the closed future), and secondarily through lived body.
The resulting finding is not "participants experienced a loss of spontaneity." It is closer to: living with early-onset Parkinson disease is experienced as an inversion of the ordinary relation between body and world, in which the world is pre-lived as a field of anticipated failures and the future is lived as return rather than departure. That sentence is what phenomenology exists to produce.
Bracketing in practice, and the honest critique of it
Descriptive phenomenologists operationalize the epoche through a written pre-understandings statement, a bracketing journal kept throughout the study, and interviewing that resists the urge to interpret aloud. Critics - hermeneutic phenomenologists among them - argue that bracketing is at best partial and at worst a claim to a neutrality nobody possesses, and that a researcher's fore-understanding is precisely what makes a phenomenon visible in the first place. Both positions are defensible; what is not defensible is claiming to bracket while writing an analysis saturated with an unexamined theoretical frame. Whichever branch you choose, the reader needs to see the researcher's prior understanding on the page - either as what was set aside, or as what was put to work.
Common misconceptions
Phenomenology means asking people how they feel. It means asking for concrete, pre-reflective description. "How did that make you feel?" invites a summary judgment; "Take me back to that moment - where were you, what did you notice first?" invites the experience.
Larger samples make a phenomenology stronger. Beyond a point they weaken it, because the analytic depth per case falls. IPA studies commonly use six to ten participants, and single-case IPA is published.
Essence means a universal truth about all humans. The essence is invariant with respect to this phenomenon as experienced by those who have lived it, not a metaphysical universal. Overclaiming here is the fastest route to a hostile review.
Recap
- Phenomenology seeks the structure of lived experience; its findings are structures, not themes.
- Husserl's descriptive branch aims at essence through bracketing and imaginative variation; Heidegger's interpretive branch treats fore-understanding as an unavoidable resource worked through the hermeneutic circle.
- Giorgi, Moustakas, van Manen, and IPA are distinct procedures with different outputs; name the one you are using and follow it.
- Data must be concrete first-person description of specific moments, not opinion or generalization.
- Small homogeneous samples are appropriate; depth per case is the currency.
Sources
- Neubauer, B. E., Witkop, C. T., & Varpio, L. (2019). How phenomenology can help us learn from the experiences of others. Perspectives on Medical Education, 8(2), 90-97. pmc.ncbi.nlm.nih.gov
- Starks, H., & Trinidad, S. B. (2007). Choose your method: A comparison of phenomenology, discourse analysis, and grounded theory. Qualitative Health Research, 17(10), 1372-1380. pubmed.ncbi.nlm.nih.gov
- Smith, D. W. (2018). Phenomenology. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Beyer, C. (2020). Edmund Husserl. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Wheeler, M. (2020). Martin Heidegger. Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Aguas, P. P. (2022). Fusing approaches in educational research: Data collection and data analysis in phenomenological research. The Qualitative Report, 27(1), 1-20. nsuworks.nova.edu
- van Manen, M. (1990). Researching lived experience: Human science for an action sensitive pedagogy. State University of New York Press. find source ↗
- Key terms
- Phenomenology
- The study of the essential structure of a lived experience as it presents itself to consciousness.
- Lifeworld
- The world as immediately and pre-reflectively experienced, prior to scientific abstraction.
- Bracketing (epoche)
- Setting aside one's preconceptions so a phenomenon can appear on its own terms.
- Essence
- The invariant structures without which an experience would not be that experience.
- Hermeneutic circle
- Iterative interpretation moving between parts and the whole, each revising the other.
- Fore-understanding
- The prior understanding a hermeneutic researcher brings and uses reflectively rather than eliminating.
Grounded Theory: Building Theory from Data
- State the aim of grounded theory and describe its core procedures at a design level.
- Distinguish the Glaser, Strauss and Corbin, and Charmaz variants and their assumptions.
Grounded theory is the tradition whose explicit goal is to generate theory - a set of concepts and their relationships that explains a social process - directly from data, rather than to test a theory imported from elsewhere. Introduced by Barney Glaser and Anselm Strauss in 1967, it was in part a rebellion against a sociology that endlessly tested grand theories while producing few new ones. Its guiding question is: what is the main process going on here, and how do participants manage or resolve it?
The founding text, The Discovery of Grounded Theory (Glaser and Strauss, 1967), grew out of the authors' fieldwork on dying in American hospitals, published as Awareness of Dying (1965). That study did not report themes about death; it produced a theory of awareness contexts - closed, suspected, mutual pretense, and open - and explained how staff and patients moved between them and what each context made possible. Read that sentence again as a template. A grounded theory names a core problem participants are working on, names the strategies by which they work on it, and specifies the conditions under which strategies shift.
Key idea: Grounded theory licenses claims about process and its explanation at a conceptual level. It does not license claims about lived essence (that is phenomenology), about culture (ethnography), or about prevalence. Its output is a theory that could, in principle, be tested by other means.
The core procedures
Grounded theory is defined less by a topic than by a distinctive package of interlocking procedures. This lesson introduces them at the level of design; a later analysis lesson works through the coding in detail.
- Constant comparison. Every new piece of data is continually compared with existing data and emerging categories - incident to incident, incident to category, category to category - so that concepts are refined and their properties clarified as analysis proceeds. Analysis and data collection are interwoven, not sequential.
- Theoretical sampling. Who or what to study next is decided by the emerging theory, not fixed in advance. Having developed a tentative category, the researcher seeks new cases that will test, elaborate, or challenge it. Sampling is thus directed by analytic need.
- Theoretical saturation. Sampling and analysis continue until new data no longer yield new properties of the categories - the point of saturation, which signals that the category is sufficiently developed.
- Memoing. Throughout, the analyst writes analytic memos capturing ideas about codes and their relationships. Memos are the bridge from coding to theory; the written theory is substantially assembled from them.
Three variants
Grounded theory split into distinct schools, and a doctoral researcher must state which they follow.
| Variant | Epistemology | Distinctive emphasis |
|---|---|---|
| Glaser (classic) | Objectivist, discovery | Theory 'emerges'; avoid forcing data into preconceived frames, including elaborate coding paradigms |
| Strauss and Corbin | Post-positivist | Structured procedures, including a coding paradigm (conditions, actions, consequences) and axial coding |
| Charmaz (constructivist) | Constructivist | Theory is co-constructed; the researcher's role in shaping data and categories is acknowledged, not hidden |
The Glaser-Strauss split was substantive: Glaser objected that Strauss and Corbin's structured paradigm forced data into a preset template rather than letting categories emerge. Kathy Charmaz later reframed the whole enterprise: since the researcher is inevitably part of what is studied, the resulting theory is an interpretive construction, not a discovery of something lying in wait. Constructivist grounded theory has become especially influential because it reconciles the method's rigor with contemporary skepticism about a neutral observer.
The split is worth understanding in more detail than a table can convey, because doctoral committees ask about it. Glaser and Strauss came from opposite intellectual worlds - Glaser from Columbia's quantitative tradition under Paul Lazarsfeld, Strauss from Chicago symbolic interactionism under Herbert Blumer - and their joint method concealed a tension that surfaced when Strauss and Juliet Corbin published Basics of Qualitative Research in 1990. Glaser's reply, Basics of Grounded Theory Analysis: Emergence vs. Forcing (1992), argued that the new coding paradigm - specifying causal conditions, context, intervening conditions, action strategies, and consequences for every category - was a preconceived template that produced "full conceptual description" rather than emergent theory. Udo Kelle's later analysis of the dispute makes the deeper point that neither position is coherent in its strong form: pure emergence is impossible because no researcher approaches data without concepts, while a fixed paradigm risks exactly the forcing Glaser named. What resolves the tension in practice is Herbert Blumer's notion of sensitizing concepts - prior ideas that suggest where to look without dictating what will be found - held loosely and revised against data.
Charmaz's constructivist turn changes what the theory claims to be. Where Glaser writes of categories that emerge from data as though the analyst were uncovering them, Charmaz argues that data are constructed in the interaction and that categories are constructed by an analyst positioned in a time, place, and biography. Practically, constructivist grounded theory keeps constant comparison, theoretical sampling, and memoing, but adds explicit reflexivity, prefers gerunds in coding to keep action in view, and reports the theory as an interpretation offered from a standpoint. Antony Bryant and Adele Clarke's situational analysis extend the tradition further toward postmodern and mapping-based approaches.
A fourth position deserves mention. Abductive analysis, articulated by Stefan Timmermans and Iddo Tavory, argues that theory generation actually works neither by induction nor by hypothesis testing but by abduction: an anomalous finding puzzles a researcher who is deeply read in theory, and the creative act is proposing the explanation that would, if true, make the surprise unsurprising. On this view, extensive theoretical knowledge is an asset rather than a contaminant - a direct challenge to the Glaserian injunction to delay the literature review.
What counts as a result
The product of a grounded-theory study is not a set of themes but an explanatory framework: a core category and the related concepts that together account for how a process unfolds and is managed. If a study labeled "grounded theory" ends with a list of descriptive themes and no integrated theory of a process, it has used the coding techniques without delivering the tradition's defining outcome. The test is whether you have produced a theory that explains, at a conceptual level, what is going on.
A worked illustration: theoretical sampling in motion
A study begins with the question of how family carers of people with dementia decide to seek residential care. The first six interviews are with carers who made the move within the past year.
After interview 4, a memo records a candidate category: holding the line - carers describe a threshold they had privately set ("as long as she knows me," "as long as he isn't wandering at night") and describe the decision as the moment the line moved rather than the moment it was crossed. This is more interesting than the researcher's original hunch about caregiver burden scores.
Theoretical sampling decision 1. If holding the line is a real process, it should be visible in carers who have not yet placed a relative. The researcher recruits four carers still at home. Two describe explicit lines; two deny having any, saying they will "know when." The category gains a property: articulacy of the threshold, ranging from explicit to inchoate.
Theoretical sampling decision 2. A negative case appears - a carer whose relative was admitted after a hospital fall, with no decision at all. Rather than discarding it, the analyst asks what it shows: the process of holding the line can be pre-empted by an external event, and those carers describe a distinctive aftermath of guilt precisely because they were denied the decision. The category is now conditional: holding the line is the process when the timing is the carer's to control, and its absence generates a different trajectory.
Theoretical sampling decision 3. Do lines move differently when a second family member is involved? The researcher recruits three carers with siblings actively involved, and finds a new strategy - outsourcing the line - in which the decision is displaced onto a professional's recommendation so that no family member owns it.
The emerging core category is not "caregiver burden" but something like managing the ownership of an unbearable decision, with holding the line, pre-emption, and outsourcing as its modes. Notice three things about how this happened. Sampling was driven by the analysis at each step. The negative case sharpened rather than threatened the theory. And the category is stated as a process with conditions, which is what distinguishes a theory from a theme.
Common misconceptions
Grounded theory means you must not read the literature first. This is the Glaserian position, and it is contested. Strauss and Corbin, Charmaz, and the abductive analysts all treat prior reading as a resource. In practice, doctoral programs require a literature review anyway; the workable discipline is to read widely, write down what you expect to find, and treat those expectations as sensitizing concepts to be tested rather than as findings.
Any inductive coding is grounded theory. No. Without theoretical sampling, constant comparison, memoing, and an integrated core category, what you have is qualitative content or thematic analysis using some grounded-theory vocabulary. Reviewers call this "grounded theory lite," and it is one of the most common criticisms of published work claiming the label.
Saturation is a number. Theoretical saturation is a property of categories, not of a sample. You can interview forty people and leave a category thin, or reach adequate development of a well-specified category in fifteen. The claim must be made category by category.
The theory must be grand. Glaser and Strauss distinguished substantive theory, which explains a process in a specific area, from formal theory, which abstracts across areas. Nearly all doctoral grounded theory is substantive, and that is the appropriate ambition.
Recap
- Grounded theory generates an explanatory theory of a social process; its output is a core category with conditions and strategies, not a theme list.
- Constant comparison, theoretical sampling, theoretical saturation, and memoing are the interlocking procedures that define it.
- Glaser emphasizes emergence, Strauss and Corbin structured procedures and a coding paradigm, Charmaz co-construction and reflexivity; state which you follow and follow it.
- Kelle's analysis of emergence versus forcing, and Blumer's sensitizing concepts, offer a workable middle position.
- Abductive analysis reframes theory generation as inference to the best explanation, making prior theoretical reading an asset.
Sources
- Chun Tie, Y., Birks, M., & Francis, K. (2019). Grounded theory research: A design framework for novice researchers. SAGE Open Medicine, 7. pmc.ncbi.nlm.nih.gov
- Kelle, U. (2005). "Emergence" vs. "forcing" of empirical data? A crucial problem of "grounded theory" reconsidered. Forum: Qualitative Social Research, 6(2). qualitative-research.net
- Charmaz, K., & Keller, R. (2016). A personal journey with grounded theory methodology: Kathy Charmaz in conversation with Reiner Keller. Forum: Qualitative Social Research, 17(1). qualitative-research.net
- Heath, H., & Cowley, S. (2004). Developing a grounded theory approach: A comparison of Glaser and Strauss. International Journal of Nursing Studies, 41(2), 141-150. pubmed.ncbi.nlm.nih.gov
- Charmaz, K. (1990). "Discovering" chronic illness: Using grounded theory. Social Science & Medicine, 30(11), 1161-1172. pubmed.ncbi.nlm.nih.gov
- Foley, G., & Timonen, V. (2015). Using grounded theory method to capture and analyze health care experiences. Health Services Research, 50(4), 1195-1210. pmc.ncbi.nlm.nih.gov
- Glaser, B. G., & Strauss, A. L. (1967). The discovery of grounded theory: Strategies for qualitative research. Aldine. find source ↗
- Key terms
- Grounded theory
- A tradition that generates explanatory theory of a social process directly from data.
- Constant comparison
- Continually comparing new data with existing data and categories to refine concepts.
- Theoretical sampling
- Selecting subsequent cases based on what the emerging theory needs, not a fixed plan.
- Theoretical saturation
- The point at which new data yield no new properties of the categories.
- Analytic memo
- Written reflection on codes and their relationships that bridges coding and theory.
- Core category
- The central concept around which a grounded theory's explanatory framework is integrated.
Case Study: Bounded Systems in Depth
- Define the case as a bounded system and distinguish single from multiple and intrinsic from instrumental case studies.
- Explain analytic (not statistical) generalization and the role of multiple evidence sources.
A case study is an in-depth investigation of a bounded system - a single instance of a phenomenon studied in its real-life context, using multiple sources of evidence. The case might be a person, a program, an event, a decision, a classroom, an organization, or a policy. What unites case studies is not a data-collection technique but the object: a specific, bounded case examined intensively and holistically.
Key idea: "Case study" is the most loosely used label in qualitative research, and that looseness is a hazard. A study is not a case study because it has few participants, nor because it took place in one hospital. It is a case study when the case itself - the bounded system - is the unit of analysis, when it is studied in context, and when evidence from multiple sources is brought to bear on it. Nerida Hyett and colleagues reviewed published qualitative case studies and found that many failed to define the case, to justify the design methodologically, or to use more than one data source; the label had been applied to what were really small interview studies.
Defining and bounding the case
The first analytic act is to bound the case: to say clearly what is inside it and what is context, and over what time period. "The implementation of a new triage protocol in one emergency department during its first year" is bounded by place, phenomenon, and time. Without clear boundaries, a case study sprawls into an unfocused description of everything. Robert Stake and Robert Yin, the two most cited methodologists here, agree on this even as they differ in emphasis - Stake more interpretive and holistic, Yin more structured and drawn toward propositions and rival explanations.
Bounding has a second, subtler function: it fixes the unit of analysis, and confusing the unit is a recurrent failure. If the case is the emergency department, then individual clinicians are sources of evidence about the department, and a finding is a statement about the department. If the case is each clinician, then you have a multiple-case design with many cases and the department is context. Yin adds the notion of embedded units: a single case (the hospital) may contain sub-units (three wards) analyzed in their own right, provided the analysis returns to the case level. A study that collects data from embedded units and then reports only cross-cutting themes about individuals has, in Yin's phrase, drifted from its unit of analysis.
Yin and Stake: two different animals under one name
Robert Yin's approach is broadly post-positivist. It begins from a study proposition - a tentative explanatory statement derived from theory - and treats the case study as a design that can test it. Its apparatus includes a written case study protocol, a chain of evidence linking question to data to conclusion, a case study database separate from the report, explicit consideration of rival explanations, and analytic techniques such as pattern matching (does the observed pattern match the predicted one?), explanation building, and time-series analysis. Multiple cases follow a replication logic: literal replication predicts similar results for theoretically similar cases, and theoretical replication predicts contrasting results for predictable reasons. Yin speaks comfortably of construct validity, internal validity, external validity, and reliability, and prescribes tactics for each.
Robert Stake's approach is constructivist. The case is a functioning system to be understood in its particularity; the researcher's job is interpretation, and the report aims at naturalistic generalization - the reader's tacit recognition when a richly rendered case resonates with their own experience. Stake organizes inquiry around issues rather than propositions, uses "issue questions" to keep the study focused on tensions within the case, and treats vicarious experience conveyed through narrative as a legitimate contribution. Where Yin wants triangulation to converge on what happened, Stake wants triangulation to check that an interpretation is not idiosyncratic.
The practical consequence is that these two cannot be blended casually. A proposal that cites Stake for its philosophy and then promises pattern matching against propositions and inter-rater reliability has fused two incompatible logics. Sharan Merriam offers a third position, closer to Stake but with a more explicit analytic procedure, and Helena Harrison and colleagues provide a useful account of how the orientations diverged.
Types of case study
- An intrinsic case study is undertaken because this particular case is of interest in its own right - a unique program, an unusual patient. The goal is to understand the case itself, not to generalize from it.
- An instrumental case study uses the case as a means to illuminate a broader issue or theory. The case is chosen because it offers analytic leverage on a question that extends beyond it.
- A collective (multiple) case study examines several cases to compare and contrast, strengthening analytic claims by showing how a process plays out across contexts.
Multiple sources and convergence
A hallmark of strong case study is the use of multiple sources of evidence - interviews, documents, observations, records, artifacts - that are brought into conversation. When independent sources converge on the same interpretation (triangulation), confidence rises; when they diverge, the discrepancy itself becomes a finding to be explained. This is what gives a good case study its density and credibility: the account is corroborated from several angles rather than resting on a single informant's word.
Analytic, not statistical, generalization
The most common misunderstanding of case study is the objection "you cannot generalize from n = 1." That objection assumes statistical generalization - inferring frequencies from a sample to a population - which case study never claims. What case study offers instead is analytic generalization: the findings speak to and refine theory, so that lessons transfer to other cases the theory covers, not to a population by frequency.
A single, carefully chosen case can falsify a theory that claims universality (one black swan is enough), or extend a concept to a new context, or generate an explanation later tested more broadly. Yin's logic is explicit here: a case study is generalizable to theoretical propositions, in the way an experiment is, and not to populations, in the way a survey is. Judged by that correct standard, the depth of a single bounded system is a strength, not a defect.
Bent Flyvbjerg's widely cited article dismantles five misunderstandings about case study research, of which two matter most here. The first is that context-independent knowledge is more valuable than context-dependent knowledge; Flyvbjerg argues that in the study of human affairs, context-dependent knowledge is what expertise actually consists of. The second is that one cannot generalize from a single case; he shows that single cases have repeatedly overturned general propositions - Galileo's refutation of Aristotelian free fall did not require a random sample - and that the force of a case depends on how it was selected. A critical case ("if it does not hold here, it holds nowhere") and a deviant case (an outlier whose mechanism illuminates the ordinary) carry far more inferential weight than a case chosen because access was easy.
A worked illustration: bounding and sourcing one case
A doctoral student wants to study "why our university's peer mentoring program failed." Watch the design decisions.
Bounding. The case is defined as the peer mentoring program in one college of the university, from its approval in March of year one to its suspension in June of year two. Inside the case: the program's design documents, its staff, its mentors and mentees, its budget line, its steering group. Outside, as context: the university's wider student success strategy, the national policy environment, other colleges' programs.
Unit of analysis. The program, not the students. A mentee interview is evidence about the program.
Type. Instrumental, since the interest is in what the case reveals about how peer support programs fail, with the program itself of secondary interest.
Propositions and rivals (Yin) or issues (Stake). Working in Yin's idiom, the student states a proposition drawn from implementation literature: programs fail when the role is ambiguous to the people performing it. Then, crucially, she writes down rival explanations to be tested rather than assumed away - insufficient funding, poor recruitment timing, leadership turnover, and the possibility that the program did not fail at all but was defunded for unrelated budget reasons.
Sources. Steering group minutes across sixteen months; the two versions of the mentor role description; training slides; the enrollment and attendance spreadsheet; eleven interviews across mentors, mentees, coordinators, and the dean; observation of two remaining sessions.
Convergence and divergence. Minutes and interviews converge on ambiguity: the role description changed in month five without retraining, and mentors describe drifting into either counseling or tutoring depending on temperament. But the attendance data diverge from the staff account - attendance was stable until month eleven, when it collapsed in a single month, which nobody mentioned. Chasing that divergence uncovers a room change that moved sessions off the main campus. The final explanation is therefore layered: chronic role ambiguity weakened commitment, and an acute logistical shock finished it. Neither source alone would have produced that account, and the honest reporting of the divergence is what makes the case persuasive.
Common misconceptions
Case study is a data-collection method. It is a design and a choice of unit of analysis. Interviews, documents, and observation are the methods used within it.
More cases are always better. In multiple-case designs, cases are selected for replication logic, not sampled for representativeness, and each additional case costs depth. Two or three well-contrasted cases usually beat eight thin ones.
Triangulation means agreement. Divergence between sources is information, not failure. The strongest case studies report where sources disagreed and what explained it.
A case study cannot test anything. Under Yin's logic it can: pattern matching against a stated proposition, with rivals specified in advance, is a genuine test, and a critical case can falsify a universal claim.
Recap
- The case is a bounded system, and bounding fixes the unit of analysis; embedded sub-units must be returned to the case level.
- Yin's post-positivist design uses propositions, protocols, rival explanations, pattern matching, and replication logic; Stake's constructivist design uses issues, interpretation, and naturalistic generalization.
- Intrinsic, instrumental, and collective designs answer different questions.
- Multiple sources of evidence are constitutive; convergence raises confidence and divergence generates findings.
- Case studies generalize analytically to theory; critical and deviant cases carry the most inferential weight.
Sources
- Crowe, S., Cresswell, K., Robertson, A., Huby, G., Avery, A., & Sheikh, A. (2011). The case study approach. BMC Medical Research Methodology, 11, 100. pmc.ncbi.nlm.nih.gov
- Hyett, N., Kenny, A., & Dickson-Swift, V. (2014). Methodology or method? A critical review of qualitative case study reports. International Journal of Qualitative Studies on Health and Well-being, 9, 23606. pmc.ncbi.nlm.nih.gov
- Harrison, H., Birks, M., Franklin, R., & Mills, J. (2017). Case study research: Foundations and methodological orientations. Forum: Qualitative Social Research, 18(1). qualitative-research.net
- Baxter, P., & Jack, S. (2008). Qualitative case study methodology: Study design and implementation for novice researchers. The Qualitative Report, 13(4), 544-559. nsuworks.nova.edu
- Flyvbjerg, B. (2006). Five misunderstandings about case-study research. Qualitative Inquiry, 12(2), 219-245. arxiv.org
- Schlunegger, M. C., Zumstein-Shaha, M., & Palm, R. (2024). Methodologic and data-analysis triangulation in case studies: A scoping review. Western Journal of Nursing Research, 46(8), 611-622. pmc.ncbi.nlm.nih.gov
- Yin, R. K. (2018). Case study research and applications: Design and methods (6th ed.). Sage Publications. find source ↗
- Key terms
- Case study
- In-depth study of a bounded system in its real-life context using multiple sources of evidence.
- Bounded system
- The specified case with clear limits of what is inside it, what is context, and over what time.
- Intrinsic case study
- A case studied because that particular case is of interest in its own right.
- Instrumental case study
- A case used as a means to illuminate a broader issue or theory.
- Analytic generalization
- Generalizing findings to theory rather than to a population by frequency.
- Triangulation
- Using multiple sources or methods so that convergence raises confidence and divergence prompts inquiry.
Narrative Inquiry: Analyzing Stories
- State the aim of narrative inquiry and distinguish analysis of narratives from narrative analysis.
- Describe major approaches to narrative analysis, including thematic, structural, and dialogic/performance.
Narrative inquiry takes the story as both its data and its object. It rests on a simple but powerful premise: people make sense of their lives and identities by telling stories, and by studying those stories closely we learn how experience is organized, given meaning, and connected across time. The question is some version of: how do people story their experience, and what does the shape of the story reveal?
Key idea: Narrative inquiry licenses claims about how experience is composed and communicated - about form, sequence, positioning, and the cultural resources a teller draws on. It does not license claims that the story is an accurate report of events. A narrative analyst who writes "the participant was mistreated by her supervisor" has confused the account with the world; what the analyst can say is how she constructs an account of mistreatment, and what that construction accomplishes.
Two lineages feed the tradition and they behave differently. The first, associated with D. Jean Clandinin and F. Michael Connelly and rooted in John Dewey's philosophy of experience, treats narrative inquiry as a whole methodology - a way of being with participants over time rather than a technique applied to transcripts. Its signature is the three-dimensional narrative inquiry space: every account is examined along temporality (past, present, future), sociality (personal and social conditions, including the researcher's own), and place (the specific physical and topological settings). Field texts - conversations, journals, photographs, artifacts - are composed with participants and then rendered into research texts, with attention to the relational ethics of representing someone's life. The second lineage, associated with Catherine Kohler Riessman and with sociolinguistics, is a family of analytic techniques for working on narrative material regardless of how it was gathered. Both are legitimate; naming which you are doing prevents a common confusion in which a study borrows Clandinin's relational language while performing Riessman's structural analysis.
Why stories, specifically
A narrative is not just any talk; it is an account with a temporal structure - a beginning, middle, and end - in which events are emplotted, connected causally or thematically into a meaningful whole. Because identity itself is substantially narrative (we are, in part, the stories we tell about who we are), narrative inquiry is especially suited to studying identity, life transitions, illness experience, and the way people make continuity out of disruption. Where phenomenology brackets to reach an essence, narrative inquiry stays with the particular, the temporal, and the emplotted.
A crucial distinction
Methodologists distinguish two things that sound alike:
- Analysis of narratives treats stories as data from which the researcher extracts themes across many accounts - a broadly thematic move that happens to use stories as its raw material.
- Narrative analysis keeps each story intact and asks how it works as a story - its structure, its plot, its function - rather than fragmenting it into cross-cutting codes. Preserving the whole is the point, because meaning lives in the configuration, not only in the parts.
Approaches to analyzing a story
Within narrative analysis several lenses are common, and studies often combine them:
- Thematic: attends to what is told - the content and meaning of the story - while still respecting the account as a whole.
- Structural: attends to how the story is built. A classic scheme (Labov) identifies functional parts - an abstract, orientation, complicating action, evaluation, resolution, and coda - and asks how the teller uses them to make a point.
- Dialogic / performance: attends to the story as a performance produced for an audience in interaction. It asks not only what the story says but what it does - how the teller positions themselves, to whom the story is addressed, and how the listener (often the interviewer) shapes it.
- Visual and multimodal: extends the same questions to photographs, drawings, and objects that participants use to tell, on the view that narrative is not confined to speech.
To these, several typologies of narrative form are worth knowing because they give an analyst something specific to look for. Arthur Frank's account of illness narratives distinguishes the restitution narrative ("I was healthy, I fell ill, I will be healthy again"), the chaos narrative (which has no plot, resists sequence, and is hard to listen to), and the quest narrative (in which illness is taken up as a journey yielding something). Gergen and Gergen distinguish progressive, regressive, and stable narrative arcs by the direction of movement toward a valued end. Michael Bamberg's positioning analysis asks three questions of any telling: how characters are positioned relative to one another within the story, how the teller positions themselves relative to the audience in the act of telling, and what identity claim the telling makes in the wider discursive landscape. Each typology is a lens, not a truth; the analytic value comes from noticing when a story resists the category you brought.
Co-construction and craft
Narrative inquiry is candid that the researcher does not merely collect stories but participates in producing them: the interview is a relationship, and the account is co-constructed between teller and listener. Reporting often preserves long stretches of a participant's own words and may organize the findings around a small number of richly rendered cases rather than many fragments.
This makes narrative inquiry unusually intimate and interpretively demanding: the analyst must honor the integrity of a person's story while still making an analytic argument about it. Done well, it shows not just what happened to someone, but how they came to understand what happened - which is often the deeper finding.
A worked illustration: a Labovian reading
A former nurse, three years after leaving the profession, says:
"I'll tell you the exact moment I knew. (1) It was a Tuesday in February, the second winter of the pandemic, and we were four short on nights. (2) A woman in bay three kept asking for her daughter and I kept saying I'd find out, and I did not find out, because I had eleven other people. Around four in the morning she stopped asking. (3) And I thought, that's it, I have become someone who lets people stop asking. (4) I handed my notice in on the Friday. (5) People say I burned out. I didn't burn out. I saw myself clearly."
A thematic coder would file this under "moral distress" and "workforce attrition." A structural reading using Labov's functional parts finds more. The abstract is the opening line, which tells the listener what kind of story is coming and claims the teller's authority over its meaning. The orientation (1) supplies time, place, and the critical condition of understaffing. The complicating action (2) is spare and cumulative, and note the repeated "kept," which builds duration without stating it. The evaluation (3) is where the teller tells us why the story matters, and it is here, not in the events, that the analytic payload sits: the transformation claimed is not exhaustion but a change in the kind of person she has become. The resolution (4) is one short sentence, deliberately undramatic. The coda (5) returns us to the present and does explicit work against a competing account - "people say I burned out" - so the story is also an argument.
Now layer positioning analysis. Within the story she positions herself against an institution that made a certain kind of failure inevitable, and positions the patient as someone who was abandoned by a system rather than by a person. In the telling, she positions herself against an interviewer she expects to hold the burnout frame, which is why the coda pre-empts it. At the broadest level, the telling claims membership in a moral category - one who left because she saw clearly - rather than a clinical category. Frank's typology helps too: this is neither restitution nor chaos but a quest narrative in which departure is the thing gained.
The analytic finding that emerges is not "nurses left because of staffing." It is closer to: former nurses construct departure as an act of moral perception rather than a failure of endurance, and they do so in explicit opposition to the burnout vocabulary available to them - a positioning that reclaims agency the institutional account denies. That is a claim about narrative work, and it is defensible from the text.
Generating narrative data
The interview that produces usable narrative looks different from the semi-structured interview built around a topic guide. Its opening move is an invitation to tell rather than a question to answer: "Tell me the story of how you came to leave nursing, beginning wherever the story begins for you." The interviewer then does something counterintuitive - stays quiet, and does not interrupt with clarifying questions, because clarification imposes the researcher's sense of what matters and breaks the teller's own emplotment. Fritz Schutze's biographical narrative method formalizes this into three phases: an uninterrupted main narration, then internal questions that ask the teller to expand parts of their own story using their own words, and only then external questions bringing in the researcher's interests. Recording matters more than usual, since pauses, false starts, laughter, and repair are analytic material rather than transcription noise; narrative transcripts are therefore transcribed at a finer grain than thematic ones.
Ethics and craft in narrative work
Narrative inquiry carries a distinctive ethical load because its unit is a recognizable life. Standard anonymization is weak protection when a study reports one person's story at length: colleagues will know who it is. Practical responses include composite construction (with disclosure that composites were used), negotiated redaction with participants, deliberate alteration of non-essential identifying detail, and the practice, common in the Clandinin lineage, of sharing draft research texts with participants and treating their objections as substantive rather than cosmetic. Relational ethics also extends past publication: the participant will read this, and will still be living the life you have narrated.
Common misconceptions
Narrative inquiry means reporting quotes at length. Length is not analysis. The tradition requires an argument about how the story works, not a curated transcript.
If a participant contradicts themselves, the data are unreliable. Inconsistency across tellings is a finding. People tell different stories to different audiences and at different points in a life, and the variation is often the most informative thing in the corpus.
Narrative analysis and thematic analysis are interchangeable. Fragmenting stories into cross-cutting codes is analysis of narratives; it discards exactly the configuration that narrative analysis exists to examine. Both are legitimate; only one is narrative analysis.
Recap
- Narrative inquiry studies how experience is composed and told, not whether the account is factually accurate.
- Clandinin and Connelly's relational methodology and Riessman's analytic techniques are distinct lineages; say which you are using.
- The three-dimensional space - temporality, sociality, place - structures inquiry in the Deweyan lineage.
- Thematic, structural, dialogic/performance, and visual lenses can be combined; Labov's parts and Bamberg's positioning give concrete analytic purchase.
- Keeping stories intact distinguishes narrative analysis from analysis of narratives.
- Reporting a single life at length raises confidentiality and relational-ethics demands that ordinary anonymization does not meet.
Sources
- Clandinin, D. J., Cave, M. T., & Berendonk, C. (2017). Narrative inquiry: A relational research methodology for medical education. Medical Education, 51(1), 89-96. pubmed.ncbi.nlm.nih.gov
- Butina, M. (2015). A narrative approach to qualitative inquiry. Clinical Laboratory Science, 28(3), 190-196. clsjournal.ascls.org
- Trahar, S. (2009). Beyond the story itself: Narrative inquiry and autoethnography in intercultural research in higher education. Forum: Qualitative Social Research, 10(1). qualitative-research.net
- Lucius-Hoene, G. (2000). Constructing and reconstructing narrative identity. Forum: Qualitative Social Research, 1(2). qualitative-research.net
- Kaiser, K. (2009). Protecting respondent confidentiality in qualitative research. Qualitative Health Research, 19(11), 1632-1641. pmc.ncbi.nlm.nih.gov
- Riessman, C. K. (2008). Narrative methods for the human sciences. Sage Publications. find source ↗
- Clandinin, D. J., & Connelly, F. M. (2000). Narrative inquiry: Experience and story in qualitative research. Jossey-Bass. find source ↗
- Key terms
- Narrative inquiry
- Qualitative study that takes the story as both data and object to learn how experience is given meaning.
- Emplotment
- The connecting of events into a meaningful whole with a temporal, causal, or thematic structure.
- Analysis of narratives
- Extracting themes across many stories, using stories as raw material.
- Narrative analysis
- Keeping each story intact and analyzing how it works as a story.
- Structural analysis
- Attending to how a story is built, such as Labov's functional parts of a narrative.
- Co-construction
- The idea that a story emerges from the relationship between teller and listener, not from the teller alone.
Module 3: Sampling and Access
Choosing whom and what to study through purposeful strategies, judging sample adequacy, and negotiating entry and positionality in the field.
Purposeful Sampling and Sample Adequacy
- Explain why qualitative sampling is purposeful rather than random and name major purposeful strategies.
- Distinguish saturation from information power as criteria for sample adequacy.
Qualitative sampling follows a different logic from quantitative sampling, and importing the wrong logic is a frequent and damaging error. A survey samples randomly so results generalize to a population by frequency. A qualitative study samples purposefully - selecting information-rich cases that best illuminate the phenomenon - because its goal is insight, not representativeness. Asking a qualitative study to justify its sample by random selection, or a purposeful study to report a margin of error, is asking it to be what it is not.
Key idea: The unit of a qualitative sampling decision is not the person but the information the person can supply about the phenomenon. That is why the operative distinction is between participants and information-rich cases: someone who has lived the phenomenon intensely, reflected on it, and can articulate it contributes more than three people who have brushed against it. It is also why "we recruited whoever responded to the flyer" is not a sampling strategy but the absence of one.
Purposeful sampling strategies
"Purposeful" is not a single technique but a family, each matched to a different analytic goal:
- Maximum variation: deliberately select cases that differ widely, so that any patterns holding across that diversity are especially robust, and the range of variation is documented.
- Homogeneous: select cases that are similar, to study a particular subgroup in depth (common for focus groups).
- Typical case: select cases that illustrate what is normal or average, to portray the ordinary rather than the extreme.
- Extreme or deviant case: select unusual cases - notable successes or failures - because the outliers often reveal what is hidden in the ordinary.
- Critical case: select a case for which, if the phenomenon holds (or fails) here, it likely holds (or fails) elsewhere - a strategic 'if it happens anywhere' choice.
- Snowball (chain): ask participants to refer others, essential for reaching hidden or hard-to-access populations.
- Theoretical: as in grounded theory, let the emerging analysis dictate the next cases.
- Criterion: include every case meeting a defined criterion, the standard logic when the phenomenon itself defines eligibility (everyone who received a diagnosis by telephone; every incident review in the period).
- Confirming and disconfirming case: once a pattern is developing, deliberately seek cases that should and should not fit, as a test rather than an illustration.
- Stratified purposeful: sample purposefully within predefined strata (early-, mid-, late-career) to guarantee coverage of a dimension the analysis needs.
Purposive and theoretical sampling are not synonyms, and conflating them is one of the most common errors in doctoral proposals. Purposive sampling is decided in advance on criteria derived from the research question: you know before you start whom you want and why. Theoretical sampling is decided during analysis on criteria derived from the developing categories: you cannot state in advance whom you will recruit at case fifteen, because that depends on what cases one through fourteen produced. A proposal that promises "theoretical sampling" and then lists a complete recruitment matrix in the same paragraph has described purposive sampling. Both are respectable; only one can be planned.
Snowball sampling deserves a warning as well as a place on the list. It is often indispensable for hidden populations, but referrals travel along existing social ties, so a chain started from one node can reproduce a single network's perspective and quietly exclude the isolated - who are frequently the people whose experience differs most. The standard mitigations are to seed multiple independent chains, to cap the number of referrals from any one participant, and to report the chain structure so a reader can see the shape of what was reached.
How large is large enough?
There is no fixed number, and a number alone never justifies a qualitative sample. Two connected principles do the work.
Saturation is the traditional criterion: continue sampling and analyzing until new data stop yielding new codes, categories, or themes - the point at which additional cases add repetition rather than insight. Saturation is a claim about the data and the analysis together, not a target set in advance, and it must be demonstrated (by showing that later cases produced little that was new), not merely asserted.
Information power is a more recent and more precise formulation: the more information the sample holds relevant to the study, the fewer participants are needed. Information power is higher - so fewer participants suffice - when the aim is narrow, the sample is highly specific to the question, established theory supports the analysis, the dialogue is strong, and the analysis focuses on cases rather than cross-case breadth. Conversely, a broad aim, a sparse sample, and a cross-case comparative analysis demand more participants. Information power reframes the question from "how many?" to "how much relevant information does each case carry, and how much do I need?"
Reporting the logic
What a committee scrutinizes is not whether you interviewed some canonical number but whether your sampling strategy fits your question and whether you can defend your stopping point. State the strategy, explain why it suits the aim, and justify adequacy through saturation or information power. A defended sample of nine can be stronger than an unexamined thirty.
The critique of saturation
Saturation is the most invoked and least examined concept in qualitative methods, and doctoral work should engage the critique rather than repeat the slogan. Four problems recur.
First, the term has drifted from its origin. In Glaser and Strauss it was theoretical saturation - the adequacy of a category developed through theoretical sampling. Most contemporary uses are data or code saturation - no new codes appearing - which is a different and weaker claim, and one that can be satisfied simply by asking narrow questions. Benjamin Saunders and colleagues distinguish these usages and show that authors frequently claim one while operationalizing another. Braun and Clarke argue that saturation is conceptually incoherent within reflexive thematic analysis, where meaning is generated by the analyst rather than discovered in data, so the idea of data being exhausted does not apply.
Second, it is usually asserted rather than demonstrated. "Saturation was reached at interview 17" appears without any evidence about what interviews 15 to 20 added. Where it is demonstrated, the method is straightforward: report the number of new codes per interview or per batch, state the stopping rule set in advance (for instance, no new codes across three consecutive interviews), and show the curve.
Third, empirical work suggests it happens earlier for some purposes than for others. Monique Hennink and colleagues distinguish code saturation - the point at which the range of issues has been identified, often within about nine interviews in their study - from meaning saturation, the point at which each issue is richly understood, which took considerably longer. Hennink and Kaiser's systematic review of empirical tests found saturation commonly reported in ranges around nine to seventeen interviews, with wide variation by homogeneity and scope. These are useful reference points, not targets.
Fourth, saturation cannot be promised in advance, which creates a practical problem for ethics applications and funding budgets that demand a number. The defensible move is to state a planned range with the reasoning behind it, commit to a stopping rule, and report what actually happened.
Kirsti Malterud's information power model is attractive partly because it can be argued prospectively. Its five dimensions - study aim, sample specificity, use of established theory, quality of dialogue, and analysis strategy - can each be assessed before recruitment, producing a reasoned estimate rather than a ritual invocation.
A worked illustration: defending a sample of eleven
A study asks how patients who have had a stroke before age fifty experience returning to work. The methods section might read:
"Sampling was purposive with a criterion of stroke before age fifty and at least one attempt to return to paid work, stratified to include participants who returned successfully, returned and left again, and did not return. Information power was judged high: the aim is narrow, the sample is highly specific to it, the analysis is supported by existing role-transition theory, and interviews averaged 82 minutes with strong dialogue. We therefore planned 10 to 14 participants. Recruitment continued to 11. New codes per interview fell from a mean of 9 in interviews 1-4 to 2 in interviews 8-9 and 0 in interviews 10-11; we treat this as adequate code development rather than as exhaustive saturation, and we note that the 'returned and left again' stratum contains only two participants, so claims about that trajectory are correspondingly tentative."
Notice what that paragraph does. It names the strategy, gives the criterion, shows the stratification and why, argues information power dimension by dimension, states a planned range, reports what happened with numbers, declines to overclaim, and flags the thinnest part of the sample. It is roughly 150 words and it is worth more than three pages of general commentary on qualitative sampling.
Common misconceptions
Saturation proves the sample was big enough. It supports the claim only if it was operationalized and demonstrated. Otherwise it is a rhetorical move.
More interviews always strengthen a study. Beyond the analytic capacity of the researcher, extra interviews thin the analysis. A corpus of forty transcripts analyzed superficially is weaker than fifteen analyzed deeply.
Maximum variation sampling makes findings generalizable. It makes shared patterns more striking and documents the range, but it does not confer statistical generalizability, and with a small n it can leave every subgroup too thin to say anything about.
Recap
- Qualitative sampling seeks information-rich cases; the strategy must be named and matched to the analytic goal.
- Purposive sampling is planned in advance from the research question; theoretical sampling is directed during analysis by developing categories.
- Snowball chains reproduce networks and systematically miss the isolated; seed multiple chains and report the structure.
- Theoretical, code, and meaning saturation are different claims; state which you mean and demonstrate it with evidence.
- Information power can be argued prospectively across aim, specificity, theory, dialogue quality, and analysis strategy.
Sources
- Palinkas, L. A., Horwitz, S. M., Green, C. A., Wisdom, J. P., Duan, N., & Hoagwood, K. (2015). Purposeful sampling for qualitative data collection and analysis in mixed method implementation research. Administration and Policy in Mental Health, 42(5), 533-544. pmc.ncbi.nlm.nih.gov
- Malterud, K., Siersma, V. D., & Guassora, A. D. (2016). Sample size in qualitative interview studies: Guided by information power. Qualitative Health Research, 26(13), 1753-1760. pubmed.ncbi.nlm.nih.gov
- Saunders, B., Sim, J., Kingstone, T., Baker, S., Waterfield, J., Bartlam, B., Burroughs, H., & Jinks, C. (2018). Saturation in qualitative research: Exploring its conceptualization and operationalization. Quality & Quantity, 52(4), 1893-1907. pmc.ncbi.nlm.nih.gov
- Hennink, M. M., Kaiser, B. N., & Marconi, V. C. (2017). Code saturation versus meaning saturation: How many interviews are enough? Qualitative Health Research, 27(4), 591-608. pmc.ncbi.nlm.nih.gov
- Hennink, M., & Kaiser, B. N. (2022). Sample sizes for saturation in qualitative research: A systematic review of empirical tests. Social Science & Medicine, 292, 114523. pubmed.ncbi.nlm.nih.gov
- Mason, M. (2010). Sample size and saturation in PhD studies using qualitative interviews. Forum: Qualitative Social Research, 11(3). qualitative-research.net
- Rahimi, S., & Khatooni, M. (2024). Saturation in qualitative research: An evolutionary concept analysis. International Journal of Nursing Studies Advances, 6, 100174. pmc.ncbi.nlm.nih.gov
- Key terms
- Purposeful sampling
- Selecting information-rich cases that best illuminate the phenomenon rather than sampling at random.
- Maximum variation sampling
- Selecting widely differing cases so that patterns holding across them are especially robust.
- Extreme/deviant case sampling
- Selecting unusual cases whose outliers reveal what is hidden in ordinary ones.
- Snowball sampling
- Reaching further participants through referrals from existing ones, useful for hidden populations.
- Saturation
- The point at which new data yield no new codes, categories, or themes.
- Information power
- The principle that samples holding more study-relevant information require fewer participants.
Access, Gatekeepers, and Positionality
- Plan realistic access to a research setting and describe the roles of gatekeepers and key informants.
- Analyze how the researcher's positionality and insider/outsider status shape data and rapport.
A brilliant design is worthless if you cannot get in. Access - securing entry to a setting and the trust of the people in it - is a practical and ethical achievement that shapes what data become possible, and it is where fieldwork most often stalls. This lesson treats access as a skill and then examines how who you are conditions what you can learn.
Key idea: Access is not a logistical preliminary that precedes the methodology; it is part of the methodology, because the route by which you entered determines whose accounts you can obtain and what those accounts will contain. Two studies of the same ward, one entered through the director of nursing and one through the union representative, will produce different data, and neither is contaminated - but a study that does not report its route of entry has withheld information the reader needs.
Gatekeepers and key informants
A gatekeeper is a person with the authority to grant or deny entry to a setting - a principal, a ward manager, a gang leader, a community elder. Formal permission from a gatekeeper is often necessary but rarely sufficient; the gatekeeper's blessing can also taint access if participants see you as the boss's agent, so you must manage the perception that your presence was imposed from above.
A key informant, by contrast, is a knowledgeable insider who helps you understand the setting, introduces you to others, and interprets local meanings. Key informants are invaluable, but reliance on them carries a hazard: they occupy a particular position in the group, and seeing the setting only through their eyes can skew your understanding toward their faction or perspective.
Gatekeeping is also an ethical problem, not only a practical one. A gatekeeper who selects which participants you may approach has, in effect, sampled for you, and the selection is rarely random: managers introduce their best performers, clinicians refer their most articulate patients, and community leaders offer families who will represent the community well. Worse, gatekeeper permission can undermine voluntariness - staff told by a manager that "we are supporting this study" may feel that declining is not really an option. The standard safeguards are to obtain organizational permission and individual consent separately and visibly, to make recruitment routes available that do not run through the gatekeeper (a poster, an independent contact address), to state explicitly that participation and non-participation will not be reported back, and to record in the audit trail who was offered and who was not. Where a gatekeeper insists on nominating participants, that constraint belongs in the limitations section, described plainly.
Getting in: what actually works
Access negotiations succeed on a small number of unglamorous things. Approach through an existing relationship where one exists, since a warm introduction outperforms a cold letter by a wide margin. Offer something the setting values and can use - a summary presentation, an anonymized report of themes, help with a literature search - without straying into inducement. Ask for a small first commitment (an hour of observation, one interview) rather than for the whole study at once, because organizations agree in increments. Be specific about burden in hours, because vagueness reads as risk. Anticipate the gatekeeper's real question, which is almost never "is this good science" but "what could this cost me if it goes wrong," and answer it directly by naming what you will and will not report. And build in time: a realistic timeline for institutional access in a health system is months, not weeks, and doctoral projects fail on this more often than on analysis.
Access is ongoing, not a one-time gate
It is a beginner's error to think of access as a door opened once at the start. In practice access is continuous and negotiated: initial permission gets you in, but deeper access - to candid talk, to backstage settings, to sensitive documents - is earned gradually as trust accrues, and it can be withdrawn if trust is broken. Building rapport, the relationship of trust and ease that makes honest disclosure possible, is therefore not a preliminary courtesy but continuous work throughout the study.
Positionality
Positionality refers to the researcher's social location - characteristics such as race, gender, age, class, profession, and life experience - and how that location relates to participants and shapes the research. Positionality is not a bias to be eliminated (it cannot be) but a condition to be analyzed. A researcher's identity affects who will talk to them, what participants will disclose, and how accounts are interpreted.
Insider and outsider
The insider (emic) researcher shares membership or identity with the participants; the outsider (etic) researcher does not. Each position carries a trade-off, and neither is simply better:
| Advantages | Risks | |
|---|---|---|
| Insider | Easier access and rapport; grasps tacit meaning; shared language | May take the familiar for granted; assumed shared views can suppress explanation; role conflict |
| Outsider | Fresh eyes notice the taken-for-granted; participants may explain more fully to a naive listener | Slower access and trust; may misread local meaning; visible difference can inhibit disclosure |
The insider's danger is familiarity blindness: what is obvious to a member is exactly what never gets articulated, so the insider may fail to ask about the very things an outsider would notice. The outsider's danger is misinterpretation and slower trust. The methodological response is the same in both cases - reflexivity: to examine, and to disclose in the write-up, how your position shaped access, rapport, disclosure, and interpretation, so that readers can weigh your account with that knowledge in hand.
The insider/outsider binary is itself too crude for most real studies. Contemporary treatments describe positionality as multiple and situational: you may be an insider by profession and an outsider by race, an insider by gender and an outsider by seniority, and which of those becomes salient can change between one interview and the next. Some writers describe the resulting stance as the "space between," in which the researcher is neither wholly in nor wholly out and must track which membership is operating at a given moment. Positionality also has a temporal dimension - a nurse researcher who left practice five years ago is a lapsing insider, credible enough to be told things but no longer accountable to the ward's daily politics.
When access is denied, or withdrawn
Refusal is data, and doctoral students routinely waste it. If three of five clinics decline, the pattern of who declined and the reasons given tell you something about the phenomenon - clinics under regulatory scrutiny may refuse precisely because the topic is live for them, which is itself a finding about the field even if it is a problem for the study. Record refusals and stated reasons in the audit trail. When access collapses mid-study, the options in order of preference are to renegotiate scope with the gatekeeper (often a narrower ask is acceptable when a broad one is not), to shift to a comparable second site while treating the truncated first site as a case in its own right, to redesign around data that remain available such as documents and public records, or to reframe the study around the access failure itself, which some strong organizational research has done. What is not defensible is quietly continuing to collect data after permission has lapsed.
A worked illustration: writing a positionality statement that does work
Most positionality statements are decorative. Compare two.
Weak: "As a white, middle-class female researcher, I acknowledge that my background may have influenced the research process. I attempted to remain objective throughout."
Strong: "I am a former hospice social worker interviewing hospice nurses. My professional membership secured access quickly - the clinical manager agreed within a week - and shaped disclosure in two identifiable ways. Participants used clinical shorthand without explanation and, when I asked them to unpack it, three said versions of 'you know what it's like.' I therefore adopted a deliberate stance of asking for description as though naive, and in later interviews I opened by saying that I had been out of practice for six years and would need things spelled out. Second, my former role is aligned with a discipline that nurses on this unit describe as under-resourced; two participants raised staffing grievances early, and I noted that they may have read me as a potential ally in an internal dispute. I have marked in the analysis where staffing appears as an unelicited complaint versus a response to my probes. Finally, I did not share a race or a first language with four of the eleven participants, and those interviews were, on average, twenty minutes shorter; I do not treat their brevity as reduced salience of the topic."
The second statement is longer, but every sentence names a mechanism, a consequence, and a response. That is the test to apply to your own: does each claim about your position connect to something specific in the data or the procedure? If not, cut it.
Common misconceptions
Being an insider is an advantage. It is a trade. Insiders get in faster and read tacit meaning, but they suffer familiarity blindness, face role conflict (what do you do when you observe unsafe practice?), and may be told less about things participants assume they already condone.
A positionality statement is a confession of bias. It is an account of how a knower was positioned, which is information a reader needs in order to interpret findings, not an apology.
Access ends when the ethics approval arrives. Formal permission is the beginning. Deeper access - backstage talk, sensitive documents, the second interview in which a participant says what they meant the first time - accrues over months and can be lost in an afternoon.
Reciprocity is optional. Participants who give time to research reasonably expect something back. What is offered must be proportionate and not coercive, and the relevant question for an ethics committee is whether the offer would induce someone to participate against their better judgment.
Recap
- The route of entry shapes the data and belongs in the methods section.
- Gatekeeper permission and individual consent are separate; gatekeeper nomination of participants is a sampling constraint to report.
- Key informants are invaluable and partial; triangulate across positions in the setting.
- Access is continuous, incremental, and revocable; rapport is ongoing work, not a preliminary courtesy.
- Positionality is multiple and situational rather than a fixed insider/outsider label.
- A positionality statement earns its place only when each element names a mechanism, a consequence, and a response.
Sources
- Holmes, A. G. D. (2020). Researcher positionality: A consideration of its influence and place in qualitative research. Shanlax International Journal of Education, 8(4), 1-10. eric.ed.gov
- Bourke, B. (2014). Positionality: Reflecting on the research process. The Qualitative Report, 19(33), 1-9. nsuworks.nova.edu
- Olmos-Vega, F. M., Stalmeijer, R. E., Varpio, L., & Kahlke, R. (2023). A practical guide to reflexivity in qualitative research: AMEE Guide No. 149. Medical Teacher, 45(3), 241-251. pubmed.ncbi.nlm.nih.gov
- Rohm, M. (2024). An exploration of practical reflexivity: Navigating categories in research encounters. Forum: Qualitative Social Research, 25(3). qualitative-research.net
- Kawulich, B. B. (2005). Participant observation as a data collection method. Forum: Qualitative Social Research, 6(2). qualitative-research.net
- Sanjari, M., Bahramnezhad, F., Fomani, F. K., Shoghi, M., & Cheraghi, M. A. (2014). Ethical challenges of researchers in qualitative studies: The necessity to develop a specific guideline. Journal of Medical Ethics and History of Medicine, 7, 14. pmc.ncbi.nlm.nih.gov
- Hammersley, M., & Atkinson, P. (2019). Ethnography: Principles in practice (4th ed.). Routledge. find source ↗
- Key terms
- Access
- Securing entry to a setting and the trust of its members, a practical and ethical achievement.
- Gatekeeper
- A person with authority to grant or deny entry to a research setting.
- Key informant
- A knowledgeable insider who helps the researcher understand the setting and reach others.
- Rapport
- The relationship of trust and ease that makes honest disclosure possible.
- Positionality
- The researcher's social location and how it relates to participants and shapes the research.
- Insider/outsider
- Whether the researcher shares membership with participants, each status carrying distinct advantages and risks.
Module 4: Generating Data
The core techniques for producing qualitative data - in-depth interviews, focus groups, observation and fieldnotes, and the use of documents and artifacts.
In-Depth Interviewing
- Distinguish structured, semi-structured, and unstructured interviews and design a semi-structured guide.
- Apply core interviewing skills - open questions, probing, and active listening - and avoid common pitfalls.
The in-depth interview is the workhorse of qualitative research: a purposeful conversation designed to elicit a participant's perspective in depth and in their own words. It looks deceptively easy - just talking - but skilled interviewing is difficult, and the difference between a rich transcript and a barren one is almost entirely the interviewer's craft.
Key idea: Nothing downstream can repair a thin interview. No coding scheme, no software, no analytic framework will recover depth that was never elicited. This is the point in a qualitative study where the largest quality gains are available for the least methodological sophistication, and it is the part doctoral students most often skip practicing.
Three degrees of structure
- Structured interviews ask every participant the same fixed questions in the same order, approaching a spoken questionnaire. They maximize comparability but suppress the open-ended exploration that is the point of qualitative work; they are the least characteristically qualitative form.
- Semi-structured interviews - the most common qualitative form - use an interview guide of open questions and topics as a flexible framework, while allowing the interviewer to follow the participant's lead, reorder questions, and probe what emerges. This balances coverage of key topics with responsiveness to the individual.
- Unstructured interviews begin from one or two broad openings and follow wherever the participant goes, approaching a guided conversation. Common in ethnography and life-history work, they yield depth at the cost of comparability across participants.
Designing a semi-structured guide
A good guide is a small set of open, non-leading questions organized to move from easier, rapport-building topics toward more sensitive ones, with planned probes ready beneath each. Grand-tour questions ("walk me through a typical day...") invite narrative; specific follow-ups pursue detail. The guide is a servant, not a script: in a good interview you depart from its order constantly.
Hanna Kallio and colleagues, reviewing the methodological literature, describe guide development as a five-phase process that is worth following explicitly: identify the prerequisites for using a semi-structured design at all, retrieve and use previous knowledge, formulate a preliminary guide, pilot it, and present the completed guide. The pilot phase is the one most often skipped and the one that most reliably improves data. Two or three internal pilots (with colleagues who know the topic) surface questions that are unanswerable as worded; one or two field pilots (with people like your participants) surface questions that are answerable but produce nothing.
A workable guide for a 60- to 90-minute interview has roughly five to eight main questions, not twenty. A common structure runs: an opening that establishes the participant as the expert and you as the learner; a grand-tour question that generates narrative and gives you material to probe for the rest of the session; two or three core questions addressing the research sub-questions; one question that deliberately invites disconfirmation ("Have there been times when it worked differently, or when what you have described did not hold?"); a question inviting the participant to raise what you have not asked; and a closing that returns the participant to the present and checks on their state after a possibly difficult conversation.
Core skills
- Ask open, not closed, questions. "What was that like for you?" opens; "Were you upset?" closes and leads. Closed and leading questions are the most common cause of thin data.
- Probe. The follow-up is where depth lives. Silence (waiting), echoing a key word, and neutral prompts ("tell me more about that," "what do you mean by...?") draw out elaboration without steering it.
- Listen actively. The interviewer should talk far less than the participant. Resist the urge to fill pauses, to finish sentences, or to insert your own views. A pause is often followed by the most important thing the participant says.
- Do not lead or judge. Signaling a preferred answer or reacting with visible approval or disapproval contaminates the data through social desirability.
Common pitfalls
Beginners characteristically ask double-barreled questions (two questions at once, so you cannot tell which was answered), pose leading questions that plant the answer, talk too much, retreat to safe closed questions when a topic gets uncomfortable, and abandon a promising thread instead of probing it.
A subtler error is treating the interview as extraction of pre-existing facts rather than as a co-construction of meaning: what a participant says is shaped by the relationship and the moment, which is why rapport, neutrality, and skilled probing matter so much. Record and transcribe whenever consent allows, because analysis depends on the exact words, and note the nonverbal and contextual detail that a recording misses.
A worked illustration: the same moment, handled two ways
A participant in a study of parents of children with a rare disease says: "After the diagnosis, everything with our friends changed."
Weak continuation: "That must have been very isolating. Did you find that people didn't know what to say, or was it more that you withdrew?" This does three damaging things in one breath. It supplies the interpretation ("isolating") before the participant has offered one, which most participants will politely adopt. It is double-barreled, so the answer cannot be attributed to either half. And it converts an open account into a choice between two researcher-supplied options.
Strong continuation: "Changed how?" - then silence. Suppose the participant says: "People stopped inviting us. Or they'd invite us and then check first whether it was 'a good week.'" The interviewer echoes the participant's own phrase: "A good week." The participant elaborates: "That's what my sister calls it. I hate it. It makes it sound like there's a bad week and then it passes." Now something has surfaced that no guide could have anticipated: a conflict over the vocabulary others use to describe the child's condition, and a temporal framing (episodes that pass) that the parent rejects. A follow-up - "Tell me about a time someone used a phrase like that" - moves from general characterization to a specific incident, which is where analyzable detail lives.
Three transferable moves are visible here. Ask "how" or "in what way" rather than offering candidate causes. Echo the participant's own words rather than paraphrasing into your vocabulary, because paraphrase substitutes your category for theirs. And convert generalizations into episodes, since "tell me about a specific time" is the single highest-yield probe in qualitative interviewing.
Sensitive topics, distress, and power
Interviews about illness, loss, discrimination, or violence require additional preparation. Agree in advance what you will do if a participant becomes distressed - typically pause the recording, stay present without pushing, offer to stop or reschedule, and never proceed on the assumption that continuing shows resilience. Have a written list of local support resources to leave behind, and know your own reporting obligations before you begin, so that any limits to confidentiality (disclosure of risk to a child, for instance) are stated in the consent process rather than discovered mid-interview. Distress is not automatically harm; many participants value being asked, and treating a difficult topic as unaskable can itself be paternalistic. The judgment to make is whether the interview is being conducted in a way that leaves the participant with more control than they had at the start.
Power runs through every interview. The researcher sets the topic, controls the recording, and will write the account; the participant may be a patient, an employee, or a student in relation to the institution the researcher represents. Practical counterweights include letting the participant choose the location and time, being explicit that any question can be declined, offering to stop the recording at any point, and telling participants how they can withdraw data afterward and until when.
Remote and asynchronous interviewing
Video and telephone interviews are now standard rather than second-best, and the evidence broadly suggests they yield comparable depth for most topics while expanding geographic reach and suiting participants with mobility or caregiving constraints. They do change the encounter: nonverbal cues are attenuated, silences feel longer and are more often filled, and the researcher loses the contextual information a home or workplace supplies. Practical adjustments include tolerating longer pauses deliberately, checking connection quality before beginning, having a fallback channel agreed in advance, and recording an observational note about the setting the participant chose. Consent and data security need explicit attention: name which platform is used, where the recording is stored, and who can access it.
Transcription as an analytic decision
Transcription is not clerical. A denaturalized transcript cleans up stutters, repetitions, and fillers to foreground content, and suits thematic and content analysis. A naturalized transcript preserves them, along with pauses, overlaps, and emphasis, and is required for narrative, discourse, and conversation-analytic work. Choosing one is choosing what can later be analyzed, and the choice should be stated. If a transcription service is used, the researcher should still listen to each recording against the transcript, both to correct errors and because that listening is the first pass of analysis. Where interviews are conducted in one language and reported in another, translation decisions - who translated, whether back-translation was used, how untranslatable terms were handled - belong in the methods section.
Common misconceptions
A longer interview is a better interview. Ninety minutes of unfocused talk yields less than fifty minutes of well-probed specifics. Length correlates with depth only when probing is good.
Rapport means the participant likes you. Rapport means the participant believes you will listen without judging and represent them fairly. Excessive warmth can suppress disagreement and produce accounts designed to please.
Neutrality means showing no response. A blank interviewer produces guarded participants. Neutrality means not signaling which answer you prefer, while remaining visibly attentive and human.
Recap
- Interview quality caps study quality; no analysis recovers depth that was not elicited.
- Semi-structured guides should be short, piloted, and organized from rapport-building toward sensitive material, with a disconfirmation question built in.
- Echo the participant's words, ask how rather than offering causes, and convert generalizations into specific episodes.
- Plan in advance for distress, state limits to confidentiality before beginning, and use practical counterweights to the power asymmetry.
- Remote interviewing is broadly comparable in depth but requires deliberate handling of silence, context, and data security.
- Naturalized versus denaturalized transcription determines what analyses remain available; state the choice.
Sources
- DeJonckheere, M., & Vaughn, L. M. (2019). Semistructured interviewing in primary care research: A balance of relationship and rigour. Family Medicine and Community Health, 7(2), e000057. pmc.ncbi.nlm.nih.gov
- Kallio, H., Pietila, A. M., Johnson, M., & Kangasniemi, M. (2016). Systematic methodological review: Developing a framework for a qualitative semi-structured interview guide. Journal of Advanced Nursing, 72(12), 2954-2965. pubmed.ncbi.nlm.nih.gov
- Britten, N. (1995). Qualitative interviews in medical research. BMJ, 311(6999), 251-253. pmc.ncbi.nlm.nih.gov
- Gill, P., Stewart, K., Treasure, E., & Chadwick, B. (2008). Methods of data collection in qualitative research: Interviews and focus groups. British Dental Journal, 204(6), 291-295. nature.com
- Roberts, R. E. (2020). Qualitative interview questions: Guidance for novice researchers. The Qualitative Report, 25(9), 3185-3203. nsuworks.nova.edu
- Moser, A., & Korstjens, I. (2018). Series: Practical guidance to qualitative research. Part 3: Sampling, data collection and analysis. European Journal of General Practice, 24(1), 9-18. pmc.ncbi.nlm.nih.gov
- Kvale, S., & Brinkmann, S. (2015). InterViews: Learning the craft of qualitative research interviewing (3rd ed.). Sage Publications. find source ↗
- Key terms
- In-depth interview
- A purposeful conversation designed to elicit a participant's perspective in depth and in their own words.
- Semi-structured interview
- The common qualitative form using a flexible guide of open questions while following the participant's lead.
- Interview guide
- A small set of open questions and topics that frames a semi-structured interview without scripting it.
- Probe
- A neutral follow-up (silence, echo, or prompt) that draws out elaboration without steering it.
- Leading question
- A question that signals or plants a preferred answer, contaminating the response.
- Double-barreled question
- A single question containing two questions, so the answer cannot be attributed to either.
Focus Groups
- Explain what focus groups add beyond individual interviews and when to choose them.
- Describe composition, moderation, and the analytic significance of group interaction.
A focus group is a facilitated discussion among a small set of participants, brought together to explore a topic through their interaction with one another. It is not simply a time-saving way to interview several people at once; its distinctive value lies precisely in the group interaction - what people say in response to, agreement with, and challenge from each other - which surfaces meanings that individual interviews cannot reach.
Key idea: The test of whether a study genuinely used focus groups is whether the findings could have been produced from the same people interviewed separately. If they could, the group was a logistical convenience and should be described as such. If they could not - if the finding concerns how a position was jointly built, contested, or policed - then the method did its work. Jenny Kitzinger's foundational account makes exactly this point: the method exists to exploit interaction, and studies that ignore it waste it.
What the group adds
In a good focus group participants build on, qualify, and contest one another's contributions. This dynamic makes visible how views are formed and defended in a social context, reveals the range of perspectives and the points of consensus and disagreement within a community, and can prompt people to articulate assumptions they would not have volunteered alone. Focus groups are therefore especially useful for exploring shared meanings, social norms, and community language, and for generating a breadth of ideas early in a project. The interaction is not noise around the data; it is the data.
Composition and number
Design choices follow from that logic:
- Size. Groups of roughly six to ten balance diversity of input against everyone's chance to speak; too large and quieter voices are lost, too small and interaction thins.
- Homogeneity. Groups are usually composed to be homogeneous enough that participants feel comfortable speaking candidly - people are franker among perceived peers - while retaining enough internal variety to spark discussion. Mixing sharply unequal statuses (for example, staff and their managers) can silence the less powerful.
- Number of groups. One group is never enough, because a single group's dynamic is idiosyncratic. Researchers run several groups, often segmented by a key characteristic, and continue until themes recur across groups (a saturation logic at the group level).
- Pre-existing versus assembled groups. Groups drawn from people who already know each other (a ward team, a support group) speak in their established idiom, refer to shared history, and will challenge each other more readily - but they also carry their existing hierarchies and cannot be assumed to speak freely about their own group. Assembled groups of strangers are freer of local politics but slower to reach candor and more likely to produce socially acceptable talk.
- Segmentation logic. Segment on the dimension you expect to matter, then compare across segments; segmenting nurses from physicians, or newer from longer-serving staff, converts a limitation into an analytic design. Report how you segmented and why.
On the number of groups, Monique Hennink and colleagues examined saturation empirically in focus group research and found that a relatively small number of groups - in their study, the range that produced most codes was in single figures, with fuller meaning development requiring more - accounted for the great majority of issues raised. The practical guidance that follows is to plan a range (commonly four to eight groups, or three to four per segment), state a stopping rule, and report what later groups added.
Moderation
The moderator's job is to facilitate interaction, not to interview each person in turn around the circle. Skilled moderation opens with easy, inclusive questions; poses open prompts to the group rather than to individuals; draws out quieter participants and gently contains dominant ones; and encourages participants to respond to each other ("does anyone see it differently?") rather than routing every comment through the moderator. The aim is a genuine conversation the moderator steers lightly, keeping it on topic while letting the interaction breathe.
Limits and analysis
Focus groups have real constraints. Group dynamics can distort what is said: a forceful participant can steer the group, and social pressure can push responses toward a perceived norm - a conformity effect that suppresses minority or unpopular views. They are poorly suited to highly sensitive or private topics, where individual interviews protect confidentiality and candor.
And because talk is public and interactive, you cannot treat a focus-group statement as an individual's settled private belief. Analytically, this means the unit of analysis includes the interaction itself - who responds to whom, where consensus builds or breaks, how the group jointly constructs a position - not merely a tally of individual remarks. Reading a focus group as if it were several separate interviews discards exactly what makes the method valuable.
A worked illustration: coding the interaction, not just the content
An excerpt from a group of six community pharmacists discussing a new prescribing role:
P3: Honestly, I don't think most of us are ready for it.
P1: I'd push back on that. I've been doing it in practice for two years.
P3: In a GP practice, though. That's different from a high street shop with one pharmacist and a queue.
P5: (laughs) There it is. It's always the practice pharmacists.
P1: That's not fair.
P4: No, but it is a bit true, isn't it? The training assumes you've got a room and a door.
P3: And ten minutes.
P1: ...alright, yes. The setting matters more than I said.
A content-only analysis would code this as "readiness concerns" and "training inadequate," lose most of the value, and then quote P3's opening line as if it were an individual opinion. An interaction-aware analysis records something more useful.
First, the trajectory: an initial claim about competence is reframed by the group into a claim about setting. That reframing is the finding - the group collectively relocates the problem from individual readiness to structural conditions, and it does so in four turns.
Second, the mechanism: P5's laughter and "there it is" invokes a shared category (practice pharmacists as a privileged type) that nobody had to explain, indicating an established distinction in this occupational community. P4's "no, but it is a bit true" performs the classic softening move that lets a challenge land without rupture.
Third, the concession: P1 revises her position aloud. In an individual interview, positions rarely move; here the movement is visible, and the analyst can say something about what kind of argument was persuasive to a peer.
Fourth, what is absent: nobody defends the training. Silence where disagreement would be expected is itself analyzable, provided it is treated as a hypothesis to check in later groups rather than as a conclusion.
Practically, this means transcripts must identify speakers and preserve turn order, overlaps, and laughter, and that fieldnotes taken by an assistant moderator - who spoke after whom, who never spoke, seating, nonverbal agreement - are part of the data rather than a courtesy.
Running the group: the practical layer
Recruit more than you need, because attrition in focus groups is high; over-recruiting by roughly a third is standard. Use two staff where possible, a moderator and an assistant who handles recording, notes the turn sequence, and watches for participants trying to enter the conversation. Open with a round of low-stakes introductions so every voice has been heard early, which measurably increases later participation from quieter members. Agree ground rules explicitly, including the point that confidentiality within the group cannot be guaranteed by the researcher and depends on the participants themselves - an important limitation to state in the consent process. Use stimulus material (a vignette, a photograph, a proposed policy statement, a ranking task) when a topic is abstract, since something to react to generates more interaction than an open question. Close by summarizing what you heard and inviting correction, which serves both as a courtesy and as an immediate credibility check.
Online and asynchronous groups
Synchronous video groups are now routine and bring genuine advantages - participants dispersed across a region can meet, travel costs disappear, and people who would not attend a room in a hospital sometimes will join from home. They also degrade the very thing the method exists for. Overlapping talk, the medium of spontaneous challenge, is technically punished; platform turn-taking pushes the discussion back toward a series of individual answers routed through the moderator; and side comments and body language largely vanish. Mitigations include capping the group at five or six, using the chat channel deliberately and treating it as data, asking participants to respond directly by name to one another, and having the assistant moderator watch for raised hands and unmuting. Asynchronous formats - a moderated discussion board or messaging thread over several days - trade immediacy for reflection and are well suited to geographically dispersed professionals, but they produce written, edited talk and should be analyzed as such rather than as speech.
Common misconceptions
Focus groups are a cheap substitute for interviews. They are cheaper per participant but produce different data. Use them when interaction is the point, not when scheduling is.
Consensus in a group means agreement in the population. Group consensus can be a product of conformity pressure, dominant speakers, or the absence of anyone positioned to disagree. Report the trajectory by which agreement was reached, not just its endpoint.
You should count how many participants endorsed each theme. In a group, silence is ambiguous - agreement, disengagement, or unwillingness to contradict a colleague - so tallies are misleading in a way they are not in individual interviews.
Sensitive topics are always unsuitable. Usually individual interviews are safer, but shared-experience groups (bereaved parents, people in recovery) can provide solidarity that enables disclosure. The decision turns on whether participants risk exposure to people who have power over them.
Recap
- The distinctive data are interactional; if the findings could have come from separate interviews, the method was not used.
- Compose groups for candor through perceived peer status, segment on the dimension expected to matter, and never rely on one group.
- Pre-existing groups bring shared idiom and existing hierarchies; assembled groups bring freedom and slower candor.
- Moderation routes talk between participants rather than through the moderator.
- Analysis codes trajectory, challenge, concession, and silence; transcripts must preserve speakers and turn order.
- Confidentiality within the group depends on participants and must be disclosed as a limit during consent.
Sources
- Kitzinger, J. (1995). Qualitative research: Introducing focus groups. BMJ, 311(7000), 299-302. pmc.ncbi.nlm.nih.gov
- Hennink, M. M., Kaiser, B. N., & Weber, M. B. (2019). What influences saturation? Estimating sample sizes in focus group research. Qualitative Health Research, 29(10), 1483-1496. pmc.ncbi.nlm.nih.gov
- Gill, P., Stewart, K., Treasure, E., & Chadwick, B. (2008). Methods of data collection in qualitative research: Interviews and focus groups. British Dental Journal, 204(6), 291-295. nature.com
- Moser, A., & Korstjens, I. (2018). Series: Practical guidance to qualitative research. Part 3: Sampling, data collection and analysis. European Journal of General Practice, 24(1), 9-18. pmc.ncbi.nlm.nih.gov
- Tong, A., Sainsbury, P., & Craig, J. (2007). Consolidated criteria for reporting qualitative research (COREQ): A 32-item checklist for interviews and focus groups. International Journal for Quality in Health Care, 19(6), 349-357. pubmed.ncbi.nlm.nih.gov
- Kaiser, K. (2009). Protecting respondent confidentiality in qualitative research. Qualitative Health Research, 19(11), 1632-1641. pmc.ncbi.nlm.nih.gov
- Krueger, R. A., & Casey, M. A. (2015). Focus groups: A practical guide for applied research (5th ed.). Sage Publications. find source ↗
- Key terms
- Focus group
- A facilitated group discussion that explores a topic through participants' interaction with one another.
- Group interaction
- The building on, qualifying, and contesting among participants that constitutes the focus group's distinctive data.
- Moderator
- The facilitator who steers a focus group lightly and encourages participants to respond to each other.
- Homogeneous grouping
- Composing a group of perceived peers so participants speak candidly, while keeping enough variety to spark discussion.
- Segmentation
- Running separate groups divided by a key characteristic to compare across them.
- Conformity effect
- Social pressure that pushes group responses toward a perceived norm and can suppress minority views.
Observation and Fieldnotes
- Distinguish participant-observer roles along the participation-observation continuum.
- Write descriptive fieldnotes that separate observation from inference and support later analysis.
Interviews tell you what people say they do; observation lets you see what they actually do, in context, as it happens. Because talk and action often diverge, direct observation is an indispensable source, and in ethnography it is central. But observation is not passive looking; it is disciplined, purposeful, and recorded, and its quality depends on where the researcher stands and how faithfully they write it down.
Key idea: The gap between report and practice is not usually dishonesty. People genuinely cannot report accurately on routines that have become automatic, on norms they have never had to articulate, or on the discrepancy between the procedure they believe they follow and the shortcut they have long since normalized. Observation is the method for reaching what participants cannot tell you because they no longer see it.
The participation-observation continuum
Gold's classic typology arranges the observer's stance along a continuum by how much the researcher participates versus merely watches:
| Role | Participation | Trade-off |
|---|---|---|
| Complete participant | Full member; role often covert | Deep access to insider experience; ethical concerns and going native |
| Participant-as-observer | Participates, role known | Good rapport and access; risk of influencing the setting |
| Observer-as-participant | Mainly observes, some interaction | More detachment; thinner immersion in meaning |
| Complete observer | Detached, unobtrusive watching | Minimal reactivity; no access to insider meaning |
Two hazards bracket the continuum. At the immersed end lies going native: over-identifying with the group until critical, analytic distance is lost and the researcher can no longer see the setting as a researcher. At the detached end lies reactivity (the observer effect): people behave differently because they know they are watched, though this typically fades as a researcher's presence becomes routine over time.
What to observe
Purposeful observation attends to more than dramatic events: the physical space and its arrangement, the actors present and their roles, the activities and their sequence, interactions and who initiates them, informal language and local terms, and, crucially, the routine and the absent - what always happens and what never happens. The taken-for-granted is exactly what an observer is positioned to notice.
Spradley's nine dimensions give the same idea a checklist you can carry: space, actor, activity, object, act, event, time, goal, and feeling. Their value is that they can be crossed against one another to generate questions you would not otherwise think to ask - how does time structure activity here? which objects are handled only by certain actors? A more focused alternative is to observe sequences: pick a recurring event (an admission, a handover, a lesson opening) and record it repeatedly until you can predict its steps, then attend to the occasions when it deviates. Deviation from a well-mapped routine is one of the most productive things a fieldworker can catch.
Sampling applies to observation as much as to interviewing. Time sampling selects blocks (early mornings, weekends, the last hour of a shift) to ensure temporal coverage. Event sampling selects occasions of a defined type wherever they occur. Person-following shadows one actor through a period, showing how a role is experienced across settings. Each yields a different picture, and a study that observed only Tuesday afternoons because that was when the researcher was free should say so.
Fieldnotes
Observation becomes data only when written as fieldnotes. Good practice records brief jottings in the moment (a word or phrase to trigger memory) and expands them into full, detailed notes as soon as possible afterward, before memory decays - ideally the same day. The cardinal rule is to separate description from inference. Descriptive notes render what was concretely seen and heard in low-inference language; interpretation and hunches go in a clearly marked reflective or observer's-comment column.
Compare: "she was angry" (an inference) versus "she raised her voice, pointed at the door, and left without closing it" (a description from which anger might later be inferred). Keeping the two apart lets you revisit the raw record when your early interpretation turns out to be wrong - and it often does. Alongside descriptive and reflective notes, many fieldworkers keep methodological notes (decisions about how to proceed) and a running record of emerging analytic ideas, so that the fieldnote corpus supports rigorous analysis rather than nostalgic recollection.
A worked illustration: jotting to expanded note to analytic memo
Jotting (written in a pocket notebook, 08:47): "gloves box empty again - K uses B's stash - laughs - 'don't tell stores'"
Expanded descriptive note (written 18:20 the same day): "At 08:47 K (HCA, approx. 40s, on this ward two years) went to the glove dispenser outside bay 2 and found it empty. She looked at it for about two seconds, said nothing, walked twelve paces to the linen store, opened the second drawer down, and took a box of medium gloves. She said to me, without my asking, 'Don't tell stores.' She laughed on the word stores. She then returned and put the box on top of the empty dispenser rather than refilling it. Two other staff passed during this and neither commented. The dispenser was still empty and the loose box still on top of it when I left at 16:10."
Observer's comment (marked OC, kept separate): "OC: My first reading was rule-breaking, but the laugh and the unprompted 'don't tell' suggest she assumed I already knew this was routine and was inviting me into it. Nobody reacted, which points to normalization rather than transgression. Check: is 'stores' a person, a department, or a shorthand for an audit process? Look for other instances of private stashes."
Methodological note: "MN: I have been observing mornings only for three weeks. Supplies may behave differently on nights when stores is closed. Extend to two night shifts."
Analytic memo, two weeks later: "Six instances now of what I am calling parallel provisioning: informal caches of consumables maintained by individual staff outside the official supply system, unremarked by colleagues, and referenced with humour that presumes shared knowledge. The humour matters - it marks the practice as an open secret rather than a concealed violation, which distinguishes it from actual rule-breaking. Emerging property: caches are maintained by staff with long tenure on the unit and are shared readily with juniors, which makes access to them a marker of belonging. Disconfirming case: on 14 March, an agency nurse asked openly where the gloves were and was directed to the official store, not the cache. That suggests the cache is bounded by membership. Next: watch what agency and bank staff are and are not told."
Notice the labor division. The descriptive note contains times, positions, counts of paces, and exact words, and no interpretation at all - it will still be usable in a year when the researcher's early reading has been abandoned. The observer's comment holds the interpretation and, crucially, states what would check it. The methodological note repairs a sampling gap. The memo does the conceptual work, names a candidate category, and engages a negative case. Only the last of these looks like analysis, but none of it is possible without the first.
Structured and unstructured observation
Observation is not a single technique. Unstructured observation, the ethnographic default, enters with sensitizing concerns rather than categories and lets the relevant units of behavior become apparent through immersion. Structured observation uses a pre-specified schedule - defined behaviors, defined intervals, tallies - and is a quantitative instrument that happens to involve watching. Between them sits focused observation, in which early open observation identifies what matters and later sessions concentrate on a narrowed set of practices, often with a simple recording template. Most doctoral fieldwork should follow that trajectory deliberately rather than drifting into it: begin wide, write memos identifying what is worth watching, then narrow and say in the methods section when and why the narrowing happened. Announcing at the outset that you will observe "everything" is a plan to produce notes you cannot analyze.
Ethics and practicalities of observing
Consent in observational research is genuinely hard, because a ward, a classroom, or a public meeting contains people who did not sign anything and who move in and out. Standard practice combines organizational approval, visible notification (posters, a briefing at handover, an introduction at the start of a meeting), individual consent for anyone who is a focus of sustained observation or is quoted, and an opt-out route that does not require confronting the researcher. Observation in genuinely public space may not require individual consent, but "public" is narrower than it appears - a hospital corridor is not a public square. Covert observation is defensible only in rare cases where the research question cannot otherwise be answered and the setting is one in which no reasonable expectation of privacy exists, and it requires explicit ethics committee approval rather than a researcher's own judgment.
Practically: write fieldnotes the same day, without exception, because detail decays within hours and a week of unwritten observation is a week wasted. Expect roughly two to four hours of writing for every hour observed at the start, falling as you learn what matters. Store notes securely and pseudonymize as you write rather than later. And keep the jottings themselves - the scrappy in-the-moment record often contains a phrase you will not have reproduced in the expansion.
Common misconceptions
Reactivity makes observation invalid. People cannot sustain performance indefinitely, and reactivity typically decays as the researcher becomes routine. More importantly, what people choose to perform for an observer is itself informative about what they think ought to happen.
Fieldnotes should capture everything. They cannot, and trying produces unusable volume. Notes are selective by design; the discipline is to make the selection principled and to record what governed it.
Description is atheoretical. Even "she raised her voice" reflects a decision about what was worth recording. The separation of description from inference is a working practice that keeps the record revisable, not a claim to have observed without a perspective.
Recap
- Observation reaches what participants cannot report: automatic routine, unarticulated norms, and the gap between stated and enacted practice.
- Gold's continuum trades depth against distance; going native and reactivity bracket its two ends.
- Spradley's dimensions and repeated-sequence observation give structure to what would otherwise be undirected watching.
- Sample time, events, and people deliberately, and report the sampling.
- Separate descriptive notes, observer's comments, methodological notes, and analytic memos; write the same day.
- Observational consent is layered - organizational approval, visible notification, individual consent for focal participants, and a usable opt-out.
Sources
- Mulhall, A. (2003). In the field: Notes on observation in qualitative research. Journal of Advanced Nursing, 41(3), 306-313. pubmed.ncbi.nlm.nih.gov
- Phillippi, J., & Lauderdale, J. (2018). A guide to field notes for qualitative research: Context and conversation. Qualitative Health Research, 28(3), 381-388. pubmed.ncbi.nlm.nih.gov
- Kawulich, B. B. (2005). Participant observation as a data collection method. Forum: Qualitative Social Research, 6(2). qualitative-research.net
- Ciesielska, M., Bostrom, K. W., & Ohlander, M. (2018). Observation methods. In M. Ciesielska & D. Jemielniak (Eds.), Qualitative methodologies in organization studies (Vol. 2, pp. 33-52). Palgrave Macmillan. link.springer.com
- Savage, J. (2000). Ethnography and health care. BMJ, 321(7273), 1400-1402. pmc.ncbi.nlm.nih.gov
- Sanjari, M., Bahramnezhad, F., Fomani, F. K., Shoghi, M., & Cheraghi, M. A. (2014). Ethical challenges of researchers in qualitative studies: The necessity to develop a specific guideline. Journal of Medical Ethics and History of Medicine, 7, 14. pmc.ncbi.nlm.nih.gov
- Emerson, R. M., Fretz, R. I., & Shaw, L. L. (2011). Writing ethnographic fieldnotes (2nd ed.). University of Chicago Press. find source ↗
- Key terms
- Observation
- Disciplined, purposeful, recorded watching of behavior in its natural context.
- Participant-as-observer
- A role in which the researcher participates in the setting while their research role is known.
- Going native
- Over-identifying with the group until analytic distance is lost.
- Reactivity
- The observer effect: people behaving differently because they know they are watched.
- Fieldnotes
- The written record of observation, expanded from in-the-moment jottings as soon as possible.
- Description versus inference
- The rule of recording concretely what was seen and heard separately from one's interpretation of it.
Documents and Artifacts
- Classify documentary and material sources and explain their advantages as unobtrusive data.
- Critically appraise a document's authenticity, credibility, representativeness, and meaning.
Not all qualitative data are generated by the researcher through talk or observation. A vast body of evidence already exists in the world as documents and artifacts - texts and objects produced for purposes other than your research - and learning to use them extends and corroborates what interviews and fieldwork provide.
Key idea: Lindsay Prior's central argument is worth internalizing early: documents should be studied as actors in a setting, not only as containers of content. A care plan is not merely a record of a decision; it is a thing that gets filled in, ignored, copied forward, audited, and cited in disputes. Asking what a document does - who is required to produce it, who reads it, what happens when it is absent - often yields more than asking what it says.
What counts as a document or artifact
- Public and official records: policies, meeting minutes, reports, laws, organizational charts, statistical returns.
- Personal documents: letters, diaries, emails, social-media posts, photographs.
- Media and popular texts: newspapers, advertisements, websites, broadcasts.
- Material artifacts: physical objects, tools, buildings, and the arrangement of spaces, which carry meaning about the culture that made and used them.
A useful distinction separates sources the researcher elicits (a diary you ask a participant to keep) from those that exist independently of the study. The latter are especially valuable as unobtrusive data.
Why documents are valuable
Documentary sources are non-reactive: because they were not produced for your study, they are unaffected by the observer effect that can distort interviews and observation. They provide historical depth, letting you study a past you could not observe and trace change over time. They are often efficient and stable - you can return to the same text repeatedly - and they excel at corroboration, allowing triangulation against what people tell you. When a manager's account of a decision conflicts with the contemporaneous minutes, the discrepancy is itself a finding.
Glenn Bowen's account of document analysis identifies several specific functions worth naming in a methods section rather than lumping under "we also reviewed documents." Documents can supply context that participants assume; they can generate questions for interviews that would not otherwise arise; they can supply supplementary data unavailable by other means; they can track change and development across versions; and they can verify or challenge findings from other sources. Saying which of these your documents are doing forces a discipline that a vague appendix listing does not.
Critical appraisal: four questions
Documents are not transparent windows onto fact; every one was made by someone, for some purpose, from some perspective. Scott's widely used framework appraises any source along four criteria:
- Authenticity. Is the document genuine and of unquestioned origin - is it what it purports to be, and who actually wrote it?
- Credibility. Is it free from error and distortion - was the author sincere and in a position to know, or did their interests shape the content?
- Representativeness. Is the document typical of its kind, and if not, is its untypicality known? Surviving records are a biased remnant; what was discarded or never written skews the archive.
- Meaning. Is the document clear and comprehensible, and do you understand it as its makers and original audience would have - its terms, conventions, and context?
Analyzing documents
Once appraised, documents are analyzed with the same interpretive tools as other qualitative data - close reading, coding, thematic and narrative analysis - always attending to the silences as well as the content: what a document omits, and whose voice is absent from the record, is often as telling as what it says. Treated critically, documents and artifacts are not a second-rate substitute for talking to people but a distinct and powerful source that anchors interview and field data in the durable traces a social world leaves behind.
A systematic procedure: the READ approach
Sarah Dalglish and colleagues offer a four-step procedure that turns document work from an informal gathering into a defensible method. R - Ready your materials. Define the corpus explicitly: what types of document, from what period, from which sources, with what inclusion and exclusion criteria, and what you did when a document could not be obtained. E - Extract data. Use a structured extraction sheet capturing, for every document, its type, author, intended audience, date, provenance, and the content relevant to your questions - so that the corpus can be interrogated systematically rather than remembered impressionistically. A - Analyze the data. Apply your analytic approach (thematic, content, discourse) to the extracted material, and analyze the corpus as a set, attending to what changes across versions and what is systematically missing. D - Distill your findings. Integrate with other data sources, state which claims rest on documents alone, and be explicit about the archive's limits.
Two habits improve this further. Keep a document log recording how each item was obtained, since provenance is part of the evidence and an unsourced PDF is not citable. And version-control your corpus: policy documents are frequently revised silently, and a study that analyzed the June version should say so.
A worked illustration: reading three versions of one policy
A study of how a university responded to reports of harassment assembles the procedure document across three revisions.
2019 version: "A complaint should normally be raised within three months of the incident. The Investigating Officer will interview the complainant, the respondent, and any witnesses."
2021 version: "A complaint must normally be raised within three months. The Investigating Officer may interview witnesses where this is proportionate."
2023 version: "Complaints raised more than three months after the incident will not normally be progressed. Informal resolution should be attempted in the first instance."
A content summary would say: the policy sets a three-month limit and describes an investigation. The analytic reading notices the grammar. The complainant's obligation hardens from should to must to an outright bar, while the institution's obligation softens from will to may and then is displaced by an informal route. The modality has moved in opposite directions for the two parties across four years. The passive construction in 2023 ("will not normally be progressed") removes the actor from the refusal entirely. And an absence is visible across all three: no version specifies what happens if the respondent holds power over the complainant.
That reading generates interview questions that could not have come from elsewhere: Was the 2021 change to witness interviews discussed? What prompted the move to informal resolution first? When interviews reveal that staff believe the three-month rule is discretionary while the 2023 text is close to absolute, the divergence between the text and its lived interpretation becomes a finding in its own right. Notice how much of the analysis lives in modal verbs, voice, and omission rather than in topics - documentary analysis rewards attention to form.
Elicited documents and material artifacts
Documents you ask participants to produce sit in a different category and have different strengths. Solicited diaries capture experience close to the moment rather than in retrospect, which matters greatly for phenomena that are episodic, fluctuating, or hard to recall accurately - pain, mood, symptom management, shift work. They also shift some control to the participant, who decides what to record. Their weaknesses are attrition, reactivity (recording changes the thing recorded), and unevenness in what different participants produce. The usual mitigations are a short structured prompt rather than a blank page, a bounded period of one to three weeks, a mid-point check-in, and a follow-up diary-interview in which participant and researcher read the diary together, which both clarifies entries and surfaces what was deliberately left out.
Material artifacts are underused in doctoral work and repay attention. The arrangement of a waiting room, the height of a reception counter, the presence or absence of a lock on a staff-room door, the laminated sign taped over the official sign - each encodes a decision about who is expected to do what. Photo-elicitation, in which participants photograph aspects of their setting and then discuss the images, gives access to material environments the researcher cannot enter and lets participants direct attention to what matters to them. When artifacts are used, describe them with the same low-inference discipline that governs fieldnotes: record what is physically there before recording what you think it means.
Digital documents and their traps
Social media posts, forum threads, and app data are documents, and they raise problems the classic categories do not fully cover. Publicly accessible does not mean public in the ethically relevant sense: users posting in a support forum have an expectation about audience that a research publication violates, and verbatim quotation of a public post is often traceable back to its author through a search engine, defeating pseudonymization. Common responses include seeking platform-level and, where feasible, individual permission, paraphrasing rather than quoting where the analysis permits, and treating small or closed communities as private regardless of technical accessibility. Provenance is also harder: content is edited and deleted, so archive what you analyze with a retrieval date, and note that the corpus you built is not reproducible from the live platform.
Common misconceptions
Documents are objective because nobody made them for your study. Non-reactivity removes the observer effect; it does not remove the author's purpose. An incident report is written by someone anticipating how it will be read.
An archive is a record of what happened. It is a record of what was written down and then survived. Both filters are systematic, and the second usually favors the powerful.
Document analysis is a supplementary method. It can be the primary or sole method, and there are strong document-only studies. Where documents are primary, the corpus definition and appraisal must be correspondingly rigorous.
Recap
- Study documents as actors that do things in a setting, not only as containers of content.
- Name the function documents serve in your design: context, question generation, supplementary data, tracking change, or verification.
- Appraise every source for authenticity, credibility, representativeness, and meaning.
- The READ approach - ready, extract, analyze, distill - makes documentary work systematic and reportable.
- Analyze form as well as content: modality, voice, and omission often carry the finding.
- Digital documents complicate consent, traceability, and provenance; accessibility is not consent.
Sources
- Dalglish, S. L., Khalid, H., & McMahon, S. A. (2021). Document analysis in health policy research: The READ approach. Health Policy and Planning, 35(10), 1424-1431. pmc.ncbi.nlm.nih.gov
- Morgan, H. (2022). Conducting a qualitative document analysis. The Qualitative Report, 27(1), 64-77. nsuworks.nova.edu
- Elo, S., & Kyngas, H. (2008). The qualitative content analysis process. Journal of Advanced Nursing, 62(1), 107-115. pubmed.ncbi.nlm.nih.gov
- Crowe, S., Cresswell, K., Robertson, A., Huby, G., Avery, A., & Sheikh, A. (2011). The case study approach. BMC Medical Research Methodology, 11, 100. pmc.ncbi.nlm.nih.gov
- Ruiz Ruiz, J. (2009). Sociological discourse analysis: Methods and logic. Forum: Qualitative Social Research, 10(2). qualitative-research.net
- Bowen, G. A. (2009). Document analysis as a qualitative research method. Qualitative Research Journal, 9(2), 27-40. find source ↗
- Scott, J. (1990). A matter of record: Documentary sources in social research. Polity Press. find source ↗
- Key terms
- Documents and artifacts
- Texts and objects produced for purposes other than the research, used as qualitative data.
- Unobtrusive data
- Data whose production was not affected by the research, avoiding the observer effect.
- Authenticity (documents)
- Whether a document is genuine, of sound origin, and what it purports to be.
- Credibility (documents)
- Whether a document is free from error and distortion given its author's sincerity and position.
- Representativeness (documents)
- Whether a document is typical of its kind, given that surviving records are a biased remnant.
- Meaning (documents)
- Whether the document is understood as its makers and original audience would have understood it.
Module 5: Qualitative Analysis
Turning raw data into defensible interpretation through systematic coding, thematic analysis, grounded-theory procedures, and disciplined use of memos and software.
Coding and Thematic Analysis
- Explain what a code is and distinguish inductive from deductive and descriptive from interpretive coding.
- Walk through Braun and Clarke's phases of reflexive thematic analysis and separate a theme from a code.
Analysis is where qualitative research is won or lost, and it is the phase most often done badly - reduced to plucking a few vivid quotations that confirm what the researcher already believed. Rigorous analysis is systematic: it works through the whole corpus, assigns meaning transparently, and builds interpretations that others could follow. The foundational technique is coding, and the most widely used framework built on it is thematic analysis.
Key idea: Coding is not analysis; it is the preparation that makes analysis possible. A coded corpus is an indexed corpus. The analytic work happens when you ask what the codes, taken together, allow you to say - and that work cannot be delegated to software, to a coding frame, or to a second coder.
What a code is
A code is a short label - a word or phrase - that captures the essence or salient meaning of a segment of data. Coding is the process of attaching such labels systematically across the data, so that the many pages of a corpus become organized by meaning rather than by their original order. Codes can be distinguished along two axes:
- Inductive (data-driven) codes arise from the data themselves, capturing what is there without a prior template. Deductive (theory-driven) codes come from a pre-existing framework applied to the data. Most studies blend the two.
- Descriptive (or semantic) codes label the surface, explicit content ("delayed diagnosis"). Interpretive (or latent) codes capture underlying meanings, assumptions, or ideas the analyst reads beneath the surface ("erosion of trust in the system").
Coding usually proceeds in cycles: a first cycle assigns many initial codes close to the data; a second cycle groups, merges, and organizes those codes into higher-order categories. The transcript is typically coded line by line or segment by segment so that nothing is skimmed.
Johnny Saldana's coding manual catalogues dozens of first-cycle methods, and knowing a handful by name lets you choose deliberately rather than defaulting. In vivo coding uses the participant's own words as the code label, preserving their vocabulary and useful early in any study. Process coding uses gerunds - "waiting," "deflecting blame," "rationing attention" - to keep action and sequence in view, and is standard in grounded theory. Descriptive coding labels the topic of a passage in a noun, which is useful for indexing a large corpus but produces topics rather than meanings if used alone. Values coding tags expressions of values, attitudes, and beliefs. Emotion coding tags stated or inferred feeling. Versus coding names oppositions ("us versus management") and is particularly productive in conflict settings. Second-cycle methods - pattern coding, focused coding, axial coding - reorganize first-cycle output into categories. Announcing "we used in vivo and process coding in the first cycle and pattern coding in the second" tells a reader far more than "the data were coded."
How much to code is a real decision. Complete coding works through every line, which suits inductive work and prevents the analyst from noticing only what confirms early impressions; it is slow and generates codes that later prove irrelevant. Selective coding works only on material relevant to the research question, which is defensible when the corpus is large and the question narrow, provided the selection rule is stated. What is not defensible is coding until you have enough for the themes you already intended to write.
Codes are not themes
A pervasive confusion equates codes with themes. They are different in kind. A code is a label on a data segment; a theme is a broader pattern of shared meaning, organized around a central idea, that the analyst constructs from clustered codes to say something significant about the data in relation to the question. A list of codes is not findings; a theme is an analytic claim. A common weakness is to present a "theme" that is really just a topic summary or a single code renamed. A genuine theme has a central organizing concept and is evidenced across multiple participants or data segments.
Reflexive thematic analysis
Braun and Clarke's widely adopted approach lays out six recursive phases (you move back and forth, not straight through):
- Familiarization: read and re-read the whole data set, noting initial ideas.
- Generating initial codes: systematically code interesting features across the entire corpus.
- Constructing themes: cluster codes into candidate themes organized by shared meaning.
- Reviewing themes: check candidate themes against the coded extracts and the whole data set; split, merge, or discard.
- Defining and naming themes: pin down the essence and scope of each theme and give it a clear name.
- Producing the report: weave themes into an analytic narrative supported by vivid, well-chosen extracts and tied back to the question and literature.
Two points give this approach its rigor. First, it is reflexive: themes are understood as actively constructed by the analyst engaging the data, not passively "emerging" from it as if lying in wait - so the analyst's interpretive role is owned, not disguised. Second, the constant checking in phase four disciplines the analysis against the temptation to keep only confirming evidence. Worked honestly and transparently, thematic analysis turns a mass of text into a defensible, evidenced argument about meaning.
Braun and Clarke have become insistent that "thematic analysis" is not one method but a family with incompatible assumptions, and that studies frequently mash them together. Coding reliability approaches (associated with Boyatzis and with Guest and colleagues) treat coding as a measurement problem, use a structured codebook, employ multiple coders, and report agreement statistics. Codebook approaches (framework analysis, template analysis) use a structured codebook developed early but treat coding as interpretive rather than as measurement, and are well suited to applied and team-based work with deadlines. Reflexive TA rejects codebooks and coder agreement outright, on the ground that if meaning is generated by an interpreting analyst, two analysts converging proves only that they were trained alike. Citing Braun and Clarke while reporting a kappa coefficient is therefore a specific, identifiable error, and their recent critical reviews of published work document how common it is.
Neighbouring approaches worth distinguishing
Thematic analysis is not the only way to work systematically through a corpus, and naming the alternative you did not choose sharpens a methods chapter. Qualitative content analysis, as set out by Satu Elo and Helvi Kyngas, works either inductively (open coding, grouping, abstraction) or deductively (applying a categorization matrix derived from prior theory), stays closer to manifest content, and is comfortable reporting frequencies within a defined corpus. Framework analysis, developed for applied policy research, moves through familiarization, identifying a thematic framework, indexing, charting into a matrix of cases by themes, and mapping and interpretation; its charting stage makes case-by-case comparison unusually visible and it suits team-based work against a deadline. Template analysis begins with a small set of a priori themes developed on a subset of data and then revises the template as coding proceeds, which is useful when some constructs are known in advance but the study remains open to what it did not anticipate. Each of these produces a different object, and none is a lesser version of the others.
A worked illustration: one excerpt, coded and clustered
From a study of newly qualified teachers. The participant says:
"I asked my mentor about behaviour in period five and she said 'you'll find your own way.' Which sounds supportive. But I've started not asking, because asking twice means you didn't find it. So now I go home and google it, and I'm learning to teach from strangers on the internet at eleven at night, and nobody knows that's happening."
First-cycle codes, with method labels attached:
- "you'll find your own way" (in vivo) - the mentor's formulation, preserved because it recurs across three other participants
- reinterpreting support as evaluation (process)
- withholding questions to manage impression (process)
- substituting anonymous sources for institutional help (process)
- invisibility of the substitution (process)
- autonomy as an expected competence (values)
- shame at needing to ask (emotion, partly inferred - flagged in a memo as inference)
Notice what would be lost by descriptive coding alone. "Mentoring" and "behaviour management" are accurate topic labels and analytically inert; they would file this excerpt alongside a passage about a different phenomenon entirely.
Now the clustering. Across the corpus, "withholding questions to manage impression" co-occurs with codes such as rehearsing questions before asking, timing questions for when the staffroom is empty, and asking peers rather than seniors. Grouped, these become a candidate theme. A weak name would be "Mentoring" (a topic) or "Barriers to seeking help" (a category, and one that locates the problem in the newcomer). A stronger name states the central organizing concept: "Asking twice means you didn't find it": help-seeking as an admission of unfitness. That formulation names a shared meaning, is evidenced across participants, and makes an analytic claim - the institution's language of autonomy converts a support relationship into an assessment relationship, and the resulting learning happens invisibly and unsupported.
Then the phase-four discipline. Return to the full corpus and look for cases that do not fit. Two participants describe mentors who normalized repeated questions, and both describe asking freely. That does not defeat the theme; it specifies it - the mechanism depends on how the mentor frames repetition, which is a modifiable condition and therefore a more useful finding than a blanket claim about newcomers.
Common misconceptions
Themes emerged from the data. Braun and Clarke single this phrase out as a passive-voice disavowal of the analyst's role. Themes are constructed, and saying so is more honest and more defensible.
A high inter-rater kappa means the analysis is trustworthy. It means two people trained on the same codebook applied it similarly. Within reflexive TA it is a category error; within coding-reliability TA it is appropriate. The problem is claiming one framework and using the other's criteria.
Themes should be reported by frequency. Counting how many participants a theme appears in can be informative, but prevalence is not importance; a pattern voiced by three participants may be the most consequential finding in the study.
A theme is a topic that came up a lot. A topic summary answers "what did people talk about." A theme answers "what does this mean, and what is my claim about it." The test is whether the theme name contains an idea or only a subject.
Recap
- Coding indexes the corpus; analysis is the argument you build from it.
- Choose and name coding methods deliberately - in vivo, process, values, emotion, versus - rather than defaulting to descriptive topic labels.
- Codes label segments; themes are patterns of shared meaning with a central organizing concept, evidenced across the data.
- Reflexive, codebook, and coding-reliability thematic analysis rest on different assumptions; mixing their criteria is a common and identifiable error.
- Phase four - checking candidate themes against the whole corpus, including disconfirming cases - is where credibility is earned.
- Prevalence is not importance; a theme name should carry an idea, not just a subject.
Sources
- Braun, V., & Clarke, V. (2023). Toward good practice in thematic analysis: Avoiding common problems and be(com)ing a knowing researcher. International Journal of Transgender Health, 24(1), 1-6. pmc.ncbi.nlm.nih.gov
- Braun, V., & Clarke, V. (2024). A critical review of the reporting of reflexive thematic analysis in Health Promotion International. Health Promotion International, 39(3), daae049. pmc.ncbi.nlm.nih.gov
- Byrne, D. (2022). A worked example of Braun and Clarke's approach to reflexive thematic analysis. Quality & Quantity, 56(3), 1391-1412. link.springer.com
- Kiger, M. E., & Varpio, L. (2020). Thematic analysis of qualitative data: AMEE Guide No. 131. Medical Teacher, 42(8), 846-854. pubmed.ncbi.nlm.nih.gov
- Vaismoradi, M., Turunen, H., & Bondas, T. (2013). Content analysis and thematic analysis: Implications for conducting a qualitative descriptive study. Nursing & Health Sciences, 15(3), 398-405. pubmed.ncbi.nlm.nih.gov
- Braun, V., & Clarke, V. (2022). Thematic analysis: A practical resource. The University of Auckland. thematicanalysis.net
- Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77-101. find source ↗
- Key terms
- Code
- A short label capturing the salient meaning of a segment of qualitative data.
- Inductive vs deductive coding
- Codes arising from the data versus codes applied from a pre-existing framework.
- Semantic vs latent code
- A code of explicit surface content versus one of underlying meaning read beneath the surface.
- Theme
- A broader pattern of shared meaning organized around a central idea, constructed from clustered codes.
- Reflexive thematic analysis
- Braun and Clarke's six-phase approach treating themes as actively constructed by the analyst.
- Familiarization
- The first phase of thematic analysis: immersive reading and re-reading of the whole data set.
Grounded-Theory Analysis in Practice
- Work through open, axial/focused, and theoretical/selective coding in grounded-theory analysis.
- Explain how memoing and constant comparison build an integrated theory around a core category.
Module 2 introduced grounded theory as a tradition; this lesson works through its analytic engine, because grounded-theory coding is more specific and more theory-directed than the thematic coding of the previous lesson. Its whole purpose is to move systematically from raw incidents to an integrated theory that explains a process. The coding proceeds in phases, with terminology differing across schools; the logic is shared.
Key idea: Grounded-theory coding is not a more thorough version of thematic coding. It is coding aimed at a specific output - a set of categories with specified properties and dimensions, integrated around a core process. Every coding decision should be interrogated with the question: does this move me toward explaining what is going on, or does it only file the data more neatly?
A vocabulary note, because the schools use the same words differently and committees notice. Glaser's classic scheme runs open coding, then selective coding (coding only for the core category and what relates to it), then theoretical coding (specifying the relationships using theoretical codes such as causes, conditions, consequences, and types). Strauss and Corbin run open, axial, and selective coding, where axial coding uses their paradigm and selective coding integrates around the core. Charmaz runs initial coding and focused coding, followed by theoretical coding, and treats axial coding as optional rather than obligatory. Use one vocabulary consistently and say which.
Open (initial) coding
Analysis begins with open coding: examining the data closely - often line by line - and attaching provisional codes to incidents, actions, and meanings, staying close to what is there. A distinctive Glaserian move is to code with gerunds (words ending in "-ing": negotiating, resisting, reassuring), because naming actions and processes rather than static topics keeps the analysis oriented toward the process that grounded theory seeks. Throughout, constant comparison operates: each new incident is compared with earlier ones and with the codes already made, so codes are continually tested and refined. In vivo codes - using participants' own striking words as code labels - help preserve their meanings.
Axial or focused coding
The many open codes are then consolidated. In the Strauss-and-Corbin tradition this is axial coding: reassembling the data by making connections between categories and subcategories, often using a coding paradigm that relates conditions, actions/interactions, and consequences to show how a category operates. In Charmaz's constructivist version the parallel step is focused coding: selecting the most significant and frequent earlier codes and using them to sift and organize larger amounts of data. Either way, the analyst is raising the level of abstraction - from many small codes toward a smaller number of well-developed categories with specified properties.
Two terms deserve precision here because they are what "well-developed" means. A category's properties are its characteristics - the attributes that describe it. Its dimensions are the ranges along which those properties vary. If the category is rationing attention, its properties might include the basis of rationing (clinical, relational, defensive), its visibility to others (open, concealed), and its emotional cost (negligible to severe). Each of those is a dimension with a range, and specifying them is what makes a category capable of explaining variation rather than merely naming a phenomenon. A study that reports six categories with no properties or dimensions has produced labels, not a theory.
Axial coding has been contested. Judith Kendall's much-cited critique argued that the coding paradigm can crowd out the analyst's own thinking, and Glaser's objection is the same in stronger terms. The defensible practice is to treat the paradigm as a set of prompts - what conditions make this happen, what do people do, with what consequences - rather than as a form to complete for every category.
Selective or theoretical coding and the core category
The final phase integrates the categories into a coherent theory. The analyst identifies a core category - the central concept that appears frequently, connects to the most other categories, and best accounts for the main process going on - and relates the other categories to it, so the analysis coheres around a single explanatory story rather than a loose set of themes. This is selective or theoretical coding. The recurring test is whether you can state, in a sentence or two, the core process and how the surrounding categories condition and flow from it.
Memoing: the theory is written in the memos
Binding all three phases together is memoing. From the first codes onward, the analyst writes analytic memos - freewritten notes that define a code, compare it with others, speculate about relationships, and record puzzles and hunches. Memos are not administrative notes; they are where the theorizing actually happens. By the end, the memos, sorted and sequenced, form the skeleton of the written theory.
A grounded-theory study whose author coded diligently but never memoed typically ends with categories and no theory, because the conceptual work that turns categories into an explanation lives in the memos. Constant comparison, theoretical sampling (from Module 2), memoing, and integration around a core category are the interlocking parts of one machine whose output is theory, not description.
Sorting: the step nobody describes
Between memoing and writing sits a physical step that method texts mention and rarely explain. Memo sorting means laying the memos out - literally, on a table or a wall, or in a document outline - and arranging them into the sequence in which the theory will be told. The act of sorting is analytic: memos that will not sit anywhere reveal categories that do not yet connect to the core; memos that keep wanting to sit in two places usually indicate that one category is doing two jobs and needs splitting; and a pile that has no memos in it exposes a gap where you assumed a relationship you never evidenced. Glaser regarded sorting as indispensable and warned against skipping it in favour of writing straight from a code list, on the grounds that the code list preserves the order of discovery rather than the order of explanation. Practically, sort when you have thirty or forty substantive memos, do it in one sitting, and write down the sequence you arrive at before the arrangement is disturbed. The resulting outline is the theory's structure, and the writing that follows is largely a matter of connecting memos with prose.
A worked illustration: from line-by-line codes to a core category
From a study of how people manage type 1 diabetes at work. A participant says:
"I test in the car. Not because I need to - I could do it at my desk in about eight seconds. But then it's a thing. Then Karen asks if I'm alright, every time, and she means it kindly, and then I'm the diabetic one. So I go to the car. It costs me maybe four minutes each time and I do it three times a day."
Line-by-line open codes (gerunds): relocating routine care; distinguishing need from choice; anticipating a colleague's concern; classifying kindness as a cost; resisting a category; accepting a time penalty; quantifying the penalty.
Memo, written immediately: "The interesting move is that concern is experienced as a cost. She is not hiding a stigmatized condition from hostility - Karen is kind. She is managing being categorized. Compare with P2 (announces it loudly on day one) and P6 (told only her line manager, on a form). All three are managing the same thing differently. Candidate category: governing disclosure. Properties: audience (individual, team, institutional), timing (pre-emptive, reactive, indefinite deferral), and framing (medical fact, personal information, non-event). Dimension to check: does the strategy change when a colleague also has a chronic condition?"
Theoretical sampling that follows: recruit two participants who work alongside someone with a visible chronic condition, and two who work alone or remotely, since the category should behave differently where there is no audience.
Focused coding across the corpus promotes "governing disclosure" and, from other transcripts, "pre-empting the story" and "trading privacy for allowance." Constant comparison shows all three concern who gets to author the meaning of the condition.
Core category: authoring the condition - the ongoing work of retaining control over what one's illness means to others at work, of which disclosure decisions, concealment routines, and pre-emptive narration are strategies. Conditions that intensify it include a small team, high visibility of the role, and a workplace culture of expressed personal concern. Consequences include time and energy costs that are invisible to employers and absent from adjustment policies.
Now apply the test. Can the theory be stated in two sentences? Yes: people with type 1 diabetes at work engage in continuous authoring of what their condition means to colleagues, using disclosure, concealment, and pre-emptive narration as strategies; the intensity of this work rises with team intimacy and role visibility, and its costs are invisible to the organization because they fall outside anything the organization measures. Does it explain variation rather than describe experiences? Yes - it predicts that the isolated remote worker and the member of a close team will differ systematically, and that prediction can be checked. Would it survive the removal of the illustrative quotations? Yes, which is the mark of a theory rather than a collection of vivid extracts.
Common misconceptions
The core category is the most frequent code. Frequency is one signal among several. The core category must account for variation, relate to most other categories, and have explanatory reach. A frequent but inert code is not a core category.
Line-by-line coding must be used throughout. It is most valuable early, on the first few transcripts, to prevent skimming and to generate a rich code set. Later transcripts are usually coded more selectively, guided by the developing categories - which is theoretical sampling operating on the data you already have.
Memos are for later. A memo written three weeks after the insight has lost the insight. Memo when the thought occurs, even at two sentences, and date every one.
Theoretical sensitivity means having a theory in advance. Glaser's term means the analyst's capacity to see conceptual significance in data - developed through disciplinary reading, professional experience, and analytic practice - not the imposition of a chosen framework.
Recap
- The schools use overlapping vocabulary differently; adopt one consistently and name it.
- Gerund and in vivo coding early keeps the analysis on process and in participants' terms.
- Well-developed categories have specified properties and dimensions; without them you have labels.
- Axial coding's paradigm is a prompt, not a form to complete; its critics warn it can displace the analyst's thinking.
- The core category must explain variation and connect the others, not merely occur often.
- Memos written at the moment of insight are where the theory is actually built.
Sources
- Chun Tie, Y., Birks, M., & Francis, K. (2019). Grounded theory research: A design framework for novice researchers. SAGE Open Medicine, 7. pmc.ncbi.nlm.nih.gov
- Charmaz, K. (2015). Teaching theory construction with initial grounded theory tools: A reflection on lessons and learning. Qualitative Health Research, 25(12), 1610-1622. pubmed.ncbi.nlm.nih.gov
- Conlon, C., Timonen, V., Elliott-O'Dare, C., O'Keeffe, S., & Foley, G. (2020). Confused about theoretical sampling? Engaging theoretical sampling in diverse grounded theory studies. Qualitative Health Research, 30(6), 947-959. pubmed.ncbi.nlm.nih.gov
- Foley, G., & Timonen, V. (2015). Using grounded theory method to capture and analyze health care experiences. Health Services Research, 50(4), 1195-1210. pmc.ncbi.nlm.nih.gov
- Kelle, U. (2005). "Emergence" vs. "forcing" of empirical data? A crucial problem of "grounded theory" reconsidered. Forum: Qualitative Social Research, 6(2). qualitative-research.net
- Birks, M., Chapman, Y., & Francis, K. (2025). Memoing in qualitative research: Two decades on. Journal of Research in Nursing. pmc.ncbi.nlm.nih.gov
- Corbin, J., & Strauss, A. (2015). Basics of qualitative research: Techniques and procedures for developing grounded theory (4th ed.). Sage Publications. find source ↗
- Key terms
- Open coding
- Initial close, often line-by-line coding of incidents and actions, staying near the data.
- Gerund coding
- Coding with '-ing' action words to keep the analysis oriented toward process.
- In vivo code
- A code that uses a participant's own striking words as its label.
- Axial/focused coding
- Consolidating open codes and connecting categories to raise the level of abstraction.
- Selective/theoretical coding
- Integrating categories around a core category into a coherent explanatory theory.
- Core category
- The central concept that connects to the most others and best explains the main process.
Analytic Tools: Memos, Displays, and CAQDAS
- Use analytic memos and data displays to move from coding to interpretation.
- State accurately what qualitative analysis software does and does not do.
Coding organizes data, but organization is not yet interpretation. This lesson covers the craft tools that bridge the gap - memos and displays - and clarifies the proper role of the software that many students mistake for an analysis engine.
Key idea: The gap between a coded corpus and a finding is crossed by writing and by looking. Memos are how you write your way to an idea; displays are how you make the corpus visible enough to notice what you would otherwise miss. Both leave a durable record, which is why they double as the evidentiary basis for claims about dependability and confirmability.
Analytic memos across all traditions
Although memoing is most codified in grounded theory, analytic memoing serves every qualitative approach. A memo is a dated, informal piece of analytic writing in which you think on the page: What does this code really mean? How does it relate to that one? What surprised me here, and why? What might explain this pattern, and what would count against my explanation? Writing forces half-formed intuitions into explicit claims that can be examined and, importantly, questioned. Keeping memos throughout also builds the audit trail that later evidences the study's dependability, by making the analytic reasoning visible rather than locked in the researcher's head.
It helps to distinguish types of memo, because "keep memos" is advice too vague to act on. Code memos define a code, give its inclusion and exclusion criteria, and cite an anchor example and a borderline one. Theoretical memos propose relationships between categories and speculate about mechanisms. Methodological memos record decisions about sampling, guide revisions, and analytic procedure, and are the raw material of a defensible methods chapter. Reflexive memos record the researcher's reactions, discomforts, and shifting assumptions, and are what a reflexivity section is later written from. Operational memos hold to-do items so they do not contaminate the analytic ones. Some analysts also keep integrative memos that attempt, at intervals, to state the whole argument in a page - an exercise that reliably exposes which parts are not yet thought through.
A worked memo. Here is what a usable code memo actually looks like, from a study of interpreters in mental health settings:
"Code: EDITING FOR THE CLINICIAN. Dated 3 May. Definition: instances where the interpreter reports altering, condensing, or reordering a patient's speech in order to make it usable to the clinician, as distinct from ordinary linguistic transformation. Include: omitting repetition the interpreter judges clinically irrelevant; converting circular narrative into chronological order; softening insult. Exclude: acknowledged errors; omissions due to not hearing. Anchor: I7, 'he was going round and round and she had four minutes, so I gave her the middle bit.' Borderline: I3, 'I don't change anything, I just make it make sense in English' - the denial and the description contradict each other, which may be a property of the code (unrecognized editing) rather than a reason to exclude. Relation to other codes: overlaps with PROTECTING THE PATIENT but differs in beneficiary - here the clinician's time is the object of care. Question for next interviews: do interpreters distinguish editing done for time from editing done for face? Counts against my reading: two interpreters describe refusing to condense, both in forensic settings where verbatim rendering is required - so setting may be a condition."
Roughly two hundred words, written in ten minutes, and it does five things at once: it fixes the code so that it can be applied consistently, it records a contradiction rather than smoothing it, it connects to a neighbouring code, it generates the next interview question, and it names disconfirming evidence. Fifty memos of this quality are a dissertation's analytic chapter in raw form.
Data displays
Miles and Huberman argued that extended prose is a weak format for seeing patterns, and that compressing data into displays - matrices, networks, and diagrams - aids valid analysis by putting relevant information where the eye can grasp it at once. Common displays include:
- Matrices: rows and columns that cross, say, participants against themes, so presence, absence, and variation become visible at a glance.
- Networks: nodes and links that map how concepts relate, useful for building and checking a process model.
- Conceptually ordered displays: tables that arrange cases or codes by an analytic dimension to reveal gradients and contrasts.
The value of a display is not decorative; assembling one forces analytic decisions (what belongs in each cell, what is missing) and exposes cases that break an emerging pattern - the negative cases that sharpen a theory.
Building a display that earns its keep. A common and immediately useful one is the case-ordered predictor-outcome matrix: rows are cases ordered by an outcome of interest, columns are the conditions you suspect matter, and cells hold condensed evidence rather than ticks. In a study of unit-level adoption of a new safety checklist, rows might be twelve wards ordered from highest to lowest sustained use, with columns for whether a local champion existed, whether the checklist was integrated into an existing routine, whether senior clinicians were observed using it, and how the change was first communicated. Filling this in does three things at once. It forces you to state what counts as evidence for each cell, which surfaces vagueness immediately. It makes gradients visible - if integration into an existing routine is present in the top five wards and absent in the bottom four, the eye catches it in a way that prose never would. And it makes the exceptions unmissable: the ward with every favourable condition and low use is the case you must now explain, and explaining it is usually where the study's contribution comes from. Keep the raw matrix in the audit trail even if a simplified version appears in the thesis.
What CAQDAS does - and does not - do
CAQDAS stands for Computer-Assisted Qualitative Data Analysis Software - packages that help manage and analyze qualitative data. The single most important thing to understand about it is a boundary: the software manages the analysis; it does not do the analysis. It is not the qualitative equivalent of a statistics package that computes a result.
What CAQDAS genuinely provides is powerful data management: storing and organizing large volumes of text, attaching codes the researcher decides on, retrieving all segments bearing a code in seconds, running queries about co-occurrence, linking memos to data, and displaying code structures. These capacities make analysis of a large corpus far more systematic, thorough, and transparent than paper-and-highlighter methods, and they create a retrievable record that strengthens the audit trail.
But the interpretive acts that constitute analysis - deciding what a segment means, naming a code, judging that several codes form a theme, discerning the core process - remain irreducibly the researcher's. A frequent novice error is to expect the software to reveal themes, or to believe that using a respected package by itself confers rigor. It does neither. Rigor comes from the quality of the analytic thinking, which the software can support and record but never supply. Choose the tool for its fit to your project and your data, and keep clearly in mind that the mind doing the interpreting is yours.
Two documented risks are worth naming. The first is code proliferation: because creating a code costs one click, software-based projects routinely accumulate several hundred codes, most applied once, producing an unusable hierarchy that has to be rebuilt. The discipline is to consolidate deliberately at intervals and to write a code memo before creating any new code, which raises the cost enough to make you think. The second is distance from the data: retrieving all 63 segments coded "trust" and reading them as a list detaches them from the transcripts that gave them meaning, and analysts working this way frequently produce accounts that are technically supported and interpretively thin. The remedy is to return to whole transcripts regularly, and to treat retrieval output as a prompt to reread rather than as the material itself. There is also a quieter risk that software's affordances shape the analysis toward what it does easily - counting code frequencies, drawing hierarchies - and away from what it does not, such as attending to the sequence within a single interview.
Choosing a package. The major options - NVivo, ATLAS.ti, MAXQDA, Dedoose, and the open-source Taguette and QualCoder - differ less in analytic capability than in collaboration model, cost, platform, and handling of non-text data. Practical selection criteria: does your institution license it and support it; can your supervisor open your project file; does it handle the data types you have (video, PDFs with layout, non-Latin scripts); can you export your work in an open format if you leave the institution; and does the collaboration model fit your team. For a single researcher with thirty interviews, a well-organized set of documents and a spreadsheet can be entirely adequate, and saying so honestly is better than adopting a package for the appearance of rigor.
Reporting software use. Name the package and version, state what it was used for (data management, coding, retrieval, matrix queries), and state explicitly that interpretive decisions were made by the researchers. Do not write "NVivo was used to analyze the data" - it is both inaccurate and, to a knowledgeable reviewer, a signal that the analytic process was not well understood.
Common misconceptions
Memos are notes to yourself and need not be kept. Memos are data about your reasoning and form the audit trail; they are also the fastest route to a first draft, since sorted memos are already most of an analytic chapter.
A matrix of participants by themes reduces qualitative data to counts. Only if you fill the cells with ticks. Fill them with condensed content or short quotations and the matrix supports comparison while preserving meaning.
Using CAQDAS makes an analysis systematic. It makes retrieval systematic. Whether the analysis is systematic depends on whether you coded the whole corpus, engaged negative cases, and can show your reasoning.
Displays are for the final report. Most displays are working tools that never appear in the thesis. Their value is in the analytic decisions building them forces.
Recap
- Memos and displays are the bridge from a coded corpus to an interpretation, and they double as audit trail evidence.
- Distinguish code, theoretical, methodological, reflexive, operational, and integrative memos; a good code memo defines, anchors, contradicts, connects, and generates.
- Matrices, networks, and conceptually ordered displays force analytic decisions and expose negative cases.
- CAQDAS manages data; it does not interpret. Code proliferation and distance from whole transcripts are its documented risks.
- Choose software on licensing, collaboration, data types, and export, and report its use precisely rather than crediting it with the analysis.
Sources
- Birks, M., Chapman, Y., & Francis, K. (2025). Memoing in qualitative research: Two decades on. Journal of Research in Nursing. pmc.ncbi.nlm.nih.gov
- Zamawe, F. C. (2015). The implication of using NVivo software in qualitative data analysis: Evidence-based reflections. Malawi Medical Journal, 27(1), 13-15. pmc.ncbi.nlm.nih.gov
- Cope, D. G. (2014). Computer-assisted qualitative data analysis software. Oncology Nursing Forum, 41(3), 322-323. pubmed.ncbi.nlm.nih.gov
- Korstjens, I., & Moser, A. (2018). Series: Practical guidance to qualitative research. Part 4: Trustworthiness and publishing. European Journal of General Practice, 24(1), 120-124. pmc.ncbi.nlm.nih.gov
- Chun Tie, Y., Birks, M., & Francis, K. (2019). Grounded theory research: A design framework for novice researchers. SAGE Open Medicine, 7. pmc.ncbi.nlm.nih.gov
- Schlunegger, M. C., Zumstein-Shaha, M., & Palm, R. (2024). Methodologic and data-analysis triangulation in case studies: A scoping review. Western Journal of Nursing Research, 46(8), 611-622. pmc.ncbi.nlm.nih.gov
- Miles, M. B., Huberman, A. M., & Saldana, J. (2020). Qualitative data analysis: A methods sourcebook (4th ed.). Sage Publications. find source ↗
- Key terms
- Analytic memo
- Dated informal analytic writing in which the researcher thinks through codes, relationships, and explanations.
- Data display
- A compressed visual format - matrix, network, or diagram - that aids valid analysis of patterns.
- Matrix display
- A rows-and-columns table (for example participants by themes) revealing presence, absence, and variation.
- Negative case
- A case that breaks an emerging pattern, prompting the analyst to refine the theory.
- CAQDAS
- Computer-Assisted Qualitative Data Analysis Software that manages, but does not perform, analysis.
- Audit trail
- The documented analytic record, strengthened by memos and software, that evidences dependability.
Module 6: Rigor, Reflexivity, Ethics, and Writing
Establishing trustworthiness, practicing reflexivity and qualitative ethics, and communicating findings in persuasive, transparent scholarly prose.
Trustworthiness and Rigor
- State Lincoln and Guba's four trustworthiness criteria and their quantitative parallels.
- Match concrete strategies - triangulation, member checking, thick description, audit trail - to each criterion.
How do you know a qualitative study is any good? Applying quantitative yardsticks - internal validity, reliability, objectivity - misfits work that never claimed to measure a stable reality from a detached standpoint. Lincoln and Guba proposed a parallel framework of trustworthiness with four criteria, each answering, in qualitative terms, a concern that validity and reliability address in quantitative terms. Together they constitute the standard against which rigor is judged.
The framework has a history worth knowing. Egon Guba set out the four criteria in 1981, and he and Yvonna Lincoln developed them in Naturalistic Inquiry (1985), adding in 1989 a fifth set of authenticity criteria for constructivist evaluation - fairness (are all stakeholder constructions represented?), ontological authenticity (did participants' own understanding become more sophisticated?), educative authenticity (did they come to understand others' constructions?), catalytic authenticity (did the inquiry prompt action?), and tactical authenticity (were participants empowered to act?). Doctoral students almost always cite the four and almost never the five, which is a missed opportunity in participatory and critical work where the authenticity criteria are precisely the relevant ones.
Key idea: Naming a criterion is not meeting it. The sentence "credibility was established through member checking and triangulation" is worth nothing on its own. What is worth something is: which findings were returned to whom, in what form, what they said, and what changed as a result.
The four criteria
| Trustworthiness criterion | Question it answers | Quantitative parallel |
|---|---|---|
| Credibility | Are the findings a faithful interpretation of participants' realities? | Internal validity |
| Transferability | Could the findings apply to other contexts? | External validity |
| Dependability | Is the process consistent, traceable, and documented? | Reliability |
| Confirmability | Are the findings grounded in the data rather than the researcher's bias? | Objectivity |
Credibility and its strategies
Credibility is the qualitative counterpart of internal validity: the fit between participants' realities and the researcher's representation of them. Several strategies build it:
- Triangulation: corroborating findings across multiple data sources, methods, investigators, or theories, so a conclusion rests on more than one footing.
- Member checking (respondent validation): returning findings or interpretations to participants to ask whether they ring true to their experience.
- Prolonged engagement and persistent observation: spending enough time in the field to understand it and to distinguish the central from the incidental.
- Negative case analysis: actively seeking data that contradict the emerging interpretation and revising it to fit.
- Peer debriefing: exposing the analysis to a disinterested peer who probes assumptions and alternatives.
Two of these need unpacking, because both are routinely invoked and rarely understood.
Triangulation is not corroboration. Norman Denzin distinguished four types: data triangulation (different sources, times, or settings), investigator triangulation (multiple researchers), theory triangulation (multiple frameworks brought to the same data), and methodological triangulation (multiple methods). The naive reading is that agreement across sources proves a finding true, which quietly imports a realist assumption that there is one thing to be right about. The stronger reading, and the one most contemporary methodologists hold, is that triangulation gives a fuller and more textured account, and that divergence between sources is at least as informative as convergence - it tells you that different vantage points produce different constructions, which is usually the more interesting result. A study that reports only where sources agreed has probably suppressed its best material.
Member checking has real limits. Linda Birt and colleagues examined the practice closely and found it often functions as a ritual rather than a check. The problems are concrete. Participants are being asked to validate an analysis that operates at a level of abstraction above their own account, and they are not positioned to assess a cross-case theoretical claim. They may agree out of politeness, deference, or fatigue. They may disagree because the analysis is unflattering rather than inaccurate - and disagreement of that kind is data, not a verdict. Their views may have changed since the interview. And in critical or psychodynamic work, an interpretation participants reject may be exactly the finding.
The practical response is to be specific about what is being checked. Returning a transcript for correction checks the record. Returning a summary of that participant's own account checks whether you understood them, which is a legitimate and useful check. Returning developed themes or a theory is better framed as member reflection - a dialogue that generates further data and may complicate the analysis - than as validation. Birt and colleagues propose synthesized member checking, in which anonymized synthesized findings are returned so participants can respond to the analysis without being asked to adjudicate it. Whatever you do, report the response rate, what participants said, and what you changed. "All participants agreed with the findings" is, on its own, a warning sign rather than a reassurance.
Transferability
Transferability parallels external validity but reassigns the responsibility. The qualitative researcher does not claim their findings generalize; instead they provide thick, rich description of the context, participants, and phenomenon so that readers can judge whether the findings might transfer to their own settings. The burden of the generalizing judgment shifts from the author to the reader, who is given enough contextual detail to make it. This is why thick description is a rigor strategy, not merely a stylistic virtue.
Dependability and confirmability
Dependability parallels reliability: it asks whether the inquiry is consistent and could be traced. It is supported chiefly by an audit trail - a documented record of methodological decisions, changes, and analytic steps detailed enough that an external auditor could follow the logic from data to findings. Confirmability parallels objectivity: it asks whether the findings flow from the data and participants rather than the researcher's preferences.
It is supported by the same audit trail (now showing how each interpretation traces to evidence) and by reflexivity, in which the researcher discloses the assumptions and positions that could have shaped the work. Note how several strategies serve more than one criterion, and how the audit trail underwrites both dependability and confirmability. A rigorous qualitative study does not gesture at these criteria in a paragraph; it builds specific, reported practices for each into the design from the start, so that trustworthiness is demonstrated rather than asserted.
A worked illustration: two trustworthiness paragraphs
Weak: "Rigor was ensured through credibility, transferability, dependability, and confirmability (Lincoln & Guba, 1985). Triangulation, member checking, and an audit trail were used. Thick description supports transferability. The researcher engaged in reflexivity throughout."
Strong: "Credibility. Interview findings were triangulated against 16 months of team meeting minutes and against observation of nine handovers; where the minutes and interviews diverged on when the protocol changed, we pursued the discrepancy and report it in Finding 3. Two negative cases - participants whose accounts contradicted the developing theme of role ambiguity - are analyzed in the Findings rather than set aside, and their presence narrowed the theme's scope to units without a designated lead. Synthesized findings were returned to 11 of 14 participants; 7 responded, 5 endorsed the account, and 2 objected that the emphasis on ambiguity understated deliberate managerial withholding, which prompted a revision to the theme name and a new subsection. Transferability. We describe unit size, staffing ratio, case mix, the regulatory context, and the specific electronic record system, since colleagues have noted the last of these strongly conditions the practice described. Dependability. An audit trail comprising the decision log, 71 dated memos, three successive versions of the coding frame, and all interview guides is held in the project repository and was reviewed by a supervisor not involved in data collection. Confirmability. Each theme statement is linked in the audit trail to the coded extracts supporting it, including extracts that qualify it. The first author's prior employment in a similar unit is described in the Reflexivity section, with two specific instances where it shaped probing."
The second paragraph is longer, but every clause is checkable. That is the difference between a demonstrated claim and a recited list.
The critique: has trustworthiness become a checklist?
Janice Morse and others have argued that the Lincoln and Guba framework, as commonly used, has degenerated into a post hoc list of techniques applied after data collection, when the real determinants of rigor - an appropriate design, adequate and appropriate sampling, iterative analysis interwoven with data collection, and genuine analytic thinking - operate during the study. On this view, verification strategies belong in the conduct of the research, not in a paragraph appended to it.
Sarah Tracy offers a different response with her eight "big-tent" criteria - worthy topic, rich rigor, sincerity, credibility, resonance, significant contribution, ethics, and meaningful coherence - which broaden the question from procedural adequacy to whether the work matters and hangs together. Reporting guidelines have taken a third route: COREQ, a 32-item checklist for interview and focus group studies, and SRQR, a 21-item synthesis for qualitative research generally, standardize what must be disclosed. Their value is real, but so is a documented risk - that checklist compliance is mistaken for quality, and that studies which report every item can still be analytically thin.
The doctoral position to hold is that trustworthiness criteria are necessary and not sufficient. Build the practices into the design, report them concretely, and understand that the ultimate test is whether a knowledgeable reader finds your interpretation warranted by the evidence you show them.
Common misconceptions
Credibility, transferability, dependability, and confirmability are interchangeable ways of saying "rigorous." They answer distinct questions - faithfulness of interpretation, applicability elsewhere, traceability of process, and grounding in data - and require different evidence.
Participants agreeing with your findings validates them. Agreement may reflect politeness or deference, and disagreement may reflect discomfort rather than error. Report the response and what changed.
An audit trail is a folder of files. It is a folder of files organized so that an external reader could trace a specific finding back to specific evidence and see the decisions in between.
Reliability applies to qualitative work if you use two coders. Dependability is about traceable, documented process, not about replication of coding. Coder agreement is a criterion within coding-reliability approaches only.
Recap
- Guba and Lincoln's four criteria map onto, but do not import, the quantitative concerns of internal validity, external validity, reliability, and objectivity; the authenticity criteria extend them for participatory and critical work.
- Triangulation yields fullness rather than proof, and divergence between sources is a finding.
- Member checking has documented limits; distinguish checking a transcript, checking your understanding of one account, and inviting reflection on a theory.
- Transferability is the reader's judgment, enabled by contextual detail the researcher supplies.
- The audit trail underwrites both dependability and confirmability and must permit tracing a finding back to evidence.
- Rigor is built during the study; Morse's critique warns against treating trustworthiness as a post hoc checklist, and reporting guidelines standardize disclosure without guaranteeing quality.
Sources
- Guba, E. G. (1981). Criteria for assessing the trustworthiness of naturalistic inquiries. Educational Communication and Technology Journal, 29(2), 75-91. link.springer.com
- Korstjens, I., & Moser, A. (2018). Series: Practical guidance to qualitative research. Part 4: Trustworthiness and publishing. European Journal of General Practice, 24(1), 120-124. pmc.ncbi.nlm.nih.gov
- Birt, L., Scott, S., Cavers, D., Campbell, C., & Walter, F. (2016). Member checking: A tool to enhance trustworthiness or merely a nod to validation? Qualitative Health Research, 26(13), 1802-1811. pubmed.ncbi.nlm.nih.gov
- Mays, N., & Pope, C. (2000). Assessing quality in qualitative research. BMJ, 320(7226), 50-52. pmc.ncbi.nlm.nih.gov
- Noble, H., & Smith, J. (2015). Issues of validity and reliability in qualitative research. Evidence-Based Nursing, 18(2), 34-35. pubmed.ncbi.nlm.nih.gov
- Schlunegger, M. C., Zumstein-Shaha, M., & Palm, R. (2024). Methodologic and data-analysis triangulation in case studies: A scoping review. Western Journal of Nursing Research, 46(8), 611-622. pmc.ncbi.nlm.nih.gov
- Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. Sage Publications. find source ↗
- Key terms
- Trustworthiness
- Lincoln and Guba's overarching standard for qualitative rigor, comprising four criteria.
- Credibility
- The fit between participants' realities and the researcher's representation; parallels internal validity.
- Transferability
- Whether findings might apply elsewhere, supported by thick description for the reader to judge; parallels external validity.
- Dependability
- Consistency and traceability of the inquiry, supported by an audit trail; parallels reliability.
- Confirmability
- Grounding of findings in data rather than bias, supported by audit trail and reflexivity; parallels objectivity.
- Member checking
- Returning findings to participants to ask whether they ring true, a credibility strategy.
Reflexivity and Qualitative Ethics
- Distinguish types of reflexivity and explain reflexivity as a methodological practice.
- Analyze ethical issues distinctive to qualitative research beyond the standard consent framework.
Because the qualitative researcher is the instrument, and because fieldwork forms real relationships with people over time, two commitments run deeper here than in quantitative work: reflexivity and a form of ethics that does not end when the consent form is signed.
Key idea: Both commitments fail in the same way - by being satisfied on paper. A positionality paragraph that names demographic categories without tracing consequences, and an ethics section that recites approval numbers, are the reflexive and ethical equivalents of asserting saturation. The test in both cases is whether something specific about the study can be shown to have changed.
Reflexivity as method
Reflexivity is the researcher's disciplined self-examination of how their own assumptions, values, social position, and presence shape every phase of the study - what questions seemed worth asking, who talked to them and how candidly, what they noticed and missed, and how they interpreted it. In a paradigm that denies a neutral view from nowhere, reflexivity is not an optional confession but a methodological practice: it is how a constructivist or critical study accounts for the researcher's inevitable influence instead of pretending it away. Several forms are distinguished:
- Personal reflexivity: examining how one's own identity, experiences, and beliefs shape the research.
- Interpersonal reflexivity: examining how the relationship and dynamics between researcher and participants shaped what was said and done.
- Methodological reflexivity: examining how one's methodological choices shaped the findings.
The usual vehicle is a reflexive journal kept throughout, and a positionality statement in the write-up. The point is not self-absorption; it is to give readers the information they need to weigh the account, and to catch the ways one's standpoint might be distorting interpretation while there is still time to correct it.
Olmos-Vega and colleagues add a fourth type that doctoral students often miss. Contextual reflexivity examines how the cultural, institutional, and disciplinary setting of the research shapes what counts as a question, a finding, and a legitimate method - including the pressures of funding, publication, and a supervisor's own commitments. They also make a distinction worth carrying: reflection looks back at what happened, while reflexivity examines the researcher's constitutive role in producing it. A researcher who writes "on reflection, the interview could have gone better" has reflected; one who writes "my framing of the opening question as being about 'barriers' presupposed that the practice was difficult, and three participants pushed back on that framing before I changed it" has been reflexive.
A worked reflexive journal entry. From a study of school exclusion:
"11 Oct, after interview 6. I noticed I was relieved when M described the exclusion as unfair. I wrote 'finally' in my jottings. That is worth examining. I came into this from six years of advocacy work, and I have a story I want the data to tell. Where has this already shaped things? (a) My guide asks 'what happened when the school decided to exclude?' - agentive framing that positions the school as actor. A neutral phrasing would be 'what happened around the exclusion.' Changing for interview 7. (b) In interviews 2 and 4, both parents partly blamed their child, and I moved past those passages quickly. Rereading them now, both contain a more complicated account of responsibility than I coded for. Recoding. (c) Action: ask my critical friend to read interviews 2 and 4 specifically for material I have under-coded, without telling her why."
That entry does the work: it identifies an emotional reaction, traces it to a biographical source, finds two concrete places where it has already operated, and installs a correction, including an external check. It is also the raw material from which a strong reflexivity section is later written.
A practical caution. Reflexivity can become its own excess - pages of self-narration that displace the participants, sometimes called navel-gazing or confessional reflexivity. Linda Finlay's discussion of "negotiating the swamp" is the classic treatment of this hazard. The discipline is to ask of every reflexive sentence whether it helps a reader interpret the findings. If it does not, it belongs in the journal, not the thesis.
Ethics beyond the consent form
The formal apparatus of research ethics - informed consent, review board approval, confidentiality - applies fully to qualitative work. But several issues are heightened or distinctive:
- Consent as ongoing, not one-time. Because designs are emergent and fieldwork is prolonged, what a participant agreed to at the outset may not match where the study goes. Ethical practice treats consent as process consent, renewed as the research evolves, rather than a signature obtained once.
- Confidentiality is harder to guarantee. Rich, contextual description - the very thing that gives qualitative work its power and supports transferability - can make individuals or a small community identifiable even when names are removed. Protecting participants may require altering non-essential details, and sometimes trading a little descriptive richness for anonymity. In a small or bounded setting, internal confidentiality is a special worry: members may recognize one another in the account.
- The relationship carries its own duties. Sustained rapport can blur into friendship, creating the risk that participants disclose more than they would have chosen to, or feel used when the researcher departs. The researcher holds real power in representing others' lives, and that power must be exercised with care.
- Representation is an ethical act. How you portray participants in writing - whose voice is centered, whether people would recognize themselves without harm - is not merely a craft question but a moral one, felt most acutely in critical and community-based work.
The through-line is that qualitative ethics is relational and situational. It cannot be discharged by a form at the start; it requires ongoing ethical judgment as relationships develop and the study changes. Reflexivity and ethics meet here: honest attention to one's own position is itself part of treating participants, and their stories, with the respect they are owed.
Procedural ethics and ethics in practice
Marilys Guillemin and Lynn Gillam drew a distinction that has become standard vocabulary. Procedural ethics is what an ethics committee reviews in advance: consent documents, risk assessment, data security, recruitment materials. Ethics in practice is what happens in the room, and it turns on what they call ethically important moments - the small, unforeseeable junctures where a decision with moral weight has to be made immediately and no protocol covers it.
Concrete examples make the category real. A participant discloses, mid-interview, that she has not told her family about her diagnosis, and then asks whether you think she should. A carer describes practices you believe amount to neglect, but which fall short of a mandatory reporting threshold. A teacher, having agreed to be interviewed, spends forty minutes describing a colleague in identifiable and damaging terms. A participant in a study of homelessness asks you for a lift, or for money. Each is a moment where the researcher's response shapes the relationship, the data, and the participant's welfare, and where the consent form is silent.
What helps is not a rule but preparation. Think through the likely moments for your topic in advance and write down your intended response. Know precisely what your reporting obligations are and have disclosed them at consent. Have a boundary you can state kindly and consistently ("I'm not able to advise on that, but I can tell you who could"). Debrief with a supervisor after difficult interviews rather than carrying them alone. And record ethically important moments in your reflexive journal, because they are simultaneously ethical events and analytic ones - what a participant asks you for often reveals what they think the encounter is.
Anonymization and its limits
Karen Kaiser's analysis of confidentiality in qualitative research identifies a tension that has no clean solution: the contextual detail that gives an account its evidential force is the same detail that identifies people. Removing names is the least of it. A participant described as "the only male midwife at a rural unit in the north" is named as surely as if you had used his name, and internal confidentiality - recognition by colleagues rather than strangers - is usually the more realistic threat.
Practical options, each with costs: alter non-essential identifying details and say that you have done so; use composites, disclosing their construction, at the price of no longer being able to trace a claim to one person; report roles rather than individuals; return the specific passages in which a participant appears and negotiate what may be published, which shifts some control to them; and, in small settings, consider whether some findings can be reported only in aggregate. Kaiser argues that the right time to address this is at consent, by discussing with participants how their data will be presented rather than promising a confidentiality that cannot be delivered.
Common misconceptions
Reflexivity means declaring your identity categories. Categories are a starting point. What matters is the traceable consequence: what you asked, what you were told, what you noticed, and what you did about it.
Ethics approval means the study is ethical. Approval addresses procedural ethics. The ethically important moments that determine whether participants were treated well occur afterwards and are the researcher's responsibility.
Anonymity can be guaranteed. It cannot in small or bounded settings, and promising it is itself an ethical failure. Describe accurately what you will and will not do.
Distress means harm has occurred. Many participants find being asked about difficult experiences valuable, and treating a topic as unaskable can be paternalistic. The relevant question is whether the participant retained control.
Recap
- Reflexivity is a methodological practice with personal, interpersonal, methodological, and contextual dimensions; reflection looks back, reflexivity examines the researcher's constitutive role.
- A reflexive entry earns its place by identifying a reaction, tracing its source, locating where it has already operated, and installing a correction.
- Excessive self-narration displaces participants; keep in the thesis only what helps a reader interpret findings.
- Procedural ethics is reviewed in advance; ethics in practice turns on ethically important moments that no protocol anticipates.
- Consent in emergent, prolonged designs is a process, renewed as the study changes.
- Rich description and confidentiality are in tension; address presentation of data at consent rather than promising anonymity you cannot deliver.
Sources
- Olmos-Vega, F. M., Stalmeijer, R. E., Varpio, L., & Kahlke, R. (2023). A practical guide to reflexivity in qualitative research: AMEE Guide No. 149. Medical Teacher, 45(3), 241-251. pubmed.ncbi.nlm.nih.gov
- Rohm, M. (2024). An exploration of practical reflexivity: Navigating categories in research encounters. Forum: Qualitative Social Research, 25(3). qualitative-research.net
- Kaiser, K. (2009). Protecting respondent confidentiality in qualitative research. Qualitative Health Research, 19(11), 1632-1641. pmc.ncbi.nlm.nih.gov
- Sanjari, M., Bahramnezhad, F., Fomani, F. K., Shoghi, M., & Cheraghi, M. A. (2014). Ethical challenges of researchers in qualitative studies: The necessity to develop a specific guideline. Journal of Medical Ethics and History of Medicine, 7, 14. pmc.ncbi.nlm.nih.gov
- Holmes, A. G. D. (2020). Researcher positionality: A consideration of its influence and place in qualitative research. Shanlax International Journal of Education, 8(4), 1-10. eric.ed.gov
- Korstjens, I., & Moser, A. (2018). Series: Practical guidance to qualitative research. Part 4: Trustworthiness and publishing. European Journal of General Practice, 24(1), 120-124. pmc.ncbi.nlm.nih.gov
- Guillemin, M., & Gillam, L. (2004). Ethics, reflexivity, and "ethically important moments" in research. Qualitative Inquiry, 10(2), 261-280. find source ↗
- Key terms
- Reflexivity
- Disciplined self-examination of how the researcher's assumptions, position, and presence shape the study.
- Personal reflexivity
- Examining how one's own identity, experiences, and beliefs shape the research.
- Reflexive journal
- A journal kept throughout a study to record and examine the researcher's influence.
- Process consent
- Consent treated as ongoing and renewed as an emergent study evolves, not obtained once.
- Internal confidentiality
- The risk that members of a small or bounded setting recognize one another in the account.
- Representation (ethics of)
- The moral dimension of how participants are portrayed in the written report.
Writing Up Qualitative Findings
- Structure a qualitative findings section and integrate evidence with interpretation.
- Use participant quotations effectively and represent voice responsibly, avoiding common write-up failures.
Qualitative writing is not a transparent report of results that speak for themselves; it is an argument, built from evidence, that persuades a reader of an interpretation. The write-up is a genuine part of the analysis - meaning is often clarified in the act of composing it - and it is where much otherwise sound work disappoints, either by drowning the reader in raw quotation or by asserting conclusions the data are never shown to support.
Key idea: The reader of a quantitative paper can, in principle, check your inference by inspecting your numbers. The reader of a qualitative paper can only check your inference against the evidence you choose to show. That asymmetry places the entire burden of demonstration on the writing. A qualitative findings section is not a report of what you found; it is the apparatus by which a reader is enabled to evaluate whether you found it.
The shape of qualitative findings
Qualitative findings are usually organized thematically or conceptually rather than by interview question or by participant. Each major section presents a theme or category, developed as a claim and supported by evidence. Two conventions differ from quantitative writing. First, the boundary between "results" and "discussion" is often softer: because presenting a finding already involves interpreting it, many qualitative reports weave description and interpretation together rather than quarantining them. Second, first-person and reflexive commentary is accepted and often expected, consistent with the researcher-as-instrument stance.
Using quotations well
Participant quotations are the evidence base, and using them is a skill:
- Quote to illustrate a point you have made, not in place of making it. A quotation supports an analytic claim; it does not substitute for one. Data cannot speak for themselves - the analyst must say what a quotation shows and why it matters.
- Balance evidence and interpretation. Two opposite failures recur. An under-analyzed write-up strings quotations together with little commentary, leaving the reader to do the analysis. An over-claimed write-up asserts sweeping interpretations with too little grounding in shown data. Aim for the interplay: claim, evidence, and the analytic connection between them.
- Choose quotations well and contextualize them. Select excerpts that are vivid and representative of a pattern (and sometimes a telling negative case), and give enough context that the reader can interpret them. Indicate edits honestly with ellipses and brackets.
- Attribute without exposing. Use pseudonyms and role descriptors, being careful that the accumulation of quoted details does not de-anonymize a participant.
Lorelei Lingard names the most common structural failure the default colon: an analytic sentence, a colon, and a quotation dropped in to prove it, repeated down the page. The pattern reads as evidence but functions as decoration, because the quotation is never worked on. Her alternative is to give quotations analytic jobs. A quotation can illustrate a claim already made; it can demonstrate something the prose could not efficiently paraphrase, such as a participant's exact phrasing; it can be interrogated, with the analyst pointing to specific words and saying what they do; it can complicate the argument by showing a case that does not fit; or it can evoke, conveying texture that summary would flatten. Before including any quotation, be able to name which of these it is doing. A related test: read the section with all quotations removed. If the argument still stands, the quotations are illustrative and probably too numerous. If the argument collapses into assertion, you have been letting data do work you should have done.
Practical conventions worth settling early. Decide on a quotation-attribution format and use it consistently - typically a pseudonym plus one or two role or context descriptors, chosen so that the descriptors used across the paper cannot be accumulated to identify anyone. Decide whether to clean up disfluencies, state your decision, and apply it uniformly; unedited transcription can make articulate people appear inarticulate, which is a representational harm as well as a stylistic problem. Reserve block quotations for extracts you are going to analyze closely, and integrate short phrases into your own sentences, which usually reads better and forces you to say something. And avoid the practice of listing three or four quotations under a theme heading with no intervening text - this is the clearest available signal to a reviewer that analysis was not completed.
Voice and representation
How you handle voice is both a craft and an ethical decision. Whose words are quoted and how often, whether participants are rendered as rounded people or reduced to fragments, and whether they would recognize themselves in your portrayal - these choices shape both the persuasiveness and the integrity of the work. Some traditions foreground participants' own words extensively; others integrate them more tightly into the analyst's argument. Either way, the writer wields real power in representing others' lives and should exercise it with care.
Demonstrating rigor in the prose
Finally, the write-up is where trustworthiness becomes visible. A strong qualitative report does not merely list rigor strategies in the methods section; it lets the reader see credibility in the density and fit of the evidence, transferability in the thickness of the contextual description, and confirmability in the clear line from data to claim.
When a reader finishes and thinks, "yes, I can see how they got from what people said to what they concluded, and I understand the world they studied well enough to judge it," the writing has done its work. That traceable, evidenced, honestly interpreted account is the final deliverable of qualitative inquiry - and the ultimate test of everything that came before it.
A worked illustration: the same theme written twice
Weak version. "Theme 3: Communication. Participants discussed communication. As one nurse said: 'Nobody tells you anything until it's already decided.' Another commented: 'The email comes out and that's the first you hear.' A third stated: 'We find out from the patients sometimes.' Communication was therefore an important issue for staff."
Three quotations, no analysis, a theme name that is a topic, and a concluding sentence that adds nothing. The reader is left to construct the finding.
Strong version. "Finding 3: Learning of decisions from those they were made about. Staff described a consistent inversion of the expected information order, in which they learned of operational decisions not from managers but from the people the decisions affected. The inversion was described as routine rather than exceptional. A charge nurse put it in temporal terms: 'Nobody tells you anything until it's already decided' - locating the grievance not in secrecy but in sequence, since the information does arrive, only after the point at which it could be used. Others described a further step, in which patients became the channel: 'We find out from the patients sometimes.' This is more than an inconvenience. Being informed by a patient about a change to that patient's own care requires the clinician to perform composure while lacking the information the patient assumes they hold, and three participants described managing that moment explicitly. The pattern was not universal. Two participants on the smaller unit described being consulted before changes, and both attributed this to a manager who attended handover daily - which suggests the mechanism is proximity of decision-maker to clinical routine rather than communication policy as such, since both units operated under the same policy."
Note what the second version does. It names the finding as an idea rather than a topic. It states the pattern before the evidence. It interrogates specific language - until it's already decided - rather than letting a quotation stand alone. It escalates from inconvenience to a consequence for practice. It reports a disconfirming case, and uses it to specify a mechanism rather than to hedge. And it makes a claim that another researcher could test.
Reporting standards
Two instruments now shape expectations. COREQ (Tong, Sainsbury, and Craig) is a 32-item checklist for interview and focus group studies, covering the research team and reflexivity, study design, and analysis and reporting; many health journals require it. SRQR (O'Brien and colleagues) is a 21-item synthesis intended for qualitative research more broadly and is less prescriptive about method. In psychology, the APA's JARS-Qual standards, developed by Levitt and colleagues, set out what should be reported for qualitative, qualitative meta-analytic, and mixed-methods work. Use them as prompts while drafting rather than as a form completed at submission, and be aware of the standing criticism: a study can satisfy every item and still be analytically thin, because the checklists govern disclosure rather than quality of thought.
Common misconceptions
Findings and discussion must be strictly separated. In qualitative work presenting a finding is already interpretive, and many strong reports integrate them. What must be separated is what participants said from what you concluded, and that separation is achieved by showing evidence, not by section headings.
More quotations mean stronger evidence. Beyond a point they signal the opposite. What persuades is the fit between claim and evidence and the visible work done on the evidence.
You should quote every participant equally. Some participants are more articulate, and quoting them disproportionately is defensible if disclosed; what is not defensible is presenting one articulate participant's view as the pattern. Report how many participants contributed to a finding without converting that into a prevalence claim.
Negative cases weaken the write-up. Reporting them is one of the strongest available signals of credibility, and using them to specify conditions turns an apparent problem into a sharper finding.
Recap
- The reader can only evaluate your inference against the evidence you show, so the write-up carries the whole burden of demonstration.
- Organize by theme or concept, name findings as ideas rather than topics, and state the claim before the evidence.
- Give every quotation an analytic job - illustrate, demonstrate, interrogate, complicate, or evoke - and avoid Lingard's default colon.
- Report disconfirming cases and use them to specify conditions and mechanisms.
- Handle attribution, editing, and accumulation of detail with confidentiality in mind, and apply conventions consistently.
- COREQ, SRQR, and JARS-Qual standardize disclosure; use them while drafting, and remember they do not measure analytic quality.
Sources
- Lingard, L. (2019). Beyond the default colon: Effective use of quotes in qualitative research. Perspectives on Medical Education, 8(6), 360-364. pmc.ncbi.nlm.nih.gov
- O'Brien, B. C., Harris, I. B., Beckman, T. J., Reed, D. A., & Cook, D. A. (2014). Standards for reporting qualitative research: A synthesis of recommendations. Academic Medicine, 89(9), 1245-1251. pubmed.ncbi.nlm.nih.gov
- Tong, A., Sainsbury, P., & Craig, J. (2007). Consolidated criteria for reporting qualitative research (COREQ): A 32-item checklist for interviews and focus groups. International Journal for Quality in Health Care, 19(6), 349-357. pubmed.ncbi.nlm.nih.gov
- Levitt, H. M., Bamberg, M., Creswell, J. W., Frost, D. M., Josselson, R., & Suarez-Orozco, C. (2018). Journal article reporting standards for qualitative primary, qualitative meta-analytic, and mixed methods research in psychology. American Psychologist, 73(1), 26-46. pubmed.ncbi.nlm.nih.gov
- Sandelowski, M. (1998). Writing a good read: Strategies for re-presenting qualitative data. Research in Nursing & Health, 21(4), 375-382. pubmed.ncbi.nlm.nih.gov
- Korstjens, I., & Moser, A. (2018). Series: Practical guidance to qualitative research. Part 4: Trustworthiness and publishing. European Journal of General Practice, 24(1), 120-124. pmc.ncbi.nlm.nih.gov
- Wolcott, H. F. (2009). Writing up qualitative research (3rd ed.). Sage Publications. find source ↗
- Key terms
- Thematic organization
- Structuring a findings section by theme or concept rather than by question or participant.
- Claim-evidence-connection
- The interplay of an analytic claim, supporting data, and the stated link between them.
- Under-analysis
- A write-up that strings quotations together with too little interpretive commentary.
- Over-claiming
- A write-up that asserts sweeping interpretations with too little grounding in shown data.
- Voice
- The craft-and-ethics choices about whose words are quoted and how participants are represented.
- Pseudonym
- A false name used to attribute quotations while protecting a participant's identity.