🔬 Research & Scholarship · Doctoral & Postdoctoral · RES 740

Research Ethics & Integrity

A graduate treatment of the ethical obligations that bind working researchers, from the protection of human and animal participants to the integrity of the scholarly record itself. You will study the historical abuses that produced modern regulation, the machinery of ethical review, and the norms - authorship, disclosure, data stewardship, transparency, and mentoring - that hold a research…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 17 lessons. The full lesson text is below so you can read it right here.

Module 1: Why Research Ethics Matter

The historical abuses that produced modern research regulation, and the foundational principles that grew out of them.

The Case for Research Ethics

  • Explain why the special power and trust of research create distinctive ethical obligations.
  • Distinguish the three broad domains of research ethics: protecting participants, ensuring integrity, and serving the public.

It is tempting to treat research ethics as a set of forms to complete before the real work begins - a checkpoint to clear rather than a discipline to master. That attitude is exactly what this course is designed to dislodge. Ethics is not external to good research; it is a condition of it. A study that harms its participants, or that reports results the data do not support, has failed as research, not merely as paperwork.

Keep one person in mind for the whole lesson. Priya Raman is a second-year doctoral student in a sleep laboratory. She runs a study on memory and overnight rest. Late one Friday her supervisor looks at the first forty participants and says the effect is "almost there." He asks her to drop the six people whose scores sit furthest from the trend. He calls them outliers, and he does not tell her to invent anything or misrepresent a procedure. He asks for one small deletion, and the manuscript is due in a week. Priya has never been taught a rule that covers this exact moment. She will meet versions of it for the rest of her career, and so will you.

Key idea: Most ethical trouble in research does not arrive labeled as a crime - it arrives as a small, reasonable-sounding request from someone you respect.

In plain terms

Research ethics answers three plain questions. Who might my work hurt? Is what I am reporting true? What do I owe the people around me? That is the whole subject. The forms and the review boards exist to force those questions into the open while there is still time to act on them. Many people treat the forms as the subject and miss the point. A signed consent form is not consent. A stamped approval is not safety. Each is a record that someone asked the right question at the right time. Learn only the paperwork and you can follow every rule and still do harm, because no set of rules covers every case. Learn the questions and you can handle cases no rule anticipated. That is why this course spends more time on judgment than on procedure.

Why research is a special case

Three features of research generate obligations that ordinary conduct does not carry.

  • Power over participants. Research routinely exposes people to risk - a new drug, an invasive question, an experimental manipulation - in the pursuit of generalizable knowledge rather than the participant's own benefit. That distinguishes research from clinical care and raises the moral stakes of consent and protection.
  • A privileged claim on public trust. Science speaks with authority. Physicians prescribe on the strength of trials; governments legislate on the strength of studies; the public funds the whole enterprise. That authority is a loan against integrity. When a fabricated result enters the literature, it is not one lab that suffers but every downstream decision built on it.
  • Self-regulation. Much of research is checked not by outside inspectors but by the researchers themselves, through peer review, replication, and honest reporting. A profession that polices itself must internalize the standards, because no external system can catch every lapse.

Compare Priya's position with a physician's. A doctor who prescribes a drug is acting for the patient in front of her, and the patient can eventually judge whether the intervention helped. Priya's forty volunteers derive no therapeutic benefit whatever; they surrendered a night of sleep so that strangers might eventually learn something about memory consolidation. That structural gap is the reason consent carries such extraordinary weight in research: the individual accepting the burden is not the individual collecting the benefit. Clinicians have a name for the confusion of those two roles. They call it the therapeutic misconception, and a participant who believes an experiment is treatment has not really understood what they agreed to.

The trust point is just as concrete. Suppose Priya deletes the six participants; the published effect will then be substantially larger than the actual one. Some other laboratory will attempt to build on that inflated estimate, fail, and probably attribute the failure to its own incompetence rather than to the original paper. A reviewer three years from now will cite the same number in a funding decision. None of those people can open her raw files, so they are relying entirely on her having told the truth, and no practical mechanism exists that would let them verify it independently.

Key idea: Research borrows against a trust it cannot repay with evidence, because almost nobody downstream can check your raw data - which is why honesty has to be a habit rather than a policy.

Three domains of obligation

This course is organized around three families of ethical duty that together define responsible conduct of research (often abbreviated RCR).

  1. Protecting participants. Duties owed to the humans and animals who make research possible: consent, risk minimization, privacy, and fair treatment. Modules 2 and 3 address these.
  2. Integrity of the record. Duties owed to the truth and to fellow scientists: honest data, fair credit, disclosed conflicts, transparent methods. Modules 4 and 5 address these.
  3. Responsibility to the community. Duties owed to trainees and to the health of the field itself: good mentoring, fair peer review, and a culture in which raising a concern is safe. Module 6 addresses these.

The three domains are not ranked, and clearing one says nothing about the others. A study can pass every human-subjects review and still be dishonest in its reporting. A perfectly honest paper can rest on participants who were never really told what they were joining. Priya's problem sits squarely in the second domain. That is worth noticing, because the ethics training many graduate students receive covers only the first.

Key idea: Passing ethical review protects your participants and says nothing at all about whether your results are honest - those are separate duties with separate machinery.

Three systems, not one, hold you to this

Newcomers often picture a single ethics authority sitting above the field. There is no such thing. Three loosely connected systems act on a working researcher, and each responds to a different kind of failure.

  • Human-subjects regulation. In the United States this is the Common Rule, codified at 45 CFR 46 and enforced locally by an Institutional Review Board. It governs how you treat participants. It has almost nothing to say about whether you later analyze the data honestly.
  • Research-misconduct policy. For Public Health Service funding this runs through the Office of Research Integrity under 42 CFR 93, with parallel offices at other agencies. It handles fabrication, falsification, and plagiarism after the fact, and only for a narrow band of serious cases.
  • Professional and community norms. Journals, funders, learned societies, and your own colleagues police the wide middle ground: authorship, disclosure, data sharing, reviewing, mentoring. Most of what you actually owe lives here, and almost none of it carries a legal penalty.

Priya's Friday afternoon falls in the third band, and arguably in the second. No review board will ever hear about it. If she deletes the six participants and reports the study as though all forty were analyzed, she has misrepresented the work. Whether any office would ever formally call that misconduct is a separate question from whether it is true.

Key idea: The formal enforcement systems catch a small fraction of ethical failure - the rest is held up by norms, which means by people like you deciding what counts as normal.

Ethics is judgment, not a checklist

A recurring theme is that principles frequently conflict. A study that would benefit many might most easily recruit a desperate few; a commitment to openness might collide with a duty of confidentiality; the drive to publish might tempt a researcher to overstate a finding. Ethical competence is not the memorization of rules but the capacity to recognize these tensions and reason through them defensibly. Regulations set a floor. The rest is professional judgment - which is precisely why it must be trained, discussed, and practiced rather than assumed.

One caution as you begin: it is comforting to imagine that unethical research is done only by villains. The historical record and the modern retraction record both say otherwise. Most misconduct is committed by intelligent, well-intentioned people under pressure, who talked themselves into it one small step at a time. The purpose of studying ethics is to install the habits and the vocabulary that let you notice the first small step - in your own work as much as in others'.

The slow slide, and what stops it

Survey evidence supports that picture. Pooling anonymous self-report studies across fields, Daniele Fanelli estimated that about 2 percent of researchers admitted having fabricated or falsified data at least once. Around a third admitted other questionable practices, such as dropping data points on a hunch. The second number is the alarming one. Deliberate fabrication is comparatively rare, whereas undisclosed analytic distortion appears to be ordinary behavior among ordinary researchers meeting ordinary deadlines.

What separates Priya from a fraud case is not character but vocabulary. If the only description available to her is "my supervisor asked me to clean the data," she will probably comply without further reflection. If she can instead formulate the thought "he is asking me to remove genuine observations after seeing which direction they push the result, and that is an analytic choice we are obliged to disclose," then she has something articulable to say out loud. The vocabulary is itself the intervention, which is why this course insists on precise terminology rather than general appeals to integrity.

A practical move is available to her. She can propose reporting both analyses - the full forty participants and the trimmed subset - with the exclusion rule stated explicitly in the methods section. That is not a compromise between honesty and pressure; it is simply the honest version, and it typically survives peer review considerably better than the trimmed result presented alone.

Key idea: Nearly every questionable request has an honest version - do the analysis and report the choice - and learning to name that version out loud is the core skill this course trains.

Where people get stuck

Three confusions cause most of the trouble in this opening material.

  • Treating ethics as compliance. Compliance asks whether a rule was broken. Ethics asks whether the conduct was defensible. Priya can satisfy every rule on her campus and still publish a number she knows is inflated. Compliance is the floor of the building, not the building.
  • Assuming misconduct requires a bad person. The useful question is never "would I do something evil?" It is "what would I do at six on a Friday evening, tired, behind schedule, when someone senior asks?" Build your habits for that version of yourself.
  • Confusing research with practice. A hospital that tracks its own infection rates to fix its own wards is doing quality improvement. A hospital that tracks them to publish a general claim about infection control is doing research. The second triggers duties the first does not. The dividing line is the aim of generalizable knowledge, not the activity itself.

Key idea: Ask "is this defensible to a stranger who watched everything I did" rather than "is this against the rules", and most hard cases start to sort themselves out.

Recap

  • Ethics is a condition of good research rather than a checkpoint before it, because a harmful or dishonest study has failed as research.
  • Research is a special case for three reasons: it imposes risk for other people's benefit, it borrows heavily against public trust, and it is largely self-policed.
  • The duties fall into three domains - protecting participants, protecting the integrity of the record, and sustaining the community - and clearing one says nothing about the others.
  • Three separate systems apply pressure: human-subjects regulation, misconduct policy, and professional norms; the last covers the most ground and carries the least force.
  • Principles conflict routinely, so ethical competence is the trained ability to name the tension and reason through it rather than the memorization of rules.
  • Self-report surveys suggest outright fabrication is rare while smaller distortions are common, which is where the real risk to the literature sits.
  • Priya's outlier request has an honest version - report both analyses and state the exclusion rule - and finding that version is the habit this course builds.

Sources

  1. National Academies of Sciences, Engineering, and Medicine. (2017). Foundations of integrity in research: Core values and guiding norms. In Fostering integrity in research. National Academies Press. ncbi.nlm.nih.gov
  2. Fanelli, D. (2009). How many scientists fabricate and falsify research? A systematic review and meta-analysis of survey data. PLoS ONE, 4(5), e5738. pmc.ncbi.nlm.nih.gov
  3. Office of Research Integrity. (n.d.). Definition of research misconduct. U.S. Department of Health and Human Services. ori.hhs.gov
  4. U.S. Department of Health and Human Services. (2018). Federal policy for the protection of human subjects, 45 C.F.R. Part 46. Electronic Code of Federal Regulations. ecfr.gov
  5. Barrow, J. M., Brannan, G. D., & Khandhar, P. B. (2024). Research ethics. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
  6. ALLEA. (2023). The European code of conduct for research integrity (rev. ed.). All European Academies. allea.org
  7. World Conferences on Research Integrity Foundation. (2010). Singapore statement on research integrity. singaporestatement.org
Key terms
Responsible conduct of research (RCR)
The body of professional norms governing honest, ethical, and accountable research practice.
Generalizable knowledge
Knowledge intended to extend beyond the specific participants studied, the defining aim of research.
Self-regulation
The system by which researchers police their own standards through peer review, replication, and honest reporting.
Human subject
A living individual about whom a researcher obtains data through intervention or interaction, or identifiable private information.
Research integrity
Adherence to honesty and rigor in proposing, performing, and reporting research.

Historical Abuses and the Birth of Regulation

  • Trace the specific historical cases that produced modern research-ethics regulation.
  • Connect each landmark document to the abuse it was written to prevent.

Every major protection in research ethics was written in blood. The regulations can feel abstract until you learn the cases that produced them; then each requirement reads as a direct answer to a specific harm. This lesson walks through that history so the rules will never again seem arbitrary.

One person will anchor the whole account. Peter Buxtun was a venereal-disease investigator hired by the U.S. Public Health Service in San Francisco in 1965. He was in his twenties, junior, and new. In 1966 he learned from colleagues about a study his own employer was running in Alabama, and what he heard did not sound like medicine. He wrote to the division that ran it. A review panel considered his objection and decided the study should continue. He wrote again in 1968, and the answer was the same. In 1972 he gave his documents to a reporter, Jean Heller of the Associated Press, and the story ran nationally on 26 July of that year. The study stopped within months. Congress passed a new law within two years. Buxtun had no authority, no seniority, and no committee behind him. What he had was the willingness to keep saying that something was wrong.

Key idea: Every protection in this course exists because a specific harm happened to specific people, and in several cases because one junior person refused to accept the official answer.

In plain terms

The story of research regulation runs in a loop. Someone does research that hurts people. The harm becomes public. Rules are written to stop that exact harm from happening again. Then a new kind of harm appears that the rules did not anticipate, and the loop starts over. That is why the documents in this lesson stack rather than replace one another. Nuremberg answered coerced experiments. Helsinki answered ordinary clinical trials that Nuremberg had not addressed. Belmont answered a long American scandal. Each was written after the damage, not before it. You should read this history for two reasons. It tells you what each rule is actually for, which makes the rule easier to apply in a case it does not name. It also tells you what the current rules probably still miss.

Nazi experimentation and the Nuremberg Code (1947)

During the Second World War, physicians in Nazi Germany conducted lethal experiments on concentration-camp prisoners without consent - immersion in freezing water, deliberate infection, high-altitude exposure to the point of death. At the subsequent Nuremberg Medical Trial, judges articulated the Nuremberg Code, ten principles for permissible human experimentation. Its opening line is the foundation of everything that follows: "The voluntary consent of the human subject is absolutely essential." The Code also required that experiments yield fruitful results unprocurable by other means, avoid unnecessary suffering, and allow the subject to withdraw.

The setting matters for understanding what the Code is. It was not a treaty, a statute, or a professional guideline. It appeared inside a criminal judgment. The trial, formally United States of America v. Karl Brandt et al., ran from December 1946 to August 1947 before a U.S. military tribunal at Nuremberg. Twenty-three defendants, most of them physicians, faced charges arising from experiments and killings carried out in the camps. The judges needed a standard against which to measure the defense argument that all wartime research is brutal, so they set out ten conditions under which human experimentation could be legitimate. That is where the Code comes from.

Its origin explains both its force and its limits. Because the ten points were written to condemn atrocity, they are absolute in tone, carry no enforcement machinery, and say nothing about the routine questions a hospital investigator faces. Many physicians read the Code as a rule for barbarians rather than a standard for respectable medicine, which is precisely the misreading the next thirty years would correct.

Key idea: The Nuremberg Code is a criminal court's list of conditions for legitimate experimentation, which is why it is absolute about consent and silent about the ordinary judgment calls of clinical research.

The Declaration of Helsinki (1964)

The Nuremberg Code was aimed at atrocity and said little about ordinary clinical research. In 1964 the World Medical Association adopted the Declaration of Helsinki, a set of ethical principles for medical research that has been revised many times since. It introduced ideas the Code lacked: the distinction between therapeutic and non-therapeutic research, the requirement for independent review of protocols, and the enduring principle that the well-being of the individual research subject must take precedence over the interests of science and society.

Helsinki is a living document, and its revision history is a compressed history of the field's arguments. The 1975 Tokyo revision introduced the requirement that protocols go to an independent committee, which is the ancestor of the modern review board. Later revisions took positions on placebo controls, on what sponsors owe communities after a trial ends, and on the duty to register trials publicly and report results whether or not they are favorable. The 2024 revision, adopted sixty years after the original, replaced the phrase "human subjects" with "human participants" throughout, a change of vocabulary meant to signal a change of relationship.

Key idea: Helsinki carried research ethics out of the courtroom and into everyday clinical practice, and its repeated revisions show that the standards are argued over rather than settled.

Tuskegee and the U.S. reckoning

The United States did not need to look abroad for abuse. The Tuskegee syphilis study, run by the U.S. Public Health Service from 1932 to 1972, followed the untreated progression of syphilis in about 600 Black men in Alabama - most of them poor sharecroppers - who were never told they had the disease and were actively prevented from getting treatment. Even after penicillin became the standard cure in the 1940s, it was withheld so investigators could continue to observe the "natural history" of the illness. When a journalist exposed the study in 1972, the public revulsion forced reform.

Several details of the case are worth holding precisely, because they are the details the regulations answer. The men were told they were being treated for "bad blood," a local term covering several conditions. Painful diagnostic spinal taps were presented to them as treatment. Investigators arranged with local draft boards to keep enrolled men from being treated during military screening. The study was not secret from the profession; findings appeared in medical journals over four decades, and no reader stopped it. That last point is the one most students miss. Tuskegee was not hidden. It was published, and the field did not object.

This is where Buxtun's role becomes instructive. His letters went through proper channels, and proper channels reviewed them and said continue; the study ended only when the facts left the institution entirely. It was terminated after the Associated Press story, a class action was settled in 1974, and a U.S. president formally apologized to the surviving men in 1997. Researchers have since documented lasting effects on willingness to join medical research among Black Americans, which is the practical shape a broken promise takes fifty years later.

Key idea: Tuskegee was reviewed internally and allowed to continue, which is why the regulatory answer was independent, external review rather than better internal review.

Beecher's list: abuse was not exotic

Two years before Buxtun wrote his first letter, an anesthesiologist at Harvard published an article that made the same point from inside the profession. In 1966 Henry Beecher described twenty-two studies drawn from mainstream medical journals in which patients had been exposed to risk without meaningful consent. He deliberately withheld the citations so that readers would attend to the pattern rather than to the individuals. His point was structural. These were not fringe operators. They were well-regarded investigators publishing in respected journals, and the profession had raised no objection.

Two of the cases Beecher discussed are now standard teaching examples. At Willowbrook State School on Staten Island, children with intellectual disabilities were deliberately infected with hepatitis virus in a study of the disease's course and of gamma globulin as protection. The investigators argued that infection was near-universal in the institution anyway and that parents had given permission. Critics answered that parental permission obtained when institutional places were scarce is not free, and that the conditions used to justify the study were conditions the state was obliged to fix. At the Jewish Chronic Disease Hospital in Brooklyn in 1963, live cancer cells were injected into chronically ill elderly patients without their knowledge, in a study of immune response.

Key idea: Beecher's contribution was to show that ethically indefensible research was ordinary, mainstream, and publishable - not the work of a criminal fringe.

The National Research Act and the Belmont Report

Congress responded with the National Research Act of 1974, which created a national commission and mandated review boards for federally funded research. The commission's landmark product, the Belmont Report (1979), distilled research ethics into three principles - respect for persons, beneficence, and justice - that remain the ethical backbone of U.S. regulation and the subject of the next lesson.

Trace the timeline and the loop is obvious. The Associated Press story ran in July 1972. The National Research Act was signed in July 1974, two years later, and it created the National Commission that produced Belmont in 1979. Federal regulations followed in 1981. In 1991 a common core of those rules was adopted across fifteen federal departments and agencies, which is why it is called the Common Rule. A substantial revision published in 2017 took general effect in January 2019, adding a required key-information summary at the front of consent forms and reducing continuing review for lower-risk studies. Roughly half a century separates Buxtun's first letter from the consent form on your desk, and the line between them is direct.

Key idea: Modern U.S. human-subjects law is a two-year legislative reaction to one newspaper story, elaborated over the following fifty years.

Case / abuseDocument producedCentral lesson
Nazi medical experimentsNuremberg Code (1947)Voluntary consent is essential
Need for clinical-research normsDeclaration of Helsinki (1964)Individual welfare over science; independent review
Tuskegee syphilis studyNational Research Act (1974); Belmont Report (1979)Consent, honesty, and justice for the vulnerable

Other cases reinforced the pattern. In the Willowbrook hepatitis studies, children with intellectual disabilities at a state institution were deliberately infected. At the Jewish Chronic Disease Hospital, live cancer cells were injected into elderly patients without their knowledge. Each scandal added weight to the same conclusion: researchers cannot be trusted to be the sole judges of what they may do to participants, and independent, principled review is not optional. Hold that history in mind. When a review board asks you to justify a risk or rewrite a consent form, it is enforcing a promise made to the people these cases harmed.

Where people get stuck

Four misreadings of this history recur, and each one leads to a practical error later.

  • Treating the cases as ancient history. The Tuskegee study ended in 1972, within the working lifetime of many current senior faculty. In 2010 a historian documented that U.S. Public Health Service researchers had deliberately exposed Guatemalan prisoners, soldiers, and psychiatric patients to sexually transmitted infections between 1946 and 1948, a study unknown to the public until her archival work surfaced it. The archive is not closed.
  • Assuming the perpetrators knew they were villains. The Willowbrook investigators published detailed ethical defenses of their work and believed them. Sincerity is not a defense, and it is not a reliable warning sign either.
  • Confusing the documents. Nuremberg came from a criminal tribunal, Helsinki from a physicians' association, and Belmont from a U.S. national commission. They differ in authority, scope, and enforceability. Only the last has direct force in American law, and only through the regulations built on it.
  • Reading review as bureaucratic drag. Internal review existed at Tuskegee and repeatedly approved continuation. That is exactly why the reform was external and independent review, with membership requirements designed to break professional consensus.

Key idea: The recurring failure in every case was not ignorance of ethics but the absence of anyone outside the project with the standing to stop it.

Recap

  • Regulation follows harm: each landmark document was written after a specific abuse became public, which is why the documents stack rather than replace one another.
  • The Nuremberg Code came out of a 1946 to 1947 criminal judgment and opens with the absolute necessity of voluntary consent.
  • The Declaration of Helsinki extended ethics to routine clinical research, added independent protocol review in 1975, and has been revised repeatedly since, most recently in 2024.
  • The Tuskegee syphilis study ran from 1932 to 1972, deceived participants about their diagnosis, and withheld penicillin after it became standard care.
  • Peter Buxtun raised concerns internally in 1966 and 1968 without effect; the study ended only after press coverage in July 1972.
  • Henry Beecher's 1966 article showed that consent failures were common in mainstream journals, with Willowbrook and the Jewish Chronic Disease Hospital as standard examples.
  • The National Research Act of 1974 produced the Belmont Report in 1979, federal regulations in 1981, the multi-agency Common Rule in 1991, and a revision effective in 2019.

Sources

  1. United States Holocaust Memorial Museum. (n.d.). The Doctors Trial: The medical case of the subsequent Nuremberg proceedings. Holocaust Encyclopedia. ushmm.org
  2. World Medical Association. (2024). WMA Declaration of Helsinki: Ethical principles for medical research involving human participants. wma.net
  3. Beecher, H. K. (1966). Ethics and clinical research. New England Journal of Medicine, 274(24), 1354-1360. pubmed.ncbi.nlm.nih.gov
  4. Krugman, S. (1986). The Willowbrook hepatitis studies revisited: Ethical aspects. Reviews of Infectious Diseases, 8(1), 157-162. pubmed.ncbi.nlm.nih.gov
  5. Mays, V. M. (2012). The legacy of the U.S. Public Health Service study of untreated syphilis in African American men at Tuskegee. Ethics & Behavior, 22(6), 411-418. pmc.ncbi.nlm.nih.gov
  6. Katz, R. V., Green, B. L., Kressin, N. R., Kegeles, S. S., Wang, M. Q., James, S. A., Russell, S. L., Claudio, C., & McCallum, J. M. (2008). The legacy of the Tuskegee syphilis study: Assessing its impact on willingness to participate in biomedical studies. Journal of Health Care for the Poor and Underserved, 19(4), 1168-1180. pmc.ncbi.nlm.nih.gov
  7. National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (2014). The Belmont Report: Ethical principles and guidelines for the protection of human subjects of research. Journal of the American College of Dentists, 81(3), 4-13. (Original work published 1979) pubmed.ncbi.nlm.nih.gov
Key terms
Nuremberg Code
The 1947 set of ten principles for human experimentation, opening with the necessity of voluntary consent.
Declaration of Helsinki
The World Medical Association's ethical principles for medical research, first adopted in 1964 and revised repeatedly.
Tuskegee syphilis study
The 1932 to 1972 U.S. study that deceived Black men and withheld syphilis treatment, triggering U.S. reform.
National Research Act (1974)
The U.S. law that mandated review boards and created the commission behind the Belmont Report.
Belmont Report
The 1979 report establishing respect for persons, beneficence, and justice as core U.S. research principles.
World Medical Association
The international body of physicians that authored and maintains the Declaration of Helsinki.

The Belmont Principles

  • State the three Belmont principles and the commitment each expresses.
  • Map each principle to the concrete regulatory practice it grounds, and recognize how the principles conflict.

The Belmont Report reduced a tangle of rules to three principles that a researcher can actually hold in mind. Nearly every requirement you will meet in human-subjects work is an application of one of them. Learn the three, learn what each grounds in practice, and you will be able to reconstruct most of the regulation from first principles.

One protocol will carry the lesson. Nadia Halim is an assistant professor of public health. She wants to test a smartphone application that sends tailored support messages to people recovering from opioid use disorder, measuring whether it reduces relapse over six months. She has an obvious recruitment source. The county drug court sends roughly two hundred people a year into supervised treatment, they all have phones, and the court coordinator is willing to hand out flyers. The study is well designed, the need is real, and the population is exactly the one the app is meant to serve. Every one of the three Belmont principles has something uncomfortable to say about it, and none of them says simply no.

Key idea: Belmont does not sort studies into permitted and forbidden - it names three things you owe, and a strong protocol shows how it satisfies all three at once.

In plain terms

Three questions, in three plain forms. Did this person really choose? Is the harm we risk worth the good we expect? And are the people taking the risk the same people who stand to gain? That is Belmont. The first question is about the individual in front of you. The second is about the study as a whole. The third is about groups, and it is the one most researchers skip. It is easy to see a person and ask whether they agreed. It is harder to step back and ask why your sample looks the way it does. Nadia's app might work. It might also end up tested on court-supervised people because they were easy to reach, then sold to insured patients who never took the risk. Nothing in the consent form would catch that. Only the third question does.

Respect for persons

This principle has two parts. First, individuals are autonomous agents entitled to make their own choices about whether to participate. Second, persons with diminished autonomy - children, prisoners, those with cognitive impairment - are entitled to additional protection. In practice, respect for persons grounds informed consent: the requirement that participation be voluntary, informed, and revocable. It also grounds the special safeguards for vulnerable groups covered later in this course.

Notice that the principle has a protective half and a respectful half, and they pull in opposite directions. The respectful half says do not decide for people who can decide for themselves. The protective half says step in when they cannot. Get the balance wrong in one direction and you are paternalistic; get it wrong in the other and you have abandoned someone who needed help. The judgment is about this person, in this study, at this moment, not about a category.

Apply it to Nadia's protocol and two problems surface immediately. First, capacity varies over time. Someone recruited during acute withdrawal may not be able to weigh a six-month commitment, though the same person could three weeks later. Second, and more serious, the recruitment channel is the court. A participant may reasonably believe that joining will look good to the judge, or that declining will look bad, whether or not anyone ever says so. That belief alone compromises voluntariness. The standard fixes are structural rather than verbal. Recruit through the treatment clinic rather than the courtroom. Have someone with no role in supervision obtain the consent. State in writing that participation has no effect on legal status. And delay enrollment until the person is medically stable.

Key idea: Voluntariness is destroyed by the setting more often than by anything anyone says, so the remedy is usually to change who asks and where, not to add a sentence to the form.

Beneficence

Beneficence obliges the researcher to maximize possible benefits and minimize possible harms. It is often paired with the older maxim of non-maleficence (do no harm), but Belmont frames it positively: not merely avoiding harm but actively securing well-being. In practice this principle grounds the risk-benefit analysis at the heart of ethical review. A protocol is acceptable only when its risks are reasonable in relation to its anticipated benefits, and when risks have been minimized as far as sound design allows.

Two working rules follow from the way Belmont frames it. The first is that risk must be minimized before it is weighed. A board should not accept a risky procedure as given and then ask whether the payoff justifies it. It should first ask whether the same question could be answered with less exposure. The second is that the anticipated benefit that counts is mostly benefit to society, not to the individual participant. Research is not treatment. A study may be entirely ethical while offering its participants nothing at all, provided the risks are low and honestly described.

Nadia's risks are not physical, which is exactly why they are easy to underrate. Her app will hold a record that a named person is in treatment for opioid use disorder. In a small county, a breach could affect employment, custody, housing, or a pending case. That is a serious harm, and it is a design problem rather than a consent problem. Beneficence obliges her to hold identifiers on encrypted servers, keep the linking key separate, strip the app of any visible label, and shorten the retention period. There is also a scientific risk. If the control group receives nothing, participants who wanted the app may drop out of treatment altogether. An active comparator, such as standard text reminders, both reduces that harm and improves the study.

Key idea: Minimize the risk first and weigh what remains second, and remember that in most research the benefit being weighed belongs to future patients rather than to the person in the chair.

Justice

Justice concerns the fair distribution of the benefits and burdens of research. It asks: who bears the risks, and who reaps the rewards? It is unjust for one group - the poor, the institutionalized, the marginalized - to shoulder the dangers of research while the benefits flow to others. In practice, justice grounds equitable selection of participants: people should not be recruited simply because they are convenient, powerless, or easy to pressure, and populations that bear research risk should stand to benefit from the results.

Justice was the principle Tuskegee most obviously violated, and the commission wrote it into Belmont with that history explicitly in view. The men in Alabama were selected because they were poor, rural, and unlikely to obtain care elsewhere, and the knowledge produced was never directed back to them. That is the classic pattern: burden concentrated in a group chosen for its powerlessness, benefit distributed to everyone else.

The principle has a second edge that people often miss. Justice is violated by systematic exclusion as well as by exploitation. If a drug is tested only in men, women inherit prescriptions based on evidence never gathered in them. If trials recruit only English speakers, the resulting guidance fits a narrower population than it claims to. Fair distribution runs in both directions.

For Nadia, the test is not whether drug-court participants may be studied. They plainly may, and the intervention is designed for them. The test is whether they are being chosen for a reason connected to the science or merely for convenience. If she can also recruit from community clinics and methadone programs, the court becomes one source rather than the source. Justice also raises the question of what happens afterward. If the app works, will it be available to the people who took the risk of testing it, or only to those who can pay?

Key idea: The justice question is "why does my sample look like this", and the wrong answer is "because they were the easiest people to reach."

PrincipleCore commitmentGrounds in practice
Respect for personsAutonomy; protection of the vulnerableInformed consent
BeneficenceMaximize benefit, minimize harmRisk-benefit assessment
JusticeFair distribution of benefits and burdensEquitable participant selection

From principle to regulation

Belmont is not law. It is a philosophical document that federal regulation then translated into checkable requirements. Seeing the translation makes the regulation far easier to remember, because each rule is a principle wearing working clothes.

  • Respect for persons becomes consent law. The general requirements for informed consent at 45 CFR 46.116 spell out what must be disclosed, in what language, and under what conditions consent may be waived. The subparts adding protections for children, prisoners, and pregnant women are the second half of the principle written out.
  • Beneficence becomes the approval criteria. Section 46.111 requires that risks be minimized, that risks be reasonable in relation to anticipated benefits, and that the study include adequate provisions for monitoring data and protecting privacy. That is beneficence turned into a list a committee can vote on.
  • Justice becomes equitable selection. The same section requires that selection of subjects be equitable, and directs boards to consider the purposes of the research and the setting in which it will take place - which is precisely the question Nadia has to answer about the drug court.

Key idea: Almost every clause of 45 CFR 46.111 is one of the three principles restated as something a committee can check, which is why learning the principles lets you predict the rules.

When principles collide

The principles are not a checklist to be satisfied independently; they routinely pull against one another, and ethical review is the disciplined negotiation of the tension.

  • A study promising great social benefit (beneficence) might most easily recruit a captive population such as prisoners (violating justice).
  • Full disclosure (respect for persons) can ruin the science of a study whose validity depends on participants not knowing its true purpose, forcing a careful, reviewed use of deception.
  • Protecting an individual from all risk (a strict reading of beneficence) can be unjust if it excludes a whole group - say, pregnant women - from research that would ultimately benefit them.

There is no formula that resolves these conflicts automatically. The Belmont framework does something more modest and more useful: it names the competing goods precisely, so that a researcher and a review board can argue about the right balance in the open, on the record, rather than letting one value silently override the others. When you write a protocol, address all three principles explicitly. A reviewer trained on Belmont will be looking for exactly that.

Nadia's protocol is the collision in miniature. Beneficence favors reaching the people most likely to relapse. Justice warns against loading risk onto a court-supervised group. Respect for persons questions whether anyone under court supervision can freely decline. A weak protocol picks one principle and ignores the rest. A strong one shows its work. Recruitment comes from three sources rather than one. Consent is taken by staff outside the supervision chain. A written assurance states that participation does not affect legal status. Enrollment waits until participants are medically stable. The control arm gets an active comparator. And the team commits to keeping the app available locally if it proves effective. Nothing there is a loophole. It is what taking all three principles seriously actually looks like on paper.

Key idea: A protocol that answers only one principle well is usually failing another, and a reviewer's job is to notice which one has gone quiet.

Where people get stuck

Four confusions account for most misapplications of the framework.

  • Reading Belmont as a checklist. The three principles are not boxes to tick in sequence. They are three lenses over the same protocol, and a change that improves one often worsens another. The document's value is that it forces the trade-off into the open.
  • Equating beneficence with benefit to the participant. Most research offers participants nothing directly, and that is perfectly ethical. What beneficence requires is that risks be minimized and be reasonable against the knowledge expected, not that every volunteer gain something.
  • Treating justice as a synonym for diversity targets. Enrolment numbers are evidence, not the principle. Justice asks who bears the burden and who gets the benefit. A study can hit every demographic target and still concentrate risk on people chosen for their powerlessness.
  • Assuming protection always means exclusion. Barring a group from research protects its individuals today and harms the group tomorrow, because clinicians end up guessing. The Belmont answer is inclusion with stronger safeguards, not avoidance.

Key idea: When two principles conflict, the ethical move is to redesign until both are better served - and to say in writing what you could not fully resolve.

Recap

  • Belmont reduces human-subjects ethics to three principles: respect for persons, beneficence, and justice.
  • Respect for persons has two halves - honor autonomy where it exists, protect where it is diminished - and it grounds informed consent.
  • Beneficence requires that risks be minimized first and then weighed against anticipated benefits, which are usually benefits to society rather than to the participant.
  • Justice asks who bears the burdens and who receives the benefits, and it is violated by exclusion as well as by exploitation.
  • Federal regulation translates the principles into checkable requirements, chiefly the consent rules at 45 CFR 46.116 and the approval criteria at 45 CFR 46.111.
  • The principles routinely conflict; Belmont names the competing goods rather than ranking them or supplying a formula.
  • Nadia's drug-court recruitment illustrates all three at once, and the fixes are structural: change who recruits, who consents, when, and what happens after the trial.

Sources

  1. National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (2014). The Belmont Report: Ethical principles and guidelines for the protection of human subjects of research. Journal of the American College of Dentists, 81(3), 4-13. (Original work published 1979) pubmed.ncbi.nlm.nih.gov
  2. U.S. Department of Health and Human Services. (2018). Criteria for IRB approval of research, 45 C.F.R. 46.111. Electronic Code of Federal Regulations. ecfr.gov
  3. U.S. Department of Health and Human Services. (2018). General requirements for informed consent, 45 C.F.R. 46.116. Electronic Code of Federal Regulations. ecfr.gov
  4. Friesen, P., Kearns, L., Redman, B., & Caplan, A. L. (2017). Rethinking the Belmont Report? American Journal of Bioethics, 17(7), 15-21. pubmed.ncbi.nlm.nih.gov
  5. World Medical Association. (2024). WMA Declaration of Helsinki: Ethical principles for medical research involving human participants. wma.net
  6. Mays, V. M. (2012). The legacy of the U.S. Public Health Service study of untreated syphilis in African American men at Tuskegee. Ethics & Behavior, 22(6), 411-418. pmc.ncbi.nlm.nih.gov
  7. Barrow, J. M., Brannan, G. D., & Khandhar, P. B. (2024). Research ethics. In StatPearls. StatPearls Publishing. ncbi.nlm.nih.gov
Key terms
Respect for persons
The Belmont principle honoring individual autonomy and protecting those with diminished autonomy.
Autonomy
The capacity and right of an individual to make informed, uncoerced decisions about their own participation.
Beneficence
The Belmont principle of maximizing benefits and minimizing harms, grounding risk-benefit analysis.
Non-maleficence
The duty to avoid causing harm, closely related to but narrower than beneficence.
Justice (research ethics)
The Belmont principle that the benefits and burdens of research be distributed fairly across groups.
Risk-benefit analysis
The weighing of a study's foreseeable risks against its anticipated benefits, central to ethical review.

Module 2: Protecting Human Subjects

The operational machinery of human-subjects protection: review boards, consent, privacy, and safeguards for the vulnerable.

The IRB and Human-Subjects Review

  • Describe the role and composition of an Institutional Review Board and the regulatory framework it enforces.
  • Distinguish exempt, expedited, and full-board review and identify what determines the appropriate tier.

In the United States, the Belmont principles are enforced through a system of local committees called Institutional Review Boards (IRBs), operating under a federal regulation known as the Common Rule (codified at 45 CFR 46). Most other countries run an equivalent system under names such as Research Ethics Committee or Ethics Review Board. Wherever you work, some independent body must approve research with human participants before it begins. This lesson explains what that body does and how it calibrates its scrutiny to risk.

Follow one submission through the system. Marcus Webb is a doctoral student in education studying why teachers leave the profession. His plan has three parts. He will survey about 300 middle-school teachers anonymously about workload and burnout. He will interview 40 seventh-graders about classroom climate. And he will link the school-level results to the district's existing discipline and attendance records. He believes the whole thing is exempt, because nobody is being given a drug and nothing hurts. He is wrong about the tier, wrong about which parts even count as human-subjects research, and right that no part of it is dangerous. His confusion is the ordinary confusion, and working through it teaches the system better than a summary of it would.

Key idea: The review tier is set by regulatory category and risk, not by how harmless the study feels to the person who designed it.

In plain terms

A review board answers four questions in order. Is this activity research? Does it involve human subjects? If yes to both, how much scrutiny does the risk deserve? And do the specific protections match the specific risks? Each question can end the process. If the answer to the first is no, the board has no jurisdiction and says so. If the answer to the third is "very little," a single trained reviewer handles it in a week. Most of the frustration students report comes from skipping straight to the paperwork without answering the first two questions, then arguing about a tier that was never in dispute. Answer them in order and the process is usually short. One rule overrides all of this: you do not decide any of it yourself.

What an IRB is

An IRB is a standing committee charged with reviewing research protocols to protect the rights and welfare of human participants. Federal rules require a minimum composition designed to prevent groupthink and conflict of interest: at least five members of varying backgrounds, at least one scientist, at least one non-scientist, and at least one member unaffiliated with the institution (a community member). This mix ensures that a protocol is judged not only on its science but on its acceptability to the wider community whose members might enroll.

The membership rules at 45 CFR 46.107 are more specific than most researchers realize, and each clause has a purpose. The board may not consist entirely of one profession, because a room of one discipline shares one set of blind spots. It must be diverse in background and sensitive to community attitudes. It must include someone with relevant expertise when it reviews research involving a vulnerable group. And a member with a conflicting interest in a study may not participate in its review, except to answer questions the board asks. That last clause is why your own department chair cannot quietly approve your protocol.

Key idea: Every membership rule exists to make sure at least one person in the room does not share the investigator's assumptions.

The threshold questions

Two definitions decide whether the IRB has jurisdiction at all. First, is it research - a systematic investigation designed to contribute to generalizable knowledge? A hospital collecting data purely to improve its own operations may not meet this bar. Second, does it involve human subjects - living individuals about whom the researcher obtains data through intervention or interaction, or whose identifiable private information is used? If both answers are yes, the study needs review.

Both definitions have edges worth learning. The 2018 Common Rule explicitly deems four activities not to be research. The first is certain scholarly and journalistic work focused on specific individuals, such as oral history and biography. The second is public health surveillance carried out by a public health authority. The third is collection and analysis for criminal justice purposes. The fourth is authorized national security activity. Meanwhile the human-subject definition turns on two routes: either intervention or interaction with a living person, or the use of identifiable private information. Research on people who have died, or on a dataset that is genuinely not identifiable, falls outside the definition even though it plainly involves people.

Marcus now has three different answers in one project. The teacher survey is research with human subjects, by interaction. The student interviews are research with human subjects, by interaction, with the added complication that the participants are minors. The district records may or may not qualify, depending entirely on whether he receives identifiers. If the district gives him a file already stripped of names and student identifiers, and he has no key and no realistic path back to individuals, that arm is likely not human-subjects research at all. He still cannot decide that himself. He asks the board for a determination and files the answer.

Key idea: One project can contain arms that are research and arms that are not, so the threshold questions apply to each component rather than to the study as a whole.

Three tiers of review

Not all research carries equal risk, so review is calibrated in three ascending levels of scrutiny.

  1. Exempt. Certain low-risk categories - some anonymous surveys, research in normal educational settings, secondary analysis of publicly available de-identified data - may be exempt from continuing review. The decisive rule: the investigator does not decide exemption unilaterally. The IRB or its designated official makes that determination, precisely because researchers cannot be trusted to judge their own studies harmless.
  2. Expedited. Research that poses no more than minimal risk and fits defined categories may be reviewed by the chair or a single experienced member rather than the convened board. Minimal risk is a term of art: the probability and magnitude of harm are no greater than those ordinarily encountered in daily life or during routine physical or psychological examinations.
  3. Full board. Research that exceeds minimal risk - a new drug, a stressful manipulation, work with a vulnerable population - must be reviewed at a convened meeting with a majority of members present and a favorable vote.

Two details separate people who understand the system from people who guess at it. First, the exempt categories at 45 CFR 46.104 are a closed list, not a judgment of harmlessness. They cover things like normal educational practices, certain surveys and interviews with adults, benign behavioral interventions, and secondary research on already-collected data. Several of them require a limited IRB review, in which someone checks the privacy safeguards even though the study is exempt from everything else. Second, a reviewer acting under expedited procedures holds all the board's authority except one: an expedited reviewer may approve a study but may not disapprove it. Only the convened board can reject a protocol, which is a deliberate check on a single person's power.

Marcus fits the pattern in three pieces. The anonymous teacher survey plausibly falls in an exempt category for surveys of adults. The interviews with seventh-graders do not, because the survey and interview exemption is narrowed for children; that arm needs at least expedited review, with parental permission and child assent. The record-linkage arm may return a determination of not human-subjects research. One project, three answers, and a competent board will give all three in a single letter.

Key idea: Exempt does not mean harmless and expedited does not mean rubber-stamped - each is a defined regulatory category with its own conditions and its own limits on reviewer power.

TierRisk levelWho reviews
ExemptLow, in a defined categoryIRB official confirms exemption
ExpeditedNo more than minimal riskChair or single member
Full boardGreater than minimal riskConvened committee, majority vote

What the board actually votes on

Reviewers are not free to approve on general impressions. Section 46.111 lists the criteria that must all be satisfied, and a protocol that addresses them directly moves faster than one that does not. In condensed form, the board must find seven things. Risks are minimized. Risks are reasonable against anticipated benefits. Selection of subjects is equitable. Consent will be sought and documented appropriately. The plan monitors data for safety where that is needed. Privacy and confidentiality are protected. And additional safeguards are in place for vulnerable participants. You can write a protocol straight against that list.

Marcus's weakest point under those criteria is confidentiality, not physical risk. If a teacher describes her principal in a recorded interview and the recording is stored in a district-owned cloud account, she can be identified and her employment can be affected. The board will ask where recordings live, who holds the key, when transcripts are stripped of names, and what he will do if a student discloses harm during an interview. Those are the questions that decide the outcome.

Key idea: Write the protocol against the seven approval criteria in 45 CFR 46.111 and the board has nothing left to ask you for.

Ongoing obligations

Approval is not a one-time event. Investigators must report adverse events and unanticipated problems, submit amendments for any change to an approved protocol before implementing it, and undergo continuing review for higher-risk studies. The IRB can suspend or terminate research that endangers participants. The lesson to carry forward: the IRB is not an adversary to be outmaneuvered but a partner in a promise. Build time for review into your project from the start, describe risks honestly, and never begin data collection on human participants before you hold written approval.

The 2018 revision changed some of this in a direction worth knowing. Continuing review is no longer automatically required for studies eligible for expedited review, or for those approved under limited IRB review, or for studies that have progressed to data analysis only. Where continuing review does apply, approval periods still run no longer than a year. Separately, most federally funded multi-site studies in the United States must now rely on a single IRB of record rather than collecting a separate approval at every site, a change meant to cut duplicate review of the identical protocol.

Key idea: Approval is a state you maintain, not an event you complete - amendments before you act, prompt reports of unanticipated problems, and continuing review where it still applies.

Where people get stuck

Five errors account for most rejected or delayed submissions.

  • Self-certifying exemption. The single most common mistake, and the easiest to avoid. The investigator has an interest in the answer, so the investigator does not get to give it. Submit and let the board or its designee decide.
  • Confusing exempt with unregulated. Exempt research is still research and still carries obligations of honesty, privacy protection, and often limited review. Exemption removes continuing review, not ethics.
  • Applying one tier to a whole project. Marcus has three arms and three answers. Describe each component and let the board sort them.
  • Treating minimal risk as intuition. Minimal risk is a defined benchmark - the probability and magnitude of harm found in daily life or routine examinations - not a feeling that nothing bad will happen. Interviewing children about their classroom is not automatically minimal risk once the questions touch on treatment by adults.
  • Starting before approval. Data collected before written approval usually cannot be rescued by approving the study afterward. Boards can rarely authorize research retroactively, and journals increasingly ask for the approval date.

Key idea: Nearly every delay traces to one of two habits - deciding a regulatory question yourself, or describing the project in a lump instead of by component.

Recap

  • Institutional review boards enforce the Belmont principles under the Common Rule at 45 CFR 46, with equivalents in other countries.
  • Membership rules require at least five members of varying backgrounds, a scientist, a nonscientist, and an unaffiliated member, and they bar conflicted members from reviewing their own work.
  • Jurisdiction turns on two definitions: research aimed at generalizable knowledge, and human subjects reached by intervention, interaction, or identifiable private information.
  • The 2018 rule deems certain activities not research, including oral history, public health surveillance, and criminal justice collection.
  • Review comes in three tiers - exempt, expedited, and full board - and the investigator never determines the tier.
  • An expedited reviewer may approve but may not disapprove; only the convened board can reject a protocol.
  • Section 46.111 lists the criteria the board must find satisfied, and writing the protocol against that list is the fastest route to approval.
  • Obligations continue after approval through amendments, reporting of unanticipated problems, and continuing review where still required.

Sources

  1. U.S. Department of Health and Human Services. (2018). Definitions for purposes of this policy, 45 C.F.R. 46.102. Electronic Code of Federal Regulations. ecfr.gov
  2. U.S. Department of Health and Human Services. (2018). Exempt research, 45 C.F.R. 46.104. Electronic Code of Federal Regulations. ecfr.gov
  3. U.S. Department of Health and Human Services. (2018). IRB membership, 45 C.F.R. 46.107. Electronic Code of Federal Regulations. ecfr.gov
  4. U.S. Department of Health and Human Services. (2018). Expedited review procedures, 45 C.F.R. 46.110. Electronic Code of Federal Regulations. ecfr.gov
  5. U.S. Department of Health and Human Services. (2018). Criteria for IRB approval of research, 45 C.F.R. 46.111. Electronic Code of Federal Regulations. ecfr.gov
  6. Federal Register. (2017). Federal policy for the protection of human subjects: Final rule. 82 Fed. Reg. 7149. federalregister.gov
  7. National Institutes of Health. (2016). Final NIH policy on the use of a single institutional review board for multi-site research (NOT-OD-16-094). grants.nih.gov
Key terms
Institutional Review Board (IRB)
A committee that reviews and monitors research to protect the rights and welfare of human participants.
Common Rule
The U.S. federal policy (45 CFR 46) governing human-subjects research across federal agencies.
Minimal risk
Risk no greater in probability or magnitude than that of daily life or routine examinations.
Exempt research
Low-risk research in defined categories that the IRB, not the investigator, certifies as exempt from continuing review.
Expedited review
IRB review by the chair or a single member for minimal-risk studies in defined categories.
Continuing review
The periodic re-evaluation of an approved higher-risk study to confirm ongoing protection of participants.

Privacy, Confidentiality, and Data Protection

  • Distinguish privacy, confidentiality, and anonymity and the protections each requires.
  • Identify practical safeguards for identifiable data, including de-identification and access controls.

Even a study that never touches a participant's body can harm them - through a breach of the sensitive information they shared. A leaked record of a mental-health diagnosis, an immigration status, or an illegal behavior can cost someone a job, a relationship, or their liberty. Protecting information is therefore a core ethical duty, and it rests on three terms that students routinely blur.

One study will carry the argument. Leah Ostrom studies why people who inject drugs drop out of HIV care in a city of 90,000. Her dataset holds, for each of 240 participants, an HIV test result, a drug-use history, recent arrests, current housing status, and a set of GPS-tagged interview locations. Nobody in her study will be physically harmed. Yet if that file leaked, participants could lose housing, custody of children, employment, or their liberty. The entire risk profile of her project is informational, and every safeguard she designs is either about who may see the file or about what the file contains in the first place.

Key idea: In most social, behavioral, and health-records research the whole risk is informational, which means the protocol is really a data-security document with questions attached.

In plain terms

Three words get used as if they meant the same thing. They do not. Privacy is about the person: how much of themselves they choose to hand over. Confidentiality is about the file: what you do with what they handed over. Anonymity is about the link: whether anyone, including you, can trace a row back to a human being. You can respect privacy and then break confidentiality by leaving a laptop in a car. You can guard confidentiality perfectly while violating privacy by asking questions you had no business asking. And you can only claim anonymity if you truly cannot reconnect the data to people, which is rarer than most researchers assume. Get the three words straight and the practical decisions follow. Blur them and you will promise something you cannot deliver.

Three distinct concepts

  • Privacy is about people: a person's control over the extent to which they share themselves - their body, thoughts, and information - with others. Respecting privacy means, for example, not observing or contacting people in ways they would find intrusive, and not asking for more sensitive information than the research genuinely needs.
  • Confidentiality is about data: the researcher's obligation to control who can access identifiable information a participant has shared, and to use it only as agreed. Confidentiality is a promise about the handling of data that has already been collected.
  • Anonymity is the strongest protection: data are collected with no identifiers at all, so that not even the researcher can link a response to a person. True anonymity makes a confidentiality breach impossible, but it also rules out follow-up and longitudinal linkage.

Ostrom's design forces the trade-off into the open. She wants to interview each participant three times over a year, so she needs to know who is who. That rules out anonymity from the start. What she can do is separate the identity from the content. The interview file carries a coded study number and nothing else. A separate file links that number to a name and a phone contact. Different people hold each one. Researchers often call such a study anonymous in the consent form. It is not, and saying so is a false promise. The honest phrasing is that responses are confidential and stored under a code.

Key idea: If you can contact a participant again, your data are not anonymous - and telling them otherwise is a promise you have already broken.

Identifiers and de-identification

Data are identifiable when they contain, or can be readily linked to, information naming an individual. Direct identifiers are obvious: name, address, medical-record number. Indirect (quasi-) identifiers are subtler: a combination of birth date, sex, and postal code can uniquely single out a person even with names removed. This is why de-identification - stripping identifiers so a record can no longer reasonably be traced to an individual - must consider combinations, not just obvious fields. A related technique, pseudonymization, replaces identifiers with a code while a separate, secured key allows re-linkage when necessary; it is reversible and so offers weaker protection than true de-identification.

U.S. health-privacy law gives the problem two concrete solutions, and they are worth knowing even outside clinical settings because they are the clearest available standards. Under the HIPAA Privacy Rule at 45 CFR 164.514, data count as de-identified by one of two routes. The Safe Harbor method removes eighteen specified categories of identifier. Names go. So do all geographic units smaller than a state, apart from limited ZIP-code prefixes. So do all date elements finer than the year. Contact details, record and account numbers, device and vehicle identifiers, web and network addresses, biometric data, full-face images, and any other unique code all go too. One more condition applies. The holder must have no actual knowledge that what remains could identify anyone. The expert determination method works differently. A qualified statistician certifies that the risk of re-identification is very small, which lets you keep more detail under a documented analysis.

Both routes exist because removing names accomplishes very little on its own. Work on re-identification has repeatedly shown how thin the protection is. One widely cited analysis estimated that the great majority of people in an incomplete, heavily sampled dataset could still be correctly re-identified from a modest set of demographic attributes. Ostrom's file is a case study in the problem. In a city of 90,000, a single row reading HIV-positive, age 34, unstably housed, arrested twice in the past year, interviewed near a specific intersection is very likely unique, even though it contains no name.

Key idea: Removing names is not de-identification - what identifies people is the combination of ordinary details, so protection has to be judged on the whole record at once.

Practical safeguards

  1. Collect the minimum. The best-protected datum is the one never gathered. Do not collect identifiers you do not need.
  2. Separate and secure. Store the key that links codes to identities separately from the research data, under access controls and encryption.
  3. Limit access. Restrict identifiable data to team members who need it, and train them in their obligations.
  4. Plan the end. Specify how long identifiable data will be retained and how it will be destroyed or fully de-identified afterward.

Applied to Ostrom's project, these become specific commitments a board can check. She records interviews on encrypted devices and uploads them to institutional storage the same day rather than keeping them on a phone. She reports location by neighborhood rather than by GPS coordinate. She replaces exact dates of arrest with month and year. She stores the name-to-code key on a separate encrypted drive to which only she and one senior colleague have access, and she destroys it once the third wave of interviews is complete. She trains every interviewer on what may be discussed outside the team, which is nothing. None of that is exotic. It is the ordinary craft of handling data about people who can be hurt by it.

Key idea: Confidentiality is engineered rather than promised - the protection lives in what you collect, where it sits, who holds the key, and when you destroy it.

Limits and legal tools

Confidentiality is a strong promise but not always an absolute one. Certain disclosures may be legally mandated - for example, credible threats of serious harm or, in many jurisdictions, suspected child abuse - and participants must be told of such limits during consent so the promise you make is one you can keep.

There is also a legal tool for the hardest case. Suppose a court issues a subpoena for your files. Researchers studying sensitive topics can sometimes obtain an instrument that shields identifiable research data from forced release. In the United States it is called a Certificate of Confidentiality. The overarching principle is honesty. Promise only the protection you can actually deliver. Describe its limits plainly. Then engineer your data handling to honor it.

Certificates of Confidentiality repay a closer look, because researchers routinely overestimate them. Since the 21st Century Cures Act, research funded by the National Institutes of Health that collects identifiable, sensitive information is covered automatically rather than by application. The protection is real: it bars compelled disclosure of identifiable research information in federal, state, and local civil, criminal, administrative, and legislative proceedings. It has three important limits. It does not override other legal duties such as mandated reporting of child abuse in many jurisdictions. It does not stop a researcher from disclosing voluntarily where the participant has consented. And it protects the information rather than the participant, so it cannot prevent harm from a careless breach.

Outside the United States the frame differs but the obligations converge. Under the European General Data Protection Regulation, health data, sex life, and ethnic origin are special categories whose processing is prohibited unless a specific exception applies. A separate article allows derogations for scientific research provided appropriate safeguards are in place, and it names data minimization and pseudonymization as the expected measures. In other words, the same two instincts - collect less, separate the key - appear in both systems.

Key idea: Legal instruments strengthen a promise you can already keep; none of them repairs a dataset you should never have collected in that form.

Where people get stuck

Four errors recur, and each one shows up in consent forms.

  • Calling a coded study anonymous. If a key exists anywhere, or if you can recontact participants, the study is confidential rather than anonymous. Use the accurate word in the consent form; participants deserve to know which promise they are receiving.
  • Believing name removal equals de-identification. Quasi-identifiers do the identifying. Small geography, exact dates, rare conditions, and unusual occupations single people out with no name in sight.
  • Promising absolute confidentiality. Mandated reporting duties, court orders in some contexts, and institutional audit rights all exist. State the limits during consent, in plain language, before anyone discloses anything.
  • Treating security as the technology office's problem. Most breaches in research are mundane: an unencrypted laptop, a shared drive with open permissions, a transcript emailed to a personal account. The controls that matter are habits, not products.

Key idea: Say exactly what you can protect, name the exceptions before anyone speaks, and then build the pipeline that makes the statement true.

Recap

  • Privacy concerns a person's control over what they share, confidentiality concerns the handling of data already shared, and anonymity means no link exists at all.
  • Longitudinal designs cannot be anonymous, so honest consent forms describe them as confidential and coded.
  • Direct identifiers are obvious; quasi-identifiers such as small geography, exact dates, and rare attributes do most of the real identifying.
  • The HIPAA Privacy Rule offers two de-identification routes: Safe Harbor removal of eighteen identifier categories, or expert determination that re-identification risk is very small.
  • Re-identification research shows that small combinations of ordinary attributes are often unique, especially in small geographies.
  • Pseudonymization is reversible and therefore weaker than true de-identification, but it is the practical compromise for follow-up studies.
  • Certificates of Confidentiality bar compelled disclosure but do not override mandated reporting, consented disclosure, or ordinary carelessness.
  • The GDPR treats health, sex life, and ethnic origin as special categories and permits research use only with safeguards such as minimization and pseudonymization.

Sources

  1. U.S. Department of Health and Human Services. (2013). Other requirements relating to uses and disclosures of protected health information, 45 C.F.R. 164.514. Electronic Code of Federal Regulations. ecfr.gov
  2. U.S. Department of Health and Human Services. (2013). Privacy of individually identifiable health information, 45 C.F.R. Part 164, Subpart E. Electronic Code of Federal Regulations. ecfr.gov
  3. Rocher, L., Hendrickx, J. M., & de Montjoye, Y.-A. (2019). Estimating the success of re-identifications in incomplete datasets using generative models. Nature Communications, 10, 3069. nature.com
  4. National Institutes of Health. (n.d.). Certificates of Confidentiality (CoC). grants.nih.gov
  5. European Union. (2016). Processing of special categories of personal data, Article 9. General Data Protection Regulation. gdpr-info.eu
  6. European Union. (2016). Safeguards and derogations relating to processing for archiving purposes in the public interest, scientific or historical research purposes, Article 89. General Data Protection Regulation. gdpr-info.eu
  7. U.S. Department of Health and Human Services. (2018). Criteria for IRB approval of research, 45 C.F.R. 46.111. Electronic Code of Federal Regulations. ecfr.gov
Key terms
Privacy
A person's control over the extent to which they share their body, thoughts, and information with others.
Confidentiality
The researcher's obligation to control access to identifiable data and use it only as agreed.
Anonymity
The condition in which data carry no identifiers, so no response can be linked to an individual.
Quasi-identifier
A combination of non-obvious fields (such as birth date, sex, and postal code) that can uniquely identify a person.
De-identification
Stripping direct and indirect identifiers so a record can no longer reasonably be traced to an individual.
Certificate of Confidentiality
A U.S. legal instrument that helps shield identifiable research data from compelled disclosure.

Vulnerable Populations

  • Explain what makes a population vulnerable in the research context.
  • Identify the specific safeguards owed to major vulnerable groups under federal regulation.

The Belmont principle of respect for persons requires extra protection for those with diminished autonomy, and justice requires that the vulnerable not be exploited simply because they are available. Federal regulation translates these duties into specific safeguards for named groups. Understanding why each group is vulnerable clarifies why its protections take the shape they do.

Amina Sow's study will show how tangled this gets in practice. She is testing whether a text-message support tool improves medication adherence among adolescents with sickle cell disease. Her intended sample is 180 participants aged 13 to 19. Within that one sample sit at least five distinct regulatory situations. Some participants are minors and some are legal adults. Two are pregnant. Four live in foster care as wards of the state. Several have measurable cognitive effects from silent strokes, a known complication of the disease. And the condition disproportionately affects a population with well-documented reasons to distrust medical research. One protocol, five different sets of obligations, and no single label that covers them all.

Key idea: Vulnerability is not a property of people but of situations, so the same sample can require several different protections at once.

In plain terms

Ask two questions about anyone you plan to enroll. Can this person understand and decide? And is this person free to say no? A young child fails the first. A prisoner may pass the first and fail the second. Someone in acute pain may fail both this morning and pass both next week. The regulations sort people into named groups because rules need categories, but the underlying judgment is always about those two capacities in this situation. There is a third question that people forget. If we protect this group by excluding it, who pays? Usually the group itself, because clinicians end up guessing at doses and treatments that were never studied in them. Protection that always means exclusion is not protection.

What vulnerability means here

A population is vulnerable when its members have a reduced ability to protect their own interests in the research relationship. Three things can cause it. There may be compromised capacity to give informed consent, as with young children or a person with dementia. There may be a situation of dependency or constraint that undermines free choice, as with prisoners, employees, or a researcher's own students. Or there may be heightened susceptibility to harm, as with a fetus. The same person can be vulnerable in one respect and fully autonomous in another. So the analysis is contextual rather than a fixed label.

The regulation reflects that contextual view. Section 46.111(b) covers participants who are likely to be vulnerable to coercion or undue influence. It names four examples: children, prisoners, individuals with impaired decision-making capacity, and economically or educationally disadvantaged persons. Where such participants are enrolled, the board must find that extra safeguards are built into the study.

The phrasing matters. It says "likely to be vulnerable to coercion or undue influence." That points at the relationship, not at a diagnosis. It also leaves the list open. Boards therefore apply extra scrutiny to groups the regulations never mention. Students recruited by their own instructor. Employees recruited by their employer. Undocumented immigrants. People in acute crisis. Members of small communities where nobody can really be anonymous.

Key idea: The named groups are examples rather than a complete list, and the operative test is whether the person's circumstances compromise capacity or free choice.

Groups with specific protections

  • Children. Because minors cannot give legally valid consent, research requires parental permission plus the child's assent, and the regulations limit the level of risk permissible in relation to the potential benefit. Research offering no direct benefit to the child is capped at low levels of risk.
  • Prisoners. Incarceration inherently compromises voluntariness: the promise of better conditions or early consideration can be powerfully coercive. Special rules restrict prison research to categories that either carry minimal risk or stand to benefit prisoners, and require an IRB member with relevant expertise (often a prisoner representative).
  • Pregnant women, fetuses, and neonates. Because interventions can affect two parties, one of whom cannot consent, regulation imposes additional conditions balancing potential benefit against risk to the fetus, while cautioning against reflexive, blanket exclusion that would leave the group under-studied.
  • Persons with impaired decision-making capacity. For those who cannot consent - due to cognitive disability, acute illness, or unconsciousness in emergency research - a legally authorized representative may consent, subject to tighter risk limits and, where possible, the person's own assent.

How the risk ceilings actually work

The children's rules are the most systematic, and they repay learning because they show the general logic. Subpart D sorts pediatric research into four tiers by risk and benefit. Research at no more than minimal risk is approvable with permission and assent. Research above minimal risk is approvable when it holds out a prospect of direct benefit to the individual child and the risk is justified by that benefit. Research above minimal risk with no prospect of direct benefit is harder. Three conditions must all hold. The increase over minimal risk must be minor. The procedures must be reasonably commensurate with the child's own medical or social experience. And the study must be likely to yield knowledge about the child's own disorder or condition. Anything beyond those tiers requires review at the federal level.

That third tier is the one Sow has to reason about. Adding an extra blood draw for a research biomarker offers her participants nothing directly. It is defensible under the narrow tier because these adolescents already undergo frequent blood draws for their condition, so the procedure is commensurate with their experience, and the knowledge concerns sickle cell disease itself. The same blood draw in healthy adolescents would not clear the same bar.

Her four participants in foster care add a further requirement. Children who are wards of the state may join research without direct benefit only under extra conditions. Chief among them is the appointment of an advocate for each child. That advocate must be independent of the research team, the agency, and the guardian.

Her two pregnant participants bring Subpart B into play. It imposes its own conditions. Risk to the fetus must be minimal where there is no direct benefit. And people involved in the study may take no part in decisions about terminating a pregnancy.

Key idea: Pediatric protections are not a single rule but a ladder - the higher the risk and the weaker the direct benefit, the narrower the conditions and the higher the approval authority.

Prisoners: a category built on setting, not capacity

Prisoner protections work differently, because the problem is not understanding but freedom. Subpart C therefore restricts prison research to four permissible categories rather than adjusting the consent process. Research may study the causes and effects of incarceration and criminal behavior, at minimal risk. It may study prisons as institutions, at minimal risk. It may address conditions that particularly affect prisoners as a class, which requires federal consultation. Or it may study practices likely to improve the health or well-being of the participant.

The subpart also reshapes the board itself. At least one member must be a prisoner or a prisoner representative. And a majority of the reviewing members must have no association with the prison beyond serving on the board.

Two practical points follow. Someone can become a prisoner mid-study, which triggers the subpart's requirements for that participant even though the protocol was never written for prison research. And the regulations bar advantages of incarceration - better food, better quarters, earlier consideration - from being so great that they impair judgment about the risks.

Key idea: Where the threat is to voluntariness rather than to capacity, the regulatory answer is to limit what may be studied at all, not to write a better consent form.

The two-edged nature of protection

There is a tension worth naming. Overprotection can itself be an injustice. If pregnant women, children, or the seriously ill are routinely excluded, then medicine is left prescribing to them on the basis of studies never done in them - guessing at doses and effects. The modern stance is therefore not exclusion but appropriate inclusion with heightened safeguards: bring vulnerable groups into research when the knowledge will serve them, but wrap that participation in the strongest protections.

GroupSource of vulnerabilityKey safeguard
ChildrenCannot legally consentPermission plus assent; risk limits
PrisonersConstrained voluntarinessRestricted categories; prisoner representative
Pregnant women / fetusesRisk to a party who cannot consentBenefit-risk conditions; avoid blanket exclusion
Impaired capacityCannot give valid consentAuthorized representative; tighter risk limits

When you design a study touching any of these groups, expect and welcome tighter review. The extra scrutiny is the system keeping the promise made after Willowbrook and Tuskegee - that those least able to protect themselves will be protected most carefully.

Think about what the old rule cost. For decades, drug labels for children said the drug had not been studied in children. Doctors still had to treat sick children. So they guessed at the dose from adult data and body weight. Some of those guesses were wrong, and children were hurt by them. The same thing happened in pregnancy. A woman with epilepsy still needs a drug when she is pregnant. If no one has studied it, her doctor is working blind. Keeping a group out of research does not keep them safe. It just moves the risk somewhere no one is watching.

There is also an implementation gap worth knowing about. Surveys of institutional policies have found wide variation in how U.S. universities define vulnerability and which additional safeguards they require, with many institutions naming groups that federal regulation does not and applying inconsistent standards to the same situation. That variation is not necessarily wrong, since local knowledge matters, but it means you cannot assume a protocol approved at one institution will be treated identically at another.

Key idea: Federal regulation sets a floor and institutions build above it inconsistently, so check your own board's definitions rather than reasoning from the regulations alone.

Where people get stuck

Four errors show up repeatedly in protocols involving these groups.

  • Treating vulnerability as a permanent label. A person with early dementia may have capacity for a low-burden observational study and lack it for a surgical trial. Capacity is judged against the specific decision, not certified once for all purposes.
  • Forgetting the unnamed vulnerable. Students recruited by their instructor and employees recruited by their employer are classic voluntariness problems, and no subpart covers them. The fix is structural: someone outside the power relationship recruits and consents.
  • Assuming a guardian's signature settles everything. Permission from a parent or legally authorized representative is necessary but rarely sufficient. Assent is usually required, sustained refusal is generally respected, and risk ceilings still apply regardless of who signed.
  • Reading protection as a reason to exclude. Excluding pregnant participants from a study of a drug they will inevitably be prescribed does not eliminate risk; it moves the risk from a monitored trial into unmonitored clinical practice.

Key idea: The recurring question is not "is this group vulnerable" but "which specific capacity is compromised here, and what design change repairs it."

Recap

  • Vulnerability means a reduced ability to protect one's own interests, arising from compromised capacity, constrained voluntariness, or heightened susceptibility to harm.
  • Section 46.111(b) names children, prisoners, individuals with impaired decision-making capacity, and economically or educationally disadvantaged persons as examples rather than as a closed list.
  • Subpart D sorts pediatric research into risk tiers, allowing above-minimal-risk research without direct benefit only under narrow, commensurate-experience conditions.
  • Children who are wards of the state require an independent advocate before enrollment in research without direct benefit.
  • Subpart C restricts prison research to four permissible categories and requires a prisoner or prisoner representative on the reviewing board.
  • Subpart B adds conditions for pregnant participants and fetuses, including limits on fetal risk and separation from decisions about terminating a pregnancy.
  • Institutional definitions of vulnerability vary considerably, so local policy must be checked rather than inferred from federal text.
  • Blanket exclusion is itself an injustice; the modern standard is appropriate inclusion wrapped in stronger safeguards.

Sources

  1. U.S. Department of Health and Human Services. (2018). Additional protections for children involved as subjects in research, 45 C.F.R. Part 46, Subpart D. Electronic Code of Federal Regulations. ecfr.gov
  2. U.S. Department of Health and Human Services. (2018). Requirements for permission by parents or guardians and for assent by children, 45 C.F.R. 46.408. Electronic Code of Federal Regulations. ecfr.gov
  3. U.S. Department of Health and Human Services. (2018). Additional protections pertaining to biomedical and behavioral research involving prisoners as subjects, 45 C.F.R. Part 46, Subpart C. Electronic Code of Federal Regulations. ecfr.gov
  4. U.S. Department of Health and Human Services. (2018). Additional protections for pregnant women, human fetuses and neonates involved in research, 45 C.F.R. Part 46, Subpart B. Electronic Code of Federal Regulations. ecfr.gov
  5. U.S. Department of Health and Human Services. (2018). Criteria for IRB approval of research, 45 C.F.R. 46.111. Electronic Code of Federal Regulations. ecfr.gov
  6. Schwenzer, K. J. (2008). Protecting vulnerable subjects in clinical research: Children, pregnant women, prisoners, and employees. Respiratory Care, 53(10), 1342-1349. pubmed.ncbi.nlm.nih.gov
  7. Jonathan, I., Akers, E., Shi, M., & Resnik, D. B. (2024). Vulnerable research participant policies at U.S. academic institutions. Journal of Empirical Research on Human Research Ethics, 19(4-5), 220-225. pmc.ncbi.nlm.nih.gov
Key terms
Vulnerable population
A group whose members have reduced ability to protect their own interests in the research relationship.
Compromised capacity
A diminished ability to give informed consent, as in young children or persons with dementia.
Legally authorized representative
A person empowered to consent to research on behalf of someone who cannot consent for themselves.
Parental permission
A guardian's authorization for a child's research participation, paired with the child's assent.
Prisoner representative
An IRB member with relevant expertise required when a board reviews research involving prisoners.
Appropriate inclusion
The principle of enrolling vulnerable groups when research benefits them, under heightened safeguards, rather than excluding them.

Module 3: Animal Research Ethics

The ethical framework, principles, and oversight governing the use of animals in research.

The Ethics of Animal Research and the 3Rs

  • Summarize the ethical positions that frame the debate over animal research.
  • Apply the 3Rs framework and describe the oversight role of the animal-care committee.

A great deal of biomedical and behavioral knowledge - and nearly every modern medicine - rests on research using animals. That fact sits atop a genuine moral problem: animals can suffer but cannot consent, and they are used for human ends. Serious ethics does not resolve this by pretending the problem away in either direction. It acknowledges the moral weight of animal welfare and constructs a framework to ensure that when animals are used, their use is justified, minimized, and as humane as possible.

Hannah Reiss will make the problem concrete. She studies chronic nerve pain. Her model requires a surgical injury to a nerve in the hind leg of a mouse, after which the animal develops a lasting hypersensitivity that resembles human neuropathic pain. She then tests whether a candidate drug reduces it. Her protocol asks for 96 mice. The work is not gratuitous. Neuropathic pain affects millions of people and responds poorly to existing drugs. It is also, unavoidably, the deliberate infliction of pain on animals who cannot agree to it. Every part of the framework in this lesson exists to answer the question her protocol raises: what would make this justified?

Key idea: Animal research ethics does not pretend the moral cost is zero - it asks whether the cost is necessary, minimized, and outweighed.

In plain terms

Three questions, asked in order. Do you need an animal at all? If yes, how few can you use? And how much of the suffering can you remove without destroying the science? That is the whole framework. The order matters. People often start at the third question, because refining a procedure feels like progress. But if a cell model would answer the question, the kindest surgery in the world is still one that never needed to happen. Notice also that the second question cuts both ways. Using too many animals is waste. Using too few is worse waste, because a study too small to give an answer has caused suffering for nothing. Good statistics is an animal-welfare tool.

The ethical landscape

Positions on animal research span a spectrum. A strict animal-rights view holds that sentient animals have interests that may not be sacrificed for others' benefit, which would forbid most research. A utilitarian view - influential in shaping actual policy - weighs the suffering caused against the benefits gained, permitting research when the expected benefit is large and the suffering is minimized. A welfarist consensus, embodied in law and oversight, accepts that animals may be used but insists that their capacity to suffer imposes strict duties to protect their welfare. Most modern regulation is welfarist in structure: it does not ban animal research, but it heavily constrains it.

One concept underlies all the regulation: sentience, the capacity to have experiences that can go well or badly. Moral weight tracks sentience rather than species membership or usefulness. That is why oversight scales with the evidence for rich experience. Non-human primates face the strictest limits, mammals face substantial ones, and invertebrates such as fruit flies and nematodes fall largely outside the frameworks. It also explains why the debate keeps moving. As evidence accumulates about the capacities of octopuses, decapod crustaceans, and fish, several jurisdictions have extended protections to cover them.

Reiss cannot escape this by arguing that pain research is important. Importance affects the benefit side of the ledger, and the benefit side is only half the analysis. A reviewer will still ask why a mouse, why 96, and what she will do when an animal reaches a level of distress she did not anticipate.

Key idea: Regulation tracks the capacity to suffer rather than species rank or scientific importance, which is why the same experiment faces different limits in different animals.

The 3Rs

The organizing framework for humane animal research, proposed by Russell and Burch in 1959, is the 3Rs. It has become the internationally accepted standard and is worth memorizing precisely.

  • Replacement. Wherever possible, replace animals with alternatives that achieve the scientific aim without them - cell and tissue cultures, computer models, or studies in organisms of lower sentience. The first ethical question is always whether an animal is needed at all.
  • Reduction. Use the fewest animals necessary to obtain valid, statistically meaningful results. This is where good experimental design and power analysis become an ethical matter: an underpowered study wastes animals by being unable to answer its question, and an oversized one wastes them needlessly.
  • Refinement. Modify procedures and husbandry to minimize pain, distress, and lasting harm and to improve welfare - better analgesia and anesthesia, humane endpoints that stop a study before suffering becomes severe, and enriched living conditions.

Run Reiss's protocol through each R and the framework stops being abstract. On Replacement, she must document a real search. Cultured neurons can show whether her drug binds its target, and that work should come first. What no dish can currently model is a whole nervous system generating a pain state over weeks. That is a defensible answer, and it is defensible only because she looked.

On Reduction, 96 is not a round number chosen for convenience. A power calculation based on pilot variance sets the group size needed to detect the smallest effect worth finding. Two further moves reduce the total. She can share tissue from each animal with a colleague studying a different question, and she can use a within-animal design where each mouse serves as its own baseline before injury.

On Refinement, she faces the genuine dilemma of her field. Analgesia relieves suffering, and analgesia also suppresses the very outcome she is measuring. The honest resolution is not to abandon pain relief. It is to give analgesia during and immediately after surgery, when the pain is surgical rather than experimental, then to define humane endpoints for the experimental phase - a weight-loss threshold, a limit on self-mutilation of the affected paw, a maximum duration - and to remove any animal that reaches one. She also shortens the observation window to the minimum that answers the question.

Key idea: The 3Rs are questions you must answer with evidence, not values you may assert - a Replacement search, a power calculation, and predefined endpoints are all documents a reviewer can read.

Severity, and why it is scored in advance

Modern frameworks ask investigators to classify expected suffering before the work begins. European law, for example, sorts procedures into non-recovery, mild, moderate, and severe, and requires that the actual severity experienced be assessed afterward and compared with the prediction. That retrospective step is what converts severity classification from paperwork into learning. It also makes the harm side of the harm-benefit analysis explicit rather than implied.

Reiss would classify her nerve-injury model as moderate. That classification carries consequences: more frequent monitoring, veterinary review, and a stronger justification requirement. If her retrospective assessment shows animals actually experienced severe distress, the next version of the protocol must change. The system assumes prediction will sometimes be wrong and builds in a correction.

Key idea: Predicting severity in advance and checking it afterward turns welfare from an intention into a measurement.

Oversight and the committee

Just as human research requires an IRB, animal research requires review by an institutional animal-care and use committee (in the United States, the IACUC). Before work begins, investigators submit a protocol that must justify the species and numbers, demonstrate that alternatives were seriously considered (the Replacement search), describe how pain will be assessed and relieved, and specify humane endpoints. The committee - which must include a veterinarian and a member of the public - can require changes or reject the protocol. Ongoing duties include veterinary care, inspection of facilities, and reporting.

RQuestion it forcesExample measure
ReplacementIs an animal needed at all?Cell culture or computer model
ReductionWhat is the fewest that yields valid results?Sound power analysis and design
RefinementHow is suffering minimized?Analgesia and humane endpoints

The 3Rs do more than soften a hard practice; they align good ethics with good science. A well-designed study that uses the right number of animals under low-stress conditions produces cleaner, more reproducible data than a wasteful, stressful one. As with human research, the deepest point is that ethical rigor and scientific rigor are not opponents. They are the same discipline applied to different stakeholders.

Reporting closes the loop. The ARRIVE guidelines set out what an animal study must state in print: species, strain, sex, and age; how many animals were used and how the number was decided; how animals were allocated to groups and whether assessors were blinded; the exact procedures and any analgesia; and how many animals were excluded and why. Journals increasingly require it. The reason is not bureaucratic. A paper that omits sample-size reasoning and blinding cannot be replicated, which means the animals produced no usable knowledge. Under-reporting is a Reduction failure disguised as a writing problem.

Key idea: An animal study that cannot be replicated from its own methods section has wasted the animals it used, which is why reporting standards belong inside the 3Rs rather than beside them.

Where people get stuck

Four confusions recur in this material.

  • Ranking the 3Rs by convenience. Replacement comes first, and it is the one most often answered with a sentence rather than a search. A documented literature search for alternatives is the expected evidence.
  • Thinking Reduction always means fewer. It means the right number. An underpowered study is the worst outcome for welfare, because every animal in it suffered for a result that cannot answer the question.
  • Confusing Refinement with Replacement. Switching from mice to a cell line is Replacement. Switching from mice to zebrafish is also usually treated as Replacement, since it moves to a less sentient organism. Adding analgesia to a mouse study is Refinement. The test is whether the animal changed or the procedure did.
  • Assuming all research animals are covered by one law. In the United States the Animal Welfare Act regulations exclude purpose-bred mice and rats, which are the overwhelming majority of research animals. Those animals are covered instead through the Public Health Service Policy and the Guide when the work is federally funded, and standards therefore depend on the funding source as well as the species.

Key idea: Ask "did the animal change or did the procedure change" to sort Replacement from Refinement, and never assume one statute covers every species in your facility.

Recap

  • Animal research raises a real moral cost, and the framework aims to justify, minimize, and mitigate it rather than deny it.
  • Positions range from animal rights through utilitarian weighing to the welfarist consensus embodied in law, which permits use under strict constraint.
  • Regulation scales with evidence of sentience, which is why primates, mammals, and invertebrates face different levels of oversight.
  • Russell and Burch's 3Rs - Replacement, Reduction, Refinement - are answered with documents: an alternatives search, a power calculation, and predefined endpoints.
  • Reduction cuts both ways, since an underpowered study wastes animals as surely as an oversized one.
  • Analgesia in a pain model illustrates a genuine Refinement tension, resolved through perioperative pain relief plus humane endpoints rather than through abandoning either goal.
  • Severity is classified before the work and reassessed afterward, which makes the harm side of the analysis explicit and correctable.
  • Reporting standards such as ARRIVE belong to the 3Rs, because an unreplicable study converts animal suffering into no knowledge at all.

Sources

  1. Russell, W. M. S., & Burch, R. L. (1959). The principles of humane experimental technique. Methuen. (Full text hosted by Norecopa) norecopa.no
  2. National Research Council. (2011). Guide for the care and use of laboratory animals (8th ed.). National Academies Press. ncbi.nlm.nih.gov
  3. National Centre for the Replacement, Refinement and Reduction of Animals in Research. (n.d.). The 3Rs. nc3rs.org.uk
  4. Percie du Sert, N., Hurst, V., Ahluwalia, A., Alam, S., Avey, M. T., Baker, M., et al. (2020). The ARRIVE guidelines 2.0. NC3Rs. arriveguidelines.org
  5. European Parliament and Council. (2010). Directive 2010/63/EU on the protection of animals used for scientific purposes. EUR-Lex. eur-lex.europa.eu
  6. Ferdowsian, H. R., & Gluck, J. P. (2015). The ethical challenges of animal research. Cambridge Quarterly of Healthcare Ethics, 24(4), 391-406. pubmed.ncbi.nlm.nih.gov
  7. Office of Laboratory Animal Welfare. (n.d.). PHS policy on humane care and use of laboratory animals. National Institutes of Health. olaw.nih.gov
Key terms
3Rs
The framework of Replacement, Reduction, and Refinement for humane animal research, proposed by Russell and Burch in 1959.
Replacement
Using non-animal alternatives, or less sentient organisms, whenever the scientific aim allows.
Reduction
Using the fewest animals necessary to obtain valid, statistically meaningful results.
Refinement
Modifying procedures and husbandry to minimize pain, distress, and lasting harm.
IACUC
The U.S. Institutional Animal Care and Use Committee that reviews and oversees animal research protocols.
Humane endpoint
A predefined point at which an animal is removed from a study to prevent unnecessary suffering.

Oversight, Regulation, and Protocol Review

  • Describe the regulatory and oversight structures that govern animal research.
  • Explain what an animal-use protocol must justify and how welfare is monitored across a study.

The 3Rs supply the ethical logic of humane animal research; regulation and institutional oversight turn that logic into enforceable practice. Just as the IRB protects human participants, a network of law, committees, and professional standards protects research animals. This lesson describes that machinery and the protocol at its center, so that the requirements read as safeguards rather than hurdles.

Watch one protocol move through a committee. Claire Boateng is the attending veterinarian on her university's animal-care committee. On her desk is the nerve-injury pain study from the previous lesson: 96 mice, a surgical model, a candidate analgesic drug. Boateng is not there to decide whether pain research is worth doing. She is there to answer a narrower set of questions. Are the numbers justified? Was an alternatives search actually performed? Who will perform the surgery, and have they done it before? What happens at three in the morning when an animal is in distress and the investigator is asleep? Her review is mostly about the gap between what a protocol promises and what a laboratory does.

Key idea: Protocol review is not a debate about whether animal research is acceptable - it is a check on whether this specific team can deliver the welfare standards it has written down.

In plain terms

Think of oversight as four things happening at once. A law sets the floor. A committee approves each project before it starts. A veterinarian watches the animals every day. And someone inspects the rooms twice a year. None of the four can be skipped, because each catches a different kind of failure. A good protocol on paper does not stop a cage from being overcrowded. A clean facility does not stop a badly designed experiment. The system is deliberately redundant. That is also why it feels heavy to a first-time investigator. The heaviness is the point, and the way through it is to write the protocol as if a stranger will have to run your study from it, because in an emergency someone will.

Layers of governance

Animal research is typically governed by several overlapping layers. Law sets baseline requirements - in the United States the Animal Welfare Act and, for federally funded work, the Public Health Service Policy, with parallel statutes elsewhere. Institutional oversight is exercised by a standing committee, in the United States the Institutional Animal Care and Use Committee (IACUC). It must include a veterinarian with relevant expertise. It must include at least one practicing scientist. And it must include at least one member unaffiliated with the institution, to represent community concerns. Professional norms and accreditation add a further layer, as do the veterinary staff who provide daily care.

The layers do not overlap neatly, and the seams matter. In the United States the Animal Welfare Act regulations set a minimum committee of three, including the veterinarian and an unaffiliated member. Institutions holding a Public Health Service assurance must meet a higher bar of at least five members, adding a practicing scientist and a nonscientist. Many universities therefore run one committee built to satisfy both. On top of that, roughly a thousand institutions worldwide hold voluntary accreditation from an external body that conducts its own site visits. Accreditation is not required by law anywhere, which is precisely why holding it signals something.

Key idea: Your obligations depend on your funding source and your species mix as much as on your country, so identify which frameworks apply to your facility before you write anything.

The animal-use protocol

Before any animal work begins, investigators submit a protocol that the committee must approve. A rigorous protocol is essentially the 3Rs written out and defended. It must:

  • Justify the species and numbers. Explain why animals are necessary at all, which is a Replacement analysis. Explain why this species fits the question. And show that the number requested is the minimum that yields valid results, which is a Reduction analysis and usually rests on a power calculation.
  • Document a search for alternatives. Most frameworks require a literature search demonstrating that non-animal methods and less-sentient models were seriously considered and found inadequate for the aim.
  • Describe pain and its relief. Classify procedures by their potential to cause pain or distress, and specify the anesthesia, analgesia, and humane endpoints (Refinement) that will limit suffering.
  • Ensure competence. Confirm that personnel are trained to perform the procedures skillfully, since unskilled technique is itself a welfare hazard.

Two further requirements catch investigators by surprise. The first is non-duplication. A protocol must state that the work does not unnecessarily repeat previous experiments. That is a welfare rule, not an originality rule, and it does not forbid replication where replication is the scientific point. The second is pain and distress categorization. Procedures are sorted by whether they cause more than momentary pain and, if so, whether relieving drugs will be used. The hardest category covers procedures where pain relief would compromise the scientific results. Using it is permitted, but it requires written scientific justification and attracts the closest scrutiny. Reiss's experimental phase sits in that category, which is why Boateng reads her endpoint criteria line by line.

Key idea: Declaring that pain relief would confound your results is allowed, and it is also the single clause guaranteed to be read most carefully by everyone who reviews your protocol.

How a committee actually reviews

Committees use two procedures, and the difference is worth knowing before you submit. Under full committee review, the protocol goes to a convened meeting and requires approval by a majority of a quorum. Under designated member review, the protocol is circulated to all members first, and if no member requests full review, one or more designated reviewers handle it. The safeguard is asymmetric and deliberate. A designated reviewer may approve a protocol, require modifications, or refer it to the full committee. A designated reviewer may not withhold approval. Only the convened committee can do that, which mirrors the corresponding rule for human-subjects review.

There is a simple way to think about why the rule is built that way. One person can say yes to a study that is clearly fine. That saves everyone time and costs nothing. One person should not be able to end a study on their own. A rejection can sink a grant, a thesis, or a career. So the power to say no stays with the whole group, where it can be argued about in the open. Say yes alone. Say no together.

Boateng will therefore not reject Reiss's protocol on her own. If she cannot resolve her concerns about the endpoint criteria, she refers it to the full committee. In practice most protocols are approved after a round of modifications, and most modifications concern monitoring frequency, endpoint definitions, and personnel training records rather than the science.

Key idea: One reviewer can approve but never reject, so the fastest route through review is to pre-empt the modifications a reviewer would otherwise have to ask for.

Monitoring across the study

Approval is the beginning of oversight, not the end. Veterinary staff monitor animal health, the committee inspects facilities and programs on a recurring schedule, and investigators must report unexpected outcomes and seek approval before changing an approved protocol. When welfare problems arise that cannot be corrected, the committee has the authority to suspend the activity, mirroring the IRB's power over human research. This continuity of oversight recognizes that welfare depends on daily practice, not merely on a well-written plan.

LayerFunction
Law and policyBaseline legal requirements for care and use
Oversight committee (IACUC)Protocol review, inspection, authority to suspend
Veterinary staffDaily health monitoring and clinical care
Professional standardsAccreditation and discipline-specific norms

Ethics and validity together

A theme worth restating is that this oversight serves science as much as it serves animals. Stressed, poorly cared-for animals yield noisy, less reproducible data; the same practices that reduce suffering - skilled handling, low-stress housing, careful design - also improve the quality and generalizability of results. A protocol that takes welfare seriously is not paying a tax on good research; it is doing good research. When a committee asks you to justify a number or refine a procedure, it is enforcing both an ethical duty to the animals and a scientific duty to the reliability of your findings.

It helps to picture what daily oversight looks like. A technician walks the room each morning and checks every cage. She notes which animals have lost weight and which are not moving well. If something looks wrong, she calls the vet. The vet can treat an animal, and can stop a procedure, without asking the investigator first. That is the point of giving her authority. A scientist who has spent two years on a study is the worst person to decide whether it should pause today.

The monitoring rhythm is specified rather than left to judgment. Committees must review the institution's animal care and use program and inspect its animal facilities at least once every six months, and must report findings, including any significant deficiencies, to the institutional official. Serious or continuing noncompliance, and any suspension of an activity, must be reported to the federal oversight office. That reporting duty is what gives a committee's authority teeth, because a suspension is not an internal matter that can be quietly resolved.

Key idea: Semiannual inspection and mandatory external reporting are what stop institutional oversight from becoming an institution reviewing itself.

Where people get stuck

Four errors account for most protocol delays and most compliance findings.

  • Writing the alternatives search as a sentence. "No alternatives exist" is not a search. Committees expect named databases, the search dates, the terms used, and a short account of why the retrieved alternatives do not answer the question.
  • Vague humane endpoints. "The animal will be euthanized if it appears to be suffering" is unusable at three in the morning by a technician who has never met you. Usable endpoints are numeric and observable: percentage weight loss, a scored body-condition threshold, immobility beyond a stated duration.
  • Drifting from the approved protocol. Adding a strain, changing a dose, or extending an observation window requires an approved amendment first. Doing it and reporting later is noncompliance even when the change reduces harm.
  • Treating training records as an afterthought. Personnel qualification is an explicit regulatory requirement. A skilled surgeon causes less suffering than an unskilled one, so who holds the instruments is a welfare question, not an administrative one.

Key idea: Every recurring compliance problem is the same problem - a protocol that describes intentions instead of specifying observable actions.

Recap

  • Oversight has four layers: law, committee approval, daily veterinary care, and periodic facility inspection, each catching a different failure.
  • Committee composition requirements differ by framework, with a minimum of three members under the Animal Welfare Act regulations and at least five under Public Health Service Policy.
  • Voluntary external accreditation adds a further layer and is not required by law in any jurisdiction.
  • An approvable protocol justifies species and numbers, documents an alternatives search, classifies pain and distress, specifies relief, and confirms personnel competence.
  • Procedures where pain relief would compromise results are permitted only with written scientific justification and receive the closest scrutiny.
  • Committees review by full committee or by designated member; a designated reviewer may approve or require changes but may never withhold approval.
  • Programs and facilities must be reviewed at least every six months, with significant deficiencies reported to the institutional official.
  • Suspensions and serious noncompliance must be reported externally, which is what makes committee authority real.

Sources

  1. U.S. Department of Agriculture. (n.d.). Institutional Animal Care and Use Committee (IACUC), 9 C.F.R. 2.31. Electronic Code of Federal Regulations. ecfr.gov
  2. U.S. Department of Agriculture. (n.d.). Personnel qualifications, 9 C.F.R. 2.32. Electronic Code of Federal Regulations. ecfr.gov
  3. U.S. Department of Agriculture. (n.d.). Animal welfare, 9 C.F.R. Chapter I, Subchapter A. Electronic Code of Federal Regulations. ecfr.gov
  4. Office of Laboratory Animal Welfare. (n.d.). PHS policy on humane care and use of laboratory animals. National Institutes of Health. olaw.nih.gov
  5. Office of Laboratory Animal Welfare. (n.d.). Frequently asked questions. National Institutes of Health. olaw.nih.gov
  6. National Research Council. (2011). Animal care and use program. In Guide for the care and use of laboratory animals (8th ed.). National Academies Press. ncbi.nlm.nih.gov
  7. AAALAC International. (n.d.). Accreditation program. aaalac.org
Key terms
Animal Welfare Act
The principal U.S. federal law setting minimum standards for the care and use of many research animals.
Animal-use protocol
The document, reviewed before work begins, that justifies species, numbers, alternatives, and pain relief.
Search for alternatives
A required demonstration that non-animal or less-sentient methods were considered and found inadequate.
Pain and distress classification
Categorizing procedures by their potential to cause suffering, to guide required relief measures.
Facility inspection
The oversight committee's recurring review of animal-care programs and housing conditions.
Institutional oversight
The committee-based system that reviews, monitors, and can suspend animal research to protect welfare.

Module 4: Integrity of the Scholarly Record

Research misconduct, questionable practices, authorship, and the disclosure of conflicts of interest.

Research Misconduct: Fabrication, Falsification, Plagiarism

  • Define the three categories of research misconduct and distinguish them from honest error.
  • Explain the standard of proof for a finding of misconduct and the harm each type inflicts on the record.

The integrity of science depends on a simple premise: that what researchers report actually happened. When that premise is broken deliberately, the offense is research misconduct, the most serious violation in the profession. U.S. federal policy defines it precisely as fabrication, falsification, or plagiarism - the three are often abbreviated FFP - in proposing, performing, reviewing, or reporting research.

Here is the situation to keep in mind. Yusuf Adeyemi is a second-year postdoc assembling a figure for a new manuscript. He pulls up his laboratory's 2019 paper to match the formatting of a control panel. Something looks familiar. He opens the laboratory's 2021 paper and compares. The two control blots appear to be the same image, cropped slightly differently, but the 2019 paper describes it as a liver sample and the 2021 paper as a kidney sample. Yusuf has not found fraud. He has found a discrepancy. What he does in the next week matters enormously, both for the record and for his own career, and almost nothing in his training has prepared him for it.

Key idea: Most misconduct concerns begin as an unexplained discrepancy, and the person who notices is usually junior, uncertain, and dependent on the person they would be reporting.

In plain terms

Three offenses, one question each. Did the data ever exist? That is fabrication. Did real data get bent? That is falsification. Did someone else make this and get no credit? That is plagiarism. Then a fourth question sits over all three, and it decides everything: did the person mean it? A researcher who mislabels a file and publishes the wrong image has made a mistake. A researcher who reuses an image knowing it came from a different experiment has falsified the record. The picture on the page can look identical in both cases. That is why misconduct findings take months, involve lawyers, and rest on notebooks, file dates, and email rather than on the published figure alone. Yusuf cannot tell from the paper which one he is looking at, and he should not try to.

The three categories

  • Fabrication is making up data or results and recording or reporting them. The fabricator did not run the experiment, or ran it and invented the numbers. It is the manufacture of evidence from nothing.
  • Falsification is manipulating research materials, equipment, or processes, or changing or omitting data or results, such that the research record does not accurately represent the work. This includes altering images, selectively deleting inconvenient data points, or misrepresenting a method. The data existed but were distorted.
  • Plagiarism is the appropriation of another person's ideas, processes, results, or words without giving appropriate credit. It is theft of intellectual property and a deception of the reader about the origin of the work. A related practice, self-plagiarism or redundant publication - reusing one's own previously published text or data as if new - deceives readers and editors even though no one else is robbed. Note the regulatory boundary carefully: because the federal definition speaks of another person's work, reusing your own text is generally handled by journals and institutions as a publication-ethics violation rather than as a federal misconduct finding.

Image handling deserves its own paragraph, because it is where the line is crossed most often and most quietly. Adjusting brightness and contrast across a whole image is normally fine if you disclose it. Adjusting one region so a faint band becomes convincing is falsification. Splicing lanes from separate gels into one apparent gel, without marking the join, is falsification. Reusing a control panel across experiments is falsification if you present it as a fresh control. One screen looked at more than twenty thousand papers in forty biomedical journals. It found problematic image duplications in a few percent of them. About half of those showed features that suggested deliberate manipulation rather than error. Yusuf's discovery is therefore neither rare nor automatically damning.

Key idea: Fabrication invents, falsification distorts, plagiarism steals - and the most common modern form of falsification is a reused or edited image rather than a made-up number.

The boundary with honest error

Crucially, the federal definition contains a mental-state requirement: misconduct must be committed intentionally, knowingly, or recklessly, and it does not include honest error or honest differences of opinion. Science advances by making and correcting mistakes; a researcher who reports a result later found to be wrong, through no deception, has not committed misconduct. The dividing line is not whether the work turned out to be correct but whether the researcher was honest. This is why a good-faith correction or even a retraction of an honest mistake is a mark of integrity, not an admission of guilt.

The regulation names three mental states, in descending order of certainty. Acting intentionally means doing it on purpose. Acting knowingly means being aware that the act is wrong while doing it. Acting recklessly means proceeding with conscious disregard for a substantial risk that the record is false. That third category is the one people underestimate. A principal investigator who signs off on figures she never checked, in a laboratory where she knows records are chaotic, may meet the recklessness standard even though she invented nothing herself. Recklessness is how supervision failures become findings.

Key idea: You do not have to intend a lie to be found responsible - conscious disregard for whether your published record is true is enough.

The standard and the process

Because the accusation is grave, findings of misconduct follow a defined process, typically requiring that the act be a significant departure from accepted practices, that it was committed with a culpable state of mind, and that it be proven by a preponderance of the evidence. Institutions conduct an inquiry and, if warranted, a formal investigation, with protections for both the accused and the person who raised the concern (the whistleblower). Confirmed misconduct can end careers, trigger retraction of publications, and require repayment of grant funds.

The sequence is worth knowing before you ever need it. An allegation goes to the institution's research integrity officer, not to a journal and not to social media. The institution first assesses whether the allegation falls within the definition and has substance. If it does, an inquiry determines whether an investigation is warranted. If it is, a formal investigation examines records, interviews witnesses, and produces a report. The institution sends that report to the federal oversight office, which conducts its own review and may make findings and impose administrative actions ranging from supervised research to debarment from federal funding. Time limits apply, with a general six-year window and defined exceptions.

Two protections run through the whole process. Research records are sequestered early, so that notebooks and files cannot be revised once an allegation is made. And retaliation against a person who makes a good-faith allegation is itself prohibited and separately actionable. Yusuf's correct first step is to write down what he observed, keep copies, and take it to the research integrity officer or an ombuds office. His incorrect first steps are equally clear: confronting his supervisor alone, posting the comparison publicly, or saying nothing and hoping someone else notices.

Key idea: Reporting means handing a documented observation to the person whose job it is to assess it - not investigating yourself, and not accusing anyone of anything.

TypeWhat is doneCore deception
FabricationInventing data or resultsThe work was never done as reported
FalsificationDistorting data, images, or methodsReal work is misrepresented
PlagiarismTaking others' work without creditFalse claim of originality

Why it matters beyond the offender

A single fabricated paper is not a self-contained crime. Other labs waste years and resources building on a result that was never real; meta-analyses are poisoned; in clinical fields, patients can be harmed by treatments endorsed on fraudulent evidence. Because science is cumulative, misconduct contaminates everything downstream. That is why the profession treats FFP not as a private failing but as an assault on the shared enterprise, and why the obligation to report credible suspicions - carefully and through proper channels - is itself an ethical duty.

The scale is measurable. One analysis examined more than two thousand retracted biomedical and life-science articles. Only about a fifth had been withdrawn for honest error. Roughly two-thirds were attributable to misconduct. Fraud or suspected fraud was the largest single category, followed by duplicate publication and plagiarism. Retraction totals have risen sharply since, passing ten thousand in a single year for the first time in 2023, driven substantially by mass retractions of papers linked to organized paper-mill operations. Some of that rise reflects more fraud. Much of it reflects better detection, including image-forensics work and post-publication commenting platforms that did not exist twenty years ago.

Key idea: Retractions are rising partly because misconduct is being found rather than only because it is increasing, and detection now happens after publication as often as before it.

Where people get stuck

Four distinctions do most of the work in this material.

  • Fabrication versus falsification. Ask whether the underlying observation ever existed. Invented numbers are fabrication. Real numbers that were trimmed, spliced, or relabeled are falsification.
  • Misconduct versus questionable practice. Federal misconduct is a closed list of three offenses with an intent requirement. Dropping an inconvenient condition and not mentioning it corrupts the literature but usually falls outside the definition, which is the subject of the next lesson.
  • Plagiarism versus self-plagiarism. The federal definition covers another person's work. Recycling your own text is a real breach of publication ethics, handled by journals and institutions, but it is generally not a federal misconduct finding.
  • Retraction versus guilt. Papers are retracted for honest error too, and a correction issued voluntarily is evidence of integrity. The signal is not that a paper was withdrawn but why.

Key idea: Every misconduct question reduces to two: what exactly was misrepresented, and what did the person know when they did it.

Recap

  • Research misconduct is defined as fabrication, falsification, or plagiarism in proposing, performing, reviewing, or reporting research.
  • The definition excludes honest error and honest differences of opinion, so intent is part of the offense rather than a mitigating factor.
  • Culpability covers intentional, knowing, and reckless conduct, which is how failures of supervision can produce findings.
  • Image manipulation is the most common modern form of falsification, and screening studies find problematic duplications in a few percent of biomedical papers.
  • A finding requires a significant departure from accepted practices of the relevant research community, a culpable state of mind, and proof by a preponderance of the evidence.
  • The process runs assessment, inquiry, investigation, institutional report, and federal oversight review, with record sequestration and protection against retaliation.
  • Self-plagiarism is a publication-ethics violation rather than a federal misconduct category, because the definition speaks of another person's work.
  • Misconduct harms the whole enterprise because science is cumulative, which is why reporting a documented concern through proper channels is a duty rather than a betrayal.

Sources

  1. U.S. Public Health Service. (2024). Requirements for findings of research misconduct, 42 C.F.R. 93.103. Electronic Code of Federal Regulations. ecfr.gov
  2. U.S. Public Health Service. (2024). Research misconduct, 42 C.F.R. 93.234. Electronic Code of Federal Regulations. ecfr.gov
  3. Office of Research Integrity. (n.d.). Handling misconduct. U.S. Department of Health and Human Services. ori.hhs.gov
  4. Fang, F. C., Steen, R. G., & Casadevall, A. (2012). Misconduct accounts for the majority of retracted scientific publications. Proceedings of the National Academy of Sciences, 109(42), 17028-17033. pmc.ncbi.nlm.nih.gov
  5. Bik, E. M., Casadevall, A., & Fang, F. C. (2016). The prevalence of inappropriate image duplication in biomedical research publications. mBio, 7(3), e00809-16. pmc.ncbi.nlm.nih.gov
  6. Van Noorden, R. (2023). More than 10,000 research papers were retracted in 2023 - a new record. Nature, 624, 479-481. nature.com
  7. National Academy of Sciences, National Academy of Engineering, & Institute of Medicine. (2009). Research misconduct. In On being a scientist: A guide to responsible conduct in research (3rd ed.). National Academies Press. ncbi.nlm.nih.gov
Key terms
Research misconduct
Fabrication, falsification, or plagiarism in proposing, performing, reviewing, or reporting research.
Fabrication
Making up data or results and recording or reporting them.
Falsification
Manipulating or omitting data, materials, or processes so the record misrepresents the actual work.
Plagiarism
Appropriating another's ideas, results, or words without appropriate credit.
Preponderance of the evidence
The standard of proof for misconduct: more likely than not that the act occurred.
Whistleblower
A person who reports a good-faith concern about suspected misconduct, entitled to protection from retaliation.

Questionable Research Practices and the Gray Zone

  • Define questionable research practices and explain why they threaten the record despite falling short of misconduct.
  • Identify common QRPs such as p-hacking and HARKing and the practices that prevent them.

Fabrication, falsification, and plagiarism are the bright-line offenses, and they are relatively rare. Far more corrosive to science, in aggregate, is a broad gray zone of questionable research practices (QRPs): choices that are not outright fraud, that many researchers have made without feeling like cheats, but that systematically bias the literature toward false-positive findings. Precisely because they feel normal, QRPs may do more cumulative damage than the dramatic cases of FFP.

Follow one analysis. Marta Kowalczyk has finished a study testing whether a four-week mindfulness program reduces exam anxiety in undergraduates. She has 120 participants, four anxiety measures, two follow-up time points, and a handful of plausible covariates including sleep, prior grades, and baseline anxiety. Her preferred outcome shows nothing. She tries the second measure. Nothing. She tries the second measure at the later time point, controlling for baseline, and gets p equals 0.048. Along the way she noticed five participants who reported no meditation practice at all and excluded them as non-compliant. She now has a publishable result. Nothing she did was invented, and any single step could be defended out loud. The paper she is about to write will nevertheless report a statistic that does not mean what it says.

Key idea: Every individual step in a p-hacked analysis can be defensible; what corrupts the inference is the set of steps taken and not reported.

In plain terms

A p-value is a statement about a procedure, not about a number. It says: if there were no real effect, results this extreme would turn up this rarely, given the analysis I planned. Change the analysis after seeing the data and the sentence stops being true. It is like firing an arrow and then painting the target around it. You did hit the middle. The claim about your aim is simply false. Marta ran roughly twenty defensible analyses and reported one. At that rate, one result crossing the 0.05 line is exactly what chance predicts. She did not lie about any measurement. She just did not tell anyone how many arrows she fired.

What makes a practice questionable

A QRP exploits the flexibility, or researcher degrees of freedom, in how data are collected, analyzed, and reported. In any real analysis there are dozens of defensible choices - which outliers to exclude, which covariates to include, when to stop collecting data, which of several measures to report. If a researcher tries many of these and reports only the combination that "worked," the result looks far more convincing than it is, and the reported statistics no longer mean what they claim to mean.

The size of that flexibility is easy to underestimate. A well-known simulation study examined just four ordinary degrees of freedom - choosing among two dependent measures, adding participants after a first look, including or excluding a covariate, and dropping one of three conditions - and showed that exploiting them together can push the false-positive rate above 60 percent, more than twelve times the nominal 5 percent. The authors demonstrated the point by producing real experimental evidence that listening to a particular song made participants younger, which is impossible, and that was precisely their argument.

Marta's situation is worse than the simulation, not better. Four measures times two time points times two covariate choices times two exclusion rules gives thirty-two analyses she could have run. The probability that at least one crosses 0.05 by chance alone is very high. Her result is not evidence of an effect. It is evidence that she looked in many places.

Key idea: A handful of ordinary analytic choices, used freely and reported selectively, can raise the false-positive rate above 60 percent while every reported number stays technically correct.

A catalog of common QRPs

  • p-hacking. Trying analyses, exclusions, and subgroups until a p-value crosses the 0.05 threshold, then reporting only that path. Because each unreported test inflates the chance of a false positive, p-hacking manufactures significance that would not survive an honest accounting.
  • HARKing (Hypothesizing After the Results are Known). Observing a pattern in the data and then presenting the hypothesis it suggested as if it had been predicted in advance. This disguises an exploratory, hypothesis-generating finding as a confirmatory test, stripping away the skepticism it deserves.
  • Optional stopping. Repeatedly checking results as data accumulate and stopping the moment significance appears, which badly inflates the false-positive rate unless the analysis is designed for it.
  • Selective reporting. Publishing only the measures, conditions, or studies that produced favorable results and quietly dropping the rest - the individual-lab version of publication bias.

Two further practices belong on the list. Outcome switching occurs when a study registers one primary outcome and then publishes a different one as primary because it reached significance. In clinical trials this is common enough that projects exist solely to compare registered protocols with published papers. Undisclosed exclusion rules, Marta's five non-compliant participants, are the everyday version. Excluding them may be entirely correct. Deciding to exclude them after seeing that it helps, and not stating the rule, is not.

Key idea: Ask when the decision was made relative to seeing the outcome - before, and it is design; after and undisclosed, and it is a questionable practice.

How common is this?

Not rare. In a large survey of academic psychologists using an incentive scheme designed to encourage truthful answers, a majority admitted having failed to report all of a study's dependent measures, and more than half admitted having decided whether to collect more data after checking whether existing results were significant. Around a third admitted excluding data after looking at the effect of doing so. In the same survey, admitted outright falsification of data was well under 1 percent. Later surveys in other countries and in industry have found broadly similar patterns.

It is worth asking why so many people do this without feeling like cheats. Part of the answer is that each step has a good cover story. Dropping a participant who clearly did not follow instructions is defensible. Adding twenty more people because the effect looks promising is defensible. Reporting the measure that worked, because the others were noisy, is defensible. Nobody wakes up planning to mislead. They make a series of small local decisions that each point the same way. The bias is in the direction, not in any one step, and you can only see the direction by looking at all of them together. Which is exactly what nobody does, because nobody has to.

Those numbers explain why the gray zone matters more than the bright line. If two in a hundred researchers fabricate and half report selectively, the second behavior touches far more of the literature. It also explains why the reform movement targets incentives and defaults rather than individuals. You cannot audit your way out of a practice that most people consider normal.

Key idea: Self-report surveys consistently find questionable practices in the majority and fabrication in a tiny minority, which is why reform focuses on changing defaults rather than catching offenders.

The remedy: pre-specification and transparency

The common thread of QRPs is that undisclosed flexibility, applied after seeing the data, corrupts the logic of statistical inference. The remedy is to fix the key decisions before seeing the outcomes and to disclose them.

  • Pre-registration records the hypotheses, design, and analysis plan in a time-stamped public archive before data collection, so the line between confirmatory and exploratory work cannot later be blurred.
  • Registered Reports go further: a journal reviews and accepts the plan before results exist, so publication cannot depend on how the results turn out.
  • Full reporting of all measures, conditions, exclusions, and analyses - not just the flattering ones - lets readers judge the evidence honestly.

Notice that none of these reforms forbids exploration. Exploratory, data-driven discovery is vital; it is simply a different kind of claim than a confirmatory test, and it must be labeled as such. The sin in a QRP is not looking at the data - it is looking, deciding, and then pretending you decided in advance. Keeping that distinction honest, in your own work and in what you demand of others, is one of the most consequential integrity habits a modern researcher can build.

Marta's honest paper is not a worse paper. It reports the pre-specified primary outcome, which was null. It reports the other measures and time points in a table so nothing is hidden. It labels the significant result as exploratory and states the number of comparisons behind it. It states the exclusion rule and shows the analysis both ways. And it ends by proposing a direct replication with the promising measure as the pre-registered primary outcome. That paper is publishable, honest, and more useful to the next researcher than the version that buried thirty-one analyses.

Key idea: The honest write-up of a p-hacked analysis is not silence - it is the full set of analyses, an exploratory label, and a proposed confirmatory test.

Where people get stuck

Four confusions recur.

  • Believing exploration is the sin. It is not. Exploratory analysis is how hypotheses are born. The offense is relabeling exploration as confirmation after the fact.
  • Treating QRPs as misconduct. They generally are not, under any federal definition. That is precisely the problem: the behavior most damaging in aggregate sits outside the enforcement system and must be handled by norms, journals, and defaults.
  • Thinking pre-registration locks you in. It does not prevent deviation; it makes deviation visible. You may change your analysis and say so, and readers can then judge.
  • Assuming a large sample fixes it. Sample size addresses power, not flexibility. Twenty analyses of a huge dataset still produce one spurious hit at the usual threshold.

Key idea: Transparency, not restriction, is the remedy - a reader who knows every choice you made can discount appropriately, and a reader who does not cannot.

Recap

  • Questionable research practices are analytic and reporting choices that fall short of fraud but bias the literature toward false positives.
  • They exploit researcher degrees of freedom: which outliers to exclude, which covariates to include, when to stop collecting, which measure to report.
  • Simulation work shows that exploiting four common degrees of freedom can raise the false-positive rate above 60 percent.
  • Common forms include p-hacking, HARKing, optional stopping, selective reporting, outcome switching, and undisclosed exclusion rules.
  • Incentivized surveys find that majorities of researchers admit selective reporting and data-dependent stopping, while admitted fabrication is well under 1 percent.
  • The diagnostic question is whether a decision was made before or after seeing the outcome, and whether it was disclosed.
  • Pre-registration and Registered Reports fix decisions in advance and separate confirmatory from exploratory claims.
  • The honest version of a p-hacked analysis reports every analysis, labels the finding exploratory, and proposes a confirmatory replication.

Sources

  1. John, L. K., Loewenstein, G., & Prelec, D. (2012). Measuring the prevalence of questionable research practices with incentives for truth telling. Psychological Science, 23(5), 524-532. pubmed.ncbi.nlm.nih.gov
  2. Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. pubmed.ncbi.nlm.nih.gov
  3. Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review, 2(3), 196-217. pubmed.ncbi.nlm.nih.gov
  4. Munafo, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, 0021. pmc.ncbi.nlm.nih.gov
  5. Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. pmc.ncbi.nlm.nih.gov
  6. Center for Open Science. (n.d.). Registered Reports. cos.io
  7. Godecharle, S., Fieuws, S., Nemery, B., & Dierickx, K. (2018). Scientists still behaving badly? A survey within industry and universities. Science and Engineering Ethics, 24(6), 1697-1717. pubmed.ncbi.nlm.nih.gov
Key terms
Questionable research practices (QRPs)
Choices short of outright fraud that bias the literature toward false-positive results.
Researcher degrees of freedom
The many defensible analytic and reporting choices that, if exploited selectively, distort inference.
p-hacking
Trying analyses until a result crosses the significance threshold, then reporting only that path.
HARKing
Hypothesizing after results are known and presenting the post hoc hypothesis as if predicted in advance.
Pre-registration
Recording hypotheses and analysis plans publicly before data collection to separate confirmatory from exploratory work.
Registered Report
A publication format in which a journal accepts a study based on its plan, before the results are known.

Authorship and the Allocation of Credit

  • Apply recognized criteria for authorship and distinguish legitimate authors from those who should be acknowledged.
  • Identify dishonest authorship practices and understand contributions-based approaches to credit.

Authorship is the currency of academic science: it confers credit, establishes priority, and drives hiring, promotion, and funding. Because so much rides on it, authorship is also a frequent source of dispute and dishonesty. Ethical authorship rests on a single idea - that credit should track genuine intellectual contribution and be coupled to accountability.

One manuscript will make this concrete. Wei Chen is a fourth-year chemistry doctoral student. She designed the synthesis, ran most of the experiments, and wrote the paper. Three other names are in question. Dario Esposito is the laboratory technician who prepared and ran all 200 samples over eight months and who suggested the change in solvent that made the reaction work. Her supervisor, who conceived the project and revised the manuscript twice, is uncontroversial. And last week her supervisor said, casually, that they should add the department chair, because "he supports the group and it helps at renewal time." The chair has never seen the data. Wei is a year from the job market and does not want a fight.

Key idea: Authorship disputes are rarely arguments about principle - they are arguments in which one party has far more power than the other, and the principle is what the weaker party has instead of leverage.

In plain terms

An author is someone who helped make the ideas and who will answer for the result. That is the whole standard. Two failures follow from it. Adding people who did not do the work makes the credit meaningless. Leaving out people who did makes the record false. Both are common, and neither feels like fraud to the person doing it. The chair's name costs Wei nothing she can point to. That is exactly why it happens. The test that cuts through most cases is simple. If this paper were retracted tomorrow for a data problem, would you expect this person to be answerable for it? If the honest answer is no, they are not an author. They may still deserve thanks, and thanks has its own section.

Criteria for authorship

The most widely cited standard, from the International Committee of Medical Journal Editors (the ICMJE criteria), holds that an author should satisfy all four of the following:

  1. Substantial contribution to the conception or design of the work, or the acquisition, analysis, or interpretation of data;
  2. Drafting the work or revising it critically for important intellectual content;
  3. Final approval of the version to be published;
  4. Agreement to be accountable for all aspects of the work, including the integrity and accuracy of the whole.

The fourth criterion is the moral core: authorship is not only credit but responsibility. Someone whose name appears on a paper is vouching for it. Contributions that do not meet all four criteria - securing funding alone, general supervision of the group, technical assistance, or providing materials - merit acknowledgment, not authorship.

Two features of the standard are routinely misread. First, the criteria are cumulative, so failing any one of them disqualifies. Second, and less obviously, the standard runs in both directions. ICMJE states that everyone who meets the criteria should be identified as an author, and that those who do not should not be listed. A team cannot decide to leave out a qualifying contributor for convenience any more than it can add a non-qualifying one for politics.

There is also an obligation that follows from the first criterion. A person who contributed substantially to acquisition or analysis of data should be given the opportunity to participate in drafting or revising, so that they can meet the remaining criteria. Withholding that opportunity and then citing their failure to draft as grounds for exclusion is a circular argument.

Apply this to Wei's paper. Dario contributed substantially to the acquisition of data and made an intellectual contribution to the method. He therefore has a claim under the first criterion and should be invited to review and approve the manuscript, which would satisfy the rest. The department chair meets none of the four criteria and cannot be made to meet them by generosity. Wei's supervisor meets all four.

Key idea: The criteria exclude free riders and also protect contributors, because a qualifying colleague must be offered the chance to complete the remaining criteria rather than quietly dropped.

Dishonest practices

  • Guest (or honorary) authorship lists someone who did not meet the criteria, often a senior figure added to lend prestige or in a quid pro quo. It dilutes the meaning of authorship and falsely assigns accountability.
  • Gift authorship is closely related: adding a colleague or friend as a favor.
  • Ghost authorship is the inverse and arguably worse: a substantial contributor is omitted, most notoriously when an industry writer drafts a paper that appears under an academic's name alone, hiding a conflict of interest from readers.
  • Coercion authorship occurs when a senior person uses power to demand undeserved credit from a junior one.

These are not rare edge cases. A survey of corresponding authors of nearly 900 articles in six high-impact general medical journals found evidence of honorary or ghost authorship, or both, in about one article in five. Honorary authorship alone appeared in roughly 18 percent of articles and was highest in original research reports. Ghost authorship appeared in about 8 percent. Both figures had declined from a similar survey twelve years earlier, but neither had disappeared. Wei's situation is therefore statistically ordinary, which is precisely the problem.

The reason ghost authorship is treated as the graver offense is worth stating plainly. Guest authorship dilutes credit. Ghost authorship falsifies the provenance of a scientific claim. When a manuscript drafted by a company's medical writer appears under an academic's name alone, readers cannot see who framed the argument, chose the emphasis, or selected which analyses to feature. The conflict of interest disappears from view along with the author.

Key idea: Guest authorship inflates a CV, but ghost authorship hides who actually shaped the claim - which is why it is the more serious deception.

Order, corresponding author, and contributorship

Conventions for author order vary by field, but in many disciplines the first author did the most work and the last author is the senior investigator who directed the project; the corresponding author takes responsibility for communication and for the integrity of the submission. Because these conventions are ambiguous across fields, there is a strong movement toward explicit contributorship statements. A widely adopted taxonomy, CRediT (Contributor Roles Taxonomy), lists specific roles - conceptualization, methodology, formal analysis, writing, and so on - so each person's actual contribution is stated rather than inferred from position.

The variation across fields is larger than newcomers expect, and misreading it causes real damage on hiring committees. In much of biomedicine and the laboratory sciences, first author did the work and last author ran the group. In mathematics, theoretical economics, and parts of particle physics, authors are listed alphabetically and position carries no information at all. In large collaborations, author lists can run to thousands of names and individual contribution is documented separately. A reviewer who applies biomedical conventions to an alphabetical field will systematically undervalue candidates whose surnames begin late in the alphabet.

Contributorship statements are the response. The fourteen roles in the widely used taxonomy - among them conceptualization, methodology, formal analysis, investigation, data curation, supervision, funding acquisition, and the two writing roles - are recorded per author and published with the paper. Two consequences follow. Free riders become visible, because someone must be listed as having done something. And contributions that were previously invisible, such as data curation and software, become citable work.

Key idea: Author position means completely different things in different fields, so contributorship statements exist to replace a convention that does not travel.

PracticeWhat happensWhy it is wrong
Guest / gift authorshipNon-contributor is listedFalse credit and false accountability
Ghost authorshipReal contributor is omittedHides contribution and often a conflict
Coercion authorshipPower extracts undeserved creditExploits a junior researcher

The practical safeguard is to discuss authorship early and openly, ideally when a project begins, and to revisit it as contributions evolve. Most authorship disputes grow from silence and assumption. Naming expectations in advance - who will do what, and what that earns - prevents the great majority of them, and models the fairness you will one day owe your own trainees.

What should Wei actually do? Nothing about her situation requires a confrontation. She can propose, in writing and framed as a journal requirement rather than an accusation, that the team complete the contributorship statement the journal asks for. That single move does most of the work, because someone has to write down what the chair did. She can pair it with a positive proposal: Dario should be an author, since he acquired the data and contributed the solvent change, and he should be sent the draft for review and approval. If her supervisor insists on the chair regardless, she has documented the standard, raised it professionally, and can consult a graduate director or research integrity officer without having accused anyone.

Key idea: The strongest move in an authorship dispute is procedural, not confrontational - ask the team to write down who did what, and the answer usually resolves itself.

Where people get stuck

Four confusions recur, and three of them appear in exam questions.

  • Guest versus ghost. Guest, honorary, or gift authorship adds someone who did not qualify. Ghost authorship omits someone who did. One inflates, the other conceals; the direction is the whole distinction.
  • Funding or supervision as sufficient. Obtaining the grant, running the laboratory, providing a reagent, or collecting data under instruction do not by themselves qualify anyone. They belong in acknowledgments or in a contributorship role that is not authorship.
  • Treating acknowledgment as a consolation prize. It is the correct destination for real contributions that do not meet all four criteria, and journals generally expect written permission from the people named.
  • Waiting until submission. Nearly every dispute traces to a conversation that never happened at the start. Agree the likely author list, the order, and the criteria for change at the project's first meeting, and revisit it when contributions shift.

Key idea: Ask "would this person be answerable if the paper were retracted" and "was this agreed at the start" - those two questions prevent most authorship trouble.

Recap

  • Authorship is credit coupled to accountability, and the widely used ICMJE standard requires all four criteria together.
  • The criteria exclude non-contributors and also protect contributors, who must be given the chance to draft or revise so they can qualify.
  • Guest, honorary, and gift authorship add people who did not qualify; ghost authorship omits people who did, and is the graver deception because it hides provenance and conflicts.
  • A survey of six high-impact medical journals found honorary or ghost authorship in roughly one article in five, so these practices remain common.
  • Funding, general supervision, technical assistance, and provision of materials merit acknowledgment rather than authorship on their own.
  • Author-order conventions differ sharply across disciplines, including fields that list authors alphabetically, so position is not a reliable signal.
  • Contributorship taxonomies record specific roles per author, making free riders visible and previously invisible work citable.
  • Discussing authorship at the start of a project, in writing, prevents the great majority of disputes.

Sources

  1. International Committee of Medical Journal Editors. (n.d.). Defining the role of authors and contributors. icmje.org
  2. International Committee of Medical Journal Editors. (n.d.). Recommendations for the conduct, reporting, editing, and publication of scholarly work in medical journals. icmje.org
  3. Wislar, J. S., Flanagin, A., Fontanarosa, P. B., & DeAngelis, C. D. (2011). Honorary and ghost authorship in high impact biomedical journals: A cross sectional survey. BMJ, 343, d6128. pmc.ncbi.nlm.nih.gov
  4. NISO. (2022). CRediT: Contributor Roles Taxonomy. National Information Standards Organization. credit.niso.org
  5. Committee on Publication Ethics. (2019). COPE core practices. COPE. (COPE materials are hosted behind a bot-protection layer that blocks automated access; consult publicationethics.org ↗ directly.)
  6. ALLEA. (2023). The European code of conduct for research integrity (rev. ed.). All European Academies. allea.org
  7. National Academy of Sciences, National Academy of Engineering, & Institute of Medicine. (2009). On being a scientist: A guide to responsible conduct in research (3rd ed.). National Academies Press. ncbi.nlm.nih.gov
Key terms
ICMJE criteria
Four conditions - substantial contribution, drafting or revising, final approval, and accountability - all required for authorship.
Accountability (authorship)
An author's responsibility to vouch for the integrity and accuracy of the published work.
Guest authorship
Listing a person who did not meet authorship criteria, often for prestige, falsely assigning credit.
Ghost authorship
Omitting a substantial contributor, often to conceal an industry writer's role and a conflict of interest.
Acknowledgment
Recognition for contributions that do not meet all authorship criteria, such as funding or technical help.
CRediT
A contributor-roles taxonomy that states each person's specific contributions to a work.

Conflicts of Interest

  • Define conflicts of interest and distinguish financial from non-financial forms.
  • Explain why disclosure and management, rather than the mere existence of a conflict, are the ethical crux.

A conflict of interest (COI) exists when a secondary interest - money, career advancement, personal loyalty - has the potential to compromise, or appear to compromise, a researcher's judgment about a primary interest such as the validity of the research or the welfare of participants. Note the structure carefully: a conflict is a situation, not an act. Having a conflict is not itself wrong; the ethical questions are whether it is disclosed and how it is managed.

Take a case that is entirely ordinary in modern academic medicine. Ravi Menon is a gastroenterologist who spent nine years developing a stool-based screening test for early colorectal cancer. Four years ago he co-founded a company to commercialize it, and he holds equity in that company along with a share of the patent. He is now the principal investigator on the multi-site trial evaluating whether the test outperforms the current standard. He believes in the test. That belief is why it exists. If the trial reports a positive result, his equity becomes worth a great deal, and if it reports a negative one, it becomes worth very little. Nothing described here is misconduct. All of it needs to be handled.

Key idea: The people with the deepest expertise in a technology are usually the people with a financial stake in it, so conflict management exists to keep that expertise usable rather than to exclude it.

In plain terms

You have a job to do, and something else is pulling at you while you do it. That is a conflict of interest. It does not mean you are corrupt. It means a reasonable person could not tell, from the outside, whether your judgment was clean. Notice what the standard is not. Nobody has to prove you were actually swayed. The reason is that bias mostly does not feel like bias. Menon will not consciously decide to bury a bad result. He will simply find the negative subgroup analysis less convincing, wonder whether a site ran the assay properly, and want to check one more thing before reporting. That is how a financial stake works. It shifts the burden of proof inside your own head, and you cannot detect it by introspection. Which is exactly why disclosure exists.

Actual, potential, and perceived

Conflicts come in three shades, and all three matter. An actual conflict is currently compromising judgment; a potential conflict could do so if circumstances change; a perceived conflict is one that a reasonable observer might suspect, whether or not judgment is in fact affected. Perception matters because science runs on trust: even the appearance of a compromised judgment can undermine confidence in a result, which is why disclosure standards deliberately reach beyond conflicts that have already caused harm.

A fourth distinction belongs beside these three, and it is a classic examination item. A conflict of commitment concerns time and effort rather than judgment: an investigator whose consulting days, company board seat, or outside laboratory competes with the obligations owed to the employing institution. It is possible to have a conflict of commitment with no financial conflict, as when a researcher spends most of the week on an unpaid advocacy role. It is equally possible to have a financial conflict with no conflict of commitment, as when someone holds shares in a company that requires none of their time. The two travel together often enough that people confuse them, and institutions handle them under different policies.

Key idea: Conflict of interest is about whether your judgment can be trusted; conflict of commitment is about where your working hours go, and an institution can have a problem with either alone.

Financial and non-financial forms

  • Financial conflicts are the most scrutinized: research funding from a company whose product is being tested, personal consulting fees or speaker honoraria, equity or stock options, patents and royalties. The concern is well founded - studies funded by an interested sponsor are, on average, more likely to report results favorable to that sponsor, an effect sometimes called the funding effect.
  • Non-financial conflicts are subtler but real: the drive to confirm one's own prior theory, the desire to help a friend or harm a rival, ideological commitment, or the pressure of one's own reputation riding on a particular outcome. These are harder to disclose on a form but can bias judgment just as powerfully.

The evidence behind the funding effect is unusually strong, because it has been studied systematically rather than anecdotally. A large methodology review pooled studies across drug and device research. It found two things. Industry-sponsored trials more often reported efficacy results and conclusions that favored the sponsor. And the agreement between a trial's own results and its stated conclusions was weaker in sponsored work.

The next part is the surprise. The difference was not explained by the usual risk-of-bias measures, such as randomization and blinding. Sponsored trials are often built well. So where does the bias enter? It enters in the choices around the numbers. Which drug do you compare yours against, and at what dose? Which outcome do you make primary? Which subgroup do you feature? And what conclusion do you draw from a modest result? None of those steps is a lie. All of them lean.

Key idea: The funding effect is not sloppy methods but favorable framing - choice of comparator, choice of dose, choice of outcome, and the gap between results and the conclusion drawn from them.

Disclosure and management

The governing principle is not prohibition but transparency plus management. Researchers disclose relevant interests to their institution, to journals when they publish, and often to participants during consent. The institution or journal then decides how to manage the conflict along a graduated scale:

  1. Disclose the interest so readers and oversight bodies can weigh it.
  2. Manage it - for example, by adding independent oversight, or by having someone without the conflict analyze the data or consent the participants.
  3. Prohibit the arrangement or recuse the conflicted person entirely when the conflict is severe and unmanageable, such as a researcher who stands to profit directly evaluating their own product in participants at risk.

In the United States, federally funded research operates under a specific regulatory scheme rather than general good intentions. Under the rules promoting objectivity in research, investigators must disclose significant financial interests related to their institutional responsibilities. The thresholds are concrete. Payments and equity in a publicly traded company count once they pass a set dollar amount over the preceding year. Any equity at all in a private company counts, with no threshold. So does income from intellectual property rights. The institution, not the investigator, then decides whether a disclosed interest is a conflict for a specific project. Where it is, the institution must report it to the funding agency and put a written management plan in place. Training is required periodically.

Menon's management plan would draw on a standard toolkit. An independent data and safety monitoring board holds the unblinding key. Statistical analysis is performed by a team with no equity, working from a pre-specified analysis plan. He does not personally recruit or consent participants, and his interest is disclosed in the consent form so participants can weigh it. He does not control the decision to publish. And his interest appears in the paper, in the trial registration, and in any talk he gives. Notice that none of this bars him from the study. It moves each decision that money could touch into hands that money does not touch.

Key idea: A good management plan does not test the conflicted person's virtue - it relocates the specific decisions the conflict could reach.

Institutional and reviewer conflicts

Two further layers matter and are frequently overlooked. Institutional conflicts arise when the organization itself has a stake. Picture a university that holds equity in a spinout company. The spinout's device is being trialed in the university's own hospital. And the board that reviews the trial reports to the university's own leadership. No individual has done anything wrong, and the whole structure still leans one way. The remedies are structural. Use an external review board. Build a firewall between the technology-transfer office and the oversight committee. And disclose the institutional interest in the consent form.

Reviewer and editorial conflicts are the ones most researchers will encounter personally. Journals ask reviewers to decline where they compete directly, collaborate closely, share an institution, or have a personal stake in the outcome. Editors face the same duties with more power. The standard here is deliberately conservative, because a reviewer holds an unpublished manuscript and the discretion to delay it.

Key idea: Conflicts scale up to institutions and sideways to reviewers and editors, and in both cases the fix is structural separation rather than a promise of impartiality.

Two settings deserve special vigilance. In clinical research, a financial stake in a positive outcome can distort the risk-benefit judgment owed to patients, so the strictest management applies. In peer review and editing, a reviewer who competes with or collaborates closely with the authors has a conflict and should recuse rather than exploit privileged access to an unpublished manuscript. The uniting lesson is disarmingly simple: you cannot always avoid having interests, but you can always be transparent about them. Concealment converts an ordinary, manageable conflict into misconduct.

Where people get stuck

Four confusions cause most of the trouble.

  • Treating a conflict as an accusation. Disclosing an interest is a routine professional act, not a confession. Researchers who feel accused tend to under-disclose, which is the outcome the system most wants to avoid.
  • Confusing interest with commitment. Interest concerns judgment; commitment concerns time and effort. Each can exist without the other, and institutions govern them under separate policies.
  • Believing disclosure is sufficient. Disclosure is the first step of three. Where the conflict is severe, management or recusal is required, and a conflicted investigator who discloses and then controls the analysis has satisfied only the easiest requirement.
  • Ignoring non-financial conflicts because no form asks. Attachment to your own prior theory, rivalry, and reputational stakes bias judgment without appearing on any disclosure form. The remedy is the same as for financial conflicts: independent analysis, blinding, and pre-specification.

Key idea: Ask what a reasonable outsider would want to know and who is making each decision - those two questions cover nearly every conflict scenario you will meet.

Recap

  • A conflict of interest is a situation in which a secondary interest could compromise, or appear to compromise, judgment about a primary interest.
  • Conflicts come in actual, potential, and perceived forms, and disclosure requirements deliberately reach the merely perceived because trust is the asset at risk.
  • A conflict of commitment concerns time and effort rather than judgment, and it can exist entirely apart from any financial interest.
  • Non-financial conflicts, including theory attachment and rivalry, bias judgment without appearing on disclosure forms.
  • Systematic reviews find that industry-sponsored trials more often report sponsor-favorable results and conclusions, and that the effect is not explained by standard risk-of-bias measures.
  • U.S. rules require disclosure of significant financial interests to the institution, which then determines whether a project-specific conflict exists and imposes a management plan.
  • Management relocates decisions rather than testing virtue: independent monitoring, independent analysis, someone else consenting participants, and disclosure to participants and readers.
  • Institutions and peer reviewers hold conflicts too, and the remedies there are structural separation and recusal.

Sources

  1. U.S. Public Health Service. (2011). Promoting objectivity in research, 42 C.F.R. Part 50, Subpart F. Electronic Code of Federal Regulations. ecfr.gov
  2. U.S. Public Health Service. (2011). Definitions, 42 C.F.R. 50.603. Electronic Code of Federal Regulations. ecfr.gov
  3. National Institutes of Health. (n.d.). Financial conflict of interest. grants.nih.gov
  4. Lundh, A., Lexchin, J., Mintzes, B., Schroll, J. B., & Bero, L. (2017). Industry sponsorship and research outcome. Cochrane Database of Systematic Reviews, 2(2), MR000033. pmc.ncbi.nlm.nih.gov
  5. International Committee of Medical Journal Editors. (n.d.). Author responsibilities: Disclosure of financial and non-financial relationships and activities. icmje.org
  6. International Committee of Medical Journal Editors. (n.d.). Disclosure of interest. icmje.org
  7. Centers for Medicare & Medicaid Services. (n.d.). Open Payments data. openpaymentsdata.cms.gov
Key terms
Conflict of interest (COI)
A situation in which a secondary interest could compromise, or appear to compromise, judgment about a primary interest.
Perceived conflict
A conflict a reasonable observer might suspect, which erodes trust even absent actual bias.
Funding effect
The tendency for industry-sponsored studies, on average, to report results favorable to the sponsor.
Non-financial conflict
A bias from theory-confirmation, loyalty, ideology, or reputation rather than money.
Disclosure
The reporting of relevant interests to institutions, journals, and sometimes participants.
Recusal
Withdrawing from a decision or review because of an unmanageable conflict of interest.

Module 5: Data Stewardship and Open Science

Managing and sharing research data, and the reproducibility reforms reshaping how science is done and reported.

Data Management and Responsible Sharing

  • Explain the components of a data management plan and the FAIR principles.
  • Balance the duty to share data against obligations of confidentiality and consent.

Data are the evidentiary foundation of every empirical claim, and how they are recorded, stored, and shared is an ethical matter, not merely a technical one. Sloppy data management can make honest work irreproducible and indistinguishable from fraud; good stewardship protects both the integrity of the record and the people the data describe. Increasingly, funders and journals require researchers to plan for data before collecting any.

Ingrid Halvorsen is finishing a twelve-year birth-cohort study. She has 3,400 mother-child pairs, annual questionnaires, cord-blood genotypes, school records obtained under agreement with two districts, and 180 hours of interview recordings with a subset of families. Her funder now requires her to deposit the data. Some of it can go into an open repository this month. Some of it can only ever be released under a data-use agreement to approved researchers. Some of it, specifically the interview recordings, probably cannot be shared at all, because participants were promised the recordings would be heard only by the study team. Her problem is not whether to share. It is that she made different promises about different parts of one dataset, twelve years ago, and now has to honor all of them.

Key idea: What you may share later is fixed by what you promised at consent, which means data sharing is a decision you make years before the data exist.

In plain terms

Two questions govern everything here. Can someone else understand this dataset without me in the room? And am I allowed to give it to them? The first is about documentation. A file of numbers with column headers like q17b is worthless to a stranger and, within about a year, worthless to you as well. The second is about permission. It combines what participants agreed to, what the law allows, and what any third party who supplied data will accept. Most researchers worry only about the second and then discover, when someone finally asks for the data, that they cannot answer the first. Both are solved the same way: by writing things down while you still remember them.

The data management plan

A data management plan (DMP) is a document, usually written at the proposal stage, that specifies how data will be handled across its life cycle. A serviceable DMP addresses:

  • What data will be generated, in what formats, and how much.
  • Documentation and metadata - the contextual information (variable definitions, collection procedures, units) without which a dataset is uninterpretable to anyone else, and often to the researcher a year later.
  • Storage, backup, and security during the project, including protection of any identifiable data.
  • Retention - how long data are kept after the project (institutions often require several years to allow verification) and how they are eventually destroyed or archived.
  • Sharing and access - whether, how, and under what conditions others may obtain the data.

Funder expectations have hardened considerably. Since 2023, the largest U.S. biomedical funder has required a data management and sharing plan with every applicable application. The plan must cover six things. What data will be produced. What tools and code a reader needs to interpret them. Which standards will be used. Where and when the data will be preserved and made accessible. What factors limit access. And who is responsible for compliance. The default timing is strict. Data must be available no later than the associated publication or the end of the award, whichever comes first. Sharing at the end of a career is no longer an option.

Documentation deserves more attention than it gets. A usable deposit needs five things. A codebook defining every variable and its permitted values. A note on how missing data are coded. The questionnaires or instruments as they were actually administered. A written account of any derived variables, with the code that produced them. And a README explaining the file structure.

Halvorsen's cohort shows why this is not busywork. It has twelve waves and several instrument changes. The sleep item was reworded in wave 7. Nobody looking at the file could tell. A future analyst would compare wave 6 with wave 8 and read a change in wording as a change in sleep. One line in a codebook prevents that. Nothing else does.

Key idea: Metadata is not administrative overhead - it is the difference between a dataset that can be reused correctly and one that will be reused wrongly.

The FAIR principles

The modern standard for good data stewardship is captured by the acronym FAIR, which holds that data should be:

LetterPrincipleIn practice
FFindableDeposited with a persistent identifier and rich metadata
AAccessibleRetrievable by a clear protocol, even if access is controlled
IInteroperableIn standard formats and vocabularies others' tools can use
RReusableWell documented and clearly licensed for reuse

A subtle but important point: FAIR is not the same as "open." Accessible means retrievable under a defined, transparent protocol - which may legitimately include authentication and approval. Sensitive human data can be FAIR while remaining controlled-access, released only to qualified researchers under a data-use agreement. The slogan of responsible sharing is therefore "as open as possible, as closed as necessary."

Findability has a concrete requirement that is easy to satisfy and often skipped: a persistent identifier. Deposit your data in a repository that mints a digital object identifier. The dataset then has a stable address. It survives a broken laboratory website and a move to a new institution. It can also be cited like a paper. That last point matters for incentives. A dataset with its own identifier collects citations, and citations are the currency in which effort is repaid.

A second framework has grown up beside FAIR rather than inside it. The CARE principles for Indigenous data governance are collective benefit, authority to control, responsibility, and ethics. They address a gap that FAIR leaves open. Data can be perfectly findable and reusable while the community it describes has no say in how it is used. Halvorsen's cohort includes families from a Native community that negotiated its own agreement at enrolment, requiring community review before any publication using their data. That agreement binds her regardless of what the funder's default policy says.

Key idea: FAIR governs whether data can be used; CARE asks who decides and who benefits, and the two frameworks answer different questions.

Sharing versus protecting

The duty to share collides, at times, with duties of confidentiality and the limits of consent. Three guardrails keep sharing ethical.

  1. Consent must anticipate sharing. If you intend to deposit data for reuse, the consent form should say so; you cannot ethically release data in ways participants never agreed to.
  2. De-identify before sharing human data, attending to quasi-identifiers, and use controlled access for data that cannot be safely anonymized.
  3. Honor legitimate constraints - genuine privacy risks, Indigenous data-sovereignty rights, and certain proprietary or security limits are valid reasons to restrict, though never a cover for hiding inconvenient results.

Done well, data sharing multiplies the value of research: others can verify findings, ask new questions of existing data, and avoid duplicating costly collection. The ethical craft is to maximize that public good while keeping faith with the individuals whose lives the data represent.

Retention is the part everyone forgets until an auditor asks. Two rules run in parallel. Funders and institutions set a floor for how long records must be kept after an award closes, commonly around three years from the final report, with longer periods where a dispute or audit is open. Journals and professional bodies often expect the underlying data to be available for longer, since a question about a figure can surface a decade later. The practical answer is to keep the analytic dataset and code indefinitely in a repository, and to keep the identifiable material only as long as the consent and the protocol allow. Those are different clocks, and a good plan states both.

Halvorsen's dataset therefore splits into tiers, and that is the normal outcome rather than a failure. The questionnaire data, aggregated and de-identified, goes to an open repository with a digital object identifier and a full codebook. The genotype data goes to a controlled-access archive, released to approved investigators under a data-use agreement that forbids re-identification attempts and onward transfer. The school records stay closed, because the districts' agreements do not permit redistribution, and she says so publicly with a contact route for anyone wanting to negotiate their own agreement. The interview recordings are not shared, though anonymized transcripts with identifying details removed may be, if the original consent supports it. Every tier is documented in the plan and stated in the paper.

Key idea: The right answer is almost never all-open or all-closed - it is a tiered release with the tier and its reason stated publicly.

Where people get stuck

Four errors turn a good intention into an unusable deposit.

  • Writing consent that forecloses sharing. A form promising that data will be used "only for this study" and then destroyed makes later deposit impossible, however valuable. Write consent that anticipates de-identified sharing, and say what will and will not be shared.
  • Confusing FAIR with open. Controlled-access data can be fully FAIR. The Accessible principle requires a clear retrieval protocol, not the absence of one.
  • Depositing data without documentation. A CSV with no codebook satisfies a funder's checkbox and helps nobody. If a competent stranger cannot reproduce your main table from the deposit, the deposit has failed.
  • Using privacy as a shield. Genuine privacy risk, third-party agreements, and community governance are legitimate limits. "The data are too sensitive" offered without specifics, when a controlled-access route exists, is not a reason but an excuse.

Key idea: Ask whether a competent stranger could reproduce your headline result from what you deposited - if not, you have archived files rather than shared data.

Recap

  • Data stewardship is an ethical duty, because poor records make honest work indistinguishable from fraud and expose the people the data describe.
  • A data management plan specifies what data will exist, how they will be documented, where they will be stored and secured, how long they are retained, and how they may be shared.
  • Major funders now require sharing plans and set default timing at publication or the end of the award, whichever comes first.
  • FAIR means Findable, Accessible, Interoperable, and Reusable, and Accessible means a clear protocol rather than unrestricted release.
  • Persistent identifiers make datasets citable, which is how the effort of documentation gets repaid.
  • The CARE principles address governance and benefit for Indigenous data, answering questions that FAIR does not ask.
  • Consent must anticipate sharing, data must be de-identified with attention to quasi-identifiers, and legitimate constraints must be honored without becoming a cover.
  • Tiered release - open, controlled, and closed components within one project - is the normal and defensible outcome.

Sources

  1. National Institutes of Health. (2020). Final NIH policy for data management and sharing (NOT-OD-21-013). grants.nih.gov
  2. National Institutes of Health. (n.d.). Data management and sharing policy. grants.nih.gov
  3. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., et al. (2016). The FAIR guiding principles for scientific data management and stewardship. Scientific Data, 3, 160018. nature.com
  4. GO FAIR Initiative. (n.d.). FAIR principles. go-fair.org
  5. Global Indigenous Data Alliance. (n.d.). The CARE principles for Indigenous data governance. gida-global.org
  6. DataCite. (n.d.). Connecting research, advancing knowledge. datacite.org
  7. Office of Management and Budget. (n.d.). Record retention requirements, 2 C.F.R. 200.334. Electronic Code of Federal Regulations. ecfr.gov
Key terms
Data management plan (DMP)
A document specifying how data will be generated, documented, stored, retained, and shared across its life cycle.
Metadata
Contextual documentation (variable definitions, procedures, units) that makes a dataset interpretable and reusable.
FAIR principles
The standard that data be Findable, Accessible, Interoperable, and Reusable.
Controlled access
Release of sensitive data only to qualified users under defined conditions such as a data-use agreement.
Data retention
The period for which data are preserved after a project to allow verification, before archiving or destruction.
Data-use agreement
A contract governing how shared, often sensitive, data may be accessed and used.

The Reproducibility Crisis and Open Science

  • Distinguish reproducibility from replicability and describe the evidence of a replication crisis.
  • Explain the open-science reforms designed to strengthen the reliability of published findings.

Over the past two decades, science has confronted an uncomfortable discovery about itself: a substantial fraction of published findings, across fields from psychology to cancer biology, do not hold up when others try to reproduce them. This is the reproducibility (or replication) crisis, and the reform movement it triggered - open science - is reshaping research practice. Understanding both is now part of basic research literacy.

Nina Petrov started her doctorate expecting to extend a famous finding, not to interrogate it. Her supervisor gave her a well-cited effect from the literature and told her to build on it. She ran it with 40 participants and found nothing. She assumed she had made a mistake, checked her materials, and ran it again with 80. Nothing. She ran it a third time with 200 participants, a pre-registered protocol, and the original authors' own stimuli. Still nothing. Eighteen months into her degree she has three null results, no publication, and a persistent private worry that she is simply bad at research. Nina is not bad at research. She has run headlong into a structural problem, and understanding it is the difference between quitting and doing something valuable.

Key idea: A failed replication feels like personal incompetence to the person running it, which is one reason the problem stayed invisible for so long.

In plain terms

Two different questions hide behind one word. Can someone rerun your analysis on your data and get your numbers? Can someone run your study again on new people and see the same effect? The first is a bookkeeping question. The second is the scientific one. Large coordinated projects have now asked the second question systematically, and the answers were uncomfortable across several fields. That does not mean most published science is fake. It means the system rewarded finding things and barely rewarded checking them, so nobody checked. Small studies produce noisy results. Journals published the exciting ones. Nobody published the quiet ones. Repeat for thirty years and you get a literature with more positive findings in it than the world actually contains.

Two terms that are often confused

Precision matters here.

  • Reproducibility, in the strict sense, means obtaining the same results from the same data and analysis. If a colleague runs your code on your dataset and gets different numbers, the work is not reproducible - a failure of computational transparency.
  • Replicability means obtaining consistent results from new data collected in a new study of the same question. Replication tests whether a finding is real and general, not merely whether the original arithmetic was correct.

A finding can be perfectly reproducible (the analysis is faithfully repeatable) yet fail to replicate (a fresh study finds no effect), which typically signals that the original result was a false positive.

A U.S. national academies report formalized this split. Reproducibility, in that report's usage, means obtaining consistent results using the same input data, computational steps, methods, and code as the original study. Replicability means obtaining consistent results across studies that address the same question, each collecting its own data. Fields differ in vocabulary, and some use the two words in the opposite sense, so a careful writer defines the terms before using them. What matters is the underlying distinction: same data and code, versus new data.

Key idea: Same data and same code is a bookkeeping test; new data is the scientific test, and only the second tells you whether the effect is real.

The evidence and its causes

Large coordinated efforts have attempted to replicate published studies and found that a worrying share do not replicate, or replicate with markedly smaller effects than originally reported. The causes are, by now, familiar from earlier lessons and reinforce one another:

  1. Publication bias - journals favor novel, positive results, so the literature over-represents false positives and hides null findings in the "file drawer."
  2. Questionable research practices - p-hacking, HARKing, and optional stopping inflate the rate of spurious significant results.
  3. Underpowered studies - small samples yield unstable estimates and, when they do reach significance, often exaggerate the true effect size.
  4. Insufficient transparency - when data, materials, and code are unavailable, errors cannot be caught and replications cannot be attempted.

The numbers are worth carrying in your head. A coordinated project repeated 100 experimental and correlational studies drawn from three psychology journals, using original materials and, where possible, the original authors' input. Ninety-seven percent of the original studies had reported statistically significant results. Thirty-six percent of the replications did. Mean replication effect sizes were roughly half the originals. Separately, a survey of more than 1,500 researchers across disciplines found that over 70 percent had tried and failed to reproduce another scientist's experiment, and more than half had failed to reproduce one of their own.

It helps to see how publication bias builds a false picture. Imagine twenty labs test the same idea, and the idea is wrong. Nineteen find nothing. One finds a positive result by chance. The nineteen null studies go in a drawer, because journals did not want them and nobody was rewarded for writing them up. The one positive study gets published. A reader five years later sees a literature with one paper in it, and that paper says the effect is real. Nothing dishonest happened at any step. The record is still false.

The mechanism behind the halved effect sizes has a name. In a low-powered study, an effect can only reach significance if the noise happens to run in your favor, so the published estimate is systematically too large. This is the winner's curse, and it explains why a well-powered replication of a real effect can still look like a failure: the original number was never achievable. Nina's third attempt, with 200 participants, may well be measuring the true effect accurately. The true effect is simply much smaller than the paper she was given.

Key idea: Underpowered studies do not merely miss real effects - when they do reach significance they overstate the size, so honest replications routinely look like failures.

The open-science response

Open science is the umbrella term for practices that make research more transparent, and therefore more self-correcting, at every stage.

PracticeProblem it targets
Pre-registration and Registered Reportsp-hacking, HARKing, publication bias
Open data and open materialsReproducibility, error detection
Open-source analysis codeComputational reproducibility
Preprints and open accessSlow, restricted dissemination
Registered replication and multi-lab studiesTesting whether effects are real and general

Two levers have proved more powerful than exhortation. The first is journal policy. A widely adopted framework sets graded standards across eight dimensions - citation, data transparency, analytic-code transparency, materials transparency, design and analysis reporting, study pre-registration, analysis-plan pre-registration, and replication - with three levels of stringency at each. A journal picks a level and enforces it at submission, which changes behavior far faster than asking authors to be virtuous. The second is the Registered Report format, where acceptance depends on the question and the design rather than the result. Studies published this way report null findings at dramatically higher rates than the conventional literature, which is exactly what you would expect if publication bias had been doing the work all along.

Key idea: Reform works by changing what journals require and when they decide, because incentives shape behavior more reliably than exhortation does.

Two clarifications guard against overreaction. First, a failed replication does not by itself prove the original was fraudulent or even wrong; effects can be real but fragile, context-dependent, or smaller than first estimated. The crisis is less about lies than about a system that rewarded novelty over reliability.

Second, replication is not a threat to be resented but the mechanism by which science earns its authority; a field that welcomes replication and shares its materials is not admitting weakness but demonstrating strength. The trajectory of reform is toward valuing credible findings over merely novel ones - a cultural shift that individual researchers advance every time they pre-register a study, share their data, or attempt a replication.

So what should Nina do? Her three null results are not a failed project; they are a finding, and the modern literature has a place for it. She can write up all three attempts as a pre-registered replication series, report the effect size with a confidence interval rather than a verdict, and submit it as a Registered Report or to a journal that publishes replications. She should contact the original authors and share her materials, since procedural differences are a real and interesting possibility. And she should state plainly what her data can and cannot support: not that the effect does not exist, but that if it exists it is much smaller than the original estimate.

Key idea: A well-powered null result is publishable evidence about the size of an effect, not a failure to find one - and reporting it that way is the whole reform in miniature.

Where people get stuck

Four confusions recur in this material.

  • Reading a failed replication as an accusation. Fragile, context-dependent, and overestimated effects all produce failed replications without anyone behaving badly. Treating replication as an attack is what makes fields defensive rather than self-correcting.
  • Confusing the two terms. Same data and code is reproducibility in the national-academies usage; new data is replicability. Because usage varies by field, define your terms in writing.
  • Believing bigger samples alone will fix it. Power helps enormously, but a huge underpowered-in-disguise study with undisclosed analytic flexibility still generates false positives. Power and pre-specification are complements.
  • Treating open science as a burden added to research. Pre-registration, shared code, and shared materials mostly relocate work you would have done anyway to an earlier point, where it is cheaper and where it also protects you.

Key idea: The crisis was less a story about liars than about a system that rewarded novelty and never paid anyone to check.

Recap

  • Reproducibility means the same results from the same data and code; replicability means consistent results from new data addressing the same question.
  • A large coordinated replication project found that 97 percent of original psychology studies reported significant results while 36 percent of replications did, with effect sizes roughly halved.
  • A cross-disciplinary survey found that over 70 percent of researchers had failed to reproduce someone else's experiment and more than half had failed to reproduce their own.
  • The main causes are publication bias, questionable research practices, underpowered designs, and insufficient transparency, and they reinforce one another.
  • The winner's curse means underpowered studies that do reach significance systematically overstate effect size, so honest replications often look like failures.
  • Open-science practices target specific failures: pre-registration for flexibility, open data and code for verification, replication for generality.
  • Graded journal transparency standards and the Registered Report format change incentives rather than relying on exhortation.
  • A failed replication is evidence about the size of an effect, not proof of fraud, and publishing it is a contribution rather than a defeat.

Sources

  1. Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. pubmed.ncbi.nlm.nih.gov
  2. Baker, M. (2016). 1,500 scientists lift the lid on reproducibility. Nature, 533, 452-454. nature.com
  3. National Academies of Sciences, Engineering, and Medicine. (2019). Understanding reproducibility and replicability. In Reproducibility and replicability in science. National Academies Press. ncbi.nlm.nih.gov
  4. Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. pmc.ncbi.nlm.nih.gov
  5. Munafo, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, 0021. pmc.ncbi.nlm.nih.gov
  6. Center for Open Science. (n.d.). TOP guidelines. cos.io
  7. Center for Open Science. (n.d.). Preregistration. cos.io
Key terms
Reproducibility
Obtaining the same results from the same data and analysis; a test of computational transparency.
Replicability
Obtaining consistent results from new data in a new study of the same question.
Replication crisis
The finding that a substantial fraction of published results fail to replicate in fresh studies.
File drawer problem
The hiding of null or non-significant results, biasing the published literature toward positive findings.
Open science
Practices that make research transparent and self-correcting, from pre-registration to open data and code.
Preprint
A manuscript shared publicly before formal peer review to speed and open dissemination.

Module 6: The Research Community

The relational duties that sustain a healthy research culture: mentoring, peer review, and collective integrity.

The Ethics of Peer Review and Publication

  • Explain the purpose of peer review and the duties a reviewer owes to authors, editors, and the field.
  • Identify unethical publication practices and the reviewer's obligations of confidentiality and impartiality.

Peer review is the quality-control mechanism at the heart of scholarly publishing: before a manuscript becomes part of the record, independent experts evaluate its rigor, novelty, and validity. Because reviewers wield real power over careers and over what counts as knowledge, the role carries strict ethical duties. Most researchers will spend far more hours reviewing others' work than they expect, so learning to do it honorably is a core professional skill.

An email arrives for Farida Nasser, an associate professor of ecology. An editor is inviting her to review a manuscript on drought tolerance in a grass species. She opens the abstract and her stomach drops. The study is uncomfortably close to work her own laboratory has been running for two years and has not yet submitted. The methods are different, the site is different, but the central claim is the one her postdoc is about to write up. Nasser now holds a document nobody else outside the journal has seen, describing results that could scoop her group by six months. Every duty in this lesson is live in that inbox, and the first decision is the one she makes in the next ten minutes.

Key idea: Peer review hands you privileged access to unpublished work by people who may be your competitors, and the entire ethics of the role follows from that fact.

In plain terms

You are being lent something valuable and secret. Three rules follow. Do not keep it. Do not use it. Do not judge it if you cannot judge it fairly. That is nearly the whole job. The hard cases are not about whether to steal an idea, because almost nobody decides to do that. They are about the blurry things. Can you mention the finding to your student? Can you start the experiment it suggested, since you were going to do it anyway? Can you take three months because you are busy, knowing the delay helps you? Each of those feels defensible from the inside. The test that cuts through them is uncomfortable and reliable. Would you be content for the authors to see exactly what you did and when? If not, do not do it.

What peer review is for

Review serves the field, not the reviewer. Its aims are to help editors decide what to publish, to catch errors and overstatement before they mislead others, and to improve a manuscript through constructive critique. It is not a gate for enforcing personal preferences, settling scores, or delaying a competitor. A good review is rigorous but fair: it judges the work against appropriate standards, distinguishes fatal flaws from matters of taste, and offers criticism the authors can act on.

Core duties of a reviewer

  • Confidentiality. A manuscript under review is a privileged, unpublished document. A reviewer must not share it, discuss it beyond what the journal permits, or - most seriously - use its ideas, data, or results in their own work before publication. Doing so is a form of theft of unpublished intellectual property.
  • Competence and diligence. Accept only manuscripts you are qualified to judge, and evaluate them carefully within the agreed time. Declining a review you cannot do well is itself responsible conduct.
  • Impartiality and disclosure of conflicts. A reviewer who competes closely with, collaborates with, or has a personal stake against the authors has a conflict and should recuse, or at minimum disclose it to the editor rather than exploit privileged access.
  • Constructive, evidence-based critique. Ground criticisms in specifics, remain civil, and separate the quality of the work from the identity or reputation of its authors.

Three practical questions arise constantly and have settled answers. May you involve a student? Yes, but only with the editor's agreement, with the trainee bound by the same confidentiality, and with the trainee named in your report so the editor knows who read the manuscript. Co-reviewing is a legitimate and valuable form of training; doing it secretly is a confidentiality breach. May you upload the manuscript to a generative artificial-intelligence service to help draft your comments? Generally no, because doing so transmits confidential unpublished work to a third party, and journals increasingly prohibit it explicitly. May you cite the manuscript's finding in your own paper before it appears? No. It is not public, and you learned it only because you were trusted.

Nasser's correct action is to decline the review and tell the editor why, in one sentence: her group has closely related unpublished work. That single email protects her, the authors, and the journal. It also preserves her ability to publish her own study without any later suggestion that she saw theirs first. Note the second-order point. If she keeps the review, then even a scrupulous review leaves her exposed, because the appearance of a conflict is itself damaging. Declining costs her nothing.

Key idea: Declining a review is cheap and protects everyone, so when in doubt about a conflict the answer is almost always to decline and say why.

Writing a review that is worth reading

Ethics and craft converge here. A useful review separates three tiers of comment. Fatal problems are claims the data cannot support, a design that cannot answer the question, or an analysis that is wrong. Substantive fixes are missing controls, unreported analyses, overstated language, and absent limitations. Preferences are the things you would have done differently and the citations to your own work. Reviewers who present tier three as tier one waste everyone's time and damage careers. Reviewers who never distinguish them force the editor to do their job.

Tone is not a courtesy, it is a function. "The authors appear unfamiliar with the field" tells nobody anything actionable. "The analysis in Table 2 does not account for spatial autocorrelation, which the sampling design makes likely; a mixed model with site as a random effect would test this" is the same objection made useful. A good rule: write every comment as though the authors will read it, the editor will weigh it, and your own name will eventually be attached to it, because under open review models it increasingly will be.

Key idea: Sort your comments into fatal, substantive, and preference, and say which is which - that single habit makes a review honest and useful at the same time.

Unethical publication practices

Beyond the review itself, the publication system has recognized abuses that every researcher should be able to name.

PracticeWhat it is
Duplicate publicationPublishing the same work in more than one venue as if each were new
Salami slicingSplitting one study into several thin papers to inflate a publication count
Redundant self-plagiarismReusing one's own text or data without disclosure, deceiving readers
Predatory publishingJournals that charge fees but provide no genuine peer review or quality control

Each of these corrupts the record in a different way - by double-counting, by fragmenting knowledge, by hiding reuse, or by lending a veneer of legitimacy to unvetted claims. The editor, too, has duties: to judge manuscripts on merit, to manage conflicts, to protect confidentiality, and to correct the record through corrections or retractions when serious errors emerge.

Two of these deserve more detail. Predatory publishing is not simply a matter of low quality. The defining feature is charging article-processing fees while providing no genuine editorial or review service, often accompanied by fake editorial boards, impossibly short review times, and aggressive solicitation emails. Research linking reviewer records to journal lists has found that people who review for such venues tend to be earlier in their careers and less experienced, which suggests the problem is partly one of training rather than of bad faith. The practical defenses are unglamorous: check whether the journal is indexed where your field's journals are indexed, look up the editorial board members and see whether they know they are listed, and ask a librarian.

The newer and larger problem is the paper mill: an organized operation that sells authorship slots on fabricated or plagiarized manuscripts, often produced at scale with templated text and manipulated images. Mass retractions traceable to such operations were a substantial driver of the record retraction totals of recent years, and integrity-screening organizations now flag suspect papers by the thousand. This shifts part of the burden onto readers and reviewers, since a mill-produced paper is designed to survive a superficial review.

Key idea: The modern threat to the record is industrial rather than individual - fake journals and paper mills produce volume that no traditional review system was designed to filter.

Openness in review

Traditional review is often anonymous, which protects candor but can shield unfair or careless reviewing from accountability. In response, some venues now use open peer review, publishing reviewers' reports or names, and post-publication review, in which scrutiny continues in the open after a paper appears. These experiments reflect the same impulse driving open science generally: that transparency, applied to evaluation as well as to data, makes the whole system more trustworthy. Whatever the model, the reviewer's underlying obligation is constant - to serve the integrity of the shared record with the same honesty you would want applied to your own work.

Correction of the record deserves a final word, because most researchers misread the signals. A correction or erratum fixes a specific error while the paper's conclusions stand. An expression of concern is an editor's public notice that a serious question exists and is unresolved. A retraction withdraws the paper from the record, and the notice should say who requested it and why, distinguishing honest error from misconduct. Retracted papers remain visible and citable, marked as retracted, precisely so the record shows what happened rather than pretending the paper never existed. Authors who discover a serious error in their own work and write to the editor are doing the hardest and most admirable thing in this lesson.

Key idea: Retraction is how the literature corrects itself, so read a retraction notice for its stated reason rather than treating the word as a verdict on the authors.

Where people get stuck

Four errors are common among new reviewers.

  • Reviewing outside your competence. Accepting because you are flattered, then judging methods you do not understand, harms authors more than declining would. Decline, and suggest someone better placed.
  • Sharing the manuscript informally. Forwarding to a colleague for a second opinion, or discussing it at a group meeting, breaches confidentiality unless the editor has agreed and the participants are named.
  • Confusing duplicate publication with salami slicing. Duplicate publication reports the same work twice as if each were new. Salami slicing splits one coherent study into several thin papers. One repeats, the other fragments.
  • Using delay as a weapon. Sitting on a competitor's manuscript is a quiet abuse that is nearly impossible to prove and entirely visible to the person doing it. If you cannot review promptly, decline promptly.

Key idea: Ask whether you would be comfortable with the authors seeing exactly what you did and when, and nearly every reviewing dilemma resolves itself.

Recap

  • Peer review serves the field by helping editors decide, catching error and overstatement, and improving manuscripts through actionable critique.
  • A manuscript under review is privileged and unpublished, so it may not be shared, kept, or used before publication.
  • Co-reviewing with a trainee is legitimate only with the editor's agreement, the same confidentiality obligations, and the trainee named in the report.
  • Uploading a manuscript to a third-party generative service breaches confidentiality and is increasingly prohibited outright.
  • Competing or closely collaborating reviewers should decline and say why, since even the appearance of conflict is damaging and declining costs nothing.
  • Useful reviews separate fatal problems, substantive fixes, and personal preferences, and phrase objections so that authors can act on them.
  • Duplicate publication, salami slicing, undisclosed reuse, predatory publishing, and paper mills each corrupt the record in a different way.
  • Corrections, expressions of concern, and retractions are the mechanisms of self-correction, and the reason stated in a notice matters more than the label.

Sources

  1. International Committee of Medical Journal Editors. (n.d.). Responsibilities in the submission and peer-review process. icmje.org
  2. International Committee of Medical Journal Editors. (n.d.). Recommendations for the conduct, reporting, editing, and publication of scholarly work in medical journals. icmje.org
  3. Committee on Publication Ethics. (2017). COPE ethical guidelines for peer reviewers. COPE. (COPE materials sit behind a bot-protection layer that blocks automated retrieval; consult publicationethics.org ↗ directly.)
  4. Severin, A., Strinzel, M., Egger, M., Domingo, M., & Barros, T. (2021). Characteristics of scholars who review for predatory and legitimate journals. BMJ Open, 11(7), e050270. pmc.ncbi.nlm.nih.gov
  5. Van Noorden, R. (2023). More than 10,000 research papers were retracted in 2023 - a new record. Nature, 624, 479-481. nature.com
  6. Van Noorden, R. (2024). Journals with high rates of suspicious papers flagged by science-integrity start-up. Nature. nature.com
  7. World Association of Medical Editors. (n.d.). WAME: A global association of editors of peer-reviewed medical journals. wame.org
Key terms
Peer review
The evaluation of a manuscript by independent experts to assess its rigor, validity, and contribution before publication.
Reviewer confidentiality
The duty not to share or exploit a privileged, unpublished manuscript under review.
Duplicate publication
Publishing the same work in more than one venue as though each were original.
Salami slicing
Fragmenting one study into several minimal papers to inflate a publication count.
Predatory publishing
Journals that collect fees but provide no genuine peer review or quality control.
Open peer review
Review models that increase transparency by publishing reviewers' reports or identities.

Responsible Mentoring and the Culture of Integrity

  • Describe the mentor's responsibilities to trainees and to the integrity of the field.
  • Explain how power dynamics and institutional pressures shape ethical behavior, and what protects it.

Research ethics is not only a set of individual choices; it is sustained or corroded by the culture and relationships in which researchers work. The single most formative relationship in a scientist's development is with a mentor - and mentoring is therefore an ethical responsibility in its own right, the channel through which the norms of an entire field are transmitted or lost. This final lesson turns from rules to the human environment that makes rules stick.

Sofia Delgado has just started as an assistant professor. She has a five-year clock, a start-up package, one postdoc, and two incoming doctoral students. Nobody has told her how to run a laboratory, and she will mostly reproduce what was done to her. Her own doctoral supervisor was brilliant, absent, and quick to anger when results were disappointing, and Sofia learned early to bring him only the findings that worked. She now has to decide whether that is what she teaches, because she will teach it either way. Everything in this lesson is a choice she is making in her first six months, mostly without noticing that she is making it.

Key idea: Nobody trains you to be a mentor, so you default to reproducing the laboratory you came from - which makes the decision to examine it the whole ethical move.

In plain terms

People do what the people above them reward. That is nearly the entire content of research culture. If a supervisor lights up at a positive result and goes quiet at a null one, everyone in the group learns to produce positive results. Nobody has to say anything dishonest for that to happen. The trainee just runs one more analysis, or quietly stops mentioning the study that did not work. Sofia's real influence is not in what she says about integrity at the lab meeting. It is in how she reacts, in the first year, the first time a student says the experiment did not work. Everyone will be watching that moment, and they will calibrate to it for years.

What a mentor owes a trainee

A research mentor is more than a supervisor of tasks. Recognized responsibilities include:

  • Training in responsible conduct - teaching not only techniques but the ethical standards of the discipline, by explicit instruction and, more powerfully, by example. Trainees learn integrity mostly by watching what their mentor tolerates.
  • Fair credit and support of independence - giving trainees deserved authorship, promoting their work, and steadily fostering their capacity to work on their own, rather than exploiting their labor to advance the mentor.
  • Realistic guidance and honest feedback - including candid advice about careers and the limits of a chosen path.
  • Reasonable protection - shielding trainees, so far as possible, from undue pressure, and never demanding participation in questionable practices.

A national consensus report on mentorship in science and engineering makes a further point that is easy to miss. Effective mentorship is a set of learnable behaviors, not a personality trait, and it is best understood as a network rather than a dyad. No single person can supply technical training, career guidance, psychosocial support, sponsorship, and community for one trainee. Groups that assume otherwise produce trainees who are dependent on one relationship and stranded when it fails. Sofia's practical move is to help each student build a committee and a set of outside contacts early, which also reduces the danger her own power creates.

Two lightweight instruments do most of the work. A mentoring compact is a short written statement of mutual expectations - meeting frequency, response times on drafts, authorship principles, vacation, what happens when someone is stuck - agreed at the start and revisited annually. An individual development plan records the trainee's own goals and the skills needed to reach them, and it is now expected by major funders for graduate students and postdoctoral researchers. Neither document is magic. Both force conversations that otherwise happen only after something has gone wrong.

Key idea: Mentoring is a learnable practice supported by a network and two short written documents, not a talent some supervisors happen to have.

The peril of power asymmetry

The mentor-trainee relationship is marked by a steep power differential. A supervisor typically controls a student's funding, authorship, recommendation letters, and future prospects. That dependency is what makes certain abuses possible - coercion authorship, pressure to produce a desired result, or exploitation of labor - and what makes trainees reluctant to report problems they observe. Recognizing this asymmetry is the first step to using power responsibly rather than carelessly.

The asymmetry is worse for some trainees than others. International students whose visa status depends on continued enrolment, students with caring responsibilities, and those without family financial cushioning have materially less capacity to walk away from a bad situation. A supervisor who says "my door is always open" has not addressed any of that. What reduces the asymmetry is structural: a thesis committee that meets without the supervisor present at least once a year, a departmental graduate director with an independent relationship to each student, portable funding where it can be arranged, and a written statement that a student may change advisors without penalty.

Key idea: Good intentions do not reduce a power differential; independent relationships and documented exit routes do.

Pressure and the conditions for integrity

Individual virtue is necessary but not sufficient; environments powerfully shape behavior. The intense pressures of academic life - "publish or perish," competition for scarce funding, the premium on novel positive results - can push otherwise honest people toward corner-cutting. A healthy research culture counteracts these pressures deliberately:

  1. Psychological safety - an environment in which a junior member can question a result, admit a mistake, or report a concern without fear of retaliation. Where raising a problem is dangerous, problems are hidden rather than fixed.
  2. Modeling by leaders - senior researchers who visibly value rigor and honesty over sheer output set the norm the whole group follows.
  3. Protected channels - ombuds offices, research-integrity officers, and clear procedures that let concerns be raised and whistleblowers be protected from reprisal.

There is empirical support for the claim that climate matters more than character. Survey research on academic scientists has found that perceptions of organizational justice - whether resources, credit, and decisions are distributed fairly, and whether procedures are applied consistently - are associated with self-reported misbehavior. Researchers who felt unfairly treated reported more problematic behavior than those who did not. That finding reframes the whole problem. If misbehavior tracks perceived unfairness, then departmental practices about space, funding, promotion, and credit are integrity interventions, whether or not anyone labels them that way.

Sofia can act on this at her own scale. She can state how authorship will be decided before projects start. She can distribute the tedious work visibly rather than assigning it to whoever complains least. She can say out loud, in the first month, that a null result is a result and that she wants to hear about failed experiments quickly. And she can do the one thing that costs the most: when a student brings her a result that undermines a grant she has already written, she can thank them, in front of other people.

Key idea: Misbehavior tracks perceived unfairness, so how a department allocates credit, space, and money is part of its integrity system.

The obligation to respond

Finally, membership in a research community carries a duty not to look away. If you have a good-faith, well-founded concern about serious wrongdoing, you have an obligation to raise it through appropriate channels - discussing it first with a trusted advisor or an integrity officer, documenting what you observed, and acting on evidence rather than rumor.

This is hard, especially against a power differential, which is exactly why whistleblower protections exist and why a mature institution treats a good-faith report as a service rather than a betrayal. The through-line of this entire course returns here: ethics is not a set of external constraints imposed on science from outside, but the very fabric of trust that makes cumulative, collaborative knowledge possible. You inherit that fabric from those who trained you, and you will hand it, strengthened or frayed, to those you train.

Retaliation is not merely discouraged; it is defined and prohibited. Federal misconduct policy defines retaliation as an adverse action taken against a complainant, witness, or committee member in response to a good-faith allegation or cooperation with a proceeding, and institutions are required to protect against it. The protection is real and it is also imperfect, which is why the practical advice for a trainee has three parts: document contemporaneously, consult before acting, and use the channels that exist rather than improvising. Consulting an ombuds office, which is typically confidential and off the record, is not the same as filing an allegation and is usually the right first step.

Key idea: A confidential conversation with an ombuds office is not an accusation, and treating it as the first step lets people get advice before anything becomes irreversible.

Where people get stuck

Four confusions close out the course.

  • Treating mentoring as a personality. It is a set of behaviors that can be learned, described in writing, and improved with feedback. "I am just not a nurturing person" is a description of a starting point, not an exemption.
  • Confusing psychological safety with low standards. They are independent. The strongest groups combine high expectations with genuine safety in admitting error; the dangerous combination is high expectations with fear.
  • Believing an open door is a structure. An invitation is not a mechanism. Independent committees, portable funding, an ombuds office, and written policies are mechanisms, and they work when the supervisor is the problem.
  • Assuming reporting means going public. The sequence is document, consult confidentially, then use the institution's channel. Public accusation is neither the first step nor usually the effective one.

Key idea: Culture is not what a laboratory says about integrity; it is what happens the first time someone brings bad news.

Recap

  • Mentoring transmits a field's norms mostly by example, which makes it an ethical responsibility rather than an administrative one.
  • A mentor owes training in responsible conduct, fair credit, support for independence, honest guidance, and protection from undue pressure.
  • Effective mentorship is a learnable set of behaviors and works best as a network rather than a single relationship.
  • Mentoring compacts and individual development plans force the conversations that otherwise happen only after something goes wrong.
  • The power differential is sharper for trainees whose visa status, finances, or caring responsibilities limit their ability to leave, and it is reduced by structures rather than by good intentions.
  • Perceptions of organizational justice are associated with self-reported misbehavior, so fair allocation of credit and resources is an integrity intervention.
  • Psychological safety and high standards are independent; the dangerous combination is high pressure with fear of admitting error.
  • Retaliation against good-faith complainants is defined and prohibited, and the practical sequence is document, consult confidentially, then use established channels.

Sources

  1. National Academies of Sciences, Engineering, and Medicine. (2019). The science of effective mentorship in STEMM. National Academies Press. ncbi.nlm.nih.gov
  2. National Academies of Sciences, Engineering, and Medicine. (2019). The science of mentoring relationships: What is mentorship? In The science of effective mentorship in STEMM. National Academies Press. ncbi.nlm.nih.gov
  3. Martinson, B. C., Anderson, M. S., & de Vries, R. (2005). Scientists behaving badly. Nature, 435, 737-738. nature.com
  4. Martinson, B. C., Anderson, M. S., Crain, A. L., & de Vries, R. (2006). Scientists' perceptions of organizational justice and self-reported misbehaviors. Journal of Empirical Research on Human Research Ethics, 1(1), 51-66. pmc.ncbi.nlm.nih.gov
  5. Martinson, B. C., Crain, A. L., & de Vries, R. (2010). The importance of organizational justice in ensuring research integrity. Journal of Empirical Research on Human Research Ethics, 5(3), 67-83. pmc.ncbi.nlm.nih.gov
  6. U.S. Public Health Service. (2024). Retaliation, 42 C.F.R. 93.238. Electronic Code of Federal Regulations. ecfr.gov
  7. National Institutes of Health. (2014). Revised policy: Descriptions on the use of individual development plans (IDPs) for graduate students and postdoctoral researchers (NOT-OD-14-113). grants.nih.gov
Key terms
Mentor
A senior researcher responsible for a trainee's technical, professional, and ethical development.
Power differential
The asymmetry of control over funding, authorship, and prospects that shapes the mentor-trainee relationship.
Publish or perish
The career pressure to produce frequent publications, which can incentivize corner-cutting.
Psychological safety
An environment where members can question, admit error, or report concerns without fear of retaliation.
Research-integrity officer
An institutional official who oversees integrity education and the handling of misconduct concerns.
Whistleblower protection
Safeguards shielding those who report good-faith concerns from retaliation.

Open the interactive version with quizzes and progress →