Module 1: What Would Make It Science
The demarcation problem argued from a real court case, Hume's argument that induction has no non-circular justification, and the two paradoxes that broke the simplest theory of how evidence supports a hypothesis.
Demarcation: The Problem That Will Not Stay Solved
- State the demarcation problem and explain why courts, funding bodies and regulators need an answer to it.
- Evaluate the verificationist and falsificationist criteria and identify the objection that sinks each.
- Assess Laudan's argument that the problem should be abandoned, and the case for the cluster accounts that replaced it.
A judge writes a definition of science
On 5 January 1982, Judge William Overton of the United States District Court for the Eastern District of Arkansas struck down Act 590, a state law requiring public schools to give balanced treatment to creation science alongside evolution. To decide the case he needed to know whether creation science was science, and so his opinion set out five essential characteristics: that it is guided by natural law, that it explains by reference to natural law, that it is testable against the empirical world, that its conclusions are tentative, and that it is falsifiable. Creation science, he ruled, met none of them.
The philosopher Michael Ruse had testified for the plaintiffs and the five criteria closely tracked his testimony. Within months another philosopher, Larry Laudan, published a short and blistering commentary arguing that the decision, though it reached the right result, had reached it by bad reasoning. Creationism, Laudan pointed out, does make testable claims. It claims the earth is young and that a global flood laid down the geological column. Those claims have been tested. They are false. To call them untestable is to hand creationists the true reply that they have been tested, and to obscure the actual objection, which is not that creationism is unfalsifiable but that it has been falsified and its proponents will not say so.
Two philosophers, one courtroom, opposite views about what the court should have said. That disagreement is the demarcation problem in its most concrete form, and it is not settled.
Key idea: The demarcation problem is not an academic exercise; a court had to answer it in order to decide a case, and two competent philosophers disagreed sharply about the answer it should give.
Why anyone needs a criterion
The question of what separates science from non-science looks like the kind of thing you could leave unresolved. In practice institutions cannot.
- Courts. In 1993 the United States Supreme Court, in Daubert v. Merrell Dow Pharmaceuticals, replaced the older test of general acceptance for expert testimony with a set of factors that judges should weigh, among them whether a theory can be and has been tested, whether it has been peer reviewed, its known error rate, and its acceptance in the relevant community. The opinion cited Popper and Hempel by name. Judges have been doing philosophy of science ever since, whether or not they wanted to.
- Public funding and public health. Deciding what a health service will pay for, or what a research council will fund, requires distinguishing claims that have earned credence from claims that have not.
- Education. The Arkansas case is one of a long series; what a curriculum presents as established knowledge is a demarcation decision.
So the question is forced. What has proved hard is answering it with a criterion that classifies the uncontroversial cases correctly and does not simply restate the prejudice it was meant to justify.
The verificationist attempt
The first serious modern attempt came from the Vienna Circle in the 1920s and 1930s and reached English readers most influentially through A. J. Ayer's Language, Truth and Logic in 1936. Its criterion was verifiability: a statement is meaningful, and belongs to the empirical part of knowledge, only if it can be verified by observation, or is a matter of definition and logic. Metaphysics, theology and much of ethics were held to fail the test not merely as bad science but as literally meaningless.
Three objections finished it, and they are worth having in order.
- Universal statements cannot be verified. All copper conducts electricity covers indefinitely many cases, and no finite set of observations establishes it. A criterion demanding conclusive verification therefore expels the laws of physics along with metaphysics, which is the opposite of the intended result.
- The criterion fails its own test. The claim that only verifiable statements are meaningful is neither a truth of logic nor something anyone could observe. On its own terms it is meaningless.
- Weakened versions let everything in. Ayer's revised criterion, requiring only that a statement help entail some observation when combined with auxiliary premises, was shown by Alonzo Church in 1949 to admit essentially any sentence whatever, given the right auxiliaries.
The point: Verificationism failed not because it was too strict about religion but because it was too strict about physics, and every attempt to loosen it lost the ability to exclude anything.
The falsificationist attempt
Karl Popper proposed the alternative that has dominated the popular understanding ever since. He was struck, as a young man in Vienna, by a contrast. Einstein's general relativity predicted that starlight passing the sun would be deflected by a specific amount, and if the measurement had come out otherwise the theory would have been finished. Marxist history and psychoanalysis, by contrast, seemed able to accommodate any outcome. A man pushes a child into a river and a man saves a child from drowning; the Freudian explains both, the first by repression and the second by sublimation. Popper concluded that the strength of a theory lies in what it forbids.
So the criterion is falsifiability: a theory belongs to empirical science if some possible observation is inconsistent with it. Lesson 4 gives Popper's position in full and in its strongest form, because it is far more sophisticated than the slogan. Here the question is only whether falsifiability works as a demarcation line, and the answer is that it does not quite.
- Astrology is falsifiable. Newspaper astrology makes predictions, and studies have tested them and found no effect. If falsifiability were the criterion, astrology would qualify as science and merely be bad science. Yet almost nobody wants that verdict, which suggests the criterion is not tracking what we care about.
- Statistical hypotheses are not strictly falsifiable. A claim that a coin is fair is not contradicted by any finite run of heads. Real sciences are full of such hypotheses, and rejecting them requires a decision about thresholds rather than a deduction.
- Nothing is tested alone. This is the Duhem-Quine point, and Lesson 5 works it through. Any prediction depends on auxiliary assumptions, so a failed prediction never tells you which component to blame.
Laudan's demolition
In 1983 Laudan published a paper arguing that the demarcation problem should be abandoned. His case had three parts.
First, no proposed criterion works. Each either excludes uncontroversially scientific work or admits uncontroversially unscientific work, and usually both.
Second, the criteria that have been offered are not doing the work people think. What we actually care about is whether a particular claim is well supported, and that is a question about evidence, not about membership in a category. Astrology's problem is that its claims are unsupported and, where tested, false. Nothing is added by also calling it unscientific.
Third, the label is doing rhetorical rather than epistemic work. Calling something pseudoscience is a way of dismissing it without engaging the evidence, and Laudan thought this was corrupting.
The reply, made forcefully since, is that the third point is where the argument is weakest. Whether or not a global criterion exists, societies have to make institutional decisions about which claimants to expertise get standing, and they cannot re-adjudicate every claim from first principles. A regulator cannot test every therapy; a school board cannot referee every controversy. The category does practical work that case-by-case evidential assessment cannot do at scale.
What survives
Most current work has given up on a single necessary and sufficient condition and treats science as a cluster concept: a family of features, none required, that together make a practice recognizable. Typical members of the cluster include the following.
| Feature | What it looks like when present | What its absence looks like |
|---|---|---|
| Testability with risk | The theory forbids outcomes, and someone checks | Every outcome is accommodated after the fact |
| Response to disconfirmation | Failed predictions change the theory or the field | Failures are absorbed and the doctrine is unchanged for a century |
| Connection to the rest of knowledge | Claims cohere with established results in neighbouring fields | The claim requires everything else to be quietly wrong |
| Error-correcting institutions | Replication, peer review, adversarial checking | Findings circulate without any mechanism for being overturned |
| Fecundity | The theory generates new problems and new results | The theory generates only restatements and defences |
Run astrology through the table and the diagnosis is precise in a way that a single criterion could not be. The predictions have been tested and failed; the practice has not changed in response; the mechanism it requires is inconsistent with well-established physics; there is no error-correcting community; and it has produced no new problems in centuries. That is a much better account of what is wrong with it than the claim that it is unfalsifiable, which is simply untrue.
What matters here: Replacing a single criterion with a cluster of features gives up the clean line but buys accurate diagnoses, which is a trade most philosophers of science now think worth making.
What would settle it
State the two positions at their strongest and then say what evidence bears on them.
The eliminativist, following Laudan, holds that there is no interesting property shared by all and only the sciences, and that every question worth asking is a question about the support for a particular claim. What would settle this in their favour is the continued failure of proposed criteria, and half a century of failures is real evidence.
The reformer holds that a cluster account can do the work, because the features listed above are not arbitrary; they are the features that make error-correction possible, and error-correction is what warrants trust. What would settle this in their favour is showing that the cluster classifies contested cases in a way that does not merely reproduce the classifier's prior opinion. That is a testable claim about the criterion, and it is the right place to press.
Notice that neither side needs the other to be stupid. The disagreement is about whether an epistemically respectable general category exists, and both sides agree that particular claims are assessed by evidence.
Common misconceptions
- Popper solved the demarcation problem. Popper offered the most influential proposal, and it captures something real about risky prediction. It does not survive as a necessary and sufficient condition, as astrology and statistical hypotheses show.
- Pseudoscience means untestable. Many pseudosciences make testable claims. What marks them out is what happens after the tests fail.
- The demarcation problem is about whether religion is true. It is about the boundaries of a method. A demarcation criterion is compatible with religious claims being true, false, or of a kind that the criterion does not address.
- If there is no sharp line, anything goes. The absence of a sharp boundary is compatible with clear cases on either side. There is no sharp line between day and night, and midnight is still dark.
- Peer review is what makes something science. Peer review is a recent institutional arrangement, absent from most of the history of science, and it is one error-correcting mechanism among several rather than the defining feature.
Where this leaves us
- The demarcation problem is forced on courts, regulators and schools, so declining to answer it is not an available option institutionally.
- Verificationism failed because universal laws are not conclusively verifiable, because the criterion fails its own test, and because weakened versions admit everything.
- Falsifiability captures the value of risky prediction but admits astrology, excludes statistical hypotheses, and runs into the Duhem-Quine problem.
- Laudan argued the problem should be abandoned because no criterion works and the label does rhetorical rather than epistemic work.
- The reply is that institutions need a workable category, since they cannot adjudicate every claim from first principles.
- Current accounts treat science as a cluster of features centred on testability, response to disconfirmation, coherence with neighbouring knowledge, and error-correcting institutions.
- What would settle the dispute is whether a cluster account classifies contested cases without simply reproducing the classifier's prior judgment.
Sources
- Hansson, S. O. (n.d.). Science and pseudo-science. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Andersen, H., & Hepburn, B. (n.d.). Scientific method. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Thornton, S. (n.d.). Karl Popper: philosophy of science. Internet Encyclopedia of Philosophy. iep.utm.edu
- Wikipedia contributors. (n.d.). McLean v. Arkansas. Wikipedia. en.wikipedia.org
- Laudan, L. (1982). Commentary: Science at the bar, causes for concern. Science, Technology, and Human Values, 7(41), 16-19.
- Ayer, A. J. (1936). Language, truth and logic. Gollancz.
- Key terms
- Demarcation problem
- The problem of stating what distinguishes science from non-science or pseudoscience, and whether any such general criterion exists.
- Verificationism
- The Vienna Circle doctrine that a statement is meaningful only if it can be verified by observation or is true by definition.
- Falsifiability
- Popper's proposed criterion: a theory is empirical if some possible observation would be inconsistent with it.
- Pseudoscience
- A practice claiming scientific standing while lacking the features, especially responsiveness to disconfirmation, that warrant it.
- Cluster concept
- A concept defined by a family of typical features, none of which is individually necessary or jointly sufficient.
- Daubert standard
- The 1993 United States Supreme Court test for admitting expert testimony, weighing testability, peer review, error rate and acceptance.
- Fecundity
- The capacity of a theory to generate new problems, predictions and lines of research rather than only defences of itself.
Hume's Problem of Induction
- Reconstruct Hume's argument as a numbered set of steps and identify the premise each reply attacks.
- Evaluate the dissolutionist, pragmatic, falsificationist and Bayesian responses and state what each one costs.
- Distinguish the claim that induction is unjustified from the claim that it is unreasonable to use.
A chicken and a farmer
In 1912 Bertrand Russell offered an image that has never been improved on. A chicken is fed by the same man every morning of its life. Each feeding increases its confidence in the generalization that this man brings food. The inference is exactly the kind we make constantly, and it is supported by a long and unbroken run of confirming instances. Then one morning the man wrings its neck.
The chicken did nothing wrong by the standards of the reasoning it was doing. It generalized from a large, consistent sample. What Russell was illustrating, and what David Hume had argued a hundred and seventy years earlier, is that no number of past instances entails anything about the next one, and that the gap cannot be closed by any further evidence of the same kind.
The unsettling part is not that induction sometimes fails. Everybody knows that. The unsettling part is that when you try to say why it is reasonable at all, every route out appears to be blocked.
Key idea: The problem of induction is not that inductive inferences are sometimes wrong; it is that no non-circular argument seems available for thinking they are ever any good.
The argument, in steps
Hume set the argument out in A Treatise of Human Nature in 1739 and in a shorter, sharper version in the Enquiry Concerning Human Understanding in 1748. Numbered, it runs like this.
- Every inference from what we have observed to what we have not observed depends on an assumption: that the unobserved cases will resemble the observed ones. Call it the principle of the uniformity of nature.
- Any belief we are entitled to must be supported either by demonstrative reasoning, which concerns relations of ideas, or by experience, which concerns matters of fact. Hume allowed no third source.
- The uniformity principle cannot be established by demonstration, because its denial involves no contradiction. That the course of nature might change tonight is perfectly conceivable, and what is conceivable is not ruled out by logic alone.
- It cannot be established by experience either. All the experience we have is of the past. To move from the past success of uniformity to its future success is itself an inference from observed to unobserved, which is precisely what the principle was needed to license. The argument would assume what it set out to prove.
- Therefore the uniformity principle has no justification, and neither does any inference resting on it.
Read step 4 again slowly, because that is where the whole thing turns. The point is not that induction has a poor track record. Its track record is excellent. The point is that appealing to that track record is a use of induction, and so cannot serve as its foundation.
Hume's own answer
Hume did not conclude that we should stop reasoning inductively, and he did not think we could. His account is psychological rather than justificatory. Repeated experience of one kind of event followed by another produces in the mind a habit, so that the appearance of the first carries the thought to the second. He called it custom. It is a fact about how creatures like us are built, not a reason.
That leaves an awkward but consistent position: our most successful practice has a cause but no justification. Hume was untroubled by this in a way many of his readers have not been, and he pointed out that nature has not left something so essential to the uncertain outcome of our reasoning.
Worth holding on to: Hume's conclusion is about justification, not about advisability. He continued to reason inductively, and thought it inevitable that we should.
Six replies, and what each costs
| Reply | The move | The cost |
|---|---|---|
| Inductive justification | Induction has worked, so it will work | Circular in exactly the way step 4 describes, unless a special defence is mounted |
| Rule-circularity defence | Circularity in the premises is vicious; circularity in the rule used may not be | Reassures only those already using the rule, and the same defence would vindicate a bad rule that endorsed itself |
| Strawson's dissolution | Being reasonable partly means being supported by evidence in the inductive way; the question is ill-formed | Concedes the substantive question: whether the practice tends to produce true beliefs |
| Reichenbach's vindication | If any method of prediction succeeds, induction succeeds in the long run, so it is a no-loss bet | Says nothing about the short run, and infinitely many rival rules converge in the limit too |
| Popper's denial | Science does not use induction; it conjectures and tries to refute | If corroboration guides what we rely on, induction has returned under a new name; if it does not, the theory offers no advice |
| Bayesian updating | Rationality is coherent probability updating on evidence | The prior probabilities are not fixed by the evidence, so the problem reappears as the problem of the priors |
Two of these deserve a closer look, because they are the ones that most often persuade people on first hearing.
Reichenbach's argument is genuinely ingenious. It does not claim induction will work. It claims that if anything works, induction will, so we lose nothing by using it. The difficulty is that the guarantee holds only in the limit of infinite evidence, and in the limit so do countless alternative rules that give wildly different advice in the short run, which is the only run anybody lives in.
Popper's denial is the boldest response: accept Hume entirely and deny that science ever needed induction. Wesley Salmon pressed the obvious question. When an engineer chooses the well-corroborated bridge design over the refuted one, is that choice rational? If yes, then corroboration is being used as a guide to future performance, which is an inductive step. If no, then the theory has nothing to say about the practical use of science, which is a great deal to give up. Lesson 4 returns to how Popperians answer this.
Why this is not scepticism about science
It is easy to draw the wrong moral. Hume's argument does not show that scientific conclusions are unreliable, that all beliefs are equally good, or that you should stop expecting the sun to rise. It shows something narrower and stranger: that our best practice cannot be given a non-circular justification of the kind philosophers wanted.
Notice also that the argument applies to itself in an interesting way. If you concluded from Hume that induction will probably fail us in the future, you would have made an inductive inference. The argument establishes an absence of justification in both directions, not a prediction of disaster.
Most working scientists, told about the problem, respond that induction works, which is true and is not an answer. Most philosophers accept that no non-circular justification has been produced and disagree about what follows. The interesting question is whether that verdict is a scandal or simply the discovery of where justification runs out.
What would settle it
Here the honest answer is unusually clean. A non-circular argument from premises a sceptic would accept, ending in the conclusion that inductive inference is more likely than not to yield truths, would settle it. Nobody has produced one in nearly three centuries, and there are reasons to think none can exist, since any such argument must either assume a substantive claim about the world's uniformity or derive it from experience.
What is genuinely live is narrower. Is rule-circularity acceptable, given that a defence of deduction faces the same structure? Does the combinatorial argument, that large samples are very likely to be representative of the population they are drawn from, help, or does it merely relocate the assumption into the claim that the future is drawn from the same population? Both are technical questions with real work being done on them, and both are places where a determined student can still make progress.
Common misconceptions
- Hume showed that induction is unreliable. He showed it is unjustified in a particular sense. Reliability is a further claim, and one that could only be established or refuted inductively.
- The problem is solved by noting that science uses probability rather than certainty. The problem applies to probabilistic inference too: why should observed frequencies constrain unobserved ones? That is the same question in new clothing.
- Popper solved it by abandoning induction. He proposed a system without inductive support. Whether it can also account for rational reliance on corroborated theories is exactly what is disputed.
- It is merely a verbal problem, as Strawson said. Strawson's move dissolves one version of the question, but the substantive question, whether the practice is truth-conducive, survives the dissolution.
- Nobody takes this seriously outside philosophy classes. The structure recurs whenever a method must be defended by its own outputs, which is a live issue in machine learning generalization, in statistical model selection, and in the justification of statistical significance thresholds.
The short version
- Russell's chicken generalizes correctly by inductive standards and is still wrong, illustrating that instances never entail the next case.
- Hume's argument runs in five steps and turns on step 4: justifying uniformity by its past success uses the very inference in question.
- Hume's own account is psychological: custom produces the expectation, and no reason underwrites it.
- The main replies are inductive justification, rule-circularity, Strawson's dissolution, Reichenbach's vindication, Popper's denial and Bayesian updating, and each pays a cost.
- Reichenbach's guarantee holds only in the limit, where rival rules also converge; Popper's denial faces Salmon's question about rational reliance.
- The conclusion is about justification, not about reliability or advisability, and drawing a pessimistic prediction from it would itself be an induction.
- A non-circular argument for the truth-conduciveness of induction would settle it; the live questions concern whether rule-circularity is acceptable.
Sources
- Henderson, L. (n.d.). The problem of induction. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Morris, W., & Brown, C. (n.d.). David Hume. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Fieser, J. (n.d.). David Hume. Internet Encyclopedia of Philosophy. iep.utm.edu
- Hume, D. (1748). An enquiry concerning human understanding. Project Gutenberg. gutenberg.org
- Russell, B. (1912). The problems of philosophy. Williams and Norgate. Chapter 6, On induction.
- Key terms
- Induction
- Inference from observed cases to unobserved ones, including generalization from a sample and prediction of the next instance.
- Uniformity of nature
- The assumption that unobserved cases resemble observed ones, which Hume argued every inductive inference presupposes.
- Hume's fork
- The division of all knowable claims into relations of ideas, established by demonstration, and matters of fact, established by experience.
- Circularity
- The defect in an argument whose justification of a principle relies on that same principle, as an inductive defence of induction does.
- Custom
- Hume's term for the habit of expectation produced by repeated experience, offered as the cause of inductive inference rather than its justification.
- Rule-circularity
- The situation in which an argument uses a rule to support that same rule without assuming its conclusion as a premise.
- Pragmatic vindication
- Reichenbach's defence that induction will succeed in the long run if any method will, making it a no-loss policy.
- Problem of the priors
- The Bayesian version of the difficulty: probabilistic updating is well defined, but the initial probabilities are not fixed by evidence.
Confirmation: Ravens, Grue, and the Trouble with Instances
- Derive the raven paradox from Nicod's criterion and the equivalence condition, and identify which premise each proposed solution rejects.
- Work the Bayesian treatment of the ravens numerically and state what it shows.
- Explain Goodman's new riddle of induction, including why the obvious objection to grue fails.
A green apple and a black raven
Take the hypothesis that all ravens are black. What would confirm it? The obvious answer is that observing a black raven does. Observing a white raven would refute it. And observing a shoe, a comet or a green apple would be irrelevant.
Carl Hempel showed in 1945 that this obvious answer cannot be right, using two premises that look impossible to give up.
The first is Nicod's criterion, named after Jean Nicod: a universal statement of the form all F are G is confirmed by observing something that is both F and G. Black ravens confirm the raven hypothesis; that is what a positive instance means.
The second is the equivalence condition: if a piece of evidence confirms a hypothesis, it confirms anything logically equivalent to that hypothesis. This is hard to deny, because logically equivalent statements are true in exactly the same circumstances. They are, in the sense that matters for evidence, the same claim written two ways.
Now put them together. All ravens are black is logically equivalent to all non-black things are non-ravens. By Nicod's criterion, the second is confirmed by observing something that is both non-black and a non-raven. A green apple is non-black and is a non-raven. So a green apple confirms that all non-black things are non-ravens, and by the equivalence condition it therefore confirms that all ravens are black.
You can, it appears, do ornithology indoors, by examining the contents of your fruit bowl.
The point: The paradox is not a puzzle about birds. It is a demonstration that the simplest account of what counts as evidence for a generalization has an unacceptable consequence, derived from premises that individually look undeniable.
Three ways out, and what each costs
Since the conclusion follows validly, the only options are to reject a premise or to accept the conclusion.
- Reject the equivalence condition. This means holding that evidence can bear differently on two statements that are true in exactly the same possible situations. That is a large commitment: it implies that confirmation is sensitive to how a hypothesis is written, which makes it a relation between evidence and sentences rather than between evidence and claims about the world. Few have been willing to pay it.
- Restrict Nicod's criterion. Perhaps positive instances confirm only when the hypothesis is stated in its natural form. But now the account needs a theory of which form is natural, and there is no obvious one. Without such a theory the restriction is simply the paradox written as a rule.
- Accept the conclusion and explain away the discomfort. This was Hempel's own line. The apple really does confirm, he said, and our resistance comes from smuggling in background knowledge: we already know the apple is not a raven, so the observation feels like it tells us nothing new. Strip the background knowledge away and the confirmation is genuine.
Hempel's response is right as far as it goes but leaves something out. It says the apple confirms; it does not say how much, and the intuition it is trying to explain away is not really that the apple confirms nothing but that it confirms far less than a raven does. That gap is what the Bayesian treatment fills.
The Bayesian treatment, with numbers
Here is a calculation you can check. Take a domain with 1,000,000 objects, of which 10,000 are ravens and 100,000 are non-black. Compare two hypotheses:
- H: every raven is black.
- K: exactly one raven is not black, and all the rest are.
Case one. You pick a raven at random and find it is black. Under H the probability of that outcome is 1, since all 10,000 ravens are black. Under K it is 9,999 divided by 10,000, or 0.9999. The likelihood ratio favouring H is 1 divided by 0.9999, which is about 1.0001.
Case two. You pick a non-black object at random and find it is not a raven. Under H the probability is 1, since under H no non-black thing is a raven. Under K it is 99,999 divided by 100,000, or 0.99999. The likelihood ratio favouring H is about 1.00001.
Both observations confirm H, exactly as the paradox says. But the black raven's likelihood ratio is ten times further from 1 than the apple's, and ten is precisely the ratio of non-black things to ravens in the domain. Generalize: the relative evidential weight of a positive instance to a contrapositive instance equals the ratio of the size of the complement class to the size of the subject class. Ravens are rare and non-black things are abundant, so a raven is enormously more informative than an apple.
This is a satisfying result because it vindicates both the logic and the intuition. The logic said the apple confirms; the intuition said it hardly matters. Both are right, and the Bayesian framework says exactly how much.
One caveat, and it is important. The result depends on the assumed sizes of the classes, which is background knowledge. If you knew that ravens vastly outnumbered non-black objects, the weights would reverse. So the resolution does not come from logic alone, which is itself a lesson: confirmation is not a purely formal relation between sentences.
Goodman's grue
Nelson Goodman published a harder problem in 1955, in Fact, Fiction, and Forecast. Define a predicate:
An object is grue if it is examined before some future time t and is green, or is not examined before t and is blue.
Now consider the evidence. Every emerald anyone has ever examined has been green. Every one of those emeralds was examined before t, and every one was green. So every emerald examined so far is also grue.
By Nicod's criterion the same body of evidence confirms two hypotheses:
- All emeralds are green.
- All emeralds are grue.
These agree about every emerald examined so far and disagree completely about emeralds first examined after t. One predicts green, the other blue. The evidence supports both equally well, and no amount of further evidence collected before t can separate them.
Why the obvious objection fails
Everyone's first response is the same: grue is a cheat, because it is defined using a time and green is not. Goodman anticipated this and showed the objection collapses.
Introduce a companion predicate: an object is bleen if it is examined before t and is blue, or is not so examined and is green. Now suppose a community whose basic vocabulary is grue and bleen. In their language, green is the defined term: something is green if it is examined before t and grue, or is not so examined and bleen. From their side, our predicate is the one with the time built into it.
The definitional dependence is symmetric. Which predicate looks gerrymandered depends entirely on which vocabulary you start from, and nothing in logic or in the evidence picks a starting point. That is why this is a new riddle and not a joke.
What matters here: Grue shows that the problem is not whether past instances support future claims, but which of the infinitely many patterns consistent with the past we are entitled to project.
Entrenchment, naturalness, and the state of play
Goodman's own answer was entrenchment. Green has a long history of successful projection; grue has none. Predicates earn projectibility by having been used in past inductions that worked out. On this view the standards of induction are, like the standards of deductive logic, codifications of practices we are unwilling to give up.
The objection writes itself: this says we should project what we have projected, which sounds like a description of a habit rather than a justification. Goodman would reply that a description of a settled practice, brought into reflective equilibrium, is the only justification available for any inferential rule, deduction included. Whether that is deep or evasive is a genuine and unresolved question.
Two rival programmes deserve mention. W. V. O. Quine argued in 1969 that projectible predicates are those corresponding to natural kinds, and that our innate sense of similarity, shaped by natural selection, is roughly aligned with the kinds nature contains. David Lewis and others developed an account of natural properties as a metaphysical primitive: some classes carve the world at its joints and some do not, independently of us. Both moves face the same challenge, which is to say what naturalness is without either circularity or an unexplained posit.
A third and increasingly common approach declines the syntactic framing altogether: what makes green projectible and grue not is that green figures in causal and nomic generalizations about light, pigment and crystal structure, and grue does not. This has the advantage of connecting projectibility to the rest of science and the disadvantage of presupposing that we already know which generalizations are lawlike, which is the subject of Lesson 10.
What would settle it
For the ravens, the dispute is essentially over. Almost everyone accepts the Bayesian analysis, and the residual disagreement concerns whether the appeal to background knowledge about class sizes shows that confirmation cannot be a purely formal relation. It does, and that is a result rather than a defeat.
For grue, nothing has settled it. What would is an account of projectibility that is independently motivated, that yields the right verdicts on cases not used to construct it, and that does not simply record which predicates we happen to favour. The causal and nomic approach is the most promising candidate and the most obviously in debt to other parts of philosophy of science, which is why the rest of this course keeps returning to it.
Common misconceptions
- The raven paradox is solved by saying the apple is irrelevant. Irrelevance is exactly what the two premises jointly rule out. Saying it without rejecting a premise is not an answer.
- Grue is defined in terms of time and green is not, so grue is illegitimate. The definitional relation is symmetric, as bleen demonstrates. This objection is the one Goodman wrote the book to defeat.
- Grue is about emeralds actually changing colour at time t. Nothing changes. A grue emerald examined after t is blue and always was; the predicate simply classifies objects by a disjunction involving examination time.
- These are word games with no bearing on science. Every curve fitted to a data set is a choice among infinitely many curves consistent with the data, and every choice of variables to measure is a choice of predicates to project. The riddle is the abstract form of a decision made daily.
- The Bayesian resolution derives everything from logic. It requires assumed class sizes. Change the background knowledge and the verdict changes, which is part of the finding.
What to carry forward
- Nicod's criterion plus the equivalence condition entails that a green apple confirms that all ravens are black, and the derivation is valid.
- Rejecting the equivalence condition makes confirmation sensitive to how a hypothesis is written; restricting Nicod's criterion requires a theory of natural formulations.
- Hempel accepted the conclusion and blamed the discomfort on imported background knowledge, which is right but says nothing about degree.
- The Bayesian treatment shows both observations confirm, with relative weight equal to the ratio of the complement class to the subject class.
- That resolution depends on background knowledge about class sizes, so confirmation is not a purely formal relation between sentences.
- Grue shows that any body of evidence supports incompatible generalizations, and the objection that grue smuggles in time fails because bleen makes the dependence symmetric.
- Goodman's entrenchment, Quine's natural kinds, Lewis's natural properties and causal-nomic accounts are the main candidate solutions, and none commands agreement.
Sources
- Crupi, V. (n.d.). Confirmation. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Cohnitz, D., & Rossberg, M. (n.d.). Nelson Goodman. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Internet Encyclopedia of Philosophy. (n.d.). Confirmation and induction. iep.utm.edu
- Wikipedia contributors. (n.d.). Raven paradox. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Grue and bleen. Wikipedia. en.wikipedia.org
- Goodman, N. (1955). Fact, fiction, and forecast. Harvard University Press.
- Key terms
- Nicod's criterion
- The principle that a universal generalization is confirmed by observing an object that satisfies both its antecedent and its consequent.
- Equivalence condition
- The principle that evidence confirming a hypothesis also confirms any logically equivalent hypothesis.
- Raven paradox
- Hempel's result that Nicod's criterion and the equivalence condition jointly imply that a green apple confirms that all ravens are black.
- Likelihood ratio
- The ratio of the probability of an observation under one hypothesis to its probability under a rival, which measures the evidence's discriminating power.
- Grue
- Goodman's predicate applying to things examined before a future time and green, or not so examined and blue.
- Bleen
- The companion predicate applying to things examined before that time and blue, or not so examined and green, used to show the definitional symmetry.
- Projectibility
- The property of a predicate that makes it suitable for inductive generalization from observed to unobserved cases.
- Entrenchment
- Goodman's proposal that a predicate is projectible in virtue of its history of successful use in past inductions.
- Natural kind
- A grouping held to correspond to a real division in the world rather than to a human classification, invoked by Quine to explain projectibility.
Module 2: Testing, and Why It Is Never Clean
Popper's falsificationism given its strongest form and then pressed hard, the Duhem-Quine thesis worked through two famous predictions from the same theory, and Kuhn's account of what actually happens when a science changes its mind.
Popper: Conjectures, Refutations, and the Case for Falsification
- State the logical asymmetry between verification and falsification and show how it drives Popper's methodology.
- Explain why Popper prefers improbable theories and why corroboration is not a measure of probable truth.
- Assess the five main objections to falsificationism and the Popperian replies to each.
Six November 1919
At a joint meeting of the Royal Society and the Royal Astronomical Society in London, Frank Dyson announced the results of two expeditions sent to observe the total solar eclipse of 29 May 1919, one to Principe off the coast of West Africa under Arthur Eddington and one to Sobral in Brazil. They had photographed stars near the eclipsed sun to measure whether their light was deflected by its gravity.
Two numbers were on the table before the plates were developed. Newtonian gravitation, treating light as corpuscular, gave a deflection at the sun's limb of about 0.87 seconds of arc. Einstein's general theory of relativity gave about 1.75, twice as much. The measurements came out near the larger figure.
What impressed the young Karl Popper, then in Vienna, was not that Einstein was right. It was that Einstein had said in advance what would show him wrong. If the deflection had come out at 0.87, general relativity was finished, and Einstein said so. Popper contrasted this with the theories he saw being applied around him. A Marxist historian could explain any political development; a psychoanalyst could explain a man who drowned a child and a man who saved one, using different mechanisms from the same theory. Their explanatory power looked like strength and was, Popper concluded, the opposite.
Key idea: A theory's scientific value lies in what it forbids, so a theory compatible with every possible observation tells you nothing about the world.
The asymmetry, stated exactly
Underneath the slogan is a point of elementary logic. Consider the universal statement that all swans are white.
- No finite number of white swans establishes it. The next one might be black, and one thousand observations leave that possibility exactly as open as ten did. This is Hume's problem from Lesson 2, in a specific form.
- One black swan refutes it, deductively and conclusively. The inference is modus tollens: if the theory is true then the observation must not occur; the observation occurred; therefore the theory is false.
So the two directions are not symmetric. Confirmation is inductive and, by Hume's argument, without justification. Refutation is deductive and needs no inductive principle at all. Popper's programme is to build a philosophy of science that uses only the deductive half.
This is a far better idea than it is usually given credit for. It takes Hume's problem entirely seriously, accepts the conclusion, and then asks what science could be if induction were unavailable. The answer: a sequence of conjectures, each held tentatively, each tested as severely as possible, each replaced when it fails.
What Popper actually recommends
Three prescriptions follow, and they are more specific than the slogan suggests.
Make bold conjectures. Do not aim for cautious hypotheses close to the data. Aim for theories that say a great deal, cover a wide range, and stick their necks out.
Prefer the more falsifiable theory. Popper measured a theory's merit by the size of the class of observations that would refute it. A theory forbidding more is better, because it tells you more.
Here is where he says something that startles most readers. A theory that forbids more is, for that very reason, less probable. The claim that all planetary orbits are ellipses is less probable than the claim that all planetary orbits are closed curves, since the first is a special case of the second and could fail where the second holds. Popper concluded that in science we should prefer the less probable hypothesis, which puts him in direct opposition to the confirmation theorists of Lesson 3 and to the Bayesians of Lesson 12. He was not confused; he was denying that high probability is what we should want.
Test severely and refuse ad hoc rescues. A severe test is one the theory would very likely fail if it were false. And when a prediction fails, you may modify the theory, but not in a way that merely blocks the refutation while forbidding nothing new. Adding an epicycle to save a theory reduces its content, which by Popper's own measure makes it worse.
Corroboration is not confirmation
What, then, do we say about a theory that has survived many severe tests? Popper's answer is deliberately austere. We say it is corroborated, and corroboration is a report on past performance. It is not a measure of probable truth, and it licenses no prediction that the theory will continue to succeed. To say otherwise would be to reintroduce induction, which the whole programme was built to avoid.
Popper also tried to say that some false theories are closer to the truth than others, defining a notion of verisimilitude in 1963. The attempt failed technically: in 1974 David Miller and Pavel Tichy showed independently that on Popper's definition no false theory can come out closer to the truth than any other false theory. Since all interesting scientific theories are presumably false, the definition delivered nothing. Later reconstructions exist, but the original is a genuine casualty and it should be reported as such.
The upshot: Popper's system buys freedom from Hume's problem at the price of being unable to say that a well-tested theory is likely to be true, and that price is what most of the objections press on.
The strongest case
Before the objections, state the position at its best, because it is easy to caricature.
- It takes the logic seriously. The asymmetry between verification and falsification is real, and no account of science can ignore it.
- It explains something confirmation theories struggle with: why a risky prediction that succeeds impresses us far more than a safe one. If your theory forbids most outcomes and the world delivers the one it permits, that is remarkable. Lesson 12 shows the Bayesian machinery agrees, which is a striking convergence between rivals.
- It gives a diagnosis of a real pathology. Theories that accommodate everything after the fact are common, and Popper explains exactly what is wrong with them.
- It is a genuinely non-inductive account of scientific progress, and it is the only serious one anybody has built.
- Its methodological advice is usable. Specify in advance what would count against you. That instruction, taken seriously, is most of what preregistration accomplishes, as Lesson 15 shows.
Five problems
| Objection | The difficulty |
|---|---|
| Nothing is tested alone | A prediction follows from the theory plus auxiliary assumptions, so a failure tells you the conjunction is false, not which conjunct. Lesson 5. |
| Probabilistic hypotheses | No finite sequence of outcomes contradicts a statistical claim, so rejection is a decision about thresholds, not a deduction. |
| Basic statements are fallible | The observation report that does the refuting is itself defeasible and theory-laden. Popper conceded this and said such statements are accepted by decision, like a jury's verdict, which makes the deductive purity partly conventional. |
| Scientists retain refuted theories, and are sometimes right | Newtonian mechanics faced an unexplained anomaly in Mercury's orbit for decades and was not abandoned. Retaining it was correct until a better theory arrived. |
| Corroboration cannot guide action | If it is rational to build a bridge on a well-corroborated theory rather than a refuted one, corroboration is being used as a guide to the future, which is induction. |
The last two are the heavy ones. Together they say that Popper's account describes a logic that no scientist could live by: it forbids the retention of anomalous theories that turned out to be worth retaining, and it forbids the reliance on tested theories that everyone including Popper relies on.
How Popperians answer
Three lines of reply are worth knowing, because none is foolish.
Methodological rules, not logic alone. Popper always insisted that falsification is a matter of decision as well as deduction. The scientist decides which statements to treat as basic, decides whether a modification is ad hoc, and decides when to abandon a theory. The rules are conventions adopted because they make criticism effective, not truths derived from logic. This concedes a good deal but preserves the core.
Move to programmes. Imre Lakatos accepted the historical objection and rebuilt falsificationism at the level of research programmes rather than individual theories, which is Lesson 7.
Formalize severity. Deborah Mayo's error-statistical philosophy keeps Popper's central intuition, that what matters is whether a claim has passed a test it would probably have failed if false, and gives it content using error probabilities from statistics. On this account corroboration is replaced by severe testing with a measurable meaning, and the vagueness in Popper's notion is what gets fixed. This is the most active living descendant of the programme.
What would settle it
Two questions do the work. First, can severity be given a measure that is not question-begging? If yes, the Popperian core survives in error-statistical form, and the objection that corroboration is empty loses its force. Mayo's programme is a serious attempt and its adequacy is genuinely disputed.
Second, does the history of science show scientists abandoning theories on refutation, or protecting them? The historical record is the evidence, and it does not favour naive falsificationism. Lesson 5 shows why the very same protective move was correct once and disastrous once, which is the deepest reason a purely logical criterion cannot be enough.
Common misconceptions
- Popper said a theory is scientific if it can be proved false. He said it is scientific if some observation statement is inconsistent with it. Whether that observation ever occurs, and whether we would accept it if it did, are separate matters.
- Popper thought one anomaly should kill a theory. He explicitly warned against treating isolated anomalies as decisive, and required reproducible falsifying effects.
- Popper was an inductivist about corroboration. He denied it repeatedly and consistently. Whether he could afford the denial is the objection, not a misreading.
- Falsificationism is the official method of science. It is one philosophical account. Scientists' actual practice includes a great deal of confirmation-seeking, and the two are not easily reconciled.
- Verisimilitude gives Popper a way to say science approaches truth. His 1963 definition was shown to fail in 1974. Later accounts exist but the original claim did not survive.
Pulling it together
- The 1919 eclipse impressed Popper because Einstein specified in advance an outcome that would refute him.
- The asymmetry is that universal statements cannot be verified by finitely many instances but can be refuted by one, using modus tollens.
- Popper recommends bold, highly falsifiable conjectures, severe tests, and a ban on modifications that only block refutation.
- He prefers less probable theories, because content and improbability go together, which puts him at odds with confirmation theory.
- Corroboration reports past performance and is not a probability of truth; the verisimilitude definition meant to fill the gap failed in 1974.
- The five main objections concern auxiliary assumptions, statistical hypotheses, fallible basic statements, the rational retention of anomalous theories, and the guidance of action.
- Popperians reply with methodological conventions, with Lakatos's programmes, and with Mayo's error-statistical account of severity.
Sources
- Thornton, S. (n.d.). Karl Popper. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Thornton, S. (n.d.). Karl Popper: philosophy of science. Internet Encyclopedia of Philosophy. iep.utm.edu
- Wikipedia contributors. (n.d.). Falsifiability. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Eddington experiment. Wikipedia. en.wikipedia.org
- Popper, K. (1959). The logic of scientific discovery. Hutchinson. (German original 1934.)
- Mayo, D. G. (1996). Error and the growth of experimental knowledge. University of Chicago Press.
- Key terms
- Logical asymmetry
- The fact that a universal statement can be conclusively refuted by a single counterinstance but never conclusively verified by finitely many instances.
- Modus tollens
- The valid deductive form used in refutation: if the theory then the prediction; not the prediction; therefore not the theory.
- Bold conjecture
- A hypothesis with high content that forbids many possible observations, preferred by Popper precisely because it risks more.
- Severe test
- A test the theory would very probably fail if it were false, which is what makes surviving it significant.
- Ad hoc modification
- A change made only to block a refutation, which reduces the theory's content and forbids nothing new.
- Corroboration
- Popper's term for a report on how a theory has withstood tests so far, explicitly not a measure of probable future success.
- Basic statement
- A singular observation report used to refute a theory, which Popper acknowledged is accepted by decision rather than given by experience.
- Verisimilitude
- Closeness to the truth, which Popper attempted to define formally in 1963 and which Miller and Tichy showed the definition could not deliver.
The Duhem-Quine Thesis: Nothing Faces the Evidence Alone
- Use the Neptune and Vulcan episodes to show that the same methodological move can be right once and wrong once.
- Distinguish Duhem's thesis about physics from Quine's global holism and say what each licenses.
- Explain why crucial experiments are not decisive, and evaluate the proposals for constraining where blame is placed.
A planet found within one degree of a prediction
On the night of 23 September 1846, Johann Gottfried Galle at the Berlin Observatory pointed his telescope at a set of coordinates that had arrived in the post from Paris. Within an hour he and his assistant Heinrich d'Arrest had identified an object that was not on their star chart. It was less than a degree from where Urbain Le Verrier had said it would be.
Le Verrier had not seen anything. He had worked from a problem: Uranus, discovered in 1781, was not where Newtonian calculation said it should be. Its orbit was off by amounts too large to blame on measurement error. In the terms of the last lesson, Newtonian gravitation had made a prediction and the prediction had failed.
Le Verrier did not abandon Newtonian gravitation. He kept the theory and questioned an assumption instead, namely that the solar system contained only the known planets. If an eighth planet lay beyond Uranus with the right mass and orbit, the discrepancy would vanish. He calculated where it would have to be, and it was there.
Key idea: When a prediction fails, the theory and its auxiliary assumptions are both in the dock, and the scientist has to decide which to blame. Le Verrier blamed an assumption about the contents of the solar system, and was spectacularly right.
The same move, made twice
Thirteen years later Le Verrier turned to another anomaly. Mercury's perihelion, the point of closest approach to the sun, advances slightly with each orbit. Newtonian calculation accounts for most of that advance, but about 43 seconds of arc per century were left over. In 1859 Le Verrier announced the discrepancy and proposed the same solution that had worked before: an undiscovered planet, this time inside Mercury's orbit. It was named Vulcan.
Vulcan was not there. Astronomers searched for over half a century. Occasional sightings were reported and none survived checking. In 1915 Einstein applied his newly completed general theory of relativity to Mercury and derived the missing 43 seconds of arc, and the anomaly was resolved by abandoning Newtonian gravitation rather than by adding a planet.
| Uranus, 1846 | Mercury, 1859 onward | |
|---|---|---|
| Anomaly | Orbit deviates from prediction | Perihelion advances 43 arcseconds per century too fast |
| Response | Keep the theory, posit a new planet | Keep the theory, posit a new planet |
| Outcome | Neptune found in 1846 | Vulcan never found |
| Correct diagnosis | An auxiliary assumption was wrong | The theory was wrong |
Read the middle row. The methodological move is identical. In one case it produced one of the great triumphs of nineteenth-century science and in the other a half-century of wasted searching. And nothing available at the time of the decision distinguished them. This is the problem, made as concrete as it can be made.
Duhem's thesis
Pierre Duhem, a physicist and historian of science, set out the general point in 1906 in The Aim and Structure of Physical Theory. A physicist, he argued, never submits a single hypothesis to the test of experiment. What is tested is the hypothesis together with the entire theoretical apparatus needed to derive an observable consequence from it: the auxiliary hypotheses, the theory of the instruments, the assumptions about the conditions.
When the prediction fails, logic tells you that the conjunction is false. It does not tell you which conjunct. The experiment, in Duhem's phrase, tells you that something in the group is wrong, and leaves you to find out what.
Duhem drew a specific and much-cited consequence about the experimentum crucis, the crucial experiment that was supposed to decide between two rival theories. He examined Foucault's 1850 measurement of the speed of light in water, widely presented as deciding between the emission theory of light and the wave theory. Duhem argued that it refuted the emission theory only in conjunction with a set of further assumptions, and that even a decisive refutation would not establish the wave theory, since the two theories are not the only possibilities. A crucial experiment would have to be a disjunctive syllogism over an exhaustive list of rivals, and nobody ever has one.
Two limits Duhem himself placed on the thesis are often forgotten. He restricted it to physics, explicitly saying that a physiologist or a chemist can sometimes isolate a hypothesis in a way a physicist cannot. And he held that the decision about where to place blame is made by what he called good sense, a trained judgment that is not reducible to a rule but is not arbitrary either.
Quine's stronger version
In 1951 W. V. O. Quine published Two Dogmas of Empiricism, which took the point much further. Our beliefs, he argued, form a web that meets experience only at its edges. A conflict at the periphery can be accommodated by revising almost anything in the interior, and the choice of what to revise is made on grounds of conservatism and simplicity rather than logic.
Two claims follow that Duhem never made. First, any statement can be held true come what may, provided we make drastic enough adjustments elsewhere in the system. Second, no statement is immune from revision, and Quine included the laws of logic, suggesting that a revision of the principle of the excluded middle had been proposed as a way of simplifying quantum mechanics.
So the two theses differ in scope and in kind. Duhem's is a claim about confirmation in physics: hypotheses are tested in groups. Quine's is a claim about all of knowledge, with a semantic component: meanings are not distributed sentence by sentence, so there is no principled line between revising a belief and changing the subject.
What matters here: The label Duhem-Quine covers two theses of very different strength, and arguments that work against the global version often leave the local version untouched.
What follows
- Falsification is never conclusive. This is the objection to Popper from the last lesson, now with a name and a mechanism. Modus tollens is valid, but its conclusion is the negation of a conjunction.
- There are no crucial experiments in the strict sense. Experiments can be decisive in practice, and the eclipse of 1919 was, but their decisiveness rests on judgments about which auxiliaries are secure, not on logic alone.
- Theory is underdetermined by evidence. For any body of observations there are, in principle, incompatible theories consistent with them. Whether this is a live problem or a philosopher's abstraction is the disputed part.
What stops it collapsing into anything goes
Taken at full strength, the thesis seems to say that you can always save any theory, which would make evidence powerless. Everybody, including Quine, thought that conclusion wrong. Three constraints are offered.
Duhem's good sense. Trained physicists agree, more often than not, about where the blame lies. This is descriptively true and explanatorily thin, but it is not nothing: a community with shared standards can converge without a rule.
Bayesian blame assignment. This is the most precise proposal, and it is easy to see once stated. Suppose your prediction rests on a theory you assign probability 0.9 and an auxiliary assumption you assign probability 0.99, and the prediction fails. The posterior blame falls in proportion to the prior improbability of each component: the component you were least confident in takes most of the hit. Le Verrier's confidence in Newtonian gravitation was, in 1846, far higher than his confidence that the solar system's inventory was complete, so blaming the inventory was rational. It was still rational in 1859, and it was wrong. Rationality does not guarantee truth, which is exactly the right thing for the model to say.
Empirical equivalence is not permanent. Larry Laudan and Jarrett Leplin argued in 1991 that two theories empirically equivalent today may not be tomorrow, because what counts as an observable consequence depends on auxiliary theories that themselves change. Moreover evidence can bear on a hypothesis indirectly, through a broader theory that entails it. So underdetermination is a much weaker constraint than it looks.
What would settle it
The interesting question is not whether the logical point holds; it plainly does. It is whether the practical constraints are strong enough that the logical point rarely matters. That is an empirical question about the history of science, and the evidence is mixed in an informative way. Blame is usually assigned in a manner that later evidence vindicates, which supports the constraints. But the Vulcan case shows that a perfectly reasonable assignment can be wrong for fifty years, which supports the sceptic. What would settle it is a systematic study of anomalies over a long period, asking how often the initial blame assignment survived. Some historians have begun exactly that work, and it is a place where the argument can still be advanced by evidence rather than by cleverness.
Common misconceptions
- Duhem and Quine argued for the same thesis. Duhem's applies to physics and to auxiliary hypotheses; Quine's applies to all of knowledge including logic, and carries a semantic claim Duhem would have rejected.
- The thesis shows that evidence cannot refute anything. It shows that evidence refutes conjunctions rather than isolated statements. Which conjunct to give up is then decided by considerations that are defeasible but not arbitrary.
- Any theory can be saved, so all theories are equally good. Saving a theory has costs: implausible auxiliaries, lost content, and a growing list of unexplained coincidences. Lesson 7 makes these costs explicit.
- The 1919 eclipse was a crucial experiment. It was decisive in practice and rightly persuaded people, but it rested on trusting the instruments, the reduction procedures and the star positions, all of which were questioned at the time.
- Underdetermination is merely hypothetical. Real cases exist, including empirically equivalent formulations in physics, though whether they are the sort of rivals that matter is contested.
What to remember
- Le Verrier responded to two anomalies with the same move, positing an unseen planet; it produced Neptune in 1846 and nothing at all after 1859.
- Duhem argued that a physicist tests a hypothesis only together with a whole group of auxiliaries, so a failed prediction convicts the group.
- He concluded there are no strictly crucial experiments, since refuting one theory does not establish its rival unless the alternatives are exhaustive.
- Quine generalized the point to all of knowledge, holding that any statement can be retained given enough adjustment elsewhere and none is immune from revision.
- The two theses differ in scope, and objections to the global version often leave the local one standing.
- Blame assignment is constrained by trained judgment, by prior probabilities that make the least secure component the likeliest culprit, and by the impermanence of empirical equivalence.
- The Vulcan episode shows that a rational blame assignment can be sustained for decades and still be wrong, which is the strongest evidence for the sceptical reading.
Sources
- Needham, P. (n.d.). Pierre Duhem. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Stanford, K. (n.d.). Underdetermination of scientific theory. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Hylton, P., & Kemp, G. (n.d.). Willard Van Orman Quine. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Wikipedia contributors. (n.d.). Discovery of Neptune. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Vulcan (hypothetical planet). Wikipedia. en.wikipedia.org
- Duhem, P. (1954). The aim and structure of physical theory (P. P. Wiener, Trans.). Princeton University Press. (Original work published 1906.)
- Key terms
- Auxiliary hypothesis
- An assumption needed alongside a theory to derive an observable prediction, such as a claim about instruments or about background conditions.
- Duhem's thesis
- The claim that a physicist can never test an isolated hypothesis, only a hypothesis together with a whole group of assumptions.
- Quinean holism
- The stronger view that all of knowledge faces experience as a corporate body, so any statement may be retained given enough revision elsewhere.
- Crucial experiment
- An experiment supposed to decide conclusively between two rival theories, which Duhem argued cannot exist unless the rivals are exhaustive.
- Underdetermination
- The situation in which more than one incompatible theory is consistent with all the available evidence.
- Good sense
- Duhem's name for the trained judgment by which physicists locate the blame for a failed prediction, not reducible to a rule.
- Blame assignment
- The decision about which component of a tested conjunction to reject when a prediction fails.
- Empirical equivalence
- The relation between theories with the same observable consequences, which Laudan and Leplin argued is neither permanent nor evidentially decisive.
Kuhn: Paradigms, and What He Did Not Claim
- Describe the cycle of normal science, anomaly, crisis and revolution, and the two senses of paradigm Kuhn later distinguished.
- Explain semantic, methodological and observational incommensurability, and Kuhn's later restriction of the thesis.
- Separate Kuhn's actual claims about rationality and progress from the relativist reading, and assess the criticisms that stuck.
A physicist reads Aristotle and finds him absurd
In the summer of 1947 Thomas Kuhn was a graduate student in theoretical physics at Harvard, asked by James Conant to help teach a course in the history of science for non-scientists. Preparing a case study on the origins of seventeenth-century mechanics, he sat down with Aristotle's Physics.
What he found appeared to be nonsense. Aristotle's statements about motion were not merely superseded; read as physics they were obviously and elementarily wrong, wrong in ways that a competent observer could not have failed to notice. Kuhn could not reconcile this with Aristotle's evident brilliance elsewhere.
Then, as he described it, the text suddenly cohered. Aristotle was not doing bad Newtonian mechanics. He was doing something else, with a different concept of motion, one that covered the growth of an oak from an acorn as readily as a stone falling, and within which his claims were reasonable and often perceptive. What had made Aristotle look stupid was reading him with Newtonian expectations.
That experience is the seed of The Structure of Scientific Revolutions, published in 1962. There is an irony in its publication history worth noticing: it appeared as a volume in the International Encyclopedia of Unified Science, the flagship project of the logical positivists whose picture of science it did most to dismantle.
Key idea: An earlier scientific tradition can look absurd when read through the concepts of a later one, and the appearance of absurdity is often a sign that the reader has imported the wrong framework rather than that the earlier thinker was careless.
Normal science as puzzle-solving
Kuhn's most durable contribution is his description of ordinary scientific work, which philosophers before him had largely ignored in favour of theory choice at moments of crisis.
Most scientists, most of the time, are not testing fundamental theories. They are doing what Kuhn called normal science: extending the reach of an accepted framework, measuring its constants more precisely, applying it to new cases, and cleaning up the discrepancies. The work has the character of puzzle-solving, in a specific sense. A puzzle, unlike a problem, comes with the guarantee that a solution exists and with rules about what counts as one. A crossword clue that cannot be solved reflects on the solver, not on the crossword.
That structure is what makes normal science productive. Because the framework is not in question, effort can be concentrated on detail, and instruments and techniques can be refined to a precision that would be pointless if the whole enterprise were up for grabs. Kuhn's provocative claim was that this dogmatism is not a defect but the condition of scientific depth. A field that reopens its foundations every year never gets past its foundations.
It also has a consequence Kuhn drew out carefully. When a normal-science puzzle resists solution, the first assumption is that the scientist has gone wrong, not the framework. That assumption is usually correct, which is precisely why anomalies have to accumulate before anyone concludes otherwise.
Paradigm: the word that caused the trouble
Kuhn's term for the framework was paradigm, and it became one of the most used and least precise words in intellectual life. The criticism was fair. Margaret Masterman went through Structure and counted twenty-one distinct senses in which Kuhn used it.
In the 1969 postscript to the second edition, Kuhn distinguished two, and the distinction is genuinely clarifying.
| Sense | What it covers |
|---|---|
| Disciplinary matrix | The whole set of shared commitments of a community: its symbolic generalizations, its preferred models, its values, and its standards for what counts as a solution |
| Exemplar | A concrete problem solution, of the kind students work through in textbooks, which teaches by pattern rather than by rule how to apply the theory to new cases |
The second is the more original idea and the more neglected one. Kuhn's claim is that scientists do not learn a theory as a set of explicit rules and then apply them. They learn a stock of worked examples and acquire the ability to see new situations as similar to one of them. This is why textbook problem sets matter, and why two physicists can agree on all the equations and still disagree about how to model a case. It is also, Kuhn thought, why paradigm shifts are hard to articulate: much of what changes was never written down as a rule.
Anomaly, crisis, revolution
The historical cycle runs as follows. A mature field practises normal science within a paradigm. Anomalies accumulate: results that resist assimilation despite serious effort by competent people. Most anomalies are eventually absorbed and are not crises. But when an anomaly touches the framework's central commitments, or resists for long enough, or blocks an application the community cares about, confidence erodes and the field enters a crisis.
In crisis, the rules loosen. Scientists become willing to consider alternatives they would previously have dismissed, foundational debate revives, and philosophy of science reappears in the journals, which Kuhn regarded as a diagnostic sign. A revolution occurs when a new candidate framework attracts enough of the community to become the basis of a new normal science.
Kuhn stressed that the transition is not a straightforward matter of the new theory explaining more. The new paradigm typically solves the crisis-provoking anomaly but loses some problems the old one had solved, which he called Kuhn loss. Early Copernican astronomy was not more accurate than Ptolemaic astronomy at predicting planetary positions. Oxygen chemistry initially could not explain some things phlogiston chemistry had handled. Choosing the new framework therefore involves judging that its promise outweighs its present deficits, which is a judgment and not a calculation.
The upshot: Because a new paradigm typically loses solved problems as well as gaining them, theory choice at a revolution cannot be settled by counting successes, which is the fact that generates everything controversial in Kuhn.
Incommensurability, in three senses
The most contested idea in the book is that competing paradigms are incommensurable: there is no common measure by which to compare them directly. Kuhn meant at least three things, and separating them defuses much of the confusion.
- Semantic. Terms shift meaning across the divide. Mass in Newtonian mechanics is conserved and is not interconvertible with energy; mass in relativistic mechanics is neither. The word survives; the concept does not. Similarly, whether the earth is a planet is not a question with one answer across the Copernican divide, because the term itself was redrawn.
- Methodological. Paradigms carry their own standards for what counts as a legitimate problem and an acceptable solution. Newtonian physicists rejected action at a distance as occult until they accepted it; the standards themselves were among the things that changed.
- Observational. What scientists see is shaped by training. Kuhn used the gestalt-switch analogy, and drew on Hanson's work on the theory-ladenness of observation, to argue that there is no theory-neutral observation language in which the two paradigms could be compared.
Kuhn spent much of the following thirty years narrowing the thesis. By the 1980s he defended only local incommensurability: a small cluster of interdefined terms fails to translate, while the rest of the two languages translates fine. That is a far weaker and far more defensible claim, and it is the one his mature work defends.
What Kuhn did not claim
Structure was widely read as saying that science is irrational, that paradigm change is a conversion experience like a religious one, and that truth is whatever a community agrees on. Kuhn spent the rest of his life denying all three, with visible frustration. What he actually held is more interesting.
On rationality: in a 1973 lecture published as Objectivity, Value Judgment, and Theory Choice, Kuhn listed five criteria that scientists across paradigms actually share, namely accuracy, consistency, scope, simplicity and fruitfulness. His claim was not that these are absent but that they function as values rather than as rules. They can conflict, they can be weighted differently by reasonable people, and they do not jointly determine a unique choice. So two competent scientists can disagree about which theory to back without either being irrational. That is a claim about the underdetermination of choice by shared standards, not a denial of standards.
On progress: Kuhn held firmly that science progresses, and that later paradigms are better puzzle-solvers than earlier ones. What he denied was progress toward a goal, toward a true account of what nature is really like. His analogy was evolution: a development from, driven by problems, rather than a development toward, drawn by a target. Whether that distinction is coherent is a fair question, and Lesson 11 takes it up as the realism debate.
On reality: his remarks that after a revolution scientists work in a different world were about practice and experience, not metaphysics. He was not claiming that a change of theory rearranges the furniture of the universe.
The criticisms that stuck
- The history is too tidy. Six decades of subsequent historical work has found far more continuity across supposed revolutions than Structure suggests, and few episodes fit the crisis-to-revolution pattern cleanly.
- Normal and revolutionary science are not cleanly separable. Most mature fields contain competing frameworks continuously rather than in punctuated bursts.
- The term is too elastic. Masterman's twenty-one senses were a real problem, and the 1969 clarification came seven years after the damage.
- Incommensurability was overstated. Historians reconstruct superseded theories all the time, and Kuhn's own recovery of Aristotle is the clearest example. If comparison were impossible, that achievement would be impossible too.
- The model fits the social sciences badly, and they adopted it hardest. Kuhn thought the social sciences pre-paradigmatic, which is not how they used the word.
What would settle it
Kuhn's claims are historical, so history is the evidence, and this is one of the few disputes in this course where that is straightforwardly true. The question is whether detailed studies of episodes find the pattern Structure describes. The verdict of the last sixty years is a partial one: the description of normal science and the role of exemplars have held up very well and are now common ground, while the sharp punctuated cycle and strong incommensurability have not. That is an ordinary scientific outcome for a theory, which Kuhn would probably have enjoyed.
Common misconceptions
- Kuhn said theory choice is irrational. He said it is not determined by shared rules, and identified five shared values that reasonable scientists weigh differently.
- Kuhn said there is no scientific progress. He said there is progress in puzzle-solving power, and denied only progress toward a final true account.
- A paradigm shift is any change of opinion. In Kuhn's sense it is the replacement of a community's whole disciplinary matrix, which happens rarely and slowly.
- Incommensurable means incomparable. It means there is no common measure, not that no comparison is possible. Kuhn's mature position restricts the failure of translation to a few interdefined terms.
- Kuhn thought scientists should be open-minded about foundations. He argued the opposite: the dogmatism of normal science is what makes precision and depth possible.
Summing up
- Kuhn's project began with the discovery that Aristotle's physics reads as nonsense only if you approach it with Newtonian expectations.
- Normal science is puzzle-solving within an unquestioned framework, and its dogmatism is what allows precision and depth.
- Paradigm covers both the disciplinary matrix of shared commitments and the exemplars, the worked problems by which scientists learn to apply a theory.
- Anomalies accumulate into crisis and occasionally revolution, and the new framework usually loses some problems the old one solved.
- Incommensurability comes in semantic, methodological and observational forms, and Kuhn's mature position restricts it to a few interdefined terms.
- Kuhn affirmed rational theory choice governed by shared values that underdetermine the outcome, and affirmed progress in puzzle-solving while denying progress toward a final truth.
- Later historical work supports his account of normal science and exemplars and does not support the sharp revolutionary cycle or strong incommensurability.
Sources
- Bird, A. (n.d.). Thomas Kuhn. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Internet Encyclopedia of Philosophy. (n.d.). Thomas Kuhn. iep.utm.edu
- Oberheim, E., & Hoyningen-Huene, P. (n.d.). The incommensurability of scientific theories. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Wikipedia contributors. (n.d.). The Structure of Scientific Revolutions. Wikipedia. en.wikipedia.org
- Kuhn, T. S. (1962). The structure of scientific revolutions. University of Chicago Press.
- Kuhn, T. S. (1977). Objectivity, value judgment, and theory choice. In The essential tension. University of Chicago Press.
- Key terms
- Normal science
- Research conducted within an accepted framework whose foundations are not in question, having the character of puzzle-solving.
- Puzzle
- A problem that comes with the assurance that a solution exists and with shared rules about what counts as one, so failure reflects on the solver.
- Disciplinary matrix
- The full set of commitments shared by a scientific community: generalizations, models, values and standards of solution.
- Exemplar
- A concrete worked problem that teaches scientists, by pattern rather than rule, how to apply a theory to new situations.
- Anomaly
- A result that resists assimilation into the paradigm despite competent effort, and whose accumulation can produce a crisis.
- Kuhn loss
- The problems a superseded paradigm had solved that its successor initially cannot, which is why revolutions are not simple gains.
- Incommensurability
- The absence of a common measure for comparing paradigms, arising from shifts in meaning, in standards and in trained perception.
- Local incommensurability
- Kuhn's later, weaker thesis that only a small cluster of interdefined terms fails to translate across a revolution.
Module 3: After Kuhn: Rationality and Method
Two responses to the historical challenge: Lakatos's attempt to rescue a rational criterion by shifting the unit of appraisal, and Feyerabend's argument that no universal methodological rule survives contact with the history of science.
Lakatos: Research Programmes and What Makes One Degenerate
- Describe a research programme in terms of hard core, protective belt, and negative and positive heuristics.
- Apply the progressive and degenerating distinction to the Neptune and Vulcan episodes.
- Assess the objections that novelty is ill-defined and that appraisal only works in hindsight.
A colloquium in London, 1965
In July 1965, at Bedford College in London, an International Colloquium in the Philosophy of Science brought Popper's followers and Thomas Kuhn into the same room. The papers and replies appeared five years later as Criticism and the Growth of Knowledge, edited by Imre Lakatos and Alan Musgrave, and the longest contribution in it was Lakatos's own: Falsification and the Methodology of Scientific Research Programmes.
Lakatos had a specific problem to solve. He accepted the historical evidence Kuhn had marshalled: scientists do not abandon theories when a prediction fails, and they are frequently right not to. He also refused Kuhn's conclusion that theory choice cannot be governed by a criterion. His question was whether falsificationism could be rebuilt so that it fitted the history without giving up the idea that some choices are rationally better than others.
His answer was to change what gets appraised.
Key idea: The unit of scientific appraisal is not a single theory tested against a single experiment, but a sequence of theories developing over time, which Lakatos called a research programme.
The anatomy of a programme
A research programme has four parts, and the structure is worth learning because it maps onto real scientific practice more closely than any of the accounts so far.
| Component | What it is | In Newtonian astronomy |
|---|---|---|
| Hard core | The fundamental assumptions the programme is defined by, treated as irrefutable by decision | The three laws of motion and the law of universal gravitation |
| Protective belt | Auxiliary hypotheses, initial conditions and observational assumptions, which absorb the anomalies | The number and masses of the planets, the theory of the telescope, assumptions about refraction |
| Negative heuristic | The instruction not to direct refutations at the core | When a prediction fails, do not question the inverse square law first |
| Positive heuristic | A partly articulated plan for developing the belt, so anomalies are anticipated rather than merely absorbed | Model the planets as point masses, then add oblateness, then add perturbations from other planets, and so on |
The positive heuristic is the piece most often skipped, and it is the one that answers the obvious complaint. If the core is protected by decision, is this not just licensed dogmatism? Lakatos's reply is that a healthy programme has a research plan laid out in advance, so its modifications are not reactions to embarrassment but steps in a programme its practitioners could have described before the anomaly arose. A theorist working in a good programme knows what the next three models look like before the next three anomalies appear.
Progressive and degenerating, defined precisely
The criterion has two levels, and keeping them apart is essential.
- A programme is theoretically progressive if each new version predicts some novel fact that its predecessor did not predict. That is a condition on the theory alone, checkable as soon as the modification is written down.
- It is empirically progressive if some of those novel predictions are subsequently corroborated.
- It is degenerating if its modifications only accommodate facts already known, adding nothing that could be checked independently.
Notice what this does to the Duhem-Quine problem of Lesson 5. That lesson showed that a theory can always be saved by adjusting the auxiliaries, and asked what stops science from becoming a game of endless rescues. Lakatos gives an answer: rescues are always possible, but not all rescues are equal. A rescue that predicts something new is a contribution; a rescue that only blocks the anomaly is a debt. The distinction is not about whether you may modify a theory but about what the modification is required to earn.
Lakatos also revised what falsification means. On his account a theory is not falsified by a recalcitrant experiment. It is falsified when a rival exists that explains everything the old one did, predicts novel facts besides, and has some of those predictions corroborated. Refutation is thus a three-cornered contest between two theories and the evidence, never a duel between one theory and a fact.
What matters here: A modification earns its place by forbidding something new, and a programme whose modifications only ever explain what is already known is paying out without taking in.
Neptune and Vulcan, appraised
Return to Lesson 5's pair, which is Lakatos's own worked example.
Uranus, 1846. The anomaly threatens the core. The negative heuristic forbids blaming gravitation, so Le Verrier modifies the belt: there is an eighth planet. This modification is theoretically progressive, and dramatically so. It does not merely absorb the anomaly; it predicts a novel fact, namely that a telescope pointed at a specific set of coordinates on a specific night will show a previously uncharted object. That prediction was corroborated within an hour of being tested. Theoretically and empirically progressive.
Mercury, from 1859. The identical move. Le Verrier modifies the belt: there is an intra-Mercurial planet. This is also theoretically progressive, for exactly the same reason. It predicts a novel and checkable fact. The difference appears only afterwards: the prediction was not corroborated, and as searches failed the modifications required to keep Vulcan alive became increasingly specific and increasingly unable to generate new predictions. It might be too small to see, too close to the sun, a ring of asteroids rather than a planet. Each version explained the failure to find Vulcan and predicted nothing further. That sequence is what degeneration looks like.
This appraisal has a property worth appreciating. It gets both cases right without hindsight bias about the outcome, because at the moment of the first modification both were progressive, and Lakatos says exactly that. The programmes diverge only later, and only through the pattern of their subsequent modifications. That is a genuine advance over saying that Le Verrier was brilliant once and foolish once.
What the account buys
- It fits the history. Scientists really do protect fundamental assumptions and really do modify auxiliaries, and Lakatos explains why this is reasonable rather than deplorable.
- It rescues a rational criterion from Kuhn's challenge without denying Kuhn's evidence.
- It explains the difference between a good rescue and a bad one, which naive falsificationism could not, and which the Duhem-Quine thesis left open.
- It licenses a scientist to continue in a temporarily unsuccessful programme, which is descriptively accurate and which naive falsificationism forbids.
Three objections
What counts as a novel fact? If a fact was already known when the modification was made, is it novel? If not, then general relativity's account of Mercury's perihelion, which was known for fifty-six years before Einstein derived it, would not count in its favour, which is absurd. Elie Zahar proposed a repair in 1973: a fact is novel relative to a hypothesis if it was not used in constructing that hypothesis. On this use-novelty criterion Mercury counts, because Einstein did not build the field equations to fit it. The repair is widely adopted and is still argued about, since it makes novelty depend on the psychology and history of the theory's construction rather than on its content.
Appraisal comes too late. The criterion tells you a programme was degenerating over the past decade. It does not tell a scientist today which programme to join. Lakatos accepted this and said explicitly that there is no instant rationality, and that it can be rational to work in a degenerating programme provided one is honest about its condition. That is a defensible position and it is also an admission that the theory offers appraisal rather than advice.
The hard core is drawn retrospectively. Scientists do not publish a list of the propositions they have decided to protect. The historian identifies the core by seeing what was in fact never revised, which risks making the account unfalsifiable in the way Lakatos himself objected to elsewhere.
Feyerabend, who was Lakatos's close friend and sharpest critic, pressed the second objection into a general charge: since Lakatos sets no time limit on how long a degenerating programme may be pursued, his methodology forbids nothing, and is anarchism wearing the costume of rationality. Lesson 8 takes that argument seriously.
What would settle it
The central question is whether the progressive and degenerating distinction can be applied prospectively, with agreement among people who do not already know the outcome. That is empirically testable and has not been tested systematically. One could take a set of programmes at a chosen date, have historians blind to the outcome classify them, and check the classifications against what happened. Nothing like this has been done at scale, and it is a genuine opportunity rather than a rhetorical flourish.
A second question is whether use-novelty is the right criterion. If novelty depends on how a theory was constructed, then two theories with identical content can differ in their support, which strikes many as the wrong result. Lesson 12 shows that Bayesian confirmation theory has its own version of this difficulty, called the problem of old evidence, and the two problems are recognizably the same problem.
Common misconceptions
- The hard core is unfalsifiable in principle. It is protected by a methodological decision that can be revoked. A degenerating programme is abandoned core and all.
- Degenerating means false. It means the sequence of modifications has stopped generating checkable novelty. A degenerating programme can stage a recovery, and Lakatos allowed for it.
- Lakatos rejected Popper. He regarded himself as developing Popper, and the account keeps the emphasis on content, risk and novel prediction. What he abandoned was appraisal of single theories against single experiments.
- The criterion tells scientists what to do. It appraises what has been done. Lakatos said there is no instant rationality, which concedes the point.
- Protecting a core is dogmatism. It is only dogmatism if the positive heuristic is empty. A programme with a research plan is protecting a core in order to develop it, not to avoid testing.
The takeaway
- Lakatos shifted the unit of appraisal from a single theory to a research programme unfolding over time.
- A programme has a hard core protected by a negative heuristic, a protective belt of auxiliaries, and a positive heuristic that plans the belt's development.
- A programme is theoretically progressive if each version predicts novel facts, empirically progressive if some are corroborated, and degenerating if modifications only accommodate what is known.
- Falsification is reconstrued as a three-cornered contest: a theory falls when a better-performing rival exists, not when an experiment goes wrong.
- Neptune and Vulcan began as equally progressive modifications and diverged only through the later pattern of their revisions.
- The main objections concern the definition of novelty, answered by Zahar's use-novelty proposal, and the fact that appraisal is retrospective, which Lakatos conceded.
- Whether the distinction can be applied prospectively is an open and genuinely testable question.
Sources
- Musgrave, A., & Pigden, C. (n.d.). Imre Lakatos. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Niiniluoto, I. (n.d.). Scientific progress. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Wikipedia contributors. (n.d.). Imre Lakatos. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Vulcan (hypothetical planet). Wikipedia. en.wikipedia.org
- Lakatos, I. (1970). Falsification and the methodology of scientific research programmes. In I. Lakatos & A. Musgrave (Eds.), Criticism and the growth of knowledge. Cambridge University Press.
- Key terms
- Research programme
- A sequence of theories sharing a common core, which Lakatos took to be the proper unit of scientific appraisal.
- Hard core
- The fundamental assumptions that define a programme, treated as irrefutable by methodological decision rather than by logic.
- Protective belt
- The auxiliary hypotheses, initial conditions and observational assumptions that are modified when predictions fail.
- Negative heuristic
- The instruction not to direct refutations at the hard core of a programme.
- Positive heuristic
- A partly articulated plan for developing the protective belt, which distinguishes a research programme from mere evasion.
- Theoretically progressive
- The property of a programme whose each new version predicts some novel fact its predecessor did not.
- Empirically progressive
- The property of a programme some of whose novel predictions have been corroborated.
- Degenerating programme
- A programme whose modifications only accommodate facts already known and generate no checkable novelty.
- Use-novelty
- Zahar's criterion on which a fact counts as novel relative to a hypothesis if it was not used in constructing that hypothesis.
Feyerabend: Against Method, Taken as an Argument
- State Feyerabend's argument as premises and conclusion, and explain what anything goes is actually doing in it.
- Work the Galileo case as Feyerabend presents it, including the telescope and the tower argument.
- Identify the strongest objections and say which parts of the position survive them.
A professor who would not look
In 1610 Galileo Galilei was showing his telescope to colleagues in Padua. Cesare Cremonini, professor of natural philosophy at the university and a serious Aristotelian, declined to look through it. The episode has been retold for four centuries as a parable about closed minds.
Paul Feyerabend asked what Cremonini's reasons might have been, and the answer is more uncomfortable than the parable allows. There was in 1610 no accepted theory of the telescope. Nobody could explain from optical principles why a tube with two ground glass discs should produce a faithful image rather than an artefact. The instrument demonstrably produced spurious effects on terrestrial objects: haloes, colour fringes, doubled images. And the standard way of certifying an instrument, checking it against things you can inspect directly, could not be extended to the heavens, because the whole question at issue was whether celestial objects are the same kind of thing as terrestrial ones. Aristotelian physics said they are not. So a check on a distant church tower would not establish reliability for Jupiter.
Galileo was asking his contemporaries to accept the testimony of an uncertified instrument against the direct evidence of their eyes, on the strength of a theory the instrument was being used to support. He was right. He was also, by the methodological standards available at the time, doing something no rule of method would have licensed.
Key idea: The best-known triumph of early modern science involved violating the methodological standards of its day, and Feyerabend's whole argument is built from that observation repeated across cases.
The argument, as premises and a conclusion
Against Method appeared in 1975 and is usually remembered for a slogan. The slogan is the least interesting part, and misreading it wastes the book. Here is the structure.
- For any methodological rule that has been seriously proposed, there is an episode in the history of science in which violating that rule was necessary for progress.
- A rule that must be violated in order to make progress is not a universally binding rule.
- Therefore no proposed methodological rule is universally binding.
- Therefore the only principle that survives as universal is the empty one: anything goes.
Step 4 is not advice. Feyerabend was explicit that anything goes is not a principle he endorses, and that a scientist who tried to follow it would be paralysed. It is the conclusion a rationalist committed to universal, exceptionless rules is forced to accept, offered as the terminus of that commitment. The argument is a reductio directed at the demand for universal method, not a programme for research.
That reading matters, because on the popular reading the position is self-refuting nonsense and on the actual reading it is a serious historical challenge: name a rule, and be shown the case that breaks it.
Working the Galileo case
Feyerabend's central example is developed at length, and two strands of it are worth following.
The tower argument. Aristotelians argued against a moving earth as follows: drop a stone from a tower, and if the earth were moving beneath it the stone would land some distance to the west. It lands at the foot of the tower. Therefore the earth is at rest. This is a good argument, and it is empirical.
Galileo's reply required denying something that seemed to be given in observation: that the motion you see is all the motion there is. He introduced what Feyerabend calls a new natural interpretation, that motion shared by the observer and the observed is not perceptible, so the stone retains the earth's motion and the observer cannot see it. This is now obvious to us because we were taught it as children. In 1632 it was not derived from observation; it was introduced to save a theory, and it changed what the observation was taken to show.
Feyerabend's point is not that Galileo cheated. It is that the observation and the theory could not be separated: the observational evidence against Copernicus was evidence only under the old interpretation, and the new interpretation was adopted because the theory required it. Facts and theories are not independent enough for the simple picture of testing to apply.
Copernicanism was empirically worse. This is the part students find hardest to believe. In 1610 the Copernican system was, on several counts, in worse shape than the Ptolemaic. It predicted a stellar parallax, an annual apparent shift in the positions of fixed stars, which nobody could detect. Galileo's answer was that the stars must be enormously farther away than anyone had supposed, which was an unsupported auxiliary hypothesis introduced to protect the theory. Parallax was not actually measured until 1838. So for over two centuries the Copernican programme carried a known and unexplained observational failure.
Counterinduction and proliferation
The positive part of Feyerabend's position, and the part most often lost, is counterinduction: the recommendation to develop hypotheses that are inconsistent with well-confirmed theories and even with apparently established facts.
The reasoning is not perverse. If observation is theory-laden, then the assumptions built into the incumbent theory are invisible from inside it; they show up as facts. The only way to expose them is to construct an alternative in which they are not assumed, and see what the evidence looks like from there. Galileo could not have found the assumption about perceptible motion by examining it more carefully. He found it by needing an alternative.
From this comes the case for proliferation: a healthy science maintains multiple incompatible frameworks rather than converging early. This part of Feyerabend's position has been quietly absorbed into mainstream philosophy of science. It reappears in Lesson 13 as the argument for using multiple incompatible models of the same system, and in Lesson 15 as the argument for multi-laboratory replication with different methods.
The upshot: Counterinduction follows from theory-ladenness rather than from mischief: if the incumbent theory shapes what counts as a fact, only a rival can reveal what it has been assuming.
The political turn
Science in a Free Society, published in 1978, took the argument somewhere more provocative. If science has no unique method, Feyerabend argued, then its institutional privilege cannot be justified by appeal to method, and citizens in a democracy should have a say over what science does and what it teaches. He proposed a separation of science and state analogous to the separation of church and state.
It is important to be accurate about what he did and did not claim here. He did not claim that astrology works or that traditional medicine is as effective as modern medicine at what modern medicine does well. His famous discussion of a 1975 manifesto against astrology signed by 186 scientists attacks the manifesto's authoritarian tone and the signatories' evident unfamiliarity with what they were condemning, while agreeing that astrology is in poor intellectual condition. The target is the appeal to scientific authority as such, not the content of the science.
Whether the political conclusion follows from the methodological premises is a separate question, and most readers think it does not. Science could lack a unique method and still be, by a large margin, the most reliable available source of knowledge about the natural world, which would justify a good deal of privilege on straightforwardly consequentialist grounds.
Four objections
- The slide from no exceptionless rule to no rules. This is the strongest objection and it is decisive against the popular version. Rules with known exceptions guide behaviour perfectly well: prefer larger samples, control for confounds, do not analyze until the data are in. That each has been violated productively once does not make it worthless as a default.
- The history is contested. Historians of the Galileo affair have disputed Feyerabend's reconstruction on several points, including how much of the Copernican case rested on the telescope and how the tower argument was actually received.
- Self-undermining. The book argues by marshalling historical evidence and demanding consistency, which are exactly the standards it says are not binding. Feyerabend would say he is arguing on his opponents' terms, which is a fair reply but a limited one.
- The political conclusion does not follow. A defence of scientific authority based on its track record rather than on its method is untouched by the whole argument.
What would settle it, and what survives
The empirical part of Feyerabend's argument is genuinely testable: propose a rule and check the historical record for violations that were productive. That work has been done piecemeal and the results support premise 1 more often than methodologists would like. What it does not support is step 4, because premise 2 is doing illegitimate work: from the fact that a rule has exceptions it does not follow that it is not a rule, only that it is not exceptionless.
What survives is substantial: the theory-ladenness of observation, the case for proliferating incompatible theories, the demand that methodological prescriptions be checked against history rather than derived a priori, and a permanent caution against treating any codified method as a guarantee. That last point returns with force in Lesson 15, where a codified method followed conscientiously turned out to produce a large body of unreliable results.
Common misconceptions
- Feyerabend said all methods are equally good. He said no method is universally correct. Comparative judgments in particular cases are untouched.
- Anything goes is his advice to scientists. It is the conclusion he says a rationalist demanding universal rules must accept, presented as a reductio.
- He thought science is no better than astrology. He denied that science's authority rests on method, and criticized the manner in which scientists dismissed astrology, while agreeing that astrology is intellectually weak.
- He was hostile to science. He was trained in physics, wrote seriously about quantum mechanics, and directed his hostility at philosophers' reconstructions of scientific method rather than at scientific practice.
- Counterinduction means ignoring evidence. It means developing rivals to well-supported theories in order to expose what those theories are assuming, which is an argument for more evidence rather than less.
Looking back
- Cremonini's refusal to look through the telescope had reasons: there was no theory of the instrument and no way to certify it for celestial objects.
- Feyerabend's argument runs from the historical claim that every proposed rule has been productively violated to the conclusion that no rule is universally binding.
- Anything goes is the terminus of a demand for exceptionless rules, offered as a reductio rather than as advice.
- Galileo's reply to the tower argument required a new interpretation of what motion is observable, introduced to save the theory rather than derived from observation.
- Copernicanism carried a known observational failure, the missing stellar parallax, for more than two centuries.
- Counterinduction and proliferation follow from theory-ladenness and have been absorbed into mainstream practice.
- The decisive objection is the slide from rules having exceptions to there being no rules; defaults with known exceptions still guide research.
Sources
- Preston, J. (n.d.). Paul Feyerabend. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Bogen, J. (n.d.). Theory and observation in science. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Wikipedia contributors. (n.d.). Against Method. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Theory-ladenness. Wikipedia. en.wikipedia.org
- Feyerabend, P. (1975). Against method: Outline of an anarchistic theory of knowledge. New Left Books.
- Hanson, N. R. (1958). Patterns of discovery. Cambridge University Press.
- Key terms
- Epistemological anarchism
- Feyerabend's position that no methodological rule holds universally, since every proposed rule has been productively violated.
- Anything goes
- The empty principle Feyerabend says a rationalist demanding exceptionless rules must end with, offered as a reductio rather than as advice.
- Natural interpretation
- An assumption so embedded in how observation is described that it is mistaken for part of the observation itself.
- Counterinduction
- The recommendation to develop hypotheses inconsistent with well-confirmed theories, in order to expose what those theories assume.
- Proliferation
- The maintenance of multiple incompatible frameworks in a field, argued for as a condition of effective criticism.
- Theory-ladenness of observation
- The claim that what counts as an observation depends on the theoretical framework of the observer.
- Stellar parallax
- The apparent annual shift in stellar positions predicted by a moving earth, undetected until 1838 and a standing anomaly for Copernicanism.
- Tower argument
- The Aristotelian argument that a stone dropped from a tower lands at its base, which was taken to show that the earth does not move.
Module 4: Explanation, Law, and the World
What it takes to explain something rather than merely predict it, what makes a generalization a law rather than an accident, and whether the unobservable entities our best theories describe are real.
Explanation: From Covering Laws to Causes and Mechanisms
- State the four conditions of the deductive-nomological model and construct an explanation that satisfies them.
- Work the flagpole, barometer, contraceptive and paresis counterexamples and say what each shows the model is missing.
- Compare causal, interventionist, mechanistic and unificationist accounts against those cases.
Why the radiator cracked
A car is left outside overnight in a hard frost and in the morning the radiator has split. Carl Hempel used this case to show what an explanation is, and the reconstruction is worth working through in full because everything that follows is a reaction to it.
The explanation has two kinds of ingredient. There are particular facts: the radiator was filled with water, it was sealed, the ambient temperature fell from well above freezing to well below it during the night, the casing was of a certain material and thickness. And there are general laws: water expands as it freezes, the pressure exerted by ice in a closed container rises as temperature falls, a container fails when internal pressure exceeds the tensile strength of its walls.
Put the particular facts and the general laws together and the cracking of the radiator follows as a matter of deduction. Given that this was the situation and given that those are the laws, the radiator had to crack. Anyone who understood the explanation could have predicted the outcome the previous evening.
Key idea: On Hempel's account, to explain an event is to show that it was to be expected, by deducing it from statements of law together with statements of the particular conditions.
The covering-law model, stated as conditions
Hempel and Paul Oppenheim set the account out formally in 1948. The explanation divides into the explanandum, the statement of what is to be explained, and the explanans, the statements that do the explaining, and four conditions must hold.
- The explanandum must be a logical consequence of the explanans, so that the explanation is a valid deduction.
- The explanans must contain at least one general law, and that law must be essential to the derivation rather than an idle passenger.
- The explanans must have empirical content, meaning it must be capable of test.
- The statements of the explanans must be true.
The model has real virtues that get forgotten in the rush to its counterexamples. It is precise. It explains why explanation and prediction feel like the same activity pointed in opposite directions, which Hempel called the structural identity thesis. It makes clear why laws matter to explanation. And it says something true: a great many genuine explanations do have this form.
Hempel added a second model for cases where the laws are statistical. In inductive-statistical explanation the explanandum does not follow deductively but is made highly probable by the explanans. Why did the patient recover? Because he had a streptococcus infection, was given penicillin, and the great majority of such patients recover. The relation is inductive support rather than entailment.
That extension immediately raised a difficulty Hempel worked on for years: the same person can belong to two reference classes with different probabilities, one making recovery likely and another, say a penicillin-resistant strain, making it unlikely. Hempel's response was a requirement of maximal specificity: use the narrowest reference class about which you have relevant information. It patches the case and it concedes that the model needs a notion of relevance that the deductive machinery does not supply.
Four counterexamples
The model was taken apart over the following two decades by a series of cases, each of which satisfies all four conditions and is manifestly not an explanation.
| Case | The derivation | What goes wrong |
|---|---|---|
| The flagpole and the shadow | From the length of the shadow, the sun's elevation and the laws of optics, deduce the height of the pole | Valid, and satisfies every condition, but the shadow does not explain the pole. The pole explains the shadow. Deduction is symmetric; explanation is not. |
| The barometer and the storm | From the sharp fall in the barometer and a well-confirmed lawlike correlation, deduce that a storm is coming | The barometer does not explain the storm. Both are effects of a drop in atmospheric pressure, which explains both. |
| The contraceptive pills | John, a man, has taken his wife's oral contraceptives regularly, and no one who takes oral contraceptives becomes pregnant, so John did not become pregnant | Valid, lawlike, true, and absurd. The pills are explanatorily irrelevant to the outcome. |
| Paresis and syphilis | Why did this patient develop paresis? Because he had untreated latent syphilis | Only a small minority of untreated syphilitics develop paresis, so the explanans does not make the explanandum probable. Yet it is the only known cause and is a genuine explanation. |
The first three show the model is too permissive: it lets in derivations that are not explanations. The fourth shows it is also too restrictive: it excludes an explanation that everyone accepts, because it demands high probability where the world supplies low probability. Michael Scriven's paresis case is particularly awkward because there is nothing more to say about why this patient rather than another developed the condition. The explanation is correct and the probability is low, and any model requiring high probability has to call it a non-explanation.
What the counterexamples have in common
Look at the first three together and a pattern emerges. In each, the derivation runs in a direction that does not match the direction of dependence. The shadow depends on the pole, not the other way round. The storm and the barometer both depend on the pressure. John's non-pregnancy does not depend on the pills at all.
Hempel's model captures nomic expectability: given the explanans, the explanandum was to be expected. What it misses is that explanation tracks a relation of dependence with a direction, and logical consequence has no direction. That diagnosis points straight at the family of accounts that replaced it.
What matters here: Explanation is asymmetric and relevance-sensitive, and deductive consequence is neither, so no purely logical account of explanation can succeed.
Causal accounts, and the interventionist repair
Wesley Salmon proposed that explanation is a matter of exhibiting the causal structure that produced the event, first through a statistical relevance model and later through a causal-mechanical account built on causal processes and their interactions. David Lewis put it compactly: to explain an event is to give information about its causal history.
The most influential current version is James Woodward's interventionist account, set out in 2003. Its central idea is that explanation answers what-if-things-had-been-different questions. X explains Y when there is a possible intervention on X that would change Y, or change the probability distribution of Y, with everything else held fixed.
Watch what this does to the counterexamples.
- Flagpole. Intervene on the pole's height, by cutting it, and the shadow changes. Intervene on the shadow, by holding a screen so that it falls elsewhere, and the pole does not change. The asymmetry falls out of the account rather than being stipulated.
- Barometer. Intervene on the barometer, by tapping the glass or moving the needle, and no storm arrives. Intervene on atmospheric pressure and both change. The common cause is identified as the explainer.
- Contraceptives. Intervene on John's pill-taking, in either direction, and his pregnancy status does not change. So the pills do not explain.
- Paresis. Intervene to remove the syphilis and the paresis does not occur. Low probability is no obstacle, because the account asks about dependence rather than about expectability.
All four cases come out right, from one principle, without adjustment. That is as clean a success as this course will show you, and it is why interventionist accounts dominate the current literature on causal explanation.
Mechanisms
A parallel development came from philosophers looking at biology and neuroscience rather than physics. Peter Machamer, Lindley Darden and Carl Craver argued in 2000 that explanation in these fields typically works by describing a mechanism: a set of entities and activities organized so as to produce a regular change from set-up conditions to termination conditions.
Consider how a neuron fires. The explanation given by Alan Hodgkin and Andrew Huxley in 1952 does not subsume the action potential under a covering law. It describes components and what they do: voltage-gated sodium channels that open when the membrane depolarizes, sodium ions flowing in, potassium channels opening more slowly, potassium flowing out, the membrane repolarizing, channels inactivating. The explanation works by decomposing the system into parts and showing how their organized activities produce the phenomenon.
This matters for a reason connected to the next lesson. Biology has few if any exceptionless laws. If explanation required covering laws, most of biology would not explain anything, which is an unacceptable result. Mechanistic accounts say what biological explanation is actually doing, and they connect naturally to the interventionist account, since manipulating a component is exactly how mechanisms are discovered.
Unification, and the honest score
A rival tradition holds that explanation is unification. Michael Friedman in 1974 and Philip Kitcher in the 1980s argued that a theory explains by reducing the number of independent facts we have to accept: Newton explained by showing that falling apples, tides and planetary orbits are instances of one pattern rather than three brute facts. On this view the explanatory power of a theory is a global property of how much it systematizes, not a local relation between an event and its causes.
Unification captures something the causal accounts strain over. Some explanations do not seem to be causal at all. Why do cicadas of certain species have life cycles of thirteen and seventeen years? Because prime periods minimize overlap with the periods of predators and competitors, which is a mathematical fact about least common multiples, not a causal process. Why can no walk cross each of the seven bridges of Konigsberg exactly once? Because of a theorem about graphs. Intervening on nothing would change these, and yet they explain.
The honest score, then, is that no single account covers the ground. Interventionist accounts handle causal explanation and the classic counterexamples best. Mechanistic accounts describe what the life sciences actually do. Unificationist accounts capture structural and mathematical explanation. Most philosophers now accept some form of pluralism, which is a real conclusion rather than a failure of nerve, though it leaves open whether the varieties share anything beyond the name.
What would settle it
Two tests are doing the work. The first is the counterexample bench: any proposed account must handle the flagpole, the barometer, the contraceptives and paresis without special pleading. Interventionism passes; the covering-law model does not; unificationism handles the flagpole only with argument that many find strained.
The second is the non-causal cases. If an account of explanation as causal dependence cannot accommodate mathematical explanations, either it is incomplete or those cases are not really explanations. Deciding which requires looking at scientific practice: do working scientists treat the cicada answer as an explanation? They do, consistently, which is evidence that the causal account is incomplete rather than that the practice is confused.
Common misconceptions
- The covering-law model was simply wrong. It correctly describes many explanations and it identified a real feature, nomic expectability. Its failure is that expectability is not sufficient and, as paresis shows, not necessary.
- Explanation is just prediction run backwards. That was Hempel's structural identity thesis, and the barometer refutes it: the barometer predicts the storm superbly and explains nothing.
- Statistical explanations need high probability. Paresis shows otherwise. What matters is that the explanans made a difference, not that it made the outcome likely.
- A mechanism is just a chain of causes. Mechanistic accounts emphasize organization, including feedback and hierarchical arrangement, not merely a sequence.
- All explanation is causal. Mathematical and structural explanations are hard to fit into that mould, and scientists offer them freely.
Putting it together
- The covering-law model explains an event by deducing it from laws plus particular conditions, under four stated requirements.
- Inductive-statistical explanation extends this to probabilistic laws and needs a maximal specificity requirement to handle competing reference classes.
- The flagpole, barometer and contraceptive cases show the model is too permissive; paresis shows it is also too restrictive.
- The common diagnosis is that explanation tracks a directed relation of dependence, while logical consequence has no direction.
- Woodward's interventionist account resolves all four cases from a single principle about what changes under intervention.
- Mechanistic accounts describe explanation in biology and neuroscience as decomposition into organized entities and activities, as with the action potential.
- Unificationist accounts capture mathematical and structural explanation, which causal accounts struggle with, and most philosophers now accept some pluralism.
Sources
- Woodward, J., & Ross, L. (n.d.). Scientific explanation. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Craver, C., & Tabery, J. (n.d.). Mechanisms in science. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Internet Encyclopedia of Philosophy. (n.d.). Explanation in science. iep.utm.edu
- Hempel, C. G., & Oppenheim, P. (1948). Studies in the logic of explanation. Philosophy of Science, 15(2), 135-175. doi.org
- Machamer, P., Darden, L., & Craver, C. F. (2000). Thinking about mechanisms. Philosophy of Science, 67(1), 1-25. doi.org
- Woodward, J. (2003). Making things happen: A theory of causal explanation. Oxford University Press.
- Key terms
- Explanandum
- The statement describing what is to be explained.
- Explanans
- The statements that do the explaining, comprising laws and particular conditions.
- Deductive-nomological model
- Hempel and Oppenheim's account on which an explanation is a valid deduction of the explanandum from true statements including an essential law.
- Inductive-statistical explanation
- Hempel's extension covering probabilistic laws, in which the explanans confers high probability on the explanandum.
- Maximal specificity
- The requirement that a statistical explanation use the narrowest reference class about which relevant information is available.
- Explanatory asymmetry
- The fact that if A explains B then B does not explain A, although the corresponding derivations may run both ways.
- Explanatory irrelevance
- The defect of a derivation that is valid and lawlike but cites a factor that makes no difference to the outcome.
- Interventionist account
- Woodward's view that X explains Y if an intervention on X would change Y or its probability distribution, other things being held fixed.
- Mechanism
- A set of entities and activities organized so as to produce a regular change from set-up to termination conditions.
- Unificationist account
- The view that a theory explains by reducing the number of independent phenomena that must be accepted as brute.
Laws of Nature
- Distinguish lawlike from accidental generalizations using counterfactual support and projectibility.
- Compare the best-system analysis and the necessitarian account, and state the main objection to each.
- Assess Cartwright's argument that fundamental laws are false and the difficulties raised by ceteris paribus clauses.
Two spheres that do not exist
Consider two statements, both of which are true.
- No solid sphere of enriched uranium has a diameter of one mile.
- No solid sphere of gold has a diameter of one mile.
As far as anyone knows, both are true. There is no such sphere of either metal anywhere in the universe. And yet they are true in strikingly different ways, as Hans Reichenbach pointed out.
The uranium statement is true because it could not be otherwise. Long before you assembled that much fissile material you would pass the critical mass and the sphere would destroy itself. The statement follows from nuclear physics, and if you asked what would happen if someone tried, the answer is definite.
The gold statement is true because nobody has bothered, and because there is not enough accessible gold. Nothing in physics forbids it. If the solar system had condensed differently, or if some civilization decided to spend its resources that way, there could be such a sphere. The generalization is true by accident.
So among all the true universal generalizations, some are laws and some are accidents, and the difference is not visible in their logical form. Saying what the difference consists in is one of the hardest open problems in the philosophy of science.
Key idea: Two true universal generalizations can be identical in logical form while one states a law and the other records a coincidence, so lawhood cannot be a matter of form alone.
The marks of a law
Before theorizing, collect the differences that need explaining. Laws differ from accidents in at least four respects.
| Mark | Law | Accident |
|---|---|---|
| Counterfactual support | If this had been a uranium sphere of that size, it would have exploded | If this had been gold, it would have been fine; nothing follows about what would have happened |
| Confirmation by instances | Examining some uranium samples supports the generalization about all of them | Examining some gold objects gives no reason to expect the pattern to hold for unexamined ones |
| Explanatory role | Cite the law and you have explained why no such sphere exists | Citing the accident explains nothing; it is what needs explaining |
| Support for intervention | Tells you what would happen if you tried, which is Lesson 9's interventionist point | Tells you nothing about what would happen if you tried |
The first mark connects this lesson to Lesson 3. A generalization that is confirmed by its instances is one whose predicates are projectible, so the problem of lawhood and the problem of grue are two views of one puzzle. Green is projectible and grue is not, in the end, because the generalizations green figures in are lawlike and the ones grue figures in are not, which is progress only if lawhood can be explained without appealing back to projectibility.
The regularity theory, and why the simple version fails
The Humean starting point is that a law is nothing over and above a regularity. There is no necessity in nature; there are only patterns, and calling one a law records that it is a pattern with no exceptions.
The naive version of this is refuted by the gold sphere in one line: it is a true exceptionless universal generalization and it is not a law. Any Humean account needs to say which regularities count, without appealing to a necessity the Humean has denied.
The best attempt is the best-system analysis, developed from suggestions by John Stuart Mill and Frank Ramsey and worked out by David Lewis. Consider all the true statements about the world. Now consider the deductive systems you could construct from some of them, taking some as axioms and deriving the rest. Such systems can be evaluated on two competing virtues: strength, meaning how much they tell you, and simplicity, meaning how economically they do it. A system can always be made stronger by adding facts, at the cost of simplicity, and simpler by dropping them, at the cost of strength. The laws of nature are the generalizations that appear in the system that strikes the best balance.
Run the two spheres through it. The uranium generalization is a consequence of nuclear physics, which any strong and simple system will contain, so it earns its place. The gold generalization would have to be added as a separate axiom, buying a trivial gain in strength for a real loss in simplicity, so it does not. The account gets the case right.
Two objections press hard. First, simplicity is relative to a language: a system can be made to look simple by choosing predicates that make it so, and the grue predicates of Lesson 3 show how far that can be pushed. Lewis answered by restricting the vocabulary to predicates for perfectly natural properties, which resolves the problem by taking naturalness as a primitive. Second, and more troubling, the account seems to make lawhood depend on standards of simplicity that are ours, which threatens to make the laws of nature partly a fact about us. Lewis was aware of the worry and hoped that nature is kind enough that one system stands out by any reasonable standard. That is a hope rather than an argument.
The necessitarian alternative
In the late 1970s Fred Dretske, Michael Tooley and David Armstrong independently proposed a different answer. A law is not a regularity at all. It is a relation between universals, a relation of nomic necessitation holding between the property of being F and the property of being G. When that relation holds, it follows that all Fs are Gs, but the law is the relation and the regularity is its consequence.
The attraction is immediate. It explains the regularity rather than merely restating it, which is Armstrong's central complaint against the Humean: on a regularity view the law simply is the pattern, so it cannot be what accounts for the pattern. It explains counterfactual support, since a necessitation relation holds whether or not any instances occur. And it separates laws from accidents by a real difference in the world rather than by our systematizing preferences.
Bas van Fraassen pressed two objections that have not been fully answered. The identification problem: which relation among universals is the necessitation relation, and how would we know we had found it? The inference problem: why does the holding of this relation between two universals entail that every particular F is G? Naming the relation necessitation does not by itself establish that it necessitates anything, and Lewis made this point sharply. The necessitarian owes an account, and the accounts offered tend to reintroduce the very notion they were meant to explain.
What matters here: The Humean can say what makes a generalization a law but struggles to say why laws explain; the necessitarian can say why laws explain but struggles to say what the necessitation relation is.
Cartwright: the fundamental laws are false
A third position denies a premise both sides share. In How the Laws of Physics Lie, published in 1983, Nancy Cartwright argued that the fundamental laws of physics are not true descriptions of actual systems at all.
Her argument is concrete. Newton's law of universal gravitation says the force between two masses is proportional to the product of the masses and inversely proportional to the square of the distance. Now take two bodies that are also electrically charged. Is the force between them what the law of gravitation says? No: the actual force is the resultant of gravitational and electrostatic contributions. So either the law is false as stated, or it is not describing the actual force but a component that is never the whole story and is never directly measured.
Cartwright drew a trade-off from this. Phenomenological laws, the messy, restricted, empirically fitted generalizations of engineering and applied physics, are more nearly true and explain less. Fundamental laws explain a great deal and buy that power by abstracting away from everything that actually operates. The more explanatory a law is, the less accurately it describes any real system. Lesson 13 develops the same thought about models.
Ronald Giere pushed further in 1999, arguing that science can be described without invoking laws at all: what science offers is models with specified domains of application and specified degrees of fit, and calling some of them laws adds nothing.
The ceteris paribus problem
Outside fundamental physics, generalizations almost always come with an implicit qualification: other things being equal. Aspirin relieves headaches, other things being equal. An increase in the money supply raises prices, other things being equal. Predators reduce prey populations, other things being equal.
This creates a difficulty with a familiar shape. If any apparent counterexample can be attributed to other things not having been equal, the generalization forbids nothing, and Lesson 5's problem returns as a problem about laws. Jerry Fodor and others have defended such laws on the ground that the qualification is not open-ended: it can be filled in by the science that studies the interfering factors, and the special sciences would otherwise be left with no generalizations at all. Critics reply that until the qualification is filled in, the law's content is unspecified.
The honest position is that ceteris paribus generalizations do real predictive and explanatory work and that nobody has given a satisfactory account of their truth conditions.
What would settle it
The dispute has a clear structure, and each side has a specific debt.
For the Humean the debt is explanatory: give an account on which laws explain regularities without the law simply being the regularity. Some best-system theorists deny the debt, holding that unification is explanation, which connects this to the unificationist account of Lesson 9 and makes the two disputes one.
For the necessitarian the debt is the inference problem: say what the necessitation relation is, in terms that do not presuppose lawhood. If that could be done, the position would be clearly superior, which is why it remains attractive despite the difficulty.
For Cartwright the debt is to explain the extraordinary predictive success of the fundamental laws she calls false. Her answer, that they are true of models rather than of the world, is exactly the answer Lesson 13 examines.
Common misconceptions
- A law is just a well-confirmed generalization. The gold sphere generalization is exceptionlessly true and is not a law, so confirmation is not the criterion.
- Laws govern nature in the sense of compelling it. On the Humean view nothing compels anything; the metaphor of legislation is a relic of a theological picture and both camps treat it with caution.
- Every science has laws. Exceptionless laws are largely confined to physics. Biology and the social sciences work with hedged generalizations, models and mechanisms, which is why Lesson 9's mechanistic account matters.
- Cartwright denies that physics is successful. She denies that its fundamental laws are true descriptions of real systems, and her explanation of the success is that they are true of models.
- The best-system analysis makes laws subjective. It makes them depend on standards of simplicity and strength, which is a real worry, but the facts systematized are entirely objective and Lewis hoped one system would stand out on any reasonable standard.
Recap
- The uranium and gold sphere generalizations are both true and identical in form, but only one is a law.
- Laws support counterfactuals, are confirmed by their instances, explain, and tell you what would happen under intervention; accidents do none of these.
- The naive regularity theory fails because accidental generalizations are exceptionless regularities.
- The best-system analysis identifies laws as the generalizations in the deductive system that best balances simplicity and strength, and it gets the sphere case right.
- Its difficulties are the language-relativity of simplicity, addressed by appeal to natural properties, and the worry that lawhood comes to depend on our standards.
- The necessitarian view takes a law to be a relation between universals, which explains regularity and counterfactual support but faces the identification and inference problems.
- Cartwright argues fundamental laws are false of real systems because they describe component contributions, trading descriptive accuracy for explanatory power.
- Ceteris paribus generalizations do real work in the special sciences, and no satisfactory account of their truth conditions exists.
Sources
- Carroll, J. W. (n.d.). Laws of nature. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Reutlinger, A., Schurz, G., & Huttemann, A. (n.d.). Ceteris paribus laws. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Bird, A., & Tobin, E. (n.d.). Natural kinds. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Baker, A. (n.d.). Simplicity. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Cartwright, N. (1983). How the laws of physics lie. Oxford University Press.
- Armstrong, D. M. (1983). What is a law of nature? Cambridge University Press.
- Key terms
- Accidental generalization
- A true universal statement that holds by coincidence, supporting no counterfactuals and not confirmed by its instances.
- Counterfactual support
- The capacity of a generalization to underwrite claims about what would have happened in cases that did not occur.
- Regularity theory
- The Humean view that a law is nothing more than an exceptionless pattern, with no necessity in nature.
- Best-system analysis
- The account on which the laws are the generalizations appearing in the deductive system that best balances simplicity and strength.
- Nomic necessitation
- The relation between universals that Dretske, Tooley and Armstrong take a law to consist in, from which the corresponding regularity follows.
- Inference problem
- Van Fraassen's objection that it has not been shown why a relation between universals entails that every particular of one kind is of the other.
- Phenomenological law
- A restricted, empirically fitted generalization that describes actual systems accurately and explains little.
- Ceteris paribus law
- A generalization hedged by an other-things-equal clause, common in the special sciences and lacking a settled account of its truth conditions.
Realism against Instrumentalism
- State scientific realism and constructive empiricism precisely, and identify what each says about belief in unobservables.
- Give the no-miracles argument and the pessimistic induction in their strongest forms.
- Evaluate structural realism, deployment realism and the base-rate objection, and say what evidence bears on the dispute.
Thirteen roads to one number
In 1913 Jean Perrin published Les Atomes, in which he assembled determinations of Avogadro's number, the count of molecules in a mole of substance, obtained by radically different experimental routes. He measured the vertical distribution of tiny resin particles suspended in water and read the number off the sedimentation equilibrium. He tracked the random displacements of individual particles under the microscope. He measured their rotations. He drew on the viscosity of gases, on the blue of the sky, on critical opalescence, on blackbody radiation through Planck's constant, and on the counting of alpha particles emitted by radioactive decay.
These methods share almost no assumptions. They involve different apparatus, different branches of physics and different sources of error. Perrin listed around a dozen of them, and they converged on a value near six times ten to the twenty-third.
The effect on the scientific community was decisive. Wilhelm Ostwald, one of the most distinguished opponents of atomism, had held for decades that atoms were a useful fiction and that thermodynamics needed no such hypothesis. In 1909 he publicly conceded that the experimental evidence had convinced him that matter has an atomic constitution. Perrin received the Nobel Prize in Physics in 1926.
Now the philosophical question. Nobody saw an atom. What Perrin observed were resin particles, scale readings and photographic plates. Does the convergence of thirteen independent methods on one number give you reason to believe that molecules exist, or only reason to expect the instruments to keep agreeing?
Key idea: The realism debate is not about whether science works. Both sides agree it works. It is about whether its success licenses belief in what it says about things nobody can observe.
The two positions, stated carefully
Scientific realism is usually stated as three claims. Metaphysically, the world has a structure independent of what we think about it. Semantically, theoretical statements are to be read literally, so that a claim about electrons is a claim about electrons and not a shorthand for a claim about meter readings. Epistemically, our best current theories are approximately true and their central theoretical terms genuinely refer.
Instrumentalism in its older form denied the semantic claim: theoretical talk was to be read as a device for organizing observations, not as description. That position has few defenders now, because it fits scientific practice badly.
The serious contemporary alternative is Bas van Fraassen's constructive empiricism, set out in 1980. Van Fraassen accepts the semantic claim in full: theories are to be read literally, and if a theory says there are electrons then that is what it says. What he denies is the epistemic claim. Science aims, he argues, at empirical adequacy: at theories that get everything right about observable things. To accept a theory is to believe it is empirically adequate and to commit yourself to working within it, and that involves no belief about the unobservable parts.
The distinction matters because it makes the debate about evidence rather than about meaning. Van Fraassen is not saying electron talk is meaningless or that it should be paraphrased away. He is saying that the evidence which supports a theory supports it only as far as the observable consequences go, and that believing more is believing beyond what you have.
The no-miracles argument
The main realist argument was given its familiar form by Hilary Putnam in 1975, following a suggestion of J. J. C. Smart. Realism, he said, is the only philosophy that does not make the success of science a miracle.
Stated at its strongest, the argument turns on novel predictive success. A theory whose theoretical claims were entirely wrong might, by luck, save the phenomena it was built to save. What is hard to attribute to luck is a theory predicting something nobody had thought of, in a domain it was not designed for, and being right. Fresnel's wave theory of light entailed that a bright spot should appear at the centre of the shadow of a circular disc, a consequence Poisson pointed out in 1818 as an absurdity and which Arago then observed. Perrin's convergence is the same argument in a different form: thirteen wrong theories agreeing on the same wrong number would be a coincidence of an entirely different order from one theory fitting the data it was fitted to.
Put as an inference, the argument is an inference to the best explanation. The best explanation of a theory's novel predictive success is that the theory is at least approximately true about the entities it posits.
The pessimistic induction
The strongest reply is historical, and Larry Laudan gave it in 1981 in a paper called A Confutation of Convergent Realism. He assembled a list of theories that were empirically successful in their day, sometimes for a very long time, and whose central theoretical terms refer to nothing.
| Theory | Central posit | Status now |
|---|---|---|
| Phlogiston chemistry | Phlogiston, released in combustion | No such substance |
| Caloric theory of heat | Caloric, a conserved fluid | No such fluid |
| Nineteenth-century optics | The luminiferous ether | No such medium |
| Humoral medicine | Four bodily humours | No such humours |
| Pre-Copernican astronomy | Crystalline spheres | No such spheres |
The argument is then an induction over the history of science. If successful theories have repeatedly turned out to posit non-existent entities, then success is not a reliable indicator of truth or of reference, and we have no more reason to think our current successful theories refer than the caloric theorists had.
Notice how the two arguments mirror each other. The realist argues from success to truth; the antirealist argues from the history of successful falsehoods that the first inference has a bad track record. Each is using the other's favourite kind of reasoning.
What matters here: Both the no-miracles argument and the pessimistic induction are inductions, and each side is using the pattern of past cases to predict how present theories will fare.
Three realist replies
Structural realism. John Worrall argued in 1989 that the historical record shows continuity, but of a specific kind. When the ether theory of light gave way to electromagnetism, Fresnel's equations for the intensity of reflected and refracted light carried over essentially unchanged into Maxwell's theory. What was retained was mathematical structure; what was discarded was the story about what the structure was a structure of. So believe the structure, and suspend judgment about the underlying nature. This concedes the pessimistic induction's evidence while denying its conclusion, which is why it has attracted so much attention.
Deployment realism. Philip Kitcher and Stathis Psillos argue that a realist need not defend every part of a successful theory, only the parts that were doing the work in generating the successful novel predictions. The specific nature of the ether was not essentially deployed in Fresnel's derivations; the wave equations were. Believe what was essentially deployed. The obvious objection is that essentially deployed is defined with hindsight, which risks circularity, and answering it requires an independent criterion.
Entity realism. Ian Hacking argued that the reason to believe in electrons is not that a theory about them is well confirmed but that experimenters use them as tools: they spray them at other things in order to investigate those things. If you can manipulate something reliably to produce effects elsewhere, doubting its existence becomes idle. This links the realism debate to the interventionist account of Lesson 9 and shifts the evidence from theory to laboratory practice.
Three antirealist replies
The base-rate objection. Colin Howson pointed out that the no-miracles argument moves from the claim that success is very likely if a theory is true to the claim that a theory is very likely true given its success. That inference requires a prior probability that the theory is true, and if that prior is low, high success rates for true theories will not make the posterior high. This is the base-rate fallacy, and Lesson 12 works a numerical example of exactly the same structure.
The bad lot objection. Van Fraassen's point is that inference to the best explanation selects the best of the hypotheses we have thought of. That gives no assurance that the true hypothesis is among them. Unless you have some reason to think your list of candidates includes the truth, being told that one of them explains best tells you only that it is the best of a possibly bad lot.
Where to draw the line. Grover Maxwell pressed the classic objection to the observable and unobservable distinction: there is a continuum from seeing with the naked eye through spectacles, a magnifying glass, a light microscope and an electron microscope, and drawing a line anywhere on it is arbitrary. Van Fraassen's reply is that the distinction is vague but real, that observable means what an unaided human could perceive, and that although this is anthropocentric, so is epistemology, because it concerns what we can check.
What would settle it
Both sides agree on where the evidence lies, which makes this dispute unusually tractable to state and unusually hard to resolve.
The crux is continuity across theory change. If detailed historical study shows that the elements of past theories responsible for their novel predictive successes were preserved, in content or in structure, through the revolutions that overthrew them, the realist wins the historical argument and the pessimistic induction collapses into a claim about discarded ornament. If instead those working elements were themselves discarded, the antirealist wins.
This work has been done for particular cases, and the results are genuinely mixed. Fresnel's equations survived and support the structural realist. Caloric theory's successful results about heat engines survived in thermodynamics, which also helps the realist. But the ether was not ornament for nineteenth-century physicists; it was the centre of major research programmes, and the phlogiston case is harder to reconstruct as continuity. What would settle the dispute is a systematic survey rather than duelling examples, and both sides know it.
Common misconceptions
- Antirealists think science is unreliable. Van Fraassen holds that our best theories are empirically adequate, which is a strong claim about reliability. He declines to believe more than the evidence reaches.
- Realism is the default and needs no argument. The no-miracles argument is an argument, and it is an inference to the best explanation, which is exactly the pattern antirealists question.
- The pessimistic induction shows current theories are false. It shows that success is a poor guide to truth. Present theories may well be true; the argument is that their success does not establish it.
- Observable means detectable by any means. On van Fraassen's usage it means perceivable by an unaided human, which is why microscopes are contested and telescopes, whose objects could in principle be reached, are less so.
- Structural realism is a compromise with no content. It makes a definite, checkable historical claim: that mathematical structure is preserved across theory change even when ontology is not.
What you now know
- Perrin obtained Avogadro's number by around a dozen independent routes converging on one value, and the convergence persuaded even committed opponents of atomism.
- Realism combines a metaphysical, a semantic and an epistemic claim; constructive empiricism accepts the first two and denies the third.
- The no-miracles argument holds that novel predictive success would be a coincidence unless theories were approximately true.
- The pessimistic induction lists successful theories whose central terms did not refer, and concludes that success does not track truth.
- Structural realism preserves mathematical structure across theory change, deployment realism preserves only the parts essential to the successes, and entity realism grounds belief in manipulation.
- The base-rate objection, the bad lot objection and the observability line are the strongest antirealist replies.
- The dispute turns on whether the elements responsible for past successes survived the theory changes that discarded their ontologies, which is a historical question both sides accept as decisive.
Sources
- Chakravartty, A. (n.d.). Scientific realism. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Ladyman, J. (n.d.). Structural realism. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Internet Encyclopedia of Philosophy. (n.d.). Scientific realism and antirealism. iep.utm.edu
- Laudan, L. (1981). A confutation of convergent realism. Philosophy of Science, 48(1), 19-49. doi.org
- The Nobel Foundation. (n.d.). Jean Baptiste Perrin: facts. nobelprize.org
- van Fraassen, B. C. (1980). The scientific image. Oxford University Press.
- Key terms
- Scientific realism
- The view that the world has a mind-independent structure, that theories are to be read literally, and that our best theories are approximately true.
- Constructive empiricism
- Van Fraassen's view that science aims at empirical adequacy and that accepting a theory involves belief only about observables.
- Empirical adequacy
- The property of a theory that everything it says about observable things is true, whatever its unobservable posits.
- No-miracles argument
- The realist inference that the novel predictive success of science would be an implausible coincidence unless theories were approximately true.
- Novel prediction
- A prediction of something not used in constructing the theory, which realists take to be the success that most needs explaining.
- Pessimistic induction
- Laudan's argument from a list of successful theories whose central terms did not refer to the conclusion that success does not track truth.
- Structural realism
- The view that what survives theory change is mathematical structure rather than claims about the nature of the entities involved.
- Deployment realism
- The view that only the theoretical constituents essential to generating a theory's novel successes need be believed.
- Bad lot objection
- Van Fraassen's argument that inference to the best explanation selects the best available hypothesis without assurance that the truth is among the candidates.
Module 5: Evidence and Representation
The probabilistic machinery that makes talk of evidential support precise, worked with real numbers, and the awkward fact that most of what science offers as a picture of the world is a model known to be false.
Bayesian Confirmation, Worked Numerically
- Derive Bayes' theorem and apply it to a base-rate problem, obtaining the correct posterior.
- Show numerically why a risky successful prediction confirms far more than a safe one.
- State the problem of the priors and the problem of old evidence, and evaluate the responses to each.
A question most doctors got wrong
In 1978 three researchers put a question to sixty physicians, house officers and students at four Harvard Medical School teaching hospitals. A disease occurs in one person in a thousand. A test for it has a false positive rate of five percent. Someone with no symptoms tests positive. What is the chance that they have the disease?
Nearly half of the respondents answered ninety-five percent. About one in five gave the correct answer, which is roughly two percent.
The correct answer is not a trick and requires no advanced mathematics. Take ten thousand people. Ten of them have the disease, and assume the test catches all ten. Of the nine thousand nine hundred and ninety who do not, five percent test positive anyway, which is about five hundred people. So roughly five hundred and ten positive results come back, of which ten are genuine. Ten out of five hundred and ten is about two percent.
The reason so many trained people got this wrong is that they attended to the test's accuracy and ignored how rare the disease is. That neglected quantity, the prior probability, is the whole subject of this lesson, and it is the quantity that carries the philosophical weight in every argument about evidence in this course.
Key idea: How much a piece of evidence should move you depends not only on how diagnostic the evidence is but on how likely the hypothesis was before you got it, and ignoring the second quantity produces answers wrong by a factor of fifty.
Bayes' theorem, derived in two lines
The theorem is not a substantive assumption. It follows from the definition of conditional probability in two steps.
By definition, the probability of H given E is the probability of H and E together, divided by the probability of E. Equally, the probability of E given H is the probability of H and E together, divided by the probability of H. Rearranging the second gives the probability of H and E as the probability of E given H times the probability of H. Substituting into the first:
P(H given E) equals P(E given H) times P(H), divided by P(E).
The four quantities have names. P(H) is the prior, your probability for the hypothesis before the evidence. P(E given H) is the likelihood, how expected the evidence is if the hypothesis is true. P(E) is the marginal or total probability of the evidence, computed as P(E given H) times P(H), plus P(E given not-H) times P(not-H). And P(H given E) is the posterior, what you should believe afterwards.
The medical case, in symbols
Now do the opening problem properly. Let D be having the disease and let plus be a positive test.
- P(D) equals 0.001, the prevalence.
- P(plus given D) equals 1, assuming the test misses nobody.
- P(plus given not-D) equals 0.05, the false positive rate.
The total probability of a positive test is 0.001 times 1, plus 0.999 times 0.05, which is 0.001 plus 0.04995, giving 0.05095.
The posterior is 0.001 divided by 0.05095, which is 0.0196, or about two percent.
Change one number and watch what happens. If the disease has a prevalence of one in ten rather than one in a thousand, the total probability of a positive test is 0.1 plus 0.9 times 0.05, which is 0.145, and the posterior is 0.1 divided by 0.145, or sixty-nine percent. The test has not changed at all. Only the population has, and the same result now means something entirely different. This is why screening a low-risk population with a good test produces mostly false alarms, and it is a fact about arithmetic rather than about medicine.
Confirmation as probability-raising
The Bayesian account of evidence is then simple to state. Evidence E confirms hypothesis H just in case the posterior exceeds the prior, that is, when learning E raises your probability for H. It disconfirms when the posterior is lower, and is neutral when they are equal.
Note that confirmation in this sense does not mean establishing. Evidence can confirm a hypothesis while leaving it very improbable, as in the medical case, where a positive test raises the probability of disease twentyfold and still leaves it at two percent. Almost every dispute about statistical evidence in public life involves this distinction, and having it precisely is worth the whole lesson.
The quantity that measures how much the evidence discriminates, independently of the prior, is the likelihood ratio: P(E given H) divided by P(E given not-H). A ratio far above one means the evidence strongly favours H; a ratio near one means the evidence barely discriminates. Priors and likelihoods are separable, which is what lets two people who disagree about priors still agree about how strong a piece of evidence is.
Why risky predictions confirm more, with numbers
Lesson 4 reported Popper's claim that a bold theory forbidding a great deal is preferable to a cautious one, and that surviving a risky test is worth more than surviving a safe one. Popper regarded himself as an opponent of probabilistic confirmation theory. The irony is that Bayes' theorem gives his insight a precise numerical form.
Take a hypothesis you consider fairly unlikely, with prior 0.3, which makes a prediction E that follows from it, so P(E given H) equals 1. Compare two cases.
| Risky prediction | Safe prediction | |
|---|---|---|
| P(H), the prior | 0.3 | 0.3 |
| P(E given H) | 1 | 1 |
| P(E given not-H) | 0.05, because almost nothing else predicts it | 0.9, because most rivals predict it too |
| P(E), total | 0.3 plus 0.7 times 0.05, which is 0.335 | 0.3 plus 0.7 times 0.9, which is 0.93 |
| P(H given E), the posterior | 0.3 divided by 0.335, which is 0.90 | 0.3 divided by 0.93, which is 0.32 |
The same hypothesis, the same prior, the same successful prediction. In the risky case belief goes from 30 percent to 90 percent. In the safe case it goes from 30 percent to 32 percent. What made the difference is the denominator: how surprising the evidence would be if the hypothesis were false.
This is the whole content of the intuition that a theory should stick its neck out, and it explains the Eddington case of Lesson 4 exactly. A deflection of 1.75 arcseconds was near-certain if general relativity was right and very unlikely otherwise, so the observation moved belief enormously. It also explains why fitting an existing dataset with a flexible model impresses nobody: a model that could have fitted anything has a high P(E given not-H) and therefore a likelihood ratio near one.
The upshot: Evidence is powerful in proportion to how improbable it would be if the hypothesis were false, which is Popper's insight recovered inside the framework he rejected.
The ravens, once more
Lesson 3 left the raven paradox with a numerical treatment; now it can be stated in the vocabulary of this lesson. Observing a black raven and observing a non-black non-raven both raise the probability that all ravens are black, so the paradoxical conclusion stands. But their likelihood ratios differ by the ratio of the size of the non-black class to the size of the raven class. Where non-black things outnumber ravens by a hundred to one, a raven is a hundred times more informative than an apple.
Both observations confirm, one negligibly. The logic and the intuition are both vindicated, and the reconciliation required probability, which is the strongest single argument for the framework.
Where do the priors come from
This is the hard question and there is no consensus.
Subjective Bayesians, following Frank Ramsey and Bruno de Finetti, hold that priors are degrees of belief, constrained only by coherence. The argument for coherence is the Dutch book: if your degrees of belief violate the probability axioms, there is a set of bets, each of which you regard as fair by your own lights, that guarantees you a loss whatever happens. Being immune to that is a minimal standard of rationality, and it turns out to be equivalent to obeying the probability calculus.
That leaves priors otherwise unconstrained, which looks like a licence for wishful thinking. The standard reply is convergence. Theorems due to Savage, Blackwell and Dubins show that agents who start with different priors, provided nobody assigns zero or one to the hypothesis, will be brought arbitrarily close together by enough shared evidence. Disagreement washes out.
Three limits on that reply must be stated honestly. Convergence can be arbitrarily slow, and there is no guarantee it happens within the lifetime of a research programme. It requires agreement on the likelihoods, which in contested science is often exactly where people disagree. And it does nothing for priors of zero, since a hypothesis given probability zero can never be revived by evidence, which is why dogmatism is representable in the framework as an unrevisable prior.
Two standing problems
Old evidence. Clark Glymour pointed out in 1980 that if you already know E, then P(E) equals 1, and the theorem gives a posterior equal to the prior. Known evidence cannot confirm anything. Yet Mercury's perihelion advance had been known since 1859, and when Einstein derived it from general relativity in 1915 it counted heavily in the theory's favour. Everyone agrees it confirmed; the framework says it could not. Proposed repairs include computing with a counterfactual probability that pretends you did not know E, and holding that what is learned is not E but the logical fact that the theory entails E. Neither is fully satisfactory, and this is the same difficulty that Lesson 7's use-novelty criterion was invented to handle, which is worth noticing: two frameworks, one problem.
The catch-all. Computing P(E given not-H) requires considering every alternative to H, including alternatives nobody has yet conceived. In practice one computes against the rivals currently on the table, which quietly assumes the truth is among them. This is van Fraassen's bad lot objection from Lesson 11 arriving in numerical dress.
What would settle it
Two questions are tractable and worth stating.
First, do priors in fact wash out in real scientific episodes? This is answerable by case study: take a controversy, reconstruct the positions as prior distributions, and check whether the accumulating evidence brought the parties together and how fast. Work of this kind exists and generally supports convergence in mature disputes and not in immature ones.
Second, does the old evidence problem admit a principled solution? A repair that handles Mercury and does not licence retrofitting any theory to any known fact would settle a great deal, including the novelty question in Lesson 7. Nobody has one that satisfies everybody.
Common misconceptions
- A p-value is the probability that the hypothesis is false. It is the probability of data at least as extreme as observed, given the null hypothesis. Reading it as a posterior is exactly the base-rate error of the opening problem, and it is central to Lesson 15.
- Bayesian methods are subjective and therefore unscientific. The likelihood ratio is not subjective, and it does most of the evidential work. What is subjective is where you started, which is a fact about you that the framework makes explicit rather than hides.
- Confirming a hypothesis means showing it is probably true. Confirmation is probability-raising. The medical case raises the probability twentyfold and leaves it at two percent.
- Bayesianism solves the problem of induction. It gives a precise account of updating and relocates Hume's problem into the choice of priors, which is why Lesson 2 listed it as a reply with a cost.
- A theory that fits the data well is thereby well supported. A theory flexible enough to fit anything has a high probability of fitting whatever occurred, so its likelihood ratio is near one and the fit is nearly worthless as evidence.
The short version
- Most physicians asked the base-rate question answered ninety-five percent when the correct answer was about two percent, because they ignored the prior.
- Bayes' theorem follows in two lines from the definition of conditional probability and combines prior, likelihood and the total probability of the evidence.
- With a prevalence of one in a thousand and a five percent false positive rate, a positive test gives a posterior near two percent; raise the prevalence to one in ten and it becomes sixty-nine percent.
- Confirmation is probability-raising, and evidence can confirm strongly while leaving a hypothesis improbable.
- A risky prediction takes the same hypothesis from 0.30 to 0.90 where a safe one takes it from 0.30 to 0.32, which recovers Popper's insight numerically.
- Priors are constrained by coherence through the Dutch book argument and are otherwise free, with convergence theorems offering partial and slow relief.
- The problem of old evidence and the catch-all problem are the two standing difficulties, and each has a counterpart elsewhere in this course.
Sources
- Joyce, J. (n.d.). Bayes' theorem. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Talbott, W. (n.d.). Bayesian epistemology. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Hajek, A. (n.d.). Dutch book arguments. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Wikipedia contributors. (n.d.). Base rate fallacy. Wikipedia. en.wikipedia.org
- Casscells, W., Schoenberger, A., & Graboys, T. B. (1978). Interpretation by physicians of clinical laboratory results. New England Journal of Medicine, 299(18), 999-1001.
- Glymour, C. (1980). Theory and evidence. Princeton University Press.
- Key terms
- Prior probability
- The probability assigned to a hypothesis before the evidence in question is taken into account.
- Likelihood
- The probability of the observed evidence on the assumption that the hypothesis is true.
- Posterior probability
- The probability of the hypothesis after conditioning on the evidence, given by Bayes' theorem.
- Likelihood ratio
- The probability of the evidence given the hypothesis divided by its probability given the negation, measuring discriminating power independently of the prior.
- Base rate fallacy
- The error of judging a posterior probability from the accuracy of a test while ignoring how common the condition is.
- Dutch book argument
- The demonstration that degrees of belief violating the probability axioms permit a set of individually fair bets that guarantees a loss.
- Convergence theorem
- A result showing that agents with different non-extreme priors are brought arbitrarily close together by sufficient shared evidence.
- Problem of old evidence
- Glymour's objection that already known evidence has probability one and so cannot raise the probability of any hypothesis.
- Catch-all hypothesis
- The negation of the hypothesis under test, whose likelihood requires considering alternatives including ones not yet conceived.
Models and Idealisation
- Distinguish Galilean, minimalist and multiple-models idealisation with examples.
- State the paradox of how a model known to be false can explain, and assess the main answers.
- Explain robustness analysis, the objection to it, and how both bear on climate model ensembles.
An economy made of water
In 1949 a New Zealander named Bill Phillips, then a student at the London School of Economics, built a machine out of perspex tanks, pipes, valves and coloured water. It stood about two metres high. Water pumped up to a top reservoir flowed down through channels representing income, and was diverted into tanks for taxation, savings, imports and consumption. Valves set the marginal propensity to consume and the tax rate. Change a valve and the water levels shifted and settled at new equilibria, which were plotted on charts by pens attached to floats.
The machine, called the Monetary National Income Analogue Computer, worked. Around fourteen were built and several were used for teaching for years. Students who could not follow the equations could watch the water.
Nobody thought the British economy was made of water. Everybody thought the machine taught something true about it. That combination, an object known to be unlike its target and nevertheless used to reason about it, is what a model is, and it raises a question that runs through all of modern science: how can something you know to be false be a source of knowledge?
Key idea: Models are not descriptions offered as true. They are objects, physical or mathematical, deliberately unlike their targets in some respects and used to reason about them anyway.
What models are, and what kinds there are
Modern science works far more with models than with laws or theories directly. The category is broad.
| Kind | Example | What is doing the work |
|---|---|---|
| Physical or scale | The Phillips machine; Watson and Crick's metal and wire double helix | A concrete object whose behaviour or structure parallels the target |
| Mathematical | The Lotka-Volterra predator-prey equations | A set of equations whose solutions correspond to trajectories of the system |
| Computational | A general circulation model of the atmosphere | A simulation run forward in time on a discretized grid |
| Analogical | Treating the atom as a solar system, or a gas as billiard balls | A known system used as a source of expectations about an unknown one |
What unites them is that in each case the scientist works on the model and then transfers conclusions to the target. The philosophical questions are all about that transfer.
Three kinds of idealisation
Michael Weisberg drew a useful threefold distinction in 2007, and it clears up a good deal of confusion because the three kinds have different justifications.
Galilean idealisation introduces distortions for the sake of tractability, with the intention of removing them later. Frictionless planes, point masses, perfectly rigid bodies. The justification is pragmatic: the problem cannot be solved otherwise, and as computational power and mathematical technique improve, the idealisations are relaxed. The name comes from Galileo's practice of considering motion without resistance and then adding resistance back.
Minimalist idealisation keeps only the factors that make a difference to the phenomenon and deliberately omits everything else, with no intention of adding it back. A model of why a population goes extinct might include only the death rate and the carrying capacity. The justification is explanatory: including the non-difference-makers would obscure the explanation rather than improve it. On this view the falsehoods are not a regrettable compromise but part of what makes the model explanatory.
Multiple-models idealisation abandons the aim of a single complete model. Several incompatible models are maintained, each capturing different aspects, and no attempt is made to merge them. Weather and climate work like this, and so does much of biology. The justification is that the trade-offs are unavoidable, which is the next section's topic.
Two worked cases
Hardy-Weinberg equilibrium is the foundation of population genetics. It states how genotype frequencies relate to allele frequencies in a population, and it holds under five assumptions: the population is infinite, mating is random, and there is no mutation, no migration and no selection.
Every one of those assumptions is false of every real population. No population is infinite; mating is never random; mutation, migration and selection are always occurring. And yet the principle is not a poor description that geneticists tolerate. It is the baseline against which real populations are measured. Deviation from Hardy-Weinberg proportions is the signal that something interesting is happening, and quantifying the deviation is how selection and other forces are detected. The model earns its place by being false in specified ways.
The ideal gas law relates pressure, volume, temperature and quantity in a gas, on the assumptions that the molecules have no volume and exert no forces on one another except during collisions. Both are false. Real gases deviate measurably, especially at high pressure and low temperature, where the neglected molecular volume and attractive forces matter most. The van der Waals equation adds correction terms for exactly those two neglected factors, which is Galilean idealisation working as advertised: the distortion was introduced knowingly and removed when it mattered.
The paradox, stated plainly
Lesson 9 required that the statements in an explanans be true. Lesson 10 saw Nancy Cartwright argue that fundamental laws are false of real systems. Now put those together with the practice just described and the difficulty is sharp.
Scientists explain the behaviour of real gases using a law about gases that do not exist. They explain real allele frequencies using a principle about populations that could not exist. They predict the behaviour of the British economy using a machine made of water. If explanation requires truth, none of this is explanation. If none of this is explanation, very little in science is.
What matters here: Either the requirement that explanations be true has to be weakened, or the relation between a model and its target has to be something other than description.
Three answers
Similarity. Ronald Giere's proposal is that the model itself makes no claims at all. It is an abstract or concrete object, and it is neither true nor false. What carries a truth value is a separate theoretical hypothesis, asserted by the scientist, of the form: this model is similar to that target in these respects and to these degrees. That hypothesis can be straightforwardly true. The Phillips machine is not a claim about anything; the claim is that its equilibrium behaviour resembles the economy's in specified ways.
The obvious objection is that similarity is cheap. Everything resembles everything else in some respect. The account is only as good as the specification of the relevant respects and degrees, and specifying those requires knowing what you want to use the model for, which makes the relation three-place rather than two-place: model, target, purpose. Most defenders accept this and regard it as a feature.
Difference-making. A second answer says that a model explains when it captures the factors that made a difference to the phenomenon and abstracts away the rest. On this view the idealisations are not lies tolerated for convenience but a positive contribution, because they display which factors matter by omitting the ones that do not. Hardy-Weinberg explains because it isolates the effect of random mating in a large population, and it would explain less if it were more complete.
Mediation. Mary Morgan and Margaret Morrison argued in 1999 that models are neither derived from theory nor read off from data. They are constructed partly from each and partly from craft knowledge, and they function as autonomous instruments that mediate between theory and world. This is less a solution to the paradox than a redescription of the practice, but it fits what modelers actually do better than the alternatives.
Robustness, and its limits
If every model is false, how do you tell a result from an artefact of your simplifying assumptions? The standard answer is robustness analysis: build several models that make different simplifying assumptions and see which results survive across all of them. A conclusion that appears in every model is unlikely to be an artefact of any particular idealisation.
Richard Levins put the strategy memorably in 1966, writing that "our truth is the intersection of independent lies." He also argued that a model cannot simultaneously maximize generality, realism and precision, so modelers must sacrifice one, and that different sacrifices produce different and complementary models of the same system. That claim is the origin of multiple-models idealisation.
Steven Orzack and Elliott Sober raised the decisive objection in 1993. Robustness is not confirmation. If all your models share a false assumption, or if none of them is close enough to the target, agreement among them tells you only that they agree. Robustness is evidence for a conclusion only given some independent reason to think at least one of the models is adequate. This is a real constraint and it is easy to forget, because agreement across models feels like evidence in a way that it is not automatically.
Climate models, where this becomes practical
Climate projection is where these questions stop being abstract. A general circulation model divides the atmosphere and ocean into a grid of cells and steps the physics forward in time. The grid cells are far larger than clouds, convection cells and many other processes that matter, so those processes cannot be simulated directly. They are parameterized: represented by formulas relating their aggregate effect to variables the model does resolve, with parameters fitted to observation.
Projections are therefore made with ensembles: many runs of one model with different initial conditions and parameter settings, and runs of many models built by different groups. Where the models agree, the result is treated as robust; where they disagree, the spread is reported as a measure of uncertainty.
Both the strength and the weakness of this practice follow from the last two sections. The strength is a genuine robustness argument: conclusions that survive across models with different parameterizations are unlikely to be artefacts of any one scheme. The weakness is the Orzack and Sober point, sharpened by a specific fact: climate models are not independent. They share code, share published parameterization schemes, are tuned against overlapping observations, and are built by a community with shared training. Model agreement therefore overstates independence, and the spread of an ensemble is not a probability distribution over outcomes, however often it is read as one. Climate scientists say this in print, and it is a good example of a philosophical point being made from inside a science rather than about it.
What would settle it
The central question is whether the similarity relation can be specified without circularity. If the relevant respects in which a model must resemble its target are simply the respects needed for the inference you want to draw, the account risks saying only that a model works when it works. Progress here would consist in criteria for relevance derived from the target system rather than from the modeler's intentions, and difference-making accounts are the most promising attempt because causal relevance is a fact about the target.
The second question is empirical and tractable: does robustness across models predict success? One could take a body of past model ensembles, identify the robust conclusions at some date, and check them against what was later observed. Some of this has been done in meteorology and in epidemiology, and doing more of it is how the Orzack and Sober objection gets an answer rather than a rebuttal.
Common misconceptions
- A model is a simplified theory. Models are often constructed with materials from several theories plus assumptions from neither, which is Morgan and Morrison's point about autonomy.
- Idealisations are approximations that will eventually be removed. Only Galilean idealisations are meant to be removed. Minimalist idealisations are permanent and deliberate, and removing them would make the model explain less.
- If a model's assumptions are false, its conclusions are unreliable. Hardy-Weinberg has five false assumptions and is indispensable. What matters is whether the falsehoods concern factors that make a difference to the question being asked.
- Agreement among models is confirmation. Only if there is independent reason to think at least one is adequate, and only to the extent the models are genuinely independent.
- The spread of a climate ensemble is a probability distribution over outcomes. It is a spread over the models that happen to have been built, which are not a random sample of possible models.
Where this leaves us
- The Phillips machine modelled an economy with water and taught something true about it, which is the shape of the problem models pose.
- Models come as physical objects, equations, simulations and analogies, and in each case conclusions are drawn on the model and transferred to the target.
- Galilean idealisation distorts for tractability and is meant to be undone; minimalist idealisation keeps only difference-makers permanently; multiple-models idealisation abandons the aim of a single complete model.
- Hardy-Weinberg rests on five assumptions that are always false and functions as the baseline against which real populations are measured.
- The paradox is that models known to be false explain and predict, which forces either a weaker truth requirement or a non-descriptive account of the model-target relation.
- Similarity accounts, difference-making accounts and the mediating-models view are the three main answers, and each owes something.
- Robustness across models is evidence only if at least one model is independently credible and the models are genuinely independent, which climate ensembles only partly are.
Sources
- Frigg, R., & Hartmann, S. (n.d.). Models in science. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Potochnik, A. (n.d.). Idealization in science. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Frigg, R., & Nguyen, J. (n.d.). Scientific representation. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Wikipedia contributors. (n.d.). MONIAC. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (n.d.). Hardy-Weinberg principle. Wikipedia. en.wikipedia.org
- Levins, R. (1966). The strategy of model building in population biology. American Scientist, 54(4), 421-431.
- Morgan, M. S., & Morrison, M. (Eds.). (1999). Models as mediators. Cambridge University Press.
- Key terms
- Model
- An object, physical or mathematical, constructed to be reasoned about in place of a target system it resembles in specified respects.
- Galilean idealisation
- Deliberate distortion introduced for tractability and intended to be relaxed as technique improves, such as frictionless planes.
- Minimalist idealisation
- The inclusion of only the factors that make a difference to the phenomenon, with the omissions treated as permanent and explanatory.
- Multiple-models idealisation
- The maintenance of several incompatible models capturing different aspects, with no aim of merging them into one.
- Hardy-Weinberg equilibrium
- The relation between allele and genotype frequencies under five assumptions that no real population satisfies, used as a baseline for detecting evolutionary forces.
- Theoretical hypothesis
- Giere's term for the truth-evaluable claim that a model resembles a target in specified respects and degrees, distinct from the model itself.
- Robustness analysis
- The strategy of accepting results that survive across models making different simplifying assumptions.
- Parameterization
- The representation of a process too small for a simulation's grid by a formula relating its aggregate effect to resolved variables.
- Model ensemble
- A collection of runs across models and settings whose spread is used to indicate uncertainty, and which is not a random sample of possible models.
Module 6: Science as a Human Institution
Whether science can or should be free of moral and social values, and what a large-scale failure of published results reveals about every account of method in the preceding modules.
Values in Science and the Value-Free Ideal
- State the value-free ideal precisely and distinguish the stages at which values may enter.
- Reconstruct Rudner's argument from inductive risk and explain why Jeffrey's reply is now thought to fail.
- Assess Douglas's distinction between direct and indirect roles for values and Longino's account of objectivity as social.
Pills and belt buckles
In 1953 Richard Rudner published a six-page paper with a title that states its thesis: The Scientist Qua Scientist Makes Value Judgments. His argument turns on a contrast that anybody can follow.
Suppose the hypothesis under test is that a batch of pharmaceutical tablets does not contain a toxic ingredient in lethal quantity. Now suppose instead that the hypothesis is that a batch of machine-stamped belt buckles is not defective. In both cases the scientist has evidence, and in both cases the evidence is short of conclusive, because empirical evidence always is.
How strong does the evidence have to be before you accept the hypothesis? Rudner's point is that the answer is obviously different in the two cases, and that the difference has nothing to do with the evidence. It has to do with what happens if you are wrong. Being wrong about the buckles costs a customer a small annoyance. Being wrong about the tablets kills people. So the threshold for acceptance depends on the seriousness of the consequences of error, and judging seriousness is a value judgment.
If that is right, then value judgments are not confined to choosing what to study or to applying the results. They are inside the act of accepting a hypothesis, which was supposed to be the purely scientific part.
Key idea: Because evidence is never conclusive, accepting a hypothesis requires deciding that the evidence is sufficient, and how much is sufficient depends on the cost of being wrong.
The value-free ideal, stated fairly
The position Rudner attacks is worth stating in its best form, because it is not naive.
The value-free ideal does not claim that scientists are or should be indifferent to anything. It draws a line between two kinds of value. Epistemic or cognitive values are properties that make a theory a better candidate for truth. Kuhn's list from Lesson 6 is the standard one: accuracy, consistency, scope, simplicity and fruitfulness. These are supposed to be legitimate grounds for preferring one theory to another. Non-epistemic values, meaning moral, political, religious and economic ones, are supposed to have no place in judging whether a claim is true.
The rationale is straightforward and forceful. If a researcher may accept a hypothesis because accepting it serves a cause they favour, then evidence has lost its authority and the whole enterprise collapses into advocacy. The ideal is a barrier against wishful thinking, and it also underwrites public trust: a finding is worth something to people who disagree with the researcher politically only if the researcher's politics did not produce it.
Where values enter without controversy
Everyone in this debate agrees that values enter science at several points. Separating those points is the first step to seeing where the disagreement actually lies.
| Stage | Do values enter? | Contested? |
|---|---|---|
| Choice of research question and funding priorities | Yes, unavoidably | No; nobody thinks this compromises objectivity |
| Ethical constraints on method | Yes, and rightly | No; consent requirements are not epistemic failings |
| Application of results in policy | Yes | No |
| Choice of evidential thresholds and error trade-offs | This is the question | Yes |
| Characterization of ambiguous data | This is the question | Yes |
| Accepting or rejecting a hypothesis | This is the question | Yes |
So the dispute is narrow and sharp. It concerns whether non-epistemic values legitimately enter the internal stages, the ones the value-free ideal was designed to protect.
Jeffrey's reply, and why it is now thought to fail
Richard Jeffrey replied to Rudner in 1956 with an elegant move. Rudner assumes that scientists accept and reject hypotheses. Suppose they do not. Suppose the scientist's job is to assign and report probabilities, and to leave acceptance to whoever bears the consequences: the regulator, the physician, the manufacturer. Then the value judgment about how much evidence is enough is made by the decision-maker, not by the scientist, and the internal stages remain value-free.
This is a genuinely good reply and it is the basis of the division of labour many institutions actually use. Its difficulty, established most thoroughly by Heather Douglas in 2000 and in a book in 2009, is that inductive risk does not wait until the end. It enters at points where nothing can be handed off, because the choices are internal to producing the number the scientist is going to report.
- Where to set the error trade-off. Any statistical test trades false positives against false negatives. Setting a conventional threshold at five percent is a choice to tolerate one kind of error twenty times more readily than the current power tolerates the other. That trade-off cannot be reported as a probability and passed on; it is built into the analysis.
- How to characterize ambiguous data. Douglas's case study concerns toxicology. Pathologists scoring rat liver slides for evidence of tumours must classify borderline slides, and the classifications are genuine judgment calls. Score the borderline slides one way and a dose-response relationship appears; score them the other way and it does not. There is no probability to report until those calls are made.
- Which model to extrapolate with. Dose-response data are collected at high doses and used to estimate risk at low doses. Whether you fit a model with a threshold below which risk is zero, or one on which risk declines linearly to zero, changes the regulatory conclusion enormously, and the data at high doses often cannot distinguish them.
What matters here: Inductive risk enters wherever a judgment call affects the reported result, which means it cannot be quarantined at the point of acceptance and handed to a decision-maker.
Douglas's distinction, which is the key move
If values legitimately enter the internal stages, what stops science becoming advocacy? Douglas's answer is a distinction between two roles a value can play, and it is the most useful single idea in this literature.
A value plays a direct role when it functions as a reason in itself for accepting or rejecting a claim: I believe this because I want it to be true, or because believing it serves a cause. Direct roles are illegitimate, always, and this is exactly the wishful thinking the value-free ideal was protecting against.
A value plays an indirect role when it does not count as evidence but sets how much evidence is required, given what is at stake in being wrong. That is what Rudner's pills and buckles illustrate, and it is unavoidable rather than optional.
The distinction preserves what was right about the value-free ideal while conceding what Rudner and Douglas showed. Values may never substitute for evidence. They may, and must, calibrate how much evidence is demanded. And the practical upshot is a requirement of transparency: since the calibration is happening anyway, it should be stated rather than hidden, so that a reader who weighs the consequences differently can see exactly where they would part company.
Longino: objectivity as something a community achieves
Helen Longino approached the problem from a different direction in 1990. Her starting point is the gap identified in Lesson 5: evidence does not bear on a hypothesis by itself, but only given background assumptions, and those assumptions are not themselves entailed by the evidence. That gap is where values enter, often invisibly, because the assumptions can be so widely shared that nobody notices them as assumptions.
If that is right, then an individual scientist cannot achieve objectivity by scrutinizing their own reasoning, because the problematic assumptions are precisely the ones they do not see. Objectivity has to be a property of a community, and it is achieved when the community's practices expose background assumptions to criticism. Longino gives four conditions such a community must meet.
- There must be recognized venues for criticism: journals, conferences, review, in which objections can actually be raised.
- The community must show uptake, meaning that criticism changes what people believe rather than being absorbed and ignored.
- There must be publicly recognized standards by which criticism can be judged.
- There must be tempered equality of intellectual authority, so that qualified voices are not excluded, which matters because a community that all shares one background will not detect its shared assumptions.
The fourth condition is the one that connects the epistemology to questions about who is in the room. It is not a claim that diverse communities are nicer. It is a claim that a homogeneous community is worse at finding its own assumptions, which is an epistemic claim and is checkable.
The strongest case for the ideal, and the strategic complication
Defenders of the value-free ideal have a serious reply. Once you concede that values legitimately shape evidential thresholds, you have given every partisan a respectable-sounding argument for demanding more evidence when a finding is inconvenient and less when it is welcome. The ideal's virtue is that it is hard to game.
That reply is strengthened by a specific historical record. Naomi Oreskes and Erik Conway documented in 2010 how the demand for greater certainty was used strategically over decades by industries facing regulation on tobacco, acid rain, ozone depletion and climate, with the same small set of actors recurring. The tactic was not to argue that the science was wrong but that it was not yet certain enough, which is a demand the value-free ideal appears to license absolutely.
The irony is worth sitting with. A doctrine designed to keep interests out of science was used as an instrument by interests. Douglas's answer is that this is an argument for her position rather than against it: if the threshold is going to be set by someone, better that it be set explicitly, with the stakes named, than implicitly under cover of a demand for certainty that nobody could ever meet.
What would settle it
The core question is whether an evidential threshold can be set without any weighing of consequences. The argument that it cannot is close to a proof, and it is short: any threshold trades one kind of error against another, the errors have different costs, and choosing a trade-off without regard to the costs is not neutrality but an arbitrary implicit weighting. Almost everyone in the current literature accepts this.
What remains genuinely open is downstream. Which values, chosen by whom, with what accountability? Douglas's transparency requirement is one answer, Longino's community conditions are another, and there are proposals for democratic input into risk thresholds in regulatory science. The question also has an empirical part: do explicit value statements in scientific reports improve or damage public trust? That is measurable, and the results so far are mixed enough to be interesting.
Common misconceptions
- The value-free ideal says scientists should have no values. It says non-epistemic values should not determine judgments about what is true, which is a much narrower claim.
- Accepting that values enter means science is just politics. Douglas's direct and indirect distinction blocks exactly this. Values calibrating how much evidence is needed is very different from values counting as evidence.
- Jeffrey's reply was refuted because scientists do accept hypotheses. It was undermined by showing that inductive risk enters at earlier points where nothing can be handed off, whether or not there is a final act of acceptance.
- Longino's diversity condition is a political requirement. It is an epistemic one: a community sharing all its background assumptions cannot criticize them, so the argument is about error detection.
- If values are unavoidable, transparency is pointless. Transparency is what allows a reader who weighs the consequences differently to see precisely which judgment they would change and what follows.
What to carry forward
- Rudner's pills and belt buckles show that the threshold for accepting a hypothesis depends on the cost of being wrong.
- The value-free ideal permits epistemic values such as accuracy and simplicity and excludes moral and political ones from judgments of truth.
- Values uncontroversially enter in choosing research questions, in ethical constraints and in application; the dispute is about thresholds, data characterization and acceptance.
- Jeffrey proposed that scientists report probabilities rather than accept hypotheses, which fails because inductive risk enters earlier.
- Douglas showed inductive risk in error trade-offs, in the classification of ambiguous data such as borderline rat liver slides, and in extrapolation models.
- Her direct and indirect distinction bars values from serving as evidence while allowing them to set how much evidence is required.
- Longino relocates objectivity to communities, with venues for criticism, uptake, shared standards and tempered equality of authority.
- The strategic use of demands for certainty is the strongest practical argument for making the threshold explicit rather than leaving it implicit.
Sources
- Reiss, J., & Sprenger, J. (n.d.). Scientific objectivity. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Anderson, E., & Willett, C. (n.d.). Feminist epistemology and philosophy of science. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Longino, H. (n.d.). The social dimensions of scientific knowledge. The Stanford Encyclopedia of Philosophy. plato.stanford.edu
- Rudner, R. (1953). The scientist qua scientist makes value judgments. Philosophy of Science, 20(1), 1-6. doi.org
- Douglas, H. (2000). Inductive risk and values in science. Philosophy of Science, 67(4), 559-579. doi.org
- Longino, H. E. (1990). Science as social knowledge. Princeton University Press.
- Oreskes, N., & Conway, E. M. (2010). Merchants of doubt. Bloomsbury Press.
- Key terms
- Value-free ideal
- The view that non-epistemic values should play no role in judging whether a scientific claim is true, though they may shape what is studied and how results are used.
- Epistemic value
- A property such as accuracy, consistency, scope, simplicity or fruitfulness, taken to bear on a theory's likelihood of being true.
- Inductive risk
- The risk of error involved in accepting or rejecting a hypothesis on evidence that is less than conclusive.
- Direct role
- A value functioning as a reason in itself for accepting a claim, which Douglas holds is always illegitimate.
- Indirect role
- A value setting how much evidence is required before a claim is accepted, given the consequences of being wrong.
- Type I and type II errors
- Accepting a false claim and rejecting a true one; any threshold trades one against the other.
- Background assumption
- A premise needed to connect evidence to a hypothesis, not itself entailed by the evidence, and a common route by which values enter unnoticed.
- Tempered equality of intellectual authority
- Longino's condition that qualified voices not be excluded, on the epistemic ground that a homogeneous community cannot detect its shared assumptions.
- Manufactured doubt
- The strategic use of demands for greater certainty to delay regulatory action, documented across tobacco, acid rain, ozone and climate.
The Replication Crisis and What It Implies About Method
- State the main replication findings with their numbers, and the limits of what a failed replication shows.
- Explain how low prior plausibility, low power, analytic flexibility and publication bias combine to produce unreliable literatures.
- Connect the crisis to Popper, Duhem-Quine, Kuhn, Lakatos, Bayes and inductive risk, and evaluate the reforms.
Ninety-seven percent, then thirty-six
In 2015 the Open Science Collaboration published the result of an effort involving some 270 researchers. They had selected one hundred studies from three leading psychology journals and attempted to replicate each one, working from the original materials and, where possible, with the original authors' input.
Of the hundred original studies, ninety-seven had reported a statistically significant result. Of the hundred replications, thirty-six did. The average effect size in the replications was about half the size reported in the originals.
Three years later a second project, led by Colin Camerer, attempted the same thing for twenty-one social science experiments published in Nature and Science between 2010 and 2015, using sample sizes far larger than the originals. Thirteen of the twenty-one produced a significant effect in the same direction, and the effects that did replicate were on average about half the original size.
Those numbers are the starting point. What they mean, and what they imply about the accounts of method in the preceding modules, is the rest of the lesson.
Key idea: A large, systematic, preregistered attempt to reproduce published findings recovered roughly a third to two thirds of them, with effect sizes about half as large, in fields that were following the accepted methodology.
What a failed replication does and does not show
Begin with the caveats, because the finding is often overstated in both directions.
A failed replication does not establish that the original finding was false. It is consistent with the original being a false positive, with the replication being a false negative, with the effect being real but smaller than reported, and with the effect being real but dependent on conditions that differed between the two studies.
The Open Science Collaboration itself reported five different criteria for what counts as a successful replication, and the number varies with the criterion. Daniel Gilbert and colleagues published a comment in Science in 2016 arguing that the replication rate had been substantially underestimated, on the grounds that some replications had low power of their own, that several departed from the original protocols in ways likely to matter, and that ordinary sampling error accounts for more of the gap than the headline suggests. The original authors replied. Both exchanges are worth reading, and the honest summary is that the exact rate is contested while the direction of the finding is not.
Note also that this is not a uniform result across all of science. Replication rates differ by field and by subfield, and some areas of experimental physics and of clinical medicine with large multi-site trials look very different.
Why a literature can be mostly wrong while everyone follows the rules
The most important contribution to understanding this came in 2005, when John Ioannidis published a paper in PLoS Medicine arguing that most published research findings are false. The argument is Bayes' theorem applied to a research field rather than to a single test, and Lesson 12 supplies everything needed to follow it.
Consider a field in which researchers test hypotheses that have some prior probability of being true. The chance that a published significant finding is actually true depends on four quantities.
| Quantity | Effect on reliability |
|---|---|
| Prior plausibility of the hypotheses tested | Testing many implausible ideas guarantees that most significant results are false positives, exactly as in the medical screening case |
| Statistical power | Low power means true effects are usually missed, so a smaller share of the significant results are the true ones |
| Analytic flexibility | Every additional undeclared choice raises the effective false positive rate above the nominal one |
| Number of teams testing the same question | With publication bias, more teams means a higher chance that somebody's noise gets published |
Each of these has been measured. Katherine Button and colleagues estimated in 2013 that the median statistical power of studies in neuroscience was around twenty percent, meaning that a study of a real effect would fail to detect it four times in five. On analytic flexibility, Joseph Simmons, Leif Nelson and Uri Simonsohn showed in 2011 that four commonplace researcher choices, made after seeing the data, could push the false positive rate above sixty percent. They also ran a real experiment that appeared to show that listening to a particular song made participants younger, which is a result nobody can accept and which was produced by the same permissible-looking choices.
Now do the arithmetic with the medical case from Lesson 12 in mind. A field testing hypotheses with a prior of, say, one in ten, at twenty percent power, with an effective false positive rate of twenty percent rather than five, will publish more false findings than true ones. No fraud is required. Nobody in this story does anything they were taught was wrong.
The episode that started it
The immediate trigger was a paper by Daryl Bem published in 2011 in the Journal of Personality and Social Psychology, reporting nine experiments that appeared to show evidence of precognition, the ability to be influenced by events that have not yet happened. The methods were conventional, the statistics were the ones the field used, and the paper passed peer review at a leading journal.
The reaction was instructive because of what it revealed. Almost nobody believed the conclusion. But the paper was not obviously worse than a great deal of published work, and that was the problem: if the standard methods, applied competently, could produce this, what else had they produced? The prior probability was doing all the work in the community's rejection of the finding, and Lesson 12 explains why it should. When a hypothesis has an extremely low prior, even a p-value below the conventional threshold leaves the posterior tiny.
The upshot: The Bem episode showed that a literature could not be assessed by whether individual papers followed the rules, because the rules were producing this result too.
What this means for every earlier module
The crisis is a stress test for the accounts of method in this course, and each responds differently.
- Popper, from Lesson 4. The logic of falsification was never the bottleneck. Failed predictions did refute things, exactly as the account says. What was missing was that anyone should attempt the refutation and be rewarded for it, and journals for decades declined to publish replications. Popper's account of the logic survives and his silence about incentives is exposed.
- Duhem and Quine, from Lesson 5. A failed replication is itself a compound test. The standard defence, that some unmeasured moderator differed between the original and the replication, is the auxiliary hypothesis rescue in modern dress. Sometimes it is right. Stated without specifying the moderator in advance, it forbids nothing, and the response has been preregistered adversarial collaborations in which both sides specify beforehand what would count.
- Kuhn, from Lesson 6. Normal science supplies the assurance that a puzzle has a solution and that failure to find it reflects on the researcher. Applied to a false effect, that assurance produces persistence, reanalysis and eventually a published positive result. The dogmatism Kuhn identified as productive has a failure mode.
- Lakatos, from Lesson 7. The distinction between progressive and degenerating programmes applies almost directly. When each failed replication is met by a new post hoc moderator that predicts nothing further, that sequence is degeneration, and Lakatos's criterion identifies it without needing to know who is right.
- Bayes, from Lesson 12. Ioannidis's argument simply is Bayes' theorem applied to a literature. The whole phenomenon is a base-rate problem, and the profession's habit of reading a p-value as a posterior probability is the same error the Harvard physicians made.
- Values, from Lesson 14. Setting the significance threshold at five percent and tolerating twenty percent power is a decision about the relative cost of false positives and false negatives. It is an inductive risk judgment, and it was made implicitly and inherited by convention rather than argued.
The reforms, and what they are betting on
The response has been substantial and mostly institutional rather than logical, which is itself a result.
- Preregistration. Recording the hypothesis, the sample size and the analysis plan before collecting data, in a public registry with a timestamp. This directly targets analytic flexibility by fixing the choices before the data can influence them.
- Registered Reports. A stronger version in which the journal reviews the plan and commits to publishing the result whatever it shows. This targets publication bias at its source, since acceptance no longer depends on the outcome.
- Larger samples and power analysis. Raising power raises the proportion of significant findings that are true, directly.
- Multi-site replication. Coordinated projects running the same protocol across many laboratories, which both increases power and tests whether effects survive the variation that Feyerabend and Levins would both have expected to matter.
- Open data and materials. Which makes reanalysis possible and makes the analytic choices visible.
Every one of these is a bet that the problem is structural rather than a matter of individual carelessness, which is what the evidence in this lesson supports.
What would settle it
The question of how reliable a literature is is now empirically tractable in a way it was not twenty years ago, and this is the most encouraging thing in the lesson. Prospective, preregistered, adequately powered replication programmes measure the rate directly rather than inferring it, and comparing fields that have adopted reforms with fields that have not gives a test of whether the reforms work.
There is also a philosophical point worth ending on. The crisis was found from inside. The methods that revealed it, meta-analysis, power calculation, preregistration and large-scale replication, are scientific methods, and the people who applied them were members of the fields being criticized. That is exactly what the cluster account in Lesson 1 identified as the mark worth caring about: not that a field never produces error, but that it contains the machinery to detect and correct it. Whether it does so quickly enough is a separate and fair question, and it is one of the few in this course that the next twenty years will actually answer.
Common misconceptions
- A failed replication proves the original was wrong. It is consistent with a false positive, a false negative, a smaller true effect, or a genuine dependence on conditions. Only patterns across many replications support strong conclusions.
- The crisis is caused by fraud. Fraud exists and is rare. The mechanisms documented here require no dishonesty at all, which is what makes them serious.
- A p-value below the threshold means the finding is probably true. That depends on the prior and the power, as Lesson 12's arithmetic shows, and this misreading is at the centre of the problem.
- The crisis shows science does not work. The crisis was identified, quantified and addressed by scientists using scientific methods, which is the error-correcting capacity working, if slowly.
- Preregistration prevents exploratory research. It distinguishes exploratory from confirmatory work rather than banning either. Exploratory findings are simply reported as what they are, hypotheses generated rather than tested.
Pulling it together
- Of a hundred psychology studies replicated in 2015, ninety-seven originals were significant and thirty-six replications were, with effect sizes about half as large.
- A 2018 project replicated thirteen of twenty-one social science experiments from Nature and Science, again at about half the original effect size.
- A failed replication is consistent with several explanations, and the exact rate was contested by Gilbert and colleagues in 2016.
- Ioannidis showed that low prior plausibility, low power, analytic flexibility and multiple teams can make most published findings false with no dishonesty involved.
- Median power in neuroscience was estimated at around twenty percent, and four ordinary analytic choices can raise the false positive rate above sixty percent.
- The Bem precognition paper showed that the accepted methods could produce a result nobody believed, which made the methods themselves the issue.
- The crisis stresses every earlier module: falsification needed incentives, moderator defences are auxiliary rescues, normal science can persist with artefacts, and the whole pattern is a base-rate problem.
- Preregistration, Registered Reports, higher power, multi-site replication and open data are institutional responses to a structural problem.
Sources
- Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. doi.org
- Camerer, C. F., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2, 637-644. doi.org
- Button, K. S., et al. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14, 365-376. doi.org
- Wikipedia contributors. (n.d.). Replication crisis. Wikipedia. en.wikipedia.org
- Center for Open Science. (n.d.). Registered Reports. cos.io
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology. Psychological Science, 22(11), 1359-1366.
- Key terms
- Replication
- An independent repetition of a study's procedure intended to test whether the reported effect appears again.
- Statistical power
- The probability that a study will detect an effect of a given size if the effect is real; low power means true effects are usually missed.
- Researcher degrees of freedom
- Undeclared analytic choices made after seeing data, which raise the effective false positive rate above the nominal level.
- Publication bias
- The tendency for significant findings to be published and null findings to remain unpublished, distorting the literature.
- Positive predictive value
- The proportion of significant published findings that are true, determined by prior plausibility, power and bias.
- Preregistration
- Public timestamped recording of hypotheses, sample size and analysis plan before data are collected.
- Registered Report
- A publication format in which the study plan is peer reviewed and accepted before results exist, removing outcome-dependent publication.
- Hidden moderator defence
- The claim that an unmeasured condition differed between an original study and a failed replication, which forbids nothing unless specified in advance.
- Adversarial collaboration
- A preregistered study designed jointly by researchers who disagree, specifying in advance what result each would accept as decisive.