S1-401 – All models are wrong

But some are useful

In 1976, the statistician George Box published a paper in the Journal of the American Statistical Association that contained a sentence so concise and so true that it has since become one of the most quoted in the history of the philosophy of science: “All models are wrong, but some are useful.” Box was writing about statistical models specifically (about the regression equations and probability distributions that statisticians use to describe data), but the observation applies with equal force to every model this series has been examining: the Underground map that is wrong in every cartographic dimension but right for navigating the tube, the economic model of rational man that is wrong about human psychology but useful for understanding certain market dynamics, the Newtonian mechanics that are wrong at relativistic speeds and quantum scales but right for calculating the trajectory of a cricket ball. The universality of the observation is its most important feature. It is not merely that some models are wrong. It is that all models are wrong. And understanding what this means, precisely and without despair, is the project of this article and of the pillar it opens.

The claim requires careful unpacking, because it is easy to hear it in two ways that are both wrong. The first is the nihilist reading: if all models are wrong, then no model is better than any other, and the distinction between science and superstition collapses. This reading is false. The second is the trivial reading: of course all models are wrong in some sense, since no finite representation can capture an infinite reality, but so what? This reading misses the force of the claim. What Box meant, and what this article will argue, is something more specific and more practically important than either: that models are wrong in ways that are themselves systematic, knowable, and relevant to how the model should be used, and that the practice of science is, in significant part, the practice of understanding how models are wrong rather than merely that they are.

What it means for a model to be wrong

A model is wrong in the relevant sense when it makes predictions that differ from what actually occurs in the domain the model is supposed to describe. This sounds simple, but the implications are subtle and important.

The first implication is that being wrong is a property that requires specification: wrong about what, in what conditions, by how much, in which direction. Newton’s laws of motion are wrong about the behavior of objects moving at speeds approaching the speed of light: they predict trajectories that differ measurably from what special relativity predicts and from what observation confirms. They are not wrong about the behavior of objects moving at everyday speeds in everyday conditions, where the predictions of Newtonian mechanics and special relativity are indistinguishable to any available measurement. The wrongness of Newton’s laws is specific: it is quantified, it is directional (Newtonian mechanics underestimates the increase in effective mass at high velocities), and it is condition-dependent (it only matters when velocities are a significant fraction of the speed of light). A model that is wrong in this specific, quantified, condition-dependent way is a very different kind of wrong from a model that is wrong randomly, unpredictably, and in all conditions. Newton’s laws are wrong in a way that can be precisely described, that is entirely predictable from within the successor theory (special relativity), and that is practically irrelevant in the conditions where the model is applied.¹

The second implication is that being wrong and being useful are not mutually exclusive. The wrong-but-useful structure, which is Box’s formula, is not a paradox. It is the normal condition of every model that has ever been practically applied. The usefulness of a model is a function of whether its predictions are accurate enough, in the conditions where it is applied, for the purposes for which it is applied. The Newtonian model of planetary motion is accurate to a very high degree for the purpose of calculating where a planet will be visible from Earth on a given night, accurate to a slightly lower degree for calculating the orbit of Mercury (where general relativistic corrections are measurable), and accurate to zero for calculating the behavior of a proton in a particle accelerator. The model is useful for the first purpose, marginally useful for the second, and useless for the third, but these verdicts apply to specific conditions and purposes, not to the model in general.

The third implication, the one that generates the most consequential practical errors, is that a model’s usefulness in one set of conditions does not guarantee its usefulness in different conditions. This is the extrapolation problem: the systematic tendency to apply a model beyond the conditions within which its accuracy has been established, producing predictions that are wrong in ways that the model’s previous successes provided no warning of. The financial risk model that performed reliably through two decades of moderately volatile markets and failed catastrophically in the conditions of 2008 was not wrong in the sense that it had been producing wrong predictions throughout its history. It was wrong in the sense that its predictions were accurate only within the range of conditions it had been calibrated on, and the conditions of 2008 fell outside that range. The wrongness was latent, invisible from within the model’s domain of valid application, until the conditions changed.

The philosophy of science: what scientists think they are doing

The question of what scientists are doing when they build and test models is one of the central questions of the philosophy of science, and the answers to it have changed significantly over the past century in ways that are directly relevant to understanding why Box’s formula is true.

The naive picture, still implicit in much public communication about science, is what the philosopher of science Peter Godfrey-Smith calls naive inductivism: scientists observe the world, accumulate observations, generalize from the observations to theories, and progressively build up a body of knowledge that converges on truth. On this picture, science is essentially a process of truth accumulation. Its outputs, the theories and laws that result from the process, are true, or at least increasingly close to truth, and the scientific method is the guarantee of that truth.

This picture is wrong in several respects that the philosophy of science has established clearly. The most fundamental objection was raised by David Hume in the eighteenth century and has never been satisfactorily answered: no finite number of observations can logically guarantee the truth of a general claim. The observation of a million white swans does not establish that all swans are white, because the next swan might be black, as indeed it turned out to be when Europeans reached Australia. Induction, the inference from specific observations to general claims, is not logically valid, and the truth of a scientific theory cannot therefore be logically established by the observations that support it.²

Karl Popper’s response to the induction problem, developed in the 1930s and highly influential in the subsequent philosophy of science, was to propose falsifiability as the criterion of scientific claims and falsification as the mechanism of scientific progress. A scientific claim is not one that can be confirmed by observation (no claim can be confirmed by finite observation, for Hume’s reasons) but one that can be refuted by observation: a claim that specifies in advance what observations would be incompatible with it, and that is therefore genuinely at risk of being wrong. Scientific progress, on Popper’s account, is not the accumulation of confirmed truths but the elimination of falsified falsehoods: science advances by testing bold conjectures against observation and retaining those that survive while discarding those that do not.

Popper’s account captures something important: the role of falsifiability in distinguishing genuine empirical claims from claims that are insulated from evidence. But it has been criticized from several directions that are equally relevant to this article’s argument. The most important criticism, associated with the philosopher Imre Lakatos, is that individual scientific theories are never falsified by individual observations in the clean way that Popper’s account describes. Every scientific theory is embedded in a web of auxiliary hypotheses (assumptions about the reliability of instruments, the absence of confounding factors, the correctness of the background theories on which the test depends), and when a prediction fails, the failure can always be attributed to one of the auxiliary hypotheses rather than to the core theory. A negative experimental result does not automatically falsify the theory being tested. It falsifies the conjunction of the theory and all the auxiliary hypotheses on which the test depends, and the scientist must decide which element of the conjunction to revise.³

This decision is not made by logic alone. It is made by a complex combination of evidential, theoretical, social, and pragmatic considerations that Thomas Kuhn’s account of scientific revolutions illuminated with unusual clarity, an account that is the subject of the next section.

Kuhn and the structure of scientific revolutions

Thomas Kuhn’s The Structure of Scientific Revolutions, published in 1962, is one of the most influential books in the philosophy of science, and one of the most misunderstood. The book introduced the concept of the paradigm: the framework of shared assumptions, methods, exemplary problems and solutions, and standards of good practice that organizes a scientific community’s work during what Kuhn called a period of normal science. And it introduced the concept of the paradigm shift: the revolutionary episode in which the existing paradigm is replaced by a new one that is incommensurable with it, not merely an extension or correction of the old paradigm but a fundamentally different way of organizing the same domain of inquiry.⁴

Kuhn’s account is directly relevant to Box’s formula, because it describes the mechanism by which models that are wrong accumulate the anomalies (the experimental results that don’t fit, the predictions that fail, the problems that resist solution) that eventually force their replacement. During a period of normal science, the scientific community works within a paradigm that is treated as given: the framework is not questioned but applied, extended, and refined. Anomalies, results that the paradigm cannot accommodate, are initially set aside, attributed to experimental error, treated as puzzles that a more careful application of the paradigm will eventually resolve. The paradigm generates a characteristic blindness to the ways in which it is wrong, not because scientists are dishonest but because the paradigm defines what counts as a problem worth solving and what counts as an acceptable solution. The anomalies that fall outside this definition are literally invisible as anomalies: they are seen as failures of technique, as poorly posed problems, as irrelevant complications.

The Kuhnian paradigm shift occurs when the anomalies accumulate to the point where they can no longer be set aside: when the protective belt of auxiliary hypotheses that has been sheltering the core theory from falsification becomes too elaborate to be credible, when a new generation of scientists finds the old framework unsatisfying in ways they cannot quite articulate, and when a new framework emerges that reorganizes the domain in a way that transforms the anomalies into predictions. The shift is not a gradual accumulation of evidence for the new framework and against the old. It is a gestalt switch, a reorganization of what is salient, what is problematic, and what counts as an explanation, that is experienced differently by scientists on different sides of it.

The practical implication of the Kuhnian account for Box’s formula is this: the wrongness of a model is typically visible, before the paradigm shift that reveals it, only from outside the paradigm. From inside, the model looks not wrong but incomplete: in need of refinement, extension, and more careful application, but not fundamentally mistaken. The history of science is a history of this pattern: confident and productive paradigms that were, from outside, wrong in systematic and specifiable ways, but that could not see their own wrongness from within. Newtonian mechanics was wrong about high velocities and strong gravitational fields. The steady-state model of the universe was wrong about cosmological expansion. The static earth model of geology was wrong about continental motion. In each case, the wrongness was latent and invisible within the paradigm until the successor paradigm made it visible.

The self-correcting mechanism: why science deserves trust

The preceding analysis might seem to leave science in a precarious position: if all models are wrong, if falsification is not a clean logical operation, and if paradigm shifts are more like gestalt switches than rational revisions, why should the outputs of science be trusted at all?

The answer is not that science produces truth. It is that science has a self-correcting mechanism: a set of institutional and methodological norms that systematically exposes models to tests they might fail and that, over time, eliminates the models that fail those tests more reliably than any alternative knowledge-producing system. The mechanism includes: the requirement that claims be stated in forms specific enough to be testable, the norm of replication (the requirement that results be reproducible by independent investigators), the norm of peer review (the requirement that work be assessed by qualified critics before publication), and the norm of openness (the requirement that methods and data be available for scrutiny). No single element of this mechanism is infallible: as the replication crisis has shown, peer review can fail, replication can be partial, and the social dynamics of scientific communities can insulate paradigms from challenge. But the mechanism as a whole is more reliable than the alternatives, because it is specifically designed to catch errors rather than to protect current beliefs.

The trust that science deserves is therefore not the trust one would give to a claim that is known to be true. It is the trust one would give to a process that is specifically designed to expose claims to tests they might fail, and that has a track record of successfully replacing models that fail such tests with models that fail fewer tests in the same conditions. This is a form of trust that is calibrated rather than absolute: that is proportional to the track record of the specific field, the strength of the evidence for the specific claim, and the degree to which the claim has been exposed to the kind of scrutiny that the self-correcting mechanism provides.

The p-value problem: a specific failure in the self-correcting mechanism

The self-correcting mechanism described in the previous section depends critically on one precondition: that the tests to which scientific claims are subjected are genuinely capable of distinguishing real effects from the accidents of random variation. In much of contemporary science, this precondition is not reliably met, not because scientists are dishonest but because the statistical tool most widely used to decide what counts as a real finding is considerably weaker than its users typically assume, and its weakness is systematic in ways that produce the replication crisis as a predictable consequence rather than a surprise.

The tool in question is the p-value, and its threshold of significance: the convention that a result is declared statistically significant (that it counts as a real finding worthy of publication and citation) if the p-value associated with the measured effect is less than 0.05. The p-value is defined as the probability of observing a result at least as extreme as the one obtained, given that there is in fact no real effect, given that the null hypothesis is true. A p-value of 0.05 therefore means: if there were no real effect, a result this extreme or more extreme would occur by chance approximately 5 percent of the time, or once in 20 experiments.

The intuitive reading of this threshold, that a p-value below 0.05 means there is only a 5 percent chance the result is a false positive, is wrong, and the wrongness is consequential. The p-value answers the question: how probable is this data if the null hypothesis is true? It does not answer the question that matters: how probable is the null hypothesis given this data? These are entirely different questions, and confusing them is not a minor statistical error. It is a systematic misreading of what the evidence shows.⁵

To understand why the distinction matters in practice, consider what happens when a research program investigates a large number of hypotheses, most of which are false, which is the normal condition of exploratory science in any field. Suppose that out of 100 hypotheses being tested, 10 are genuinely true and 90 are false. If each true hypothesis has an 80 percent probability of producing a statistically significant result (the statistical power of the test), and the significance threshold is 0.05, then the expected number of true positives is 8 and the expected number of false positives is 0.05 multiplied by 90, which is 4.5. Of the roughly 12.5 significant results produced, approximately 4.5, more than a third, are false positives, despite every individual result meeting the conventional threshold of statistical significance. The p-value threshold of 0.05 does not guarantee a 5 percent false positive rate in the published literature. It guarantees a 5 percent false positive rate per test. The false positive rate in the published literature depends additionally on the proportion of tested hypotheses that are true, the statistical power of the tests, and, critically, the publication bias that determines which results get published.⁶

Publication bias is the decisive amplifier. If positive results are much more likely to be published than negative results, which they are, consistently, across most scientific fields, then the published literature is a non-representative sample of all experiments conducted. The published literature contains the positive results from genuine effects, the positive results from chance variation (the 5 percent false positives), and very few of the negative results that would allow readers to assess the false positive rate. From inside the published literature, the false positive rate looks very low because the denominator, the number of experiments conducted, is invisible. From outside, looking at the full set of experiments including the unpublished ones, the false positive rate can be alarmingly high.

This is precisely what Ioannidis’s 2005 analysis established, and what the Open Science Collaboration’s replication project confirmed empirically: that a substantial proportion of published findings in psychology, medicine, nutrition, and economics represent chance variations that achieved significance in the original study but could not be reproduced when the experiment was repeated by independent investigators. The replication crisis is not a scandal about scientific dishonesty. It is the predictable consequence of applying a weak statistical threshold (one that was set arbitrarily in the 1920s by the statistician Ronald Fisher as a rough guide rather than as a rigorous decision criterion) to fields where the prior probability of any given hypothesis is low and publication bias is high.⁷

The threshold of 0.05 was never intended to be the universal bright line it became. Fisher himself argued that p-values should be used as informal guides in the context of a sustained research program, not as the sole criterion for declaring a result real and worthy of publication. The subsequent institutionalization of the 0.05 threshold (its codification in journal submission requirements, its embedding in research funding decisions, its use as the single criterion separating publishable from unpublishable results) transformed a rough heuristic into a mechanism that systematically overproduces false positives.

The Bayesian framework that the next section develops provides a more coherent alternative. Rather than asking whether the data are surprising enough given no effect, Bayesian analysis asks directly how the data should update our belief in the hypothesis, taking into account both the evidence the experiment provides and the prior probability of the hypothesis being true. This approach does not eliminate uncertainty. It quantifies it in a way that is honest about what the evidence actually establishes. A Bayesian analysis of a weak experimental result with a low prior probability will correctly conclude that the result, even if nominally significant, provides only modest evidence for the hypothesis. A p-value of 0.049, applied to the same result, incorrectly implies that the result has cleared a meaningful evidential threshold.

The practical consequence for how we receive scientific findings is specific and important. A single published study reporting a significant result at p less than 0.05 in a field with low prior probabilities and known publication bias (nutrition science, social psychology, certain areas of medical research) is considerably weaker evidence than its statistical significance suggests. The appropriate response is not to dismiss all findings in these fields, but to hold individual findings with calibrated uncertainty proportional to the prior probability of the hypothesis, the statistical power of the study, the degree of publication bias in the field, and the number of independent replications available. This is precisely the Bayesian updating process that the next section describes, and it is the more honest expression of what Box’s formula demands: not uniform skepticism about all scientific findings, but specific, quantified, condition-dependent assessment of how wrong any particular finding might be.

The Bayesian argument: why fundamental revision becomes increasingly unlikely

There is a more precise and more satisfying answer to the question of why science deserves trust than the self-correcting mechanism argument alone provides, and it comes from the Bayesian framework that article S1-901 of this series will develop in full. The Bayesian argument does not merely say that science works. It specifies why the outputs of mature, well-tested science are extremely unlikely to be fundamentally wrong, and it makes this claim in terms that are themselves testable and that have been consistently confirmed by the history of science.

The Bayesian framework begins with a prior probability (an initial estimate of how likely a claim is to be true, before the evidence is examined) and updates it through the likelihood ratio: the ratio of the probability of observing the available evidence if the claim is true to the probability of observing it if the claim is false. Claims that are confirmed by many independent pieces of evidence, each of which was more likely to occur if the claim is true than if it is false, accumulate posterior probabilities that approach certainty in a well-defined mathematical sense. They do not reach certainty (the induction problem means that no finite evidence can guarantee truth) but they approach it in a way that makes their practical revision increasingly improbable as the evidence accumulates.

The history of scientific paradigm shifts, examined through this lens, reveals a pattern that is both more encouraging and more precise than Kuhn’s description alone would suggest. When a new theory replaces an old one, it does not simply erase the old theory’s predictions. In virtually every well-documented case of scientific revolution, the successor theory recovers the predecessor’s successful predictions as special cases, as approximations that hold in the conditions where the old theory was calibrated, while extending coverage to conditions the old theory could not handle. Newton’s mechanics is not overturned by special relativity. It is recovered as the low-velocity limit: in the conditions where Newton’s predictions were tested and confirmed, special relativity gives the same answers, to the available precision of measurement. The germ theory of disease does not refute the clinical observations of pre-germ-theory medicine. It explains them: the treatments that worked before the mechanism was understood worked because they happened to disrupt pathogen transmission or support immune function, and the germ theory made sense of why they worked.⁸

This pattern, successor theories preserving the empirical content of predecessor theories while extending it, is what philosophers of science call the no-miracles argument for scientific realism: the extraordinary predictive success of mature scientific theories would be miraculous if those theories were not approximately tracking something real about the structure of the world. And it is the Bayesian argument against the radical skepticism that a naive reading of Box’s formula might invite: if all models are wrong, why not treat all models as equally uncertain? The answer is that they are not equally uncertain. The models that have survived the longest exposure to the widest variety of independent tests, that have been confirmed by the most diverse lines of evidence, and whose core predictions have been recovered by every successor theory that has been developed, have accumulated posterior probabilities that are not merely high but are astronomically high relative to the prior probability of any specific alternative.

The practical implication of this argument is a specific and important calibration of the series’ central claim. The frontier of scientific knowledge (the young subdisciplines, the newly developed models, the claims that have been tested in only a narrow range of conditions) is where revision is genuinely possible and where genuine humility is warranted. The interior of established science (the atomic theory of matter, the germ theory of disease, the fact of evolution by natural selection, the structure of DNA, the age of the universe) has accumulated so much evidence from so many independent directions that fundamental revision of these claims is not merely unlikely in a general sense. It is specifically improbable in a quantifiable way: any successor theory would need to not only accommodate the existing evidence but explain why the existing evidence confirmed the predecessor theory so consistently in the conditions where it was tested. That is an extraordinarily demanding constraint. It does not make revision impossible (the history of science contains genuine surprises) but it makes it very unlikely to be fundamental in the domain where the evidence is richest.

Box’s formula, all models are wrong, is therefore not a counsel of uniform skepticism. It is a statement that applies differentially: most urgently at the frontier, where models are young and evidence is sparse; least urgently in the interior, where models have been tested so thoroughly that the remaining wrongness is specific, quantified, and bounded, like the relativistic corrections to Newtonian mechanics, which are real and important in their domain but leave the Newtonian predictions entirely intact where they were established. The appropriate calibration of confidence to evidence is not equal skepticism about everything science has produced. It is high confidence in what is well-established, genuine openness at the frontier, and the specific epistemic humility of knowing which parts of the map have been drawn and redrawn so many times that their outlines are very unlikely to change, and which parts are still being drawn for the first time.

Box’s formula applied to the series

Box’s formula, all models are wrong, but some are useful, is not merely a statement about statistical models or scientific theories. It is the central epistemological claim of this entire series, stated in its most compressed form. Every article in the series has been applying this formula to a different domain: the perceptual models of article S1-201, the cognitive models of articles S1-202 through S1-213, the linguistic models of articles S1-301 through S1-307, and, in the articles that follow this one, the scientific models of physics, biology, and economics, and the social and political models of ideology, democracy, and human nature.

The consistent message is not that models should be distrusted or abandoned. It is that models should be held with the calibrated confidence that their actual epistemic status warrants: with conviction proportional to the evidence for their accuracy within their domain of application, and with genuine openness to the possibility that the domain of application is narrower than the model’s success has suggested. The model of the free market is not wrong in the way that the flat earth is wrong. It captures something real about the behavior of decentralized systems with price signals. It is wrong in the way that Newtonian mechanics is wrong: accurately and usefully wrong within a specific domain, and systematically misleading when extrapolated beyond that domain to conditions (environmental externalities, public goods, information asymmetries) that the model was not built to handle.

The diagnostic question that the series has been applying throughout, what would have to be true for this model to be wrong?, is the operational expression of Box’s formula. It is the question that keeps the model honest: that maintains the distinction between the model’s tracked accuracy within its domain and the extrapolation that applies it to conditions where the accuracy has not been established. Applied to the models of science, it is the question that drives paradigm shifts. Applied to the models of politics and economics, it is the question that prevents ideologies from becoming closed systems insulated from evidence. Applied to the models of the self, it is the question that makes genuine self-knowledge possible, or at least more possible than it would be without it.

The map is not the territory. All models are wrong. Some are useful. And the most important thing about any useful model is knowing, as precisely as possible, how it is wrong.

Further reading

George Box and Norman Draper’s Empirical Model-Building and Response Surfaces (1987) is the source of Box’s famous formula in its most fully developed context, statistical modeling, and provides the most precise available account of what it means for a model to be wrong in a quantifiable sense and how the wrongness of a model should be taken into account in its application.

Karl Popper’s The Logic of Scientific Discovery (1934, English translation 1959) is the foundational text for the falsifiability criterion: the argument that scientific claims are distinguished from non-scientific claims by their vulnerability to empirical refutation. It is demanding but rewarding, and it remains the most rigorous available statement of the position that the subsequent philosophy of science has both built on and challenged.

Thomas Kuhn’s The Structure of Scientific Revolutions (1962) is one of the most important and most frequently misread books in the philosophy of science: important because it describes the social and psychological dynamics of scientific change with unusual accuracy, and frequently misread because its account of paradigm shifts is taken to imply a relativism about scientific truth that Kuhn himself did not endorse. The second edition (1970) includes a postscript in which Kuhn addresses the relativist reading and clarifies what he did and did not mean to claim.

Peter Godfrey-Smith’s Theory and Reality: An Introduction to the Philosophy of Science (2003) is the most accessible and most intellectually honest introduction to the philosophy of science currently available, covering the major positions with equal fairness and without forcing a conclusion that the evidence does not support.

Imre Lakatos’s “Falsification and the Methodology of Scientific Research Programmes,” collected in The Methodology of Scientific Research Programmes (1978), provides the most sophisticated available response to both Popper and Kuhn: the account of how scientific theories are protected by auxiliary hypotheses during periods of productive development, and how the distinction between progressive and degenerating research programs allows rational evaluation of competing theories without the clean logic of falsification that Popper required.

Regina Nuzzo’s “Statistical Errors,” published in Nature in 2014 and freely available online, is the most readable short treatment of the p-value problem and its consequences for scientific practice: written for a scientific rather than statistical audience, and directly relevant to understanding the replication crisis as a consequence of statistical methodology rather than scientific dishonesty.

Notes

¹ The specific way in which Newtonian mechanics is wrong, and the conditions under which the wrongness becomes measurable, is one of the best-documented cases in the history of physics. The relativistic corrections to Newtonian mechanics become significant when velocities approach a substantial fraction of the speed of light (approximately 3 × 10⁸ meters per second). For everyday velocities, even including the fastest aircraft and rockets, the relativistic correction is far smaller than any available measurement can detect. The planet Mercury, the closest planet to the Sun, moves fast enough and is deep enough in the Sun’s gravitational well that the precession of its orbit deviates measurably from the Newtonian prediction, a discrepancy of 43 arcseconds per century that was known before Einstein and that general relativity explained precisely. This is the standard case: Newtonian mechanics is wrong in a specific, quantified, condition-dependent way that was predictable from within the successor theory.

² Hume’s problem of induction, first stated in A Treatise of Human Nature (1739), is the observation that no finite sequence of observations can logically establish a universal generalization, because the next observation might always be different from all preceding ones. The problem has generated an enormous philosophical literature and has not been resolved: no one has produced a logically valid argument from observations to universal claims. The most influential responses to the problem are Popper’s falsificationism (which accepts that theories cannot be confirmed by observation and proposes falsification as the mechanism of scientific progress) and Bayesian confirmation theory (which accepts that theories cannot be certain on the basis of finite evidence but proposes that they can be more or less probable, and that the probability should be updated by evidence according to Bayes’ theorem). Neither response eliminates the problem. Both manage it in different ways.

³ Lakatos, I. (1978). The Methodology of Scientific Research Programmes: Philosophical Papers, Volume 1. Cambridge University Press. The concept of the protective belt of auxiliary hypotheses (the set of subsidiary assumptions that can absorb anomalies without requiring revision of the core theory) is one of Lakatos’s most important contributions to the philosophy of science. It explains why falsification, in practice, is never the clean logical operation that Popper’s account describes: a single anomalous observation never straightforwardly refutes a theory, because the anomaly can always be attributed to a failure in the auxiliary hypotheses rather than in the core theory. Lakatos’s response was to evaluate not individual theories but research programs (sequences of theories that share a hard core of assumptions) and to distinguish progressive programs (in which the sequence of theories makes new predictions that are confirmed) from degenerating programs (in which the sequence of theories only accommodates anomalies post hoc without making new predictions).

⁴ Kuhn, T. S. (1962). The Structure of Scientific Revolutions. University of Chicago Press. The concept of incommensurability (the claim that competing paradigms cannot be evaluated against a common standard, because they disagree about what counts as a problem, what counts as evidence, and what counts as a solution) is the most controversial element of Kuhn’s account. Critics have argued that incommensurability implies relativism: that if paradigms cannot be evaluated against a common standard, then the choice between them cannot be rationally justified, and scientific progress cannot be understood as progress toward truth. Kuhn himself denied this implication but did not always succeed in making his denial convincing. The most successful attempts to reconcile Kuhn’s insights with a non-relativist account of scientific progress are found in the work of Imre Lakatos and in the scientific realist tradition associated with Hilary Putnam and Richard Boyd.

⁵ The distinction between the probability of data given the hypothesis (what the p-value measures) and the probability of the hypothesis given the data (what researchers typically want to know) is a specific instance of the base rate neglect described in article S1-210 of this series, the same cognitive error that produces the physicians’ mistakes in Tversky’s opening study. The formal name for the error is the confusion of the conditional probability P(data | hypothesis) with its inverse P(hypothesis | data). These two quantities are related by Bayes’ theorem but are not the same, and they can differ by very large factors when the prior probability of the hypothesis is low. The intuitive force of this distinction is captured in the prosecutor’s fallacy: the probability that an innocent person has blood type matching the evidence is not the same as the probability that the person is innocent given the matching blood type. Scientists who interpret a p-value of 0.05 as meaning there is a 5 percent chance the result is a false positive are making the same inferential error as the prosecutor who interprets a 1-in-1000 match probability as meaning there is a 1-in-1000 chance the defendant is innocent.

⁶ The p-value was introduced by the statistician Karl Pearson in the early twentieth century and codified as a decision threshold primarily through the influence of Ronald Fisher’s Statistical Methods for Research Workers (1925). Fisher set 0.05 as a convenient round number below which he considered results worth further investigation, not as a definitive criterion of truth. The subsequent institutionalization of 0.05 as the universal threshold of scientific significance is documented in Cowles, M., and Davis, C. (1982). On the origins of the .05 level of statistical significance. American Psychologist, 37(5), 553-558. The critique of the threshold has been developed most accessibly in Wasserstein, R. L., and Lazar, N. A. (2016). The ASA’s statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133, in which the American Statistical Association formally stated that the p-value does not measure the probability that the studied hypothesis is true, that it does not measure the probability that the data were produced by random chance alone, and that scientific conclusions should not be based only on whether a p-value passes a specific threshold.

⁷ The replication crisis has been documented most thoroughly in psychology, where the Open Science Collaboration’s large-scale replication project (2015) found that only 36 to 39 percent of published findings could be successfully replicated. Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. Similar findings have emerged in medicine, economics, and nutrition science. The structural causes of the crisis (publication bias toward positive results, underpowered studies, flexible analysis methods, and the incentive structure of academic publishing) are reviewed in Ioannidis, J. P. A. (2005). Why most published research findings are false. PLOS Medicine, 2(8), e124. The institutional responses that have been proposed and adopted include pre-registration of hypotheses, registered reports, open data requirements, and adversarial collaboration, each designed to restore the self-correcting mechanism’s reliability by reducing the specific failure modes that the p-value threshold and publication bias have introduced.

⁸ The pattern by which successor theories recover the empirical content of predecessor theories as special cases is documented across the history of physics and is one of the strongest arguments for scientific realism: the view that mature scientific theories are approximately true descriptions of the world rather than merely useful instruments for prediction. The philosopher Hilary Putnam called this pattern the convergence of reference: successive theories in a domain tend to refer to the same entities and to make increasingly accurate claims about them, even when the theories are otherwise very different. The most thorough philosophical treatment of this argument is Psillos, S. (1999). Scientific Realism: How Science Tracks Truth. Routledge. The counter-argument, that the history of science contains cases where successful predecessor theories referred to entities (phlogiston, caloric, the luminiferous ether) that the successor theory eliminated entirely rather than recovering, is the basis of the pessimistic meta-induction: the argument that if past theories turned out to be wrong about their central entities, we should expect current theories to be equally wrong. The most balanced assessment of the debate, which concludes that the convergence pattern is real and significant even if not universal, is Stanford, P. K. (2006). Exceeding Our Grasp: Science, History, and the Problem of Unconceived Alternatives. Oxford University Press.

Leave a reply

Your email address will not be published. Required fields are marked *