Science and Error Correction
Status: draft.
Opening Question
In 2015, a large collaborative effort attempted to replicate 100 published findings from top psychology journals. Fewer than half replicated. This was reported, correctly, as an embarrassment for psychology — but the more interesting fact is what happened next: the finding was published, widely discussed, and used to drive real reforms in how the field operates. A system that can publicly discover and correct its own failure rate is doing something most human institutions cannot. What makes science — imperfect, slow, and often wrong in the short run — unusually good at being wrong productively?
Historical Perspective
The institutional machinery of modern science is younger than the idea of empirical inquiry itself. The Royal Society of London, founded in 1660, pioneered peer review and the published, dated scientific paper specifically to establish priority and allow claims to be checked by others — a social technology as important to science's success as any single discovery it produced. Karl Popper's The Logic of Scientific Discovery (1934) later gave this practice a philosophical anchor: falsifiability. A claim only counts as scientific, on Popper's account, if it is structured so that some possible observation could prove it wrong. This is a design principle for knowledge claims, not merely a description of them — it defines science by its willingness to be disproven, which inverts the everyday human instinct to defend a belief once held.
Robert Merton's 1942 essay on the norms of science distilled the social side of this into four principles, since remembered by the acronym CUDOS: Communalism (findings belong to the community, not the discoverer, once published), Universalism (claims are judged by evidence, not the status of the claimant), Disinterestedness (institutional incentives should reward accurate findings over convenient ones), and Organized Skepticism (every claim is subject to structured, public challenge). Merton described these as norms, not laws of nature — meaning they hold only as long as an institution actively enforces them.
What Modern Evidence Suggests
- The Open Science Collaboration's 2015 replication study, published
in Science, found that only 36% of a sample of 100 psychology studies replicated with a statistically significant effect in the same direction, and replicated effect sizes were on average about half the originally reported size. Comparable, if less severe, replication problems have since been documented in parts of biomedicine and economics.
- The underlying causes are reasonably well understood: **publication
bias (journals favor novel, positive findings over null results or replications), p-hacking and researcher degrees of freedom** (Simmons, Nelson & Simonsohn's 2011 paper on "false-positive psychology" showed how easily flexible analysis choices, applied even in good faith, inflate false-positive rates), and career incentives that reward publication volume and novelty far more than accuracy or replication.
- Reforms following the crisis offer a live, ongoing case study in
Merton's "organized skepticism" being deliberately re-engineered: pre-registration (committing to a hypothesis and analysis plan before seeing the data), registered reports (journals agreeing to publish based on methodology before results are known, removing the incentive to only submit positive findings), open data and code requirements, and dedicated replication journals and funding lines.
Where the Principle Fails
Falsifiability and replication are not free of their own failure modes. Popper's criterion, applied too rigidly, can be used to dismiss early-stage or historical sciences (much of evolutionary biology, cosmology, and geology rely heavily on inference from evidence rather than controlled experiment) as somehow less legitimate, when in practice these fields still generate falsifiable, testable predictions — just not always via a laboratory experiment. And replication itself can be gamed: a well-funded actor can commission many replication attempts and publicize only the ones that suit them, weaponizing the language of rigor exactly as Chapter 2's truth-infrastructure failures warn against.
Civilization Design Principle
> Reward the correction of an error at least as highly as the original > discovery, and build institutions where that correction is cheap to > perform and costly to suppress.
Science's real innovation is not that its practitioners are unusually honest — Merton and later sociologists of science were explicit that individual scientists are as prone to bias, ambition, and self-deception as anyone else. The innovation is a set of institutional incentives (peer review, replication, publication norms, professional reputation tied to being provably right over time) that make error-correction survivable and even rewarding for the community as a whole, independent of any one scientist's character.
Institutional Translation
- Funding bodies that dedicate a fixed share of budget to replication
and null-result studies, not only novel research.
- Registered reports as a standard submission track across academic
publishing, not a niche option.
- Open data and code mandates as a condition of public research
funding, enabling independent verification.
- Beyond academia: applying the same logic to **corporate and government
research** — internal red-teaming, adversarial review, and published post-mortems as standard practice rather than exceptional crisis response (see 09_learning_governments.md).
Metrics
- Replication rate by field, tracked longitudinally as reforms take
effect.
- Share of funded research dedicated to replication versus novel
discovery.
- Time from a published error's discovery to formal correction or
retraction.
- Career outcomes for researchers who publish null results or
successfully challenge prior findings, as a proxy for whether organized skepticism is actually rewarded or merely praised.
Questions Still Unresolved
Replication is expensive, and no field can replicate everything — some triage is unavoidable, which means someone has to decide which findings are important enough to check, and that gatekeeping function is itself a potential point of capture. There is no settled answer in the literature reviewed here for how to allocate scarce replication effort without reintroducing the same status and incentive biases the reforms are meant to correct.