Skip to content

Human Questions

What Is the Replication Crisis?

An introduction to the replication crisis — the recent discovery that many published scientific findings fail to reproduce — covering its causes, from small samples and p-hacking to publication bias, and the reforms designed to restore trust in science.

Quick Answer

The replication crisis is the recognition, beginning around 2011, that many published scientific findings — especially in psychology, medicine, and economics — fail to reproduce when the studies are repeated. Large-scale replication projects found success rates below 50 percent in some fields. The crisis was caused by small samples, flexible analysis, selective reporting, and publication bias, and it has prompted reforms including pre-registration, open data, and replication studies.

replication-crisisreproducibilityp-hackingpublication-biasscientific-methodmeta-science

Key Takeaways

  • The replication crisis is the widespread failure of published findings to reproduce
  • Large replication projects found success rates below 50 percent in some fields
  • Small samples, p-hacking, and publication bias are the main culprits
  • Replication is the foundation of scientific trustworthiness
  • Reforms include pre-registration, open data, and incentive changes

Direct Answer

The replication crisis is the recognition, gathering force since around 2011, that a substantial fraction of published scientific findings do not reproduce when the original studies are repeated. Replication is the test of a result's robustness: independent researchers run the same procedure, with new data, and ask whether the same finding appears. When the Open Science Collaboration attempted to replicate one hundred psychology studies in 2015, fewer than half produced statistically significant effects in the same direction, and the average effect size dropped by about half. Similar patterns emerged in medicine, economics, and the biomedical sciences. The crisis does not mean that science is fraudulent or that its results are worthless; it means that the published literature contains many claims that are weaker — or false — because the research system rewarded striking results over reliable ones.

Historical Context

The ideal of replication is as old as the scientific method itself. Francis Bacon made repeatability part of experimental philosophy, and the Royal Society institutionalized it: a result is accepted when others can reproduce it. Karl Popper built the ideal into the logic of science: a falsifiable claim is one that repeated testing could refute, and a single non-replication is, in principle, a refutation. Yet for most of the twentieth century, replication was undervalued in practice. Journals preferred novel, positive findings; replication studies were hard to publish; and individual researchers had no career incentive to check others' work. The turning point came in the 2010s: first in social psychology, where several famous findings — including priming and ego-depletion effects — failed to replicate, and then across disciplines, as meta-scientific studies quantified the scale of the problem. The crisis was, in a sense, the scientific method turning its own tools of criticism and testing onto itself.

Key Distinctions

Three clusters of causes explain most failures to replicate. Statistical problems: small sample sizes make studies underpowered — too weak to detect real effects reliably — so a "significant" result may reflect noise; p-hacking, analyzing data in flexible ways until a p-value crosses the threshold, manufactures significance; and underpowered studies that do reach significance are likely to exaggerate the true effect size. Incentive problems: publication bias means journals publish mostly positive results, burying null findings; careers reward productivity, novelty, and press coverage; and there was little penalty for sloppy but honest analysis. Questionable research practices: HARKing (hypothesizing after results are known), selective reporting of the studies that worked, and undisclosed flexibility in analysis all inflate the apparent rate of real effects. None of these requires fraud; they are features of a system that rewarded discovery-signals over error-detection.

Contemporary Debates

The crisis has provoked both reform and controversy. The reform movement — often called open science — includes pre-registration (stating hypotheses and analysis plans before data collection, so confirmations cannot be manufactured afterward), registered reports (journals committing to publish results regardless of outcome), open data and materials, and routine replication studies. Large consortium efforts now replicate findings in psychology, economics, and cancer biology. At the same time, defenders of affected fields argue that the replication rate is not as dire as headlines suggest: many replications change methods, and statistical power varies by field. There is also debate about what a replication failure means — a false original, a flawed replication, or a real effect that does not generalize — which mirrors the older philosophical problem that Popper and Kuhn debated: science is a social system, and its reliability is maintained by institutions, not by individual genius.

Practical Implications

For the reader of science, the crisis is a lesson in intellectual humility and method. Do not treat a single study, however prestigious, as proof: look for converging evidence across many studies, and prefer findings that have been replicated by independent teams. Check whether the study was pre-registered, whether the sample was adequate, and whether effect sizes were modest. Distrust claims built on very small samples or surprising results that nobody has tried to reproduce. For those inside science, the crisis demands reforms that are now spreading: pre-registration, open data, and incentives that reward reliability as well as novelty. The deeper lesson is Popperian: science is not the accumulation of confirmed truths but the critical testing of claims, and a system that stops testing its own results has stopped being scientific.

Further Learning

Knowledge Network

Archive references

Sources

3 scholarly sources
  • 01
    Reproducibility of Scientific ResultsBy Stanford Encyclopedia of PhilosophyConsult source
  • 02
    Replication CrisisBy Internet Encyclopedia of PhilosophyConsult source
  • 03
    Reproducibility and Replicability in Science (OUP)By National Academies Press / Oxford University PressConsult source

ZHAIBIAN Editorial Board reviewed

Reviewed by ZHAIBIAN AI Editorial Review · 2026-08-10

Based on 3 scholarly sourcesLast updated 2026-08-10