In 2015, a collaboration of 270 researchers published one of the most uncomfortable papers in modern science: an attempt to replicate 100 psychology studies, selected from three top journals and run with the original authors' input. Only 36 to 39 percent of the replications found statistically significant results in the same direction as the originals, and where effects did appear they were typically about half the original size. The paper — 'Estimating the reproducibility of psychological science,' in Science — did not prove that the original findings were false. It proved something subtler and more corrosive: the field's standard pipeline produced far more positive results than the evidence could support.
This is a methods explainer; it evaluates research practices, not any individual treatment or claim.
What exactly failed — statistics, or incentives?
The machinery of statistics was not broken; the pressures on it were. Analyses published in the early 2010s — most influentially by Joseph Simmons, Leif Nelson and Uri Simonsohn, who demonstrated they could 'discover' that listening to a song made participants younger — showed how flexible practices could conjure significance from noise: optional stopping of data collection, multiple outcome measures, subgroup slices, all reported as if chosen in advance. These were not usually frauds. They were ordinary researchers following incentives: journals wanted novel, positive results, and flexible analysis delivered them. Simmons and colleagues' paper, published in Psychological Science in 2011, became the field's mirror.
Did other fields find the same problem?
Yes, which is what makes the crisis a lesson about method rather than about psychology alone. The Reproducibility Project: Cancer Biology, published in eLife in 2018, repeated key experiments from 53 high-profile preclinical cancer papers; effect sizes were dramatically smaller in the replications, and several could not be completed as designed because original data and protocols were unavailable. Economics ran its own large replication efforts through the 2010s and found substantial — though smaller — rates of failure. The pattern across fields: published positive findings systematically overstate effect sizes, a phenomenon economists and methodologists describe as a distorted literature produced by publication and specification choices.
How strong is the evidence for all of this?
The replication crisis is itself unusually well-evidenced, by the field's own new standards. The Science paper involved 270 researchers, pre-specified replication designs, correspondence with original authors, and public data. The cancer-biology project was likewise preregistered and peer-reviewed. Follow-up analyses debated individual cases — some originals survived scrutiny, and a 2019 re-analysis argued the Science paper's own metrics could tell different stories — which is normal scientific argument, not a repudiation. The converging picture across psychology, cancer biology and economics rests on multiple large, independent, preregistered projects: about as solid as meta-science gets.
Related stories: Correlation Is Not Causation: How Study Design Decides the Argument · Underpowered Studies Find Too Much, and Too Big.
What did the term 'crisis' get right — and wrong?
Critics of the word have a point: the disciplines involved did not collapse, they audited themselves in public, funded the audit, and published results that made many of their own famous papers look weak. That is the system working under stress. What 'crisis' captures is the scale of the correction. A literature in which the positive findings are systematically overdrawn cannot be trusted at face value — clinicians, policymakers and journalists had been reading inflated effect sizes as if they were the real thing. The crisis framing also marks a change in who bears the burden of proof: before 2015, doubting a published result required extraordinary evidence; after it, the extraordinary evidence arrived, and skepticism became the default posture for findings that had only a single, small study behind them.
What actually changed afterward?
More than skeptics predicted. Preregistration — committing to hypotheses and analyses before data collection — moved from novelty to norm in social psychology, with platforms like the Open Science Framework hosting thousands of registered studies. Journals adopted mandatory data-sharing policies; the journal Psychological Science introduced badges for open data and materials. Large multi-lab projects — Many Labs, Psychological Science Accelerator — became an established way to test whether effects generalize across samples. Reporting of exact statistics and disclosure of all measured variables spread. Grant agencies moved too: several funders began requiring explicit plans for data management and transparent analysis. None of this guarantees truth; it raises the cost of fooling yourself, and it shortens the distance between a result and the raw data behind it.
Does a failed replication mean the original was fraud?
No, and conflating the two is the most common public misreading. A replication can fail because the original was a fluke, because the effect depends on unstated context — culture, timing, exact wording — or because the replication itself was imperfect. Replication debates are now often conducted in public, on preprint servers and social media, which is noisy but transparent. Fraud is a separate, smaller phenomenon; surveys and misconduct cases suggest outright fabrication is rare, while the epidemic was honest researchers exploiting analytic degrees of freedom.
What would confirm the fix is working?
The test is longitudinal: replication rates in cohorts of newer studies should exceed the 2015 baseline. Early multi-lab replications of recent, rigorously designed studies show substantially higher success rates than the 2015 cohort, though sample sizes remain modest. The durable lesson for any reader of research — in psychology or anywhere — is to ask of every striking finding: has anyone besides the original team produced it, at larger scale, with the analysis fixed in advance? Until the answer is yes, the honest reading is 'promising, single study' — a label psychology spent a decade learning to apply to itself.




