Statistical power is the probability that a study detects a real effect of a given size — and the uncomfortable arithmetic is that a study with 20 percent power will miss a true effect four times out of five, while any 'discovery' it does report is likely to be an exaggeration. Researchers conventionally aim for 80 percent power, which means accepting a one-in-five chance of missing something real. Yet surveys of published research have suggested many studies run far below even that bar: a widely cited 2013 analysis by Katherine Button and colleagues in Nature Reviews Neuroscience estimated the median statistical power in a sample of neuroscience studies at roughly 21 percent.
This is a methods explainer; it does not evaluate any specific clinical claim or treatment.
What determines the power of a study?
Four ingredients. The size of the true effect — big effects are easy to see, subtle ones are not. The sample size. The measurement noise — unreliable instruments and messy protocols bury signals. And the significance threshold the field demands, conventionally p <0.05. Before running a study, researchers can compute the sample needed to detect the smallest effect they would care about at 80 or 90 percent power. The catch is that this calculation requires assuming the effect size, and if the assumed effect is borrowed from small, optimistic earlier studies, the 'powered' study inherits the optimism and remains underpowered in fact.
Why do underpowered studies overstate effects?
This is the counterintuitive part, and it follows directly from how discovery works under noise. A low-power study only reaches significance when random fluctuations happen to push its estimate far above the truth — the wins that get published are the outliers. Methodologists, including Andrew Gelman and John Carlin, formalized this as 'Type M' errors: exaggeration factors among significant results from underpowered designs. Combine small samples with publication preference for positive findings, and the published literature converges on a systematically inflated picture of effects, which then misleads the power calculations of the next generation of studies. The cycle is self-sustaining unless someone runs a large, well-powered replication.
How strong is the evidence on low power?
The claim rests on more than the Button analysis, though that paper is the anchor: it aggregated meta-analyses and systematic reviews across neuroscience rather than cherry-picked single studies. Its estimates were later debated — some statisticians argued the methodology overstated the problem — and that debate is part of healthy meta-science. Independent lines of evidence converge: replication projects found that original studies with smaller samples showed the largest drops when repeated; meta-analyses routinely display the 'small-study effect,' where small studies report bigger effects than large ones; and direct measurements of published power distributions, though model-dependent, consistently place a large share of studies below conventional targets. The direction of the finding is settled even where the exact numbers are argued.
Related stories: A 'Significant' Result Can Be Trivial. Effect Size Tells You Which. · Correlation Is Not Causation: How Study Design Decides the Argument.
How does power connect to the replication crisis?
Directly. The replication projects of the 2010s found that the studies least likely to reproduce were those with the smallest original samples, exactly as power arithmetic predicts: a design that barely skates past the significance threshold is sampling from the extreme tail of possible outcomes, and tails do not repeat. The same logic explains why replications, run with larger samples, usually produce estimates closer to zero than the originals. None of this required anyone to have behaved dishonestly. It is what happens when a profession's standard unit of evidence — the small, independent study — is mathematically incapable of measuring the effects it hunts, yet is asked to publish binary verdicts anyway.
Is the answer simply to demand null results?
Partly. Publication systems that accept well-designed studies reporting no effect reduce the survivorship bias that inflates literatures — if only outliers get printed, the printed record is the outliers. But null results only inform when the study was powered to detect the effect; a null from a tiny sample is uninformative rather than reassuring. That is why reform proposals pair open publication of null findings with mandatory power justification, so that 'we found nothing' carries the meaning 'a real effect of at least this size would probably have been detected.' The combination — adequate power plus outlets for unflattering results — is what turns individual studies into cumulative evidence.
Why doesn't bigger data solve everything?
Size helps power but not bias. A million-row observational dataset can be exquisitely powered for a confounded comparison, and big samples even worsen one problem: with enormous n, trivially small effects become 'significant,' which is why large-sample research must lean on effect-size interpretation rather than thresholds alone. Power thinking and design thinking are complements — randomization, blinding and pre-specified analyses determine what the estimate means; sample size determines how precisely it is measured. The rule of thumb that survives every methodological argument: no single number in a paper, however small its p-value, means much until you know how precisely the study could have measured the thing it claims to have found.
What should a reader look for?
First, the sample size and whether the paper reports a power calculation. Second, the confidence interval rather than the p-value: a wide interval means the study could not measure the effect precisely, whatever the verdict. Third, the pattern across the field — if small studies cluster at dramatic effect sizes and large studies shrink toward zero, the small ones are probably decorating noise. Fourth, whether the study was designed before data collection, which removes the flexibility that low power punishes so badly.
What would change the picture?
Larger collaborative studies, registered reports with peer review before data collection, and journals that accept well-designed null results all cut the incentive to run and publish underpowered fishing expeditions. The check is empirical: effect sizes in replications should approach the originals as samples grow. Where that convergence appears, the literature is healing; where it does not, the inflation diagnosis was wrong — and either way, the test is the same: replicate with power.




