The National Science Foundation operates on a budget of about $9 billion a year, spread across all non-medical fields of basic research, and its core transaction — the merit-reviewed research grant — follows a machinery that most of the public never sees. A proposal arrives at a program officer; it is mailed to several working scientists for written review; a panel of experts then meets, discusses batches of proposals, and scores them; the program officer weighs the scores, the reviews and the portfolio, and recommends awards upward through division management. Funding rates at NSF have recently run near or below one in five. Every stage of that pipeline is designed, and each stage's design explains a lot of what applicants experience as mystery.
This is an explainer on federal grant mechanics, not application coaching.
What are the two merit review criteria?
NSF evaluates every proposal on two axes: Intellectual Merit — the potential to advance knowledge — and Broader Impacts — the potential to benefit society, which in practice covers training, broadening participation, infrastructure, and outreach. The two-criteria structure dates to the 1990s and was codified in America COMPETES legislation of 2007. Reviewers rate each criterion on a five-step scale from Excellent to Poor. The dual criteria are distinctive: NIH reviews chiefly scientific significance and feasibility, while NSF asks its panels to price the social return explicitly. Broader Impacts is perennially misunderstood — it is not a garnish added to the last page but a second argument running through a competitive proposal, and panels routinely flag proposals that treat it as one.
Who reads the proposal, and what do they see?
Typically three to five ad hoc reviewers receive the proposal by mail and return written reviews; then a rotating panel — working scientists serving limited terms, brought to Arlington or convened virtually — reads an assigned batch. Panels compress discussion into minutes per proposal, which is why the written reviews matter so much: the panel scores reflect the reviews as much as the panelists' own reading. Program officers select the reviewers and panelists, seeking expertise without conflicts of interest, and they manage an uncomfortable constraint: the pool of qualified reviewers is the same population as the applicant pool, meaning every scientist is simultaneously judge and judged. The system's integrity rests on rotation, conflict rules and disclosure requirements rather than on any firewall that fully separates the two roles.
What happens after the panel?
The program officer writes a recommendation, weighing panel consensus, reviewer expertise, and portfolio balance across institutions, career stages and subfields — a discretion that surprises applicants who assume scores mechanically determine outcomes. Division directors and division sign-off follow, with a formal budget and grants-officer review. Declines arrive with the reviews and, on request, the panel summary. The statistics tell the structural story: overall funding rates have hovered near or below 20 percent recently, and highly ranked proposals routinely go unfunded when the division's budget runs out — which is why 'very good' reviews and a decline can coexist honestly in the same envelope.
Related stories: One Advisor, Five Years, No Recourse: Academia's Most Concentrated Job · Pay to Present: Who Actually Affords an Academic Conference.
How is NSF different from NIH, its bigger sibling?
The contrast clarifies both systems. NIH's budget is several times larger, organized around disease institutes, and its review uses standing study sections whose members serve years, with scoring on a fine numeric scale; NSF uses short-rotating panels and letter grades, spreading smaller awards across all sciences, engineering and education. An NIH R01 typically funds one laboratory's project deeply; an NSF grant more often funds a graduate-student-centered project with training and outreach attached, reflecting the Broader Impacts mandate. Proposal lengths, submission cadences and resubmission rules also differ, which is why interdisciplinary applicants frequently misjudge one agency using the other's culture. The two agencies together anchor most US academic basic research, and their structural differences — not just their budgets — determine what kinds of science get done.
How strong is the evidence that the system selects well?
Moderately strong that it selects sensibly, thin that it selects best. Grant-review reliability studies find modest but real correlations between panel scores and later research impact, and reviewers agree more within panels than across them, suggesting the process converges on shared local standards. The loudest critique — that peer review favors established researchers and conventional projects — has empirical support: funded applicants skew senior, and NSF's own experiments with fully randomized lotteries for near-threshold proposals in New Zealand and the Netherlands revealed how much near-identical quality sits on either side of the funding line. NSF has run its own variation, holding some applications to different review models, and continues to study reviewer diversity and workload. The honest verdict of that literature: peer review is better than chance and worse than its reputation.
What should an applicant actually understand?
That the proposal is written for two readerships at once: the ad hoc reviewer with weeks and the panelist with minutes — which is why a crisp first page, explicit Broader Impacts, and a budget that reads as designed rather than aspirational are load-bearing. That program officers are approachable before submission and are the single most useful source of fit questions. And that a decline with strong reviews is usually a budget event, not a verdict — resubmission to the next cycle is the system's expected rhythm, not its exception. NSF publishes its merit review process and data openly, which makes it, for all its frictions, one of the more inspectable funding machines in science.




