Skip to content
About Contact
Multi Academic JournalRESEARCH · ACADEMIA · SCIENCE EXPLAINED

Can Reviewers Use AI? What Journals Actually Allow Now

Most major publishers permit large language models to polish a reviewer's prose but forbid feeding confidential manuscripts into chatbots — and detection of secret AI use remains unreliable.

Can Reviewers Use AI? What Journals Actually Allow Now
Can Reviewers Use AI? What Journals Actually Allow Now

By 2025 every major publisher had issued a position on generative AI in peer review, and they converge on one line: reviewers may use large language models to improve their own writing, but they must not upload a confidential manuscript to an AI service, because doing so discloses unpublished work to a third party outside the review's confidentiality agreement. Elsevier, Springer Nature, and Wiley each state this in their reviewer guidelines; the World Association of Medical Editors and the Committee on Publication Ethics issued compatible advice in 2023. The rule exists because of how chatbots work — a prompt becomes data on someone else's servers, and unpublished manuscripts are, in copyright and confidentiality terms, other people's property.

This is an explainer on policy and risks; reviewers should always check the specific journal's current instructions.

Why the confidentiality rule is the bright line

Peer review runs on trust: authors hand unpublished work to strangers who compete with them. Uploading a manuscript to a chatbot recreates the classic breach — disclosure to outsiders — at zero marginal effort. The concern is not hypothetical in kind, only in scale; publishers had already warned reviewers against sharing manuscripts with colleagues, and a chatbot is simply one more unauthorized reader. Some AI vendors responded by offering enterprise agreements with training opt-outs, and a few publishers began contractually protected AI-assistance tools, but public chatbots remain outside the pale under most policies as of 2025.

What is AI actually used for in review?

Three legitimate uses recur. Language polishing for reviewers writing in a second language, which guidelines generally permit as long as the reviewer verifies every statement. Time-saving triage of a reviewer's own arguments — drafting structure, summarizing their own notes. And on the editor's side, automated screening for statistical red flags, image duplication, and undisclosed references, which several large publishers have run for years; image-integrity tools caught duplicated gels long before chatbots arrived.

Where are the real failure modes?

First, fabrication: large language models invent plausible-sounding objections, hallucinate references, and miss the subtle technical error a domain expert would catch. A 2024 wave of researchers testing chatbot-assisted reviews found the outputs fluent and often confidently wrong. Second, deskilling and volume: AI lowers the cost of producing a review that looks done, inviting reviewers to accept more manuscripts than they can genuinely judge — the same dynamic that made paper mills possible for authors. Third, detection is weak: several 2024-2025 studies found telltale AI phrases appearing in peer-review reports at measurable rates, meaning some fraction of reviews are already machine-polished or machine-written without disclosure.

Related stories: How Double-Blind Peer Review Actually Works, and Where It Fails · Special Issues: How Guest-Edited Volumes Work, and How They Get Hijacked.

How is AI screening different from AI reviewing?

The distinction carrying most of the weight in practice is between judgment and triage. Publishers have run automated screening for years with limited controversy: software that compares images across submissions to catch duplication, flags reference lists dominated by retracted papers, checks statistical reporting against common error patterns, and routes suspicious manuscripts to human editors. The human decides; the software narrows. Using a language model to draft a review reverses that direction of authority, placing generative judgment in the tool and verification on the human, which is the arrangement the confidentiality rules and reviewer guidelines were written against. Editorial commentary through 2024-2025 suggests the industry intends to hold that line while expanding the screening side considerably, and the vendors selling screening tools are marketing exactly that split: more automation upstream at triage, none at the point of judgment.

What do editors actually see?

Editors reporting on the first years of chatbot-era peer review describe a recognizable signature: reviews that are fluent, structurally tidy, generic in a particular way, and sometimes anchored to points the manuscript never makes. Detection heuristics built on phrase frequencies, the same statistical fingerprints used to find AI writing elsewhere, surfaced in internal audits and academic studies alike, and several editors noted reviews containing the characteristic vocabulary of language models at rates too high to be coincidence. The response has been procedural rather than punitive: reminder notices to reviewers, explicit attestation checkboxes on review forms, and in some venues, editor spot-checks of flagged reports. What no publisher has done is adopt AI-detection as grounds for rejecting a review automatically, because the error rates cut both ways and falsely accusing a volunteer reviewer carries its own costs.

What about authors using AI to write papers?

The same publishers bar listing AI systems as authors, on the grounds that authorship requires accountability a model cannot hold. Disclose-or-not policies differ by publisher, which leaves manuscripts in an inconsistent regulatory patchwork — one journal's required disclosure is another's irrelevance.

What would settle the debate?

Reliable evidence on whether AI-assisted review improves, or merely accelerates, the detection of real problems — trials comparing review quality with and without model assistance on identical manuscripts. Without that, the policies now in place are a confidentiality firewall, not a quality judgment. The line to watch: whether publishers move from banning uploads to licensing private models under confidentiality contracts, which would redraw the bright line entirely.

Frequently Asked Questions

Can peer reviewers use ChatGPT on a manuscript?
They may not upload the confidential manuscript to a public chatbot, per major publisher policies from Elsevier, Springer Nature, and Wiley. Using AI to polish the reviewer's own prose is generally allowed if every statement is verified and no manuscript content leaves the review environment.
Why is uploading a manuscript to an AI service a problem?
Peer review is confidential. A prompt sent to a chatbot transmits unpublished work to a third-party server, which is a disclosure outside the review agreement — the same breach as sharing the paper with an unauthorized colleague.
Can AI be listed as an author?
No. All major publishers prohibit attributing authorship to AI systems because authorship requires accountability and responsibility that a model cannot hold.