Insights and Analytics

AI Isn't Writing Peer Reviews. It's Polishing Them.

Jared Rand

A long-time reviewer announced this year that he was done with ICLR and ICML. His reason: the conferences had become "Track Chairs using AI to synthesize AI reviews of AI generated papers."

We checked the papers half of that claim last month. arXiv's flood turned out to be people, not machines: 28x more researchers, each publishing 9% more than a decade ago, producing work that is almost never duplicative.

So we went and got the reviews. 24,479 of them, from ICLR, NeurIPS, COLM and MIDL, 2023 through 2026.

Something did change. It is large, it is recent, and it is not what the complaint assumes.

Get an email when we publish a new post. No account needed, unsubscribe anytime.

The thing you can measure

Em dash prevalence in peer reviews by venue and year

An em dash requires a deliberate keystroke. Option-shift-hyphen on a Mac, an alt code on Windows. Essentially nobody produces one while typing into a review box. Curly quotation marks come from Word, from Google Docs, or from a chat interface. Neither survives being typed directly into OpenReview's plain text field.

So their presence tells you something specific: this text was written somewhere else and pasted in.

Venue 2023 2024 2025 2026
ICLR 6.1% 12.2% 36.5%
NeurIPS 5.2% 6.4% 22.4%
COLM 20.1% 27.2%
MIDL 20.3% 31.0%

More than a third of ICLR 2026 reviews now carry it, up sixfold in two years. Every year-over-year step clears p < 1e-22.

It is not one conference's culture. NeurIPS, COLM and MIDL all move the same way. The shift tracks the calendar, not the venue.

It is also not an artifact of reviews getting longer: em dash density rose 8.4x, steeper than the raw rate.

The thing that actually matters

Typography proves the text was composed elsewhere. It cannot tell you whether a reviewer pasted their own review through a model to fix the grammar, or asked a model to write the review.

That distinction is the whole argument. One is a spell-checker. The other is an abdication.

So we tested it. We had Claude Opus 5 read 899 reviews blind, never showing it the year, and score each one from 0 to 100 on how likely it was that a machine substantially drafted it. We told it to judge the writing rather than the review's quality, and to treat specific engagement with the paper as strong evidence of a human.

Distribution of blind judge scores, ICLR 2024 vs 2026

Band ICLR 2024 ICLR 2026
0-29 (clearly human) 84.3% 53.7%
30-49 6.0% 13.0%
50-69 3.0% 11.7%
70-84 4.7% 18.0%
85-100 (clearly machine) 2.0% 3.7%

Suspicion tripled. Reviews scoring 50 or above went from 9.7% to 33.3%.

And the top band did not move. 2.0% to 3.7%, p = 0.22. Not significant.

The mass moved out of "clearly human" and into the middle. It never arrived at "clearly machine."

That is what polishing looks like. A review genuinely drafted by a model, all generic praise and symmetric scaffolding over nothing, lands in the top band. The top band is precisely where a drafting story would show up, and it is flat.

One detail matters for trusting this: the judge reported keying on typography in 0.0% of cases. Its stated signals were substance (69%) and structure (31%). Reviews with em dashes score only 13 points higher than those without. It is not secretly reading punctuation back to us.

Everything else moved the wrong way

What rose and what fell in ICLR reviews, 2024 to 2026

If machines were writing these reviews, machine vocabulary should be rising alongside the typography. It fell.

Signal 2024 2026
Any paste-typography 66.2% 83.9%
Curly quotes 24.0% 47.5%
Em dash 6.1% 36.5%
Judge suspicion (50+) 9.7% 33.3%
Judge high-confidence (85+) 2.0% 3.7%
Model vocabulary (delve, showcase) 22.6% 15.9%
Non-native English markers 11.4% 8.7%
"not X, but Y" phrasing 2.5% 1.9%

The word delve, the most-cited machine-writing tell of 2023, fell from 3.1% of ICLR reviews to 0.4%.

There is an obvious reading. Newer models were trained away from the vocabulary that made them obvious, and nobody thought to train them away from their punctuation. Which means every vocabulary-based AI detector is quietly rotting, and our numbers are a floor rather than a ceiling.

Who is actually doing this

Machine polish is most valuable to people for whom writing fluent English is work.

That is a testable prediction. If the researchers reaching for these tools are disproportionately non-native English speakers, then the grammatical fingerprints polish removes should fade over the same window that polish rises.

Non-native English markers by venue and year

They do. Every marker falls, in every year, at both venues.

Marker 2024 2025 2026
Plural agreement slip 6.57% 5.87% 5.11%
Third-person agreement slip 2.83% 2.67% 2.21%
"authors" without "the" 1.31% 1.14% 0.88%
Any marker 11.38% 10.24% 8.71%

Down roughly a quarter (p = 2.4e-05). We restricted this to constructions that are actually ungrammatical, because many major languages of the machine-learning community, including Chinese, Japanese, Korean and Russian, have no articles at all. Article and agreement errors are the highest-precision signal available.

This measures text, not people. Reviewers are anonymous and nothing here identifies anyone.

But the implication is uncomfortable for the slop narrative. The tell a fluent English speaker notices, that smooth symmetric em-dashed prose, may be picking out exactly the reviewers who had the most to gain from the tool and the least to do with laziness.

A field that starts reading polish as evidence of not caring will land hardest on the people who were already working hardest to be understood.

What this means

Same people. Different incentives. The researchers writing arXiv papers and the researchers writing OpenReview reviews are largely the same population. Papers carry your name and your reputation. Anonymous reviews carry neither. Only the incentives flip between the two, and the behaviour flips with them.

Reviewers are not sending machines to do their thinking. They are sending their thinking through a machine on the way out. On this evidence the judgement is still human. What changed is the prose.

The original complaint is half right, and the half it gets wrong matters. Reviews really have changed. But the tell being reacted to is the writing, and the writing is downstream of somebody who still read the paper.

Whether that constitutes a collapse depends entirely on where you think the value of a review lives. If it lives in the judgement, this data says the judgement is intact. If it lives in the prose, it was never really there.

Can you trust the judge?

A blind judge with no calibration is just an opinion with a number attached, and a lot of this rests on it. So we checked it against writing whose origin we already know.

Judge score distributions for known-human, mainstream and known-machine content

Three sets of blog articles, same judge, same scale:

Set n Mean 70+
Known human, published on or before 2021 319 12.9 5.0%
Mainstream 2026 publications 319 57.6 48.6%
Known content-farm output, 332 denylisted domains 318 96.1 100.0%

Nothing published in 2021 was written by ChatGPT, so that first set is real ground truth rather than an assumption.

The judge separates known machine content from known human writing at AUC 0.998. At the threshold we used, it catches 100% of content-farm output while flagging 5% of pre-ChatGPT human writing. It also tells farm content apart from era-matched real publications at 0.971, so it is not simply detecting "written recently."

That matters for the peer review numbers. ICLR reviews sit far below the machine signature in both years: 2024 at a mean of 18.6, near the human baseline of 12.9, and 2026 at 37.5. Neither looks anything like genuinely generated text. And the load-bearing result was the flat top band, which is exactly where this judge is sharpest, since it puts 98.7% of real machine content above 85. If reviews were being drafted by models at scale, the band the judge is best at would have moved.

Two honest limits. Content-farm output is fully generated SEO filler, a much cruder thing than the polish-versus-drafting distinction, so passing this test does not transfer calibration across genres. And mainstream 2026 publications score 57.6, nearly half the slop score. Either real publications are heavily AI-assisted now, or the judge partly keys on stylistic conventions that arrived after 2022 regardless of who wrote them. We cannot separate those, and that ambiguity could account for some of the 18.6 to 37.5 rise. It cannot account for the top band staying flat, because drift would have pushed that up too.

One label note worth passing on: the 333-domain denylist we used contains dev.to, a legitimate developer community. We excluded it. Everyone else on the list matches the disposable-cloud-domain pattern.

What we could not measure

Four things we tried and could not make work, reported because each is a trap.

Open-source AI detectors point the wrong way. We scored reviews with the GLTR/DetectGPT approach under GPT-2. Machine-written text should be more predictable. ICLR 2026 reviews came out less predictable. We tested the obvious explanation, that GPT-2 never saw post-2022 vocabulary, by restricting to reviews with no modern terminology. The gap survived. The detector is not confounded, it is miscalibrated: "machine text has low perplexity" was calibrated against 2019-era output, and 2026 model prose is simply out of its distribution.

You cannot compare papers to reviews on typography. arXiv abstracts contain 0.0% em dashes against 36.5% in reviews, which looks like a spectacular finding and is entirely LaTeX. The TeX pipeline converts em dashes to hyphens before anyone sees the text.

Reviewer workload is unmeasurable. OpenReview anonymises reviewers per submission, so the same person carries a different identifier on every paper. Across 33,844 notes, only 19 persistent identifiers appear. The question of whether individual reviewers are drowning cannot be answered from public data at any sample size.

Meta-reviews are inconclusive. They score highest of all with our judge, but the judge's criteria include "generic, restates rather than evaluates," which is a description of what a meta-review is. That is a genre confound, not a finding.

Full methodology, data and code

Related posts

Get more posts like this

Subscribe to get new posts by email. No account needed, unsubscribe anytime.