AI Screens 2.6 Million Cancer Papers and Flags Nearly 10% as Suspected Paper Mill Output
A BERT-based classifier trained on confirmed fraud cases found that the share of suspicious cancer papers climbed from roughly 1% in the early 2000s to more than 16% by 2022, according to a study published in The BMJ.
A machine learning tool built to detect the linguistic fingerprints of fraudulent research has produced the most sweeping audit of cancer literature on record, and the numbers are unsettling.
According to a study published in The BMJ, <cite index="19-2,19-3">a BERT-based classifier, trained on 2,202 confirmed retracted paper-mill papers, screened 2.6 million cancer studies published between 1999 and 2024, ultimately flagging 261,245 papers, or 9.87%, for suspicious writing patterns.</cite> That rate didn't hold steady over time. <cite index="19-3,19-5">The share of flagged papers rose from about 1% in the early 2000s to more than 16% by 2022.</cite>
The study was led by Adrian Barnett, a biostatistician at Queensland University of Technology. <cite index="20-7">The tool reads only a paper's title and abstract to make its call.</cite> That's a real constraint worth naming: the full methods section, where fabrication often hides, isn't part of the input. And as the researchers themselves acknowledge, <cite index="21-9">the findings are not confirmed cases of research fraud and should be checked by human specialists.</cite>
The mechanism the tool is hunting isn't subtle. <cite index="18-5,18-6">The flagged papers don't get labeled fraudulent outright, but they share linguistic fingerprints with papers linked to so-called paper mills, businesses that produce or sell scientific manuscripts, sometimes using fabricated or manipulated data.</cite> <cite index="21-7">Paper mills often use recycled text, awkward phrasing, or fabricated data and images.</cite>
Where the papers are landing is the part that collapses the comfortable assumption that this is a fringe problem in low-tier journals. <cite index="20-15">Of the top 20 journals the team examined, 19 turned up flagged papers; the sole exception was Nature Cancer.</cite>
The stakes aren't abstract. <cite index="21-10,21-11">Cancer research influences clinical trials, drug development, and patient care, and if fabricated studies enter the evidence base, they can mislead real scientists and slow progress for patients.</cite> That concern connects directly to a parallel problem surfacing at the peer-review stage. A feature published August 3 by Nature documents a separate but related phenomenon: <cite index="10-1">review mills, where researchers write fake referee reports coupled with coercive citation requests, are setting off a war in academic publishing.</cite> The scope of copied review language is measurable. <cite index="10-11,10-12">One analysis found that 0.05% of reviews received by Institute of Physics Publishing journals in 2025 were 100% identical, and a separate analysis of nearly 150,000 reports across MDPI, several PeerJ journals, The BMJ, and computer-science venues found that between 0.5% and 1.3% were highly similar to other reviews.</cite>
Taken together, the two datasets describe a pipeline under pressure at both ends: manuscripts produced industrially and, in some cases, sent through referees who are themselves participants in the fraud.
The BMJ classifier's reported accuracy is 91%, but that figure deserves scrutiny. <cite index="20-11,20-12,20-13">That 91% is accuracy against yesterday's style of fraud; the fabricated papers the model trained on followed formulaic templates, and now that generative AI can produce fake papers on demand, the template itself is dissolving.</cite> The arms-race framing isn't rhetorical. <cite index="19-6">Three journals are already testing BERT-based screening</cite>, but the adversarial pressure on any classifier will only increase as the forgeries get better.
<cite index="21-8">The team plans to expand the tool to other fields of research and improve the model as more confirmed cases of paper-mill activity become available.</cite> That expansion will tell us whether cancer literature is an outlier or a preview.
Sources cited:
- The BMJ (study: Barnett et al., BMJ 2026;392:e087581) (https://www.bmj.com/content/392/bmj-2025-087581)
- ScienceDaily (Queensland University of Technology, July 16 2026) (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- Nature career feature: 'A waste of time for all of us' (August 3 2026) (https://www.nature.com/immersive/d41586-026-01360-8/index.html)
- ecancer (July 2026) (https://ecancer.org/en/news/27724-new-tool-exposes-scale-of-fake-research-flooding-cancer-science)
This release was originally distributed via ETL Newswire. Visit The BMJ (study: Barnett et al., BMJ 2026;392:e087581) for the full story, related releases, and contact information.
Visit The BMJ (study: Barnett et al., BMJ 2026;392:e087581) →