AI Screen Flags Nearly 10% of Cancer Research Papers as Suspected Paper Mill Output
A BERT-based classifier trained on retracted paper-mill articles scanned 2.6 million cancer studies and flagged 261,245 of them, with rates climbing from roughly 1% in the early 2000s to more than 16% by 2022.
A machine-learning tool built at Queensland University of Technology has put a number on a problem the research-integrity community has long suspected was large but couldn't measure: roughly one in ten cancer research papers published over the past quarter century shows writing patterns consistent with suspected paper-mill fabrication.
<cite index="18-1,18-2">A team led by biostatistician Adrian Barnett published findings in The BMJ reporting that a BERT-based language model screened 2.6 million cancer studies published from 1999 through 2024 and flagged 261,245 papers, or 9.87%, for writing patterns associated with suspected fabrication.</cite> The paper is identified in The BMJ as Scancar et al., BMJ 2026;392:e087581.
The method is deliberately lean. <cite index="17-8">The tool reads only a paper's title and abstract to judge whether it is fabricated</cite>, which makes it fast enough to process millions of records but means it can't inspect data tables, figures, or supplementary files where other fraud signatures often hide.
<cite index="18-3,18-4">The model was trained on 2,202 retracted paper-mill articles drawn from a retraction database and then validated against independent expert datasets, reaching 91% accuracy in identifying papers that matched the retracted template style used as the training signal.</cite> That's a real limitation: <cite index="17-12,17-13,17-14">the 91% figure is accuracy against yesterday's style of fraud. The fabricated papers it trained on followed formulaic templates, and the detector learned the fingerprint of that prose. Now that generative AI can produce fake papers on demand, the template itself is dissolving.</cite>
The trend line is what should trouble journal editors most. <cite index="18-5">The share of flagged papers rose over time, from about 1% of annual cancer publications in the early 2000s to more than 16% by 2022.</cite> That trajectory predates the current wave of large language models, so the 2022 figure may already be a floor, not a ceiling.
Cancer subtype matters too. <cite index="18-6">Flagging rates varied by cancer type: gastric cancer papers were flagged at about 22%, bone cancer at about 21%, and liver cancer at about 20%.</cite> The authors' preprint, posted to bioRxiv ahead of the BMJ publication, explains that high publication pressure, data that's relatively straightforward to fabricate, and limited peer-review capacity all make these subfields easier targets.
The journal-prestige assumption doesn't hold either. <cite index="17-15,17-16,17-17">Of the top 20 journals the team examined, 19 turned up flagged papers, with the sole exception being Nature Cancer. Even journals in the top 10% by impact factor were not a safe zone.</cite>
Barnett is careful about what the flag actually means. <cite index="25-8,25-9">Throughout the paper, the authors use "flagged" to refer to articles whose titles and abstracts are textually similar to retracted papers tagged as paper mills in Retraction Watch, and they are explicit that flagging is a statistical screen, not an attribution of misconduct.</cite> A flagged paper isn't automatically fraudulent; it's one that warrants human review. At 261,000 papers, there's no realistic editorial workforce to do that manually, which is precisely the trap the field is in.
<cite index="8-2,8-3">Nature reported this month that review mills, researchers who write fake referee reports with coercive citation requests, are setting off a war in academic publishing, raising the question of what happens to those caught in the middle.</cite> Barnett's BMJ study adds a quantitative spine to that coverage: the pipeline from submission to citation is compromised at a scale the field hasn't previously put a number on.
The honest read of this paper is that it's a lower-bound estimate built on a detector trained on an older fraud signature, applied to a literature that's getting harder to screen. Replication across other biomedical fields would be the logical next step, and the authors say their tool could support editorial triage. Whether journals will actually deploy it, and how they'll handle the queue of suspects if they do, is a policy question the study pointedly leaves open.
Sources cited:
- The BMJ (Scancar et al., BMJ 2026;392:e087581) (https://www.bmj.com/content/392/bmj-2025-087581)
- ScienceDaily, AI flags more than 250,000 suspicious cancer research papers (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- Pebblous AI Blog, An AI Detector Screened 2.6 Million Cancer Papers and Flagged 260,000 Fakes (https://blog.pebblous.ai/blog/cancer-paper-mill-ai-detector/en/)
- GNCrypto News, AI model flags 261,245 cancer papers for template writing (https://www.gncrypto.news/news/ai-model-flags-261245-cancer-papers-template-writing/)
- bioRxiv preprint, Revealing the Paper Mill Iceberg (https://www.biorxiv.org/content/10.1101/2025.08.29.673016.full.pdf)
- Nature News (https://www.nature.com/news)
This release was originally distributed via ETL Newswire. Visit The BMJ (Scancar et al., BMJ 2026;392:e087581) for the full story, related releases, and contact information.
Visit The BMJ (Scancar et al., BMJ 2026;392:e087581) →