Published by Emerging Technologies Laboratory · via ETL Newswire
Science· 

AI Screener Flags Nearly 10 Percent of Cancer Literature as Potential Paper Mill Output

A BERT-based tool trained on retracted studies scanned 2.6 million cancer papers and found more than 261,000 bearing the textual fingerprints of fraudulent 'paper mills,' according to a study published in The BMJ.

By Dr. Maya Iyer, Staff Reporter · Science Desk

A machine learning system built to detect scientific fraud has returned a number that's hard to sit with: roughly one in ten cancer research papers published over the past quarter-century may carry the writing patterns of fraudulent paper mills.

<cite index="14-1,14-2">A team led by Queensland University of Technology biostatistician Adrian Barnett published findings in The BMJ on August 4, 2026, reporting that a BERT-based language model screened 2.6 million cancer studies published from 1999 through 2024 and flagged 261,245 papers -- 9.87 percent -- for writing patterns associated with suspected fabrication.</cite> That's not a retraction count. It's a flag count, and the distinction matters.

<cite index="12-5,12-6">The findings don't suggest those studies are fraudulent. What they do suggest is that many share the same linguistic fingerprints as papers linked to so-called 'paper mills' -- businesses that produce or sell scientific manuscripts, sometimes using fabricated or manipulated data.</cite>

The detection method leans on a familiar logic. <cite index="11-6,11-7,11-8">The team trained a BERT model to identify subtle textual 'fingerprints' that repeatedly appear in known paper mill products. When evaluated using verified examples, the model correctly detected suspicious papers 91 percent of the time. Barnett compared it to email spam filtering: the tool flags papers matching the writing style and structure seen in retracted, fraudulent work.</cite>

<cite index="14-3,14-4">The model was trained on 2,202 retracted paper-mill articles drawn from a retraction database and then validated against independent expert datasets. In validation tests the classifier reached 91 percent accuracy in identifying papers that matched the retracted template style used as the training signal.</cite> That's a respectable validation result, though 91 percent accuracy on a curated test set isn't the same as 91 percent precision across 2.6 million papers in the wild. The false positive rate at that scale could still number in the tens of thousands.

The trend line is the more alarming figure. <cite index="11-10,11-11">Flagged papers rose sharply over the past two decades, from about 1 percent in the early 2000s to a peak of more than 16 percent in 2022, and suspicious papers appear across thousands of journals published by major companies, including journals with strong reputations and high impact.</cite>

<cite index="13-24,13-25">Paper mills target not the large, hard-to-verify clinical trials but molecular cancer biology and early-stage lab studies, where plausible results are easy to fabricate. The more a field is squeezed by publish-or-perish pressure, the more demand pools around buying authorship.</cite> That framing is worth keeping: the incentive structure is as much a part of this story as the algorithm.

The geographic distribution is stark. <cite index="17-12">Over 170,000 flagged papers were affiliated with Chinese institutions, representing 35 percent of China's cancer research output.</cite> Those numbers need context the paper itself acknowledges: a high flag rate reflects where paper mills are most active and where institutional pressure to publish is most intense, not necessarily where individual researchers are most dishonest.

The arms-race problem lurks in the methodology. <cite index="17-20,17-21,17-22">The Problematic Paper Screener, which catches 'tortured phrases' -- telltale signs of early machine paraphrasing -- has a built-in expiration date. Modern large language models don't produce tortured phrases. The fingerprint the screener relies on is disappearing as the generators improve.</cite> Barnett's BERT tool faces the same ceiling: it was trained on yesterday's fraud templates. Paper mills that adapt their output style could slip past it entirely.

<cite index="11-2">Three journals are already trialing the tool</cite> as a pre-publication screen, which is a meaningful step. But screening is not the same as retraction, and retraction is not the same as correcting the downstream science built on flagged work. The 261,000-paper figure is best understood as a lower bound on a problem the field has not yet developed the institutional machinery to address at scale.

Sources cited:
- The BMJ (Scancar, Byrne, Causeur & Barnett, 2026; 392: e087581) (https://doi.org/10.1136/bmj-2025-087581)
- ScienceDaily, AI flags more than 250,000 suspicious cancer research papers (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- TechApple Global, AI Flags More Than 250,000 Suspicious Cancer Research Papers with 91% Accuracy (https://global.techapple.com/2026/08/ai-flags-more-than-250000-suspicious-cancer-research-papers-with-91-accuracy/)
- GN Crypto News, AI model flags 261,245 cancer papers for template writing (https://www.gncrypto.news/news/ai-model-flags-261245-cancer-papers-template-writing/)
- DiaphorAI, 250,000 Cancer Studies Flagged as Fake (https://diaphorai.com/posts/ai-scientific-fraud-autoimmune-knowledge/)

Reporting by Dr. Maya Iyer, Staff Reporter, for the Science desk · ETL Newswire staff
Read more at the source

This release was originally distributed via ETL Newswire. Visit The BMJ (Scancar, Byrne, Causeur & Barnett, 2026; 392: e087581) for the full story, related releases, and contact information.

Visit The BMJ (Scancar, Byrne, Causeur & Barnett, 2026; 392: e087581) →