Published by Emerging Technologies Laboratory · via ETL Newswire
Science· 

AI Flags 260,000 Cancer Papers for Possible Paper Mill Origins

A BERT-based classifier scanned 2.64 million cancer studies published over 25 years and found roughly one in ten matching the textual fingerprint of fraudulent paper mills, according to a study published in The BMJ.

By Dr. Maya Iyer, Staff Reporter · Science Desk

A machine learning tool has flagged roughly 260,000 cancer research papers, about one in every ten examined, as bearing the textual hallmarks of so-called paper mills, according to a study published in The BMJ. The scale of the finding puts a number on what integrity researchers have long suspected: the problem is bigger than the retraction record suggests, and it has been growing for two decades.

The study was led by Professor Adrian Barnett of the Queensland University of Technology's School of Public Health and Australian Centre for Health Services and Innovation, working with an international group of collaborators. Their corpus was large: 2.6 million cancer research papers published between 1999 and 2024, drawn from PubMed. The team restricted the corpus to journal articles only, excluding reviews and clinical trial reports on the grounds that paper mills likely rely on manuscript-type-specific templates, so a single model can't fairly judge them all. That's a methodological call worth noting.

The classifier itself is a BERT language model, the same family of transformer-based models that powers much of modern text analysis. The team trained it to detect subtle textual fingerprints that repeatedly appear in known paper mill products. When evaluated against verified examples, the model correctly identified suspicious papers 91 percent of the time. Barnett compares the approach to email spam filtering: the model learns the structural and linguistic patterns of fraudulent boilerplate, then looks for those patterns at scale.

The trend line is what's hard to ignore. Flagged papers rose from roughly one percent of cancer publications in the early 2000s to more than 16 percent by 2022. That climb appeared across thousands of journals, including high-impact ones. Of the top 20 journals the team examined, 19 turned up flagged papers. Molecular cancer biology and early-stage laboratory research showed the highest concentrations, with gastric, liver, and bone cancer sub-fields clustering near the top.

But the 91 percent accuracy figure carries a significant caveat, and the authors don't hide it. The model was trained on yesterday's fraud: formulaic, template-driven manuscripts. As generative AI makes it easier to produce fake papers without those templates, the classifier's training signal may erode. The tool flags papers matching the writing style and structure seen in retracted, fraudulent work, a definition that will need updating as the fraud itself evolves.

The team and the journal are both careful on one point: a flag is not a verdict. Papers identified by the system should not automatically be treated as fraudulent. The results are warning signals, not confirmed findings of misconduct, and each flagged paper still needs review by human experts. Three journals are reportedly piloting the tool as part of their editorial triage process, which is a reasonable use case, as a first-pass screen, not a gatekeeping mechanism.

The stakes are not abstract. Cancer research feeds clinical trials, drug development, and treatment guidelines. If fabricated studies accumulate in the evidence base, they don't just waste resources; they can misdirect legitimate scientists and, downstream, affect patient care. The paper mill problem isn't new, but a classifier that can flag potential problems at this scale changes what's tractable for editors and research integrity officers, provided the tool keeps pace with how the fraud itself is changing.

Sources cited:
- The BMJ (via ScienceDaily) (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- Gulf News (https://gulfnews.com/lifestyle/health-fitness/ai-flags-250000-cancer-studies-as-scientists-warn-of-fake-research-surge-1.500610046)
- TechApple Global (https://global.techapple.com/2026/08/ai-flags-more-than-250000-suspicious-cancer-research-papers-with-91-accuracy/)
- Pebblous AI Blog (citing BMJ original) (https://blog.pebblous.ai/blog/cancer-paper-mill-ai-detector/en/)

Reporting by Dr. Maya Iyer, Staff Reporter, for the Science desk · ETL Newswire staff
Read more at the source

This release was originally distributed via ETL Newswire. Visit The BMJ (via ScienceDaily) for the full story, related releases, and contact information.

Visit The BMJ (via ScienceDaily) →