Machine Learning Flags Nearly 10% of Cancer Research Papers as Potential Paper Mill Output
A BERT-based classifier screened 2.6 million oncology studies spanning 25 years and identified more than 261,000 with textual signatures matching known fraudulent manuscripts.
A machine learning tool trained to recognize the writing patterns of fraudulent scientific manuscripts has now swept through a quarter century of cancer research and returned a number that should trouble anyone who relies on that literature.
The study, published in The BMJ as Scancar et al. (2026;392:e087581), was led by biostatistician Adrian Barnett at Queensland University of Technology and an international group of collaborators. According to the paper reviewed in full-text on PubMed Central, the classifier flagged 261,245 of 2,647,471 papers, or 9.87% (95% CI: 9.83 to 9.90), as bearing the fingerprints of so-called paper mills.
Paper mills are commercial operations that produce or broker scientific manuscripts, often with fabricated or manipulated data, for researchers who need publication records. The mechanism behind the detector is a fine-tuned BERT architecture, a language model trained on 2,202 papers already retracted for suspected paper mill origin. It reads only a paper's title and abstract, not the full text or data, and classifies based on the prose style and structure those retracted papers share. Reported accuracy against the validation set was 0.91, according to reporting reviewed in The BMJ.
The scale of what the model found is striking, but the distribution is just as important as the headline number. Flagged papers were not confined to predatory or low-impact outlets. According to analysis reported by Cancer Therapy Advisor citing the original BMJ paper, the model flagged 9.87% of papers as likely originating in a paper mill, including papers in the top 10% of journals by impact factor. Of the top 20 journals examined, 19 had flagged papers in the corpus, per reporting reviewed at The Scientist.
Certain cancer types showed especially elevated rates. According to the paper's results as reported by ScienceDaily, gastric, liver, bone, and lung cancer studies had particularly high rates of flagged papers. The geographic clustering is pronounced: more than 170,000 flagged papers carried affiliations with Chinese institutions, representing roughly 36% of China's cancer research output in the dataset, according to the BMJ paper's findings as reported by ScienceDaily and The Scientist.
The authors were careful to draw a line between a flag and a finding. As reported in ScienceDaily's coverage of the QUT release, the team emphasized that papers identified by the system should not automatically be treated as fraudulent; each case still requires human review. That caveat matters a lot at this scale. With 261,000 flagged studies, no journal office or integrity body has the capacity to manually adjudicate every case.
There's also a structural limitation in the model's design that deserves attention. The classifier learned the fingerprint of yesterday's fraud: formulaic, templated prose produced by paper mills operating before generative AI made individualized fake manuscripts trivially cheap to produce. As noted in analysis reported by the Pebblous blog reviewing the BMJ paper, the tool's 91% accuracy is accuracy against that older pattern. False negatives were disproportionately associated with gastric, liver, colorectal, and lung cancer papers, according to the methods and results reported in the full PMC text of the Scancar et al. paper.
What does a 9% false negative rate mean in practice? Applied to a corpus of 2.6 million papers, it means the tool likely misses a substantial number of fraudulent manuscripts that have learned to evade the template. The authors acknowledged this and framed the tool as a first-pass screen, not a verdict system.
The practical uptake is already moving. According to ScienceDaily's coverage of the QUT announcement, three scientific journals are testing the classifier as part of their editorial screening process. That's a meaningful early signal, though it also raises a question the paper doesn't fully resolve: what happens to the 25 years of already-published, already-cited literature the tool just flagged?
Professor Barnett put the stakes plainly in a statement covered by EurekAlert: "Paper mills are producing 'research' on an industrial scale, and our findings suggest the problem in cancer research is far larger than most people realised." Clinical guidelines, drug development pipelines, and ongoing trials all draw on this literature. A contaminated evidence base doesn't just slow science; it can redirect it. That's the part worth sitting with.
Sources cited:
- The BMJ (Scancar et al., 2026;392:e087581, via PubMed Central) (https://pmc.ncbi.nlm.nih.gov/articles/PMC12853418/)
- ScienceDaily (QUT release) (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- The Scientist (https://www.the-scientist.com/nearly-ten-percent-of-cancer-papers-flagged-as-potentially-fake-74185)
- Cancer Therapy Advisor (https://www.cancertherapyadvisor.com/news/cancer-research-literature-paper-mills/)
- EurekAlert (QUT) (https://www.eurekalert.org/news-releases/1114511)
This release was originally distributed via ETL Newswire. Visit The BMJ (Scancar et al., 2026;392:e087581, via PubMed Central) for the full story, related releases, and contact information.
Visit The BMJ (Scancar et al., 2026;392:e087581, via PubMed Central) →