AI Flags Nearly 1 in 10 Cancer Papers as Potential Paper Mill Products
A BERT-based classifier trained on Retraction Watch data scanned 2.6 million PubMed studies and identified more than 261,000 with textual patterns matching known fraudulent manuscripts.
A machine learning audit of 25 years of cancer literature has put a hard number on a problem the research integrity community has long suspected was larger than the retraction record suggests.
<cite index="9-5">The study, published in The BMJ, examined 2.6 million cancer research papers released between 1999 and 2024.</cite> <cite index="14-8,14-9">The model achieved an accuracy of 0.91, and when applied to the full corpus it flagged 261,245 of 2,647,471 papers, or 9.87% (95% CI 9.83 to 9.90), revealing a large increase in flagged papers across both the full literature and the top 10% of journals by impact factor.</cite> That last point matters: this isn't a low-tier journal problem.
The tool itself is straightforward. <cite index="10-15">The classifier, built by Adrian Barnett's team at Queensland University of Technology, reads only a paper's title and abstract to judge whether it resembles a fabricated manuscript.</cite> <cite index="14-4">The model was trained on 2,202 retracted paper mill papers and validated on independent data collected by image integrity experts.</cite> The researchers describe it the way you'd describe a spam filter: <cite index="12-3,12-4">it doesn't decide whether a paper is malicious; it spots the ones that deserve another look.</cite>
Flags aren't findings. <cite index="12-12,12-13">The findings published in The BMJ don't suggest those studies are fraudulent -- what they do suggest is that many share the same linguistic fingerprints as papers linked to so-called "paper mills," businesses that produce or sell scientific manuscripts, sometimes using fabricated or manipulated data.</cite> <cite index="12-17">Every paper flagged still needs to be reviewed by experts before any conclusions can be reached.</cite>
The trend data is the part that should keep editors up at night. <cite index="12-8">The proportion of potentially problematic cancer studies has steadily climbed over the past two decades, from around 1% in the early 2000s to more than 16% by 2022.</cite> <cite index="14-12">Flagged papers were overrepresented in fundamental research and in gastric, bone, and liver cancer.</cite> <cite index="10-7">Paper mills target not large, hard-to-verify clinical trials but molecular cancer biology and early-stage lab studies, where plausible results are easy to fabricate.</cite>
The geographic distribution is uneven and worth naming carefully. <cite index="14-10">More than 170,000 papers affiliated with Chinese institutions were flagged, accounting for 36% of Chinese cancer research articles.</cite> <cite index="10-8">The more a field is squeezed by publish-or-perish pressure, the more demand pools around buying authorship</cite> -- and that pressure is not unique to any one country, even if the flagging rates aren't uniform.
There are real limits to what the model can tell us. The classifier's training set, drawn from the Retraction Watch database, reflects yesterday's known frauds. A mill that has adapted its templates since those retractions occurred could plausibly slip through. <cite index="10-2,10-3">The tool is a BERT-based classifier, and its 91% accuracy is accuracy against yesterday's fakes -- passing is not the same as being clean.</cite>
The practical upside is that three journals are already running the system operationally. <cite index="12-5">Three scientific journals are already testing the technology as part of their editorial process.</cite> <cite index="14-13,14-14">The authors conclude that paper mills are a large and growing problem in the cancer literature and are not restricted to low-impact journals, and that collective awareness and action will be crucial to address the problem.</cite>
For anyone who has ever built a research project on a cell-line study from a subfield with a 16% suspected-fraud rate, the methodological question isn't abstract. The replication crisis conversation has largely centered on underpowered psychology studies; this data suggests oncology's foundational literature has its own structural problem, and it's getting worse over time.
Sources cited:
- The BMJ (via PMC full text) (https://pmc.ncbi.nlm.nih.gov/articles/PMC12853418/)
- ScienceDaily (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- Gulf News (https://gulfnews.com/lifestyle/health-fitness/ai-flags-250000-cancer-studies-as-scientists-warn-of-fake-research-surge-1.500610046)
This release was originally distributed via ETL Newswire. Visit The BMJ (via PMC full text) for the full story, related releases, and contact information.
Visit The BMJ (via PMC full text) →