Published by Emerging Technologies Laboratory · via ETL Newswire
Science· 

AI Tool Flags Nearly 10 Percent of Cancer Literature as Potential Paper Mill Output

A BERT-based classifier screened 2.6 million cancer studies published since 1999 and matched roughly 261,000 of them to the textual fingerprints of known fraudulent paper mills, according to a study in The BMJ.

By Dr. Maya Iyer, Staff Reporter · Science Desk

A machine learning system developed at Queensland University of Technology has flagged more than a quarter-million cancer research papers as bearing the hallmarks of organized scientific fraud, in what researchers describe as one of the largest research integrity audits ever conducted in a single medical field.

The study, published in The BMJ and led by biostatistician Adrian Barnett, trained a BERT-based language model on 2,202 retracted paper-mill articles drawn from a retraction database, then validated it against independent expert datasets before turning it loose on the full corpus. When the model screened 2.6 million cancer papers published between 1999 and 2024, it flagged 261,245 of them, or 9.87 percent of the total, for writing patterns associated with suspected fabrication, according to reporting reviewed by The Scientist.

The trend line inside those numbers is the part worth sitting with. As reported in The BMJ, flagged papers rose from roughly 1 percent of annual output in the early 2000s to a peak of more than 16 percent in 2022. That is not a slow drift. That is a structural shift in how a portion of the research pipeline is being gamed.

Barnett's model works on a principle that's easy to underestimate: paper mills, at scale, rely on templates. The classifier reads only a paper's title and abstract and looks for the subtle, repeated prose patterns those templates leave behind. When evaluated against verified examples, the model correctly identified suspicious papers 91 percent of the time, according to TechApple Global's coverage of the study. Barnett has compared it to email spam filtering, and the analogy is apt, including the arms-race implication.

That implication is the sharpest limitation in this work. The 91 percent figure is accuracy against the fraud that existed when the training set was assembled. Earlier paper mill output leaned on formulaic paraphrasing tools that produced what researchers call "tortured phrases," mangled synonyms that are easy for a classifier to catch. Modern generative AI does not produce those artifacts. As one analysis noted, the detection features used here have shelf lives. The model is, by design, retroactive.

Barnett acknowledged the ceiling directly. As he told The Scientist, "It could actually be more because we're just detecting one particular kind of template. If the mills have other templates that are more sophisticated, we would have missed them."

Geographic concentration in the flagged papers is pronounced. More than 170,000 of the suspicious papers were affiliated with Chinese institutions, representing roughly 35 to 36 percent of China's total cancer research output in the dataset, according to coverage reviewed across multiple outlets. The study's authors have been careful to note that flagging is not the same as confirmed fraud, and the findings, published in The BMJ, do not assert that flagged papers are fraudulent, only that they share linguistic characteristics with papers already retracted for paper mill activity.

Three journals are reportedly piloting the tool as a pre-publication screen, which is the more constructive angle here. The audit of 25 years of literature is striking as a number, but it's a one-time pass. The more durable use case is prevention: running the classifier at submission, before a suspect paper enters the citation record and gets built upon.

The broader context makes that urgency concrete. A March 2026 study in PNAS described paper mill operations as criminal organizations, and the Retraction Watch database now lists more than 63,000 total retractions. No single tool closes that gap. But a classifier that can triage millions of submissions is at minimum a first filter, as long as the people building it keep retraining it on whatever fraud looks like next year.

Sources cited:
- The BMJ (via ScienceDaily) (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- The Scientist (https://www.the-scientist.com/nearly-ten-percent-of-cancer-papers-flagged-as-potentially-fake-74185)
- TechApple Global (https://global.techapple.com/2026/08/ai-flags-more-than-250000-suspicious-cancer-research-papers-with-91-accuracy/)
- KFF Health Monitor (https://www.kff.org/health-information-trust/how-ai-can-both-detect-and-enable-fraudulent-research/)
- Pebblous AI Blog (https://blog.pebblous.ai/blog/cancer-paper-mill-ai-detector/en/)

Reporting by Dr. Maya Iyer, Staff Reporter, for the Science desk · ETL Newswire staff
Read more at the source

This release was originally distributed via ETL Newswire. Visit The BMJ (via ScienceDaily) for the full story, related releases, and contact information.

Visit The BMJ (via ScienceDaily) →