AI Screen Flags Nearly 10% of Cancer Papers as Potential Paper Mill Products
A machine learning tool trained on retracted paper mill manuscripts swept 2.6 million cancer studies and found more than 250,000 bearing the same textual fingerprints, according to a study published in The BMJ.
A machine learning screen of 25 years of cancer research has produced what may be the most unsettling audit the field has seen: roughly one in ten published papers shows writing patterns consistent with fraudulent paper mill output.
The study, published in The BMJ and led by statistician Professor Adrian Barnett at Queensland University of Technology (QUT), is not a verdict. It's a flag. But the scale is hard to sit with.
Researchers trained a BERT-class language model on the textual fingerprints of papers already known to have been retracted for paper mill activity. According to the QUT press release reviewed for this piece, the tool then worked through 2.6 million cancer research papers published between 1999 and 2024. It identified more than 250,000 with writing patterns similar to those retracted articles.
The number that keeps coming up in the underlying reporting is 9.87 percent. That figure, cited in The Scientist's coverage of the BMJ paper, represents the share of the entire cancer literature screened that the tool flagged. To put that in context, prior estimates for problematic articles across all of biomedical research ran close to 6 percent by 2023, according to The Scientist. Cancer, it appears, is a preferred target.
The trend line is the part that warrants the most scrutiny. As reported by Gulf News, the proportion of flagged studies climbed from roughly 1 percent in the early 2000s to more than 16 percent by 2022. That trajectory did not reverse.
The tool's performance is genuinely good for this kind of screen. Internal testing found it correctly identified known paper mill papers about 91 percent of the time, according to Gulf News, with specificity above 96 percent per additional analysis. But specificity is not the same as certainty for any individual paper, and the authors are careful to say so. As Gulf News noted, the findings published in The BMJ "don't suggest those studies are fraudulent" on their own. What they share are linguistic fingerprints, which still require expert review to interpret.
The researchers' own framing compares the tool to a spam filter: it does not decide whether a paper is fraudulent, it identifies the ones that warrant a second look. Three scientific journals are already testing it as part of their editorial workflow, according to Gulf News.
The false-negative problem is real, too. At 91 percent accuracy applied to 2.6 million papers, the math suggests a substantial number of paper mill products could clear the screen entirely. That does not undermine the tool's value as triage, but it does mean the literature's actual contamination rate could be higher than the flagged share implies.
Paper mills, as QUT's release defines them, are companies that sell fake or low-quality scientific studies. A March 2026 study in PNAS, cited in separate commentary on the BMJ findings, characterized these operations as criminal organizations whose output is doubling roughly every 18 months, far outpacing the 15-year doubling time for legitimate scientific publication.
The downstream consequences are not abstract. Studies that enter systematic reviews or inform clinical guidelines can eventually shape treatment decisions. A PNAS-cited estimate, noted in analysis published by diaphorai.com, suggests only 15 to 25 percent of paper mill products are ever retracted, meaning the rest accrue citations and stay in circulation.
This is one study, and it comes with the usual caveats: the model was trained and validated on a particular corpus of known retractions, and its performance outside that distribution is not fully characterized. Molecular cancer biology and laboratory-based research showed the highest flagging rates in the data, according to Gulf News, which may partly reflect the genres paper mills favor rather than the fields themselves.
Barnett's team has made their screening tool available for journal use. Whether editors adopt it widely enough to shift the contamination curve is the next question the data can't answer yet.
Sources cited:
- The BMJ (study) (https://www.bmj.com/content/392/bmj-2024-087581)
- QUT press release (https://www.qut.edu.au/news?id=203173)
- The Scientist (https://www.the-scientist.com/nearly-ten-percent-of-cancer-papers-flagged-as-potentially-fake-74185)
- Gulf News (https://gulfnews.com/lifestyle/health-fitness/ai-flags-250000-cancer-studies-as-scientists-warn-of-fake-research-surge-1.500610046)
- diaphorai.com (https://diaphorai.com/posts/ai-scientific-fraud-autoimmune-knowledge/)
This release was originally distributed via ETL Newswire. Visit The BMJ (study) for the full story, related releases, and contact information.
Visit The BMJ (study) →