Published by Emerging Technologies Laboratory · via ETL Newswire
Science· 

AI Screen Flags Nearly 10 Percent of Cancer Research Papers as Potential Paper Mill Products

A BMJ study analyzed 2.6 million cancer papers published between 1999 and 2024 and found more than 250,000 carrying the textual fingerprints of fraudulent paper mills, with the rate climbing sharply over two decades.

By Dr. Maya Iyer, Staff Reporter · Science Desk

A machine learning tool built by researchers at Queensland University of Technology has identified more than a quarter of a million cancer research papers that may have originated from so-called paper mills, according to a study published in The BMJ this month. The scale of the finding, nearly one in ten papers screened, is larger than most research integrity experts had estimated.

Paper mills are commercial operations that manufacture or broker scientific manuscripts, sometimes using fabricated or manipulated data, and sell authorship slots to researchers who need publication credits. The practice has been documented for years, but the new work is among the first attempts to measure its footprint systematically across an entire field.

The QUT team, led by biostatistician Professor Adrian Barnett, trained their classifier on a BERT-based language model, teaching it to recognize recurring writing patterns in papers already retracted for suspected paper mill activity. In performance testing, the tool correctly flagged problematic papers with roughly 90 percent accuracy. The team then ran it against the full corpus.

The results were stark. Among 2.6 million cancer research papers drawn from PubMed, the study published in The BMJ found that 261,245 papers showed textual similarities to retracted paper mill papers. That is 9.87 percent of the literature analyzed. The trend line is the more alarming figure: the proportion of flagged papers climbed from around 1 percent in the early 2000s to more than 16 percent by 2022, according to analysis reviewed by Gulf News. Certain cancer subtypes look worse. Gastric, liver, bone, and lung cancer research showed especially high concentrations of suspicious papers.

Three caveats matter here, and the authors state them plainly. First, a flag is not a verdict. As the Gulf News report of the study noted, every paper the AI identifies still requires human expert review before any conclusion can be reached. Second, the 90 percent accuracy figure is measured against the style of fraud that existed before generative AI made it trivial to vary prose templates. The classifier learned fingerprints from older, formulaic manuscripts; newer AI-generated fakes may not match those patterns. Third, the tool read only titles and abstracts, not full text, and it excluded literature reviews and clinical trials, which could be paper mill targets requiring separate models.

Still, the clinical stakes make the finding hard to set aside. As Professor Barnett said in a statement reviewed by ecancer, cancer research feeds directly into drug development and patient care, and fabricated studies that enter the evidence base can mislead scientists and slow progress. The team plans to expand the tool to other research areas and to refine the model as more confirmed paper mill cases become available.

The finding drops into a peer-review ecosystem already under strain. A separate report in Nature, cited in the Gulf News coverage of this study, found that cancer papers suspected of paper mill origin were attracting more citations than legitimate studies, meaning the contamination doesn't stay contained. It propagates.

The honest methodological read here is that 261,245 is a ceiling estimate with noise, not a confirmed count of fraudulent papers. But the direction of the trend is not ambiguous: whatever is happening in cancer literature has been getting worse for twenty years, and the field's existing gatekeeping mechanisms missed it. The QUT system is a screening tool, not a solution. What comes after the flag is still a human problem.

Sources cited:
- The BMJ (via ScienceDaily) (https://www.sciencedaily.com/releases/2026/07/260714225538.htm)
- The Scientist (https://www.the-scientist.com/nearly-ten-percent-of-cancer-papers-flagged-as-potentially-fake-74185)
- Gulf News (https://gulfnews.com/lifestyle/health-fitness/ai-flags-250000-cancer-studies-as-scientists-warn-of-fake-research-surge-1.500610046)
- ecancer (https://ecancer.org/en/news/27724-new-tool-exposes-scale-of-fake-research-flooding-cancer-science)
- Pebblous AI Blog (https://blog.pebblous.ai/blog/cancer-paper-mill-ai-detector/en/)

Reporting by Dr. Maya Iyer, Staff Reporter, for the Science desk · ETL Newswire staff
Read more at the source

This release was originally distributed via ETL Newswire. Visit The BMJ (via ScienceDaily) for the full story, related releases, and contact information.

Visit The BMJ (via ScienceDaily) →