Your Q1 Journals Are Not Safe

· ORBIXER AI LABS

Your Q1 Journals Are Not Safe

RESEARCH INTEGRITY · ORBIXER AI LABS

Your Q1 Journals Are Not Safe

What the BMJ Paper Mill Study Means for US and UK Research Offices

11 min read · For Research Deans, VPs for Research, Research Integrity Officers, editors and librarians

For twenty years, the working assumption in research administration has been that publication fraud is a problem of the periphery. It happens in obscure venues, at institutions under extreme publication pressure, in journals nobody serious reads. Screen for predatory publishers, keep your researchers away from the fringe, and the core literature stays clean.

In January 2026, that assumption stopped being defensible.

A study published in The BMJ by Adrian Barnett's team at Queensland University of Technology trained a machine-learning classifier on 2,202 retracted paper-mill papers and ran it across 2,647,471 original cancer research articles published between 1999 and 2024. The model flagged 261,245 papers — 9.87% of the corpus (95% CI 9.83–9.90).

That number is arresting on its own. It is roughly three times the previously cited estimate of around 3% prevalence across published biomedical research.

But the finding that should change how research offices operate is the one underneath it:

Flagged papers rose over time across the entire corpus — and within the top 10% of journals by impact factor.

Not the fringe. The core. Paper mills are not restricted to low-impact venues, and they have not been for years.

If your institutional risk model assumes that publishing in — or citing from — a well-indexed, high-quartile journal constitutes due diligence, that model is broken. This piece is about what replaces it.

Part 1: What the evidence actually says

Precision matters here, because paper-mill statistics get mangled in retelling and the mangled versions are easy to discredit. Below are the load-bearing numbers, with their real denominators.

The BMJ cancer screening study (2026)

Measure

Value

Papers screened (cancer research, 1999–2024)

2,647,471

Papers flagged as paper-mill-like

261,245 (9.87%, 95% CI 9.83–9.90)

Model accuracy

0.91 internal · 0.93 external validation

Sensitivity / specificity

87% / 96% internal · 87% / 99% external

Training set

2,202 retracted paper-mill papers

Papers with Chinese institutional affiliation flagged

>170,000 — representing 36% of Chinese cancer research articles

Prevalence in top 10% of journals by impact factor

Rising across the study period

Two clarifications that are routinely got wrong, including in widely circulated summaries:

The corpus was 2.6 million papers, not 24.8 million. The larger figure refers to the unfiltered PubMed extract before cancer-specific selection. Citing 24.8 million as the screened corpus overstates the study by roughly a factor of ten.

The 36% figure is a within-country rate, not a share of the flagged pool. It means 36% of Chinese-affiliated cancer articles were flagged — not that 36% of all flagged papers were Chinese. Any calculation that multiplies a total flagged count by 36% is combining incompatible denominators and produces a meaningless number.

We flag both errors deliberately. If your institution is going to cite this study in a policy document or an institutional statement, it needs to cite it correctly. Integrity arguments made with sloppy statistics do not survive contact with a scrutiny panel.

The Wiley–Hindawi collapse

The largest single integrity event in the history of scholarly publishing, and the clearest demonstration that scale and reputation are not protection.

Event

Detail

Wiley acquires Hindawi

January 2021 · total purchase price US98 million · 200+ open-access journals

Special issues paused

October 2022 – January 2023 · US$9 million revenue lost in one quarter

Clarivate delists 19 Hindawi journals from Web of Science

March 2023 · for failing editorial quality criteria

Quarterly revenue impact disclosed

US

8 million decline (Q2 FY2024 vs prior year)

Full-year revenue impact guided

US5–40 million (FY2024)

Total retractions, 2023–2025

11,300+ papers — the largest coordinated retraction programme on record

Brand retired

Hindawi name discontinued · four journals closed as heavily compromised

Figures from Wiley's acquisition announcement, Wiley earnings disclosures as reported by Retraction Watch, and Clarivate's delisting announcement. Some secondary sources report a US

04 million write-down and a further impairment charge; that specific figure is not confirmed in Wiley's primary disclosures reviewed here — status to be verified against Wiley's SEC filings before citation in formal documents.

Note what this case proves. Hindawi was not a predatory publisher in the Beall sense. It was an established open-access house, acquired for nearly

00 million by one of the world's largest academic publishers, with journals indexed in Web of Science. The paper mills got in through guest-edited special issues — a legitimate editorial mechanism with a weak verification perimeter.

The vulnerability was structural, not reputational. Which means reputation is not a screen.

Global retraction volume

Nature's analysis (Van Noorden, December 2023) recorded that retractions passed 10,000 in 2023 — a new annual record — with Hindawi journals alone pulling more than 8,000 articles. That concentration is itself the story: a single publisher's cleanup accounted for the overwhelming majority of a record year, meaning the underlying background rate across the rest of the literature remained largely unaddressed.

A word on the other figure in circulation. The COPE/STM analysis estimates that around 2% of submitted manuscripts may originate from paper mills. The BMJ study's 9.87% describes published papers in one field. These are different denominators measuring different things — submissions versus published output, all disciplines versus cancer research — and they should never be stacked, averaged, or presented as a trend. Independent work estimating hundreds of thousands of suspicious papers across two decades uses a third methodology again.

Part 2: Why the detection arms race matters to you

Paper mills are not static adversaries. They adapt to whatever screen the sector deploys — which is why any institution relying on a single tool is defending last year's attack.

Generation one relied on plagiarism. Text lifted from published work, detectable by similarity software. Similarity checkers were built, and worked.

Generation two used automated rewriting to defeat similarity scoring. This produced the “tortured phrases” phenomenon — technical terms mechanically substituted with synonyms until they became scientifically meaningless.

Accepted scientific term

Replacement found in suspicious papers

Artificial intelligence

Counterfeit consciousness

Random forest

Haphazard backwoods

Artificial neural network

Fake neural organization

Big data

Enormous information

Guillaume Cabanac and Cyril Labbé built the Problematic Paper Screener to detect exactly this signature, and it has flagged thousands of published papers, contributing to substantial numbers of subsequent retractions.

Current flagged totals should be checked against the live Problematic Paper Screener dashboard rather than any static figure — the count moves continuously.

Generation three manipulates the process, not just the text. Fabricated image data. Synthetic datasets that pass basic statistical sanity checks. Suggested reviewers whose contact addresses route back to the mill. Authorship slots sold on manuscripts already in review. One documented Russia-based operation was estimated to have sold co-authorship positions worth around US$6.5 million between 2019 and 2021.

Generation four is generative. Large language models have collapsed the marginal cost of producing plausible scientific prose to near zero. Tortured phrases were a detection gift — an artifact of crude tooling. That gift is gone. Text-level detection is now the weakest layer in the stack, not the strongest.

The strategic implication for a research office is straightforward and unwelcome: the signals are moving from the manuscript to the metadata. Which venue. Which editorial process. Which co-author network. Which citation trail. These are institutional-level questions, and no manuscript-level tool answers them.

Part 3: The institutional exposure nobody is measuring

Here is where most Western research administration has a blind spot.

The dominant framing of research integrity in the US and UK is a misconduct framing: an allegation arrives, a process runs, an outcome is recorded. Robust, well-governed, and almost entirely reactive.

The exposure created by industrial paper mills is different in kind. It is ambient contamination, and it lands on institutions that have done nothing wrong.

Consider four scenarios, none of which involves misconduct by your staff:

Citation inheritance. A researcher writes a systematic review. Six of the primary studies are later retracted for paper-mill origin. The review is now unreliable, the clinical guideline built on it is unreliable, and neither the author nor the institution is notified when the retractions land.

Collaborator exposure. A department signs a partnership. Two members of the partner team have publication records concentrated in journals under active integrity investigation. Nobody checked, because checking a dozen co-authors' venue histories by hand is a full day's work per partnership.

Portfolio drift. An output was published in 2021 in a well-regarded indexed journal. In 2024 that journal was delisted. The output sits in your repository, in your CRIS, and potentially in your assessment submission, carrying a venue status that changed after acceptance and that nothing in your systems monitors.

Retraction lag. Paper-mill retractions have historically taken around two years from publication to appear. Your institution is, structurally, always operating on a two-year-old picture of its own literature.

Now place that against the regulatory environment as it stands in 2026.

United States

The Office of Research Integrity's revised Final Rule (42 CFR Part 93) became applicable to institutions on 1 January 2026. Allegations received on or after that date run under the new framework, with tighter procedural and record-keeping obligations. Institutions were required to submit updated assurance of policies and procedures with their 2025 annual report, due 30 April 2026. Every US research office is now in its first full operating year under revised rules.

United Kingdom

REF 2029 released draft assessment criteria for Strategy, People and Research Environment on 31 July 2026, opening a sector-wide consultation ahead of formal guidance in autumn 2026. Reporting on the draft rules indicates the institutional statement is expected to ask universities to describe how they manage and monitor risks associated with research security and integrity. SPRE carries roughly 20% of an institution's overall score. Institutions eligible for Research England funding are already required to implement the Concordat to Support Research Integrity as a funding condition; UKRIO's Code of Practice was updated in 2025 to align with the revised Concordat and to add guidance on AI and emerging technologies.

Both regulators are converging on the same demand, from different directions: describe your monitoring, and show it produces evidence.

Most institutions can describe the policy. Very few can produce the evidence — because evidence at this scale means measurement across thousands of outputs, and almost nobody is measuring.

Part 4: Three questions to run on your own institution

These are deliberately answerable in principle. In practice, they are almost never answered.

1. How many outputs in your repository sit in venues with a documented retraction cluster or a delisting event since publication?

Not “do we have any.” The number, with the list attached.

2. When a paper your researchers cite is retracted, how long until anyone at the institution knows?

For most institutions the answer is: only if someone happens to notice.

3. Before an international collaboration is formalised, does anyone assess the publication venues of the partner team?

Research security guidance increasingly assumes this happens. Operationally, it almost never does.

If the answers are unknown, we don't, and no — that is the sector median, not an outlier. It is also precisely why the reporting requirements are tightening.

Part 5: What an integrity layer has to do

Closing this gap requires a category of tool that does not fit neatly into existing procurement lines. Not a similarity vendor. Not a bibliometrics subscription. Something narrower.

Verification with a visible chain of reasoning. A research office cannot defend a decision to a panel on the strength of a colour-coded badge. It can defend a documented pathway: which indexing bodies actually list this venue, which metrics are legitimate versus manufactured, what integrity events are on record, and how confident the assessment is. Manufactured metrics — the alphabet soup of unofficial “impact factors” sold to journals — are a first-order signal, and distinguishing them from Web of Science JIF, Scopus CiteScore, SJR and SNIP is table stakes.

Retrospective portfolio screening. Point the system at the repository, not at one manuscript. The deliverable is a ranked risk register across the whole corpus — the artefact an institutional statement or an annual report can actually be built on.

Continuous monitoring. A venue's integrity status is not a fixed property. Titles get delisted. Clusters get uncovered. Retractions land years after publication. A check performed at submission has a short shelf life. What institutions need is a watch list that updates itself and reports what changed.

Exportable evidence. Not a dashboard. A dated, sourced, reproducible document that a Research Dean can hand to a panel, a funder, or a regulator.

That last point is the one most tools miss. Compliance is not a feeling of confidence. It is a file you can hand over.

Where ORBIXER fits

ORBIXER AI LABS builds this layer.

We started with venue verification because it is the atomic unit of the problem. Every downstream question — about a citation, a collaborator, a submission, a repository — resolves back to is this venue what it claims to be. The platform separates legitimate indexing and metrics from manufactured ones, records the reasoning behind each verdict rather than emitting a bare score, and monitors its underlying data continuously rather than assuming last year's answer still holds.

The Verify engine is live and in daily use. Portfolio-scale screening and institutional reporting for US and UK research offices are in active development, with early-access partnerships opening now.

We are not claiming to have solved research integrity. We are building one specific, unglamorous, badly-neglected piece of infrastructure: the ability to prove, in writing, what your institution publishes in and cites from.

If your office is drafting an institutional statement this autumn, or operating under the Final Rule for its first full year, that gap becomes visible on a deadline.

Better to see it now.

What researchers, editors and institutions can do today

Researchers — verify the venue before submission, not after acceptance. Treat guaranteed publication, guaranteed timelines, and unofficial impact metrics as disqualifying signals. Check the retraction status of key references before a review goes to press.

Reviewers — treat suspiciously clean datasets, statistical results that do not reconcile with the reported methods, reused or duplicated images, and author-suggested reviewers with non-institutional email addresses as escalation triggers.

Editors — verification of guest editors and suggested reviewers is now the highest-leverage control point in the workflow. The Hindawi case demonstrated that special issues are the softest perimeter in scholarly publishing.

Institutions — move evaluation away from raw publication counts. Every incentive that rewards volume over verification is a demand signal that paper mills exist to supply. And begin measuring your existing corpus, because you will be asked to describe it before you are ready.

WORK WITH US

ORBIXER is opening early-access partnerships with US and UK universities, research offices, and libraries for institutional-scale verification and integrity reporting.

Get in touch: info@orbixer.in

Explore the platform: orbixer.in

ORBIXER AI LABS builds research-integrity infrastructure for universities and research institutions. Founded by an IIT Kharagpur alumnus, ORBIXER works with institutions across India and is expanding to US and UK research markets in 2026.

Sources

Barnett et al., machine learning based screening of potential paper mill publications in cancer research, The BMJ, 2026

Wiley, acquisition of Hindawi Limited — company announcement, January 2021

Wiley earnings disclosures as reported by Retraction Watch, 2023–2024

Clarivate, Web of Science delisting announcement, March 2023

Cabanac & Labbé, Problematic Paper Screener

COPE / STM, paper mills report

US Office of Research Integrity, 2024 Final Rule (42 CFR Part 93) and associated guidance documents

REF 2029 SPRE draft criteria coverage, Times Higher Education, July 2026

UK Research Integrity Office, Code of Practice for Research (2025)

Concordat to Support Research Integrity (2025)

Van Noorden, R., more than 10,000 research papers were retracted in 2023 — a new record, Nature, December 2023

Retraction Watch Database

Regulatory guidance is subject to change. Institutions should confirm current requirements directly with ORI, Research England, and the REF 2029 team. Figures marked “status to be verified” should be confirmed against primary sources before use in formal documentation.

Continue reading

Paper mills depend on journals that skip verification — do not let your own work end up beside theirs. Before you submit, learn how to check if a journal is Scopus indexed, review the ten warning signs of predatory journals, and confirm whether a journal is fake or real. The ORBIXER Journal Finder runs these checks on every journal it lists.