Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Editorial · CASRAI · integrity-compliance

BuyTheBy Dataset Puts a Price on Paper Mill Authorship

A new dataset, BuyTheBy, quantifies for the first time what paper mills actually charge for authorship: 18,710 ads, prices from roughly $56 to $5,631 for first-author slots, and just 5 of 53 matched published papers retracted so far.

Published 24 Jul 2026· 4 minute read

A new dataset gives the first systematic, quantified look at how paper mills actually price the authorship slots they sell. Called BuyTheBy, it compiles 18,710 text-based advertisements — 15,839 of them with listed prices — scraped from seven paper mills operating via Telegram channels and websites linked to India, Iraq, Uzbekistan, Latvia, Ukraine, Russia, and Kazakhstan, spanning March 2020 to early April 2026. Nature‘s news desk covered the dataset in April 2026 (‘How much for a fake authorship? Ad database reveals secrets of scientific fraud’), with corroborating coverage from Chemical & Engineering News, Times Higher Education, and Retraction Watch.

What the dataset covers

BuyTheBy was compiled by Reese Richardson and Spencer Hong of Northwestern University, working with Anna Abalkina of Freie Universität Berlin — a research-integrity investigator previously known for exposing citation-cartel and hijacked-journal schemes. The dataset catalogs 20,598 individual authorship or product positions across 5,567 unique ‘products’ in 14 product categories, yielding 51,812 timestamped price data points in total. It is posted as an arXiv preprint (arXiv:2604.24576) and archived on Zenodo.

This is a meaningfully different research object than most paper-mill coverage to date. Detection studies and case investigations (see our coverage of tortured-phrases red flags and the COPE paper mills working group relaunch) establish that paper mills exist and describe how their output is detected. BuyTheBy instead treats the paper-mill market as an actual market, with prices, product tiers, and geography, and asks what that market structure reveals.

What authorship actually costs

Across the dataset, first-author slot prices ranged from roughly $56 to $5,631, with first authorship averaging around $1,000 and a reported median closer to $788. Pricing varied sharply by which mill was selling and to whom: the cheapest operation identified, based in India, priced everything under $150, while the most expensive prices came from a Russia-based mill serving local and Kazakhstan-based clients. One Iraq-based mill advertised slots on a specific paper at $350–$600. Lower-tier authorship positions (later co-author slots rather than first authorship) were consistently cheaper than first-author slots, though granular figures for those lower tiers are less consistently reported across secondary coverage than the first-author range above.

This is distinct from the mechanism our February 2026 coverage of fake-author citation cartels described, where Chemical & Engineering News reported entirely fabricated author identities — not real researchers buying a slot — being inserted into roughly 20 chemistry papers to sell citations at $5–$10 each, partly to exploit article-processing-charge waivers tied to specific countries. BuyTheBy’s ads, by contrast, document real transactions offered to real buyers seeking a byline on an already-planned or already-written manuscript. Both are paper-mill business lines; they are not the same product.

Weak follow-through on enforcement

Perhaps the more consequential finding sits downstream of the pricing data itself. The researchers cross-referenced roughly 600 advertisements tied to about 400 articles against published literature and found 53 published papers whose titles matched an advertised slot. Of those 53 identified papers, only 5 have been retracted. That gap — between ads that appear to correspond to real, identifiable published papers and the retraction record for those same papers — is a concrete illustration of how far detection lags the market it is meant to police, and it echoes a pattern China’s National Health Commission’s 2026 misconduct disclosures touched on only in passing, where ‘co-authorship slots sold’ was listed as one enforcement category among several, without pricing detail.

Why a pricing dataset matters for research administration

For journal editors, integrity officers, and research administrators, BuyTheBy converts an anecdotal problem into something closer to a market with observable structure — price tiers, seller geography, product categories, and a visible retraction shortfall. That has practical implications: pricing and product-category data can, in principle, help integrity teams calibrate which submission patterns (unusually large multi-author teams, late author-list changes, submissions from institutions with no prior connection to the study) merit closer scrutiny, in the same way paper mill detection more broadly has moved from manual sleuthing toward systematic screening tools. It also sharpens the distinction between the authorship-integrity problems CASRAI’s dictionary already documents — gift authorship (an unearned byline given to someone with institutional standing) and ghost authorship (an uncredited contributor omitted from the byline) — from a third, transactional pattern: a real byline sold outright to a buyer with no contribution to the work at all, which is better understood as a discrete category of research misconduct and journal-policy violation than as a variant of either.

What to watch next

The dataset’s authors and its Nature coverage frame BuyTheBy as a first systematic step, not a finished picture — seven paper mills is not the whole market, and advertised prices are not necessarily transacted prices. Whether journals, integrity offices, and funders build any standing screening practice on top of pricing and product-category signals like these, rather than treating each paper-mill exposure as a one-off, is the open question the dataset leaves for the next cycle of research-integrity tooling.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →