Most of the debate over AI and scholarly publishing this year has been about how to keep AI-generated material out of the literature. aiXiv takes the opposite bet: it is a new preprint platform, described by its creators as a “structured peer-review and iterative publishing infrastructure for the AI-era of scalable scientific output,” built specifically to accept papers written by AI systems and to review them using AI agents rather than human referees. Coverage of the project in Science and the platform’s own preprint describe it as originating from a group of collaborators connected to the University of Toronto, the University of Oxford, and Tsinghua University; one of the people credited with creating it, Guowei Huang, has been described in press coverage as a PhD candidate at the University of Manchester. CASRAI has not independently verified the complete author/institution list beyond what is stated in that coverage and in the project’s own arXiv preprint (arXiv:2508.15126), and readers should treat the institutional framing as reported rather than independently confirmed by CASRAI.
What aiXiv is
aiXiv describes itself as an open platform for both human and AI “scientists” to submit research proposals and papers, have them reviewed, and revise them through iterative feedback cycles — explicitly including work where an AI system, not a person, generated the hypotheses, ran the analysis, or wrote the manuscript. That puts it in a similar conceptual space to Agents4Science, the Stanford/Together AI-organized conference that required an AI system to be listed as primary author — but the two are not the same kind of thing. Agents4science was a one-day, single-event conference with a submission window. aiXiv is positioning itself as standing publishing infrastructure: an ongoing repository and review pipeline, not a one-off program.
How the review pipeline reportedly works
According to the project’s own description and reporting on it, each submission to aiXiv is assessed by multiple large language model “agents” — reported as five — evaluating factors that reportedly include novelty, technical soundness, and potential impact. Based on the agents’ collective recommendations, a paper either proceeds to being posted publicly or is returned to the author (human or AI) for revision and resubmission, with the platform claiming iterative revision cycles measurably improve output quality. Review is reported to take on the order of minutes rather than the months typical of conventional peer review, and the platform states it includes safeguards against submissions that attempt to embed hidden instructions intended to manipulate the AI reviewers into a favorable verdict. CASRAI is not stating a specific numeric acceptance threshold (e.g., how many of the reviewing agents must recommend acceptance) as fact here — secondary coverage references a threshold, but CASRAI could not confirm the exact figure against a primary source at the time of writing, and is omitting it rather than risk restating an unconfirmed number as established fact. Readers who need the precise mechanism should consult aiXiv’s own preprint (arXiv:2508.15126) directly.
Two preprint servers, opposite bets
The timing makes the contrast hard to miss. Over the past year, arXiv — the largest and most established preprint server in physics, math, and computer science — has moved to restrict unchecked AI-generated content and has tightened its endorsement and language policies specifically in response to concerns about low-quality or AI-assisted submissions overwhelming moderators and diluting the record. aiXiv is making the opposite argument: that AI-authored and AI-reviewed work deserves its own transparent, disclosed venue rather than being kept out of preprint servers altogether or, worse, entering human-reviewed venues without disclosure. Both platforms are responding to the same underlying pressure — a rising volume of AI-assisted or AI-generated manuscripts — and have reached opposite conclusions about what preprint infrastructure is for.
The case for a dedicated AI-native venue
The argument in aiXiv’s favor does not require believing AI-written papers are as good as human-written ones. It requires only two things being true: that human peer review already does not scale to the volume of research output being produced (a strain reviewer-fatigue and reviewer-compensation debates have been documenting for a while), and that AI-generated research is going to keep being produced regardless of whether a venue exists for it. If both are true, a clearly labeled, disclosed venue for machine-generated and machine-reviewed work is arguably preferable to the alternative: that work quietly entering human-reviewed journals and conferences with AI involvement undisclosed, which is precisely the failure mode behind reports of fully AI-generated papers passing conventional peer review and AI-assisted paper-laundering through conventional venues. On this view, aiXiv is less a lowering of standards than an attempt to make an already-happening phenomenon visible and auditable.
The case against it — and the evidence that specifically undercuts an all-AI review loop
The case against aiXiv is not merely a general skepticism of AI-generated research. There is a specific, directly relevant piece of adversarial evidence: a study titled “BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?” (arXiv:2510.18003). The paper tests whether an AI agent instructed to fabricate results — using presentation and framing tactics rather than real experiments — can get a fabricated paper past an LLM-based review system. It found that fabricated papers were able to achieve meaningfully high acceptance rates, and identified what the authors call a “concern-acceptance conflict”: reviewer agents would sometimes explicitly flag integrity concerns in a submission and still assign it an acceptance-level score anyway. The paper’s proposed mitigations produced only marginal improvement, with detection performance not far above chance. In other words: the exact failure mode aiXiv’s review pipeline would need to avoid — convincing-but-unsound work getting past LLM reviewers — is one that has been demonstrated, not merely hypothesized, in a closely analogous setting. An all-agent review pipeline does not automatically inherit that vulnerability, but it cannot be assumed immune to it either, and aiXiv has not published evidence that its own five-agent process is resistant to this specific attack.
Beyond BadScientist, the structural objections are familiar from other AI-and-publishing debates but sharpen here because there is no human author of record: if an aiXiv paper turns out to be wrong, fabricated, or built on a hallucinated citation, there is no person to contact, correct, or sanction in the way a journal or institution can address a human author’s misconduct. That compounds the accountability gap already inherent to preprint posting generally, where content goes live before any formal review. And because aiXiv content is openly posted and presumably crawlable, there is a real provenance and pollution risk: AI-generated preprints of uncertain reliability entering citation graphs, training corpora, and evidence syntheses, with downstream systems unable to easily distinguish a rigorously produced preprint from a convincingly written but unsound one — a concern that connects directly to the broader debate over provenance and traceability in AI training data drawn from the research literature.
What research administrators and librarians actually need to decide
This is the question CASRAI’s audience is best positioned to act on, and it does not yet have a settled answer at most institutions. A handful of concrete decisions follow directly from aiXiv’s existence:
- CRIS and repository ingestion. Should aiXiv postings be harvestable into institutional CRIS or repository systems as research outputs at all, and if so, under what output type — treated like any other preprint, or flagged distinctly as AI-generated/AI-reviewed content pending its own metadata field?
- CVs and grant applications. Can an aiXiv posting be listed on a CV or cited as preliminary work in a grant application in the same way a conventional preprint can? Funders and institutions that have not yet addressed AI-authored outputs in their research-output policies will need to, rather than defaulting to silence.
- Systematic review and evidence synthesis. Should aiXiv content be eligible for inclusion in systematic reviews or evidence syntheses at all, given that unlike arXiv, bioRxiv, or medRxiv it explicitly accepts AI-authored and AI-reviewed material by design rather than as an unintended side effect of insufficient moderation? Screening protocols built around excluding non-peer-reviewed grey literature may need an explicit rule for AI-native venues specifically, not just a general preprint exclusion.
- Institutional policy. Many research-integrity and publication policies were written before an AI-native, AI-reviewed venue existed. Whether an institution needs a specific position on aiXiv-type platforms, or whether existing AI-disclosure and authorship policy already covers it, is itself a decision worth making deliberately rather than by default.
Open questions
Several practical questions remain unresolved and worth watching: whether major indexers — Crossref, OpenAlex, Google Scholar — will crawl and index aiXiv content the way they do established preprint servers, which shapes how discoverable and citable it becomes; whether aiXiv postings receive persistent identifiers (DOIs) at all, and if so, through what registration agency; how retraction or correction would even work for a paper with no accountable human author; and how the platform will handle versioning as papers move through its iterative agent-review-and-revise cycles. aiXiv is, by its own account, still early — reporting has described it as hosting a modest number of papers and proposals so far — and much of how it is used, indexed, and cited will depend on decisions still to be made both by the platform and by the institutions deciding whether to recognize its output.







