In 2025, two AI research-agent systems each produced a paper that went through a real, unmodified peer-review process at a legitimate venue — and passed. Intology AI’s Zochi had a paper accepted into the main proceedings of ACL 2025, a top-tier natural-language-processing conference. Sakana AI’s AI Scientist-v2 had a paper accepted into a workshop at ICLR 2025. Neither is a stunt claim without documentation — both are described in detail by the developers themselves, and the distinction between the two (main conference vs. workshop track, and how much human involvement sat between the system’s output and the submitted manuscript) matters more than the headline “AI writes a paper” framing suggests. This piece lays out what actually happened in each case, then works through what it means for authorship attribution and research-integrity oversight going forward.
Two systems, two real acceptances
Zochi at ACL 2025: a main-conference acceptance
Intology AI’s Zochi system produced a paper titled “Tempest: Automatic Multi-Turn Jailbreaking of Large Language Models with Tree Search,” submitted to the main proceedings of the Association for Computational Linguistics (ACL) 2025 conference — not a workshop track. According to Intology’s own writeup, the paper received a final meta-review score of 4 out of 5 and placed within the top 8.2% of all ACL 2025 submissions, at a conference with an acceptance rate of roughly one in five. Intology describes this as the first instance of an AI system’s research reaching a top-tier main-conference bar, distinct from the workshop-level acceptances (including Sakana’s, below) that preceded it.
Two details matter for an accurate picture of what happened. First, human involvement was not zero: Intology states the system operated with minimal human intervention limited to “figure creation, citation formatting, and minor fixes,” and that human researchers handled all reviewer communication and wrote the rebuttal manually, without AI assistance. Second, and directly relevant to the authorship question below, the human researchers at Intology remain the accountable authors of record — Intology’s own framing is that its team “remain as authors and bear responsibility for validation and ethical compliance,” with Zochi credited as the system that performed the research, not listed as a byline author standing in place of a human. Intology has said it wants to work with the academic community on “comprehensive frameworks for AI participation in research, including authorship standards” — i.e., even the team behind this result is treating authorship policy as unresolved, not as something this acceptance settles.
The AI Scientist-v2 at an ICLR 2025 workshop
Sakana AI ran an earlier, methodologically different experiment: a manuscript produced by its “AI Scientist-v2” agent was submitted, under the system’s own generated content, to the ICLR 2025 workshop “I Can’t Believe It’s Not Better: Challenges in Applied Deep Learning.” The paper — on whether an explicit compositional-regularization term improves compositional generalization in neural networks — received an average reviewer score of 6.33, which Sakana reports would have cleared the workshop’s acceptance bar after meta-review, making it (by Sakana’s account) the first fully AI-generated manuscript to pass a standard peer-review process end to end. Sakana’s writeup states the experiment was run with the workshop organizers’ knowledge and with institutional review board approval obtained through the University of British Columbia, rather than being submitted covertly to test whether reviewers could be fooled.
The distinction between the two cases is genuinely load-bearing, not a technicality: a workshop track at a major conference typically has a substantially higher acceptance rate and a lighter review bar than the main proceedings track at the same or a comparably ranked venue. Reporting on either case should specify which track was involved — “passed peer review” alone flattens a meaningful difference in what the acceptance actually demonstrates.
What “passing peer review” does and doesn’t mean here
Both results are real and both went through an unmodified, standard review process rather than a review process altered to accommodate an AI submission. Neither result means peer review has been “solved” by AI, and neither is evidence that AI-generated research is now indistinguishable from, or equivalent to, the median accepted paper at these venues. What they show, more precisely, is that current AI research-agent systems — with some human involvement in execution and, in Zochi’s case, in reviewer interaction — can produce work that clears a real editorial and review bar at specific tracks of specific conferences, at least once each, in a field (machine learning / NLP) where the training data, benchmarks, and reviewer expectations happen to overlap heavily with what these systems were built to do. Generalizing from two data points in one field to “AI can now pass peer review” as a blanket claim is not supported by what’s been reported.
The authorship question this raises
Neither system was listed as a named author replacing a human on the accepted manuscript, and that is not incidental — it lines up with where journal and conference authorship policy already stands. Under the ICMJE’s four-part authorship test, a contributor must have substantially contributed to the work, drafted or critically revised it, given final approval of the version to be published, and agreed to be accountable for its accuracy and integrity. A generative or agentic AI system cannot give final approval or be held accountable in any meaningful sense — it cannot be named in a correction, cannot respond to an allegation of misconduct, and cannot be sanctioned. CASRAI covers this policy landscape in more depth in Can AI Be Listed as an Author? ICMJE, COPE, and Publisher Positions and in the ICMJE’s 2023 rejection of AI co-authorship specifically; see also AI as author and generative-AI disclosure statements as the mechanism most publishers now use instead — disclosing the tool’s role in the acknowledgments or methods, while keeping a human as the accountable, named author.
The Zochi and AI Scientist cases fit this pattern rather than breaking it: in both, human researchers remained the named, accountable authors, and the AI system’s role was substantial enough to be described in detail (and, in Sakana’s case, disclosed transparently to the venue in advance) but not substantial enough — under current policy — to earn a byline. What’s genuinely new is the scale and independence of the system’s contribution relative to prior generative-AI-assistance cases (drafting text, generating a figure, running a literature search): these systems performed the ideation, experimentation, and initial writing largely autonomously. That widens the gap between “how much of this did a machine actually do” and “who is named as having done it,” which is precisely the gap authorship policy exists to police, and precisely the gap that gets harder to audit as agentic systems get more capable.
What this means for research-integrity oversight
The oversight implications sit downstream of peer review, not just at the authorship line:
- Screening for undisclosed AI authorship becomes harder as the underlying work gets more sophisticated. A generated figure or a fluent paragraph is one thing to flag; a fully AI-executed research pipeline that a human then edits, verifies, and submits under their own name is a different detection problem entirely, closer to catching an undisclosed ghostwriter than an undisclosed grammar tool.
- These cases are not paper-mill cases, and shouldn’t be conflated with them — both developers disclosed what they’d done and Sakana coordinated directly with the venue — but the underlying capability is the same one bad actors could use covertly. CASRAI’s coverage of paper mills and tortured-phrases detection and of AI-generated-content screening at scale (see ICML 2026’s mass desk-rejections over AI-generated reviews and the NeurIPS 2026 Pangram AI-detector controversy) is directly relevant background: venues are already contending with AI on the reviewer side of the process, and now have a documented, disclosed case of AI on the authoring side too.
- Publisher and conference policy on generative AI in manuscripts is still catching up to what these systems can now do. CASRAI tracks the current policy landscape across major publishers in Journal and Publisher Policies on Generative AI in Manuscripts; most existing policy was written with AI as a drafting or editing aid in mind, not as an autonomous experimental-research agent producing a complete, submission-ready manuscript.
- Institutional research-integrity offices should expect disclosure questions, not just detection questions. Where AI-assistance disclosure has mostly been about text generation and figures, offices should be prepared for researchers legitimately asking how to disclose use of an agentic research system that ran experiments, drafted the manuscript, and left a human to verify and submit it — a materially different case from asking ChatGPT to tighten a paragraph.
How this differs from CASRAI’s coverage of AI-generated figures
CASRAI’s guide on who’s accountable when a figure is AI-generated addresses a narrower, already-common scenario: a single AI-generated element embedded in an otherwise human-authored paper, and who takes responsibility for it under existing image-integrity and authorship policy. The Zochi and AI Scientist cases are broader in scope — a system contributing to or executing most of the research and writing process, not one figure — and are new enough that no publisher has a policy written specifically for this scenario yet. The accountability principle from the narrower case still applies (a named human author is responsible for the whole submission, whatever produced any given part of it), but the scale of what needs verifying, and the stakes if that verification is skipped, are considerably higher when an agent has run the experiments and not just generated an illustration.
Key questions this raises
Does either case mean ACL or ICLR now accepts AI-authored papers as a matter of policy? No. Both are individual results from individual submissions that went through the venues’ standard review process; neither venue has announced a policy specifically endorsing or institutionalizing AI-generated submissions, and the accepted papers were submitted with named human authors of record.
Is a workshop acceptance the same claim as a main-conference acceptance? No, and conflating the two overstates what either result shows on its own — workshop tracks generally have higher acceptance rates and a lighter review bar than main-conference proceedings at the same event.
Could a covert version of this — an AI-executed paper submitted without disclosure — get through review undetected? Both documented cases involved developers who disclosed the AI system’s role afterward (and, for Sakana, coordinated with the venue beforehand); that transparency is exactly why these cases are known at all. It does not establish anything about how reliably an undisclosed version would be caught, which is the harder oversight question these cases raise rather than answer.







