Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Editorial · CASRAI · AI and ML research outputs

NeurIPS 2026 Pangram AI-Detector Desk Rejections

NeurIPS 2026 desk-rejected 178 position papers (18.4%) via the Pangram AI detector, no appeal allowed. What is verified, what is disputed, and why it matters.

Published 23 Jul 2026· 6 minute read

NeurIPS is one of the largest and most influential venues in machine-learning research, and its 2026 Position Paper Track just became a live test case for what happens when an AI-detection tool is used not for author-facing academic-integrity screening, but as an automated desk-rejection gate in peer review itself. In June 2026, the track’s own chairs published a detailed account, on the official NeurIPS blog, of using the AI-detection company Pangram to screen submissions — and of rejecting a meaningful share of them, with no appeal, on the strength of that screening. The numbers are unusual enough, and the methodology questions real enough, that the story is worth separating carefully from the more general AI-detector debate CASRAI has covered elsewhere (see AI Detector False-Positive Controversies: What Researchers Should Know in 2026, which covers the parallel, but distinct, story of AI-writing detectors in student academic-integrity cases).

What NeurIPS actually did

The NeurIPS 2026 Position Paper Track requires that submitted papers be substantially human-written, with generative AI permitted only for narrow, peripheral uses such as copy-editing — not for drafting substantive argument or content. According to the track chairs’ own account, the track received roughly 969-971 submissions and ran all of them through Pangram (version 3.3.2), a commercial AI-text-detection tool that scores a document’s likelihood of being AI-generated by splitting it into overlapping text windows and scoring each window individually.

Based on those scores, the track chairs desk-rejected 178 submissions (18.4% of the track) outright, with no appeal process available to those authors. A further 123 submissions (12.7%) were placed into a conditional category: authors were required to supply evidence — specifically, version histories or document checkpoints showing the paper’s state before and after any AI-assisted edits — by June 15, 2026, or face rejection as well. The chairs of record on the post are Alex Lu, Seth Lazar, and David Rugamer, with Stanley Hua and Kate Metcalf credited as assistant chairs.

The numbers, and why the methodology matters as much as the headline figure

The most striking detail in the chairs’ own writeup is not the final rejection rate but how much that rate moved depending on a configuration choice invisible to submitting authors. Pangram’s default window size (roughly 250-350 words per scored segment) flagged 42.7% of submissions as scoring in the 90-100% AI-probability range on an initial pass — with 273 submissions (about 28.2% of the track) scoring a full 100%. When the track chairs switched to a narrower, 100-word “medium” window size, that same flagged share dropped to 12.7%, which is the figure that ultimately became the conditional-review population.

In other words, the reported rejection rate is a direct function of a parameter the track chairs chose after seeing the default results looked implausibly high — a legitimate calibration response, but also exactly the kind of after-the-fact methodology adjustment that invites scrutiny, especially applied retroactively to submissions that had no visibility into how they were being scored. The chairs’ post does cite external audit figures for Pangram reporting a false-positive rate below 0.1%, and notes that running the same tool against a set of already-accepted ICLR 2026 papers (a comparable machine-learning venue) flagged only about 1% of that corpus as AI-generated — offered as evidence the tool is not simply over-flagging typical ML-paper prose across the board. Independent, peer-reviewed replication of either figure was not part of what the chairs’ post itself documented.

Why it’s controversial

Three things distinguish this from a routine desk-rejection policy. First, the no-appeal design: unlike most peer-review integrity findings, which typically allow a response before a final decision, the 178 outright rejections had no appeal path attached at all — a stricter standard than most journals or conferences apply even to confirmed plagiarism. Second, the scale: 18.4% of an entire track is a large fraction of submissions to reject on a single automated signal, at a major, competitive venue where a rejection has real career consequences for the (often early-career) authors involved. Third, and most discussed in the ML research community’s own public reaction on social media following the post, is the window-size sensitivity itself — a roughly 30-percentage-point swing in flagged submissions based on a scoring-window parameter is difficult to reconcile with confidence in any single reported number, including the final 12.7%/18.4% figures the track ultimately acted on.

This lands in a research-integrity landscape that is already unsettled around AI-detector reliability more broadly. A widely cited 2023 Stanford study (Liang et al., published in Patterns) found GPT-detector tools produced an average 61.3% false-positive rate on TOEFL essays written by non-native English speakers, and CASRAI’s own guide on why detectors misfire on human writing covers the underlying perplexity/burstiness mechanics in more depth. NeurIPS’s own post is transparent that formulaic, dense, jargon-heavy academic prose — exactly the register position papers are often written in — scores differently than typical human writing, which is part of why window size moved the numbers so much.

What’s still unresolved

As of this writing, several material questions remain open and are not answered by the chairs’ own post: whether Pangram’s detection methodology has been independently peer-reviewed; how many of the 178 no-appeal rejections involved papers that used only the permitted, narrow forms of AI assistance (copy-editing) rather than substantive AI drafting; and whether NeurIPS will publish outcomes from the June 15, 2026 conditional-review deadline — how many of the 123 conditionally-flagged authors successfully demonstrated compliant human authorship, and how many were ultimately rejected. Until that follow-up data is published, the case is better understood as an unusually well-documented, still-developing example of AI-detector use in peer review than as a fully resolved episode.

Why this matters beyond one conference

Peer-review desk-rejection is a different use case from the author-facing academic-integrity detection CASRAI has covered elsewhere — there is no instructor-mediated appeal, no due-process hearing, and (in this case) initially no appeal mechanism at all, applied to career-relevant publication decisions at a top-tier venue. It is also a distinct mechanism from the editorial-integrity screening some journal publishers have begun adopting — see CASRAI’s guide to AI detection tools adoption in academic publishing for how that separate trend is unfolding. Research administrators, conference organizers, and program committees weighing whether to adopt a similar AI-detection gate for their own review processes should treat the NeurIPS 2026 episode as a live case study in the tradeoffs involved: detection tools can surface genuine policy violations, but a single automated score, applied at scale with no appeal and with results this sensitive to configuration choices, carries real risk of penalizing legitimate human-authored work. For the underlying definitions and disclosure norms this intersects with, see CASRAI’s dictionary entries on AI-generated detection tools and generative-AI disclosure statements.

Sourcing note: the figures above are drawn from the NeurIPS 2026 Position Paper Track chairs’ own account, published on the official NeurIPS blog, corroborated by independent ML-industry news coverage. Any specific percentage or outcome should be treated as reflecting the state of the process as documented in that June 2026 post; readers should check NeurIPS’s own channels directly for any subsequent update, including published outcomes of the June 15, 2026 conditional-review deadline.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →