The Machine Learning Reproducibility Challenge (MLRC) is, for the first time since it launched in 2018, an official track of NeurIPS — one of the largest and most influential machine learning conferences. NeurIPS organizers confirmed in a blog post published May 4, 2026 that MLRC 2026 accepted papers will be presented in person at NeurIPS 2026 in Sydney, Australia (December 6–13, 2026), alongside papers from the Main Track and the Evaluations & Datasets Track.
For a field where reproducibility work has historically lived on the margins — workshop slots, standalone challenge websites, informal blog write-ups — formal integration into a top-tier conference’s official track structure is a meaningful signal about how the ML research community is choosing to institutionalize reproducibility as a first-class scientific contribution rather than a secondary check on someone else’s work.
What MLRC is and what changes
MLRC invites researchers to attempt to reproduce, replicate, or extend the generalizability of previously published machine learning results and report what they find — including negative or partial results, which the challenge has consistently stated it values on equal footing with confirmations. Since 2018 it has run as an independent, community-organized effort (see reproml.org). What changes in 2026 is procedural and institutional: submissions no longer route around NeurIPS as a separate, parallel activity. Instead, per NeurIPS’s own official announcement, papers must first be accepted by the Transactions on Machine Learning Research (TMLR) journal, then undergo a lightweight compatibility review by the MLRC organizing committee before being slotted into in-person NeurIPS 2026 presentation.
Submission scope
NeurIPS’s call for reproducibility work at neurips.cc/Conferences/2026/CallForReproducibility and the MLRC 2026 call for papers describe a submission scope broader than a straightforward “redo the experiment” mandate. Eligible work includes rigorous reproductions and replications of results published in top ML venues from 2025 onward, generalizability studies that test whether published findings hold in new settings, meta-reproducibility analyses, tooling and methods that make reproducibility easier to achieve, and — notably for a track hosted inside an AI conference — studies examining AI systems themselves as the subject of reproducibility questions, alongside AI-assisted reproducibility research methods.
Timeline
Per the same primary sources, the soft deadline for expressing intent to submit was June 4, 2026 (AOE), with a hard deadline of September 30, 2026 for underlying TMLR acceptance decisions to be in the system. Author notifications for the MLRC track were set for October 7, 2026, ahead of in-person presentation at NeurIPS 2026 in Sydney in December. As with any conference process still in motion at the time of writing, the final accepted-paper slate and program details are not yet public; treat specifics beyond the dates above as subject to NeurIPS’s own updates.
Why this matters beyond the ML research community
CASRAI’s audience is mostly research administrators, data stewards, and research-integrity staff rather than ML practitioners directly — but this development is relevant to that audience for a specific reason: it is an unusually concrete example of a research field changing its incentive structure to reward reproducibility work as a publishable, career-credentialed output, rather than treating it as an unfunded, unrewarded service activity. That is the same structural problem that reproducibility-crisis literature across psychology, biomedicine, and other fields has identified for over a decade: reproduction attempts are valuable to the field but rarely rewarded to the individuals who do them. A major conference building a formal, TMLR-gated, in-person-presented pathway for exactly that work is a data point research-integrity offices and funders can point to when making the case that reproducibility deserves institutional support — data management planning, computational infrastructure, staff time — rather than remaining an unfunded expectation.
It also reinforces requirements that will look familiar to anyone working with reproducible AI experiment criteria or research data management: the MLRC review process depends on the original work having been documented well enough to attempt reproduction at all, which puts pressure back up the pipeline toward better computational reproducibility practice — code availability, environment/container specification, and clear documentation of methods — at the point of original publication, not just at the point someone tries to check the work later.
Related CASRAI resources
- Reproducible AI experiment — the operational definition of what makes an ML/AI experiment reproducible in the first place.
- Reproducibility crisis — the wider, cross-disciplinary context this track is a response to.
- Computational reproducibility and methods reproducibility — the specific reproducibility sub-types most relevant to ML work.
- Reproducibility vs. replicability: what’s the difference?
- Reproducibility infrastructure: workflows, containers, and code sharing — the practical infrastructure this kind of review depends on.
- NeurIPS 2026’s Pangram AI-detector desk-rejection controversy — another 2026 NeurIPS policy change affecting how submissions are screened, in this case for undisclosed AI-generated text rather than reproducibility.
Sources
- NeurIPS Blog, “MLRC 2026: Reproducibility as an Official Track at NeurIPS,” published May 4, 2026: blog.neurips.cc
- NeurIPS 2026 Call for Reproducibility: neurips.cc/Conferences/2026/CallForReproducibility
- MLRC 2026 official site and call for papers: reproml.org, reproml.org/call_for_papers







