Skip to main content
v2026.11,610 entries · CC-BY 4.0

Editorial category Track C

AI and ML research outputs

Model cards, system cards, datasheets, benchmarks, evaluation suites.

  • 16 August 2026

    Agentic AI Benchmarks Now Measure Real-World Task Completion — Not Just Q&A Accuracy

    New agentic leaderboards score AI models on whether they actually finish real tasks, recover from errors, and avoid inventing tools — not on multiple-choice accuracy. For institutions weighing AI agents for literature review, data cleaning, or grant administration, that shift changes what reliable enough to deploy means, and exposes governance gaps in policies built around chatbot use rather than autonomous systems.

  • 16 August 2026

    The Frontier LLM Landscape in August 2026: Why No Single Model Fits Every Institutional Use

    Five frontier models now trade the lead by task and price: Opus 5, Fable 5, GPT-5.6, Grok 4.6, Kimi K3. What that means for institutional AI policy.

  • 16 August 2026

    Falling AI Token Costs and the New Math of Institutional Budgets

    Frontier-model pricing now spans roughly two to three orders of magnitude. What that spread means for how research offices should structure AI-tool budgets and disclosure policy.

  • 16 August 2026

    Gemini 3.7 Flash and the Speed-Cost Case for Institutional AI Tools

    Gemini 3.7 Flash is mid-pack on reasoning but the fastest, cheapest model on independent benchmarks. For research-admin screening and triage at scale, that changes the calculus more than the leaderboard rank does.

  • 16 August 2026

    Claude Opus 5’s Adaptive Reasoning Tiers: What Effort-Level Configuration Means for AI Procurement

    Claude Opus 5 turns reasoning effort into a per-query dial (low through xhigh/max) rather than a per-model choice, complicating how institutions license, disclose, and govern AI tools.

  • 8 August 2026

    Argonne’s AI Agent Team Turns a Single Prompt Into a Full Atomistic Simulation

    Argonne National Laboratory and the University of Illinois Chicago built a multi-agent AI framework that plans, runs, and validates atomistic and molecular-dynamics simulations end to end from a single natural-language prompt — distinct from earlier AI screening tools because it automates the simulation pipeline itself. The paper names two DOE Office of Science user facilities in its affiliations, Argonne’s Center for Nanoscale Materials and its Leadership Computing Facility, and the full framework code is public on GitHub under an MIT license.

  • 8 August 2026

    ergoCub: The Humanoid Robot Designed Around the Human Standing Next to It

    ergoCub is a new humanoid robot whose hardware and control software were optimized together around human ergonomics, reducing measured lower-back strain for people lifting alongside it. It was developed through a three-way collaboration between the Italian Institute of Technology (IIT), the University of Manchester, and industry partner GenerativeBionics, described in a new Nature Machine Intelligence paper; specific funding for the study could not be independently confirmed.

  • 8 August 2026

    MIT’s Self-Assembling Robot Boats Snap Together Into Bridges on Command

    MIT researchers and international collaborators have published FloatForm, a fleet of small robotic boats that self-assemble into bridges, platforms, and other floating structures on command, in the open-access journal Nature Communications. The work traces back to the MIT-AMS Institute Roboat collaboration in Amsterdam, illustrating a multi-year, multi-country research lineage published openly from the outset.

  • 8 August 2026

    KAIST’s HOUND Robot Picks Its Own Gait on Stairs and Forest Trails — With a Defense Agency Listed as Co-Author

    KAIST researchers have built APT-RL, a control system that lets their HOUND quadruped robot choose its own gait — trotting or bounding — in real time across stairs, slopes, and forest terrain, reaching peak speeds of about 6 m/s. The paper’s author list, published by KAIST in Science Robotics, names both Korea University and South Korea’s Agency for Defense Development as co-author affiliations — an explicit funder/affiliation transparency case study for research-administration readers tracking dual-use disclosure.

  • 8 August 2026

    DeepMind’s New Robot AI Can Refuse a Bad Command — And Now There’s a Benchmark to Prove It

    Google DeepMind’s Gemini Robotics 2 gives humanoid robots coordinated whole-body control, but the notable part for research administrators is what shipped alongside it: a dedicated Safety Technical Report and a new ASIMOV-Agentic benchmark testing whether the AI will refuse unsafe commands and escalate to a human when uncertain. DeepMind also validated the system across independently-made hardware, including Apptronik’s Apollo 2, Franka Duo, and platforms from Dexmate, SO101, and Trossen, alongside partners Boston Dynamics and Agile Robots.

  • 8 August 2026

    DeepMind’s Open-Sourced WeatherNext Model Gives Forecasters an Extra Day’s Warning on Cyclones

    Google DeepMind’s WeatherNext Cyclones model now produces three-day storm forecasts as accurate as prior two-day forecasts, a gain validated in a peer-reviewed Nature paper, released as open-source code and weights on GitHub, built with NOAA’s National Hurricane Center, the UK Met Office and CIRA/Colorado State University, and trained on the shared IBTrACS storm-track archive. It was used operationally by the National Hurricane Center during Hurricane Melissa in the 2025 season.

  • 7 August 2026

    AlphaFold Maps the Structural Causes of CRISPR-Cas9 Off-Target Editing

    Researchers at Peking University used DeepMind’s AlphaFold3 to map which structural contacts distinguish CRISPR-Cas9’s on-target and off-target binding, then redesigned the enzyme around that map — cutting measured off-target activity from 28% to 5% and pointing gene-editing safety work toward mechanism, not just cataloguing.

  • 7 August 2026

    Stanford Team Uses AI to Design 16 Working Bacteriophages, Exposing a Biosecurity Screening Gap

    Stanford researchers used the Evo genome-language model to design 16 functional synthetic bacteriophages from scratch, and biosecurity specialists warn the AI-generated genomes evade existing DNA-synthesis screening tools built on known-pathogen databases.

  • 7 August 2026

    AI Screened Billions of Compounds — and Found Two New Superconductors

    An Aalto University-led consortium used machine-learning pre-screening plus first-principles calculations to predict two new kagome superconductors, YRu3B2 and LuRu3B2, later confirmed in the lab — a proof of concept for AI-accelerated materials discovery, published in Physical Review Research.

  • 7 August 2026

    An AI That Resizes Proteins Without Breaking

    A generative AI system called Raygun, described in Nature on 29 July 2026, can shrink, expand and edit natural proteins while aiming to preserve their fold and functional sites — reframing protein length as a designable variable rather than a fixed constraint, and sharpening a validation bottleneck that already outpaces wet-lab screening capacity.

  • 7 August 2026

    What Tsimerman’s Move From U of T to OpenAI Signals for Math Departments

    Fields Medalist Jacob Tsimerman’s move from a University of Toronto professorship to OpenAI’s AI-safety team, announced hours after his July 2026 award, is a concrete signal that frontier AI labs can now recruit at the very top of pure mathematics — and that proof-grade verification rigor is becoming a recruited AI-safety specialty.

  • 6 August 2026

    Five Tools That Check for Hallucinated Citations — and What They Still Miss

    A July 2026 arXiv study tests five hallucinated-citation checkers and finds none reliable enough to run unsupervised. What each tool gets right, and wrong.

  • 6 August 2026

    A Bad Dataset Came Down for Copyright, Not for Being Wrong

    Kaggle removed a flawed stroke-detection dataset in July 2026 after a copyright complaint — five months after a research-integrity complaint about the same data went nowhere. Three Scientific Reports papers built on it have now been retracted.

  • 6 August 2026

    NSF Opens $100M State and Regional AI Infrastructure Hubs: Who Can Apply

    NSF 26-513 opens $100M for up to 10 State and Regional AI Infrastructure Hubs (Nov. 4 deadline). NSF funds coordination and training; consortium partners must fund the compute itself. Here is who can lead, what is required, and the eligibility rules that matter most.

  • 4 August 2026

    aiXiv: A Preprint Server Where AI Writes the Papers and AI Reviews Them

    aiXiv is a new preprint platform, reportedly linked to researchers at the University of Toronto, Oxford, and Tsinghua, that accepts AI-written papers and reviews them using multiple LLM agents instead of human referees. It is a direct counterpoint to arXiv’s recent moves to restrict unchecked AI content — and a new BadScientist study shows LLM reviewers can be fooled by convincing-but-unsound fabricated papers, raising real questions for how research administrators and librarians should treat its output.

  • 30 July 2026

    FDA’s Elsa 4.0 and HALO: Sponsor Impact

    FDA shipped Elsa 4.0 and the HALO data platform in May 2026 — an internal AI upgrade, not a sponsor-facing one. Here is what changed and what did not.

  • 30 July 2026

    FDA RFI: AI-Enabled Early-Phase Trials Pilot

    FDA’s April 2026 RFI sought input on an AI-enabled early-phase trials pilot. Comment window closed June 29, 2026 — here’s what it covered.

  • 30 July 2026

    BMJ Study: AI Flags 9.9% of Cancer Papers

    A BMJ study used a BERT classifier to screen 2.6 million cancer papers, flagging 9.9% as suspected paper-mill output — a scale manual review cannot match.

  • 30 July 2026

    MLRC 2026 Becomes an Official NeurIPS Track

    NeurIPS 2026 has made the Machine Learning Reproducibility Challenge an official track, with TMLR-accepted papers presented in Sydney this December.

  • 29 July 2026

    EACL 2026: No Limitations Section, No Review

    EACL 2026’s call for papers confirms that submissions missing a dedicated Limitations section are desk rejected without review — a completeness check, not a quality judgment, that flows from ACL Rolling Review’s shared policy across the ACL, EMNLP, and NAACL submission pipeline.

  • 24 July 2026

    Agents4Science: The Stanford/Together AI Conference That Required AI Systems as First Authors

    Agents4Science 2025, organized by researchers at Stanford and Together AI, ran as a one-day virtual conference on October 22, 2025 with an unusual submission rule: an AI system had to be the paper’s primary author, generating the hypotheses, running the experiments, and writing the manuscript, with a human listed as a supervising co-author. It is one of the most concrete tests to date of how authorship credit and disclosure norms hold up when the ‘researcher’ doing the work is an AI agent.

  • 24 July 2026

    Fully AI-Generated Papers Are Now Passing Peer Review: What Zochi and Sakana Mean for Authorship

    Zochi (Intology) and Sakana’s AI Scientist each passed real peer review in 2025 — one at ACL, one at an ICLR workshop. What it means for authorship.

  • 24 July 2026

    Senate Judiciary’s July 2026 PERA Hearing: Stakes for University Diagnostics and AI Patents

    Senate Judiciary July 14, 2026 hearing on PERA (S.1546) weighed patent eligibility for diagnostics and AI — key stakes for university tech transfer.

  • 24 July 2026

    HORIZON-ZEN+: AI Curation Tools Coming to Zenodo’s EU Open Research Repository

    HORIZON-ZEN+, a CERN-coordinated EU project (2025-2027), adds AI-assisted curation and discovery tools to Zenodo’s EU Open Research Repository.

  • 23 July 2026

    OSTP’s ‘Golden Age’ Report: What Grants Teams Need

    OSTP’s July 2026 report and FY2028 memo direct agencies to fund researchers over institutions and diversify grants beyond peer review.

  • 23 July 2026

    AI Image-Integrity Screening Goes Mainstream: MDPI and ASM Adopt Proofig, Imagetwin

    MDPI’s multi-year Proofig AI deal and ASM’s Imagetwin pilot data show AI image-forensics screening becoming standard in publisher editorial workflows.

  • 23 July 2026

    Scholarly Publisher AI Licensing Deals: Inside the 2026 Numbers

    Wiley disclosed $49M in FY2026 AI licensing revenue ($110M lifetime). Taylor & Francis and Springer Nature show similar deals; small publishers see little.

  • 23 July 2026

    NeurIPS 2026 Pangram AI-Detector Desk Rejections

    NeurIPS 2026 desk-rejected 178 position papers (18.4%) via the Pangram AI detector, no appeal allowed. What is verified, what is disputed, and why it matters.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →