Editorial category Track C
AI and ML research outputs
Model cards, system cards, datasheets, benchmarks, evaluation suites.
- 16 August 2026
Agentic AI Benchmarks Now Measure Real-World Task Completion — Not Just Q&A Accuracy
New agentic leaderboards score AI models on whether they actually finish real tasks, recover from errors, and avoid inventing tools — not on multiple-choice accuracy. For institutions weighing AI agents for literature review, data cleaning, or grant administration, that shift changes what reliable enough to deploy means, and exposes governance gaps in policies built around chatbot use rather than autonomous systems.
- 16 August 2026
The Frontier LLM Landscape in August 2026: Why No Single Model Fits Every Institutional Use
Five frontier models now trade the lead by task and price: Opus 5, Fable 5, GPT-5.6, Grok 4.6, Kimi K3. What that means for institutional AI policy.
- 16 August 2026
Falling AI Token Costs and the New Math of Institutional Budgets
Frontier-model pricing now spans roughly two to three orders of magnitude. What that spread means for how research offices should structure AI-tool budgets and disclosure policy.
- 16 August 2026
Gemini 3.7 Flash and the Speed-Cost Case for Institutional AI Tools
Gemini 3.7 Flash is mid-pack on reasoning but the fastest, cheapest model on independent benchmarks. For research-admin screening and triage at scale, that changes the calculus more than the leaderboard rank does.
- 16 August 2026
Claude Opus 5’s Adaptive Reasoning Tiers: What Effort-Level Configuration Means for AI Procurement
Claude Opus 5 turns reasoning effort into a per-query dial (low through xhigh/max) rather than a per-model choice, complicating how institutions license, disclose, and govern AI tools.
- 8 August 2026
Argonne’s AI Agent Team Turns a Single Prompt Into a Full Atomistic Simulation
Argonne National Laboratory and the University of Illinois Chicago built a multi-agent AI framework that plans, runs, and validates atomistic and molecular-dynamics simulations end to end from a single natural-language prompt — distinct from earlier AI screening tools because it automates the simulation pipeline itself. The paper names two DOE Office of Science user facilities in its affiliations, Argonne’s Center for Nanoscale Materials and its Leadership Computing Facility, and the full framework code is public on GitHub under an MIT license.
- 8 August 2026
ergoCub: The Humanoid Robot Designed Around the Human Standing Next to It
ergoCub is a new humanoid robot whose hardware and control software were optimized together around human ergonomics, reducing measured lower-back strain for people lifting alongside it. It was developed through a three-way collaboration between the Italian Institute of Technology (IIT), the University of Manchester, and industry partner GenerativeBionics, described in a new Nature Machine Intelligence paper; specific funding for the study could not be independently confirmed.
- 8 August 2026
MIT’s Self-Assembling Robot Boats Snap Together Into Bridges on Command
MIT researchers and international collaborators have published FloatForm, a fleet of small robotic boats that self-assemble into bridges, platforms, and other floating structures on command, in the open-access journal Nature Communications. The work traces back to the MIT-AMS Institute Roboat collaboration in Amsterdam, illustrating a multi-year, multi-country research lineage published openly from the outset.
- 8 August 2026
KAIST’s HOUND Robot Picks Its Own Gait on Stairs and Forest Trails — With a Defense Agency Listed as Co-Author
KAIST researchers have built APT-RL, a control system that lets their HOUND quadruped robot choose its own gait — trotting or bounding — in real time across stairs, slopes, and forest terrain, reaching peak speeds of about 6 m/s. The paper’s author list, published by KAIST in Science Robotics, names both Korea University and South Korea’s Agency for Defense Development as co-author affiliations — an explicit funder/affiliation transparency case study for research-administration readers tracking dual-use disclosure.
- 8 August 2026
DeepMind’s New Robot AI Can Refuse a Bad Command — And Now There’s a Benchmark to Prove It
Google DeepMind’s Gemini Robotics 2 gives humanoid robots coordinated whole-body control, but the notable part for research administrators is what shipped alongside it: a dedicated Safety Technical Report and a new ASIMOV-Agentic benchmark testing whether the AI will refuse unsafe commands and escalate to a human when uncertain. DeepMind also validated the system across independently-made hardware, including Apptronik’s Apollo 2, Franka Duo, and platforms from Dexmate, SO101, and Trossen, alongside partners Boston Dynamics and Agile Robots.
- 8 August 2026
DeepMind’s Open-Sourced WeatherNext Model Gives Forecasters an Extra Day’s Warning on Cyclones
Google DeepMind’s WeatherNext Cyclones model now produces three-day storm forecasts as accurate as prior two-day forecasts, a gain validated in a peer-reviewed Nature paper, released as open-source code and weights on GitHub, built with NOAA’s National Hurricane Center, the UK Met Office and CIRA/Colorado State University, and trained on the shared IBTrACS storm-track archive. It was used operationally by the National Hurricane Center during Hurricane Melissa in the 2025 season.
- 7 August 2026
AlphaFold Maps the Structural Causes of CRISPR-Cas9 Off-Target Editing
Researchers at Peking University used DeepMind’s AlphaFold3 to map which structural contacts distinguish CRISPR-Cas9’s on-target and off-target binding, then redesigned the enzyme around that map — cutting measured off-target activity from 28% to 5% and pointing gene-editing safety work toward mechanism, not just cataloguing.
- 7 August 2026
Stanford Team Uses AI to Design 16 Working Bacteriophages, Exposing a Biosecurity Screening Gap
Stanford researchers used the Evo genome-language model to design 16 functional synthetic bacteriophages from scratch, and biosecurity specialists warn the AI-generated genomes evade existing DNA-synthesis screening tools built on known-pathogen databases.
- 7 August 2026
AI Screened Billions of Compounds — and Found Two New Superconductors
An Aalto University-led consortium used machine-learning pre-screening plus first-principles calculations to predict two new kagome superconductors, YRu3B2 and LuRu3B2, later confirmed in the lab — a proof of concept for AI-accelerated materials discovery, published in Physical Review Research.
- 7 August 2026
An AI That Resizes Proteins Without Breaking
A generative AI system called Raygun, described in Nature on 29 July 2026, can shrink, expand and edit natural proteins while aiming to preserve their fold and functional sites — reframing protein length as a designable variable rather than a fixed constraint, and sharpening a validation bottleneck that already outpaces wet-lab screening capacity.
- 7 August 2026
What Tsimerman’s Move From U of T to OpenAI Signals for Math Departments
Fields Medalist Jacob Tsimerman’s move from a University of Toronto professorship to OpenAI’s AI-safety team, announced hours after his July 2026 award, is a concrete signal that frontier AI labs can now recruit at the very top of pure mathematics — and that proof-grade verification rigor is becoming a recruited AI-safety specialty.
- 6 August 2026
Five Tools That Check for Hallucinated Citations — and What They Still Miss
A July 2026 arXiv study tests five hallucinated-citation checkers and finds none reliable enough to run unsupervised. What each tool gets right, and wrong.
- 6 August 2026
A Bad Dataset Came Down for Copyright, Not for Being Wrong
Kaggle removed a flawed stroke-detection dataset in July 2026 after a copyright complaint — five months after a research-integrity complaint about the same data went nowhere. Three Scientific Reports papers built on it have now been retracted.
- 6 August 2026
NSF Opens $100M State and Regional AI Infrastructure Hubs: Who Can Apply
NSF 26-513 opens $100M for up to 10 State and Regional AI Infrastructure Hubs (Nov. 4 deadline). NSF funds coordination and training; consortium partners must fund the compute itself. Here is who can lead, what is required, and the eligibility rules that matter most.
- 4 August 2026
aiXiv: A Preprint Server Where AI Writes the Papers and AI Reviews Them
aiXiv is a new preprint platform, reportedly linked to researchers at the University of Toronto, Oxford, and Tsinghua, that accepts AI-written papers and reviews them using multiple LLM agents instead of human referees. It is a direct counterpoint to arXiv’s recent moves to restrict unchecked AI content — and a new BadScientist study shows LLM reviewers can be fooled by convincing-but-unsound fabricated papers, raising real questions for how research administrators and librarians should treat its output.
- 30 July 2026
FDA’s Elsa 4.0 and HALO: Sponsor Impact
FDA shipped Elsa 4.0 and the HALO data platform in May 2026 — an internal AI upgrade, not a sponsor-facing one. Here is what changed and what did not.
- 30 July 2026
FDA RFI: AI-Enabled Early-Phase Trials Pilot
FDA’s April 2026 RFI sought input on an AI-enabled early-phase trials pilot. Comment window closed June 29, 2026 — here’s what it covered.
- 30 July 2026
BMJ Study: AI Flags 9.9% of Cancer Papers
A BMJ study used a BERT classifier to screen 2.6 million cancer papers, flagging 9.9% as suspected paper-mill output — a scale manual review cannot match.
- 30 July 2026
MLRC 2026 Becomes an Official NeurIPS Track
NeurIPS 2026 has made the Machine Learning Reproducibility Challenge an official track, with TMLR-accepted papers presented in Sydney this December.
- 29 July 2026
EACL 2026: No Limitations Section, No Review
EACL 2026’s call for papers confirms that submissions missing a dedicated Limitations section are desk rejected without review — a completeness check, not a quality judgment, that flows from ACL Rolling Review’s shared policy across the ACL, EMNLP, and NAACL submission pipeline.
- 24 July 2026
Agents4Science: The Stanford/Together AI Conference That Required AI Systems as First Authors
Agents4Science 2025, organized by researchers at Stanford and Together AI, ran as a one-day virtual conference on October 22, 2025 with an unusual submission rule: an AI system had to be the paper’s primary author, generating the hypotheses, running the experiments, and writing the manuscript, with a human listed as a supervising co-author. It is one of the most concrete tests to date of how authorship credit and disclosure norms hold up when the ‘researcher’ doing the work is an AI agent.
- 24 July 2026
Fully AI-Generated Papers Are Now Passing Peer Review: What Zochi and Sakana Mean for Authorship
Zochi (Intology) and Sakana’s AI Scientist each passed real peer review in 2025 — one at ACL, one at an ICLR workshop. What it means for authorship.
- 24 July 2026
Senate Judiciary’s July 2026 PERA Hearing: Stakes for University Diagnostics and AI Patents
Senate Judiciary July 14, 2026 hearing on PERA (S.1546) weighed patent eligibility for diagnostics and AI — key stakes for university tech transfer.
- 24 July 2026
HORIZON-ZEN+: AI Curation Tools Coming to Zenodo’s EU Open Research Repository
HORIZON-ZEN+, a CERN-coordinated EU project (2025-2027), adds AI-assisted curation and discovery tools to Zenodo’s EU Open Research Repository.
- 23 July 2026
OSTP’s ‘Golden Age’ Report: What Grants Teams Need
OSTP’s July 2026 report and FY2028 memo direct agencies to fund researchers over institutions and diversify grants beyond peer review.
- 23 July 2026
AI Image-Integrity Screening Goes Mainstream: MDPI and ASM Adopt Proofig, Imagetwin
MDPI’s multi-year Proofig AI deal and ASM’s Imagetwin pilot data show AI image-forensics screening becoming standard in publisher editorial workflows.
- 23 July 2026
Scholarly Publisher AI Licensing Deals: Inside the 2026 Numbers
Wiley disclosed $49M in FY2026 AI licensing revenue ($110M lifetime). Taylor & Francis and Springer Nature show similar deals; small publishers see little.
- 23 July 2026
NeurIPS 2026 Pangram AI-Detector Desk Rejections
NeurIPS 2026 desk-rejected 178 position papers (18.4%) via the Pangram AI detector, no appeal allowed. What is verified, what is disputed, and why it matters.







