If your study involves an AI or machine-learning system, the standard reporting guideline for your study design (CONSORT, SPIRIT, TRIPOD, STARD) is very likely no longer sufficient on its own — a family of AI-specific extensions now exists, and using the wrong one, or none, is itself a reporting-quality problem journals increasingly flag at submission. These extensions were built the same way every EQUATOR-listed guideline is built: by extending an existing, general-purpose checklist with items specific to the new risk (here, an algorithm) rather than starting over. This page is a decision guide to the AI-specific family — which one applies to which stage of study, what each one actually covers, and where the checklists live.
Which AI reporting guideline applies to your study?
| Your study is… | Use | Reporting stage | Published |
|---|---|---|---|
| A clinical trial protocol for an intervention that includes an AI/ML component, written before enrollment starts | SPIRIT-AI | Protocol | 2020 — BMJ, Nature Medicine, Lancet Digital Health (simultaneous) |
| The results of that same trial, once it has completed | CONSORT-AI | Trial results | 2020 — BMJ, Nature Medicine, Lancet Digital Health (simultaneous) |
| A study developing or validating a diagnostic or prognostic prediction model that uses regression or machine-learning methods | TRIPOD+AI | Model development / validation | 2024 — BMJ |
| A small-sample, early clinical try-out of an AI-driven decision-support tool, after development but before a full-scale trial | DECIDE-AI | Early-stage clinical evaluation | 2022 — Nature Medicine and BMJ (simultaneous) |
| An AI/ML study specifically in medical imaging (any stage) | CLAIM | Imaging-specific, any stage | 2020 — Radiology: Artificial Intelligence |
| A diagnostic-accuracy study evaluating an AI-based test against a reference standard | STARD-AI (in development — see note below) | Diagnostic accuracy | Not yet finalized |
The pattern across the first four rows is the same logic CASRAI’s EQUATOR Network guide lays out for reporting guidelines generally: the guideline is picked by what stage of the research your write-up describes, not by preference. SPIRIT and CONSORT are a matched pair (protocol vs. results) for the same trial; TRIPOD+AI belongs to prediction-model studies, which are not trials at all; and DECIDE-AI fills a gap those don’t cover — the early, small, iterative clinical try-outs that happen after an algorithm is built but before it is trial-ready.
CONSORT-AI: reporting trial results for AI interventions
CONSORT-AI is the extension of the CONSORT statement for randomized controlled trials that evaluate an intervention involving artificial intelligence. It was developed by the SPIRIT-AI and CONSORT-AI Working Group, led by Xiaoxuan Liu, Samantha Cruz Rivera, David Moher, Melanie Calvert and Alastair Denniston, and published simultaneously in September–October 2020 in the BMJ, Nature Medicine, and Lancet Digital Health.
Rather than replacing the base CONSORT checklist, CONSORT-AI adds AI-specific reporting items layered on top of it — covering things the original CONSORT items don’t anticipate: a precise description of the AI intervention and the version/build actually used, how input data was acquired and handled (including any pre-processing), how human users interacted with the system’s output during the trial, how the system’s outputs and any human-AI disagreement or override were analyzed, and whether the algorithm, its code, or its outputs are accessible for independent scrutiny. CASRAI’s CONSORT checklist guide covers the base CONSORT 2025 statement (which superseded CONSORT 2010 in April 2025) in full; CONSORT-AI is layered on top of whichever CONSORT version is current when the trial is written up, not a standalone replacement for it.
CONSORT-AI applies once a trial has been conducted and you are reporting its results. If you are still writing the protocol, the sibling guideline below applies instead.
SPIRIT-AI: reporting trial protocols for AI interventions
SPIRIT-AI is CONSORT-AI’s protocol-stage counterpart — the same working group produced both, and both were published on the same day in the same three journals in 2020. Where CONSORT-AI governs how you report what happened in a completed trial, SPIRIT-AI governs how you specify, in advance, what an AI intervention is and how it will be evaluated. In practice this means describing the intended use case, the algorithm’s inputs and outputs, the version-control and update plan for the model during the trial, and how human decision-makers will be expected to use its output — all decided and documented before a single participant is enrolled, which is the entire point of a protocol-stage guideline: it exists to be checked against, after the fact, by the CONSORT-AI report.
TRIPOD+AI: reporting prediction models that use AI/ML
TRIPOD+AI is the 2024 update to the TRIPOD statement, extended to explicitly cover prediction models built with machine-learning methods alongside traditional regression. It was published in the BMJ in 2024, led by Gary Collins, Karel Moons and a large multidisciplinary panel, and introduced an expanded 27-item checklist (plus a companion TRIPOD+AI for Abstracts checklist) with more detailed explanation of each item than the original 2015 TRIPOD statement carried.
TRIPOD+AI is not a trial-reporting guideline at all — it applies to studies that develop a diagnostic or prognostic prediction model, validate one on new data, or do both, regardless of whether the model is deployed inside a clinical trial. A model that predicts hospital readmission risk from an electronic health record, for instance, is a TRIPOD+AI study whether or not it is ever tested in a randomized trial; if it later is tested in one, that trial’s results would be reported separately under CONSORT-AI. The “+AI” naming reflects that this single, unified checklist now covers both regression-based and machine-learning-based prediction models, rather than the field maintaining a separate track for each.
DECIDE-AI: reporting early-stage clinical evaluation of AI decision support
DECIDE-AI (DEvelopmental and exploratory Clinical Investigations of DEcision support systems driven by Artificial Intelligence) is a reporting guideline for the small-scale, early clinical evaluations that happen between algorithm development and a full-scale trial. It was developed by Baptiste Vasey, Myura Nagendran, Alastair Denniston, Peter McCulloch and the DECIDE-AI expert group, and published simultaneously in Nature Medicine (2022;28:924–933) and the BMJ (2022;377:e070904).
It exists to fill a specific gap: most AI decision-support systems are piloted, in a handful of real clinical users and a small number of cases, before anyone commits to a full CONSORT-AI trial. Those early pilots have historically been reported inconsistently or not reported at all, which is exactly the problem DECIDE-AI targets — it asks for structured reporting of things like how clinicians actually used the system’s output in real (not simulated) practice, usability and workflow issues encountered, and any safety signals observed, at a stage where sample sizes are too small for the statistical machinery CONSORT-AI assumes.
The rest of the AI reporting-guideline family
Two further items are worth knowing about even though they sit slightly outside the SPIRIT-AI/CONSORT-AI/TRIPOD+AI/DECIDE-AI core:
- CLAIM (Checklist for Artificial Intelligence in Medical Imaging) — published in Radiology: Artificial Intelligence in 2020, CLAIM is widely used across radiology and other medical-imaging AI research. It isn’t part of the EQUATOR-coordinated SPIRIT-AI/CONSORT-AI/TRIPOD+AI/DECIDE-AI family and isn’t tied to a single study stage the way those are; it’s a general-purpose checklist for AI/ML studies specifically in the imaging domain, and many imaging journals ask for it directly.
- STARD-AI — an AI-specific extension to STARD (the standard reporting guideline for diagnostic-accuracy studies) has been under development, but as of this writing does not have a finalized, published checklist the way the four guidelines above do. If a diagnostic-accuracy study is your use case, check the EQUATOR Network’s Reporting Guidelines Library directly for STARD-AI’s current status before citing it as complete, and fall back to base STARD plus a transparent, item-by-item description of the AI-specific elements CONSORT-AI and TRIPOD+AI ask for elsewhere in the meantime.
Why a separate guideline exists for each stage, rather than one AI checklist
It’s a reasonable question why the field didn’t simply publish one “AI reporting guideline” instead of four-plus overlapping ones. The answer is the same reason CASRAI’s broader reporting guidelines coverage gives for the non-AI guidelines: a protocol, a completed trial, a prediction-model validation study, and a small early pilot are different documents answering different questions, written by people who often haven’t seen each other’s write-up, and a single merged checklist would either be too generic to catch stage-specific failure modes or too long to be usable. Extending CONSORT, SPIRIT, TRIPOD and inventing DECIDE-AI where no predecessor existed kept each guideline anchored to a document type reviewers and editors already recognize, while adding exactly the AI-specific items that document type was missing.
Frequently asked questions
What does CONSORT-AI add to the base CONSORT checklist?
It layers AI-specific items on top of the standard CONSORT checklist covering the AI intervention’s description and version, how input data was acquired, how the system’s output was used by human decision-makers during the trial, how errors and human-AI disagreement were analyzed, and whether the algorithm or its outputs are available for scrutiny. It does not replace the base CONSORT items — a CONSORT-AI trial report still needs to satisfy the current CONSORT statement in full.
Is TRIPOD+AI the same thing as “TRIPOD-AI”?
TRIPOD+AI, published in 2024, is the current, unified checklist covering both regression-based and machine-learning-based prediction models under a single 27-item statement. Earlier references to a separate “TRIPOD-AI” track reflect the guideline’s development history rather than a currently maintained parallel document — TRIPOD+AI is the citation to use now.
Do I need DECIDE-AI if I’m already running a full randomized controlled trial?
No — DECIDE-AI is specifically for the small-scale, early clinical evaluation stage that typically precedes a full trial. If your study is a properly powered RCT of an AI intervention, CONSORT-AI (for the results) and SPIRIT-AI (for the protocol) are the applicable guidelines instead.
Which guideline applies to a diagnostic AI algorithm evaluated against a reference standard, outside of imaging?
That’s a diagnostic-accuracy study, which base STARD covers; an AI-specific STARD extension has been in development but isn’t finalized as of this writing, so check the EQUATOR Network’s library for its current status, and be transparent in the meantime about the AI-specific details (model version, training/test data provenance, human-AI interaction) that CONSORT-AI and TRIPOD+AI ask for in their own domains.
Where can I find the official checklists themselves?
All of these are catalogued in the EQUATOR Network’s Reporting Guidelines Library, alongside the base guidelines they extend. CASRAI’s own guide to choosing a reporting guideline is a good starting point if you’re not yet sure which base guideline (CONSORT, SPIRIT, TRIPOD, STARD) applies before layering an AI extension on top of it.
Last verified: 16 August 2026, against EQUATOR Network and PubMed listings for each guideline’s original publication.







