“AI literature review tool” gets used loosely to describe two very different kinds of software: general-purpose AI research assistants that help a researcher discover and summarize papers (see AI-Powered Research Assistant Tools for that broader category), and a narrower set of tools built specifically for the mechanics of a systematic review or other evidence synthesis — title/abstract screening, full-text screening, and structured data extraction against a protocol. This guide covers the second category: what these tools actually do at each stage of a review, what independent (non-vendor) studies have found about their accuracy, and what current evidence-synthesis standards bodies require in terms of human oversight and disclosure. It is a category overview, not a product comparison or a recommendation of one vendor over another.
Where these tools sit in a review workflow
A systematic review reported to PRISMA 2020 standards moves through a defined sequence: a comprehensive, documented search; deduplication of the retrieved records; title/abstract screening against pre-specified eligibility criteria; full-text screening of the records that survive; and structured data extraction from the studies that are ultimately included. PRISMA 2020’s flow diagram groups this into three labeled phases — Identification, Screening, and Included — and AI-assisted tools have concentrated almost entirely on the middle phase (screening) and the data-extraction step that follows inclusion, not on replacing the search strategy itself. This is a meaningfully narrower scope than “AI does my literature review,” and it matters for how these tools should be evaluated and disclosed.
This narrower use case is also distinct from a literature review or scoping review conducted without a registered protocol, and from general guidance on how to write a literature review as a manuscript section — the tools discussed here are built around protocol-driven screening and extraction, and most assume (or actively support) a prospectively registered protocol, such as one filed with PROSPERO.
What screening tools actually do
Two widely used platforms in this space are Rayyan and Covidence, both built around dual-level screening (title/abstract, then full-text) with multiple independent reviewers, inter-rater agreement tracking, and PRISMA-compatible reporting exports. Both have layered machine-learning features on top of that manual workflow, and it is worth being precise about what those features do:
- Rayyan offers a relevance-ranking feature that reorders the queue of not-yet-screened records using a model trained on the review team’s own screening decisions made so far in that project, plus paid-tier features for PICO-based filtering and semi-automated data extraction.
- Covidence offers a comparable “Relevance Sorting” feature that reorders the unscreened queue the same way, an “RCT Classifier” (developed with Cochrane) that flags records likely to be randomized controlled trials, a reviewer-visible and reversible automated-exclusion option for clearly irrelevant records, and an AI-assisted data-extraction feature (available for a subset of extraction templates) that the vendor itself states must be human-verified.
The common shape across both tools is worth noting: the AI layer is marketed as decision support — reordering a queue a human reviewer still works through, or flagging a record for a human to confirm or reverse — rather than as an autonomous substitute for the screening decision itself. That is a different design choice from tools that perform the screening or extraction step directly and report an accuracy figure for having done so.
Elicit: AI-performed screening and extraction
Elicit is the most commonly cited example of the second design: rather than reordering a human reviewer’s queue, it performs abstract and full-text screening and data extraction directly, aligned to a PRISMA-style workflow, and reports its own accuracy against that output. Because Elicit’s own accuracy claims are vendor-reported, they are worth weighing against independent, peer-reviewed evaluations rather than restated as fact:
- A 2025 study in Cochrane Evidence Synthesis and Methods comparing Elicit against the original traditional searches of four published systematic reviews found average search sensitivity around 39-40% for Elicit (range roughly 25.5-69.2% across the four case studies) versus roughly 94.5% for the reviews’ traditional searches — though Elicit’s precision was higher. The authors concluded Elicit was not sensitive enough to replace a traditional systematic-review search on its own, but could be useful as a supplementary or preliminary tool.
- A 2026 feasibility study in Research Synthesis Methods tested Elicit’s data-extraction feature across 70 variables in seven systematic reviews spanning environmental and life sciences. Around 78% of variables met an 87% accuracy threshold during development, falling to about 69% when applied to new, unseen articles; re-extraction by different user accounts matched on the extracted value about 90% of the time but matched on the supporting quote only about 46% of the time and on stated reasoning only about 30% of the time. The tool could not extract data presented only in figures or tables. The authors recommended using it as a secondary reviewer or a sanity check, not as an autonomous replacement for a human extractor.
Both studies land on the same practical conclusion from two different angles (search/screening and extraction): current tools in this category are useful as a supplementary layer that can speed up a human reviewer’s work and catch things a human might miss, but their recall and reliability have not been shown to substitute for a trained reviewer working to a protocol.
Human oversight and disclosure: the RAISE recommendations
In October 2025, Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence issued a joint position statement on AI use in evidence synthesis, endorsing the RAISE (“Responsible use of AI in evidence SynthEsis”) recommendations. The core requirements are directly relevant to any team using screening or extraction AI:
- AI or automation tools must be used with human oversight — not as an unsupervised substitute for reviewer judgment.
- Authors must describe the steps taken to verify AI-generated outputs.
- Any AI use that makes or suggests a judgement call — study eligibility, data extraction, risk-of-bias assessment, or certainty-of-evidence determinations — requires disclosure, including the system’s name and version, its purpose in the review, and its known limitations.
- Basic use for spelling, grammar, or manuscript structure generally does not require this level of disclosure — a permitted/exempt distinction that mirrors how publisher manuscript-AI policies already draw the line for the writing stage, arrived at independently from the evidence-synthesis-methodology side.
This framework applies specifically to the review-conduct stage — screening, extraction, appraisal — and sits alongside, rather than replaces, the separate question of AI and authorship credit addressed by ICMJE and COPE: an AI tool cannot be listed as an author or contributor, because authorship carries an accountability an AI system cannot bear (see Can AI Be Listed as an Author?). Where a review team also uses a generative-AI disclosure statement or a CRediT-based contribution statement, the relevant AI-screening or AI-extraction use should be named there too — see How to Disclose AI Assistance in a CRediT Statement for the mechanics of writing that disclosure.
Practical guidance for using these tools in a review
- Treat AI screening/extraction as a second reviewer or sanity check, not a replacement for one. The evidence above supports using these tools to supplement dual independent human screening, not to reduce it to single-reviewer screening.
- Verify what the tool actually decided versus what it reordered or flagged. A relevance-ranking or automated-exclusion feature that a human confirms is a different level of AI involvement than a tool that performs the screening decision itself — disclose accordingly.
- Document verification steps, per the RAISE recommendations — not just that an AI tool was used, but how its outputs were checked.
- Check the specific tool’s current claims before citing them. Vendor-reported accuracy figures for actively developed commercial products change; the independent peer-reviewed studies cited above are a more stable reference point than a vendor’s own marketing page.
- Register the protocol prospectively where appropriate (e.g. via PROSPERO) before screening begins, regardless of which tools are used — AI-assisted screening does not change the case for prospective registration.
- Follow the target journal’s own AI-disclosure policy in addition to RAISE, since publisher requirements for generative-AI disclosure vary and are, per COPE and ICMJE guidance, generally more specific and more current than any general framework — see the publisher policy landscape guide.
Frequently asked questions
Can AI replace manual screening in a systematic review?
Not on current evidence. Independent studies of AI-assisted search and screening have found substantially lower sensitivity than traditional, librarian-designed systematic searches, and current joint guidance from Cochrane, Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence explicitly requires human oversight rather than treating AI as an unsupervised substitute for reviewer judgment.
Do I need to disclose AI use in a literature review or systematic review?
Under the RAISE recommendations endorsed by Cochrane and partner organizations (October 2025), any AI use that makes or suggests a judgement — screening decisions, data extraction, risk-of-bias or certainty-of-evidence assessments — requires disclosure of the system, its purpose, and its known limitations. Routine spelling/grammar assistance generally does not require this level of disclosure. Check the target journal’s own generative-AI policy as well, since publisher requirements vary.
What is the difference between AI-assisted screening tools like Rayyan or Covidence and a tool like Elicit?
Rayyan and Covidence layer AI features (relevance reordering, RCT classification, flagged automated exclusions) on top of a human-driven dual-screening workflow — the human reviewer still makes and records each screening decision. Elicit is built to perform screening and extraction steps directly and reports its own accuracy for having done so. Both categories require disclosure and human verification under current guidance, but the level and nature of human involvement differs by design.
Are these tools different from general AI research assistants like ChatGPT?
Yes. General-purpose AI chat tools are not built around a documented, protocol-driven screening and extraction workflow and are not the tools evaluated in the independent studies cited on this page. For guidance specific to broader AI research-assistant use (literature discovery, synthesis, and writing assistance more generally), see AI-Powered Research Assistant Tools and the AI Research Tool dictionary entry.







