Skip to main content
v2026.11,772 entries · CC-BY 4.0

Title and Abstract Screening: Workflow, Pilot Calibration and Disagreement Rules

How review teams screen titles and abstracts against pre-specified criteria: piloting and calibrating the criteria first, when dual independent screening is required versus a faster liberal-accelerated pass, and the Cochrane decision rules for resolving reviewer disagreement.

Ask CASRAI · included with Regulatory Radar

Ask about Title and Abstract Screening: Workflow, Pilot Calibration and Disagreement Rules

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Written and maintained by CASRAI Editorial Board

Last updated

Title and abstract screening is the first eligibility gate a systematic review applies after deduplicating its search results, and it is also where the most reviewer-hours and the most disagreement happen — long before anyone opens a full-text PDF. PRISMA 2020 frames it as its own labeled phase of the review, distinct from both the initial identification of records and the later eligibility check on full texts. This guide covers the actual mechanics review teams need in a protocol or methods section: how the pass fits into the PRISMA 2020 flow, how to pilot and calibrate screening criteria before splitting the workload, when full independent dual screening is expected versus when a faster liberal-accelerated approach is defensible, and the decision rules Cochrane methodology uses to resolve reviewer disagreement.

Where screening sits in the PRISMA 2020 flow

PRISMA 2020 restructured the older four-phase (2009) flow diagram — Identification, Screening, Eligibility, Included — into three labeled phases: Identification, Screening, and Included. The old standalone Eligibility phase was not dropped as a concept; it was folded into the Screening phase as explicit “reports assessed for eligibility” and “reports excluded, with reasons” boxes (Page MJ, McKenzie JE, Bossuyt PM, et al., “The PRISMA 2020 statement: an updated guideline for reporting systematic reviews,” BMJ 2021;372:n71). In practice this means the single word “screening” on a modern PRISMA diagram actually covers two distinct passes with different rigor: a fast title/abstract pass against broad eligibility criteria, followed by a slower full-text pass where each remaining report is checked in detail and every exclusion is logged with a specific reason. This guide is about the first of those two passes.

Checklist item 16a of PRISMA 2020 asks authors to describe their search-and-selection results “ideally using a flow diagram” — free tooling for this (the PRISMA2020 R package and its companion Shiny app) is covered in CASRAI’s PRISMA flow diagram entry and the broader PRISMA methodology guide.

Screening against pre-specified criteria, not a moving target

Title and abstract screening only works cleanly if the eligibility criteria it’s applied against were fixed before screening started — in the protocol, not invented mid-screen. That means the PICO-framed or PICOT question and the inclusion/exclusion criteria in a registered protocol (ideally one registered on PROSPERO before screening begins) are what a screener actually applies at this stage — not a fresh judgment call per record. The standard convention, taught consistently across systematic-review methods guidance, is to err toward inclusion when a title or abstract doesn’t give enough information to decide: a record that might plausibly meet the criteria goes forward to full-text review rather than being excluded on a guess, since a false exclusion at this stage is unrecoverable in a way a false inclusion (caught later, at full text) is not.

Pilot testing and calibrating the criteria before the full screen

Before splitting the full record set across reviewers, review teams commonly screen a shared pilot batch independently first — a random sample of the search results, often cited in methods guidance as somewhere in the range of 50–100 records or roughly 5–10% of the total, though there’s no single mandated figure and the right size scales with how large and how ambiguous the record set is. The point of the pilot batch isn’t to screen faster; it’s to find out, before the workload is split, where two reviewers applying the same written criteria actually disagree. Genuine disagreement at this stage almost always traces back to a criterion that reads clearly to whoever wrote it but is ambiguous in practice — a vague population definition, an outcome that’s described differently across the literature than in the protocol’s own wording. The team discusses the disagreements, tightens the criteria’s wording, and only then proceeds to screen the remaining records under the calibrated version. Skipping this step doesn’t just slow down the eventual disagreement-resolution process — it means the ambiguity gets discovered piecemeal, one disagreement at a time, well after both reviewers have already ground through most of the list.

Dual independent screening — and when a faster, liberal-accelerated pass is used instead

The default expectation for a full systematic review is independent dual screening: two reviewers each screen every title and abstract without seeing the other’s decisions, and their results are compared afterward. This mirrors the same two-independent-reviewers principle Cochrane methodology applies to data extraction — Cochrane’s Methodological Expectations of Cochrane Intervention Reviews (MECIR) standard C46 requires “(at least) two people working independently” for outcome data extraction specifically, and review teams commonly extend the identical logic to the screening stage as standard practice, even though the Cochrane Handbook documents the independent-extraction rule and the screening rule in separate sections.

Rapid reviews are the deliberate exception. The Cochrane Rapid Reviews Methods Group (RRMG) publishes interim guidance (Garritty C, et al., “Cochrane Rapid Reviews Methods Group offers evidence-informed guidance to conduct rapid reviews,” Journal of Clinical Epidemiology, 2021, updated 2024) specifically aimed at streamlining or omitting systematic-review steps — screening among them — to shorten timelines for time-sensitive questions. Under a commonly used streamlined pattern in this literature, sometimes called liberal-accelerated screening, a single reviewer screens the full record set, and a second reviewer checks only the records the first reviewer excluded, rather than both reviewers independently screening every record. [REPORTED — the exact term and mechanics are widely used in the rapid-review methods literature, but this guide could not independently pull the precise defining passage from RRMG’s own published guidance in this drafting pass; treat the label as descriptive of the common practice, not a verbatim RRMG definition.] The trade-off is explicit and should be named in the methods section of any review that uses it: liberal-accelerated screening cuts reviewer-hours roughly in half compared to full dual screening, at the cost of a real (if generally judged acceptable for time-sensitive rapid reviews) risk that the single first-pass reviewer misses an eligible record the second reviewer never sees, since the second reviewer never checks the included set. It is not considered an adequate substitute for full dual screening in a standard systematic review, only in reviews that have explicitly adopted a rapid-review design. See CASRAI’s rapid review vs. systematic review guide for how this fits the broader set of steps rapid reviews streamline.

Resolving disagreements between screeners

Whichever screening model a review uses, the decision rule for handling disagreement needs to be written into the protocol in advance, not improvised once a conflict shows up. The Cochrane Handbook’s guidance on this (Chapter 5, section 5.5.5) states it plainly: “An explicit procedure or decision rule should be specified in the protocol for identifying and resolving disagreements. Most often, the source of the disagreement is an error by one of the extractors and is easily resolved. Thus, discussion among the authors is a sensible first step. More rarely, a disagreement may require arbitration by another person.” That guidance is written for data-extraction disagreements specifically, but review teams apply the identical two-step rule at the screening stage as a matter of established convention: reviewers first meet and discuss each screening conflict (most turn out to be a simple misread of a criterion or an abstract, resolved in minutes), and only the small residue of genuine, considered disagreement goes to a third reviewer or a named arbitrator for a final call. Naming who that third person is, in the protocol, before screening starts, avoids the conflict of interest of one of the two original screeners effectively getting the tie-break.

Measuring screener agreement with kappa

The same Cochrane Handbook section notes that “agreement of coded items before reaching consensus can be quantified, for example using kappa statistics (Orwin, 1994), although this is not routinely done in Cochrane reviews.” Reporting a kappa statistic from the pilot-calibration batch (see above) is common in published methods sections precisely because it gives a reviewer-agreement number a reader can sanity-check, even though Cochrane itself doesn’t mandate it. The standard interpretation bands for Cohen’s kappa, from Landis JR & Koch GG’s original 1977 scale, are widely taught as a rough guide rather than a strict cutoff: below 0 poor, 0–0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, and 0.81–1.00 almost perfect agreement. A pilot batch landing below “substantial” agreement is the practical trigger for another round of criteria calibration before the full screen proceeds, rather than a threshold to simply note and move past. CASRAI’s inter-rater reliability guide covers how to choose between kappa, weighted kappa, and other coefficients when screening involves more than a simple include/exclude decision.

What gets logged, and what PRISMA actually asks for at this stage

A common misconception is that PRISMA 2020 expects a specific, itemized reason for every single record excluded at title/abstract screening, the same way it does at the later full-text eligibility stage. It doesn’t: the screening phase of the flow diagram reports an aggregate count of records excluded, while the granular “excluded, with reasons” breakdown belongs to the subsequent full-text eligibility phase, where each remaining report gets an individual, citable reason for exclusion. Screening logs are still worth keeping at the record level even so — most screening software (see below) records each reviewer’s include/exclude/unsure decision per record automatically, which is what makes the pilot-calibration kappa calculation and any later disagreement audit possible in the first place.

Tools that support calibration and dual review

Purpose-built screening software exists specifically to make pilot calibration and disagreement tracking mechanical rather than manual. ASReview and DistillerSR both support independent dual screening with built-in conflict views; Covidence, Rayyan, and DistillerSR and the broader AI-assisted screening tool comparison cover how these platforms differ on conflict-resolution workflow and reporting. CASRAI’s AI tools for systematic literature review guide covers where machine-learning-assisted prioritization (as in ASReview’s active-learning ranking) fits alongside, not instead of, human dual screening.

Frequently asked questions

How many reviewers should screen titles and abstracts?

Two, working independently, is the standard expectation for a full systematic review, mirroring the same two-independent-reviewers principle Cochrane methodology applies to data extraction. Rapid reviews sometimes substitute a faster liberal-accelerated pattern (see above) as an explicit, named trade-off, not a silent shortcut.

What is liberal-accelerated screening?

A streamlined screening pattern used in some rapid reviews, in which one reviewer screens the full record set and a second reviewer checks only the records excluded by the first reviewer, rather than both reviewers independently screening everything. It roughly halves reviewer-hours at the cost of the second reviewer never double-checking the included set.

Do I need a reason logged for every record excluded during title/abstract screening?

No — PRISMA 2020 expects an aggregate exclusion count at the screening phase of the flow diagram. Itemized, individually cited exclusion reasons are expected at the subsequent full-text eligibility phase, not the title/abstract pass.

How is screener agreement measured during pilot calibration?

Commonly with a kappa statistic calculated on a shared pilot batch screened independently by both reviewers before the full set is split between them. The Cochrane Handbook notes this quantification isn’t mandatory but is a recognized way to check agreement before proceeding.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.