Skip to main content
v2026.11,610 entries · CC-BY 4.0
CASRAIRegulatory RadarNever miss a regulatory change that affects your research officeA daily digest of new regulatory and compliance content, plus 150 questions/day to Ask CASRAI. Built for research administrators and compliance officers.See Regulatory Radar CASRAI · Own product

Consensus AI: What It Is and Its Real Limits

What Consensus AI (consensus.app) actually does, how the Consensus Meter works and where vote-counting across studies can mislead, verification steps, pricing, and how it compares to Elicit, SciSpace, AnswerThis, Undermind, Semantic Scholar and Scite.

Ask about Consensus AI: What It Is and Its Real Limits

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Consensus (consensus.app) is a named commercial AI-powered search engine for peer-reviewed research papers. Given a natural-language question, it retrieves papers from an indexed academic corpus and generates a synthesis of what those papers report, and for questions phrased as a testable claim it displays a Consensus Meter — a visual breakdown of how the retrieved papers agree, disagree, or split on that claim. It is one entry in CASRAI’s AI-powered research assistant tools guide, alongside comparable tools such as Elicit, SciSpace, AnswerThis, and Undermind. This guide covers what Consensus actually does, how the Consensus Meter works and where it can mislead, how to verify its output, what it is and is not suited for, published pricing, and how it compares to the alternatives researchers most often weigh it against.

What Consensus is, and what it isn’t

Consensus is not a generic descriptor for “AI that searches papers” — it is a specific product built by a company also called Consensus. It sits between three different tools researchers already know, and understanding those boundaries is most of what you need to use it well:

  • Versus keyword search (a library database, PubMed, Google Scholar): keyword search returns a ranked list of documents matching your terms; you still have to read each one and synthesize the answer yourself. Consensus adds a synthesis layer on top of retrieval — it reads the retrieved papers and generates a written summary of what they say, with each point in that summary linked back to a specific source paper.
  • Versus a general-purpose AI chatbot: a chatbot with no dedicated retrieval step against an indexed academic corpus can be asked to “summarize the research,” but its answer is not reliably grounded in retrievable, checkable sources, and carries a materially higher risk of fabricated or unverifiable citations. Consensus’s core function is running the retrieval step first, against a defined corpus of papers, and then synthesizing only across what it actually retrieved.
  • Versus a systematic review: this is the comparison that matters most and the one vendor marketing is least likely to volunteer — see the Consensus Meter section below.

What corpus it actually searches

Consensus’s search has historically been built in partnership with Semantic Scholar, the Allen Institute for AI’s academic search and citation-graph service, and the companies announced that partnership publicly at Consensus’s launch. Third-party trackers and Consensus’s own marketing describe the indexed corpus as spanning well over 200 million peer-reviewed papers, conference papers and preprints across disciplines, drawn from Semantic Scholar’s index together with other aggregated sources and publisher partnerships. Treat the exact figure as a moving, vendor-reported number rather than an independently audited one — corpora of this kind grow continuously and vendors do not publish audited counts. The practical point for a researcher is different from the exact number: Consensus is not searching a single curated collection the way a subject-specific systematic-review protocol would specify a fixed set of databases (e.g. MEDLINE, Embase, CENTRAL) with a documented search date — it is querying a broad, general aggregator, which is a different and less reproducible evidentiary basis.

The Consensus Meter — what it aggregates, and its real epistemic limits

This is the part of the product most worth being plain-spoken about, because Consensus’s own marketing has no incentive to be. For a question phrased as a testable yes/no/maybe claim (for example, whether a specific intervention affects a specific outcome), Consensus can display a Consensus Meter: a bar showing how many of the papers it retrieved and summarized for that specific query support the claim, oppose it, or report mixed results.

What the meter is actually counting is votes across a heterogeneous, retrieval-dependent sample of studies — not a weighted synthesis of evidence quality. That distinction matters more than it sounds:

  • It does not weight by study design. A well-powered randomized controlled trial and a small observational study or an animal model each count as one paper toward the bar, with no explicit weighting for where each sits in a standard evidence hierarchy.
  • It does not weight by sample size or statistical power. A study with an n of 20 and a study with an n of 20,000 can each contribute one “supports” or “opposes” vote.
  • It does not apply a risk-of-bias assessment. Tools like the Cochrane risk-of-bias tool exist specifically because study quality varies enormously even among published, peer-reviewed papers on the same question; the meter has no equivalent appraisal step.
  • It is sensitive to how the query is retrieved and phrased. The meter reflects agreement only within the specific set of papers Consensus’s retrieval step happened to surface for that exact query on that day — a differently worded question, or a query run a month later against an updated index, can surface a different paper set and a different bar.
  • It does not report effect sizes. A bar full of “supports” tells you direction of findings, not magnitude, clinical/practical significance, or confidence intervals — a claim can be technically “supported” by five studies each reporting a trivially small effect.

In short: vote-counting across heterogeneous studies is not evidence synthesis. A systematic review under a framework like PRISMA 2020 exists precisely to control for these problems — a documented, reproducible search strategy against fixed databases, explicit inclusion/exclusion criteria, formal risk-of-bias appraisal of every included study, and (in a meta-analysis) a statistically weighted pooled effect estimate, not a raw count of papers on each side. A Consensus Meter bar that looks decisively one-sided can be resting on a handful of small, low-quality, or methodologically weak studies, and there is no way to tell that from the bar itself — you have to open the underlying papers and read them. Use the meter as a fast orientation signal for scoping a question, never as a stand-in for critical appraisal or a substitute for the judgment a systematic review is built to provide.

How to verify what Consensus tells you

Every claim in a Consensus synthesis and every paper counted in a Consensus Meter is meant to be traceable to a specific source. Before citing or acting on anything Consensus produces:

  1. Open the underlying paper, not just the AI-generated snippet. Confirm the paper actually says what the synthesis claims it says — AI summarization of scientific findings can flatten nuance, drop qualifiers (“in this specific population,” “under these conditions”), or misstate direction of effect.
  2. Check the study design and sample size yourself. Since the meter doesn’t surface this, you have to look: is this a randomized trial, an observational study, a case report, an animal or in-vitro study, a preprint that hasn’t been peer reviewed?
  3. Check publication status and venue. A preprint and a paper published in a high-quality peer-reviewed journal after review should not carry equal weight in your own judgment, even though they may both show up as a single “supports” vote.
  4. Re-run the query with different phrasing. If the meter shifts meaningfully with how the question is worded, that instability is itself informative — it tells you the underlying evidence base is thin or the claim is not cleanly testable with a single retrieval query.
  5. Cite the original paper, not the Consensus synthesis. In a manuscript, grant proposal, or evidence brief, the citation belongs to the peer-reviewed source you verified, not to Consensus’s summary of it.

What Consensus is genuinely good for

  • Fast orientation on an unfamiliar question. Getting a first read on whether a claim has any published support at all, and a starting set of papers to read, before committing time to a deeper search.
  • Sanity-checking a specific claim. Quickly checking whether a statement in a draft, a proposal background section, or someone else’s argument has any peer-reviewed backing, as a first pass before deeper verification.
  • Scoping a topic before designing a formal search. Seeing roughly how much literature exists and which terms and subtopics come up, to inform a more rigorous search strategy later.
  • Non-systematic narrative writing. Background sections, blog-style science communication, or grant narrative where a fast, source-linked orientation is useful and the writing does not claim to be an exhaustive or reproducible synthesis.

What it is genuinely bad for

  • Anything protocol-driven. A registered systematic review, a Cochrane review, a health-technology assessment, or any output that requires a documented, reproducible, pre-specified search strategy against named databases. Consensus’s retrieval is not reproducible in that sense and does not log a fixed search date and query against a defined database set the way review protocols require.
  • Clinical, regulatory, or high-stakes decisions. The Consensus Meter’s vote-count is not risk-of-bias-adjusted evidence and should never substitute for guideline-level or clinician-appraised evidence.
  • Meta-analysis or effect-size questions. “Does X work” and “how large is X’s effect, and how confident are we” are different questions; Consensus answers something closer to the first, imperfectly, and does not answer the second at all.
  • Exhaustive or auditable literature coverage. Because the retrieved paper set is not fixed or fully disclosed, two people running the same query cannot be guaranteed to see the same result set indefinitely, which is disqualifying for any use case that requires an auditable search trail.

Why it does not replace a PRISMA-registered search

A systematic review conducted under PRISMA 2020 reporting standards documents an explicit, reproducible search strategy across named databases, a defined search date, explicit inclusion/exclusion criteria applied by (typically) two independent reviewers, formal risk-of-bias assessment of every included study, and — where a meta-analysis is warranted — a statistically pooled, weighted effect estimate with a stated confidence interval. Consensus’s retrieval-and-synthesis process is not designed to produce any of that: the corpus queried is a broad aggregator rather than a fixed set of named databases, the search is not logged as a reproducible protocol, inclusion is not governed by explicit criteria applied consistently by independent reviewers, and there is no risk-of-bias appraisal step. None of that makes Consensus a bad tool — it makes it a different one, built for fast orientation rather than for the auditability a formal evidence synthesis requires. Treat a Consensus Meter result as a hypothesis to check, never as a completed synthesis.

Pricing

Consensus publishes pricing on consensus.app, structured as a free tier with limited monthly AI-generated syntheses and Consensus Meter uses, plus a paid individual “Premium” plan that unlocks unlimited Pro Analyses and removes usage caps, and separate Team/Enterprise pricing for institutional or lab-wide access (including university-wide licensing in some cases). Independent pricing trackers have reported the individual Premium plan in the roughly $8-10 per month range on an annual commitment, with materially higher, custom-quoted rates for team and enterprise tiers — these figures are third-party-reported rather than confirmed directly from Consensus’s own pricing page at the time of writing, and AI-tool pricing changes frequently, so confirm the current plan structure and price directly on consensus.app/pricing before budgeting or purchasing.

How Consensus compares to the alternatives

Researchers most often weigh Consensus against these tools. All are AI-assisted, but they solve different problems.

Tool Core approach Where it’s strongest
Consensus Retrieval + claim-level synthesis with the Consensus Meter agreement bar Fast orientation on whether a specific claim has published support
Elicit Structured evidence extraction into a comparable table (methods, outcomes, sample size) across a paper set, plus systematic-review-screening support Building a structured, exportable evidence table; supporting (not replacing) the screening stage of a real review
SciSpace Paper-level AI reading assistant — chat with a specific PDF, plus a literature-review/Discovery mode for broader search Deep, close reading and question-answering against one paper or a small set you’ve already found
AnswerThis Research-gap identification — surfaces what a body of literature has not yet answered Early-stage proposal and thesis scoping, finding an unaddressed angle
Undermind Iterative “deep search” agent that reasons across multiple search rounds before returning a curated set Broad, exhaustive-feeling discovery for a complex, multi-faceted question
Semantic Scholar Free scholarly search engine and citation graph, no AI-generated synthesis by default Straightforward discovery, citation tracking, and as the underlying index several AI tools (including Consensus) build on
Scite “Smart Citations” — classifies each citing paper as supporting, contrasting, or mentioning the cited claim, based on citation context rather than a synthesized meter Checking how a specific published claim or paper has actually been treated by the papers that later cited it

The distinction worth holding onto: Elicit, SciSpace and Undermind are built to help you assemble and read a paper set; Scite is built to show you how a specific claim has been cited afterward; Consensus is built to give you a fast, claim-level directional read across whatever it retrieves. None of the seven substitutes for a registered systematic review, and none of their vendors will tell you that as directly as this page does.

Frequently Asked Questions

Is Consensus AI accurate?

Consensus links every claim in its synthesis back to a specific retrieved paper, which makes it checkable — but “checkable” is not the same as “verified.” The accuracy of what it retrieves and how it characterizes each paper still needs to be confirmed against the source, and the Consensus Meter’s agreement bar should never be read as a validated accuracy signal about the underlying science; it only reflects the paper set retrieved for that specific query.

What is the Consensus Meter, exactly?

A visual bar, shown for questions phrased as a testable yes/no/maybe claim, indicating how many of the retrieved and summarized papers support, oppose, or report mixed findings on that claim. It is a vote count across a retrieval-dependent paper set, not a study-quality-weighted synthesis — it does not account for study design, sample size, risk of bias, or effect size.

Can I use Consensus for a systematic review or meta-analysis?

No. A registered systematic review requires a documented, reproducible search strategy across named databases, explicit inclusion/exclusion criteria, independent dual screening, formal risk-of-bias appraisal, and (for a meta-analysis) a statistically weighted pooled effect estimate. Consensus’s retrieval is not reproducible or auditable in that sense, and its meter performs none of the quality-weighting a review requires. Use it to scope a question before you design a review, not to conduct one.

Is Consensus free?

Consensus offers a free tier with limited usage, alongside a paid individual plan and separate team/institutional pricing. Confirm current plan limits and pricing directly at consensus.app, since AI-tool pricing structures change frequently.

What database does Consensus search?

Consensus’s search has been built in partnership with Semantic Scholar and is reported to span well over 200 million papers aggregated across publishers, preprint servers and conference proceedings. It is a broad, continuously updated aggregator, not a fixed, named set of databases with a documented search date the way a systematic review protocol specifies.

How is Consensus different from ChatGPT or another general AI chatbot?

A general chatbot with no dedicated academic-retrieval step can be asked to summarize research, but its answer is not reliably grounded in an indexed corpus of actual papers and carries a materially higher risk of fabricated citations. Consensus runs a retrieval step against a defined academic corpus first, then synthesizes only across what it actually retrieved, with each point traceable to a source paper.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →