Examples
Worked examples
- Is an instance
A forecaster on the Replication Markets platform trading on the probability that a specific published SBS claim would successfully replicate, as part of SCORE's crowdsourced forecasting arm.
- Is an instance
A research team's automated confidence-scoring algorithm generating a probability estimate for a claim's replicability directly from the published paper's text and metadata, without any human forecasting input.
Counter-examples
Looks similar, but isn't
- Not an instance
A single independent replication study of one paper -- that is a replication study in its own right, not the SCORE program, which specifically combined forecasting/scoring with a coordinated replication-testing framework across thousands of claims.
- Not an instance
A journal-level metric such as Impact Factor or CiteScore -- these measure citation influence, not the predicted or actual replicability of individual claims, which is what SCORE's confidence scores targeted.
Editorial commentary
SCORE (Systematizing Confidence in Open Research and Evidence) was a program funded by the U.S. Defense Advanced Research Projects Agency (DARPA), managed within its Defense Sciences Office. DARPA’s rationale was operational: the Department of Defense routinely draws on social and behavioral science (SBS) research to inform planning, investment, and modeling around topics like deterrence and stability, but the broader reproducibility crisis in the social sciences meant that not every published claim was equally trustworthy. SCORE set out to build automated tools that could assign an explainable ‘confidence score’ to a research claim — a quantitative estimate of how likely it was to hold up under independent replication — at a reliability comparable to or better than the best human expert judgment.
How the program worked
SCORE sampled roughly 3,000 empirical claims drawn from about 60 journals across multiple SBS disciplines, covering publications from 2009 through 2018. Those claims were scored through several parallel channels: individual expert surveys, structured expert panels, and crowdsourced forecasting/prediction markets — most visibly the Replication Markets platform, where forecasters traded on the probability that a given claim would replicate. Separately, research teams built automated algorithms intended to generate comparable confidence scores directly from text and metadata, without human forecasting. A subsample of the original claims was then handed to independent research teams to actually attempt replication, producing ground-truth outcomes that the forecasts, expert judgments, and algorithms could be checked against. The Center for Open Science was among the organizations involved in coordinating this replication and evaluation work.
What SCORE found
Published analyses from the program (see references below) reported that aggregated human forecasts — particularly market-based and survey-based forecasts pooled across many participants — tracked actual replication outcomes better than any single automated algorithm tested during the program, and that predicted replication rates varied noticeably across academic subfields within the social and behavioral sciences. These results fed the broader open-science literature on which SBS subfields are more or less replication-prone, and on how well crowd forecasting can substitute for costly direct replication at scale.
Why it matters for research administration
SCORE is a concrete, funded example of using prediction-market and forecasting methods as a scalable proxy for direct replication when replicating every claim in a literature is impractical. For research-integrity and research-assessment offices, it is a reference point for discussions of confidence scoring, claim-level credibility assessment, and how funders might one day triage which findings warrant costly direct verification versus which can be reasonably trusted based on aggregated forecaster judgment. It sits alongside pre-registration and other reforms catalogued under the reproducibility literature as part of the broader open-science toolkit, though SCORE is distinctive in applying that toolkit specifically to the social and behavioral sciences at DoD’s initiative rather than through a journal or funder mandate.
Current status
DARPA’s own program page now describes SCORE as complete and lists it for reference purposes only; it is not an active, ongoing initiative or an open call for proposals.
References
- DARPA, ‘SCORE: Systematizing Confidence in Open Research and Evidence’ program page (darpa.mil).
- Alipourfard et al., ‘Systematizing Confidence in Open Research and Evidence (SCORE)’ and related SCORE working-group outputs.
- ‘Are replication rates the same across academic fields? Community forecasts from the DARPA SCORE programme,’ Royal Society Open Science (2020).
- Center for Open Science, program partnership announcement (cos.io).
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="SCORE (Systematizing Confidence in Open Research and Evidence)"
vocab-term-identifier="https://casrai.org/dictionary/term/score-systematizing-confidence-in-open-research-and-evidence" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/score-systematizing-confidence-in-open-research-and-evidence",
"name": "SCORE (Systematizing Confidence in Open Research and Evidence)",
"identifier": "https://casrai.org/dictionary/term/score-systematizing-confidence-in-open-research-and-evidence",
"description": "A DARPA-funded research program, run by DARPA's Defense Sciences Office from 2019 to roughly 2022, that tested whether quantitative 'confidence scores' -- generated from human forecasting/prediction markets, expert surveys, and automated algorithms -- could reliably estimate how likely a published social and behavioral science (SBS) claim was to replicate. The program sampled around 3,000 research claims from roughly 60 journals published between 2009 and 2018, had forecasters and algorithms score them before any new data collection, then commissioned independent replications on a subsample to check which forecasting method came closest to the actual outcome. It is now complete; DARPA lists it as inactive and retained for reference.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/reproducibility#set",
"url": "https://casrai.org/dictionary/term/score-systematizing-confidence-in-open-research-and-evidence",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"dateModified": "2026-07-18T06:30:47",
"inLanguage": "en"
}






