Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us
Research Data Management (RDM)

Reproducibility, Synthetic & AI-Era Research Data

This sub-cluster covers two related concerns: ensuring that research results can be independently verified, and the newer questions about data integrity and provenance that generative AI has introduced. Reproducibility and replicability are distinct concepts, a distinction the National Academies of Sciences, Engineering, and Medicine (NASEM) set out clearly in its report Reproducibility and Replicability in Science: reproducibility means obtaining consistent results using the original data and methods, while replicability means obtaining consistent results in a new study addressing the same question. Computational reproducibility in particular depends on preserving not just data but the code, software environment, and workflow used to analyze it, which is why efforts like the ACM's Artifact Review and Badging framework have emerged to formally evaluate and label the availability, functionality, and reusability of research artifacts alongside published papers. Generative AI raises newer, less settled questions for this space: synthetic data generated by AI models is increasingly used to augment or substitute for real research data, raising questions about how its provenance should be documented and disclosed, while unresolved debates over text-and-data-mining copyright exceptions in the US and EU affect whether and how research data can legally be used to train AI models in the first place. For a research administrator, this sub-cluster is where established reproducibility practice meets an unsettled policy environment. Pages here cover reproducibility standards and badging schemes, what synthetic data disclosure currently looks like in practice, and the state of the text-and-data-mining exception debates as they bear on research data use.

Guides

Data Version Control (DVC) for Research Datasets and ML Pipelines

DVC is an open-source, Git-based tool for versioning large research datasets, model artifacts, and multi-step ML pipelines. This guide explains how it works, where it fits in the research data lifecycle, and why it isn’t a substitute for repository deposit.

HathiTrust Research Center: Non-Consumptive Text and Data Mining Access

HathiTrust Research Center (HTRC) provides non-consumptive computational access to the HathiTrust text corpus for text and data mining, distinct from the public HathiTrust library catalog.

Where to Publish Negative and Null Results: Journals and Venues That Accept Them

A practical guide to publishing negative and null results: which dedicated journals are still active, which soundness-based mega-journals accept them by policy, and structural alternatives like Registered Reports.

The Carpentries: Software Carpentry, Data Carpentry, and Library Carpentry Explained

A guide to The Carpentries nonprofit and its three lesson programs (Software Carpentry, Data Carpentry, Library Carpentry): the workshop model, instructor certification pipeline, institutional membership, and how this hands-on data-skills training fits into a research data management program.

How to Write a Registered Report: A Step-by-Step Guide to the Stage 1 Protocol

A practical, step-by-step guide to drafting a Registered Report Stage 1 protocol: writing testable hypotheses, specifying methods and analysis plans, justifying sample size, and avoiding common reviewer revision requests.

Metascience: The Study of Science Itself

Metascience is the empirical field that studies research methods, reporting, reproducibility, evaluation, and incentives across science as a whole. This guide covers how the field took shape, how it differs from the reproducibility crisis and replication studies, its core methods, and what its findings mean for research administrators.

The Replication Crisis: Origins, Causes, and the Open Science Response

The replication crisis is the field-spanning finding, dated from Ioannidis’s 2005 paper through the 2015 psychology Reproducibility Project and beyond, that a large share of published results fail to replicate. This guide covers its origins, causes, and the preregistration, Registered Reports, and data-sharing reforms that followed.

The Psychology Replication Crisis: The 2015 Reproducibility Project and What Changed

In 2015, a 270-author collaboration attempted to replicate 100 published psychology studies and found only 36% held up. This guide covers what the Reproducibility Project: Psychology actually found, why psychology was especially exposed, and the preregistration, Registered Report, and open-data reforms that followed.

The “Gold Standard Science” Executive Order (EO 14303) Explained

Executive Order 14303, “Restoring Gold Standard Science” (May 2025), sets nine criteria for trustworthy federal science and directs agencies to publish underlying data, code, and models. Here is what it requires, its timeline, and how it relates to existing research data management standards.

Replication in Psychology: Two Worked Examples

What a real psychology replication study looks like: the Reproducibility Project: Psychology and the facial feedback Registered Replication Report, explained with sources.

Synthetic Data in Research: Generation, Documentation, and Sharing Policy

How synthetic data is generated, documented, and validated in research settings, and what current GDPR, ICO, EHDS, and NIH guidance actually says about using it in place of real data for funder data-sharing mandates.

Reproducibility Infrastructure: Workflows, Containers, and Code Sharing

What computational reproducibility requires in practice: workflow managers (Snakemake, Nextflow, CWL), containers (Docker, Apptainer/Singularity), and durable code archiving (Zenodo, Software Heritage) with persistent identifiers.

AI Training Data Provenance, Copyright, and TDM Exceptions for Research

How EU, UK, and US copyright/TDM law applies to AI training in research, and how to document training-data provenance and licensing in your DMP.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →