A data management plan for a clinical trial is not the same document as a general research-data DMP, even though both are called “DMPs” and both cover storage, access, and retention. A funder-facing DMP (the kind covered in CASRAI’s NIH Data Management and Sharing Plan guide) is written to satisfy a grant sponsor that research data will be findable, shareable, and preserved after a project ends. A clinical trial DMP is written to satisfy a regulator, an IRB, and a trial sponsor that the data generated inside a specific electronic data capture (EDC) system, over the life of a specific protocol, will be complete, attributable, and audit-ready before that dataset is ever locked for analysis. The two documents share a name and almost nothing else in practice.
This guide covers the operational content a clinical trial DMP actually needs: how it specifies the EDC/CTMS technology stack, how it documents the case report form (CRF) data flow from site entry to locked dataset, what a database lock procedure looks like, how source data verification (SDV) is planned and recorded, and what ICH Good Clinical Practice (GCP) requires for data retention. For the systems and roles referenced throughout, see CASRAI’s guides on Clinical Trial Management Systems (CTMS) and clinical data management, and the dictionary entry for Electronic Data Capture (EDC).
Why the general research DMP template falls short
A general DMP — the kind funders like NIH or NSF expect — is organized around data types, formats, metadata standards, and a post-project sharing and preservation commitment. It typically says relatively little about who may edit a record, how an edit is authenticated, or what happens the moment data collection formally stops. A clinical trial DMP has to say all of that, in operational detail, because the data it governs is regulated: it will support (or undermine) a marketing application, it is subject to inspection by FDA, EMA, or another national regulator, and it must meet the ALCOA+ data-integrity attributes (Attributable, Legible, Contemporaneous, Original, Accurate, plus Complete, Consistent, Enduring, and Available) that regulators and auditors use to judge whether a record can be trusted at all.
In practice this means a clinical trial DMP needs four things a generic research DMP template does not typically ask for: a named, validated EDC/CTMS system specification; a documented CRF-to-database data flow with defined query and edit-check logic; a written database lock and unlock procedure; and a source data verification plan tied to a specific monitoring approach. Each is covered below.
EDC system specification: what the DMP needs to name
Where a general DMP might say “data will be stored in an institutional repository,” a clinical trial DMP has to specify the actual system of record and its regulatory posture. At minimum, that section should identify:
- The EDC platform and version in use for the trial (e.g., a specific vendor build), and whether it is validated for its intended use — computer system validation is expected, not optional, for any system supporting a regulated submission.
- 21 CFR Part 11 conformance — FDA’s 1997 rule governing electronic records and electronic signatures, which sets the audit-trail, access-control, and e-signature requirements a system must meet for its records to be considered equivalent to paper records and handwritten signatures in a regulated submission. The DMP should state how the EDC system meets these requirements (audit trail capture, unique user accounts, e-signature manifestation) rather than simply asserting compliance.
- User roles and access control — who can enter data, who can issue and answer queries, who can view but not edit, and who holds the elevated permissions needed to unlock a locked database. Role definitions should map to real job functions (site coordinator, data manager, monitor, CTMS administrator), not generic “admin/user” tiers.
- System validation and change-control documentation — evidence the EDC build was tested against its specification before go-live, and a process for documenting any mid-study configuration change (a new edit check, a form amendment) without breaking the audit trail on already-collected data.
Note the distinction between the three systems a clinical trial DMP typically has to reference: EDC is the site-facing data-entry layer; a Clinical Data Management System (CDMS) is the broader term covering EDC plus query management, coding, edit checks, and audit trail; and a CTMS is a separate system tracking trial operations — site activation, subject visit scheduling, regulatory document tracking, and enrollment — that is often integrated with, but is not the same system as, the EDC/CDMS. A DMP that conflates these three, or fails to specify which system does what, leaves a real gap an inspector or a study team will find later.
CRF data flow: from case report form to locked dataset
The DMP should document the data flow in the order it actually happens, because each stage has its own controls and its own owner:
- CRF design — what data is collected, in what format, at what visit, mapped to the protocol and statistical analysis plan.
- Database build — configuring the EDC system with the CRF structure, edit checks, and skip logic, then validating the build against the design before the site is opened for enrollment.
- Data entry and edit checks — site staff enter data (often within a protocol-specified window of the visit); automated range, consistency, and logic checks flag implausible or incomplete entries in near-real time.
- Query management — the data manager issues queries to sites when data looks wrong or is missing; the DMP should specify expected query turnaround, who can close a query, and how query history is preserved in the audit trail.
- External data reconciliation — merging and cross-checking data from sources outside the EDC: central laboratories, ECG cores, imaging vendors, wearables, and interactive response technology (IRT) used for randomization and drug supply. The DMP should name each external source and how its data reaches the trial database.
- Medical coding — mapping free-text adverse events, medical history, and concomitant medications to standardized dictionaries (e.g., MedDRA, WHO Drug) so terms are analyzable and comparable across sites.
A DMP that only says “data will be collected via EDC and cleaned before analysis” has not actually specified this flow. Naming the CRF-to-database path, the query SLA, and the reconciliation sources is what turns a generic statement into an operational plan a monitor or auditor can check against.
Database lock: what the DMP needs to specify
Database lock is the formal, documented point at which the dataset is frozen — no further edits without a controlled unlock process — once cleaning is complete, so statistical analysis can begin against a fixed dataset. A clinical trial DMP should describe the lock procedure itself, not just state that a lock will occur. At minimum that means specifying:
- The lock readiness criteria — all outstanding queries resolved or formally closed, all serious adverse events reconciled between the safety database and the EDC, all external data (labs, imaging, IRT) reconciled and loaded, and coding finalized.
- Who signs off on lock readiness (typically the data manager, biostatistician, and sponsor/study lead jointly) and how that sign-off is documented.
- The unlock process for the rare case where a locked database must be reopened — who can authorize it, what triggers it (e.g., a post-lock data-integrity finding), and how the change is captured in the audit trail so the lock’s integrity as a historical record isn’t undermined.
- Whether the trial uses a soft lock (an interim freeze, e.g., ahead of a planned interim analysis or a Data Safety Monitoring Board review) distinct from the final hard lock, and what access restrictions apply during each.
Source data verification (SDV) and data integrity
Source data verification is the process of checking that data entered into the EDC accurately reflects the original source record (the patient chart, lab report, or other source document) at the site. A clinical trial DMP should state the SDV approach — whether it is 100% SDV, risk-based/targeted SDV focused on critical data points (primary endpoint, informed consent, eligibility, serious adverse events), or a hybrid — and who performs it (typically a clinical research associate/monitor during site visits, or via centralized/remote monitoring for decentralized elements).
Underlying all of this is the ALCOA+ data-integrity framework — Attributable, Legible, Contemporaneous, Original, Accurate, plus Complete, Consistent, Enduring, and Available — which originated from FDA’s original five ALCOA attributes and was extended by the UK MHRA’s 2018 “GXP” Data Integrity Guidance. A clinical trial DMP that references ALCOA+ explicitly, and ties its SDV/monitoring plan to it, is doing real integrity planning rather than a general storage-and-backup statement. This connects directly to the essential documents and source-data requirements set out in ICH E6 (see below), and to CASRAI’s dictionary entries on the data steward role and DMP compliance check concepts as they apply generally.
Retention: what ICH E6 and FDA regulation actually require
Retention is where a clinical trial DMP diverges most sharply from a funder DMP’s open-ended “preserve for the long term” language. ICH E6 Good Clinical Practice defines essential documents as those which, individually and collectively, permit evaluation of the conduct of a trial and the quality of the data produced — grouped by trial stage (before the clinical phase begins, during conduct, and after completion) and collectively constituting the Trial Master File (TMF) that a sponsor’s audit function and a regulatory inspector both review. Under FDA’s IND recordkeeping regulations (21 CFR 312.62), an investigator must retain trial records for at least two years after the marketing application is approved for the studied use, or, if no application is approved, until two years after the investigation is discontinued and FDA is notified — whichever applies. Sponsor and institutional retention obligations, and retention periods under other national regulators (EMA, MHRA), can run longer and are frequently set by the specific trial agreement or institutional SOP rather than by the DMP author’s own preference.
A clinical trial DMP should therefore state a concrete retention trigger and duration (tied to marketing-application status or trial discontinuation, not just a fixed calendar date), name who is responsible for retention at each site/sponsor/institution, and describe where the locked dataset and its supporting audit trail will be archived once the trial closes out — distinct from, and typically longer than, the data-sharing timelines a funder DMP specifies for de-identified research data. ICH E6(R3), the current core GCP guideline revision (finalized/adopted 6 January 2025), restates and reorganizes these principles under a risk-proportionate, quality-by-design framework without relaxing the underlying retention obligation.
Who owns the clinical trial DMP
Because a clinical trial DMP spans regulatory, operational, and technical territory, no single role owns every section. In practice: the clinical data manager typically owns the CRF design, database build, edit-check, and query-management content; the CTMS administrator or clinical operations lead owns the site/enrollment tracking content where it intersects with data flow; the sponsor’s quality or regulatory affairs function owns the 21 CFR Part 11/validation and retention content; and the principal investigator and institution’s research office are accountable for the site-level source data and essential-document retention obligations. A DMP that names an owner for each section is more likely to survive staff turnover and audit than one written by a single author and filed away.
Frequently asked questions
Is a clinical trial DMP the same as a Trial Master File (TMF)?
No. The TMF is the collection of essential documents (protocol, approvals, monitoring reports, and more) that demonstrates a trial was conducted to standard, per ICH E6 Section 8. The DMP is a planning document that describes how trial data specifically — the EDC/CTMS systems, CRF flow, lock procedure, and retention plan — will be managed. A well-run trial cross-references the two: the DMP itself is typically filed as an essential document within the TMF.
Does every clinical trial need a separate DMP from the funder’s general research DMP?
Where the trial is federally funded, the awarding agency’s general DMP requirement (for example, NIH’s Data Management and Sharing Policy) still applies and typically governs post-trial data sharing and preservation. But that funder-facing DMP does not substitute for the operational, EDC/CTMS-specific data management documentation a trial needs to run — and that a sponsor, IRB, or regulator will expect to see. Most trial teams maintain both, cross-referenced, rather than treating one as satisfying the other.
Who writes the EDC-specific sections of a clinical trial DMP?
Typically the clinical data manager or data management lead, in coordination with the EDC vendor or in-house CTMS/EDC administrator for system-validation and access-control details, and with the sponsor’s quality/regulatory function for 21 CFR Part 11 and retention language. See CASRAI’s guide on clinical data management for the broader role and workflow this sits inside.
What happens to trial data after database lock?
Once locked, the dataset moves to the biostatistics team for analysis against the pre-specified statistical analysis plan, and is transformed into submission-ready tabulation and analysis formats (see CDISC standards, covered in CASRAI’s clinical data management guide) if the trial supports a regulatory submission. The DMP’s retention plan governs what happens to the underlying database, audit trail, and source-data-verification records after that point.







