A data trust is a governance structure, not a piece of software or a single legal document: an independent trustee (an organisation or body) takes on fiduciary-style duties toward a defined group of data subjects or contributors and makes decisions about how a body of data is accessed, used, and shared on their behalf. The Open Data Institute (ODI) put forward the now-standard working definition in 2018: a data trust is ‘a legal structure that provides independent stewardship of data.’ For research administrators evaluating options for sharing sensitive data – health records, genomic data, community-contributed data, or other data carrying real risk to the people it describes – the data trust model is worth understanding as a distinct point on the governance spectrum, separate from a data-sharing agreement, a data governance policy, or a data commons.
What makes a data trust different
The defining feature of a data trust is the interposition of an independent trustee between data contributors/subjects and data users, with the trustee bound to act in the interests of the people the data is about (or who contributed it) rather than in the interests of whichever party currently wants access. That fiduciary orientation – a duty of care modelled loosely on financial or land trusts – is what separates a data trust from other governance mechanisms that also involve rules about who can use data and how:
- Data trust vs. data sharing agreement: a data sharing agreement (DSA) is a bilateral or multilateral contract between specific parties for a specific transfer or use of data. It doesn’t create an ongoing institution or a party with standing duties to the data subjects themselves. A data trust is a persistent governance structure that can adjudicate many access requests, from many requesters, over time, against a consistent set of duties – a DSA is typically a one-off (or renewable) agreement for a defined purpose.
- Data trust vs. data governance policy: an institutional data governance policy sets internal rules, roles, and accountability for data an organisation already controls (see CASRAI’s Research Data Governance guide for the general framework). A data trust, by contrast, typically sits outside any single institution and is designed precisely for cases where no single institution should unilaterally control access – for example, data pooled from multiple contributing organisations or communities.
- Data trust vs. data commons: a data commons pools data for relatively open access under a shared set of norms, without necessarily assigning an accountable trustee with fiduciary duties. A data trust adds an accountable, legally responsible intermediary; a commons emphasises openness and shared stewardship more than fiduciary accountability to specific data subjects.
- Data trust vs. data hub: a data hub is a technical aggregation point – infrastructure for storing and moving data. It says nothing on its own about governance. A data trust can use a data hub as its technical substrate, but the hub is not itself a governance model.
- Data trust vs. trusted research environment (TRE): a TRE is a secure computational environment where approved researchers can analyse sensitive data without the data leaving a controlled setting (the five safes model). A TRE answers the question of how access is technically and physically controlled once granted. A data trust answers a different, prior question: who decides whether access should be granted at all, and on whose behalf. The two are complementary – a data trust’s trustee could designate a TRE as the required access mechanism for approved projects. The UK’s Goldacre Review, which shaped much of the current TRE policy landscape for health data, is explicit that TRE architecture is about safe access controls, not about resolving who holds fiduciary authority over the data – that is the gap the data trust concept targets.
Where the concept came from
The data trust idea entered UK policy discussion through the 2017 Hall-Pesenti Review, Growing the Artificial Intelligence Industry in the UK, commissioned by the UK government, which proposed data trusts as a mechanism to increase trusted data sharing to support AI development while protecting the interests of the people data is about. The Open Data Institute took up the concept and, working with the UK’s then Office for Artificial Intelligence and Innovate UK, ran three exploratory pilots between December 2018 and March 2019 to test what a data trust would actually involve in practice: one on tackling illegal wildlife trade, one on reducing food waste, and one on improving local public services in partnership with the Royal Borough of Greenwich. None of the three pilots was a research-data project specifically, but the ODI’s resulting reports on lessons from the pilots are the most concrete, citable account of what a data trust’s governance mechanics look like in operation – repeatable terms of access, a defined beneficiary group, and an accountable intermediary – and they remain the reference point most subsequent research-sector proposals cite.
The Ada Lovelace Institute subsequently broadened the framing from data trusts specifically to data stewardship more generally, publishing reports on legal mechanisms for data stewardship and on participatory approaches to data governance. Its work treats the data trust as one option on a wider menu of stewardship models – alongside data cooperatives, data collaboratives, and civic data trusts – and emphasises that whichever legal form is chosen, meaningful participation by the people the data describes is what makes a stewardship arrangement rights-preserving rather than just administratively convenient.
Why it matters for sensitive research data
Health records, genomic sequences, and other special-category data under data protection law (see CASRAI’s Special Category Data (GDPR Article 9) entry) create a recurring governance problem for researchers and research administrators: the people the data describes usually cannot practically be asked, request by request, whether a new secondary use is acceptable, but a one-time consent captured years earlier at collection can’t anticipate every future research use either. A data trust is one structural answer to that gap. Rather than each new researcher negotiating a bespoke agreement with data holders, or a single institution deciding unilaterally on behalf of contributors it doesn’t fully represent, a trustee with a standing, publicly stated duty of care evaluates requests against agreed criteria – arguably a closer proxy for what contributors would actually want than either extreme.
This matters most in exactly the situations sensitive research data sharing tends to raise: data pooled from multiple sites or institutions where no single body has clear authority to decide for the whole dataset; data from communities (including Indigenous communities – see CASRAI’s Indigenous Data Governance entry, which sets out CARE Principles-based governance as a related but distinct approach grounded in collective rather than trust-law rights) who want an ongoing say in how their data is used, not just a one-time consent form; and data whose value depends on being combined across contributors in ways that make individual, bilateral agreements unworkable at scale.
The trustee/fiduciary model in practice
A functioning data trust generally needs to define, explicitly and in advance:
- Who the trustee is and what legal form it takes – commonly a charitable foundation, a community-interest company, or another structure with a formal duty-bound relationship to its beneficiaries, rather than an ordinary commercial entity whose primary duty runs to shareholders.
- Who the beneficiaries are – the data subjects or contributors on whose behalf the trustee acts, and how their interests are represented in the trustee’s decisions (directly, through elected representatives, through an advisory board, or otherwise).
- The terms of access – the criteria the trustee applies when deciding whether to approve a given request: permitted purposes, prohibited uses, review or ethics-approval requirements, and any conditions attached to approval (for example, a requirement to use a trusted research environment rather than exporting raw data).
- Accountability and redress – how the trustee is held to its duties: reporting obligations, an appeals mechanism for rejected requests, and recourse if the trustee’s decisions are later found not to serve beneficiaries’ interests.
- Sustainability – who funds the trustee’s ongoing operation, since a data trust that depends entirely on the goodwill or budget of one data-holding institution risks losing the independence that distinguishes it from an ordinary governance policy.
None of this is settled by a single template. The ODI’s own assessment of its 2019 pilots and independent evaluation work commissioned alongside them concluded that data trusts remain more of a governance pattern to be adapted case by case than a one-size-fits-all legal product – which is consistent with why CASRAI’s dictionary entry (linked below) describes it the same way.
Practical considerations for research administrators
If you are evaluating whether a data trust arrangement is the right fit for a sensitive dataset your institution holds, contributes to, or wants to access, weigh the following before committing to the structure:
- Scale and multiplicity of contributors: a data trust’s overhead (setting up an independent trustee, defining terms of access, building accountability mechanisms) is generally only justified where data comes from, or is used by, enough distinct parties that a single bilateral data sharing agreement per relationship becomes impractical. A two-party collaboration rarely needs a trust.
- Existing legal basis: a data trust changes who decides on access; it does not by itself supply a legal basis for processing personal data under data protection law. Confirm the applicable basis – for UK/EU contexts, commonly GDPR Article 6(1)(e) for public-task processing or the research-specific provisions under GDPR Recital 33 for broad consent – independently of the governance structure layered on top.
- Interaction with existing repository and access infrastructure: a data trust’s trustee can specify how approved requesters actually access the data – for instance, requiring use of a trusted research environment or requiring data to pass through a repository holding CoreTrustSeal certification – rather than building new technical infrastructure from scratch.
- Representation of beneficiaries: a data trust that names a trustee but gives contributors no real mechanism to shape or challenge its decisions functions, in practice, closer to an ordinary data governance policy with extra legal paperwork. The Ada Lovelace Institute’s participatory data stewardship work is a useful checklist for what genuine beneficiary representation looks like beyond a notional duty of care.
- Whether a lighter-weight option already solves the problem: if the real need is a well-drafted agreement between a small, fixed set of known parties, a data sharing agreement or an institutional data governance policy may deliver the needed protections without the cost of standing up an independent trustee.
Frequently asked questions
Is a data trust a legal requirement for sharing sensitive research data?
No. There is no jurisdiction that mandates a data trust structure specifically for research data sharing. It is one governance option among several – data sharing agreements, institutional governance policies, trusted research environments, and data commons are all more commonly used. A data trust is worth considering specifically when governance needs to serve multiple contributors or a diffuse beneficiary group that no single institution can fairly represent.
Does a data trust replace the need for ethics approval or a legal basis for processing?
No. A data trust governs who decides on access and how; it does not substitute for research ethics review or for an underlying legal basis to process personal data (such as consent, public task, or another applicable basis under data protection law). Most data trust designs build ethics review into the trustee’s approval criteria rather than replacing it.
How is a data trust different from a biobank or repository’s own access committee?
An access or data-access committee run by the data-holding institution itself is a governance mechanism, but it is not independent of the institution in the way a data trust’s trustee is designed to be – the institution retains ultimate control. A data trust deliberately places decision-making authority with a separate, duty-bound entity, which matters most when the data-holding institution’s interests may not fully align with contributors’ interests.
Has the data trust model been widely adopted for research data specifically?
Adoption remains limited and largely exploratory. The concept’s most developed pilots (ODI, 2018-2019) were not research-data projects. The Ada Lovelace Institute and others have since explored the model, and closely related variants, for health and biomedical data specifically, but as of this writing it remains a governance pattern under active development rather than an established, widely deployed standard in the way a data sharing agreement or institutional data governance policy is.







