Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us
Dictionary termTrack CProposedv2026.1

EU AI Act Article 10 (Data and Data Governance)

Article 10 of the EU AI Act (Regulation (EU) 2024/1689), in Title III, Chapter 2, sets the data-quality and data-governance obligations for the training, validation, and testing datasets used to build high-risk AI systems. A system falls under Article 10 only after it has already met the Act's separate 'high-risk' classification test (an Annex III listed use case, or a safety component covered by Annex I sectoral product law). For a high-risk system built using techniques that involve training models on data, Article 10(2) requires documented data-governance practices covering collection and origin, preparation operations (annotation, labelling, cleaning, updating, enrichment), bias examination, bias-mitigation measures, and identification of data gaps; Article 10(3)-(4) require the training, validation, and testing datasets to be relevant, sufficiently representative, and to the best extent possible free of errors and complete for the system's intended purpose and deployment context. For a high-risk system that does not use such training techniques, Article 10(6) narrows the same requirements to the testing dataset only. Article 10(5) separately permits limited, safeguarded processing of special-category personal data solely to detect and correct bias. It does not apply to AI systems or models developed and used solely for scientific research and development prior to being placed on the market or put into service, per the Article 2(6)/2(8) research exemption.

ByCASRAI Editorial Board
· Last updated 18 Jul 2026

Examples

Worked examples

  • Is an instance

    A university builds an AI-based admissions- or grant-eligibility screening tool intended for real institutional use — an Annex III 'access to education/employment' high-risk use case. Because it is trained on historical applicant data, Article 10(2)-(4) require the developers to document the data's origin and collection process, examine it for demographic bias before use, and apply mitigation measures (e.g. reweighting, removing unrepresentative segments) before the tool is deployed operationally.

  • Is an instance

    A hospital-affiliated research team develops a clinical-decision-support model purely for a peer-reviewed methods paper, with no plan to deploy it in patient care or place it on the market. This activity falls under the Article 2(8) research-and-development exemption, so Article 10 does not directly apply — though the same data-governance discipline remains good research practice and becomes mandatory the moment the tool moves toward real clinical deployment.

Counter-examples

Looks similar, but isn't

  • Not an instance

    An internal, non-deployed research text-mining pipeline that is never placed on the market or put into service, and does not fall into any Annex III high-risk use case or Annex I safety-component category, is not subject to Article 10 at all — the high-risk regime only attaches once a system meets the Act's 'high-risk' classification, not merely because it processes training data.

Editorial commentary

Article 10 (“Data and Data Governance”) is the provision of the EU AI Act — Regulation (EU) 2024/1689, Title III, Chapter 2 (“Requirements for High-Risk AI Systems”) — that governs the quality and governance of the datasets used to build a high-risk AI system. It only applies once a system has already met the Act’s separate high-risk classification test: it is either listed in Annex III (e.g. AI used in recruitment, education/training access, creditworthiness assessment, essential public/private services, biometric identification, or certain law-enforcement and migration contexts), or it is a safety component of a product already covered by EU sectoral legislation listed in Annex I (e.g. medical devices, machinery). For research administrators, the practical question is rarely “does the AI Act apply to research?” in the abstract — it is whether a specific AI tool a lab or unit is building or deploying crosses into one of those high-risk categories, and if so, whether it is intended for real institutional or market use rather than staying inside the Article 2(6)/2(8) research-and-development exemption.

What Article 10 Requires

Data governance practices (Article 10(2))

For a high-risk system built using techniques that involve training a model on data, providers must put in place data-governance and data-management practices covering: the relevant design choices; the data collection processes and, for personal data, the original purpose of collection; relevant data-preparation processing operations such as annotation, labelling, cleaning, updating, enrichment, and aggregation; the formulation of assumptions about what the data is intended to measure or represent; an assessment of the availability, quantity, and suitability of the datasets needed; an examination for possible biases likely to affect health, safety, or fundamental rights, or to lead to discrimination, especially where data outputs influence future operations; appropriate measures to detect, prevent, and mitigate any identified bias; and the identification of relevant data gaps or shortcomings that prevent compliance, and how those are addressed.

Dataset quality criteria (Article 10(3)-(4))

Training, validation, and testing datasets must be relevant, sufficiently representative, and — to the best extent possible — free of errors and complete in view of the system’s intended purpose; they must have the appropriate statistical properties for the persons or groups of persons the system is intended to be used on. These characteristics may be met at the level of individual datasets or a combination of datasets. Datasets must also account for the characteristics or elements particular to the specific geographical, contextual, behavioural, or functional setting within which the system is intended to be used.

Special-category data for bias correction (Article 10(5))

Providers may exceptionally process special categories of personal data (as defined in GDPR Article 9) strictly to detect and correct bias, subject to conditions: the bias cannot effectively be addressed by other data (including synthetic or anonymised data); the data is subject to technical limitations on re-use; state-of-the-art security and pseudonymisation measures apply; strict access controls and confidentiality obligations are in place; and the data is deleted once the bias has been corrected or the relevant retention period has ended, whichever comes first.

Narrower scope for non-machine-learning high-risk systems (Article 10(6))

Where a high-risk AI system does not use training techniques involving models trained on data — for example, a rule-based or expert-system approach — paragraphs 2 to 5 apply only to the testing dataset, not to training or validation data.

Does Article 10 Apply to Research Institutions?

Most AI activity inside a university or research institute — exploratory model development, methods papers, internally used research tools that are never placed on the market or put into service operationally — falls under the Article 2(6) and 2(8) exemptions for AI systems and models developed and used solely for scientific research and development. Article 10 does not directly bind that activity. The exemption narrows sharply, however, once a tool is intended for real institutional deployment that lands in an Annex III category: examples relevant to research administration include AI used to screen admissions or grant applications (access to education), AI-assisted recruitment or promotion-review tools (employment), and AI-based eligibility or fraud-detection systems for benefits or essential services. The European Commission has stated that clarifying exactly where this research-exemption boundary sits — particularly for pre-clinical research and product development for medicines and medical devices, where research activity can shade into commercial development before market placement — is a stated guidance priority, but as of this writing that specific guidance has not yet been published. Institutions building or procuring AI tools that could plausibly cross into high-risk, real-deployment territory should not rely on the research exemption as a default and should treat Article 10’s data-governance discipline as good practice regardless of whether it is legally mandatory yet for a given tool.

Current Implementation Timeline

The AI Act entered into force on 1 August 2024. Prohibited-practice provisions applied from 2 February 2025, and general-purpose AI model obligations from 2 August 2025. As originally enacted, obligations for stand-alone high-risk (Annex III) systems — including Article 10 — were due to apply from 2 August 2026, and obligations for high-risk AI embedded in Annex I regulated products from 2 August 2027. A “Digital Omnibus on AI” simplification package reached political agreement between the European Parliament and Council on 7 May 2026, with Parliament’s formal endorsement on 16 June 2026 and the Council’s final approval on 29 June 2026; it defers the stand-alone high-risk (Annex III) compliance deadline to 2 December 2027 and the Annex I embedded-product deadline to 2 August 2028. As of this writing, formal signature and publication in the Official Journal of the EU — the step that makes the amended dates legally binding — was still expected imminently rather than confirmed complete. Research institutions with AI tools approaching high-risk classification should verify the current status directly via the European Commission’s AI Act implementation timeline before treating either date as settled for compliance planning.

How Article 10 Relates to Other Frameworks

Article 10’s data-governance obligations sit alongside, rather than replace, existing research-data practice. Its GDPR Article 9 carve-out for bias-correction processing operates within, not instead of, ordinary GDPR obligations — see CASRAI’s GDPR entry and GDPR and data protection compliance in research guide for how those obligations apply to research data generally. Documenting dataset origin, collection process, and preparation operations under Article 10(2) overlaps substantially with existing training data provenance practice and with the kind of documentation a data management plan already asks a research team to produce; using a documented metadata schema for dataset characteristics makes the Article 10(3)-(4) representativeness and completeness assessment considerably easier to evidence. See also CASRAI’s broader guide to EU AI Act obligations and exemptions for research organizations and its guide to AI training data provenance, copyright, and TDM exceptions for adjacent obligations that often apply to the same dataset.

Practical Steps for Research Data Managers

  • Maintain a documented record of each training/validation/testing dataset’s origin, collection process, and (for personal data) original collection purpose — this is the evidentiary basis for Article 10(2) compliance.
  • Run and document a bias examination for datasets behind any AI tool that could plausibly be used in an Annex III context, even before deployment is certain, so the assessment already exists if the exemption boundary is crossed later.
  • Log known data gaps or representativeness limitations explicitly, rather than leaving them undocumented — Article 10(2) treats gap disclosure as a compliance element, not an admission of failure.
  • Coordinate any special-category data processing done specifically for bias detection/correction with the institution’s data protection office, since Article 10(5) processing sits on top of, not instead of, GDPR Article 9 obligations.
  • Track whether a research tool is moving from pure R&D toward real institutional deployment — that transition is exactly where the Article 2(6)/2(8) exemption stops applying.

Frequently Asked Questions

Does Article 10 apply to all AI systems, or only high-risk ones?

Only high-risk AI systems as defined elsewhere in the Act (Annex III listed use cases, or safety components under Annex I sectoral law). Article 10 sets requirements that apply once a system is already classified as high-risk; it is not a general data-governance rule for all AI.

Is university research AI automatically exempt from Article 10?

Not automatically. AI systems and models developed and used solely for scientific research and development, prior to market placement or being put into service, fall under the Article 2(6)/2(8) research exemption. A tool intended for real deployment in an Annex III context (e.g. admissions or grant screening) is not covered by that exemption even if it originated as a research project.

What counts as a ‘sufficiently representative’ dataset under Article 10?

The Act does not fix a numeric threshold. It requires datasets to be relevant, sufficiently representative, and to the best extent possible free of errors and complete for the system’s specific intended purpose and deployment context, with appropriate statistical properties for the persons or groups the system will be used on — an assessment made against that stated purpose, not a universal standard.

When do Article 10 obligations actually take effect?

Originally 2 August 2026 for stand-alone high-risk (Annex III) systems. A May-June 2026 Digital Omnibus agreement defers that to 2 December 2027, but formal Official Journal publication of that deferral was still pending as of this writing — verify current status before relying on either date for compliance planning.

Can a research institution use personal health or demographic data under Article 10 without extra safeguards?

Only within narrow limits. Article 10(5) allows processing special-category personal data specifically to detect and correct bias, and only where safeguards are met (no viable alternative such as synthetic data, pseudonymisation, strict access controls, and deletion once the bias correction is complete). This sits alongside, not instead of, ordinary GDPR Article 9 obligations.

References

  • Regulation (EU) 2024/1689 (EU AI Act), Title III, Chapter 2, Article 10 — eur-lex.europa.eu
  • artificialintelligenceact.eu, Article 10 summary and explainer
  • European Commission, AI Act implementation timeline — ai-act-service-desk.ec.europa.eu
  • European Commission AI Office, “Supporting the implementation of the AI Act: clear guidelines” (digital-strategy.ec.europa.eu, 4 December 2025)

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="EU AI Act Article 10 (Data and Data Governance)"
      vocab-term-identifier="https://casrai.org/dictionary/term/eu-ai-act-article-10-data-and-data-governance" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/eu-ai-act-article-10-data-and-data-governance",
  "name": "EU AI Act Article 10 (Data and Data Governance)",
  "identifier": "https://casrai.org/dictionary/term/eu-ai-act-article-10-data-and-data-governance",
  "description": "Article 10 of the EU AI Act (Regulation (EU) 2024/1689), in Title III, Chapter 2, sets the data-quality and data-governance obligations for the training, validation, and testing datasets used to build high-risk AI systems. A system falls under Article 10 only after it has already met the Act's separate 'high-risk' classification test (an Annex III listed use case, or a safety component covered by Annex I sectoral product law). For a high-risk system built using techniques that involve training models on data, Article 10(2) requires documented data-governance practices covering collection and origin, preparation operations (annotation, labelling, cleaning, updating, enrichment), bias examination, bias-mitigation measures, and identification of data gaps; Article 10(3)-(4) require the training, validation, and testing datasets to be relevant, sufficiently representative, and to the best extent possible free of errors and complete for the system's intended purpose and deployment context. For a high-risk system that does not use such training techniques, Article 10(6) narrows the same requirements to the testing dataset only. Article 10(5) separately permits limited, safeguarded processing of special-category personal data solely to detect and correct bias. It does not apply to AI systems or models developed and used solely for scientific research and development prior to being placed on the market or put into service, per the Article 2(6)/2(8) research exemption.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/eu-ai-act-article-10-data-and-data-governance",
  "sameAs": [],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "dateModified": "2026-07-18T06:30:47",
  "inLanguage": "en"
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →