Why Art. 10 is more than a data protection question
Art. 10 AI Act sets out how training, validation and testing data sets for high-risk AI systems must be sourced and managed. Under para. 1, the provision applies whenever an AI model is trained using data; for systems developed without training, para. 6 still imposes at least the obligations for testing data sets. The article therefore addresses not only data protection but the technical and organisational quality of the data basis itself. For companies providing or operating high-risk AI under Annex III, this is not an abstract requirement: the obligations for these systems apply from 2 August 2026, and Art. 10 is among the provisions that demand robust evidence to that effect.
The eight building blocks of governance procedures
Para. 2 lists eight aspects that data governance and data management procedures must cover. In practice, this means you need documentation on
- the design choices underlying the data selection (point (a)),
- collection processes, data sources and – for personal data – the original purpose of collection (point (b)),
- preparation steps such as annotation, labelling, cleaning, updating or aggregation (point (c)),
- the assumptions underlying the data (point (d)),
- an assessment of the availability, quantity and suitability of the data sets (point (e)),
- an examination of possible biases likely to affect health, safety, fundamental rights, or the prohibition of discrimination (point (f)),
- measures to detect, prevent and mitigate such biases (point (g)), and
- the identification of data gaps or shortcomings, together with a plan to address them (point (h)).
The typical gap in practice: data sets exist, but the reasoning behind them is nowhere recorded. Who selected which data source and why, which assumptions were made, which bias examination took place – all of this must be reconstructable after the fact, not merely held in the heads of individual ML leads.
Quality criteria: relevant, representative, error-free
Para. 3 requires that training, validation and testing data sets be relevant, sufficiently representative, and, to the best extent possible, free of errors and complete in view of the intended purpose. They must have the appropriate statistical properties – including, where relevant, in relation to the persons or groups of persons for whom the system is intended to be used. These properties may be met at the level of individual data sets or a combination thereof, which in practice leaves room for manoeuvre where a single data set alone is not sufficiently representative.
Para. 4 adds that, to the extent required by the intended purpose, the data sets must take into account, as appropriate, the characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting within which the system is intended to be used. A system developed for a particular market or user group but trained on data from a different context will typically not satisfy this requirement without further review and adjustment.
The special case: special categories of personal data
Para. 5 opens a narrow exception: providers may process special categories of personal data where strictly necessary for the detection and correction of biases under para. 2, points (f) and (g). This processing is subject, in addition to the requirements of the GDPR, Regulation (EU) 2018/1725 and Directive (EU) 2016/680, to six cumulative conditions: bias correction must not be effectively achievable by processing other data, including synthetic or anonymised data (point (a)); technical limitations and state-of-the-art security and privacy-preserving measures, including pseudonymisation, must be in place (point (b)); strict, documented access controls must be established (point (c)); the data must not be transmitted to third parties (point (d)); the data must be deleted once the bias has been corrected or the retention period has ended (point (e)); and records of processing activities must document the reasons why such processing was strictly necessary (point (f)).
This exception is often overlooked in practice – or, conversely, interpreted too broadly. Both are risky: without relying on the exception, many companies lack a legal basis even to examine bias in sensitive characteristics such as ethnic origin or health data. With an overly broad interpretation, there is a risk of breaching the cumulative conditions, in particular the deletion obligation under point (e) and the documentation obligation under point (f).
What this means for your documentation
Taken together, Art. 10 requires an unbroken chain of evidence: from the data source through preparation to bias testing. Anyone who has not documented this chain end to end will struggle, at the latest during a conformity assessment or market surveillance, to demonstrate compliance with the requirements of paras. 2 to 5. Precisely because the high-risk obligations for Annex III systems apply from 2 August 2026, it is worth reviewing your data basis early.
Whether, and to what extent, your AI system qualifies as high-risk, and which obligations under Art. 10 concretely apply, can be clarified in a few minutes using our free risk assessment at /einstufung.