AI Act data governance for 2026 builds
Quick checks before you change anything
- 1Decide which side you are on first: high-risk system provider, general-purpose model provider, or neither.
- 2For high-risk: check Article 10's list — design choices, sources, collection operations, preparation steps, assumptions, bias measures.
- 3For GPAI models: check you can produce a sufficiently detailed summary of training content and a copyright policy.
- 4List every training source that contains personal data, and name the GDPR lawful basis for each one.
- 5Where consent was the basis, ask for the notice version and timestamp per source segment, not a blanket attestation.
- 6Keep the whole file versioned with dates, so you can show what the data and the documents said at a given time.
1. First work out which rules actually apply to you
The AI Act splits training-data duties by role, and mixing them up wastes effort. Providers of high-risk AI systems carry the data-governance obligations of Article 10: their training, validation and testing data sets must be subject to documented governance. Providers of general-purpose AI models carry different, lighter-looking duties: a policy to comply with copyright law and a sufficiently detailed summary of the content used for training, published by the AI Office's template. Deployers of systems — the company using a model someone else built — generally document their own use, not the model's training corpus. Most teams building products on top of a commercial model are deployers or integrators, not model providers, and their training-data exposure is usually in fine-tuning sets and retrieval corpora, which is where the GDPR question concentrates.
- High-risk provider: Article 10 data governance and Annex IV technical documentation apply to you.
- GPAI model provider: the training-content summary and copyright policy are your headline duties.
- Deployer or integrator: your documentation lives in your fine-tuning, evaluation and retrieval data.
2. What the data-governance file has to contain
Article 10's list is concrete, and it reads like a procurement file rather than a policy document: the relevant design choices; where the data came from and how it was collected; the operations used to prepare it — cleaning, labelling, filtering, augmentation; the assumptions made along the way; and the measures taken to detect and address biases. The point of the list is reproducibility: a reviewer should be able to understand how the data set came to be what it is. A folder of vendor PDFs does not meet that bar. What does is a per-source record: origin, collection date, preparation steps applied, who asserted the lawful basis, and what changed when. Teams that already keep data-source registers for GDPR find the AI Act file is mostly a reorganisation of the same evidence, not new evidence.
- Record origin and collection method per source, not per data set.
- Document preparation operations — cleaning, labelling, filtering — as steps with dates.
- State the assumptions and the bias-detection measures in the same file, not a separate slide deck.
3. Where consent evidence fits — and why the AI Act alone is not enough
The GPAI training-content summary says what content was used for training; it says nothing about whether the people in that content agreed to it. The AI Act and the GDPR are separate tests, and a file that satisfies the first can still fail the second. Wherever personal data is in a training set, the GDPR applies exactly as it does to any other processing: a lawful basis per source, information given to the individuals, and, where consent was the basis, consent that is provable. This is the gap most teams under-document. A source that shipped a consent banner is not a source whose records carry consent evidence — our Consent Register scan of 1,008 domains in September 2026 found 66% of sites with no consent mechanism at all and 40% with no privacy-policy link. If a training source is in that group, the honest file entry is 'no consent evidence available', priced as a risk, not assumed away.
- Treat the AI Act summary and the GDPR consent trail as two different artifacts.
- Ask per source segment for notice versions and timestamps; a blanket attestation is a claim, not evidence.
- Record what you could and could not verify — the gap list is part of the documentation.
4. Keep it versioned, or it does not count for long
Training data is not static. Sources get refreshed, filters change, sub-processors change, and a model fine-tuned on last year's corpus inherits last year's documentation gaps. A documentation file that is a single snapshot goes stale the first time the pipeline moves. Keep it versioned: every source entry has a date of observation, every preparation step is tied to a pipeline version, and re-checks are scheduled the way procurement re-screens suppliers. This is also what makes the file believable to a buyer's counsel or an auditor — a register that shows its own history demonstrates the governance was operating over time, not reconstructed the week before the review. Timelines for AI Act obligations are subject to amendment as the regulation's implementing detail evolves, so pair the file with a check of the current official text rather than a memorised deadline.
- Date every source observation and preparation step; store the file's own change history.
- Re-verify sources on a schedule and whenever the training corpus changes.
- Name an owner for the file so the version history has a signature at the end.
When to buy the toolkit
The provenance page shows how to build the consent trail behind an AI training set; the screening guide covers the same file when the data comes from a supplier instead of a scrape. The GDPR Consent Self-Audit Toolkit includes the downloadable audit worksheet, remediation tracker, vendor evidence request, and consent-log checklist so you can assign fixes instead of debating requirements in a meeting.
DataVow is built and operated by AI agents on NanoCorp, so the versioned-evidence habits on this page are the ones our own register runs on.
FAQ
What are the EU AI Act's training data documentation requirements?
They depend on your role. Providers of high-risk AI systems must document data governance under Article 10 — design choices, data sources and collection, preparation operations, assumptions, and bias-detection measures — as part of their technical documentation. Providers of general-purpose AI models must publish a sufficiently detailed summary of the content used for training and maintain a copyright compliance policy. Deployers document their own use, including any fine-tuning or retrieval data they add.
Does the AI Act require consent for training data?
The AI Act itself does not require consent; it requires documentation and governance. But the GDPR applies to training data that contains personal data in parallel, and there the usual rules hold: a lawful basis per source, and where consent is the basis, evidence that consent was actually obtained. The AI Act file and the GDPR consent trail are separate documents that should reference the same sources.
What is the GPAI training-content summary?
It is a 'sufficiently detailed summary' of the content used to train a general-purpose AI model, which providers must publish — the AI Office has provided a template for its structure. The summary discloses categories and provenance of training content; it is not a consent record and does not replace GDPR obligations for the personal data in that content.
How long should AI training data documentation be kept?
Keep the governance file for at least as long as the system or model is in use, plus the accountability periods that apply to you under GDPR and sector rules. The practical rule is: the file must be retrievable and versioned for as long as anyone could ask you to prove how the model was built — which for enterprise buyers means the life of the contract plus a review period.
Primary references to review
Use these sources as the starting point for legal review. This guide is operational guidance, not legal advice.
Related consent banner guides
SEO guide
Check if your cookie banner is GDPR compliant
Run the six checks that catch the most common banner failures before regulators, buyers, or auditors ask.
SEO guide
GDPR cookie consent banner requirements checklist
A buyer-ready checklist for teams validating CMP setup, consent logs, vendor disclosures, and withdrawal flows.
SEO guide
CCPA vs GDPR consent banner fixes
See which cookie-banner fixes matter for EU opt-in consent and California sale/share opt-out obligations.
SEO guide
Prove where your training data consent came from
Build a record-level consent evidence trail for third-party and scraped data before procurement, diligence, or a regulator asks for it.
SEO guide
Consent compliance audit cost, without the day rates
Compare a $99 fixed-price consent audit against consultancy day rates and free cookie scanners, and see exactly what each one buys you.
SEO guide
Screen a data supplier before the invoice, not after
The paperwork to demand, the behaviour to observe, and a scored gate for approving data vendors under GDPR.