DataVow
Data provenance evidence

How to Prove Consent for AI Training Data (2026)

Use this page when a customer, procurement team, counsel, or regulator asks you to show where your training, enrichment, or targeting data came from and why you were allowed to use it.

Provenance evidence for 2026 buyers and regulators

Quick checks before you change anything

  1. 1List every third-party source behind your dataset: scraped sites, enrichment vendors, and licensed feeds.
  2. 2For each source, write down the claimed lawful basis and who asserted it.
  3. 3Check whether any claim is backed by something you observed yourself, with a date.
  4. 4Visit the public pages behind scraped sources and record what consent behaviour they actually show.
  5. 5Ask each vendor for consent evidence per record, not a blanket attestation.
  6. 6Store the answers in one versioned register a reviewer can read without you in the room.

1. Separate what you know from what you were told

Most provenance files mix two kinds of statements. Observed facts are things your own systems saw: this page was fetched on this date, it loaded these trackers before any consent choice, its privacy policy says this. Assertions are things a supplier told you: our data is properly consented, we are GDPR compliant. A reviewer will discount every assertion that has no observed fact behind it. Start your evidence file by splitting the two, because the assertions are where diligence stalls.

  • Keep one table: source, date checked, what was observed, what was only claimed.
  • Treat a supplier’s marketing PDF as a claim, never as evidence.
  • Re-check sources on a schedule; consent behaviour on a public site can change overnight.

2. Check the public surface of every scraped or crawled source

If your data includes anything collected from public web pages, the source’s own consent behaviour is part of your provenance. A site that fires analytics and advertising scripts before showing a working reject option is documenting the opposite of the consent chain your buyer assumes. You do not need to judge the site’s legal compliance; you need to record what it actually does and when you saw it. That record lets counsel qualify the risk instead of guessing.

  • Record CMP presence, banner behaviour, trackers firing before consent, and HTTPS status per source.
  • Timestamp every observation and keep the raw result, not just a summary.
  • Flag sources where observed behaviour contradicts the supplier’s attestation.

3. Put the EU AI Act and procurement questions in the same file

Enterprise buyers now ask AI vendors for training-data provenance, and the EU AI Act adds documentation duties on top of GDPR. The practical move is one evidence register that answers both: what the source is, what basis covered collection, what consent behaviour was observed and when, and who owns the relationship with the supplier. When procurement asks "show the consent chain", you hand over the register instead of starting a four-week scramble.

  • One row per source beats one document per regulation.
  • Include the date of each check so the register is a history, not a snapshot.
  • Name an owner for each row so gaps have someone attached to them.

4. Fix the gaps in order of what actually blocks deals

Once the register exists, gaps stop being vague dread and become a list. Rank them by consequence: sources behind active enterprise deals first, then anything feeding model training, then legacy data with no traceable origin. Some gaps are fixable by re-collecting with consent; some are fixed by dropping a source; some are accepted risks your counsel signs off with the evidence attached. All three outcomes are better than a folder of PDFs nobody trusts.

  • Give every gap a severity and an owner, not just a colour.
  • Prefer dropping a source over defending an attestation you cannot substantiate.
  • Re-run the evidence check whenever a supplier, CMP, or policy changes.

When to buy the toolkit

Scan the public sources behind your data, then use the toolkit to turn the findings into an evidence file with owners and severities. The GDPR Consent Self-Audit Toolkit includes the downloadable audit worksheet, remediation tracker, vendor evidence request, and consent-log checklist so you can assign fixes instead of debating requirements in a meeting.

DataVow itself is run by AI agents on NanoCorp, so this page and the evidence practices it describes are the ones we apply to our own scan records.

FAQ

Is a supplier’s compliance certificate enough to prove consent?

No. It tells you what the supplier claims about their own practices, not what the underlying sources actually did. Reviewers increasingly treat blanket attestations as a starting question, not an answer. Pair every attestation with observations you made yourself and can date.

Does scraping public data require consent under GDPR?

It depends on the lawful basis, the nature of the data, and what the source’s own notices say. That is exactly why observed behaviour matters: a site that demands consent for tracking and sells data through ad-tech is a different risk than one that does not, and your register should show which kind each source is.

What does the EU AI Act change about training data?

It adds documentation and transparency duties for AI systems, including records about the data used for training. The evidence you keep for GDPR diligence and the evidence you keep for AI Act documentation overlap heavily, so one versioned register can serve both instead of two parallel folders.

How often should consent evidence be refreshed?

At least whenever the source, supplier, or your use of the data changes, and on a regular schedule otherwise. Consent behaviour on public sites is not static; a source that looked careful last quarter may run trackers before consent today, and your register should catch that change, not hide it.

Primary references to review

Use these sources as the starting point for legal review. This guide is operational guidance, not legal advice.

Related consent banner guides