Data Marketplaces & Intermediaries

The EU AI Act just made data provenance a procurement requirement — not a nice-to-have

Source: Industry reporting · Sep 29, 2026

With the EU AI Act reaching full enforcement in August 2026, a requirement that used to live in legal footnotes is now operational: developers must disclose training-data sources, honor copyright opt-outs, and document provenance. And 'provenance' no longer means a vague attestation — it's increasingly field-level, requiring source IDs, a recorded consent basis, and a chain of custody from the original collection event all the way to the training snapshot.

That reframes how data is bought. AI due diligence now opens with the question 'what was this model trained on, and on what consent basis for each dataset?' A corpus without a reconstructable, consent-backed lineage isn't just lower quality — it's a liability a careful buyer will discount or refuse. Provenance is moving from compliance overhead to an actual procurement criterion.

This is the ground a consented, provenance-tracked dataset is built to stand on. When consent and chain of custody are captured at the source, a buyer isn't just acquiring data — they're acquiring the documentation that lets them use it under the new rules. As enforcement bites, 'can you prove where this came from and that people agreed to it?' becomes the question that decides which data is sellable at all.

Key Points

  • The EU AI Act reached full enforcement in August 2026: disclose training-data sources, respect opt-outs, document provenance
  • Provenance is now field-level — source IDs, consent basis, and a reconstructable chain of custody from collection to training snapshot
  • AI due diligence now opens with 'what was it trained on, and on what consent basis?'
  • Consent + provenance are shifting from compliance overhead to a buying criterion