Data Marketplaces & Intermediaries

The video-data gold rush reshaping foundation models

Source: Forbes · May 11, 2026

A Forbes analysis lays out how video training data has become a multi-billion-dollar bottleneck as foundation models push past text into temporal and spatial reasoning. The piece traces a clear shift underway across the industry: away from scraping video at quantity-at-all-costs, toward curated, licensed, specialized datasets with documented provenance. The specific mechanism getting attention is C2PA — the Coalition for Content Provenance and Authenticity — whose content manifests are moving from a nice-to-have standard to an operationally necessary filter, used to separate verified, human-sourced content from synthetic material and avoid the 'model collapse' that comes from training on a model's own outputs.

This is the clearest outside confirmation yet of the thesis our business is built on: as foundation-model builders exhaust easily scraped video and audio and move to specialized, higher-stakes domains, the question stops being 'how much data can we get' and becomes 'can we prove where this came from and that a real person consented to it.' Provenance isn't a compliance afterthought in this framing — it's becoming a hard technical requirement, on the same list as labeling quality and annotation accuracy.

For rights-cleared, consented, chain-of-custody multimodal interview data specifically, this is directly on-thesis: interview video is exactly the kind of high-stakes, human-verified, hard-to-synthesize content that a C2PA-literate buyer would prioritize over unverified scraped alternatives. This is the strongest external validation this run of why provenance is the pitch, not just a feature of it.

Key Points

  • Video training data has become a multi-billion-dollar bottleneck as models move beyond text
  • Industry is shifting from quantity-at-all-costs to curated, licensed, specialized datasets
  • C2PA (Coalition for Content Provenance and Authenticity) manifests are becoming operationally necessary
  • Goal: filter synthetic content and avoid 'model collapse' from unverified training data