Data Marketplaces & Intermediaries

AI training-data licensing: only 4 in 10 deals now include training rights

Source: Quartz · Jan 21, 2026

It's easy to assume that every content-licensing deal signed in the AI era automatically includes rights for model training, but the market data tells a more cautious story. Roughly four in ten licensing deals now explicitly include AI training rights — meaning a majority of deals still don't, either because the rights aren't being requested, the pricing hasn't been worked out, or the licensor isn't yet willing to grant them. That gap is itself useful market intelligence: it suggests there's still meaningful friction, whether legal, pricing, or trust-based, standing between content owners and AI labs even now, well into the current training-data licensing boom.

Where training rights are included, the deal sizes vary enormously — from around $5 million up to $250 million, with the OpenAI–News Corp licensing agreement cited as a benchmark at the top end of that range. That spread reflects just how differentiated the value of a dataset can be depending on its scale, exclusivity, freshness, and how well-suited it is to the specific training objectives a given lab is pursuing. A narrow, highly specialized dataset with strong provenance can command a premium disproportionate to its raw size, while a large but generic archive may struggle to reach comparable per-deal value.

Perhaps the more structurally important trend is the shift away from static, one-time archive licenses toward continuously-refreshed, real-time data feeds. Labs increasingly want ongoing access to new data as it's generated, not a frozen snapshot licensed once and left to age. That shift changes the economics of licensing deals meaningfully: a continuously-refreshed feed supports recurring revenue and ongoing commercial relationships rather than a single upfront payment, but it also demands sustained data production and quality control from the licensor, not just a one-time content dump.

For a company building a licensable interview-data business, this backdrop is directly relevant to how we think about deal structure. The market's clear movement toward continuously-refreshed feeds over static archives is exactly the model that suits structured, ongoing interview data collection — new conversations are generated continuously, not once. It also reinforces that dataset differentiation (scale, freshness, exclusivity, provenance) is what determines where in that $5M-to-$250M range a given deal lands, which is the core value proposition any licensing conversation on our side needs to make concrete and defensible.

Key Points

  • Only ~40% of content-licensing deals now explicitly include AI training rights
  • Deal sizes span roughly $5M to $250M (OpenAI/News Corp being a top-end example)
  • Market is shifting from static, one-time archive licenses toward continuously-refreshed, real-time data feeds