Data Marketplaces & Intermediaries

Fewer AI licensing deals now grant training rights — the market is pivoting to grounding

Source: Media and the Machine · Sep 23, 2026

A quiet but important change is running through AI content-licensing: the training right itself is disappearing from many deals. Where nearly every 2023–24 agreement bought the right to train models on the licensed content, only around 40% of 2026 deals do. OpenAI's arrangements with Getty and the Associated Press reportedly omit training rights entirely, and of Meta's eleven disclosed deals, only the News Corp deal mentions training.

The reason is a shift in what the content is for. The industry is moving from 'buy a corpus to build a smarter model' toward 'license a live source to deliver better, current answers' — retrieval and real-time grounding rather than one-time absorption into weights. That change is driven by diminishing returns on scraped web text, hardening publisher resistance, and legal exposure that makes broad training grants risky to sign.

For anyone building a data business, the signal is nuance, not retreat: licensing is alive and well, but the terms are fragmenting into specific, use-case-bound rights. Provenance and precise, well-scoped permissions — knowing exactly what a dataset may and may not be used for — become the product, not the fine print. A corpus that arrives with clean, explicit, purpose-bound consent is worth more in this market than one sold under a vague blanket grant.

Key Points

  • Only ~40% of 2026 content-licensing deals include model-training rights, down from nearly all in 2023–24
  • OpenAI's Getty and AP deals reportedly omit training; of Meta's 11 deals, only News Corp mentions it
  • The shift is from 'buy content to train a better model' to 'license content to deliver better answers' (RAG / real-time grounding)
  • Drivers: diminishing returns on scraped web data, publisher resistance, and mounting legal scrutiny