Data Marketplaces & Intermediaries
AI data licensing is shifting from one-time sales to real-time rented feeds
Source: Pebblous · Sep 22, 2026
The way AI labs buy data is changing shape. According to a new analysis, the number of disclosed licensing deals roughly doubled — from about 10 to 22 — after Cloudflare started blocking AI crawlers by default in July 2025, cutting off the free scraping that labs had relied on and pushing them toward paid, permissioned access. In the deals now visible, Perplexity has emerged as the single largest buyer at around 37 percent, with OpenAI near 29 percent.
Underneath the deal count is a structural shift: from buying a static corpus once to renting a live feed continuously. As blocking hardens and the open web closes, what labs need is not a one-time dump but ongoing, metered access to fresh, licensed data — a subscription to a stream rather than a purchase of a pile.
The timing pressure is real. Epoch AI projects that the usable stock of public human text is effectively exhausted somewhere between 2026 and 2032. Once the open reservoir runs dry, the only growth left is non-public, licensed, consented data — exactly the category that rewards clean provenance and the ability to prove a dataset was collected with permission. The market is moving toward continuous, contracted, rights-cleared supply, and away from the scrape-first era that built the first generation of models.
Key Points
- Disclosed licensing deals jumped from ~10 to ~22 after Cloudflare began blocking AI crawlers by default in July 2025
- Perplexity is now reported as the single largest buyer (~37% of known deals), with OpenAI around 29%
- Epoch AI projects the stock of public human text is effectively exhausted between 2026 and 2032
- The model is moving from a one-time corpus sale toward ongoing, metered access to live data feeds