Other

The data wall is real — and it's reshaping what ‘training data’ means

Source: Occludo AI · Sep 11, 2026

For most of the last decade, the default answer to “how do you make a language model better” was some version of “more data.” That answer is running out of runway. Researchers tracking the growth of publicly available text relative to the pace at which frontier labs consume it have been flagging the same trend for a few years now: the supply of new, high-quality, previously-unseen text on the open web is growing far more slowly than training runs are scaling. Common Crawl, Wikipedia, licensed book corpora, and code repositories are all finite, heavily reused, and increasingly picked clean. Every major lab now trains on some meaningful fraction of the same underlying pool.

Synthetic data was supposed to be the release valve. If a model can generate its own practice problems, critique its own answers, and bootstrap improvement without needing fresh human text, the data wall stops mattering. In narrow domains with a verifiable ground truth — math proofs, code that either compiles and passes tests or doesn't — that loop genuinely works, and it's a real part of how recent reasoning-focused models got better. But synthetic data has a structural ceiling outside those verifiable domains: it's sampled from the model's own distribution, so it can sharpen what the model already knows how to do, but it can't reliably hand the model a way of reasoning it never had. Model collapse research — the finding that models trained repeatedly on their own outputs drift and lose the tails of the original distribution — is the sharper version of the same problem. Synthetic data recombines; it doesn't originate.

So the interesting question isn't “where's more text”, it's “what's the thing text never captured in the first place.” A transcript of a real technical conversation — someone working through a hard problem out loud, hitting a wrong turn, getting a correction, revising, and eventually landing on a defensible answer — contains a kind of information a static document never does: process, not just output. The question that was asked, the answer that was given, the follow-up that probed a weak point, the correction when the answer was wrong, a domain expert's judgment of how good the response actually was, and, where it exists, what happened afterward when that judgment was tested against reality. That's six or seven linked pieces of signal riding on top of what looks, on the surface, like a single exchange — and almost none of the text corpora models have already been trained on preserves that chain. A forum post, a textbook, a blog explainer gives you the polished output. It doesn't give you the reasoning that produced it, the mistakes along the way, or an independent expert's assessment of whether it held up.

That's the shift worth watching in AI training-data licensing over the next few years: less emphasis on raw volume of scraped or purchased text, more emphasis on structured, consent-based capture of real human reasoning — with the full chain intact, not just the final answer. It won't replace the web-scale pretraining corpus that got the field this far. But as the marginal value of one more billion generic tokens keeps falling, and the marginal value of one more verified, expert-judged reasoning trace keeps rising, that's where a meaningful share of the next wave of data spend is headed.

Key Points

  • Public web text — the fuel that trained the last generation of frontier models — is a finite resource, and researchers have been warning for years that high-quality supervised text could run short this decade
  • Synthetic data can stretch a model's existing capabilities further, but it struggles to teach a model something genuinely outside what it already knows, because it's generated by that same model's own distribution
  • The data category getting scarcer isn't raw text volume, it's *process* — the record of how a real expert reasoned through a problem, not just the answer they landed on
  • A growing share of new data-licensing activity is shifting toward structured human interactions with a documented chain: the question, the response, the follow-up, the correction, an evaluator's judgment, and the real-world outcome