Data Marketplaces & Intermediaries

The data business is shifting from labeling to building RL environments

Source: Pebblous · Sep 18, 2026

The center of gravity in the AI-data business is moving. For the last several years the money was in labeling — annotating images, ranking model outputs, writing preference data. The newest, highest-value work is different: building reinforcement-learning environments, interactive sandboxes with expert-written rubrics that a model can actually operate inside and be scored against, rather than a static pile of labeled examples.

The incumbents still dominate the top line. The so-called Big Four — Scale, Surge, Mercor and Handshake — reportedly still account for more than three-quarters of category revenue. But a wave of environment-first challengers is climbing fast: Prime Intellect, which has cultivated a community library of 2,500-plus environments; Mechanize, reportedly paying environment builders unusually high salaries; and Snorkel, Toloka and AfterQuery, the last said to be past $100M in annualized revenue. The demand signal underneath all of this is enormous — Anthropic alone is reported to spend north of a billion dollars a year on RL environments.

It is worth holding the skepticism alongside the enthusiasm. Andrej Karpathy, among others, has openly questioned whether 'RL environments' is a durable category or a temporary label for work that will get absorbed back into model labs. Either way, the through-line matters for anyone who supplies data: buyers increasingly pay for structured, interactive, expertly-scored experiences, not for volume. The premium is on realism and rubric quality — long-form, high-context, genuinely human material — which is precisely the kind of data that is hardest to fake and most valuable to collect with consent.

Key Points

  • The valuable work is moving from static labeling to interactive RL environments — sandboxes with expert-written rubrics a model can act inside
  • The Big Four (Scale, Surge, Mercor, Handshake) still hold >75% of revenue, but environment-first challengers are rising: Prime Intellect (2,500+ community environments), Mechanize, Snorkel, Toloka, AfterQuery ($100M+ ARR)
  • Mercor's ARR went from ~$1B in February to ~$2B in June 2026; Anthropic alone reportedly spends $1B+/yr on RL environments
  • Andrej Karpathy has publicly questioned whether the category lives up to the hype