Data Marketplaces & Intermediaries

A US appeals court says AI training can be copyright infringement — and licensed, consented data just became the safe harbor

Source: Legal reporting · Oct 3, 2026

A US appellate court this cycle became the first to hold that using copyrighted material as AI training data was not fair use, at least for the system in front of it. The court was careful to note that the technology at issue was not a generative AI, which limits the ruling's direct reach — but the significance is in the precedent it hands to every creator and publisher currently suing generative-AI developers. The 'it's all fair use' assumption that underwrote a lot of scraping just took its first appellate hit.

The strategic read is that legal risk has moved from theoretical to concrete. For years the industry trained on scraped data on the working assumption that fair use would hold; a single appellate decision does not settle the question, but it changes the risk calculus for anyone building on unlicensed corpora. The downside is no longer hypothetical, and it is the kind of downside — injunctions, damages, forced retraining — that can threaten a model's entire commercial basis.

That repricing flows directly to the alternative. Licensed, consented, provenance-tracked data has always been the more expensive path and the safer one; this ruling widens the gap in its favor. A corpus you can prove you had the right to use is not just ethically cleaner — it is insurance against exactly the legal exposure the court just made real. The premium on documented data is a risk premium, and the risk just went up.

The through-line is the one this market keeps rediscovering from new directions: provenance is not a nicety, it is the asset. Regulators made it a procurement requirement; buyers made it a pricing factor; now a court has made it a liability shield. However the broader fair-use fight resolves, the direction is set — training data you can defend is worth more than training data you merely gathered, and the distance between the two is widening with every ruling.

Key Points

  • A US appellate court became the first to find that using copyrighted material as AI training data was not fair use in the case before it
  • The court limited its holding to the specific (non-generative) system at issue, but creators and publishers will cite it against generative-AI developers
  • The ruling raises the legal risk of training on scraped, unlicensed data
  • Licensed, consented, provenance-tracked corpora are the defensible alternative — and just got more valuable