Data Marketplaces & Intermediaries
A $200M clinical-data deal shows what consented, multimodal data is actually worth
Source: Industry reporting · May 21, 2026
The licensing market put real numbers on the table this year. AI companies spent well over four billion dollars acquiring licensed data in 2025, and the disclosed single deals now span a wide band — from a few million dollars at the low end to a quarter of a billion for the largest. The deals that anchor the top of that range share a profile: they are not scraped text. The headline publisher agreement sits near $250M, and one pharma-sector arrangement reportedly ran around $200M for access to a multimodal oncology dataset covering millions of patients — the kind of richly-structured, regulated corpus that simply cannot be assembled from the open web.
The strategic read is that price tracks irreproducibility. Commodity web text keeps getting cheaper because there is effectively an unlimited supply of it and a growing share is itself machine-generated. The datasets that hold their value are the ones a buyer cannot recreate at any price: longitudinal, consented, multimodal records tied to real people and real outcomes, collected under terms that make them legally usable. A clinical dataset is the extreme case — it is multimodal, it is deeply consented by necessity, and it took years and institutional trust to assemble.
That is the same structure, one tier up, that makes consented interview data valuable. A multimodal interview — voice, video, and structured response from a real, consenting participant — is not reconstructable from a crawl. It carries the two properties the $200M end of the market is paying for: it is multimodal, and its provenance is clean enough to license without the buyer inheriting a lawsuit.
The through-line is that the ceiling on a data deal is set by what the data proves, not by how much of it there is. The nine-figure deals are not paying for volume; volume is the cheap part. They are paying for consent, structure, and a chain of custody that holds up. That is exactly the asset the rest of this market is converging on, and it is the one the biggest checks are already being written for.
Key Points
- AI companies paid over $4B for licensed data in 2025, with disclosed single deals ranging from roughly $5M to $250M
- One pharma-sector deal reportedly ran ~$200M for access to a multimodal oncology dataset spanning millions of patients
- Clinical datasets command top dollar because they are multimodal, deeply consented, and nearly impossible to reconstruct from the open web
- The price premium is a direct function of provenance: regulated, consented, richly-structured data is the scarce asset