Voice, Speech & Realtime AI
Gradium's $100M seed, with Nvidia aboard, shows real-time voice is now a data problem
Source: Sifted · Sep 30, 2026
Gradium, a Paris startup spun out of the French lab Kyutai, has extended its seed round past $100 million, with Nvidia joining an investor list that already included FirstMark, Eurazeo, Xavier Niel, Rodolphe Saadé, and Eric Schmidt. Co-founded by Neil Zeghidour — a veteran of Google Brain, DeepMind, and FAIR — the company is building ultra-low-latency voice models that aim to remove the awkward pause that still gives most AI voice agents away.
A seed round at this scale, with chip and sovereign-adjacent money behind it, is a signal about where the hard problem now sits. Getting a voice model to feel genuinely real-time and natural isn't only a latency-engineering challenge; it depends on learning from large volumes of authentic human speech — how people actually take turns, interrupt, hesitate, and recover mid-sentence.
That is the thread back to our thesis. The messy, unscripted texture of real conversation is exactly what a model needs to stop sounding synthetic, and exactly what you cannot manufacture from clean studio reads. As more capital flows into natural real-time voice, the premium rises on consented recordings of real humans in genuine, unscripted dialogue — the raw material that makes a voice model sound human rather than almost-human.
Key Points
- Paris-based Gradium, spun out of the Kyutai lab, extended its seed to $100M+ with Nvidia joining
- Co-founded by Neil Zeghidour (ex-Google Brain, DeepMind, FAIR); building ultra-low-latency voice models
- Backers include FirstMark, Eurazeo, Xavier Niel, Rodolphe Saadé, and Eric Schmidt
- Natural, instant conversational voice depends on large volumes of real human speech and turn-taking