Voice, Speech & Realtime AI
The enterprise voice-AI stack keeps consolidating
Source: Speechmatics · Sep 25, 2026
The voice-AI layer is consolidating into a smaller set of full-stack enterprise providers. Speechmatics shipped Agent STT on its new Linden 1 model, claiming a 1.05% semantic error rate, 369-millisecond finalization, support for 55-plus languages, and pricing around $0.30 an hour. OVH acquired the Paris speech-to-text startup Gladia, which brings 300,000-plus developers and EU data residency. And SoundHound AI is acquiring LivePerson to fuse voice and messaging for contact centers, targeting $350–400 million in revenue by 2027.
Two forces are visible at once: accuracy and latency are improving to the point where real-time voice agents are genuinely production-ready, and ownership is concentrating as the bigger players absorb specialists to offer an end-to-end stack. Cheap, fast, accurate transcription is becoming a commodity input.
As voice agents handle more real conversations, the volume of recorded, sensitive human dialogue flowing through these systems grows sharply — and so does the value of, and responsibility around, that data. The infrastructure to capture and transcribe conversation is consolidating and cheapening; what stays scarce is conversation that has been collected with consent and cleared for reuse. Capture is solved; safe, licensed data is not.
Key Points
- Speechmatics shipped Agent STT on its new Linden 1 model: 1.05% semantic error rate, 369ms finalization, 55+ languages, ~$0.30/hr
- OVH acquired Paris speech-to-text startup Gladia (300K+ developers, EU data residency)
- SoundHound AI is acquiring LivePerson to combine voice and messaging for contact centers (targeting $350–400M revenue by 2027)
- Voice infrastructure is consolidating into a few full-stack enterprise providers