Voice, Speech & Realtime AI
Real-time video agents arrive — full-duplex, face-to-face AI, and a new frontier for consented data
Source: Industry reporting · Oct 1, 2026
Video AI crossed another line this cycle. Beyond the steady march of text-to-video models — now generating coherent clips up to around thirty seconds — a preview of a real-time, full-duplex video-to-video system showed an AI holding a live, face-to-face conversation: seeing, responding, and taking turns in real time rather than rendering a clip after the fact. Conversational video, not just generated video, is now on the table.
The strategic read is that the interaction surface is getting richer fast. A live video agent combines everything the earlier modalities had separately — the words, the voice, the face, the timing, the back-and-forth of a real exchange — into a single stream. For anything built around human interaction, from interviewing to support to coaching, that is the most information-dense channel yet, and the one that feels most like talking to a person.
It is also, for exactly the same reason, the most sensitive. A live face-to-face exchange captures biometric-grade signal — expression, gaze, vocal affect, reaction timing — all at once. Collected carelessly, that is not a richer dataset; it is a richer liability, the kind that draws regulators and lawsuits. The value and the risk scale together, and they scale steeply.
Which puts the same requirement under this frontier that sits under all the others, only more so. Real-time conversational video is worth building toward precisely because the signal is so rich — but it is only an asset if every bit of it is captured with informed consent and a clean record of permitted use. The teams that treat consent and provenance as the foundation of a video-agent product, not a patch applied later, are the ones who will get to use the richest interaction data there is. Everyone else is building the most expensive liability in the market.
Key Points
- A real-time, full-duplex video-to-video model preview lets an AI hold a live face-to-face conversation, not just generate clips
- Alongside it, text-to-video models now stretch to ~30-second coherent clips, pushing generation toward production use
- Live conversational video is the richest interaction modality yet — face, voice, timing, and turn-taking together
- That richness is only usable if captured with consent and provenance; otherwise it is the biggest liability yet