Voice, Speech & Realtime AI
Smallest.ai raises $13M Series A, tops $21M total, to build parallel-processing voice AI
Source: The AI Insider · Aug 10, 2026
Most voice-AI systems today still process conversation sequentially: listen, then think, then respond, one step after another, which is where a lot of the awkward latency in AI voice interactions comes from. San Francisco-based Smallest.ai is betting its Series A on a different architecture entirely. The company just raised $13 million in Series A funding — led by Seligman Ventures, with Sierra Ventures and 3one4 Capital participating — bringing its total funding past $21 million, to build out Hydra, a speech-to-speech model designed around parallel rather than sequential processing.
The pitch, in the words of CEO Sudarshan Kamath: "Humans don't wait for someone to finish speaking before they begin thinking. We listen, think, and respond simultaneously." Hydra is built to mirror that — running listening, reasoning, action, and response generation concurrently instead of as discrete sequential steps, which is what the company says drives its latency numbers down into milliseconds rather than the seconds that sequential pipelines typically produce.
The product line supporting this includes Pulse STT Pro, a speech-to-text layer covering 38 languages with built-in speaker diarization and PII redaction — notable for us, since automated PII redaction in voice data pipelines is directly adjacent to the kind of privacy-safe data handling our own business is built around — and Lightning V3.1, which the company positions as a top-ranked voice model on the combined dimensions of speed, quality, and cost efficiency.
On traction, Smallest.ai reports customers seeing support-cost reductions of up to 80% and agent productivity gains of up to 10x, with a client roster that includes RingCentral, Truecaller, Readymode, and several others, run by a lean team of roughly 60 employees. The target verticals — financial services, healthcare, contact centers, and business process outsourcing — are exactly the high-volume, latency-sensitive, compliance-conscious environments where architecture choices like parallel processing translate directly into measurable cost savings rather than just a technical curiosity. Worth tracking as both a potential buyer of licensed voice-interaction data and a technical signal of where real-time voice AI infrastructure is heading architecturally.
Key Points
- $13M Series A led by Seligman Ventures, with Sierra Ventures and 3one4 Capital participating; $21M total raised
- Hydra: a speech-to-speech model processing listening, reasoning, action, and response in parallel rather than sequentially
- Pulse STT Pro supports 38 languages with speaker diarization and PII redaction; Lightning V3.1 ranks highly on speed/quality/cost
- Customers report up to 80% support-cost reduction and up to 10x agent productivity gains
- Clients include RingCentral, Truecaller, Readymode; ~60 employees