Voice, Speech & Realtime AI

Fish Audio closes $52M seed round to expand voice generation models

Source: The AI Insider · Aug 5, 2026

A $52 million seed round is large by almost any standard, and Fish Audio's is a signal of just how much capital is chasing the voice-generation layer of the AI stack right now. The round was led by Coreline Ventures and Capital Today, with a long list of additional participants including 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0 — a wide syndicate for a seed-stage round, suggesting strong investor demand to get into the deal rather than a single lead carrying the round alone.

Fish Audio, founded by former Nvidia researcher Shijia Liao, has already shipped five AI models: four for speech generation and one for speech-to-text. The company's customer base splits across two distinct use cases — some clients want realistic synthetic voices for avatars and character work, while others prioritize natural, low-latency speech specifically for voice-agent applications. Serving both ends of that spectrum from one model family is a meaningful technical claim, since realism-for-avatars and latency-for-agents often pull model design in different directions.

On traction, the company reports more than 8 million users spread across its open-source and hosted platform offerings, with $21 million in annual recurring revenue. The open-source-plus-hosted dual distribution model is worth noting specifically — it's a go-to-market pattern that builds a large, low-friction user base through the free/open tier while converting a subset into paying hosted customers, and it's becoming increasingly common among AI infrastructure companies competing for developer mindshare before monetizing.

For the training-data licensing market broadly, voice-generation companies like Fish Audio are both potential customers (they need diverse, high-quality voice data to keep improving model realism and coverage across languages/accents) and, at scale, potential sources of exactly the kind of synthetic voice content that raises provenance questions if it ever gets relicensed downstream without clear labeling. Worth tracking as the voice-AI funding environment continues to run hot across both the generation side (Fish Audio) and the real-time interaction side (Smallest.ai, also in today's roundup).

Key Points

  • $52M seed round led by Coreline Ventures and Capital Today, with 359 Capital, Parable, and others participating
  • Five AI models released — four for speech generation, one for speech-to-text
  • 8M+ users across open-source and hosted platforms; $21M ARR
  • Founded by Shijia Liao, a former Nvidia researcher