Fish Audio raises $52 million seed round after reaching $21 million ARR

Fish Audio raises $52 million seed round after reaching $21 million ARR

The AI audio startup launched its S2.1 Pro voice model on July 29, saying it can clone a voice from a 5-second sample and is already being used in production by several companies.

Fact Check
The core claims are confirmed by the official Fish Audio X announcement and the TechCrunch primary article: $52M seed round, $21M ARR, and public launch of S2.1 Pro capable of cloning a voice from a 5-second sample, already used in production by companies including HeyGen, LiveKit, and Retell. The only inaccuracy is the launch date: both the official post and TechCrunch date the announcement to July 28, 2026, not July 29. Some secondary aggregators (CryptoBriefing) rounded the figure to $50M, but the primary and official sources confirm $52M. Because the funding, ARR, model capability, and production usage are all accurate and only the date is off by one day, the claim is likely true.
Summary

Fish Audio said it completed a $52 million seed round without naming investors and launched its new voice generation model, S2.1 Pro, on July 29. The company says it has more than 8 million users across its open-source and hosted text-to-speech products and has reached $21 million in annual recurring revenue. Fish Audio builds voice synthesis models that convert written text into natural-sounding speech through downloadable software and a hosted API for developers. It says its Fish-Speech and S2 model families support more than 80 languages, while the earlier S2 and S2 Pro models published in March 2026 posted benchmark word error rates of 0.54% in Chinese and 0.99% in English. Fish Audio said S2.1 Pro can clone a voice from a 5-second sample, runs about twice as fast as Cartesia, costs about one-sixth of ElevenLabs, and is already used in production by HeyGen, LiveKit, Retell, Sanas, and OpenArt. The company has also said it reached a $500 million valuation after a separate $30 million venture financing round in 2025.

Terms & Concepts
  • text-to-speech: Technology that converts written text into spoken audio using synthetic voices
  • hosted API: A cloud-based interface that lets developers integrate a company's voice software into their own applications
  • voice cloning: A process that recreates a person's voice from a short audio sample for synthetic speech generation