
The AI audio startup launched its S2.1 Pro voice model on July 29, saying it can clone a voice from a 5-second sample and is already being used in production by several companies.
Fish Audio said it completed a $52 million seed round without naming investors and launched its new voice generation model, S2.1 Pro, on July 29. The company says it has more than 8 million users across its open-source and hosted text-to-speech products and has reached $21 million in annual recurring revenue. Fish Audio builds voice synthesis models that convert written text into natural-sounding speech through downloadable software and a hosted API for developers. It says its Fish-Speech and S2 model families support more than 80 languages, while the earlier S2 and S2 Pro models published in March 2026 posted benchmark word error rates of 0.54% in Chinese and 0.99% in English. Fish Audio said S2.1 Pro can clone a voice from a 5-second sample, runs about twice as fast as Cartesia, costs about one-sixth of ElevenLabs, and is already used in production by HeyGen, LiveKit, Retell, Sanas, and OpenArt. The company has also said it reached a $500 million valuation after a separate $30 million venture financing round in 2025.