August 26, 2026 • 41 min
Fish Audio CEO Rissa Cao on TTS and Enterprise Voice AI

CX Analyst & Thought Leader
August 26, 2026
Voice AI is becoming more natural, expressive, multilingual, and fast enough for real-time applications. Fish Audio is building the underlying models for creators, developers, and enterprises across text-to-speech, speech-to-text, voice cloning, and streaming.
In this Executive Spotlight interview, Katherine Stone speaks with Fish Audio CEO and co-founder Rissa Cao about the company’s evolution from creator-focused voice generation to enterprise voice AI, its S2.1 Pro model, voice cloning, pricing, partnerships, and future model roadmap.
What This Episode Covers
- What Fish Audio is and how its voice AI models work
- How Fish Audio evolved from an open-source creator project into an enterprise platform
- How the company plans to use its $52 million seed round
- What S2.1 Pro adds across emotion control, accuracy, multilingual support, and latency
- How creator feedback helped shape Fish Audio’s enterprise offering
- How Fish Audio approaches voice cloning consent, verification, and revenue sharing
- How pricing and partnerships work across creators, developers, and enterprises
- Why Fish Audio is focused on the model layer and what comes next in audio understanding and full-duplex voice AI
Fish Audio is currently focused on building the underlying voice models rather than a complete voice-agent orchestration platform. As Rissa explains, the roadmap includes audio-language understanding and full-duplex audio-to-audio models while continuing to improve naturalness, accuracy, latency, and multilingual performance.