August 26, 202641 min

Fish Audio CEO Rissa Cao on TTS and Enterprise Voice AI

Author
Katherine Stone's profile picture

CX Analyst & Thought Leader

August 26, 2026

Voice AI is becoming more natural, expressive, multilingual, and fast enough for real-time applications. Fish Audio is building the underlying models for creators, developers, and enterprises across text-to-speech, speech-to-text, voice cloning, and streaming.

In this Executive Spotlight interview, Katherine Stone speaks with Fish Audio CEO and co-founder Rissa Cao about the company’s evolution from creator-focused voice generation to enterprise voice AI, its S2.1 Pro model, voice cloning, pricing, partnerships, and future model roadmap.

What This Episode Covers

  • What Fish Audio is and how its voice AI models work
  • How Fish Audio evolved from an open-source creator project into an enterprise platform
  • How the company plans to use its $52 million seed round
  • What S2.1 Pro adds across emotion control, accuracy, multilingual support, and latency
  • How creator feedback helped shape Fish Audio’s enterprise offering
  • How Fish Audio approaches voice cloning consent, verification, and revenue sharing
  • How pricing and partnerships work across creators, developers, and enterprises
  • Why Fish Audio is focused on the model layer and what comes next in audio understanding and full-duplex voice AI

Fish Audio is currently focused on building the underlying voice models rather than a complete voice-agent orchestration platform. As Rissa explains, the roadmap includes audio-language understanding and full-duplex audio-to-audio models while continuing to improve naturalness, accuracy, latency, and multilingual performance.

Loading articles...

Stay updated with cx news

Subscribe to our newsletter for the latest insights and updates in the CX industry.

By subscribing, you consent to our Privacy Policy and receive updates.