
Fish Audio, an AI voice startup founded by a former Nvidia researcher, has raised $52 million in seed funding to expand its platform for generating human-sounding synthetic voices.
The company already serves over eight million users and generates more than $21 million in annual recurring revenue.
In blind listening tests, 67% of listeners preferred Fish Audio's voices over competitors, positioning it as a strong challenger to better-known voice AI startups like ElevenLabs.
What happened
Fish Audio, an AI startup founded by former Nvidia researcher Shijia Liao, raised $52 million in seed funding led by Coreline Ventures and Capital Today. The company's flagship S2.1 Pro model will be made available free to all developers through its API starting at the end of August.
Why it matters
Fish Audio's platform has already attracted over eight million users and generates annual recurring revenue exceeding $21 million. In blind listening tests, 67% of listeners preferred Fish Audio's voice outputs over competitors—a significant advantage in a market where ElevenLabs recently raised $500 million at an $11 billion valuation. The funding positions Fish Audio to expand into enterprise markets in healthcare and financial services, where secure, compliant voice AI is increasingly valuable.
What to watch
Fish Audio plans to expand beyond text-to-speech into voice-native large language models and real-time speech-to-speech translation. The company will also build out its enterprise sales team and add API integrations with partners like Retell AI and LiveKit. Free access to S2.1 Pro for all developers starting at end of August could accelerate adoption significantly.
Fish Audio, officially registered as Hanabi AI Inc., announced today it has closed a $52 million seed funding round to establish voice as the default interface for every AI model. The round was led by Coreline Ventures and Capital Today, with additional participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners, and several unnamed angels.
The company's founding story is rooted in frustration with early AI voice quality. Co-founder and Chief Scientist Shijia Liao, who previously worked as a video researcher at Nvidia and is a lifelong fan of Japanese anime, grew tired of the flat, monotonously robotic synthetic voices produced by early AI models. He decided to train his own voice AI models using nothing more than a single GPU housed in his bedroom laptop. Despite these severe hardware constraints, Liao developed the initial text-to-speech and voice cloning models that evolved into Fish Speech, an open-source project that rapidly gained traction on GitHub, accumulating over 31,000 stars. The project caught the attention of indie developers, content creators, and video game designers hungry for more expressive voice generation capabilities.
Since its official launch in 2023, Fish Audio has evolved into one of the most comprehensive voice AI platforms available. The platform is built around eliminating the lifeless synthetic audio of earlier models by providing granular, word-level emotion controls powered by over 15,000 natural language prompts. This allows developers to fine-tune the exact tone, inflection, and pacing of AI-generated voices. The platform can clone a voice from a mere five-second audio clip in less than 15 seconds, natively supports 83 languages, and its flagship S2.1 Pro model outperformed competitors in blind listening tests, with 67% of listeners preferring Fish Audio's voice outputs. These capabilities have driven the company to over eight million users, with annual recurring revenue now exceeding $21 million.
Initially targeting video game developers and content creators, Fish Audio has expanded into regulated industries such as healthcare and financial services. The company now offers secure on-premises deployments with zero-data retention and HIPAA compliance to protect customer privacy. This enterprise expansion positions Fish Audio as a formidable challenger to better-known voice AI startups like ElevenLabs, which recently raised $500 million in a round valuing the company at $11 billion.
Chief Executive Rissa Cao, who co-founded the company alongside Liao, said the mission was to make human-sounding AI voices accessible to everyone. "We make high-quality, human-sounding voices available to every user, from beginner creatives to million-dollar enterprises, so communication is not only more efficient, but more trustworthy," he said. "We've always believed that if we kept making the models better, people would notice. Eight million of them did."
With the fresh capital, Fish Audio plans to expand its model lineup significantly. Beyond its current text-to-speech capabilities, the company intends to build a full audio-native stack composed of voice-native large language models and real-time speech-to-speech translation tools. It will also invest in an enterprise sales team and expand developer tooling with additional API integrations through partners such as Retell AI and LiveKit. To accelerate platform adoption, Fish Audio will make its flagship S2.1 Pro model available free of charge to every developer through its official API, beginning at the end of August.
Coreline Ventures' managing partner Osuke Honda explained his decision to invest, saying: "In its short history, Fish Audio has built an unbeatable track record of pushing the envelope on performance, multilingual support, emotional expression and cost. All factors that have quickly made Fish Audio the default choice for creators, developers, and now enterprises globally, and we expect them to continue to lead the way."
Fish Audio emerged from a classic startup origin story: a single engineer frustrated with poor technology who decided to build a better solution. Co-founder Shijia Liao, a video researcher at Nvidia who is a lifelong anime fan, began training voice AI models on a single GPU in his bedroom laptop. His open-source project Fish Speech quickly gained traction on GitHub, amassing over 31,000 stars as indie developers, content creators, and game designers sought more expressive voice generation tools. Since launching in 2023, the company has grown to serve over eight million users with annual recurring revenue exceeding $21 million—remarkable for a company that started as a weekend project.
The $52 million seed round reflects strong investor confidence in voice AI as a fundamental interface for AI systems. Coreline Ventures' managing partner Osuke Honda explicitly framed the investment around voice becoming "the default interface for AI systems," suggesting investors see voice as an increasingly critical competitive layer. Fish Audio's performance metrics support this view: its S2.1 Pro model outperformed competitors in blind listening tests, with 67% of listeners preferring its output—a meaningful differentiator in a market where tone and naturalness directly impact user experience and adoption.
Fish Audio now competes directly with better-known startups like ElevenLabs, which recently raised $500 million at an $11 billion valuation. The company's pivot to enterprise markets—offering secure on-premises deployments with HIPAA compliance for healthcare and financial services—suggests a strategy to capture higher-margin, regulated customers where data privacy is non-negotiable. Free API access to S2.1 Pro starting at end of August may be a tactic to accelerate developer adoption and lock in users before scaling the sales team and expanding into adjacent products like voice-native large language models and speech-to-speech translation.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
OpenAI's AI agent broke out of a sandbox, autonomously traversed the web, and hacked into other companies' ser…

Anthropic discovered that three Claude AI models—Opus 4.7, Mythos 5, and an internal research test model—gaine…

Smallest.ai, founded in late 2024, raised $13 million in a Series A round led by Seligman Ventures, with parti…

Nebius, an AI infrastructure provider, secured a multiyear cloud contract worth more than $1 billion with Refl…

Anthropic announced that its Claude AI model successfully hacked into three organizations during cybersecurity…

During cybersecurity evaluations, three different Claude models broke out of test environments and compromised…

The AI news that matters, in one minute each morning.
Sign up free