AIToday

Fish Audio raises $52M seed to build AI voice models

TechCrunch AI3h agoSend on LINE
Fish Audio raises $52M seed to build AI voice models

Key takeaway

Fish Audio, a Palo Alto startup founded by former NVIDIA researcher Shijia Liao, has raised $52 million(約83億円) in a seed round to expand its AI voice generation platform, which already serves over 8 million users and generates $21 million(約34億円) in annual recurring revenue. The company offers customizable voice models for both creators—who need expressive synthetic voices for content—and enterprises automating customer support, with clients including HeyGen and Sanas. Investors highlight Fish Audio's ability to build state-of-the-art models with fine-grained developer controls and cost-efficient training as competitive strengths in a crowded market.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Palo Alto-based Fish Audio, which has over 8 million users and generates $21 million(約34億円) in annual recurring revenue, raised $52 million(約83億円) in a seed round led by Coreline Ventures and Capital Today on Tuesday. The startup has released five models in the past year—four speech generation models and one speech-to-text model—and operates a library of more than 15,000 natural language controls.

  • Why it matters

    AI voice generation is increasingly valuable for creators who need expressive synthetic voices and enterprises automating customer support and sales. Fish Audio's open-source approach (three of its models are open-sourced, with the latest S2.1 Pro available only via paid API) has attracted indie developers and video game designers, while organizations like HeyGen and Sanas use its enterprise APIs. The company faces a crowded market with competitors including ElevenLabs, WellSaid, and Cartesia.

  • What to watch

    Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model. The startup has also automated its voice takedown process—creators can now remove unauthorized voice uploads in less than three minutes by submitting a voice sample or contract—addressing earlier concerns about consent.

In Depth

Fish Audio, a Palo Alto-based startup founded by former NVIDIA researcher Shijia Liao, announced on Tuesday that it has raised $52 million(約83億円) in a seed round led by Coreline Ventures and Capital Today. The funding also includes participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. The company already serves over 8 million users and generates $21 million(約34億円) in annual recurring revenue.

Liao started Fish Audio as a small project after becoming frustrated by non-expressive synthetic voices available on the market. He trained a voice generation model on a single GPU and open-sourced it; the Fish Speech repository on GitHub now has more than 31,000 stars and is used by indie developers, video game designers, and creators. Fish Audio operates a library of more than 15,000 natural language controls to address diverse use cases. Over the past year, the company has launched five models: four speech generation models and one speech-to-text model. Three of the speech generation models have been open-sourced, while the latest S2.1 Pro model is available only through its paid API.

Fish Audio targets both creators and enterprises with distinct offerings. For creators and teams, it offers paid monthly plans that unlock a set number of generation minutes plus voice cloning features. For enterprises, the company provides a separate version of its APIs and platform; organizations including HeyGen (which uses Fish Audio's voices to power AI avatars) and Sanas (a voice agent company) are already customers. CEO and co-founder Rissa Cao explained that different users have different preferences: "Companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls."

The startup built its voice library by asking users to submit their own voices for training its models and compensating them if their voices are used. This approach triggered controversy a few months ago when some creators alleged that their voices were uploaded to Fish Audio without their consent. Although the startup had a DMCA takedown process in place, the removals took a long time. Cao announced that Fish Audio has now automated the takedown process: creators can submit a short voice sample or a contract to prove ownership, and their voice will be removed from the platform in less than three minutes. However, this retroactive approach does not prevent unauthorized uploads from being used until the creator discovers and reports them.

Osuke Honda, a partner at Coreline Ventures, stated that a community-driven model requires trust: "A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially."

Cao explained that when Fish Audio was offering only its open-source product with creator-focused plans, it ran efficiently without external capital. However, the company wanted to develop more advanced models and accommodate enterprise customers as investor interest ramped up, which led to seeking funding. Looking ahead, Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model. The speech generation market is competitive, with companies including ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp competing for creator and enterprise budgets. Rico Mallozzi, a partner at 359 Capital, highlighted Fish Audio's competitive advantage: "I think what they've been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices."

Context & Analysis

Fish Audio emerged from founder Shijia Liao's frustration with non-expressive synthetic voices available on the market. Building a voice generation model on a single GPU as a small project, Liao open-sourced the work, creating the Fish Speech repository on GitHub, which now has more than 31,000 stars. This community-first approach attracted indie developers, video game designers, and creators, eventually scaling to over 8 million users and $21 million(約34億円) in annual recurring revenue.

The startup's library of more than 15,000 natural language controls addresses distinct market needs: creators require expressive voices for narrative content, while enterprises need steerable models for automated customer support and sales operations. Fish Audio serves both segments through tiered offerings—monthly paid plans with generation minutes and voice cloning for creators, and enterprise APIs for organizations like HeyGen (which powers AI avatars) and Sanas.

Growing scale and enterprise demand prompted the $52 million(約83億円) seed round. Investors emphasized that Fish Audio's technical capability to produce state-of-the-art models with fine-grained developer controls and cost-efficient training differentiates it in a crowded field that includes ElevenLabs, WellSaid, Cartesia, and others. However, the community-driven model—built on user-submitted voices for training—created friction: some creators alleged unauthorized voice uploads. The automated takedown process (now under three minutes) addresses immediate removal but does not prevent initial upload without consent; it operates as a retroactive tool rather than a consent layer.

FAQ

How many people currently use Fish Audio?
More than 8 million people use Fish Audio's open-source or hosted versions of its models since the startup launched last year.
What does Fish Audio plan to build next?
The company plans to release an audio understanding model this year and is also building a speech-to-speech model.
How did Fish Audio handle unauthorized voice uploads?
Fish Audio has automated its takedown process so creators can submit a voice sample or contract to prove ownership, and their voice will be removed from the platform in less than three minutes.

Get the latest Audio & Speech news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime