Audio & Speech
Jul 28, 2026

The Gist
Fish Audio has raised $52 million to develop AI voice models as the sector experiences rapid growth, with over 30% of new podcasts now AI-generated according to Listen Notes. Meanwhile, regulators are catching up: Japan's Justice Ministry is backing civil liability laws for unauthorized AI voice use, while startups like an ex-AWS scientist's new venture and Deepgram are advancing voice AI technology with improved cost efficiency and security features.
Today's Stories
- 1
Fish Audio raises $52M seed to build AI voice models
Palo Alto-based Fish Audio, which has over 8 million users and generates $21 million(約34億円) in annual recurring revenue, raised $52 million(約83億円) in a seed round led by Coreline Ventures and Capital Today on Tuesday. The startup has released five models in the past year—four speech generation models and one speech-to-text model—and operates a library of more than 15,000 natural language controls. AI voice generation is increasingly valuable for creators who need expressive synthetic voices and enterprises automating customer support and sales. Fish Audio's open-source approach (three of its models are open-sourced, with the latest S2.1 Pro available only via paid API) has attracted indie developers and video game designers, while organizations like HeyGen and Sanas use its enterprise APIs. The company faces a crowded market with competitors including ElevenLabs, WellSaid, and Cartesia.
Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model. The startup has also automated its voice takedown process—creators can now remove unauthorized voice uploads in less than three minutes by submitting a voice sample or contract—addressing earlier concerns about consent.
- 2
Japan's Justice Ministry backs civil liability for unauthorized AI voice use
An expert committee of Japan's Justice Ministry approved a draft report Monday establishing that unauthorized use of public figures' voices through generative AI can be subject to civil liability under the right of publicity—a legal protection that allows celebrities to control the commercial value of their identities. Individuals can demand compensation or removal of infringing online posts. Voice actors and others have long called for clarification on what constitutes illegal voice use, as no Japanese court ruling had previously addressed these rights. The report specifies that even AI-generated voices that are not identical to originals but share similar voice quality and style may constitute infringement—a meaningful standard for victims of unauthorized AI voice mimicry used for profit.
The Justice Ministry will release a final report as early as August based on expert input. The draft also clarifies that generating sexual images from an actor's portrait using AI can infringe on both the right of publicity and portrait rights, addressing a broader class of generative AI harms.
- 3
30%+ of new podcasts are AI-generated, Listen Notes finds
Listen Notes, a podcast search database, reports that over 30% of newly created podcasts consist of AI-generated audio rather than human-created content. The platform distinguishes its own database of 3,793,012 podcasts by filtering out AI-slop, deleted podcasts, and non-audio content that other services include in their counts. The proliferation of AI-generated podcasts dilutes the quality and discoverability of genuine human-created audio content. For listeners and creators, this means the true volume of authentic podcasts is substantially lower than headline numbers suggest—a distinction that affects how the industry measures growth and success.
Listen Notes uses 24/7 automated scripts and human moderators to maintain its database of genuine podcasts, setting a standard for quality curation in an industry where inflated podcast counts are now common practice.
- 4
Deepgram integrates AWS IAM temporary delegation for SageMaker support access
Deepgram has integrated IAM temporary delegation, a new AWS IAM capability, into its support workflow for customers running Deepgram speech AI models on Amazon SageMaker AI. The integration allows Deepgram engineers to request time-limited, scoped access to customer endpoints and logs directly from the support ticketing system, with customers approving requests in their own IAM console. This replaces the previous model of long-lived cross-account IAM roles, which required recurring provisioning and audit conversations. Deepgram has reduced the time for initial investigation on a SageMaker AI support ticket from days to minutes, since customers can now approve access in their IAM console rather than scheduling screen-shares. Every delegated API call is tagged in AWS CloudTrail with Deepgram's partner account ID for full auditability.
Credentials issued through the integration expire automatically after twelve hours, and customers can revoke access at any time before expiration. The integration requires an active Deepgram support contract, an enabled AWS CloudTrail trail in the region where the SageMaker AI endpoint runs, and AWS infrastructure charges for SageMaker AI endpoint hosting, Amazon CloudWatch logs, AWS CloudTrail, and networking (Deepgram offers a 14-day trial at no additional cost for its models, but AWS infrastructure costs apply from the start of deployment).
- 5
Ex-AWS scientist targets OpenAI, Meta with cheaper voice AI startup
Alex Smola, former distinguished scientist at Amazon, founded Boson AI to release Higgs RealTime, a speech-to-speech model designed to be significantly cheaper than competitors' voice systems. The Santa Clara startup has raised $70 million(約110億円) and counts Chinese entrepreneur Su Hua and Singapore-based Temasek's venture arm among its backers. Voice AI has become a primary battleground for OpenAI, Meta, and other leaders building full-duplex systems (technology enabling natural back-and-forth conversation where users can interrupt mid-sentence). Boson claims its models cost one-tenth as much as rivals, potentially reshaping the economics of deploying voice agents for customer support, sales, and other enterprise applications. Smola is targeting finance, telecommunications, healthcare, and insurance clients.
Boson faces well-funded rivals—Microsoft recently announced a new voice model iteration, OpenAI has launched GPT-Live (a family of full-duplex audio models), and Meta released Muse Spark optimized for voice across wearables. The key technical hurdle remains latency; while one-second delays are acceptable in text chat, they feel disruptive in spoken conversation.
- 6
Open-source profiler for voice agents launches with cost tracking
A developer has released VoiceGateway, an open-source tool that profiles voice agents (built on platforms like LiveKit or Pipecat) by monitoring STT (speech-to-text), LLM, and TTS (text-to-speech) calls outside the audio path, tracking latency, model IDs, and cost. Self-hosted voice agent builders currently lack built-in visibility into model-layer performance and spending. VoiceGateway fills that gap with a single function call and includes cost reconciliation against provider invoices—useful for teams optimizing inference costs and latency.
The tool is MIT licensed, runs in Docker, stores data locally in SQLite, and uses DuckDB for analytics. The creator is actively soliciting feedback on what monitoring features are missing for self-hosted AI workloads.
What to Watch
Watch for Fish Audio's audio understanding and speech-to-speech models launching this year, along with their streamlined voice protection system that lets creators reclaim unauthorized voice uses in minutes—a meaningful step toward addressing consent in generative audio. Meanwhile, real-time voice AI remains a competitive battleground where latency is the make-or-break factor, with major players like Microsoft, OpenAI, and Meta all racing to deliver seamless spoken interactions, so expect rapid advances in how natural and responsive these conversations feel.
Sources
- Fish Audio raises $52M seed to build AI voice models for creators and enterprises
- Panel backs civil liability for unauthorized AI use of public figures’ voices
- 30%+ new podcasts are AI-slop
- Deepgram enhances Amazon SageMaker AI support with AWS IAM Temporary Delegation
- How an ex-AWS scientist plans to take on OpenAI and Meta in voice AI models
- Open Source Profiler for Voice Agents - Understanding from inside
- What I learned while calculating the real cost of a Voice AI call
- ARIA – Voice-native 3D spatial AI SoC with governed autonomy (BSL 1.1)
- OpenAI’s new voice mode makes it to the ChatGPT desktop app
- Claude's voice mode now runs on Anthropic's most capable models across all platforms
Share this with a friend
Send today's roundup to anyone who wants to keep up.
Get daily AI news free with AIToday
200+ AI sources, summarized in 1 minute. Email / LINE / Slack.
Sign up free