AITodayYour daily AI briefing

Audio & Speech

Jul 29, 2026

Audio & Speech

The Gist

Google upgraded its Lyria music generation tool to allow seamless editing of song sections, while OpenAI introduced GPT-Transcribe, a new speech-to-text model that, despite showing improvements, still lags behind competitors like ElevenLabs and Google in accuracy. In other developments, Fish Audio secured $52 million in seed funding to advance AI voice technology, a private voice journal app called Echologue launched, and Japan's Justice Ministry signaled support for holding people civilly liable for unauthorized use of AI-generated voices.

Today's Stories

  1. 1

    Google's Lyria 3.5 lets users edit song sections without restarting

    Google released Lyria 3.5, a music generation model that produces more natural-sounding melodies, better lyrics, and more realistic vocals. The model is now available through Google Flow Music and includes a new feature called 'Selective Section Painting' that allows users to edit specific parts of a track or turn short melodies into full songs without starting over. The ability to edit individual sections without restarting the generation process makes music creation faster and more flexible for users. Users can also fine-tune tempo and duration of specific elements like vocals, drums, and bass, giving them more creative control over the final output.

    Tracks can run from 30 seconds to 3 minutes in length. Google has not provided details about the training data for Lyria 3.5, despite having disclosed that Lyria 3 was trained on materials YouTube and Google had the right to use under their terms of service, partner agreements, and applicable law.

  2. 2

    OpenAI's GPT Transcribe improves 0.7 points but trails ElevenLabs, Google, Mistral

    OpenAI released GPT Transcribe and GPT Live Transcribe, two speech recognition models via its API. GPT Transcribe processes pre-recorded audio about 34 times faster than real time, while GPT Live Transcribe handles real-time streaming with low latency. The new GPT Transcribe achieves a 3.31 percent word error rate—a 0.7 percentage point improvement over GPT-4o Transcribe—at a 25 percent price cut to $0.0045 per minute of audio. OpenAI's transcription models now rank #4 on the AA-WER benchmark, behind ElevenLabs Scribe v2 (2.3%), Google's Gemini 3 Pro (2.9%), and Mistral's Voxtral Small (3%). The price drop positions OpenAI competitively against Mistral's Voxtral Transcribe V2, which starts at $0.003 per minute—suggesting OpenAI is responding to margin pressure in a crowded speech recognition market.

    The models accept text as transcription context, keywords, and multiple input languages, and they complement OpenAI's recently announced Realtime model generation, which includes GPT-Realtime-Whisper. Full technical details are available in OpenAI's Transcription Guide.

  3. 3

    Developer releases Echologue, private AI voice journal app

    A developer created Echologue, a voice-first journaling app that runs locally on the user's device with AI-powered features including automatic tagging, meaning extraction, and semantic search using embeddings stored on-device. The app addresses a common journaling friction point — the need to write — by letting users speak casually for a minute when something happens, then query their journal with natural-language questions (e.g. "remind me highlights from the last 3 months" or "how was I feeling during my trip in Spain"). All processing happens on-device with zero data retention, so users remain fully anonymous and the creator has no visibility into usage.

    The app supports export to an LLM, allowing users to process their journal entries with external AI tools if they choose, while maintaining the option to keep everything private and local by default.

  4. 4

    OpenAI launches GPT-Transcribe, speech-to-text model for audio files and live sessions

    OpenAI announced GPT Transcribe, a speech-to-text model that processes completed audio files, streamed file transcripts, and committed turns in Realtime sessions over WebSocket. The model supports unstructured context, keyword hints, and multiple language hints to improve transcription of domain terms, multilingual audio, and code-switching. The model's ability to handle keyword hints and language hints means it can be tuned for specialized vocabularies and multilingual contexts—a capability that could reduce transcription errors in fields like medicine, law, and software development where precise terminology matters. Support for Realtime sessions over WebSocket suggests real-time transcription use cases alongside batch processing.

    The announcement does not include pricing, availability date, API access details, or region restrictions. No information is provided about how GPT Transcribe compares to existing transcription services or whether a second model (GPT-live-transcribe) mentioned in the title will be detailed separately.

  5. 5

    Fish Audio raises $52M seed to build AI voice models

    Palo Alto-based Fish Audio, which has over 8 million users and generates $21 million(約34億円) in annual recurring revenue, raised $52 million(約83億円) in a seed round led by Coreline Ventures and Capital Today on Tuesday. The startup has released five models in the past year—four speech generation models and one speech-to-text model—and operates a library of more than 15,000 natural language controls. AI voice generation is increasingly valuable for creators who need expressive synthetic voices and enterprises automating customer support and sales. Fish Audio's open-source approach (three of its models are open-sourced, with the latest S2.1 Pro available only via paid API) has attracted indie developers and video game designers, while organizations like HeyGen and Sanas use its enterprise APIs. The company faces a crowded market with competitors including ElevenLabs, WellSaid, and Cartesia.

    Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model. The startup has also automated its voice takedown process—creators can now remove unauthorized voice uploads in less than three minutes by submitting a voice sample or contract—addressing earlier concerns about consent.

  6. 6

    Japan's Justice Ministry backs civil liability for unauthorized AI voice use

    An expert committee of Japan's Justice Ministry approved a draft report Monday establishing that unauthorized use of public figures' voices through generative AI can be subject to civil liability under the right of publicity—a legal protection that allows celebrities to control the commercial value of their identities. Individuals can demand compensation or removal of infringing online posts. Voice actors and others have long called for clarification on what constitutes illegal voice use, as no Japanese court ruling had previously addressed these rights. The report specifies that even AI-generated voices that are not identical to originals but share similar voice quality and style may constitute infringement—a meaningful standard for victims of unauthorized AI voice mimicry used for profit.

    The Justice Ministry will release a final report as early as August based on expert input. The draft also clarifies that generating sexual images from an actor's portrait using AI can infringe on both the right of publicity and portrait rights, addressing a broader class of generative AI harms.

What to Watch

Watch for Google's disclosure of Lyria 3.5's training data sources, as questions about artist consent and copyright remain unresolved, while simultaneously monitoring OpenAI's pricing and rollout plans for GPT Transcribe and Fish Audio's upcoming audio understanding model to see how the competitive landscape for AI-powered speech and transcription tools develops. Additionally, the Justice Ministry's final report due in August could set important legal precedents for how AI-generated content infringement cases are handled globally.

Sources

Share this with a friend

Send today's roundup to anyone who wants to keep up.

Get daily AI news free with AIToday

200+ AI sources, summarized in 1 minute. Email / LINE / Slack.

Sign up free