SORABITO and Eleven Labs have built an interactive voice AI system that diagnoses equipment faults from verbal…

Music streaming platform Deezer announced that AI-generated music now represents more than 50% of daily upload…

Sony Music Entertainment filed a new lawsuit in New York against AI music generator Udio on Monday, claiming i…

Artist 1010Benja released a song called "Semiramis' Dream" on his latest EP using Suno generative AI, but in a…

A hacking incident exposed that Suno, an AI music generator, scraped millions of songs and lyrics from online…

The AI music generator Suno was hacked via a supply chain attack that exposed source code showing the platform…

Spotify is rolling out a conversational AI interface for Premium subscribers, allowing them to control playbac…

Hugging Face introduced Real World VoiceEQ, a benchmark that evaluates more than 40 voice models across 15+ di…

Telnyx released an open-source Python application (102 lines) that uses Telnyx Call Control and Llama 3.3 70B…

AWS Partner ScienceSoft has built a HIPAA-compliant AI voice scheduling assistant using Amazon Nova Sonic and…

Apple released iOS 27 as a public beta, making the revamped "Siri AI" available to iPhone users for the first…

OpenAI introduced GPT-Live, a family of voice-optimized AI models powering ChatGPT's voice mode
OpenAI released an upgraded voice mode described as a step change, and Grok 4.5 became available

Gradium, a Paris-based startup building voice AI models, closed its seed round at $100 million(約160億円) total…

OpenAI is rolling out GPT-Live-1, a new voice model for ChatGPT that can listen and speak at the same time, re…

OpenAI released two new conversational voice models—GPT-Live-1 and GPT-Live-1 mini—that can speak and listen s…

Cohere released Cohere Transcribe Arabic, a 2-billion-parameter open-source model for Arabic speech recognitio…

Researchers found that speech recognition systems waste compute processing silence instead of speech

Revin, an AI voice and SMS platform, launched its agents across Vertex Service Partners' portfolio of 27 regio…

Hugging Face and Cerebras demonstrated a speech-to-speech AI pipeline that combines open-source models—Nvidia'…

Netflix is premiering Wonka's The Golden Ticket on September 23rd, a reality competition based on the fictiona…

Jamendo SA, a subsidiary of Winamp Group, filed a federal lawsuit in Massachusetts against Suno, Inc., allegin…

AI Voice Studio is now available, allowing creators to convert scripts into studio-quality voiceovers across 7…

Whissle released a containerized voice AI system that performs automatic speech recognition (ASR), text-to-spe…

Gemini 3.5 Live Translate is available now for developers through the Gemini Live API and Google AI Studio, as…

ElevenLabs and the UK's Department for Science, Innovation and Technology (DSIT) agreed to collaborate on thre…

At its Worldwide Developers Conference, Apple announced 'Siri AI'—an updated voice assistant coming in OS upda…

Audio-Interaction, created by researchers from China, Hong Kong, and Singapore, processes continuous audio str…

ElevenLabs demonstrated a voice-powered robot at a New York pop-up that took coffee orders and prepared drinks…

AethexAI raised $3 million in pre-seed funding led by 4DX Ventures, with participation from Enza Capital, Dorm…

Suno raised $400 million at a $5.4 billion valuation, double its valuation from seven months ago

Snowflake reported that AI accounts on its platform jumped from 9,100 to 13,600 in a single quarter, product r…

The AI voice generator market is set to achieve a compound annual growth rate (CAGR) of 30.7% over the forecas…

Users in the Suno subreddit describe a pattern of consuming primarily their own AI-generated music over tradit…

James Beswick leads the Stripe Developer Relations team and was previously a Developer Advocacy leader at AWS

Starting November 2025, Amazon SageMaker AI supports bidirectional streaming for real-time inference, allowing…

Five minutes each morning. That's the whole AI news habit.
The day's essentials from 200+ sources, delivered to Email, LINE, or Slack. Free, always.
30,000+ monthly readersStability AI unveiled Stable Audio 3.0, a family of four audio models

Stability AI released four new audio models under the Stability Audio 3.0 name: small SFX (459M parameters), s…

ElevenLabs is offering professors free access to its Pro tier and the ability to provide time-bound access to…

Thinking Machines Lab, founded in February 2025 by Mira Murati and other former OpenAI researchers, published…

Rivian is releasing its Rivian Assistant to all compatible Gen 1 and Gen 2 vehicle owners who subscribe to Con…

A Bi-LSTM gender classifier model (166K parameters, 0.64 MB) designed for real-time voice AI pipelines support…

Amazon Ring evaluated more than 40 AI voice vendors before selecting Vapi to route 100% of its inbound calls t…

Wispr Flow, a Bay Area-headquartered startup building AI-powered voice input software, began beta testing a Hi…

TypeWhisper 1.4 is the current release-candidate line for macOS, featuring system-wide dictation, file transcr…

The Meeting Agent is a script that records audio from microphone and speaker outputs using PulseAudio and FFmp…

OpenAI released three streaming audio models: GPT-Realtime-2 (a native speech-to-speech model for voice agents…

OpenAI shipped three new voice models: GPT-Realtime-2 (for reasoning and real-time conversation), GPT-Realtime…

Google rolled out Gemini 3.1 to Google Home smart speakers, initially available to early access users

ElevenLabs announced additional investors in its Series D fundraise first announced in February, including ins…

InterviewDen offers live voice and text mock interviews that listen, speak, and grade responses

AWS published a blog post exploring how to convert a traditional text agent into a conversational voice assist…

Arietta Voice combines local speech-to-text (Moonshine), text-to-speech (Kokoro), turn detection (Silero VAD a…
Researchers at University of Pennsylvania and Google released PAVO-Bench, a 50,000-turn voice interaction data…

VoiceGoat is a modular platform designed for security practitioners to practice exploiting voice-based AI syst…

VibeVoice-ASR, a speech-to-text model, is now part of a Transformers release and available directly through th…
Jon Seager, VP of engineering at Canonical, shared a blog post on Monday detailing plans to add AI features to…

ComfyUI, a startup building AI creation tools, closed a $30 million funding round that values the company at $…

A LessWrong author surveyed their own AI usage patterns and found themselves using AI assistance for hours eve…

TTS.ai is a new text-to-speech service that was shared on Hacker News

PrivaKit uses transformers.js, Whisper, and WebGPU to perform AI tasks like transcription, OCR, and image proc…

Sony Music filed a lawsuit against Udio, claiming the AI music startup used stream ripping to extract audio fr…

Gemini 3.1 Flash model adds integrated text-to-speech (TTS) functionality

Gemini 3.1 Flash TTS is now available across multiple Google products and services

Fully automated system using Suno's API generates new AI songs in different genres every few minutes, with lyr…

Suno is an AI music generation platform that raises questions about the future of music creation and industry…

ChatGPT's voice interface appears to be powered by a weaker AI model compared to the standard text-based ChatG…

Seeduplex represents ByteDance's advancement in conversational AI, allowing simultaneous speaking and listenin…

Self-supervised learning (SSL) models successfully capture tone information in their latent representations, b…

Contextual Earnings-22 dataset created to address gap between academic benchmarks and actual industrial speech…

ElevenLabs now offers on-premise deployment, allowing enterprises to run voice AI solutions within their own i…

Microsoft identifies improved voice understanding as a critical gap in its AI development roadmap

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
30,000+ monthly readers
Get Started FreeFree · takes 30 seconds · unsubscribe anytime