Audio & Speech
Jul 23, 2026

The Gist
OpenAI has expanded ChatGPT's voice capabilities to desktop with hands-free task automation, while Anthropic is broadening Claude's voice mode across its Opus and Sonnet models with integrations for Gmail, Slack, and Notion. Meanwhile, Black Forest Labs released Flux 3, a new generation model capable of creating images, videos, and audio natively in a single system.
Today's Stories
- 1
OpenAI brings ChatGPT Voice to desktop, hands-free task automation
OpenAI has released ChatGPT Voice on desktop and laptop computers, extending the GPT Live voice technology it launched on mobile earlier this month. The feature lets users speak to ChatGPT in a stream-of-consciousness manner to complete multiple tasks in the background while working on other things—such as checking a calendar, reviewing emails, or drafting documents—without using their hands. The desktop rollout targets software programmers and high-volume AI token users, a lucrative segment of the AI market. The feature integrates directly with Codex, OpenAI's programming product (which merged into ChatGPT Work on July 9), and users can set a hotkey to trigger voice commands while coding in other applications, removing friction from how developers interact with AI.
OpenAI is positioning voice as a competitive advantage in the crowded AI market. The company also signaled that GPT Live may power its forthcoming hardware device—an in-home speaker and conversation partner, according to Bloomberg. Codex reached 10 million users after the July 9 ChatGPT Work launch, suggesting significant developer momentum.
- 2
Skipper: offline AI assistant for boats, remote homes, vehicles
Skipper is an AI companion designed to work without internet, offering voice interaction, custom memory, HD vision, and local knowledge storage for use in off-grid environments like boats, remote homes, vehicles, workshops, expeditions, and private installations. Users in environments with unreliable or no internet access can now access AI assistance—voice reminders, manuals, safety procedures, equipment information—without depending on cloud services, keeping all core intelligence and information stored locally within the Skipper system rather than on external servers.
Skipper's ability to integrate with onboard, household, or mobile systems and retain user-specific context and personality; the product is customizable for specialized environments and retains all data locally, appealing to maritime, remote, and field-work users who cannot rely on connectivity.
- 3
Claude's voice mode expands to Opus and Sonnet models
Anthropic has made its Opus and Sonnet models available in voice mode—until now limited to the lighter Haiku model. The company is also bringing voice mode to apps like Gmail, Slack, and Canva, and expanding language support to French, German, Spanish, Hindi, Indonesian, Italian, Japanese, Korean, and Portuguese. Haiku was designed for quick answers and kept conversations brief, but users started applying voice mode to deeper work—analyzing business problems and generating complex responses. Opus and Sonnet can handle that load: they can deliver detailed analysis, turn conversations into one-page pitches, or adjust your calendar if your train is late. For businesses and developers in Japan and other markets, this opens voice interaction to tasks that require real reasoning, not just fast replies.
Users can now switch between text and voice mid-conversation and shift between models on the fly—so a quick Haiku chat can seamlessly escalate to Opus if the problem gets harder. Language support now includes Japanese, removing a barrier for native speakers.
- 4
Claude voice mode now supports three models, integrates Gmail, Slack, Notion
Anthropic updated Claude's voice mode to let users choose between Opus, Sonnet, and Haiku models, with the mode now defaulting to the fastest version of whichever model they last used in text chat. The voice mode can now connect to Gmail, Google Calendar, Slack, Canva, and Notion, enabling tasks like updating meeting slots, drafting emails, and creating documents. Claude's voice mode was previously limited to the Haiku model and quick responses, making it unsuitable for complex work. The expansion to more powerful models (Opus and Sonnet) and integration with external apps allows users to handle longer, more substantive conversations and complete real work directly through voice—a capability OpenAI's voice mode does not yet offer.
The update is available in beta to all users across platforms, but free users are restricted to the Haiku model with only one connected app. Anthropic has not changed the underlying voice model or detailed its voice stack, so users may not experience conversational improvements like better interruption handling that OpenAI's recent voice update provided.
- 5
Black Forest Labs releases Flux 3 with native audio video generation
German AI company Black Forest Labs released Flux 3, a multimodal foundation model that learns from images, video, and audio together. The model can generate videos with native audio up to 20 seconds long for the first time, supporting text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips. In early evaluations using 10-second clips at 720p, Flux 3 was preferred over multiple rivals: 93 percent over Luma Ray 3.2, 77 percent over Runway Gen-4.5, and 69 percent over Grok Imagine Video. Against stronger competitors, it matched or came close to Seedance 2.0 and Gemini Omni Flash (each at 52 percent preference), which already serve major production workflows. BFL also developed Flux-mimic, a video-action model being tested on production tasks at Audi, suggesting the architecture extends to robotics applications.
Flux 3 Image is set to launch in early access within the next few weeks, with improvements to complex prompts and multilingual text rendering. Action prediction will initially be offered through select partners. BFL plans open-weight access to the multimodal backbone under the name 'Flux 3 Dev,' with longer-term work on a single model combining perception, action, and language prediction.
- 6
Black Forest Labs launches FLUX 3 for image, video, and audio generation
Black Forest Labs released FLUX 3, a multimodal AI model that generates images and combined audio/video clips up to 20 seconds from a single prompt. The model is jointly trained across image, video, and audio rather than combining separate models. It also extends to robotic vision and actions. FLUX 3 represents the company's first public video generation model and frames creative generation, simulation, computer use, and robotics as connected applications of a single capability. Black Forest Labs positions this as "visual intelligence"—models that can perceive, predict, and act across physical and digital environments—rather than treating each modality as a separate tool.
FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, and FLUX 3 Act. The release is limited to start, meaning broader availability or pricing details may follow.
What to Watch
As the AI voice market heats up with OpenAI preparing hardware-integrated voice assistants and Anthropic expanding its multimodal capabilities, watch for how these platforms differentiate on conversational quality, local data privacy, and seamless model switching—features that could determine which voice AI becomes the default choice for consumers and enterprises. Meanwhile, Black Forest Labs' upcoming Flux 3 suite promises to blur the lines between image generation, video creation, and real-world action prediction, so keep an eye on how quickly these tools move from early access to practical applications across industries.
Sources
- OpenAI debuts ChatGPT Voice so you can have ongoing conversations, ask the AI to complete tasks while hands-free
- Skipper – An offline, voice-enabled AI companion for off-grid environments
- Claude’s voice mode is now available for Opus and Sonnet
- Anthropic updates Claude voice mode with more capable models
- Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs
- Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start
- AI Voice Phishing Performs on Par With Human Scammers at a Fraction of the Cost
- I ran the actual numbers on AI dubbing via API (ElevenLabs + lipsync) - here's what a minute of localized video really costs
- 「アームの動きが鈍い」の口頭相談から故障診断へ。SORABITOがイレブンラボと対話型音声AIを構築・コマツで実証開始
- OpenAI unveils Presence, a new platform that lets enterprises launch and manage realtime voice agents and chatbots
Share this with a friend
Send today's roundup to anyone who wants to keep up.
Get daily AI news free with AIToday
200+ AI sources, summarized in 1 minute. Email / LINE / Slack.
Sign up free