Audio & Speech
Jul 24, 2026

The Gist
OpenAI has expanded voice capabilities across its ChatGPT platform, adding hands-free voice control to its desktop app for tasks ranging from general conversations to hands-free coding and task automation. Anthropic's Claude is similarly broadening its voice functionality, making voice mode available across both its Opus and Sonnet models on all platforms. Meanwhile, a new AI assistant called Skipper is launching as an offline voice solution designed for use in boats, remote homes, and vehicles where internet connectivity may be limited.
Today's Stories
- 1
OpenAI adds voice control to ChatGPT desktop app
OpenAI updated its ChatGPT desktop app on Thursday to support ChatGPT Voice, allowing users to speak commands and control AI agents to perform tasks on their computer. The feature uses ChatGPT-Live voice models and works with both ChatGPT Work and Codex, including the ability to access websites, apps, and on macOS, screen contents via Appshots. The desktop version is more capable than the smartphone version that launched earlier this month—it lets users dictate complex multi-step commands and interact back-and-forth with ChatGPT for clarification. In a demo, OpenAI showed a developer asking ChatGPT to create a thread, make a pull request, and debug an issue in a single voice command, suggesting voice control could reshape how developers and office workers interact with their machines.
The rollout is global as of Thursday. Users can also access ChatGPT Voice in Codex from iOS through remote access. Anthropic has similarly updated Claude's voice mode to work with its Opus, Sonnet, and Haiku models for tasks in Gmail, Calendar, Slack, Notion, and Canva.
- 2
Claude voice mode expands to Opus, Sonnet across all platforms
Anthropic expanded Claude's voice mode to its Opus and Sonnet models (previously limited to Haiku), with the ability to switch models mid-conversation on mobile, desktop, and web. The mode supports eleven languages and can integrate with tools like Gmail, Google Calendar, and Slack to compose and send emails by voice. Claude's voice mode now competes more directly with OpenAI's GPT-Live and Google's Gemini Live. While those rivals use full-duplex audio (speaking and listening simultaneously) for a more natural feel, Claude uses a turn-based system and differentiates itself through tool integration—it is currently the only provider that lets users compose and save emails directly from audio mode.
Users can now choose which Claude model to use during voice conversations and switch languages flexibly, giving them control over speed and capability trade-offs depending on their task.
- 3
OpenAI brings GPT-Live voice to desktop coding—engineers can now debug hands-free
OpenAI announced that GPT-Live, its full-duplex voice model (which listens and speaks simultaneously), now powers the ChatGPT desktop application on macOS and Windows, integrated directly with agentic coding systems like Codex and ChatGPT Work. Software engineers can now orchestrate multi-threaded coding jobs, review pull requests, and debug applications using natural voice commands instead of typing—potentially enabling hands-free software development workflows.
GPT-Live was initially launched on July 8, 2026; the desktop integration comes two weeks after that debut and represents the first integration of full-duplex voice into developer-facing tools.
- 4
OpenAI brings ChatGPT Voice to desktop, hands-free task automation
OpenAI has released ChatGPT Voice on desktop and laptop computers, extending the GPT Live voice technology it launched on mobile earlier this month. The feature lets users speak to ChatGPT in a stream-of-consciousness manner to complete multiple tasks in the background while working on other things—such as checking a calendar, reviewing emails, or drafting documents—without using their hands. The desktop rollout targets software programmers and high-volume AI token users, a lucrative segment of the AI market. The feature integrates directly with Codex, OpenAI's programming product (which merged into ChatGPT Work on July 9), and users can set a hotkey to trigger voice commands while coding in other applications, removing friction from how developers interact with AI.
OpenAI is positioning voice as a competitive advantage in the crowded AI market. The company also signaled that GPT Live may power its forthcoming hardware device—an in-home speaker and conversation partner, according to Bloomberg. Codex reached 10 million users after the July 9 ChatGPT Work launch, suggesting significant developer momentum.
- 5
Skipper: offline AI assistant for boats, remote homes, vehicles
Skipper is an AI companion designed to work without internet, offering voice interaction, custom memory, HD vision, and local knowledge storage for use in off-grid environments like boats, remote homes, vehicles, workshops, expeditions, and private installations. Users in environments with unreliable or no internet access can now access AI assistance—voice reminders, manuals, safety procedures, equipment information—without depending on cloud services, keeping all core intelligence and information stored locally within the Skipper system rather than on external servers.
Skipper's ability to integrate with onboard, household, or mobile systems and retain user-specific context and personality; the product is customizable for specialized environments and retains all data locally, appealing to maritime, remote, and field-work users who cannot rely on connectivity.
- 6
Claude's voice mode expands to Opus and Sonnet models
Anthropic has made its Opus and Sonnet models available in voice mode—until now limited to the lighter Haiku model. The company is also bringing voice mode to apps like Gmail, Slack, and Canva, and expanding language support to French, German, Spanish, Hindi, Indonesian, Italian, Japanese, Korean, and Portuguese. Haiku was designed for quick answers and kept conversations brief, but users started applying voice mode to deeper work—analyzing business problems and generating complex responses. Opus and Sonnet can handle that load: they can deliver detailed analysis, turn conversations into one-page pitches, or adjust your calendar if your train is late. For businesses and developers in Japan and other markets, this opens voice interaction to tasks that require real reasoning, not just fast replies.
Users can now switch between text and voice mid-conversation and shift between models on the fly—so a quick Haiku chat can seamlessly escalate to Opus if the problem gets harder. Language support now includes Japanese, removing a barrier for native speakers.
What to Watch
As OpenAI and Anthropic expand voice capabilities across their AI platforms—with ChatGPT Voice now globally available in Codex and Claude's voice mode supporting flexible model switching and language options like Japanese—expect voice to become a core competitive differentiator in everyday work tools, with OpenAI potentially leveraging GPT Live to power its rumored in-home speaker device. Watch how quickly developers and enterprises adopt these voice features in their workflows, particularly as seamless mid-conversation switching between text, models, and languages becomes standard, and whether specialized voice applications like Skipper can capture niche markets where offline functionality and data retention matter most.
Sources
- OpenAI’s new voice mode makes it to the ChatGPT desktop app
- Claude's voice mode now runs on Anthropic's most capable models across all platforms
- Agentic coding goes hands-free as OpenAI brings GPT-Live's full duplex voice control to Codex and ChatGPT on the desktop
- OpenAI debuts ChatGPT Voice so you can have ongoing conversations, ask the AI to complete tasks while hands-free
- Skipper – An offline, voice-enabled AI companion for off-grid environments
- Claude’s voice mode is now available for Opus and Sonnet
- Anthropic updates Claude voice mode with more capable models
- Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs
- Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start
- AI Voice Phishing Performs on Par With Human Scammers at a Fraction of the Cost
Share this with a friend
Send today's roundup to anyone who wants to keep up.
Get daily AI news free with AIToday
200+ AI sources, summarized in 1 minute. Email / LINE / Slack.
Sign up free