AITodayYour daily AI briefing

Audio & Speech

Jul 27, 2026

Audio & Speech

The Gist

Voice AI is rapidly advancing with new tools and features—OpenAI has added voice control to its ChatGPT desktop app, while a former AWS scientist is launching a cheaper voice AI alternative to compete with OpenAI and Meta. However, concerns are mounting about the true costs of voice AI services, as actual call expenses significantly exceed advertised per-minute rates, prompting developers to seek better visibility tools like the new open-source profiler for voice agents. Meanwhile, infrastructure improvements continue with Deepgram's AWS integration and the launch of ARIA's new voice-native platform.

Today's Stories

  1. 1

    Deepgram integrates AWS IAM temporary delegation for SageMaker support access

    Deepgram has integrated IAM temporary delegation, a new AWS IAM capability, into its support workflow for customers running Deepgram speech AI models on Amazon SageMaker AI. The integration allows Deepgram engineers to request time-limited, scoped access to customer endpoints and logs directly from the support ticketing system, with customers approving requests in their own IAM console. This replaces the previous model of long-lived cross-account IAM roles, which required recurring provisioning and audit conversations. Deepgram has reduced the time for initial investigation on a SageMaker AI support ticket from days to minutes, since customers can now approve access in their IAM console rather than scheduling screen-shares. Every delegated API call is tagged in AWS CloudTrail with Deepgram's partner account ID for full auditability.

    Credentials issued through the integration expire automatically after twelve hours, and customers can revoke access at any time before expiration. The integration requires an active Deepgram support contract, an enabled AWS CloudTrail trail in the region where the SageMaker AI endpoint runs, and AWS infrastructure charges for SageMaker AI endpoint hosting, Amazon CloudWatch logs, AWS CloudTrail, and networking (Deepgram offers a 14-day trial at no additional cost for its models, but AWS infrastructure costs apply from the start of deployment).

  2. 2

    Ex-AWS scientist targets OpenAI, Meta with cheaper voice AI startup

    Alex Smola, former distinguished scientist at Amazon, founded Boson AI to release Higgs RealTime, a speech-to-speech model designed to be significantly cheaper than competitors' voice systems. The Santa Clara startup has raised $70 million(約110億円) and counts Chinese entrepreneur Su Hua and Singapore-based Temasek's venture arm among its backers. Voice AI has become a primary battleground for OpenAI, Meta, and other leaders building full-duplex systems (technology enabling natural back-and-forth conversation where users can interrupt mid-sentence). Boson claims its models cost one-tenth as much as rivals, potentially reshaping the economics of deploying voice agents for customer support, sales, and other enterprise applications. Smola is targeting finance, telecommunications, healthcare, and insurance clients.

    Boson faces well-funded rivals—Microsoft recently announced a new voice model iteration, OpenAI has launched GPT-Live (a family of full-duplex audio models), and Meta released Muse Spark optimized for voice across wearables. The key technical hurdle remains latency; while one-second delays are acceptable in text chat, they feel disruptive in spoken conversation.

  3. 3

    Open-source profiler for voice agents offers visibility into model costs and latency

    A developer has released VoiceGateway, an open-source tool that profiles voice agent workloads by monitoring every speech-to-text (STT), language model (LLM), and text-to-speech (TTS) call—tracking latency splits, model IDs, and costs without requiring changes to existing agents built on LiveKit, Pipecat, or similar platforms. Voice agent builders using self-hosted frameworks currently lack built-in visibility into model-layer performance and expenses; VoiceGateway fills that gap by sitting outside the audio path and offering a single integration point, making it easier to understand where time and money are spent in production voice systems.

    The tool is MIT licensed, runs in Docker, stores data locally in SQLite, and includes a reconciliation feature (voicegw reconcile) that lets users verify recorded costs against provider invoices—key features for teams managing cost accountability in self-hosted deployments.

  4. 4

    Voice AI call costs far exceed advertised per-minute rates

    A developer calculated the true cost of Voice AI phone calls and found that advertised per-minute pricing omits platform fees, transcription, model inference, voice generation, carrier charges, recording, failed calls, and phone-number rental. Realtime speech-to-speech models add complexity because input and output audio can have separate prices, and costs vary by destination country. Voice AI products typically advertise a single per-minute price that rarely reflects actual expenses. Understanding the full cost structure—spanning infrastructure, telecom carriers, and infrastructure overhead—is essential for businesses evaluating Voice AI providers or building their own systems, as the real cost can be significantly higher than the headline figure suggests.

    The developer has published an open-source Astro project with cost calculators, country-specific SIP (Session Initiation Protocol, the standard for Voice AI routing) estimates, latency benchmarks, Asterisk (a phone-system software) configuration examples, and a local call-log diagnostic tool to help others model Voice AI expenses accurately.

  5. 5

    ARIA voice-native 3D SOC platform launches under source-available license

    A developer has released ARIA, a voice-controlled security operations cockpit (SOC) that runs on local hardware without cloud dependency. The platform uses a "trust ladder" system where AI autonomy must be earned through demonstrated outcomes and can be revoked immediately; it integrates connectors for GitHub, AWS, Okta, Snyk, Azure AD, VirusTotal, and Elastic Security. Code is source-available (not open source) under Business Source License 1.1, with an engineering audit published detailing which subsystems are production-grade and which are demo-stage. Most security tools act first and ask for trust later; ARIA inverts that, requiring human approval or tracked performance before granting autonomous action. For regulated environments where sending telemetry to third-party SaaS is illegal, the platform enforces sovereignty in code—a local model path refuses to send prompts outside loopback or private IP ranges even if cloud API keys are present. Single operators can navigate via voice, 3D spatial interface, or text, with every action logged to an audit trail.

    The product is under active solo development and explicitly not hardened for production; parts are labeled demo-grade. Key gaps include a real local text-to-speech engine (currently a cloud call), multi-tenancy isolation, and CI/CD automation. The quickstart requires only Node.js ≥ 22 and optional Ollama; the desktop app and containerized server options are available. Full module status with file paths and line numbers is published in ROADMAP_AND_LIMITATIONS.md.

  6. 6

    OpenAI adds voice control to ChatGPT desktop app

    OpenAI updated its ChatGPT desktop app on Thursday to support ChatGPT Voice, allowing users to speak commands and control AI agents to perform tasks on their computer. The feature uses ChatGPT-Live voice models and works with both ChatGPT Work and Codex, including the ability to access websites, apps, and on macOS, screen contents via Appshots. The desktop version is more capable than the smartphone version that launched earlier this month—it lets users dictate complex multi-step commands and interact back-and-forth with ChatGPT for clarification. In a demo, OpenAI showed a developer asking ChatGPT to create a thread, make a pull request, and debug an issue in a single voice command, suggesting voice control could reshape how developers and office workers interact with their machines.

    The rollout is global as of Thursday. Users can also access ChatGPT Voice in Codex from iOS through remote access. Anthropic has similarly updated Claude's voice mode to work with its Opus, Sonnet, and Haiku models for tasks in Gmail, Calendar, Slack, Notion, and Canva.

What to Watch

Watch for whether real-time voice AI can solve its latency problem—the industry's biggest technical challenge—as major players like Microsoft, OpenAI, and Meta race to ship faster full-duplex models while smaller developers explore cost-efficient alternatives through open-source tools and self-hosted deployments. Meanwhile, keep an eye on how quickly these voice assistants integrate into everyday work apps (Gmail, Slack, Notion) and whether the ecosystem can deliver production-grade reliability and cost transparency that teams need to adopt voice AI at scale.

Sources

Share this with a friend

Send today's roundup to anyone who wants to keep up.

Get daily AI news free with AIToday

200+ AI sources, summarized in 1 minute. Email / LINE / Slack.

Sign up free