
Alex Smola, a former Amazon distinguished scientist, has launched Boson AI to compete directly with OpenAI and Meta in voice AI models. His startup claims to offer speech-to-speech capabilities at one-tenth the cost of competitors by using a full-stack approach that lets enterprise clients run systems in their own data centers. Boson has raised $70 million(約110億円) and is targeting finance, telecommunications, healthcare, and insurance sectors, as voice AI emerges as a primary frontier in the race to build the next generation of conversational AI.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Alex Smola, former distinguished scientist at Amazon, founded Boson AI to release Higgs RealTime, a speech-to-speech model designed to be significantly cheaper than competitors' voice systems. The Santa Clara startup has raised $70 million(約110億円) and counts Chinese entrepreneur Su Hua and Singapore-based Temasek's venture arm among its backers.
Why it matters
Voice AI has become a primary battleground for OpenAI, Meta, and other leaders building full-duplex systems (technology enabling natural back-and-forth conversation where users can interrupt mid-sentence). Boson claims its models cost one-tenth as much as rivals, potentially reshaping the economics of deploying voice agents for customer support, sales, and other enterprise applications. Smola is targeting finance, telecommunications, healthcare, and insurance clients.
What to watch
Boson faces well-funded rivals—Microsoft recently announced a new voice model iteration, OpenAI has launched GPT-Live (a family of full-duplex audio models), and Meta released Muse Spark optimized for voice across wearables. The key technical hurdle remains latency; while one-second delays are acceptable in text chat, they feel disruptive in spoken conversation.
Alex Smola, a former distinguished scientist at Amazon and leading machine learning researcher, founded Boson AI in Santa Clara, California, to build voice AI models that compete directly with OpenAI, Meta, and Microsoft. The startup is preparing to release Higgs RealTime, a speech-to-speech model, and has raised $70 million(約110億円) from backers including Chinese entrepreneur Su Hua and Temasek, Singapore's venture investment arm.
Smola's core insight is that the industry was moving beyond text-only interfaces toward multimodal AI—systems that integrate audio and vision to deliver richer user experiences. However, he identified a market gap: existing voice models from OpenAI, Meta, and others are expensive to build and deploy. Boson claims its models cost one-tenth as much as competitors' systems. "What we have is about an order of magnitude more affordable [than competitors], and I would say it's nonetheless very competent to use," Smola told Fortune. The company's full-stack approach is key to this cost advantage: it allows enterprise customers to run systems within their own data centers, train custom voice and video models from scratch, and maintain direct control over data residency and model execution—a critical requirement for sensitive business environments in finance, telecommunications, healthcare, and insurance.
The voice AI space has become a primary battleground. OpenAI Chief Executive Sam Altman recently posted that he now talks with ChatGPT more than he texts with it, noting the company's "new voice model really crossed a threshold." OpenAI has launched GPT-Live, a family of full-duplex audio models enabling natural, fluid back-and-forth conversations where users can interrupt mid-sentence. Meta released Muse Spark, similarly optimized for voice with the ability to "talk naturally with the assistant—interrupt, switch topics, or swap languages." Microsoft announced a new voice model iteration. Each competitor is positioning differently: Microsoft around workplace productivity, OpenAI as a default for app developers, and Meta for hardware integration across wearables.
Bosons faces real technical hurdles. While a one-second response lag is acceptable in text chat, the same pause feels awkward in spoken conversation. Smola says Boson is training its AI to handle latency challenges by parsing the nuances of the "human condition," including emotional tone and rapid speech patterns. Full-duplex audio also incurs high costs because it increases graphics processing unit usage. Despite these obstacles, Smola sees immediate enterprise value: voice models can automate customer support and sales interactions, maintain total context during calls to generate real-time logs and personalized records faster than human teams, and ultimately become "strictly superior" to humans in many roles. Looking further ahead, Smola envisions embodied robots equipped with animated faces and advanced voice models as household staples, with robots combining reasoning, vision, and voice as the "ultimate technical frontier."
Voice AI has emerged as the next major frontier in conversational AI, moving beyond text-only interfaces into multimodal systems that combine audio, vision, and reasoning. OpenAI CEO Sam Altman recently noted that the company's new voice model "really crossed a threshold," signaling that the technology has matured enough to drive meaningful shifts in human-machine interaction. This has prompted intensive competition: OpenAI released GPT-Live, Meta launched Muse Spark, and Microsoft announced a new voice iteration, with each company positioning its offering differently—Microsoft around workplace productivity, OpenAI as a default for app developers, and Meta for hardware integration across wearables.
Smola's differentiator is cost and control. By offering a full-stack architecture that allows enterprise customers to run systems within their own data centers and train custom models from scratch, Boson gives clients direct control over data residency and model execution—a significant advantage for regulated industries like finance and healthcare. At one-tenth the cost of competitors, the startup could disrupt the economics of voice deployment, though building low-latency, emotionally-aware systems that feel natural in conversation remains technically challenging. The broader vision Smola articulates—embodied AI combining reasoning, vision, and voice—suggests voice models are not a standalone product but a component of a larger shift toward multimodal, context-aware agents.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime