
What happened
Tencent's Hunyuan Speech team and university researchers introduced Gander, which splits real-time conversation from background tasks between a "cerebellum" and a swappable "brain." On Full-Duplex-Bench v3, Gander interrupted users in 8 percent of cases versus 13.5 percent for GPT-Realtime.
Why it matters
Gander's lower interruption rate suggests splitting conversation from reasoning can make voice assistants feel less disruptive than systems like GPT-Realtime, though its task accuracy fell slightly below that of the weakest competitor.
What to watch
The researchers call the work early, and how to scale Gander up remains an open question. Tencent plans to publish the weights and training data once it completes the open source release process.
WHO IT HITSTeams building conversational voice and chat agents — who the article notes report latency problems particularly often — may see Gander's split-role design as a template for handling interruptions without delaying task work.
Summaries like this, in your inbox every morning.
Gander reflects a broader push to divide agent work across multiple models. The article notes that OpenAI's GPT-Live separates conversation from reasoning, Sakana AI's Fugu is a separate language model that calls others from an expandable pool, and OpenAI is testing proactive agents that create follow-up tasks. Tencent's own July release of Hy3, an open language model that reportedly narrowed the gap with rivals especially on agent tasks, already runs in WorkBuddy, Yuanbao, and WeChat — and the company is negotiating to take the largest stake in agent startup Manus after Beijing blocked Meta's acquisition. That deal is seen as a fit for its plans, including an agent embedded in WeChat.
Gander was trained on about 2.7 million examples, some of which teach it to stay quiet when there is background noise or nobody in a group is addressing it. The researchers say the work is still early: scaling it up remains an open question, and there is no standard way to evaluate systems like it. Handling interruptions and avoiding delays are practical concerns — an Anthropic analysis found experienced users interrupt Claude Code in about 9 percent of work steps, compared with roughly 5 percent for newcomers.
Whether Gander becomes more than a research showcase hinges on how well it scales and whether its task accuracy improves once the underlying brain model does, since a stronger brain should lift the whole system. For teams building conversational agents, the experiment may offer evidence that treating conversation and reasoning as separate problems is worth the complexity — but only if the accuracy gap narrows in practice.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Epoch AI and Ipsos surveys found the share of US adults using AI on at least six of seven days rose from 8 per…

AAA AI launched a multi-agent system connecting local runtimes like Ollama or cloud APIs under one orchestrati…

Cambridge researchers interviewed 27 former members of Boko Haram's two factions, ISWAP and JAS, who described…

Meta's AI assistant Muse was downloaded over 900,000 times in its first week, according to Sensor Tower, but a…

TypeSafe AI released Jev on September 15, its first model after roughly two years of development; founder Diog…

OpenAI launched Astra, which reportedly uses "recurrent depth" to make reasoning more efficient by not spellin…
