AIToday
Large Language ModelsAudio & SpeechZenn AI/MLPublished: Oct 4, 2026, 22:00 JST

Streaming STT cuts AI voice reply to p50 1.594 seconds

Streaming STT cuts AI voice reply to p50 1.594 seconds

3 Key Points

  1. What happened

    The developer added Nemotron 3.5 Streaming ASR 0.6B 560ms INT8 to the Koehaku voice runtime, and a 100-turn test measured p50 1.594 seconds from the user finishing speech to the first audio render.

  2. Why it matters

    That is close to conversational timing, so a voice AI reply may no longer feel like waiting for a text chat to load.

  3. What to watch

    The test ran on one local setup, so the result hinges on whether the same code holds up in other environments. Watch the ONNX Runtime spinning fix, since disabling it cut worker CPU time by about 78%.

WHO IT HITSDevelopers building real-time voice assistants on local hardware are the ones who can act on this — the numbers show speech recognition speed is only one lever, and runtime details like thread spinning can silently degrade long sessions.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The project started as an experiment to have an AI speak in a slow, character-style voice, built from AquesTalk for speech synthesis, whisper.cpp for recognition, and an LLM in between. That version worked, but the pause before a reply became obvious once the interaction was spoken rather than typed. The developer's answer was to rebuild the speech recognition side so text arrives while the user is still talking.

The streaming model is only part of what changed. VAD decides the user has paused, Smart Turn decides whether the turn is really over, and an LLM begins speculatively on the interim text. If the final transcript matches what was guessed, that answer is kept; if not, it is thrown away and regenerated. In the 100-turn test, 74 of 100 replies were kept this way. A separate Whisper path also uncovered a timeout that was cancelling the final recognition step, which the developer fixed by separating committed work from speculative timeouts.

The stakes now hinge on whether these gains survive outside the developer's own machine. The spinning-thread problem only appeared across a long soak test, so the runtime's value may depend on whether it can keep measuring and catching such slowdowns. Software teams building voice features — and anyone making AI characters that need to answer at the right moment — are the ones likely to feel the difference.

FAQ
What exactly was measured?
The test measured the time from the end of the user's speech to the moment the browser's AudioWorklet renders the first audio, not the instant sound leaves the speaker.
Why did the 100-turn test get slower in the later turns?
The ONNX Runtime thread pool kept spinning between inferences, keeping the CPU busy and lowering its effective performance state, which gradually slowed Nemotron's decoding.
Did the developer compare Nemotron with Whisper?
Yes. On a small set of five Japanese and five English synthetic clips, Nemotron had a 2.17% Japanese character error rate and 7.50% English word error rate, against Whisper's 11.96% and 10.00%.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleTSMC, $2.4 trillion chipmaker behind AI, reports $143 billion revenue