
What happened
The developer oboroge0 released Hayamimi, a real-time multilingual speech recognition system that runs on CPU only, without a GPU or cloud API. It achieved a 3.8% Character Error Rate on Japanese TV audio and ran 10–50× faster than real time on a 6-core desktop CPU.
Why it matters
Its accuracy on Japanese broadcast speech and its speed mean CPU-only setups can match dedicated hardware, potentially widening who can deploy live multilingual transcription without cloud services.
WHO IT HITSDevelopers and IT teams who want live transcription without GPUs or cloud APIs — for streaming, captioning, or internal tools — can now run it on ordinary CPUs with 2GB or less of memory. Broadcasters and accessibility staff may also benefit if they need multilingual subtitles on modest hardware.
Summaries like this, in your inbox every morning.
The release follows a pattern where CPU-based speech recognition has often depended on a single model. Hayamimi takes a different approach: it detects the language of each utterance and routes it to a dedicated model, which allows it to handle multiple languages without a GPU. The system uses INT8-quantized ONNX models running on sherpa-onnx, so PyTorch and CUDA are not required. On the accuracy side, a two-pass correction re-decodes the previous utterance after a 2-second silence, improving the Japanese Character Error Rate from 15.5% to 12.0%. The project also supports speaker labeling, hotwords, translation, and a browser dashboard through a local HTTP server. These features suggest the developer is targeting live production use — streaming overlays, transcription pages, and network audio input — rather than just lab experiments. The documented limitations, such as not handling mixed-language utterances or separating simultaneous speakers, indicate the system is designed for clear, single-speaker scenarios rather than complex meeting environments.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Apple told the European Commission it agreed to make employment offers to "certain employees of Huxe AI" and t…

Bragi CEO Nikolaj Hviid said AI audio is still searching for its own interface, and the audio market lacks a c…

RESONAL, a music unit run by one human and the AI vocalist 彩瀬??, released a guide showing how a song is made b…

OpenAI began offering a ChatGPT feature that accepts uploaded audio files and can transcribe them, summarize t…

NECネクサソリューションズ started offering "対話型音声AIエージェント Powered by DEN.Ai" on 2026年10月6日

Automation Anywhere announced Wednesday an agreement to acquire Boost.ai, a conversational voice AI company, f…