AIToday
Audio & SpeechGIGAZINE AIPublished: Oct 11, 2026, 13:00 JST

Hayamimi runs real-time speech recognition on CPU only, hits 3.8% error

Hayamimi runs real-time speech recognition on CPU only, hits 3.8% error

3 Key Points

  1. What happened

    The developer oboroge0 released Hayamimi, a real-time multilingual speech recognition system that runs on CPU only, without a GPU or cloud API. It achieved a 3.8% Character Error Rate on Japanese TV audio and ran 10–50× faster than real time on a 6-core desktop CPU.

  2. Why it matters

    Its accuracy on Japanese broadcast speech and its speed mean CPU-only setups can match dedicated hardware, potentially widening who can deploy live multilingual transcription without cloud services.

WHO IT HITSDevelopers and IT teams who want live transcription without GPUs or cloud APIs — for streaming, captioning, or internal tools — can now run it on ordinary CPUs with 2GB or less of memory. Broadcasters and accessibility staff may also benefit if they need multilingual subtitles on modest hardware.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The release follows a pattern where CPU-based speech recognition has often depended on a single model. Hayamimi takes a different approach: it detects the language of each utterance and routes it to a dedicated model, which allows it to handle multiple languages without a GPU. The system uses INT8-quantized ONNX models running on sherpa-onnx, so PyTorch and CUDA are not required. On the accuracy side, a two-pass correction re-decodes the previous utterance after a 2-second silence, improving the Japanese Character Error Rate from 15.5% to 12.0%. The project also supports speaker labeling, hotwords, translation, and a browser dashboard through a local HTTP server. These features suggest the developer is targeting live production use — streaming overlays, transcription pages, and network audio input — rather than just lab experiments. The documented limitations, such as not handling mixed-language utterances or separating simultaneous speakers, indicate the system is designed for clear, single-speaker scenarios rather than complex meeting environments.

FAQ
What hardware do I need to run Hayamimi?
It runs on a CPU with Python 3.10+ and ffmpeg. No GPU or CUDA is needed, and the models are INT8-quantized ONNX that run on sherpa-onnx.
How accurate is it for Japanese?
It achieved a 3.8% Character Error Rate (CER) on real Japanese TV broadcast audio. A two-pass correction step improves CER from 15.5% to 12.0% after a 2-second silence.
What languages does it support?
It has dedicated models for five major languages (Japanese, Chinese, Korean, Cantonese, English) and 24 European languages. About 1600 other languages fall back to Meta Omnilingual ASR.
Does it translate subtitles in real time?
Yes, with the --translate option it can translate Japanese subtitles into a target language such as English, Chinese, Korean, or Spanish in real time.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articletheCUBE to livestream AI Data Pipeline Forum Oct. 13