
Microsoft VibeVoice-ASR 9B now leads open-source models for medical speech-to-text with 8.34% WER, nearly matching Gemini 2.5 Pro's 8.15%
The model requires ~18GB VRAM and processes audio slowly at 97 seconds per file, compared to 6 seconds for Parakeet
Benchmark expanded from 26 to 31 models, with new entrants including ElevenLabs Scribe v2 (9.72% WER) and NVIDIA Nemotron Speech Streaming 0.6B (11.06% WER)
Researcher discovered bugs in Whisper's text normalizer that were artificially inflating WER scores by 2-3% across all tested models
All code and results are open-source for the third iteration of this medical speech recognition benchmark
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
Taoyuan is positioning itself as a northern hub for AI data centers (AIDC), citing the Tatan area and an LNG c…

SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

The U.S. Department of Defense announced on August 31 that it has deployed ChatGPT Mil, a customized version o…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider
