VibeVoice-ASR, a speech-to-text model, is now part of a Transformers release and available directly through the Hugging Face Transformers library. The model handles 60-minute long-form audio in a single pass and generates structured transcriptions identifying speaker identity, timestamps, and content, with support for customized context.
The model is natively multilingual, supporting over 50 languages. vLLM inference is now supported for faster inference. Finetuning code is available, and a Technique Report has been published.
VibeVoice uses continuous speech tokenizers (Acoustic and Semantic) operating at 7.5 Hz frame rate to preserve audio fidelity while boosting computational efficiency. It employs a next-token diffusion framework with an LLM (an AI model that understands and generates text) to understand textual context and a diffusion head to generate high-fidelity acoustic details.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Chinese large-model developer Z.ai says it can now support large-scale inference using roughly 100,000 domesti…

Broadcom announced VMware AI Factory, a software-defined foundation for VMware Private AI Cloud, at VMware Exp…

Mitsubishi Electric and its U.S

OpenClaw launched version 2.0, its largest update yet, with a version number of 2026.8.1
David Heinemeier Hansson (DHH), creator of Ruby on Rails, has released Omarchy 4.0 (Omarchy Quattro), the late…

Debian voted to allow developers to use AI tools in contributions to the Linux distribution, covering developm…
