
Apple has published technical details on the memory-efficient audio synthesis system powering Siri Expressive Voices, which runs entirely on device using Apple's AFM 3 Core Advanced foundation model. The breakthrough involves a detokenizer architecture that converts semantic audio tokens into high-quality speech while fitting within the strict memory and compute limits of the Apple Matrix Coprocessor.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Apple has released technical details on the audio synthesis architecture behind Siri Expressive Voices, which generate rich, configurable speech in real time entirely on device using AFM 3 Core Advanced, Apple's most powerful on-device foundation model.
Why it matters
The architecture uses a memory-efficient detokenizer that converts semantic audio tokens into high-fidelity audio within the tight compute and memory constraints of the Apple Matrix Coprocessor (AMX), enabling advanced voice synthesis without relying on cloud processing.
What to watch
The system converts semantic audio tokens to a residual vector quantization (RVQ) representation using a three-component design, demonstrating Apple's focus on delivering sophisticated AI capabilities entirely on consumers' devices.
Apple has unveiled the technical architecture behind Siri Expressive Voices, a capability that synthesizes rich, configurable speech in real time entirely on device. The system is powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model, which generates semantic audio tokens. The key innovation is a memory-efficient detokenizer that converts those semantic tokens into high-fidelity audio while operating within the tight compute and memory constraints of the Apple Matrix Coprocessor (AMX). The detokenizer achieves this efficiency through a three-component design that converts semantic audio tokens into a residual vector quantization (RVQ) representation. This architecture allows Siri to deliver expressive, nuanced speech synthesis locally on consumer devices without offloading work to cloud servers, preserving both user privacy and system responsiveness.
Apple's release of technical details on Siri Expressive Voices reflects the company's strategic commitment to on-device AI processing. Rather than streaming voice generation to cloud servers, the system leverages AFM 3 Core Advanced, the company's most powerful on-device foundation model, to handle the full synthesis pipeline locally. This approach prioritizes both privacy (no data leaves the device) and responsiveness (no network latency). The core innovation—the memory-efficient detokenizer—addresses a fundamental constraint in consumer hardware: the Apple Matrix Coprocessor has strict limits on both compute and memory. By converting semantic audio tokens into a residual vector quantization representation with a three-component design, Apple engineers created a pathway from high-level semantic information to natural-sounding speech without exceeding those hardware boundaries.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime