AIToday
Large Language ModelsAudio & SpeechHacker NewsPublished: Apr 29, 2026, 01:00 JST1 min read

PAVO: an 85,041-parameter router for voice pipelines cuts P95 latency 10.3% and energy 71% versus fixed-cloud, trained via PPO in 106 seconds on a 50,000-turn benchmark.

PAVO: an 85,041-parameter router for voice pipelines cuts P95 latency 10.3% and energy 71% versus fixed-cloud, trained via PPO in 106 seconds on a 50,000-turn benchmark.

3 Key Points

  1. Researchers at University of Pennsylvania and Google released PAVO-Bench, a 50,000-turn voice interaction dataset (40K train / 10K test) with complexity labels on HuggingFace, plus a trained tiny meta-controller (85,041 parameters) that decides per turn whether to route ASR → LLM → TTS calls to cloud or edge.

  2. The router characterizes inter-stage coupling: Gemma2 2B quality drops from 0.825 → 0.585 as ASR word-error rate crosses 2% (n=200 per WER level), meaning downstream LLM performance depends on upstream ASR configuration. Hard-constraint masking reduces coherence-failure rate from 7.1% → 0.9% (7.9× reduction) at +110 ms median latency cost.

  3. Against a fixed-cloud baseline on LibriSpeech, PAVO achieved P95 end-to-end latency −10.3% (−167 ms, p = 2×10⁻⁶), median latency −34%, and energy per turn −71%, measured on NVIDIA H100 and Apple M3 8 GB across Llama 3.1 8B, Mistral 7B, and Gemma2 2B.

  4. Code, trained weights, and all 5,430 coupling calibration measurements are open-sourced; full reproduction takes roughly half a day on an H100; training-only reproduction runs in ~2 minutes on a single A100.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia revives Rubin CPX chip with major redesignYahoo Finance AI · 2h ago
  • AI advice followed by 79%, but well-being unchangedITmedia AI+ · 5h ago
  • Enterprises face agent governance gapSiliconANGLE AI · 8h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleRural America mounts fierce resistance to AI data center expansion, creating political vulnerability for Trump administration