
AWS and NVIDIA cut ASR GPU needs by 75% using MPS.
The change reduces instances from 16 to 4.
Latency stays under 650 ms at 92.1 RPS per GPU.
What happened
AWS, NVIDIA, and Heidi Health published a solution that cuts GPU infrastructure for automatic speech recognition (ASR) inference by 75 percent—from 16 GPU instances down to 4—using NVIDIA CUDA Multi-Process Service (MPS) with Triton Inference Server on Amazon EC2.
Why it matters
A single ASR request typically uses only 15–20 percent of a GPU's compute capacity, leaving 80 percent idle due to default time-slicing. By partitioning the GPU into concurrent instances, the setup maintains sub-second latency at 92.1 requests per second (RPS) per GPU, which could reduce operational costs for high-throughput speech services.
What to watch
The solution is implemented on Amazon EC2 g6e.4xlarge and g7e.4xlarge instances (NVIDIA L40S, 48 GB) and requires NVIDIA drivers 535+ with CUDA 12.x. Code is available in an open-source repository, deployable with three commands.
Ask the AI about this article →
The post describes how Heidi Health, an AI Care Partner processing over 2.4 million clinical consultations weekly across 190 countries, faced high GPU costs due to inefficient GPU utilization. Low per-request GPU usage and strict latency requirements forced the company to run 16 instances. The collaboration with AWS and NVIDIA produced a solution that leverages CUDA MPS to partition the GPU, achieving a 75% reduction in infrastructure while maintaining sub-second latency.
This optimization builds on an earlier post about fine-tuning the Parakeet TDT 0.6B V2 model. The current work focuses on serving that model efficiently. By combining MPS with Triton's dynamic and sequence batching, the solution addresses the specific bottlenecks of transcription and diarization workloads. The open-source repository offers a practical path for others facing similar challenges, potentially lowering the barrier to deploying ASR at scale on AWS.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Nvidia is reported to be acquiring Hugging Face for $13 billion, following its $6 billion deal with Poolside a…

Beatport now bans tracks made entirely or mostly by AI

The Fire and Disaster Management Agency plans to launch a model project in fiscal 2027 to use AI in handling 1…

Anthropic PBC previewed the Model Hardware Standard (MHS), a unified interface that lets AI agents control sci…
NVIDIA has agreed to acquire AI platform Hugging Face for $2 trillion, according to the article

AWS announced an integration where Amazon Quick, an agentic AI workspace, connects to fal's generative media p…
