AIToday
Audio & SpeechAI Business & IndustryAmazon AI BlogPublished: Aug 28, 2026, 04:00 JST2 min read

Deepgram expands SageMaker AI observability metrics

Deepgram expands SageMaker AI observability metrics

Key takeaway

  • Deepgram launched two observability features for its speech models on SageMaker AI.

  • One gives billing and usage metrics in CloudWatch, the other engine and GPU metrics.

  • Both work under network isolation.

3 Key Points

  1. What happened

    Deepgram has introduced two capabilities for its speech-to-text and text-to-speech models on Amazon SageMaker AI: Deepgram Enhanced Metrics for billing and usage transparency, and Prometheus and OpenTelemetry support for engine-level and per-GPU visibility. Both are available today on Deepgram SageMaker AI deployments.

  2. Why it matters

    These innovations address the observability trade-off of self-hosted speech AI, giving you visibility into billing, feature usage, and engine behavior that was previously locked inside the vendor's container. They work within AWS Marketplace network isolation, so you can reconcile your AWS bill and monitor capacity without opening a network path or adding extra permissions.

  3. What to watch

    Detailed observability is on by default for newly created endpoints, publishing every 60 seconds. You can query the metrics with PromQL from CloudWatch, Grafana, or any Prometheus-compatible tool, and filter by endpoint, instance, or GPU.

Ask the AI about this article →

Context & Analysis

Deepgram's move addresses a long-standing pain point for self-hosted speech AI: the lack of visibility into what happens inside the vendor's container. Previously, you could see basic endpoint metrics like request counts, but not the billing units, feature usage, or engine behavior that drive capacity planning and cost management. By publishing these metrics into your own CloudWatch account, Deepgram lets you reconcile your AWS bill against actual traffic and understand which features your applications use, without compromising network isolation.

The two innovations complement each other. Deepgram Enhanced Metrics focus on billing and usage, answering questions like what you're billed for and how much streaming versus pre-recorded traffic you handle. The Prometheus and OpenTelemetry support goes deeper, offering per-GPU and host-level metrics that can reveal a saturated device hiding behind a summed average. Together, they provide a complete picture, from cost to engine headroom.

For teams running Deepgram on SageMaker AI, this could simplify compliance and cost governance. The billing stream cannot be disabled, ensuring a permanent audit trail, while the usage stream can be turned off if needed. As these metrics are standard CloudWatch metrics, you can build dashboards and set alarms, integrating AI operations into your existing monitoring workflows.

FAQ

How do I enable these new metrics?
Detailed observability is on by default for newly created endpoints, publishing every 60 seconds. For existing endpoints, you can enable it via the MetricsConfig on the endpoint configuration.
Can I filter the metrics to a specific endpoint or instance?
Yes, the Prometheus and OpenTelemetry metrics carry SageMaker resource labels, including endpoint name, variant, and instance ID, so you can filter to a single endpoint or instance.
What does the 'engine_estimated_stream_capacity' metric tell me?
It's the engine's own estimate of how many concurrent streams the instance can sustain. Comparing it with active requests gives you an engine-reported headroom signal for scaling decisions.
Amazon AI BlogRead Original Article

Also reported by Top Companies AI

Get the latest Audio & Speech news every morning

For example, today's edition would include:

  • Musician Detectives Hunt AI Music GriftersThe Verge AI · 16h ago
  • Beatport bans AI-generated music from DJ marketplaceTHE DECODER · 1d ago
  • Japan to test AI for 119 emergency call handlingJapan Times Tech · 1d ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleChatGPT plus critical-thinking training boosts student originality