
Deepgram launched two observability features for its speech models on SageMaker AI.
One gives billing and usage metrics in CloudWatch, the other engine and GPU metrics.
Both work under network isolation.
What happened
Deepgram has introduced two capabilities for its speech-to-text and text-to-speech models on Amazon SageMaker AI: Deepgram Enhanced Metrics for billing and usage transparency, and Prometheus and OpenTelemetry support for engine-level and per-GPU visibility. Both are available today on Deepgram SageMaker AI deployments.
Why it matters
These innovations address the observability trade-off of self-hosted speech AI, giving you visibility into billing, feature usage, and engine behavior that was previously locked inside the vendor's container. They work within AWS Marketplace network isolation, so you can reconcile your AWS bill and monitor capacity without opening a network path or adding extra permissions.
What to watch
Detailed observability is on by default for newly created endpoints, publishing every 60 seconds. You can query the metrics with PromQL from CloudWatch, Grafana, or any Prometheus-compatible tool, and filter by endpoint, instance, or GPU.
Ask the AI about this article →
Deepgram's move addresses a long-standing pain point for self-hosted speech AI: the lack of visibility into what happens inside the vendor's container. Previously, you could see basic endpoint metrics like request counts, but not the billing units, feature usage, or engine behavior that drive capacity planning and cost management. By publishing these metrics into your own CloudWatch account, Deepgram lets you reconcile your AWS bill against actual traffic and understand which features your applications use, without compromising network isolation.
The two innovations complement each other. Deepgram Enhanced Metrics focus on billing and usage, answering questions like what you're billed for and how much streaming versus pre-recorded traffic you handle. The Prometheus and OpenTelemetry support goes deeper, offering per-GPU and host-level metrics that can reveal a saturated device hiding behind a summed average. Together, they provide a complete picture, from cost to engine headroom.
For teams running Deepgram on SageMaker AI, this could simplify compliance and cost governance. The billing stream cannot be disabled, ensuring a permanent audit trail, while the usage stream can be turned off if needed. As these metrics are standard CloudWatch metrics, you can build dashboards and set alarms, integrating AI operations into your existing monitoring workflows.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
China has published its first national standard for liquid cooling in data centers, responding to AI racks app…

OpenAI announced on August 28 that it will terminate its model access contract with AI coding tool Cursor, pro…

Bernstein analysts say AI is creating a memory bottleneck, with demand expanding from high-bandwidth memory (H…

OpenAI CEO Sam Altman said in a Time magazine interview that he thinks "it is a good time to slow down" AI mod…

Nvidia's earnings on Wednesday shifted the narrative, showing its advantage extends beyond GPUs to the systems…

Intel expanded its partnership with Kasm Technologies to support compliant, local AI workloads on Intel Xeon 6…
