
Observe by Snowflake is rolling out a redesigned MCP server and new CLI that let AI agents directly access observability and telemetry data for incident investigation and debugging tasks.
The redesign removes a cost-heavy LLM intermediary, connecting agents straight to Observe's APIs, and includes prebuilt workflows for common observability tasks.
This shift reflects the growing reality that AI agents—not just humans—are now primary consumers of production telemetry.
What happened
Observe by Snowflake announced general availability of a redesigned MCP (model context protocol) server and a new CLI (command-line interface) that give AI agents direct access to production telemetry and observability data. The new MCP server architecture removes an intermediary LLM layer, connecting agents directly to Observe's APIs instead.
Why it matters
As AI agents increasingly handle tasks like alert triage, incident investigation, and error debugging—work traditionally done by humans—observability platforms must serve both human users and machine agents. The new tools enable engineers to build custom alert-triage agents and copilots that can query production data autonomously, reducing time to incident resolution and eliminating the cost of Observe's previous LLM middleman.
What to watch
The rollout is general availability now across all clusters, with eu-2 and ca-1 regions to follow shortly. Both the MCP server and CLI ship with prebuilt skills for common observability workflows (incident investigation, failure tracing, validation, n+1 detection, and outlier detection) that work immediately after configuration.
Observe by Snowflake has announced the general availability of a redesigned MCP server and a new CLI, both of which enable AI agents to access production telemetry and observability data directly. This announcement addresses a new reality in incident response and debugging: AI agents are now primary consumers of operational data. Coding assistants can investigate errors before an engineer opens a ticket; AI SREs can correlate failures across services without human intervention. As agents become more sophisticated in reasoning about production systems, observability platforms must work for both humans and machines—through UIs, APIs, and interfaces that agents can call and automate.
The previous Observe CLI covered only a narrow slice of platform functionality, with most capabilities inaccessible from the terminal. The new CLI is agent-compatible and provides programmatic access to Observe's full capabilities—the same surface available through the MCP server and the UI. It works from agent environments like Claude Code as well as interactive terminal sessions. Engineers can compose workflows that run autonomously in the background (handling routine tasks without human presence) or interactively (with an engineer guiding the investigation). The CLI is designed as a programmatic interface for composing, automating, and extending observability workflows.
Both the CLI and MCP server ship with prebuilt skills: structured workflows built on how Observe engineers solve common observability tasks. Agents and engineers can use these skills to investigate incidents, trace failures, validate changes, find n+1 issues, and detect outliers. They work immediately after configuration, requiring no additional setup. The most significant change is architectural and economic. The original MCP server had a custom LLM harness built in: it received a user query, handled the reasoning, and returned an answer. In the new design, that intermediary is removed. Agents now connect directly to Observe's APIs, which gives them direct visibility into Observe's data structures and permission to explore data sets. This removes the latency and cost of the old LLM middleman. In the old model, agents had to send queries to a single endpoint where Observe's LLM interpreted the question; they lacked direct access to data structure and couldn't explore data sets autonomously—expensive for both Observe and its customers. The new MCP server gives agents access to the same APIs that power Observe's frontend, including APM service maps, OpenTelemetry data collection setup, and active alert listing. Agents can query Observe's context graph directly, determine which data sets are most valuable, and write efficient OPAL queries to run precise observability queries against any telemetry data. General availability is rolling out now across all clusters, with eu-2 and ca-1 regions to follow shortly.
The shift Observe is announcing reflects a fundamental change in how observability platforms must operate: AI agents are now primary consumers of telemetry data alongside humans. Where observability was once a domain exclusively for engineers reading dashboards and logs, agents now investigate errors autonomously, correlate failures across services, and triage alerts before a human is paged. This architectural shift—moving from a UI-centric design to API-first access—is no longer optional.
The removal of the LLM intermediary is the most significant technical change. The old design forced all agent queries through Observe's own reasoning layer, which was costly to operate and gave agents no direct visibility into data structures. Advances in foundation models (Claude, GPT, and others) mean that modern AI agents are sophisticated enough to reason about observability data directly, without needing Observe to interpret the query. By connecting agents to the same APIs that power the Observe UI, the platform gains cost efficiency, reduces latency, and gives agents full autonomy to explore data, query the context graph, and craft precise observability queries. This is also why the CLI and MCP server have full parity: every operation available through the UI is now available programmatically, ensuring no workflow is locked behind a human interface.
The inclusion of prebuilt skills—structured workflows for incident investigation, failure tracing, validation, and outlier detection—lowers the barrier to adoption. Engineers no longer need to write custom agents from scratch; they can compose, automate, and extend common tasks immediately. The real-world use cases mentioned in the article (alert-triage agents that automatically pull production telemetry, copilots that assist during incident investigations, and custom workflows) suggest that the demand for this capability is already present among Snowflake's customer base.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Researchers from Google, the University of Chicago, and other institutions disabled the internal mechanism in…

Apple is taking a focused approach to AI by integrating Apple Intelligence into existing products rather than…

Microsoft's Vice President for Southern Europe, Charles Calestroupat, told Fortune Greece that the region—Gree…

Anthropic's biological and chemical weapons classifiers—filters designed to block dangerous knowledge extracti…

A survey by Epoch AI and Ipsos of 1,106 employed US adults (conducted July 10–19, 2026) found that 20 percent…

Artificial Analysis, known for independent LLM evaluations, has launched Optima, a platform that lets users bu…

The AI news that matters, in one minute each morning.
Sign up free