
Snowflake built an internal semantic layer that standardizes how millions of data sources are understood and queried by both humans and AI agents.
The layer acts as a single translation point between raw data and downstream tools, ensuring consistent definitions and dramatically improving AI accuracy—from 20% to over 90% in benchmark tests—while reducing token cost and latency for agent queries.
What happened
Snowflake built an internal semantic layer—a standardized data translation system that sits between raw data tables and downstream consumers (dashboards, AI agents, analytics tools). In July 2025 alone, over 400 internal users ran 5,400 queries through the product data science agent using this layer; across all Snowflake teams, over 5,600 employees ran 320,000 queries using the broader semantic layer.
Why it matters
Without a shared semantic layer, the same business question (e.g., "What is an active customer?") yields different answers depending on which raw table and metric definition a team uses. A semantic layer solves this by defining business logic once and reusing it everywhere—eliminating conflicting data sources and reducing the token overhead for AI agents. In AtScale's benchmark testing, adding semantic context increased text-to-SQL accuracy from 20% to more than 90% across 40 business questions.
What to watch
Snowflake's semantic views are interoperable with Apache Ossie (incubating), a vendor-neutral open standard for semantic models. Users can start with a single semantic view defined in a YAML or SQL file, materialize it to preaggregate metrics, and declare which dimension-metric combinations need fast performance—letting Snowflake maintain those slices automatically.
Ask the AI about this article →
Snowflake's challenge mirrors a widespread enterprise problem: as organizations scale their data collection to petabytes, data consistency becomes a bottleneck even before performance does. The core issue is semantic fragmentation—multiple teams define the same business concept differently across separate tables, forcing data scientists to act as human arbiters. This friction multiplies when AI agents enter the picture, because agents query at machine speed and scale, surfacing these conflicts faster and more frequently than humans ever could.
Snowflake's response—building a semantic layer internally and documenting the practice—treats semantics as a production software problem requiring versioning, testing, and CI/CD discipline. The July 2025 usage numbers (5,400 queries in one month from the product data science agent alone, and 320,000 across all teams) validate the approach at scale. By curating semantic views with dbt integration and materializing high-query dimension-metric combinations, Snowflake reduced latency and token consumption while enabling both structured dashboards and open-ended natural-language queries on the same trusted data API.
The interoperability with Apache Ossie signals an emerging vendor-neutral standard for semantic metadata, suggesting that this layer will become a platform-agnostic expectation rather than a Snowflake-specific feature.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
On August 20, Broadcom was reported to be negotiating more than $60 billion in fresh debt, with the deal poten…

On August 20, Nvidia denied a report by The Information claiming the company planned small-batch shipments of…

Waymo disclosed details of its onboard computing system for autonomous driving, including a purpose-built 5 nm…

OpenAI's GPT-5.6 Sol, launched July 9, drove a 35 percent revenue increase this quarter, with enterprise reven…

OpenAI is previewing transparent background support for GPT-Image-2 through its API, allowing users to generat…

Nvidia is paying $6 billion for Poolside's 'Model Factory' software system and bringing on 109 employees who w…
