
What happened
Snowflake launched dynamic model routing in Cortex AI Gateway, which automatically selects the cheapest capable model for each task, and added access to DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (private preview coming soon) to its open model portfolio alongside existing models from Anthropic, Google, OpenAI, SpaceXAI, Mistral AI and Meta.
Why it matters
Organizations deploying AI agents often pay for frontier-model power on every request even when simpler tasks don't need it. Internal testing showed dynamic routing cut token use by up to 3× on data pipeline work while maintaining comparable quality, and by approximately 25% on coding tasks while keeping pull-request throughput steady. Snowflake hosts inference inside its own governance boundary, so enterprise data never leaves the company's security perimeter.
What to watch
DeepSeek-V4-Flash 0731 scored 74.4% on ADE-bench (Snowflake's data engineering benchmark), outperforming the leading proprietary model tested; GLM-5.3 is designed for high-volume, cost-sensitive workloads and scored 66% on the same benchmark with the lowest token footprint of any model in the test.
Summaries like this, in your inbox every morning.
AI costs are climbing faster than business value for many organizations deploying multiple agents and applications. The problem is simple: most companies default to running every request through the most powerful (and most expensive) model available, even when simpler tasks—like generating a status summary—do not need frontier-level reasoning. Snowflake's answer is twofold: intelligently route each request to the smallest capable model, and expand the pool of those capable models so the router has more cost-effective options.
The routing layer itself sits within Snowflake's existing enterprise governance controls, so administrators can approve which models are eligible, enforce data residency rules, and audit every routing decision. Because Snowflake maintains the routing logic centrally, customers do not need to rebuild applications or hardcode model selection; the routing layer adapts as new models ship and pricing changes. The internal benchmarks are concrete: on a data-build (dbt) pipeline, dynamic routing achieved up to 3× token efficiency versus frontier-only; on engineering pull requests, teams cut token use by approximately 25% while holding throughput steady.
The open models Snowflake is adding—particularly DeepSeek-V4-Flash, which scored 74.4% on Snowflake's data engineering benchmark and outperformed the leading proprietary model tested—signal that the open-source frontier has caught up where it matters most for data teams. This matters because a larger pool of high-quality, efficient options gives the routing layer more leverage to lower the average cost per task. The compounding effect emerges over time: as routing decisions improve with real-world usage and new models arrive, the share of workloads handled by efficient models grows, while enterprise governance remains intact.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
OpenAI told investors it actually had $18 billion less revenue than the $68 billion it said last month, cuttin…
At SailPoint's Navigate event in Austin, CEO Mark McClain said identity security for AI agents has become a bo…
The author predicted a $5,000 investment split between Nvidia and Broadcom will triple by 2028, citing Broadco…

Nvidia is the better pick over AMD, per a Motley Fool analysis

The World Bank's biannual Economic Update argues most African economies should adopt and adapt AI rather than…

Tomek Korbak, Jasmine Wang and Mikita Balesni, fired by OpenAI, released an open letter on Oct
