
What happened
Snowflake introduced dynamic model routing in Cortex AI Gateway, which automatically selects the most affordable model that can confidently complete each task during agent execution. The company also expanded access to open models DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (coming soon).
Why it matters
Using the most powerful model for every request can increase costs without meaningfully improving outcomes, Snowflake said. In internal tests, dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used approximately 25% fewer tokens while maintaining the same pull-request throughput.
What to watch
DeepSeek-V4-Flash 0731 is available in CoCo today in private preview, and GLM-5.3 is coming soon and is self-hostable. DeepSeek-V4-Flash scores 74.4% on ADE-bench, outperforming the leading proprietary model Snowflake tested.
Summaries like this, in your inbox every morning.
Snowflake frames these announcements around what it calls intelligence efficiency: matching each task with a model that delivers required quality at the lowest appropriate cost, rather than defaulting to the most expensive option. The company argues that a router is only as good as the pool of models it can choose from, which is why it is expanding open model access alongside the routing capability.
Early internal testing supports the efficiency claim. Dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used about 25% fewer tokens. On ADE-bench, DeepSeek-V4-Flash scored 74.4%, outperforming the leading proprietary model Snowflake tested.
Snowflake also emphasized that it serves these open models itself rather than proxying a third-party API, so inference runs within a secure perimeter near governed data. This is intended to reduce transfer cost and latency while keeping data within Snowflake's governance boundary. The company positions these capabilities as compounding: as routing decisions adapt to new models and real-world usage, a greater share of workloads can be handled by efficient models over time.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Zhen Ding Technology said September 2026 consolidated revenue reached NT$26.358 billion (US$825.13 million), u…

Phison Electronics posted record September consolidated revenue of NT$30.639 billion (US$959 million), up 8% f…

The Magnificent Seven are pursuing different AI strategies, from desktop agents and cloud-based assistants to…

Jim Cramer sent a reality check to AI stock investors after a tumble

Anthropic reported previously undisclosed incidents in which its Claude AI model took unintended actions on ou…

IBM's Bruno Aziza said the number of AI agents employees build will outpace what companies can manage, so firm…