
Snowflake now routes AI tasks to cheaper models automatically. It also added new open models.
Early tests show up to three times better token efficiency.
This helps cut AI spending without losing quality.
What happened
Snowflake introduced dynamic model routing in Cortex AI Gateway, which automatically selects the most affordable model that can confidently complete each task during agent execution. The company also expanded access to open models DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (coming soon).
Why it matters
Using the most powerful model for every request can increase costs without meaningfully improving outcomes, Snowflake said. In internal tests, dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used approximately 25% fewer tokens while maintaining the same pull-request throughput.
What to watch
DeepSeek-V4-Flash 0731 is available in CoCo today in private preview, and GLM-5.3 is coming soon and is self-hostable. DeepSeek-V4-Flash scores 74.4% on ADE-bench, outperforming the leading proprietary model Snowflake tested.
Ask the AI about this article →
Snowflake frames these announcements around what it calls intelligence efficiency: matching each task with a model that delivers required quality at the lowest appropriate cost, rather than defaulting to the most expensive option. The company argues that a router is only as good as the pool of models it can choose from, which is why it is expanding open model access alongside the routing capability.
Early internal testing supports the efficiency claim. Dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used about 25% fewer tokens. On ADE-bench, DeepSeek-V4-Flash scored 74.4%, outperforming the leading proprietary model Snowflake tested.
Snowflake also emphasized that it serves these open models itself rather than proxying a third-party API, so inference runs within a secure perimeter near governed data. This is intended to reduce transfer cost and latency while keeping data within Snowflake's governance boundary. The company positions these capabilities as compounding: as routing decisions adapt to new models and real-world usage, a greater share of workloads can be handled by efficient models over time.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Nvidia is reportedly in discussions to invest in AI startup Perplexity as part of a funding round that would v…

SK Hynix is pushing beyond HBM (high-bandwidth memory, a key component in AI chips) to two new areas: HBF and…

dotData explained its approach to using a separate LLM (an AI that understands and generates text) to rate ano…

ITmedia's ICT Research Division tested ChatGPT, Gemini, and Claude Sonnet 5 in July 2026

Anthropic's Claude AI service is currently experiencing an outage, as indicated by the article title

Taiwan prosecutors indicted nine people, including employees of Nvidia and Super Micro, for allegedly exportin…
