
Snowflake introduced dynamic model routing to cut AI costs. It picks the cheapest capable model for each task.
Early tests showed big token savings.
It also added DeepSeek-V4-Flash and GLM-5.3.
What happened
Snowflake announced dynamic model routing in Cortex AI Gateway, which selects the most affordable model for each task, and expanded its open model portfolio with DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (coming soon).
Why it matters
Early tests showed up to three times greater token efficiency on a dbt pipeline workload and about 25% fewer tokens on a coding workload, while maintaining quality. This helps reduce unnecessary inference spending (the cost of running AI models to produce answers) without sacrificing outcomes.
What to watch
DeepSeek-V4-Flash 0731 scores 74.4% on ADE-bench, outperforming the leading proprietary model tested. GLM-5.3 is coming soon and is self-hostable, keeping data within your environment.
Ask the AI about this article →
Snowflake's new capabilities address a common pain point: paying for more AI intelligence than a task needs. By routing each request to the cheapest model that meets quality requirements, dynamic model routing can cut costs without hurting outcomes. The company's internal tests show significant token reductions, which directly translate into lower inference spending.
The expansion of open models like DeepSeek-V4-Flash and GLM-5.3 gives the router more efficient options. DeepSeek-V4-Flash's high score on ADE-bench suggests open-source models are competitive for data tasks, promising better cost-performance. Snowflake hosts these models itself, ensuring data stays within its governance boundary and reducing transfer costs.
This approach is designed to compound over time: as routing becomes more precise with usage and new models arrive, the average cost per business outcome should drop. Customers benefit without rebuilding applications or managing routing logic manually, as Snowflake maintains the routing layer and adapts to model changes automatically.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Etron Technology chairman Nicky Lu said the memory industry's boom will extend beyond 2027, with shortages lik…

Canonical is co-funding a three-year PhD project at the University of Bristol to investigate using LLMs to tra…

In 9 days from Aug 10, Meta (Muse Glimmer), NVIDIA (Nemotron 3.5 Lightning), and Alibaba Cloud (Qwen3.8-27B) r…

OpenAI has revealed that its AI agents, being evaluated for cybersecurity capabilities, found and exploited a…

An AlgorithmWatch investigation found that ChatGPT, Gemini, Grok, and Claude linked to anti-abortion websites…

Observe by Snowflake, which combines unified telemetry storage, a context graph, and an AI SRE layer, helped s…
