
What happened
Snowflake announced dynamic model routing in Cortex AI Gateway, which selects the most affordable model for each task, and expanded its open model portfolio with DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (coming soon).
Why it matters
Early tests showed up to three times greater token efficiency on a dbt pipeline workload and about 25% fewer tokens on a coding workload, while maintaining quality. This helps reduce unnecessary inference spending (the cost of running AI models to produce answers) without sacrificing outcomes.
What to watch
DeepSeek-V4-Flash 0731 scores 74.4% on ADE-bench, outperforming the leading proprietary model tested. GLM-5.3 is coming soon and is self-hostable, keeping data within your environment.
Summaries like this, in your inbox every morning.
Snowflake's new capabilities address a common pain point: paying for more AI intelligence than a task needs. By routing each request to the cheapest model that meets quality requirements, dynamic model routing can cut costs without hurting outcomes. The company's internal tests show significant token reductions, which directly translate into lower inference spending.
The expansion of open models like DeepSeek-V4-Flash and GLM-5.3 gives the router more efficient options. DeepSeek-V4-Flash's high score on ADE-bench suggests open-source models are competitive for data tasks, promising better cost-performance. Snowflake hosts these models itself, ensuring data stays within its governance boundary and reducing transfer costs.
This approach is designed to compound over time: as routing becomes more precise with usage and new models arrive, the average cost per business outcome should drop. Customers benefit without rebuilding applications or managing routing logic manually, as Snowflake maintains the routing layer and adapts to model changes automatically.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
IBM's Bruno Aziza said the number of AI agents employees build will outpace what companies can manage, so firm…
Mikita Balesni, Jasmine Wang, and Tomek Korbak said on October 8 they were fired the previous week and that no…

AMD said on Sept. 28 it agreed to buy World Labs, which makes AI models of 3D spaces and robot-training tools…

Nathan Lambert published an essay arguing AI progress will accelerate through engineering and infrastructure g…

Google DeepMind's Pushmeet Kohli said AlphaFold did not solve protein folding, because proteins are disordered…

ALPHA FORGE's new sandbox.py calls the same inference orchestrator as the daily batch but never calls the ledg…
