
Snowflake is lowering the cost of AI by introducing dynamic model routing—a system that automatically picks the cheapest AI model capable of completing each task—and by adding open-source models like DeepSeek-V4-Flash and GLM-5.3 to its platform.
Internal testing showed the routing layer cut token usage by up to three times on data pipelines and 25% on coding work, while maintaining output quality.
For businesses running many AI agents, this means they can scale AI applications without paying for overkill model power on routine tasks.
What happened
Snowflake introduced dynamic model routing in Cortex AI Gateway, which automatically selects the cheapest AI model capable of completing each task, and expanded access to open models including DeepSeek-V4-Flash 0731 and GLM-5.3. Internal testing showed routing achieved up to three times greater token efficiency on data pipelines and approximately 25% fewer tokens on coding workloads while maintaining quality.
Why it matters
Organizations typically default to expensive frontier models for every request, inflating costs without proportional gains in output quality. Snowflake's routing layer—which respects existing governance controls and logs every decision—lets businesses match tasks to appropriate model capability without rebuilding applications. As the pool of efficient open models grows, the routing system can direct more workloads to cheaper options, compounding cost savings over time.
What to watch
DeepSeek-V4-Flash 0731 is available now in private preview; GLM-5.3 is coming soon. Both models are self-hostable and run within Snowflake's governance boundary, keeping enterprise data on-premises. DeepSeek-V4-Flash scored 74.4% on ADE-bench (Snowflake's data engineering evaluation), outperforming the leading proprietary model Snowflake tested; GLM-5.2 achieved 66% on the same benchmark with the lowest token footprint in the test.
Ask the AI about this article →
AI costs are accelerating faster than measurable business value, creating pressure on organizations to optimize how they deploy and pay for models. Snowflake's response addresses a fundamental inefficiency: the default assumption that every task requires the most powerful (and most expensive) model available. A weekly status summary and a portfolio risk synthesis do not demand equal intelligence, yet most systems today treat them the same way.
Dynamic model routing solves this by treating model selection as a real-time optimization problem rather than a static choice. The routing layer can adapt automatically as new models emerge and pricing changes, without requiring engineers to rebuild applications or hardcode logic. Because Snowflake controls the routing layer, it can continuously refine decisions based on real-world workload performance. The governance integration is critical: routing decisions are logged and respect existing data residency and compliance boundaries, which enterprise procurement teams require.
The open-model expansion amplifies the economic benefit. DeepSeek-V4-Flash 0731's score of 74.4% on ADE-bench—a data engineering evaluation framework—matching or exceeding proprietary alternatives, signals that the open-source frontier is competitive where it matters for data teams. GLM-5.3's predecessor (GLM-5.2) achieved 66% accuracy while using fewer tokens than any other model in the benchmark, making it attractive for high-volume, cost-sensitive workloads. As more efficient models enter the pool, the routing system has better options for directing requests away from expensive frontier models, compounding savings over time.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.