AIToday
Large Language ModelsOpen-Source AIAI Business & IndustrySnowflake AI BlogPublished: Aug 24, 2026, 19:00 JST2 min read

Snowflake Launches Dynamic Model Routing to Cut AI Costs

Snowflake Launches Dynamic Model Routing to Cut AI Costs

Key takeaway

  • Snowflake introduced dynamic model routing to cut AI costs. It picks the cheapest capable model for each task.

  • Early tests showed big token savings.

  • It also added DeepSeek-V4-Flash and GLM-5.3.

3 Key Points

  1. What happened

    Snowflake announced dynamic model routing in Cortex AI Gateway, which selects the most affordable model for each task, and expanded its open model portfolio with DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (coming soon).

  2. Why it matters

    Early tests showed up to three times greater token efficiency on a dbt pipeline workload and about 25% fewer tokens on a coding workload, while maintaining quality. This helps reduce unnecessary inference spending (the cost of running AI models to produce answers) without sacrificing outcomes.

  3. What to watch

    DeepSeek-V4-Flash 0731 scores 74.4% on ADE-bench, outperforming the leading proprietary model tested. GLM-5.3 is coming soon and is self-hostable, keeping data within your environment.

Ask the AI about this article →

Context & Analysis

Snowflake's new capabilities address a common pain point: paying for more AI intelligence than a task needs. By routing each request to the cheapest model that meets quality requirements, dynamic model routing can cut costs without hurting outcomes. The company's internal tests show significant token reductions, which directly translate into lower inference spending.

The expansion of open models like DeepSeek-V4-Flash and GLM-5.3 gives the router more efficient options. DeepSeek-V4-Flash's high score on ADE-bench suggests open-source models are competitive for data tasks, promising better cost-performance. Snowflake hosts these models itself, ensuring data stays within its governance boundary and reducing transfer costs.

This approach is designed to compound over time: as routing becomes more precise with usage and new models arrive, the average cost per business outcome should drop. Customers benefit without rebuilding applications or managing routing logic manually, as Snowflake maintains the routing layer and adapts to model changes automatically.

FAQ

How does dynamic model routing work?
It selects the most affordable model that can confidently complete each task, directing simpler tasks to efficient models and complex ones to frontier models. It only considers administrator-approved models and logs each routing decision.
What is DeepSeek-V4-Flash 0731's performance?
It scores 74.4% on ADE-bench, outperforming the leading proprietary model tested. It is available in private preview in CoCo.
Snowflake AI BlogRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleTeachers Hit by AI Deepfakes as Students Cross Line