AIToday
Large Language ModelsAI Business & IndustrySnowflake AI BlogPublished: Aug 24, 2026, 16:01 JST2 min read

Snowflake launches dynamic model routing to cut AI costs

Snowflake launches dynamic model routing to cut AI costs

Key takeaway

  • Snowflake now routes AI tasks to cheaper models automatically. It also added new open models.

  • Early tests show up to three times better token efficiency.

  • This helps cut AI spending without losing quality.

3 Key Points

  1. What happened

    Snowflake introduced dynamic model routing in Cortex AI Gateway, which automatically selects the most affordable model that can confidently complete each task during agent execution. The company also expanded access to open models DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (coming soon).

  2. Why it matters

    Using the most powerful model for every request can increase costs without meaningfully improving outcomes, Snowflake said. In internal tests, dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used approximately 25% fewer tokens while maintaining the same pull-request throughput.

  3. What to watch

    DeepSeek-V4-Flash 0731 is available in CoCo today in private preview, and GLM-5.3 is coming soon and is self-hostable. DeepSeek-V4-Flash scores 74.4% on ADE-bench, outperforming the leading proprietary model Snowflake tested.

Ask the AI about this article →

Context & Analysis

Snowflake frames these announcements around what it calls intelligence efficiency: matching each task with a model that delivers required quality at the lowest appropriate cost, rather than defaulting to the most expensive option. The company argues that a router is only as good as the pool of models it can choose from, which is why it is expanding open model access alongside the routing capability.

Early internal testing supports the efficiency claim. Dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used about 25% fewer tokens. On ADE-bench, DeepSeek-V4-Flash scored 74.4%, outperforming the leading proprietary model Snowflake tested.

Snowflake also emphasized that it serves these open models itself rather than proxying a third-party API, so inference runs within a secure perimeter near governed data. This is intended to reduce transfer cost and latency while keeping data within Snowflake's governance boundary. The company positions these capabilities as compounding: as routing decisions adapt to new models and real-world usage, a greater share of workloads can be handled by efficient models over time.

FAQ

What is dynamic model routing?
It is a Snowflake Cortex AI Gateway capability that selects the most affordable model that can confidently complete a task at each step of agent execution. Lower-complexity tasks go to efficient models, while deeper reasoning goes to frontier models.
Which new open models did Snowflake announce?
Snowflake expanded access to DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (private preview coming soon). These join models from Anthropic, Google, OpenAI, SpaceXAI, Mistral AI and Meta.
Snowflake AI BlogRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCyberAgent creates AI-first unit AIX Design Office

The AI news that matters, in one minute each morning.

Sign up free