AIToday
Large Language ModelsAI Business & IndustrySnowflake AI BlogPublished: Aug 24, 2026, 16:01 JST

Snowflake launches dynamic model routing to cut AI costs

Snowflake launches dynamic model routing to cut AI costs

3 Key Points

  1. What happened

    Snowflake introduced dynamic model routing in Cortex AI Gateway, which automatically selects the most affordable model that can confidently complete each task during agent execution. The company also expanded access to open models DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (coming soon).

  2. Why it matters

    Using the most powerful model for every request can increase costs without meaningfully improving outcomes, Snowflake said. In internal tests, dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used approximately 25% fewer tokens while maintaining the same pull-request throughput.

  3. What to watch

    DeepSeek-V4-Flash 0731 is available in CoCo today in private preview, and GLM-5.3 is coming soon and is self-hostable. DeepSeek-V4-Flash scores 74.4% on ADE-bench, outperforming the leading proprietary model Snowflake tested.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Snowflake frames these announcements around what it calls intelligence efficiency: matching each task with a model that delivers required quality at the lowest appropriate cost, rather than defaulting to the most expensive option. The company argues that a router is only as good as the pool of models it can choose from, which is why it is expanding open model access alongside the routing capability.

Early internal testing supports the efficiency claim. Dynamic routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach, and a coding workload used about 25% fewer tokens. On ADE-bench, DeepSeek-V4-Flash scored 74.4%, outperforming the leading proprietary model Snowflake tested.

Snowflake also emphasized that it serves these open models itself rather than proxying a third-party API, so inference runs within a secure perimeter near governed data. This is intended to reduce transfer cost and latency while keeping data within Snowflake's governance boundary. The company positions these capabilities as compounding: as routing decisions adapt to new models and real-world usage, a greater share of workloads can be handled by efficient models over time.

FAQ
What is dynamic model routing?
It is a Snowflake Cortex AI Gateway capability that selects the most affordable model that can confidently complete a task at each step of agent execution. Lower-complexity tasks go to efficient models, while deeper reasoning goes to frontier models.
Which new open models did Snowflake announce?
Snowflake expanded access to DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (private preview coming soon). These join models from Anthropic, Google, OpenAI, SpaceXAI, Mistral AI and Meta.
Snowflake AI BlogRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCyberAgent creates AI-first unit AIX Design Office