AIToday
Large Language ModelsOpen-Source AIAI Business & IndustrySnowflake AI BlogPublished: Aug 23, 2026, 06:00 JST3 min read

Snowflake cuts AI costs with smart model routing and open models

Snowflake cuts AI costs with smart model routing and open models

Key takeaway

  • Snowflake announced dynamic model routing that automatically picks the cheapest AI model for each task.

  • Tests showed token use dropped 25% on coding and tripled on data pipelines.

  • The company also added DeepSeek-V4-Flash and GLM-5.3 open models to give routing more cost-effective options.

3 Key Points

  1. What happened

    Snowflake is rolling out dynamic model routing in Cortex AI Gateway, which automatically selects the cheapest model capable of completing each task, alongside new open-model access to DeepSeek-V4-Flash 0731 (private preview) and GLM-5.3 (coming soon). Internal tests showed the routing approach achieved up to three times greater token efficiency on a data pipeline workload and maintained the same pull-request throughput while using approximately 25% fewer tokens in coding tasks.

  2. Why it matters

    Organizations pay for AI per token used; most default to expensive frontier models even for simple tasks. Snowflake's routing selects the right-sized model for each task and adjusts as new models emerge, letting teams lower AI spend without rebuilding applications or hardcoding model choices. The routing layer also logs every decision and respects existing governance rules, keeping data compliance intact.

  3. What to watch

    DeepSeek-V4-Flash 0731 scored 74.4% on ADE-bench (Snowflake's data-engineering benchmark), outperforming the leading proprietary model tested; GLM-5.2 scored 66% on the same benchmark with the lowest token footprint, signaling where open-source models now compete on cost and quality for data teams.

Ask the AI about this article →

Context & Analysis

AI budgets are rising but business returns are not always keeping pace—a dynamic Snowflake frames as a mismatch between model capability and task need. Most organizations route every request to the most powerful (and most expensive) model available, even for straightforward tasks like generating summaries. Snowflake's response targets this waste through two levers: intelligent routing and a broader roster of open models.

Dynamic model routing sits inside Cortex AI Gateway and respects the governance rules already in place—role-based access control, data residency, audit logging—ensuring compliance and visibility. The early results are concrete: a data pipeline workload achieved three times higher token efficiency compared to always using a frontier model, and a coding workload maintained throughput while cutting token use by approximately 25%. Critically, the routing adapts as new models enter the market and existing ones improve, so organizations do not need to rewrite applications to benefit from those shifts.

The addition of capable open models—particularly DeepSeek-V4-Flash, which outperformed Snowflake's tested proprietary baseline on data-engineering tasks, and GLM-5.2, which delivered the lowest token footprint in the benchmark—expands the practical options available to the router. The system compounds: more model choices give routing more cost-effective branches to select, and as real-world usage data accumulates, routing decisions become more precise. Over time, a greater share of workloads can shift to efficient models without compromising outcomes.

FAQ

What is dynamic model routing and how does it work?
Dynamic model routing in Cortex AI Gateway automatically selects the most affordable model that can confidently complete each task. Lower-complexity and repetitive tasks go to efficient models, while deeper-reasoning workloads route to frontier models. Routing decisions adapt as models and pricing change, without requiring teams to rebuild applications.
What open models is Snowflake adding?
Snowflake is expanding access to DeepSeek-V4-Flash 0731 (available in private preview today) and GLM-5.3 (private preview coming soon). These join existing models from Anthropic, Google, OpenAI, Mistral AI, Meta, and SpaceXAI. Both new models are self-hostable and keep data within Snowflake's governance boundary.
How do these open models perform compared to proprietary ones?
DeepSeek-V4-Flash scored 74.4% on ADE-bench (Snowflake's data-engineering benchmark), outperforming the leading proprietary model tested. GLM-5.2 scored 66% on the same benchmark with the lowest token footprint of any model tested, making it cost-effective for high-volume workloads.
Snowflake AI BlogRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCanva cuts forecast as AI costs soar; enterprises shift toward cost control