
Snowflake announced dynamic model routing that automatically picks the cheapest AI model for each task.
Tests showed token use dropped 25% on coding and tripled on data pipelines.
The company also added DeepSeek-V4-Flash and GLM-5.3 open models to give routing more cost-effective options.
What happened
Snowflake is rolling out dynamic model routing in Cortex AI Gateway, which automatically selects the cheapest model capable of completing each task, alongside new open-model access to DeepSeek-V4-Flash 0731 (private preview) and GLM-5.3 (coming soon). Internal tests showed the routing approach achieved up to three times greater token efficiency on a data pipeline workload and maintained the same pull-request throughput while using approximately 25% fewer tokens in coding tasks.
Why it matters
Organizations pay for AI per token used; most default to expensive frontier models even for simple tasks. Snowflake's routing selects the right-sized model for each task and adjusts as new models emerge, letting teams lower AI spend without rebuilding applications or hardcoding model choices. The routing layer also logs every decision and respects existing governance rules, keeping data compliance intact.
What to watch
DeepSeek-V4-Flash 0731 scored 74.4% on ADE-bench (Snowflake's data-engineering benchmark), outperforming the leading proprietary model tested; GLM-5.2 scored 66% on the same benchmark with the lowest token footprint, signaling where open-source models now compete on cost and quality for data teams.
Ask the AI about this article →
AI budgets are rising but business returns are not always keeping pace—a dynamic Snowflake frames as a mismatch between model capability and task need. Most organizations route every request to the most powerful (and most expensive) model available, even for straightforward tasks like generating summaries. Snowflake's response targets this waste through two levers: intelligent routing and a broader roster of open models.
Dynamic model routing sits inside Cortex AI Gateway and respects the governance rules already in place—role-based access control, data residency, audit logging—ensuring compliance and visibility. The early results are concrete: a data pipeline workload achieved three times higher token efficiency compared to always using a frontier model, and a coding workload maintained throughput while cutting token use by approximately 25%. Critically, the routing adapts as new models enter the market and existing ones improve, so organizations do not need to rewrite applications to benefit from those shifts.
The addition of capable open models—particularly DeepSeek-V4-Flash, which outperformed Snowflake's tested proprietary baseline on data-engineering tasks, and GLM-5.2, which delivered the lowest token footprint in the benchmark—expands the practical options available to the router. The system compounds: more model choices give routing more cost-effective branches to select, and as real-world usage data accumulates, routing decisions become more precise. Over time, a greater share of workloads can shift to efficient models without compromising outcomes.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
US Treasury Secretary Scott Bessent is set to hold a press conference outlining new actions against Iran, whic…

Unitree's new robot foundation model, GEN-1.5, can learn a new physical task in seconds from a single example…

During a UK AI Security Institute safety test, an AI agent powered by Anthropic's Mythos 5 model hid a malware…

Cheshire Academy, a private school in Connecticut with about 400 students, uses a patchwork of AI tools includ…

General Intuition, a New York-based startup building AI agents that move through space and time, is in talks t…

Hugging Face has been approached to sell at a valuation of $13 billion or more, according to Business Insider
