
Snowflake is launching dynamic model routing—a system that automatically assigns each AI task to the cheapest model that can handle it—and expanding its open-model library with DeepSeek-V4-Flash and GLM-5.3.
Internal testing shows this approach cuts token usage by up to 3× on some workloads while maintaining quality.
The capability learns and adapts as new models arrive, letting enterprises reduce AI costs without rewriting code or managing routing logic themselves.
What happened
Snowflake announced dynamic model routing in Cortex AI Gateway, which automatically directs each AI request to the lowest-cost model capable of completing the task, plus added DeepSeek-V4-Flash 0731 and GLM-5.3 (coming soon) to its open model portfolio.
Why it matters
Early tests show up to 3× better token efficiency on data pipelines and ~25% fewer tokens on coding tasks—meaning organizations can cut AI inference spending without rebuilding applications or sacrificing output quality. As more capable, cheaper models become available, the routing system improves continuously without manual reconfiguration.
What to watch
DeepSeek-V4-Flash 0731 scores 74.4% on ADE-bench (outperforming the proprietary model Snowflake tested); GLM-5.3 is in private preview, designed for self-hosting with low token footprint. Both run inside Snowflake's security boundary, keeping data governed.
Ask the AI about this article →
AI spending has outpaced business value for many organizations deploying agents and applications, because defaulting every request to the most powerful model wastes budget on tasks that need far less intelligence. Snowflake's framing—"intelligence efficiency," the ability to turn compute and models into measurable business value—addresses this by matching each task to the cheapest model that still delivers required quality. Dynamic routing automates this decision-making without requiring development teams to hardcode model choices or rebuild applications as the model landscape shifts.
The economics compound because the routing layer benefits from a growing pool of capable open models. DeepSeek-V4-Flash's 74.4% score on ADE-bench (Snowflake's data-engineering benchmark) surpasses the leading proprietary model in that evaluation, while GLM-5.3 trades some reasoning power for extremely low token overhead—ideal for high-volume, cost-sensitive workloads. By hosting inference directly rather than proxying third-party APIs, Snowflake keeps data in its own governance perimeter and can optimize for the access patterns enterprise workloads actually produce. As new models arrive and existing ones improve, the system learns which tasks fit which models, progressively reducing the average cost per business outcome without requiring application redesign.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic is privately hoping to file for its initial public offering by the end of this month, targeting a ra…
Broadcom is reportedly seeking to borrow up to $100 billion in debt financing to support growth efforts at Ant…
As AI technology matures, the bottleneck in the industry is moving beyond semiconductor constraints like GPUs…

Elice Group, a South Korean AI infrastructure provider, announced the launch of the country's first AI data ce…

Tencent Holdings is receiving shipments of Nvidia H200 AI chips under a new Chinese policy that permits limite…

On August 12, AT&T's Chief Data and AI Officer said OpenAI models power about 25% of the telecom's total AI us…
