
Snowflake is cutting AI costs by routing tasks to the cheapest model that still delivers quality. A new gateway automatically matches each request to the right model, avoiding expensive over-provisioning.
Early tests show three times better token efficiency on data pipelines while keeping quality stable.
The company is also adding open models like DeepSeek-V4-Flash, letting customers pick the best fit for cost and performance.
What happened
Snowflake introduced dynamic model routing in Cortex AI Gateway, which automatically selects the lowest-cost model capable of completing each task, alongside new open models DeepSeek-V4-Flash 0731 (in private preview) and GLM-5.3 (private preview coming soon). Internal tests showed the routing approach achieved up to three times greater token efficiency on data pipeline workloads compared to always using the most powerful model, while maintaining comparable quality.
Why it matters
Organizations pay for more AI capability than most tasks require—a weekly summary does not need the same reasoning power as portfolio risk analysis. By matching task complexity to model cost, enterprises can reduce inference spending without rebuilding applications or managing routing logic themselves. Snowflake's approach also ensures open models run inside the company's data governance boundary, keeping sensitive data within existing compliance controls.
What to watch
DeepSeek-V4-Flash 0731 is available today in private preview; GLM-5.3 comes soon and will be self-hostable. On Snowflake's internal ADE-bench test, DeepSeek-V4-Flash scored 74.4%, outperforming the proprietary model tested, while GLM-5.2 (the predecessor) achieved 66% accuracy with the lowest token footprint in the benchmark, making it suited for high-volume, cost-sensitive workloads.
Ask the AI about this article →
The tension between AI performance and cost has forced companies into an inefficient choice: either overpay for powerful models on every task, or build and maintain custom routing logic to pick cheaper alternatives. Snowflake's framing—"intelligence efficiency"—redefines the question: not which model is best overall, but which model is best suited to each specific task at the lowest appropriate cost. The company's answer has two parts.
First, dynamic model routing in Cortex AI Gateway removes the burden from developers. Rather than engineering teams hardcoding model selection, Snowflake's layer observes each request and chooses the lowest-cost option that can confidently complete it. The system respects administrator-approved models and existing data residency rules, so requests stay within compliance boundaries. As the body notes, routing decisions adapt as new models arrive and relative capabilities shift—without forcing customers to rebuild applications.
Second, expanding access to capable open models directly improves the router's options. DeepSeek-V4-Flash 0731 scored 74.4% on Snowflake's internal ADE-bench test, surpassing the proprietary frontier model tested. GLM-5.2 (the predecessor to GLM-5.3) achieved 66% accuracy with the lowest token footprint of any model benchmarked, making it ideal for high-volume, cost-sensitive workloads. By hosting these models directly rather than proxying third-party APIs, Snowflake ensures inference runs inside enterprise governance and keeps data within existing controls—a meaningful advantage for regulated organizations.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Google Cloud announced a strategic partnership with Verizon to deploy its full-stack AI, including Gemini Ente…

Google Cloud announced a strategic partnership with Verizon to deploy its full-stack AI, including Gemini Ente…

Private equity firms have reportedly begun placing AI experts within their portfolio companies

Nucleai announced an ongoing translational research collaboration with Gilead Sciences, using AI-driven tissue…

BlackRock has announced an AI infrastructure workforce agreement, signaling its entry into the AI-driven inves…

Corning has signed new multiyear, multibillion dollar supply agreements with Meta and Zayo to support AI-drive…
