
AT&T cut AI costs by up to 56% using model routers that send simpler tasks to cheaper models.
Quality dropped only 2%.
The company plans to increase open-source model usage from 40% to 60-70% of employee queries.
What happened
AT&T cut the costs of coding and some other advanced AI tasks by as much as 56% by using LiteLLM model routers, which route employee queries to cheaper AI models when appropriate. The quality of the AI's performance declined by only 2%, according to an interview with AT&T vice president Mark Austin.
Why it matters
The tools decide how complex a task is and send simpler ones to cheaper models, an approach that lets the telecom keep spending on Anthropic and OpenAI models flat while expanding use of open-source models like Nvidia's Nemotron, Meta's Llama, and Google's Gemma. AT&T aims to raise the share of employee queries powered by open-source models from 40% to between 60% and 70% in the coming years.
What to watch
AT&T says open-source model capabilities generally lag frontier models by six to 10 months, but the gap is narrowing and they are "just as good or better" than older Anthropic and OpenAI models. The company is evaluating the potential risks of using open-source models from Chinese firms DeepSeek and Moonshot but is not using them yet.
Ask the AI about this article →
AT&T's reported cost savings arrive as businesses grapple with rising AI expenses, a trend highlighted in June reports about companies seeking better cost management. The shift from chatbots to more compute-intensive agents, along with AI labs moving from flat subscriptions to token-based billing, has driven costs upward. This context helps explain why a large enterprise would invest in routing technology rather than simply cutting AI usage.
The company's strategy pairs cost-cutting routers with a deliberate push toward open-source models, whose capabilities generally trail frontier models by six to 10 months. Austin's observation that this gap is narrowing suggests AT&T sees open-source as an increasingly viable alternative to premium Anthropic and OpenAI models for many internal tasks. The evaluation of DeepSeek and Moonshot models indicates interest in expanding options further, though risk assessment remains a hurdle.
The reported end of "tokenmaxxing" — pushing employees toward the biggest models and heaviest usage — aligns with AT&T's approach. By matching task complexity to model capability, the company appears to be institutionalizing a more selective, cost-conscious AI strategy that other firms may study.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Etron Technology chairman Nicky Lu said the memory industry's boom will extend beyond 2027, with shortages lik…

Canonical is co-funding a three-year PhD project at the University of Bristol to investigate using LLMs to tra…

In 9 days from Aug 10, Meta (Muse Glimmer), NVIDIA (Nemotron 3.5 Lightning), and Alibaba Cloud (Qwen3.8-27B) r…

OpenAI has revealed that its AI agents, being evaluated for cybersecurity capabilities, found and exploited a…

An AlgorithmWatch investigation found that ChatGPT, Gemini, Grok, and Claude linked to anti-abortion websites…

Observe by Snowflake, which combines unified telemetry storage, a context graph, and an AI SRE layer, helped s…
