AT&T Cuts AI Coding Costs by 56% Using Open-Source Models

iconCryptoBriefing
Share
AI summary iconSummary
AT&T slashed AI coding costs by 56% using open-source models for routine employee queries. The telecom firm uses LiteLLM to route traffic, now handling 40% of requests with open-source tools. Daily token processing hit 45 billion, with a 2% performance dip. Custom telecom-tuned models tested in 2026 cut inference costs by 90% for industry-specific tasks. This shift reflects key AI + crypto news and aligns with broader industry trends in cost optimization.

AT&T found a way to slash its AI coding costs by more than half, and the trick is almost disappointingly simple: stop using the expensive model when a cheaper one works just as well.

The telecom giant implemented model routing technology through LiteLLM that redirects routine employee queries, particularly coding-related ones, toward lower-cost open-source models. The result was a 56% reduction in AI coding costs with only a 2% decline in performance quality. For a company processing roughly 45 billion tokens daily through its internal “Ask AT&T” platform, those savings add up fast.

The routing playbook

The concept behind AT&T’s approach is what the industry calls model routing, essentially a traffic cop for AI queries. Simple questions get sent to lightweight, inexpensive models. Complex tasks still go to premium options from OpenAI and Anthropic.

Advertisement

AT&T VP Mark Austin noted that open-source models are narrowing the performance gap with their proprietary counterparts, with a difference of only 6-10 months in capabilities.

Currently, open models handle about 40% of employee AI queries at AT&T. The company is targeting 60-70% in the near term. The models doing the heavy lifting on the open-source side include Nvidia’s Nemotron, Meta’s Llama, and Google’s Gemma, with AT&T actively evaluating additional alternatives.

Telecom-tuned models push savings even further

AT&T didn’t stop at generic model routing. Between February and July 2026, the company experimented with custom telecom-tuned models designed for industry-specific tasks. Those experiments delivered up to 90% savings in inference costs at scale.

The telecom-specific models are purpose-built for the kinds of queries AT&T employees actually make: network troubleshooting, customer service scripts, internal documentation lookups.

What this means for enterprise AI spending

Goldman Sachs has flagged this trend as potentially advantageous for Big Tech firms, suggesting AT&T’s task-specific routing approach could serve as a template for cost management in enterprise AI deployment.

For other enterprises considering a similar move, the 45-billion-token-per-day figure is instructive. AT&T isn’t running a small pilot. This is production-scale deployment across a workforce of roughly 150,000 employees.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.