AT&T cuts AI coding costs 56% with model routing

AT&T has reduced AI coding costs by 56% with only a 2% decline in performance quality by routing routine employee requests to cheaper open-source and open-weight models through LiteLLM. Its internal Ask AT&T platform processes about 45 billion tokens daily, with open models handling roughly 40% of employee AI queries and the company targeting a 60% to 70% share in the near term. Models from Nvidia, Meta and Google are used for less demanding tasks, while premium systems from OpenAI and Anthropic remain available for complex work. Between February and July 2026, AT&T also tested telecom-tuned models for network troubleshooting, customer-service scripts and internal documentation, achieving up to 90% savings in inference costs at scale. The newer figures differ from earlier company estimates of a 25% open-model share, a 70% to 80% target and a workforce of roughly 100,000 employees; the latest account describes a workforce of roughly 150,000, suggesting different measurement periods or scopes.

The information on this website is generated using AI and we cannot guarantee its accuracy. Please use it as reference information only.