
AT&T cut employee AI costs by 56% using LiteLLM routers that route simpler queries to cheaper open-source models, with only a 2% quality drop. The company aims to boost open-source usage to 70%.
AT&T reduced the cost of AI-powered coding and other employee tasks by as much as 56% using model routers that send simpler queries to cheaper open-source models, The Information reported Thursday.
The performance decline was only 2%, Mark Austin, an AT&T vice president overseeing internal AI tools, told the publication. The company uses LiteLLM routers to judge a task's complexity before deciding which model should handle it.
AT&T wants to keep spending on models from Anthropic and OpenAI flat while raising the share of employee queries handled by open-source or open-weight models to 60%-70% from 40% today. The open-source models in use include Nvidia's Nemotron, Meta's Llama and Google's Gemma. AT&T is not using models from China's DeepSeek or Moonshot but is evaluating the potential risks, Austin said.
The capability gap between open-source and frontier models has generally been six to 10 months, but that is narrowing. Older Anthropic and OpenAI models are now often matched or beaten by open-source alternatives, Austin said.
The push to cap AI costs comes as companies shift from chatbots to agents, which consume more compute, and as AI labs switch from flat subscriptions to token-based billing. The era of "tokenmaxxing"–pushing employees toward the biggest models and heaviest usage–is ending, PYMNTS reported in July, as new cost-management tools emerge.
Prepared with AlphaScala editorial tooling from the source reporting linked above. Indexable analysis may include a cited Alpha Score value. Publishing checks screen each story before release. Educational coverage, not personalized advice.