
Enterprises are adopting a new AI cost-saving tactic: use expensive frontier models like Anthropic's Fable 5 only for planning, then route routine tasks to cheaper models.
Enterprises worried about AI spending are finding a new cost-saving tactic: use the most expensive frontier models only for high-level planning, then hand the routine work to cheaper models.
Ameya Kanitkar, chief technology officer of the San Francisco-based AI measurement platform Larridin, said companies should treat models like Anthropic's Fable 5 as they would a high-priced lawyer. "You wouldn't use the most expensive lawyer for filing some routine things that less expensive lawyers can," he told Business Insider in July.
Kanitkar said the advisory model works by having a powerful model create a roadmap of workflows and break problems into smaller sets. "The advisory model basically plans things, breaks down the problems into smaller sets, and has the complete context of how everything's going to work," he said. "And then sub-tasks are delegated to cheaper models."
Larridin advises companies on AI use, giving them visibility into how effectively employees use AI tools and how to get a better return on AI spending.
Michael Murphy, a partner at the Sydney-based AI consultancy Adaptovate, said it is not a good use of company dollars to send "the most powerful model out there to do something that's replacing a Google check or transcribing meeting minutes or creating a creative brief." Frontier models should be used to strategize, create initial builds of a new app or website, or handle tasks that require complex thinking, he said. Companies should then figure out which "lightweight flashlight models" are most appropriate for day-to-day tasks.
The approach has support beyond the consulting world. Coinbase CEO Brian Armstrong posted on X in June that he expects "80% of workloads will be running on 99% cheaper models within 12-18 months." The best models should be kept for "IQ maxxing," he said, referring to tasks like scientific breakthroughs or agent orchestration.
Cost-saving tactics like this are gaining traction as companies become increasingly concerned about not getting proportional returns on their AI spending. Many have abandoned tokenmaxxing, a trend where companies gave employees free rein to experiment with AI and urged them to burn as many tokens as possible. Duolingo even made AI usage a performance metric.
Now companies are being more conservative. Several budget hacks have emerged, such as model routing or using open-source Chinese models like Moonshot AI's Kimi K3 or Z.ai's GLM-5.2.
A new wave of startups is cashing in on the efficiency push. AI-routing companies that help steer developers toward different AI models and monitor for overspending are becoming investor favorites. New York-based OpenRouter announced in May that it had raised $113 million, valuing the company at $1.3 billion. Business Insider was the first to report that OpenRouter competitor Concentrate AI secured more than $5 million in funding.
The advisory model changes the cost calculus. A company that pays premium per-token rates for Fable 5 on every query could cut its bill by routing the majority of work through cheaper models. Kanitkar's firm measures those savings directly for clients. The question is whether enterprises will adopt the discipline to separate planning from execution, or keep defaulting to the most powerful model for every task.
Drafted by a large language model from the source reporting linked above, then screened by automated publishing checks. It is not read by a journalist before publication. Some articles cite our Alpha Score. Verify prices and figures against the original source. Educational coverage, not personalized advice.