top of page

The Token Trap: Why Enterprise AI Budgets Are Breaking

Evidence of the Token Budgeting Crisis


Jonathann Low, Partner and Co-founder of Predictiv Consulting, explained that the AI industry's move from flat-rate fees to token-based billing, which charges for every line of code generated, has become too expensive for many companies. Microsoft used Anthropic’s Claude Code, whose bill is priced by the token, and the costs ran way past the annual budget. This led it to terminate its license with Claude by 30th June 2026 and move its own engineers onto its own Github Coptilot CLI.  Uber's CTO acknowledged that his company had burned through its entire Claude Code budget for 2026 by April. Its president, Andrew Macdonald, then stated he couldn’t draw a clear connection between rising token consumption and an increase in useful consumer features. Duolingo reversed its decision to evaluate developer performance using AI activity metrics after discovering that usage does not reliably correlate with value creation. It shows that the first wave of enterprise AI adoption was built on economic assumptions that are already breaking.



The Root Causes of Token Overspend


Token spend has now simultaneously become a finance, engineering design, and governance problem. These different departments don't have a single, unified system that tracks the cost, so it leads to the creation of an unmanaged strategic liability. In traditional software if an application maxes out its computer processor (CPU utilization), then it alerts the engineers. However, autonomous AI agents don't work in that manner. If there is any flaw with the AI agent, then  instead of a system crash, the first warning that there is an issue comes up through a devastating monthly bill, which shows millions of tokens were being used in the background. Additionally, organizations many times have to re-prompt an LLM multiple times to get a single usable output. However, it still needs to pay for the discarded token just to get to the valuable data.


The Road Ahead


The Linux Foundation announced its intent to launch the Tokenomics Foundation, an initiative focused on establishing open industry standards, benchmarks, and best practices for the economics of AI infrastructure, working in close partnership with the FinOps Foundation.  Research from Goldman Sachs projects global token usage will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month, with the inference market expanding from roughly $106 billion in 2025 to $255 billion by 2030. Industries don't create standards bodies for temporary problems; they create them when a problem is becoming structural and when no single vendor can be trusted to define the standard.


To control rising AI costs, companies have started to use smaller models, limit prompt lengths, and cache duplicate requests. They also set up real-time spend alerts, use smarter data retrieval (RAG), and certify architects to pick budget-friendly options. There is  a requirement of outcome-based pricing models and strict financial governance over AI infrastructure. Organizations that continue treating tokens as an unlimited resource may discover that the largest obstacle to AI adoption is not technological capability, but cost management.



This article has been authored by Uddhav Gupta, LegalTech Fellow at the Indian LegalTech Network and a student at Maharashtra National Law University, Nagpur.

 

 

 

 

 

 
 
 

Comments


The Indian LegalTech Network (ILTN) connects legal innovators across India to collaborate, share, and lead the future of law and technology. Become a member now!

Email: contact@indianlegaltech.net

Phone: +91 98151 34913

Send us a message, we'll get back to you shortly!

Connect with Indian LegalTech Network (ILTN)

  • LinkedIn
  • Instagram

© 2025 by Indian LegalTech Network. 

bottom of page