Design Critique

Managing shadow tokens amid AI market surge

 ·  By Lysandr Foxglove
Managing shadow tokens amid AI market surge - shadow tokens
Managing shadow tokens amid AI market surge

Uber burned through its entire annual budget for AI in just four months. Such a phenomenon creates what some call “shadow tokens,” which are AI credits paid for by the company but largely invisible to decision-makers. Engineers make the final call on consumption, and that often leads to an all-you-can-eat attitude toward spending.

How Unlimited Access Can Blow a Budget

The trend is accelerating. Microsoft is reportedly winding down many internal licenses across key engineering teams, and one in five organizations is missing its AI spend forecast by more than 50%. Gartner predicts that by 2028, AI coding costs per developer will equal the salary companies pay that person. The shift occurs because large language models introduce a variable cost that scales with behavior rather than headcount.

Previously, enterprise leaders knew exactly what a software seat cost. Large language models flip that status quo on its head: the unit of consumption is the action itself, and the running total of the cost is exponential.

This isn’t immediately obvious during a pilot. Tools can look cheap in controlled tests, but they scale unpredictably based on session length and context window size.

The $20-per-seat enterprise plan often misleads executives, because tokens are charged separately at API rates with no ceiling.

Related: DNS Change May Cause Internet Outages

At Uber, engineers adopted Claude Code with few constraints. In agentic mode, the system autonomously read codebases, planning changes across dozens of files and opening pull requests. Each step adds up quickly.

Anthropic’s own documentation notes that agents consume approximately seven times as many tokens as standard sessions.

It resembles the early days of shadow IT, where employees added unauthorized apps to the network without IT’s knowledge.

Shifting the culture from usage to yield

It is easy to blame engineers for this, but the structure itself creates the problem.

When a company rewards developers for shipping new features fast without linking that speed to the literal cost of the electricity and compute power required to generate the code, it creates a perverse incentive.

Related: Travel Tech I Refuse to Leave at Home

The people holding the credit card are not the same people deciding whether the technology is worth the price.

This disconnect forces decision-makers into a reactive posture where they only discover the damage after the financial hit is already significant.

The culture at Uber created a system where more tokens equaled a better ranking. Such an approach leads to quantity-over-quality processes that are hard to control.

Leaders need to shift the focus from usage to yield—the measurable output generated per dollar spent.

Effective governance requires established financial metrics. Companies can set maximum spend limits dictated by spend tagging and cost-per-output benchmarks.

Leave a Comment

Your email address will not be published.