On paper, enterprise AI has never been cheaper. Token prices—the usage rate you pay for the data an AI model processes and generates—have fallen sharply over the past year due to fierce competition, the rise of economy-tier models, and the emergence of open-source alternatives. But monthly invoices tell a different story.
The falling per-unit price has occurred simultaneously with a surge in total volume. An AI agent can make dozens of requests, known as calls, to a large language model (LLM) to complete one task. Multi-agent conversations, background coding agents, and expanding context windows add to the volume, which is on track to explode. Goldman Sachs expects token consumption (enterprise and consumer) to increase 24-fold to 120 quadrillion tokens per month between 2026 and 2030.
Combined with integration and training overhead, total AI spending is outpacing what many finance teams budgeted. In the FinOps Foundation’s State of FinOps Report, AI cost management ranked as the No. 1 skill financial operations (FinOps) teams need to develop. And 98% of FinOps teams now manage AI spend directly, up from just 31% two years ago. That's a function scrambling to catch up to a bill that outgrew its owners.
Uber is perhaps the highest-profile illustration of this. The company burned through its entire 2026 budget for AI coding tools—Claude Code and Cursor—in four months, after usage spread across thousands of engineers faster than anticipated. The company's own instrumentation made the problem worse: Internal leaderboards ranked and rewarded consumption with seemingly no corresponding owner on the budget side. Uber's fix was a $1,500-per-engineer monthly cap.
Gartner says, “Organizations should introduce mechanisms such as token thresholds, escalation policies, and automated monitoring to manage usage.”
There's now an organized response to that scrambling. In June, the Linux Foundation launched the Tokenomics Foundation—backed by ServiceNow, Accenture, Google Cloud, IBM, JPMorgan Chase, and others—to build vendor-neutral standards and benchmarks for measuring AI cost and efficiency. Until now, there’s been no neutral way to benchmark token efficiency across models and tech providers.
Although total AI spending includes much more than token charges, the cost of running AI models is determined by two things:
- The price of each token: This is falling.
- The number of tokens used: This is growing faster than expected, partly because of the price drop.
Cheaper tokens have made existing workloads less expensive and new use cases economically viable.
There's a name for this, coined during the Industrial Revolution: the Jevons paradox. William Stanley Jevons observed that more efficient coal engines didn't reduce coal consumption. They increased it by making coal-powered machinery possible in places it hadn't been before.
Apollo Global Management Chief Economist Torsten Slok made the connection explicit for enterprise AI. “This is Jevons paradox in action,” he wrote in a June blog post, noting that “as tokens get cheaper, companies don’t spend less but instead run more AI agents, automate more workflows, and generate more code, pushing aggregate expenditure higher even as the unit cost of intelligence collapses.”
That's the piece finance teams miss when they benchmark this year's AI cost against last year's per-token price. The unit price was never the variable that mattered. What a cheaper unit price unlocks is a thousand simultaneous decisions to use AI that were never profitable before.
At Uber, those decisions belonged to thousands of engineers. At most enterprises, they're distributed across teams, products, and use cases, which means nobody really owns the consequences. That's where the governance problem begins.
When consumption decisions get distributed but budget ownership stays centralized, a usage cap becomes the default control lever.
A cap can be an effective backstop against runaway AI usage, but it's not a panacea. A ceiling can limit how much gets spent, but it provides no information about what that spend bought.
For example, let’s say two engineers have the same spending cap. One initiates a bunch of requests to an LLM that results in a new product feature. The other makes requests using ineffective prompts, and nothing usable comes out of it. Both cost the same, but they can’t be distinguished by their invoices.
According to J.R. Storment, executive director of the FinOps Foundation, mature cloud FinOps teams typically forecast their spend to within 1% to 3% of where they land. But with AI, that level of accuracy has been "totally blown out."
That's what happens when you have a cap but no visibility into what sits under it. The fix isn't a lower cap; it's ownership over the decisions influencing the spend.
Practitioners at the FinOpx X2026 conference say optimization requires visibility and attribution, in that order. “If you don’t know which teams are driving spend, which models are being used for which use cases, and which calls are redundant or cacheable, you’re guessing. And in a cost category that’s growing this fast, guessing is expensive.”
Most teams don’t see the attribution. Without high-quality attribution tags (which track who’s using AI and what they’re using it for), enterprises can’t allocate AI spend with precision or identify where to optimize. As a result, they risk unchecked costs as workloads scale. Organizations that skip this step lose the ability to answer a basic question: Where is the money going? While a cap can tell you how much, it can’t tell you why.
This is not one problem with one owner. Engineering teams, for example, generate costs through building and deploying AI, whereas end users and customers drive costs through prompting AI apps. Those are two different problems that need two different solutions. A cap tuned to engineering usage won't cover customer usage, for example. Combining all usage into one line item makes an invoice illegible.
This is also where the market is building. ServiceNow Cloud Cost Management includes an AI control tower. It’s not a tool to reduce your token bill (Jevons paradox makes that less likely anyway), but a way to make cost decisions auditable and owned. It surfaces which teams and workflows are driving spend, what those workflows produced, and which business priorities each spend is tied to.
"AI spend is climbing faster than most enterprises can account for," says Ofer Vaisler, vice president of AI Control Tower product management at ServiceNow. "AI Control Tower closes that gap with a comprehensive approach: seeing every dollar spent, controlling it so costs never surprise you, optimizing to get the most out of each dollar, and maximizing ROI [return on investment] by aligning spend with business priorities. That's how enterprises turn AI cost into AI value."
It's a sign of where enterprise AI governance tooling is headed: not toward tighter caps, but toward the visibility that tells organizations whether their caps are set correctly.
That’s the difference between a backstop and a legitimate strategy. The cap can stay. The differentiator is whether businesses can tell the difference between a token that created value and one that didn't.
Find out how ServiceNow can help you control AI cost and reduce wasteful spending.