Subscribe Home Conversations On AI App Development CRM Enterprise IT Ethics & Governance Futures HR Industries ServiceNow on ServiceNow Platform Foundations Products & Solutions All topics For Leaders In IT & Dev Customer Experience Finance, Operations & Strategy Employee Experience Security & Risk News & Events People & Culture My List Explore All
August 26, 2026 4 min Cheaper tokens, higher AI cost: What you need to know Budgets don’t fail on price alone. They fail when usage scales faster than accountability. AI Thought Leadership
Lisa Lee
Lisa Lee Writer, ServiceNow
Tokens that look like coins falling through the air
Top takeaways AI spending can rise even as model use gets cheaper, as lower prices make more workloads worth running. Finance and tech leaders need visibility into which teams, workflows, and use cases are driving AI costs. Usage caps can limit overruns, but lasting control depends on tying AI usage to ownership and outcomes.
Alt text

On paper, enterprise AI has never been cheaper. Token prices—the usage rate you pay for the data an AI model processes and generates—have fallen sharply over the past year due to fierce competition, the rise of economy-tier models, and the emergence of open-source alternatives. But monthly invoices tell a different story. 

The falling per-unit price has occurred simultaneously with a surge in total volume. An AI agent can make dozens of requests, known as calls, to a large language model (LLM) to complete one task. Multi-agent conversations, background coding agents, and expanding context windows add to the volume, which is on track to explode. Goldman Sachs expects token consumption (enterprise and consumer) to increase 24-fold to 120 quadrillion tokens per month between 2026 and 2030.  
 

The need for AI cost management 

Combined with integration and training overhead, total AI spending is outpacing what many finance teams budgeted. In the FinOps Foundation’s State of FinOps Report, AI cost management ranked as the No. 1 skill financial operations (FinOps) teams need to develop. And 98% of FinOps teams now manage AI spend directly, up from just 31% two years ago. That's a function scrambling to catch up to a bill that outgrew its owners. 

Uber is perhaps the highest-profile illustration of this. The company burned through its entire 2026 budget for AI coding tools—Claude Code and Cursor—in four months, after usage spread across thousands of engineers faster than anticipated. The company's own instrumentation made the problem worse: Internal leaderboards ranked and rewarded consumption with seemingly no corresponding owner on the budget side. Uber's fix was a $1,500-per-engineer monthly cap.  

Colorful lines on a dark blue digital chart

Gartner says, “Organizations should introduce mechanisms such as token thresholds, escalation policies, and automated monitoring to manage usage.” 

There's now an organized response to that scrambling. In June, the Linux Foundation launched the Tokenomics Foundation—backed by ServiceNow, Accenture, Google Cloud, IBM, JPMorgan Chase, and others—to build vendor-neutral standards and benchmarks for measuring AI cost and efficiency. Until now, there’s been no neutral way to benchmark token efficiency across models and tech providers.  
 

The budget-busting conundrum 

Although total AI spending includes much more than token charges, the cost of running AI models is determined by two things:  

  • The price of each token: This is falling. 
  • The number of tokens used: This is growing faster than expected, partly because of the price drop.  

Cheaper tokens have made existing workloads less expensive and new use cases economically viable.  

There's a name for this, coined during the Industrial Revolution: the Jevons paradox. William Stanley Jevons observed that more efficient coal engines didn't reduce coal consumption. They increased it by making coal-powered machinery possible in places it hadn't been before.  

AI spend is climbing faster than most enterprises can account for. Ofer Vaisler VP, AI Control Tower Prod Mgmt, ServiceNow

Apollo Global Management Chief Economist Torsten Slok made the connection explicit for enterprise AI. “This is Jevons paradox in action,” he wrote in a June blog post, noting that “as tokens get cheaper, companies don’t spend less but instead run more AI agents, automate more workflows, and generate more code, pushing aggregate expenditure higher even as the unit cost of intelligence collapses.”  

That's the piece finance teams miss when they benchmark this year's AI cost against last year's per-token price. The unit price was never the variable that mattered. What a cheaper unit price unlocks is a thousand simultaneous decisions to use AI that were never profitable before. 

At Uber, those decisions belonged to thousands of engineers. At most enterprises, they're distributed across teams, products, and use cases, which means nobody really owns the consequences. That's where the governance problem begins. 
 

Caps control the number but can’t read it  

When consumption decisions get distributed but budget ownership stays centralized, a usage cap becomes the default control lever. 

A cap can be an effective backstop against runaway AI usage, but it's not a panacea. A ceiling can limit how much gets spent, but it provides no information about what that spend bought.  

For example, let’s say two engineers have the same spending cap. One initiates a bunch of requests to an LLM that results in a new product feature. The other makes requests using ineffective prompts, and nothing usable comes out of it. Both cost the same, but they can’t be distinguished by their invoices.  

According to J.R. Storment, executive director of the FinOps Foundation, mature cloud FinOps teams typically forecast their spend to within 1% to 3% of where they land. But with AI, that level of accuracy has been "totally blown out."  

That's what happens when you have a cap but no visibility into what sits under it. The fix isn't a lower cap; it's ownership over the decisions influencing the spend. 

As tokens get cheaper, companies don’t spend less but instead run more AI agents, automate more workflows, and generate more code. Torsten Slok Chief Economist, Apollo Global Management

Practitioners at the FinOpx X2026 conference say optimization requires visibility and attribution, in that order. “If you don’t know which teams are driving spend, which models are being used for which use cases, and which calls are redundant or cacheable, you’re guessing. And in a cost category that’s growing this fast, guessing is expensive.” 

Most teams don’t see the attribution. Without high-quality attribution tags (which track who’s using AI and what they’re using it for), enterprises can’t allocate AI spend with precision or identify where to optimize. As a result, they risk unchecked costs as workloads scale. Organizations that skip this step lose the ability to answer a basic question: Where is the money going? While a cap can tell you how much, it can’t tell you why. 

This is not one problem with one owner. Engineering teams, for example, generate costs through building and deploying AI, whereas end users and customers drive costs through prompting AI apps. Those are two different problems that need two different solutions. A cap tuned to engineering usage won't cover customer usage, for example. Combining all usage into one line item makes an invoice illegible. 
 

Visibility is the differentiator 

This is also where the market is building. ServiceNow Cloud Cost Management includes an AI control tower. It’s not a tool to reduce your token bill (Jevons paradox makes that less likely anyway), but a way to make cost decisions auditable and owned. It surfaces which teams and workflows are driving spend, what those workflows produced, and which business priorities each spend is tied to. 

"AI spend is climbing faster than most enterprises can account for," says Ofer Vaisler, vice president of AI Control Tower product management at ServiceNow. "AI Control Tower closes that gap with a comprehensive approach: seeing every dollar spent, controlling it so costs never surprise you, optimizing to get the most out of each dollar, and maximizing ROI [return on investment] by aligning spend with business priorities. That's how enterprises turn AI cost into AI value." 

A ray of light refracting through a prism into a rainbow

It's a sign of where enterprise AI governance tooling is headed: not toward tighter caps, but toward the visibility that tells organizations whether their caps are set correctly.  

That’s the difference between a backstop and a legitimate strategy. The cap can stay. The differentiator is whether businesses can tell the difference between a token that created value and one that didn't. 

Find out how ServiceNow can help you control AI cost and reduce wasteful spending

Next up
Dive into more conversations AI App Development CRM Enterprise IT Ethics & Governance Human Resources Industries ServiceNow on ServiceNow Platform Foundations Products & Solutions All Topics
Stay in the know Join Us
stay in know image
Alt