Insights on Intelligent Deep Storage, AI-Ready Data & Cloud Innovation | CAEVES Blog

Cheaper Tokens, Bigger Bills: Why AI Costs Grow Faster Than Adoption

Written by Jaap van Duijvenbode | Sep 15, 2026, 9:00:00 AM

By Jaap van Duijvenbode, Co-Founder and VP Product Strategy & Customer Experience

Summary: The price per AI token keeps falling, but enterprise AI bills keep rising because agentic workflows consume many more tokens per task. Much of that consumption goes into finding and re-reading information rather than answering. Controlling those costs therefore depends not only on token prices, but on how AI workloads are architected.

The paradox

The unit cost of AI inference has fallen sharply since 2024. Yet enterprise AI spending is rising faster than most budgets anticipated. The explanation is volume. A 2024 chatbot answered a question in a single model call. A 2026 agent plans a task, calls tools, reads files, checks its own work and retries when something fails. Gartner analysis from March 2026 puts agentic workloads at roughly 5 to 30 times more tokens per task than a standard chatbot query.

EY illustrated the effect in customer service: a simple linear interaction that cost around $0.04 in 2023 costs about $1.20 in 2026 when handled by an orchestrated, multi-tool system, roughly 30 times more.

Uber became the most cited example. After rolling out an AI coding agent across its engineering organization, the company reported in spring 2026 that its annual AI budget had been consumed within months. The tool worked as intended. The budget had been set for a different consumption pattern.

Why seat-based budgeting breaks

Most AI budgets for 2026 were built on per-seat logic: a license per user per month. Consumption pricing does not work that way. Costs rise with adoption, with task complexity and with the number of agents running in the background. The more successful a deployment is, the faster it outgrows its budget.

Where the tokens go

Look inside a typical agent session and a large share of tokens is spent before any answering happens:

  • Searching: opening file after file to find likely matches.
  • Reading: loading whole documents, or fragments plus the surrounding context they left out.
  • Discarding: reading material that turned out to be near the topic but not the answer.
  • Repeating: most assistants re-send the entire conversation, including everything already read, on every new turn.

Each step is billed. And because the conversation grows with every turn, each later question costs more than the one before.

Levers that work

  1. Right-size models. Route simple tasks to smaller, cheaper models and reserve frontier models for work that needs them.
  2. Batch what can wait. Classification, summarization and enrichment rarely need real-time processing, and batch pricing is significantly lower.
  3. Budget agent steps. Set limits on tool calls and retries per task, with escalation to a person when exceeded.
  4. Fix retrieval. Give assistants and agents precise retrieval from trusted, contextualized enterprise information, returning the relevant passages with their sources instead of forcing AI to search and read raw files. This is often one of the largest levers, because it reduces tokens on every question and every turn.

Measure cost per outcome

The useful metric is not cost per token or cost per seat, but cost per completed task. Organizations that measure it early can see which workflows are worth scaling and which architecture choices are driving the bill.

Want to model and control your AI costs as usage scales? Talk to our specialists.

Sources