The unit cost of AI inference has fallen sharply since 2024. Yet enterprise AI spending is rising faster than most budgets anticipated. The explanation is volume. A 2024 chatbot answered a question in a single model call. A 2026 agent plans a task, calls tools, reads files, checks its own work and retries when something fails. Gartner analysis from March 2026 puts agentic workloads at roughly 5 to 30 times more tokens per task than a standard chatbot query.
EY illustrated the effect in customer service: a simple linear interaction that cost around $0.04 in 2023 costs about $1.20 in 2026 when handled by an orchestrated, multi-tool system, roughly 30 times more.
Uber became the most cited example. After rolling out an AI coding agent across its engineering organization, the company reported in spring 2026 that its annual AI budget had been consumed within months. The tool worked as intended. The budget had been set for a different consumption pattern.
Why seat-based budgeting breaks
Most AI budgets for 2026 were built on per-seat logic: a license per user per month. Consumption pricing does not work that way. Costs rise with adoption, with task complexity and with the number of agents running in the background. The more successful a deployment is, the faster it outgrows its budget.
Look inside a typical agent session and a large share of tokens is spent before any answering happens:
Each step is billed. And because the conversation grows with every turn, each later question costs more than the one before.
The useful metric is not cost per token or cost per seat, but cost per completed task. Organizations that measure it early can see which workflows are worth scaling and which architecture choices are driving the bill.
Want to model and control your AI costs as usage scales? Talk to our specialists.
Sources