Falling costs per token should make AI cheaper. The opposite is happening, according to Gartner, because complex AI agents consume far more tokens than simple chatbots.
At first glance, the trend sounds positive: The cost of processing tokens with large language models continues to decline, while each new generation of models performs computations more efficiently than its predecessors. Yet according to a recent forecast from market research and consulting firm Gartner, enterprises will have to spend significantly more on AI inference in the coming years. By 2028, inference costs per so-called agentic workflow are expected to be more than five times higher.
The reason is what Gartner calls the “inference paradox.” A better cost structure per unit of computation ultimately leads to higher overall costs, without giving companies a reliable or predictable return. In short, models are getting cheaper, but the tasks assigned to them are becoming more expensive faster than the underlying technology is becoming cheaper.
From Chatbots to Autonomous Agents
The shift away from simple, assistive AI features toward AI agents that can independently execute multi-step tasks is driving the increase. Gartner analyst Will Sommer describes the difference this way: “While a simple chatbot only needs to read and interpret a request and quickly respond with a likely suitable answer, an AI agent must continuously reason, negotiate and question its own conclusions.”
All of this requires computing power and therefore consumes tokens. According to Gartner’s calculations, routing a task to an agentic reasoning model costs providers at least five times as much in inference as a basic chatbot interaction. As task complexity increases, Gartner says the cost multiplier can become significantly higher.
Three Trends Shaping Token Costs
Gartner identifies three fundamental developments shaping what it calls the token economy:
- The cost structure of foundation models is improving rapidly.
- More efficient AI models make it possible to use even more capable, and more expensive, models for demanding use cases.
- More complex AI workflows consume significantly more tokens than simple chat interactions, driving up overall costs.
Tokens are therefore becoming cheaper, but not quickly enough to keep pace with the growth in AI capabilities and the costs associated with them. In other words, the pace of innovation is outstripping the cost curve.
No Shortcut Through “More Efficient” Models
For product leaders, Sommer says this means one thing above all: “Product leaders cannot assume that more efficient token economics will justify the cost of AI.” Every new generation of AI capabilities will require more, and often more expensive, tokens. A universal model that is both highly capable and inexpensive is not on the horizon. Companies looking to build competitive AI products will have little choice but to develop and maintain complex ecosystems comprising multiple models.
To achieve a positive return on investment from advanced AI capabilities such as reasoning agents, Gartner sees two possible approaches. Companies can either make the returns generated by these capabilities grow exponentially alongside the costs, or they can use sophisticated inference tiering, routing and orchestration to assign tasks to the most cost-effective model based on their complexity. Both approaches are feasible, Gartner says, but they require substantial effort across entire workflows.
Sommer sums up the challenge: “Organizations that default to generic autonomous intelligence will face unlimited costs, orders of magnitude higher than those of optimized product ecosystems.”
(Editorial Team)