Zendoric
← Back to the day · July 20, 2026

93% of companies blow their agentic AI budget: the problem isn't the token, it's the self-review loop

🕒 Published on Zendoric: July 20, 2026 · 00:19

McKinsey documents that 93% of corporate AI teams exceed their budget and that 60% of agentic AI spending goes to the agent itself reviewing and regenerating its answers. The price of a token has plunged more than 99% in two years and yet the bill has tripled: the problem is architecture, not rates.

By MarketScale (McKinsey & Company) · July 20, 2026.

The starting figure looks like a contradiction: the cost of inference for a model with GPT-3.5's capability fell from $20 per million tokens to $0.07 over the course of 2024, according to Stanford HAI's AI Index 2025 — a drop of more than 99%. And yet, enterprise spending on language models tripled over the following twelve months through the end of 2025, according to data from Menlo Ventures. McKinsey investigated that contrast and offers a concrete answer: 93% of enterprise AI teams are already over budget, according to its Enterprise AI FinOps survey, conducted in May 2026 across 75 organizations in five sectors.

The main cause has a name: 'response refinement,' the loop in which an agent reviews, corrects and regenerates its own answer before delivering it to the user. That process — repeating, checking and retrying — consumes 60% of total spending on agentic AI (AI systems that autonomously carry out multi-step tasks without a human intervening at each step), versus 40% for all other agentic operations. For any team that had assumed a more or less balanced split in spending, the finding forces a complete recalculation of where the problem lies.

McKinsey identifies three forces that combine to inflate the bill. First, scale: the more tasks are delegated to agents, the more tokens (the smallest units of text a model processes, and the basis of its billing) are consumed in total, even if each one costs next to nothing. Second, the shift in commercial model: providers have moved from flat rates to consumption-based pricing, which incentivizes longer, more elaborate responses rather than more efficient ones. Third, and the most correctable: many companies use frontier models — the most expensive and powerful — for routine tasks that a small model would handle just as well and far more cheaply.

The effect is already showing up in real decisions, not just spreadsheets. According to McKinsey's forthcoming State of AI 2026 survey (1,719 participants, May–June 2026), one in five organizations has already limited its use of AI specifically because of operating cost. It is a measurable brake on adoption, not a theoretical concern: it affects which model is chosen, which workflows are automated and which are left for later.

There is a line in the report, attributed by McKinsey to David Tepper, CEO of Pay-i, that sums up the shift in focus the authors call for: 'tokens are not value, tokens are the bill.' It is an important operational distinction. If the cost-cutting debate is anchored to the price per million tokens, you are looking at the wrong indicator; what matters is the relationship between the quality of the agent's output and the total cost of producing it. An agent that runs more review loops but delivers a result that never needs human correction may, in practice, come out cheaper than one with a lower price per token whose answers have to be constantly checked.

Our reading is that this story confirms something we have been observing in other analyses: the competitive edge in enterprise AI no longer lies solely in which model you choose, but in the engineering built around it. Just as we saw with data instrumentation and agentic scaffolding in production, here the decisive factor is whether an organization knows how to measure what its agent does — hit rate, need for human oversight, real cost per task completed correctly — or whether it simply pays the bill without understanding where it comes from. McKinsey makes it clear: most companies still lack that instrumentation, which turns the executive question ('is this agent worth it?') into one that almost no one can answer with data today.

In the short term, this is real friction, and it should be said plainly: cost overruns, agents deployed blindly, budgets that run out before proving a return, and organizations that are starting to slow adoption out of sheer financial caution. It is exactly the kind of mismatch that accompanies any technology moving from pilot to production at full speed, and it should not be downplayed: a fifth of companies are already scaling back the scope of their AI projects for this reason.

But the very figure that opens the report — a drop of more than 99% in the unit cost of inference in barely two years — is proof that the underlying curve keeps moving in the right direction. The problem is not that AI is getting more expensive; it is that we are using it badly: frontier models for trivial tasks, pricing incentives that reward the agent's verbosity, and an absence of quality metrics that would allow trimming spending that generates no value. These are problems of design and governance, not physical limits on the cheapening of compute. The companies that build that measurement layer first — cost per task completed correctly, not cost per token — will be the ones to capture soonest the abundance this technology promises: doing much more with far fewer resources. Those that keep paying for review loops they do not understand will keep financing, without realizing it, the most expensive and least efficient phase of the sector's collective learning curve.

🔗 Related on Zendoric

Sources & references