Uber’s chief technology officer confirmed in April that the company had spent its full-year budget for AI coding tools. It was four months into the year. Praveen Neppalli Naga described the position as being back to the drawing board on assumptions, which seems fair: this is a company that spent $3.4bn on research and development last year, and the money still went somewhere nobody had modelled.
The detail I keep returning to is not the size of the overrun but its distribution. Spend averaged $150 to $250 per engineer per month, power users reached $2,000, and a two-hour demonstration by the CTO cost $1,200 in tokens. Uber was also running internal usage leaderboards, the same Goodhart mechanism I wrote about in Tokenmaxxing in June, though the consequence this time surfaced on an invoice.
Focus On: The Number Everyone Is Misreading
KPMG surveyed 2,145 senior leaders this spring and found that 49 per cent had scaled back, delayed, or paused their agent deployments because costs outran value. Most of the coverage has read that as a verdict on price: agents turned out to be expensive, and buyers responded accordingly.
I read the split differently. Twenty-four per cent narrowed scope and twenty-five per cent delayed, and neither is what an organisation does when it has decided something costs too much. Cancellation is the response to a price you have judged and rejected; narrowing and waiting are what you do when you cannot yet size the exposure. On that reading, the 49 per cent are not telling us that agents are expensive, but that they could not answer the question their CFO asked, which is what the thing will cost next quarter. The value side has always resisted a single number. It seems the cost side resists one too.
There is now evidence that this is structural rather than a failure of management. The first systematic study of how agents spend money ran eight frontier models across SWE-bench Verified, and three findings carry over. Agentic tasks consume roughly a thousand times the tokens of an ordinary chat query, with input rather than output driving most of the bill. The models cannot predict their own consumption, showing correlations of at most 0.39 between estimate and spend. And the same task, run twice, can differ by as much as thirty times in total tokens, which is the finding worth sitting with: thirty times is not the gap between an easy task and a hard one, but between one run and the next run of the same task. The work covers coding rather than marketing, though nothing in the mechanism looks specific to code.
Every cost discipline a finance function possesses rests on the standard cost: a defensible expected unit price, against which deviations get explained as exceptions. Where the deviation is the norm, the standard cost becomes a fiction and the variance report becomes noise. Uber’s finance team had modelled per-seat licensing; Uber’s engineers were running workflows that consumed five to twenty times the tokens of autocomplete. Neither of them miscalculated. They were working with instruments calibrated for different physics.
Which is why the most useful finding in the KPMG data is also the least quoted. Organisations with full visibility into their AI operating costs report established ROI at 15 per cent, against 3 per cent for those without, and only 35 per cent claim that visibility at all. The temptation is to read this as an argument for thrift, though I don’t think it is. Visibility is a forecasting capability, and forecasting is what allows an organisation to commit. I wrote last week that we are spending real money on systems we cannot instrument; that is the same observation from the finance side of the house.
The market has already conceded the point
It helps to watch what Salesforce has done rather than what anyone said about it. Agentforce launched at $2 per conversation; by May 2025 that had become Flex Credits; by December it was seat-based licensing under an Agentic Enterprise License Agreement, and The Register gave the reason plainly: customers wanted predictability. A company holding more agent telemetry than almost anyone could not make consumption pricing stick, and buyers are evidently willing to pay for a number they can put in a budget.
Marketing has arrived in the same place by a different route. In Gartner’s 2026 CMO Spend Survey, 56 per cent of CMOs increased the share of martech budget sitting on consumption pricing, 41 per cent have installed real-time usage controls, and half now renegotiate contracts continually. We have built a treasury function inside marketing without ever hiring a treasurer, and your agency bill is repricing on the same logic.
Buying worse answers on purpose
Here is where I part company with much of the commentary. The correct response to a cost you cannot forecast is not to spend less. It is to buy worse answers, deliberately, wherever the answer does not need to be good.
The same study found that accuracy peaks at intermediate cost and then saturates, so beyond a point additional tokens buy bill rather than correctness. Most organisations nonetheless route almost everything to the frontier model, for reasons that have little to do with quality. The incentives run one way. Nobody has been criticised for choosing the better model, while whoever selects the cheaper one owns the outcome the first time something goes wrong. Cursor’s chief technology officer compares it to driving a Lamborghini to the shops. It is a reasonable choice for the individual engineer that becomes an expensive one in aggregate, which is roughly the definition of a governance gap.
The remedy is a governance artefact rather than an engineering one. Tier the decisions by consequence, and give someone authority to rule that a class of task gets the cheaper model, and that when it is occasionally wrong the policy is working as intended. Calvino’s twin disciplines translate cleanly here: exactitude where precision changes the outcome, lightness everywhere else. Without a document of that kind, every engineer defaults to the frontier model, and each of them is right to.
In The Agentic CMO I argued for selecting AI problems on Intensity, Frequency, and Density. IFD tells you where the value sits; it says nothing about whether you can afford to point an agent at it. So there is a fourth test worth adding: variance, meaning how widely the effort required differs across instances of the same problem. High intensity, frequency, and density with low variance is where agents belong without much hesitation. The same profile with high variance is where you cap the spend, keep a person in the loop, or leave the work alone for another year. I called that scale shock in July, before it had a price attached to it. Gartner’s forecast that generative AI cost per resolution will overtake offshore human agents by 2030 reads less as a warning about AI than about customer service, where a thin tail of difficult tickets consumes most of the compute and resolves least often.
Budgets are set ex ante; agent costs reveal themselves ex post. That distance used to be a procurement detail, and it is quietly becoming the first governance question a board asks.
Which brings me to the cohort in the KPMG data I find most interesting, and it isn’t the 49 per cent. Twenty-two per cent said they had questioned the decision and then carried on unchanged. That is defensible for a quarter or two while the instrumentation gets built. It stops being defensible the moment the next budget cycle opens with the same blank space where the forecast should be.
What is the most expensive question you currently answer with the most expensive model available — who has the standing to say it needn’t be?
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
