From the AI ROI Series, recorded 11 August 2026. Two headlines that week are worth translating into tokenomics. Intel asked Wall Street for $15 billion and Wall Street handed over $20 billion, its first share sale since 1971, with the reason given in the filing being general corporate purposes. Then Mark Zuckerberg, in a manifesto about abundance, called compute finite and therefore carrying an opportunity cost, which is the line most people skipped.
It is the same message from both. Compute gets more expensive from here, even while the unit price per token keeps falling on paper. Both of those are true, which is why the rate card is the wrong thing to watch.
Waste rarely looks like failure
We measure the value of AI, which means most of our time goes on measuring the waste, and the awkward thing about waste in an AI programme is that it usually looks like growth. Adoption climbs, usage climbs, and spend climbs with them, and everybody agrees it is going well. A rising bill on a good rollout and a rising bill on a wasteful one look identical from the outside. You cannot tell them apart without knowing what it costs to finish one piece of work, which is the number almost nobody has.
Eighteen months, and the number nobody in the building had seen
One of ours, no names. Eighteen months into an AI programme, adoption spreading, engineering shipping agents, coding proficiency already high, everybody pleased. Then around April the invoices started creeping as usage-based pricing kicked in. Not a spike, which somebody would have noticed. A bit more every month, always with a good reason attached, because more teams were using more of it. Somewhere upstairs, somebody finally asked what they were actually getting for it, which is the question we now get asked more than any other.
They proved it themselves, as it happens, since they are proficient users and we only assisted with the initial lift. The thesis was what it cost them to finish one piece of work. Over those eighteen months, while spend tripled and adoption climbed, the cost of finishing one piece of work went up about 40%. Remember what the rate card was doing over the same period. Every model they used got cheaper per token, and the cost to finish a job still rose 40%.
Then we split it by agent
Which nobody had done either. They had dozens running. Four of them were carrying about 70% of the measurable return. The bottom third was consuming roughly 30% of the spend and returning almost nothing. Every one of those started life as an experiment, as it should have, and that is the pattern I think is defining enterprise AI right now. Teams were told to find AI use cases, engineering did what good engineering does and built agents, tried things, shipped them. Some worked brilliantly and some did not, and not one of them was ever switched off. They did not have an AI programme so much as dozens of experiments with a budget and no way to tell which was which.
We ended the bottom tier, moved that money into the four that were working, and they got next year’s growth out of this year’s envelope. The invoices stopped creeping.
Which is the part worth keeping, because we are not here to save costs on enterprise AI. We are here to help organisations make the most out of it, and those are different jobs with different answers. This is directional and it is one company, so check it against your own estate rather than mine.
A second example, from a different engagement
An engineer built an agent to clean up a codebase. Not a production system, just a tool he wrote himself. He started it on a Thursday evening and went home for the weekend, and it ran for four days. About a quarter of the work it attempted actually finished. The rest was the same loop, over and over, on one file it was never going to fix, because nobody had told it when to stop trying. That came to $21,000 of compute across a long weekend, roughly three quarters of which bought nothing at all.
Nobody noticed for four days, and it was not negligence. I know that team, and they are robust about process and QA. Even wasting three quarters of everything it touched, that agent was still cheaper than paying a person to do the same work, by a lot. Every dashboard in the building was green and the ROI was positive the entire time it was setting money on fire, which is exactly why it ran all weekend. The returns on agents are often good enough to hide almost any amount of waste and still show you a number you are happy with.
This is not one clumsy tool, either. A paper published five days before we recorded priced every action and told agents their budget. The best one stayed inside it under 4% of the time, and doubling the budget barely changed the behaviour. Agents cannot manage their own money.
Three fixes, none of them clever: give the agent a stopping rule, cache the part of the context that never changes, and stop sending every job to the most expensive model when a cheaper one finishes it. The cost of getting one job done fell by about seven eighths, for the same amount of work out of the other end. We did not make the agent smarter. We stopped paying for the work that never finished.
Why this is a forecasting problem
Those invoices began as a visibility problem, became a cost problem, and ended as a board question, in that order. Everyone I talk to has more agents planned for next year and nobody has fewer, and most are watching the same two numbers that company was, which will look excellent right up until somebody asks what they bought. That is survivable while compute is cheap. Intel just raised $20 billion because it will not be, and when that price moves it moves under all of it, including the agents returning nothing. Cheap tokens fund waste rather than fixing it, and at the moment they are funding the fabs too.
So the question is not whether you will run more agents next year, because you will. It is whether you will be able to say which of them earned it, which is a question about what your agents actually produced rather than about how many you shipped. The same discipline that makes a build-or-buy decision answerable makes this one answerable, and it starts from the same place: a baseline you measured rather than assumed.
Four questions to take into your next review
- Do you know what is running in your enterprise?
- Are you ready for what is next?
- Are you ahead of the competition?
- Are you making the most out of your tokens?
Those are table stakes for every AI leader through to the CFO, and they need answering with data behind them rather than with confidence. If your answer to the first one is a list of tools rather than a list of outcomes, that gap is the same one behind a falling rate card and a rising bill, and it is why measured AI ROI and the metrics underneath it are worth building before the next budget cycle rather than during it.
I’m Paul, co-founder of Olakai. Measuring what AI actually costs and what it actually returns, on your own workload, is the work I spend my days on, and I am generally happy to be told I am wrong. Schedule your AI evaluation, and we will show you your own record, in your own environment.
