From the AI ROI Series, recorded 21 July 2026. Apple raised prices on Macs and iPads and pointed straight at AI, saying the data centre buildout has driven memory chip costs up faster than they have ever seen and that they can no longer shield customers from it. Microsoft hiked the Xbox. Analysts expect smartphones to rise around 20% this year. Whatever else those are, they are not normal inflation.
So the AI tax has stopped being an enterprise problem and turned up in the phone in your pocket. Which leaves the only question I actually care about here: if the cost of AI is leaking all the way into a MacBook, what is it doing to your token bill?
The arithmetic nobody is connecting
Your vendor will show you a chart of the per-token sticker price trending down, and they are all doing it, and they will call that savings. The sticker did fall, and that part is true. But the moment your enterprise plan flips to pay as you go, the same work runs five to ten times more, because you are now being metered on consumption rather than on seats. Both facts are true at once, which is why the chart and the invoice disagree so violently.
The biggest labs are monetising hard on the back of it. Anthropic just passed OpenAI in business spend, almost entirely because of Claude Code, since coding is the battlefield. And the uncomfortable part, which I do not think gets said plainly enough, is that they make more money when you burn more tokens. We have all become dependent, and we have not yet seen autonomous agents running at full scale. That is the 2027 bill, and it is still coming.
Meanwhile almost nobody can answer the basic question. You have seen your invoices, so you roughly know what AI costs you, but is it working? Only about 15% of companies forecast their AI spend within 10% of reality, and most miss by more than a quarter, which means most teams are being driven by need rather than by science. That is exactly why boards are now turning to their executives and telling them to get a handle on it.
Even Google is playing the same game
Before anyone calls Google the cheap exception, look a little closer, because Google is running the smartest version of the same play. A free tier good enough to live on, Gemini bundled straight into your Workspace seats, and an API priced below cost, all to get you embedded before the meter starts to matter. The meter is still there, sitting behind the bundle for now. And when even Google has to cap Meta’s compute and tell them to use fewer tokens, the capacity ceiling and the pricing that follows it are quite real. Nobody is exempt from that, including the people selling you the exemption.
Tokenmaxxing got us here. Tokenwising is how we make it pay.
Tokenmaxxing was the 2025 story: burn everything, more is better. It was a reasonable place to end up, honestly, because it proved the value at a point when the value was still in question. 2026 is about tokenwising, which is spending like it is your own money. Here is what that looks like on a real floor, without naming names.
The pilot
A client of ours, an engineering org. Over the last 60 days we ran a pilot with about 30 developers, roughly a third of the department. The goal was never to brag about lines of code, since shipping thousands of lines of AI-written code is a vanity metric, and I have said so about token leaderboards often enough. The goal was quality: code that passes testing, ships, and moves the business.
Today about half their code is AI-written, and they want to push past 80% in the next six months. That is aggressive, it is expensive, and it is exactly the sort of thing you should not do blind. So we measured three things: who is getting real lift and who is simply burning tokens, which model each task actually needs rather than defaulting to the most expensive one available, and where the waste hides, which is in the re-prompts, the abandoned runs, and the agents quietly looping. None of that was about slowing them down. It was about making the 80% push a science project rather than a guess.
One concrete piece of it is model routing. Match the model to the task and the subtask, and stop burning your most capable and most expensive model on mundane coding work where it produces no better result. Done well, that honestly takes 40% to 50% off the bill on those tasks, which is the difference between an 80% target that pays for itself and one that quietly bleeds. Route on the wrong unit, though, and you will cut the bill while destroying the value underneath it, which is a trap I walked into publicly and corrected later.
The part I am most excited about, which is real guardrails
Here is what is actually new. Your vendors will let you set one big account-level spend limit, and at enterprise scale that is close to useless, because it does nothing to stop a single agent going rogue and eating half the quarter’s budget over a weekend. Budgets, forecasting, and visibility with alerts baked into the workflow are largely solved at this point. The hard part is enforceable limits, real ones, at the project, team, department, and even the individual-developer level. Not a warning after the money is gone, but an actual ceiling, which is the thing that lets a chief data officer sleep. It is early, and I suspect a lot of you are quietly wrestling with the same problem, so I will go deeper on it in coming episodes. Some of the mechanics are already written up in how the budgets and alerts work.
The move
Tokenmaxxing got us here, and tokenwising is how the AI transformation actually pays. The check I would run this month is narrow enough to finish in an afternoon. Do you know your cost per unit of output, rather than your cost per seat or your total invoice? Do you know which of your developers are getting genuine lift, measured against something, and which are producing volume? And if one agent ran unattended over a weekend, is there anything in your stack that would stop it, or only something that would tell you about it on Monday?
Most organisations can answer none of the three, which is the whole reason a measured view of AI ROI and a real record of what your coding tools are doing matter more this year than they did last year. The falling rate card is going to keep making the case that things are getting cheaper, and your invoice is going to keep disagreeing, and only one of those two has your name on it.
One question to take into your next standup: how much of your code is AI-written right now, and do you know whether it is making you money or simply making more code?
I’m Paul, co-founder of Olakai. Measuring what AI actually costs and what it actually returns, on your own workload, is the work I spend my days on. Schedule your AI evaluation, and we will connect to what you already run and show you your own record, in your own environment.
