From the AI ROI Series, recorded 28 July 2026. Two things happened that week, and they are more connected than they look. Visa cut 2,600 jobs, about 7% of the company, and the cuts fell mostly on technology and product teams. The CEO’s memo said AI is accelerating the evolution of how work gets done, although Visa’s own people were careful to say AI was not the only reason, and I am not going to overstate it.
Hold on to which teams got cut, though, because technology and product is the exact group that burns almost all of an enterprise AI budget. The second thing is that Anthropic published a position on open-weight models after taking a beating for not signing the open letter. Somewhere between those two stories, half of LinkedIn decided the answer is to stop paying vendors and run the models yourself.
So let us do what we do here, and price it.
First, what Dario actually wrote
The version going around is not the version he wrote, so this part is worth getting right before any arithmetic. He did not call for a ban. He said it plainly: “Anthropic has never advocated for banning open-weight models.” He called open models without the dangerous capabilities a public good, and the safety testing he asked for would apply to Anthropic’s own models too. He asked for three things: keep advanced chips away from authoritarian governments, stop industrial-scale distillation, which is copying frontier models through the API, and test any sufficiently capable model before release, open or closed.
Is he neutral? Of course not, he sells closed models, and you should read him like an S-1. But if this were straight protectionism he would have backed the ban, because banning Chinese open models inside US companies is the single policy that most helps his revenue, and he turned it down. The sentence everyone quoted instead was the one about open models costing nothing besides the compute needed to run them. So we priced the compute.
What actually changed in the pricing
A developer on a flat subscription costs about $200 a month. No meter, no visibility, burn as much as you like. That same developer, doing identical work, billed by the token, costs $1,200 a month or more. Six times, for the same output. Nobody’s usage exploded, the visibility did, and that invoice is what sent everyone hunting for a cheaper answer in the first place.
The arithmetic, on a composite company
The company I carry through these episodes is 250 people, 40 of them developers, running about 115 billion tokens a year. The open-weight example is Kimi K3, the 2.8 trillion parameter model everybody points at. Three scenarios, with every assumption bent in favour of building.
| Scenario | Annual cost | vs buying |
|---|---|---|
| Buy it from a vendor | $620,000 | baseline |
| Build it, engineers already on payroll and reassigned, no new salaries | $748,000 | ~1.2x |
| Build it, with three specialists who can run it in production | $1.47M | ~2.4x |
On the metal alone it is nearly competitive at 1.2 times, which honestly surprised me. Then you put the people back and the arithmetic stops working, because you will need those people whether or not you have budgeted for them. The third scenario also carries $1.58 million of hardware on day one, locked to one model, in a market where something changes every month.
Where the $620,000 actually sits
Before anyone asks whether that figure is just the developers, no, it is all 250 people. The 40 developers burn $576,000 of it. Everybody else accounts for $44,000. So developers are 93% of the spend, which brings us back to Visa, because the teams getting cut are the teams generating almost the entire AI bill.
| Where the money goes when you build | Share |
|---|---|
| People | about half |
| Metal | about a third |
| Power | 3% |
Power at 3% is the number I got most wrong going in, and I expected it to be much bigger personally. I do not know about you, but if anyone is selling you self-hosting on an energy argument, they have not built one, and that holds even against an 18% year-on-year increase in US electricity prices.
The one line to take to your board
Your entire annual AI bill, all 250 people, is $620,000. The three engineers you need to run the thing yourself cost $722,000. Maybe you already have them and maybe you do not, but three engineers cost more than the whole company’s AI bill, before a single server, before power, hosting, or support. All of that assumes $1,200 a month per developer, which is the assumption most likely to be wrong for your team, so run it with your own number. Push the developer usage rate as hard as you like and the metal gets you down to about 1.1 times in my analysis. The staff number never gets there.
It is worth saying that Kimi K3 is not cheap to buy either. It prices the same as Sonnet 5, and it is very capable, but cheap is the wrong word for it at the end of the day. The Wall Street Journal ran a piece the same week arguing AI pricing had peaked and that self-hosting was the era we were entering, which I found quite surprising from that masthead.
Four reasons to build, and cost is not one of them
There are four cases where building is the right call: an air gap, sovereignty, the model being your actual product, or idle hardware of this class that you already own. Those are real, and they are edge cases. Everything else is a measurement problem wearing a procurement costume. Token prices genuinely are falling, that part is true, but consumption per task is climbing faster, and the subsidised era is over. Usage-based pricing is a cash business rather than philanthropy.
What has been bothering me for weeks is the speed of the round trip: a non-expert explains enterprise AI to other non-experts, and within seventy-two hours it is an urban legend that you should just deploy Kimi K3 yourself because it is free. That is a story your uncle tells you after a few beers at the family barbecue.
What to check before you price a build
If you are chasing open weights to save money without knowing your own baseline, you have nothing to compare against, and you will not save anything because you never knew what you were spending. So the check is the boring one. Do you know your current cost per developer per month, measured rather than assumed? Do you know what share of your total AI spend sits with the 40 or so people who generate most of it, which is a question about your own engineering usage data rather than a vendor’s console? And do you know your cost per unit of output well enough that you could tell whether a build actually beat it a year from now?
Those who do not measure end up exposed, which is the same pattern behind the AI bill crowding out other budget lines and behind most of the tool sprawl I see in engineering organisations. A baseline is what makes the build-or-buy question answerable at all, and keeping that baseline across every tool and every token is the part I would put in place before a procurement exercise rather than after one. It is the same discipline that makes routing and AI ROI measurable instead of anecdotal, and it is why the metrics layer comes first.
One question, answerable from your last quarter without looking anything up: if you self-hosted tomorrow, what number would you compare the result against?
I’m Paul, co-founder of Olakai. Measuring what AI actually costs and what it actually returns, on your own workload, is the work I spend my days on. Maybe I am completely wrong here and it works for you, in which case I would genuinely like to hear about it. Your AI is an investment, so let’s measure it like one.
