From the AI ROI Series, recorded 9 September 2026. Two stories broke this week that most people filed under separate headlines: OpenAI’s Astra, and Anthropic walking away from a six billion dollar acquisition. I want to make the case that they are the same story, and that the story is about the one number that decides who wins in enterprise AI.
Anthropic was reportedly ready to pay up to $6 billion for Decart, a lab whose optimisation engine runs agents at about eight times the industry average. Then, after full due diligence and right before its IPO, it reportedly walked. Both halves of that are reported rather than confirmed, so hold them accordingly.
The number underneath both headlines
The mechanism before the arithmetic, because the phrase doing the work here is one most buyers never see on an invoice. Cost to serve is what it costs a vendor to answer your request, the compute burned turning your tokens into their output. It sits underneath the rate card you are quoted, and it separates a model business that compounds from one that merely grows. An engine running agents at eight times the throughput moves that number directly, which is why it is worth billions to somebody who serves inference for a living.
Anthropic’s own cost to serve, halved in a year
A year ago Anthropic’s margin on serving inference was around 38%. Today it is around 70%. In plain dollars, running the AI used to eat about 62 cents of every revenue dollar and now it eats about 30, which is the same fact said twice rather than two separate findings. They cut their own cost to serve nearly in half in a single year. These are reported figures, so hold them loosely, and note what they cover: this is the margin on serving inference, before training, research, and staff, and it is not net profit.
That qualifier matters if you read what I wrote about Anthropic’s valuation a fortnight ago, where the figure in play was a company-wide gross margin of around 44%. The two are different measures and they sit comfortably together, since gross margin carries a great deal of cost that serving a token does not. The direction of travel is the interesting part. My argument then was that every enterprise pushing work down the price curve takes a point of the vendor’s margin. Here is the same vendor defending that margin from the other side by making each token cheaper to serve, one variable squeezed from both ends.
Why they walked, three honest reads
So you can see why an engine that gives you eight times the throughput is worth billions, and you can also see why they might still walk. There are three honest reads, and they do not contradict each other.
- They already proved they can cut the cost themselves, so why pay $6 billion for more of a lever they are good at pulling?
- The engine comes bundled with a whole video business, which is not their strategy, and you cannot cleanly buy just the meter.
- Walking away from your biggest deal ever, in the week the market started grading on return rather than spend, is a discipline signal, and that is good finance management.
Pick whichever you like. What survives all three is the reason Astra and Decart are the same story. The model on your desk will keep flipping, OpenAI this quarter, Anthropic the next, somebody else after that. But whoever is winning the capability race, every one of them is measuring its cost to serve down to the cent, and one of them ran $6 billion of diligence to move it. The capability race is loud. The economics race is silent, and it is permanent.
Now turn it around, because you live on the other side
You are going to switch model this year based on who is best, and most enterprises will. But can you tell me what any of them actually cost you per outcome, per task? I am not asking about invoice totals. What did one completed task, one shipped feature, one resolved ticket cost you in AI, and what did it give back? For almost everyone I talk to, the answer is no.
And it is not because the teams are not sharp. It is because that data was never captured as a record in the first place. It sits scattered across four vendor consoles that do not talk to each other and were not built to tell you very much, so you cannot even see where to spend less, and none of them know what that token was actually for. This is the same gap that lets a falling rate card sit next to a rising bill, and the same one that made a vendor’s pricing change land under an agent budget without anyone noticing for a week.
Same tokens, two lenses, and they almost never meet
The dilemma is the same one whichever chair you sit in, and it splits cleanly down the middle of most organisations. If you are in finance, you watch the invoice climb and you cannot say whether that is a problem or a good investment. If you are in engineering, you watch the workload climb and you cannot put a dollar on it. It all runs on the same tokens. They are completely different lenses on one number, and the two almost never meet in one place, which is how you end up with two teams arguing from two screens and two spreadsheets.
I do see a shift here that I like a great deal, with people building their own dashboards and their own small solutions, and some of it is genuinely good work. There is always a story about somebody saving 20 minutes on a task. It is real, and it might even be statistically significant, but a pointy saved minute is a long way from a firm-level answer, let alone a board-level one, especially once you are past the pilot and into scale.
Put the two sides together and the asymmetry is the whole point. The seller measures every token to the cent, across every model, and will spend $6 billion of diligence to move the number by a few points. The buyer measures a good afternoon on whichever model is fashionable this quarter.
That gap is the reason Olakai exists. We capture every AI interaction and every outcome across an organisation, down to the token, structure it into one attributed record, and make it usable through your own AI, so you can see cost per completed task and value per outcome by team, by agent, by model, and by vendor. It is one record read through whichever lens is yours: finance reads it as return and budget, engineering reads it as throughput and the cost of what shipped. Data is the product. The record is the instrument.
FY27 budget season, and the questions that matter
The model on your desk is going to keep changing. The one thing that should not change is your ability to measure what it is worth. We are all getting the same question this year whether we like it or not, which is what did it return, and most of us cannot answer it cleanly yet. That is a measurement gap rather than a failing on anyone’s part, and measurement gaps are fixable.
So these are the table stakes questions for this year, and they are for anyone who owns a piece of the AI budget, which means finance, engineering, and AI leaders alike.
- When you switch to Astra, or to whatever comes next, will you know whether it costs you more or less per outcome than the model it replaced?
- Did it make your agents better, and can you show the difference?
- If you are in finance, could you put that number in front of the board on Monday with the evidence behind it?
- If you are in engineering, could you show which agents and which models earned their cost, and which did not?
If any of those is a no, that is the work.
Directional as always. The data is public, so check my math and tell me if you see it differently, because I welcome that all day long. I’m Paul, co-founder of Olakai. Your AI is an investment, so let’s measure it like one.
