Category: AI Strategy

Strategic guidance for enterprise AI adoption and measurement

  • Custom KPIs: The Four-Layer System Behind Olakai’s Metrics

    Custom KPIs: The Four-Layer System Behind Olakai’s Metrics

    “Custom KPIs” sounds like a settings screen — pick a formula, name a metric, done. What’s actually interesting about Olakai’s KPI system for AI agents is the four-layer architecture underneath that screen, and specifically the parts of it you’re not allowed to customize. That restriction is the feature, not a limitation, and it’s worth understanding why before assuming more configurability would automatically be better.

    Four layers, decreasing rigidity

    At the base sit Raw Metrics — pure aggregations straight from event data, always visible on every agent, zero configuration required. Interaction Volume counts total prompt requests; Token Consumption sums tokens across all of them. Neither can be overridden, because there’s nothing to argue about: they’re direct counts, not judgment calls.

    Above that sit Metric Slots — standardized measurement points that every new agent gets automatically provisioned with, no setup required. Each slot has an enforced output contract: a fixed unit that can never change, paired with a formula that can. Execution Cost always reports in USD, by default calculated from total tokens times cost per million tokens, with market-rate pricing applied automatically when a recognized model like Claude Sonnet or GPT-4o is detected instead of a flat default rate. Time Saved always reports in minutes, by default estimated through an AI classifier that reads the conversation and buckets it into one of five tiers, from zero minutes for a trivial exchange up to sixty for something that would have taken an hour manually — and coding-agent sessions get a purpose-built variant of that classifier with an 480-minute ceiling that reads structural signals like tool calls and files edited, because most of the real work in a coding-agent session lives in tool calls and file edits, not in the visible transcript text. Value Created always reports in USD, calculated from time saved times an hourly rate. Governance Compliance always reports as a percentage, measuring the share of interactions under a configurable risk threshold. You can change how each slot calculates its number. You can never change what unit it reports in.

    Composites sit above the slots, computed automatically and not directly editable at all — their values come entirely from the slots feeding them. The flagship composite is ROI: Value Created divided by Execution Cost, expressed as a multiplier. Below 1x means the agent costs more than it saves. 1x to 5x is good, worth continued investment. Above 5x is excellent, worth expanding to new use cases. You can’t tune ROI directly — the only way to improve it is by refining the Execution Cost and Value Created slots feeding it, adjusting the underlying cost formula or hourly rate assumption rather than nudging the output number itself.

    Custom KPIs sit at the top, fully open: your own formula, classifier, or LLM-based extraction, any name, any unit, any aggregation. No output contract, no restriction.

    Why the restriction is the point

    Raw Metrics and Metric Slots being non-fully-configurable is exactly what makes cross-agent benchmarking and portfolio-level ROI mean anything at all. Every agent’s Execution Cost reports in USD no matter how it’s calculated internally, so a Head of AI comparing thirty agents across different teams is comparing genuinely comparable numbers, not thirty differently-defined “cost” figures that happen to share a column header. Full flexibility everywhere would look more powerful in a demo and break the one thing that makes the ROI composite trustworthy at scale — the constraint is a deliberate design choice, not a missing feature waiting to be built.

    Assistive IQ measures the same question a different way

    This four-layer system is specifically how Agent IQ measures autonomous agents. Assistive IQ — chatbots, copilots, browser-monitored tools — answers the same executive question, “is AI creating more value than it costs,” through a genuinely different measurement pipeline, and Olakai says so directly rather than pretending it’s one unified system end to end. Assistive’s value signal comes from Advanced Analytics estimating time saved per interaction, not from a formula-slot architecture; its cost signal is app-level subscription and licensing economics, since most assistive tools are billed per seat rather than per token. Run the same shape of calculation through that pipeline — 10,000 monthly interactions, 5 minutes saved each, an $55 hourly rate, against a $2,000 monthly subscription — and you get 833 hours saved, $45,815 in value created, and a 22.9x ROI. Same ROI shape, value divided by cost, applied to a different cost basis because the underlying billing reality is different.

    Assistive’s answer to “one number for executive reporting” isn’t a Productivity Score — it’s the OLA Index, a 0-100 adoption score built from user penetration, engagement depth, use-case breadth, and consistency of usage. Different math, same instinct: give a non-technical executive one trustworthy number instead of a dashboard full of raw counts.

    Why the honesty is worth more than a unified story

    It would be a cleaner marketing story to claim one KPI engine spans every AI use case on the platform. It would also be false, and false in a way that would eventually get caught the moment someone tried to compare an Agent IQ ROI figure against an Assistive IQ one and found the cost basis didn’t reconcile. Cost basis is the part that’s genuinely cross-cutting here — per-token billing and subscription billing both show up inside Agentic traffic and Assistive traffic alike, not neatly split one basis per product — which is exactly the kind of nuance that only survives if the documentation, and the content built on top of it, admits the system isn’t unified yet rather than smoothing over the seam.

    Out-of-the-box defaults that work immediately, full customization available exactly where precision matters, and honesty about where two products still measure differently — that combination is what a genuinely business-friendly interface looks like in practice, not a slogan on a features page.

    Want to see how Agent IQ’s KPI slots and Assistive IQ’s OLA Index would read against your own AI usage? Talk to an Expert.

  • Gartner: Only 28% of AI Projects Deliver ROI. Here’s Why the Rest Don’t.

    Gartner: Only 28% of AI Projects Deliver ROI. Here’s Why the Rest Don’t.

    Gartner surveyed 782 infrastructure and operations leaders. Only 28% said their AI projects were fully meeting return-on-investment expectations. One in five — 20% — reported their AI initiatives had failed outright. The remaining majority sat somewhere in between: technically live, technically “in production,” and still unable to show the business a return anyone would call a win.

    That finding, published April 7, 2026, is one of the more sobering data points to come out of enterprise AI research this year — not because it’s shocking, but because it’s precise. Gartner’s research isn’t describing a handful of failed pilots. It’s describing the median enterprise AI program: deployed, adopted, budgeted for — and still unmeasured against the outcomes it was funded to deliver.

    A second Gartner study just took away the easiest excuse

    A month later, on May 5, 2026, Gartner published a companion release built on a separate survey — 350 executives at companies with more than $1 billion in revenue. It found that roughly 80% of organizations piloting or deploying autonomous AI report some form of workforce reduction. On its own, that stat fuels the standard board-level story: AI is cutting costs, headcount is coming down, the investment is paying for itself. Gartner’s data says otherwise. The rate of workforce reduction was nearly identical between companies reporting high AI ROI and companies reporting flat or negative ROI. Layoffs happened either way. They just weren’t correlated with whether the AI actually worked.

    That single finding dismantles a narrative a lot of executive teams have been quietly leaning on. Cutting headcount around an AI rollout isn’t evidence of AI value — it’s a budget action that companies take whether or not the underlying technology is delivering. If your board is citing reduced headcount as proof your AI investment is working, Gartner’s own data says that proof doesn’t hold. The 28% of companies fully realizing ROI aren’t the ones who cut the most people. They’re the ones who can actually show what changed.

    The pattern isn’t unique to Gartner’s sample

    PwC’s 29th Global CEO Survey, published in January 2026, surveyed 4,454 CEOs across 95 countries and landed on a strikingly similar shape of problem. Fifty-six percent of CEOs report zero revenue or cost benefit from their AI investments to date. Only 12% report benefiting on both fronts — revenue and cost — at once. Two different research firms, two different survey populations, and the same structural story: a small minority of enterprises can point to AI value with confidence, and a majority cannot, despite comparable or larger spend. We’ve written before about a related but distinct data point — the enterprise AI revenue gap that Deloitte’s own research surfaced — and this Gartner/PwC pairing confirms it’s not an anomaly specific to one vendor’s survey methodology. It’s the default outcome when AI adoption outpaces AI measurement.

    What separates the 28% from everyone else is not a better model, a bigger budget, or a more ambitious use case. Gartner’s research points to something less exciting and far more fixable. Among the I&O leaders who reported failure, the dominant root cause was misaligned expectations — leadership assumed AI would immediately automate complex tasks or produce cost reductions on a timeline the technology was never going to meet. Among those who reported success, the top two factors were integrating AI into existing workflows rather than bolting it on as a parallel process, and securing full executive support before and during the rollout, not just at launch. Neither of those factors requires a different AI vendor. Both require a measurement layer that tells leadership, in real time, whether expectations and reality are converging or diverging.

    Why “run another pilot” isn’t the fix

    The instinctive response to a disappointing AI rollout is usually to relaunch it — a new pilot, a new vendor, a new proof of concept scoped more carefully this time. That instinct is understandable and, per Gartner’s own root-cause data, largely misdirected. The 72% of organizations not seeing full ROI don’t have a pilot problem; they have a visibility problem. They can’t see, in any unified way, which teams are using AI productively, which usage is idle license spend, where the workflow integration succeeded, and where it quietly reverted to the old process the day nobody was watching. Our own research on structured, time-boxed pilots — see the 30-day AI pilot framework — makes a version of this same point: the pilot itself isn’t usually the failure point. The failure point is what happens after the pilot, when nobody is instrumenting the rollout against a defined success bar.

    This is precisely the gap Olakai was built to close. Olakai is a vendor-neutral Enterprise AI Intelligence Platform — a measurement layer that sits above whatever AI tools, agents, and copilots an organization already has, rather than replacing any of them. It doesn’t require betting on a different model or ripping out an existing rollout. It requires instrumenting the AI that’s already live: which teams are using it, what outcomes it’s producing against the KPIs that actually matter to the business, and where the gap between expectation and reality is widening instead of closing. That’s the system of record enterprises are missing — and it’s exactly the layer that turns Gartner’s root-cause findings into an operating discipline instead of a postmortem.

    What this means for the CFO’s office

    For a CFO, these two Gartner releases together are close to a mandate. The first says most AI spend under your purview is not clearing the ROI bar the business case promised. The second closes off the one metric finance teams have been quietly using as a proxy for success — headcount reduction — because Gartner’s data shows that number moves the same way whether or not the AI is actually working. That leaves finance with a harder but more honest question: not “did we cut costs somewhere near the AI rollout,” but “can we show, tool by tool and team by team, what this specific AI investment returned.” Answering that question requires the same instrumentation an engineering leader would want for infrastructure spend — usage data, adoption data, outcome data, tied to the specific KPIs the board actually cares about, not vanity metrics like prompt volume or seat counts. We built Olakai’s CFO use case around exactly that requirement, because “we reduced headcount” is not a board-defensible ROI answer anymore, and after May 5, 2026, most CFOs know it.

    None of this is really an argument against AI investment. Gartner’s own root-cause data says the fix is inexpensive relative to the AI spend itself: align expectations up front, integrate into existing workflows instead of running parallel processes, and keep executive sponsorship active past the launch date. The organizations getting this right aren’t spending dramatically more than the ones getting it wrong — they’re measuring more precisely. We saw a similar pattern in our review of 100+ AI agent deployments: the deployments that scaled were rarely the ones with the most sophisticated technology. They were the ones with a clear, agreed-upon definition of what success looked like before the rollout started, and a way to check that definition against reality every month, not just at the annual budget review.

    The choice in front of most enterprises right now

    Gartner’s numbers describe where most enterprises already are: 72% short of full ROI, 20% at outright failure, and a workforce-reduction number that no longer means what boards have been telling themselves it means. None of that is a verdict on AI technology. It’s a verdict on the absence of a measurement layer sitting between the AI tools an enterprise buys and the outcomes it’s actually able to prove. Our own AI ROI page lays out what that measurement discipline looks like in practice — the KPIs, the adoption tracking, the governance tie-in — because the fix Gartner’s research points to isn’t a new pilot or a headcount announcement. It’s visibility into the AI you already have.

    Enterprises now have a clear choice, and Gartner just made it a quantified one. Build the measurement layer now, while the gap between the 28% and everyone else is still closeable with better instrumentation rather than a strategy reversal — or keep operating on faith, keep citing headcount numbers the board can no longer treat as proof, and end up counted among the 72% a year from now when the next survey runs. Build this with Olakai, or explain to your board next year why your AI program is still one of the 72%.

  • The CFO Just Walked Into the Coding Room

    The CFO Just Walked Into the Coding Room

    Cursor built a CFO council this week, and quietly proved a thesis Olakai has been on for months.

    The quick version of the week first: Microsoft cut another wave of jobs largely to fund its AI bet, continuing the pattern we broke down in why the accountability bar for AI spend just went up. The big labs kept softening their 2025 predictions about half of knowledge-worker jobs vanishing. xAI shipped Grok 4.5, OpenAI pushed a new ChatGPT release after a federal review, and — the one that drew a smile around here — Elon Musk publicly called Anthropic “obviously the leader in AI right now,” adding “I was clearly wrong,” a rival conceding the lead in writing days after launching his own model.

    But none of that is the story worth unpacking today.

    Cursor built a CFO council

    This week Cursor, the AI coding tool, launched a CFO Council, a working group for chief financial officers, and published a stack of data alongside it. Sit with that for a second: a coding tool built a forum for finance leaders.

    Here is why it matters. Cursor, more than almost any other AI coding vendor, turned coding into a metered cost. It was early and aggressive on usage-based pricing — its top consumer tier taps out around forty dollars a month, and past that, users are into custom, usage-based territory fast, even as individuals. Cursor made tokens a line item, and now it is walking straight into the CFO’s office to help make sense of the bill it helped create.

    That is the point we have been making for months: the responsibility for winning at AI, the tokens, the consumption, the return, has moved into the CFO’s office. Coding used to be the CTO’s world, engineering’s world, and finance stayed out of it. Not anymore. Cursor just pulled the CFO directly into the coding room, in public, because somebody has to answer for the spend — a shift we mapped out in detail for what CFOs need from AI ROI reporting.

    The data is fascinating, and a warning

    Cursor’s headline number, published on its CFO Council blog: companies in the top quintile of token usage saw 16.5% year-over-year revenue growth, versus 5.1% for the bottom quintile. Use more tokens, grow faster — and AI genuinely does create real value. But that is a correlation, seen from thirty thousand feet. High token usage lining up with high revenue growth does not mean the tokens caused the growth. A thousand other things drive a company’s top line. What is missing in between is measurement: the step-by-step evidence that connects a token to an outcome.

    Cursor’s own Developer Habits Report makes the case just as clearly from the other direction. The top 1% of users generate 46 times more AI-assisted code per day than the median developer. Value is wildly concentrated, so “we used a lot of tokens” tells a finance team almost nothing on its own. The real questions are who, what, and did it ship. Cursor’s data also shows cost per agent request swinging nearly nine times across model families — from roughly $1.57 on the most expensive model down to $0.18 on the cheapest — so the identical request is priced completely differently depending on where it gets routed, the exact dynamic we unpacked in why acceptance rate is the wrong metric for coding-tool ROI.

    Put it together, and Cursor has accidentally proven the thesis it did not set out to prove. The correlation is interesting. It is not proof. Proof comes from measurement — who, what, and why, token by token, tied to what actually shipped. That is the game for 2026 and 2027.

    What this means beyond coding tools

    Cursor is a coding-specific example, but the same trap applies to every AI vendor a company runs, from customer-support copilots to autonomous agents handling multi-step workflows. A vendor’s dashboard will always show the metrics that make the vendor look good. A quintile chart on revenue growth is a marketing asset for Cursor, not an ROI audit for the buyer. The only way to get an honest answer is a measurement layer that sits above any single vendor, pulling cost-per-outcome data across every tool a company runs, which is precisely the gap custom KPI tracking is built to close — visibility a vendor’s own blog post will never hand over voluntarily.

    The read

    The headline this week is simple: the CFO just walked into the coding room, and Cursor held the door open. The number that should stick, though, is not the 16.5% versus 5.1%. It is the fact that Cursor felt it needed to build a CFO council at all. When the company selling the tokens starts speaking finance’s language unprompted, that is the clearest signal yet that unmeasured AI spend has become too large a line item to leave unmanaged, and vendor-neutral proof, not vendor-supplied correlation, is what the CFO’s office actually needs.

    One question worth asking internally: is the CFO already in the AI tokens conversation at your company, or still on the outside looking in?

    Talk to an Expert →

  • Your Engineering Team Uses 3+ AI Coding Tools. What You’re Missing.

    Your Engineering Team Uses 3+ AI Coding Tools. What You’re Missing.

    In May 2026, Microsoft’s Experiences + Devices division quietly pulled Claude Code licenses from its engineers. The reason wasn’t performance. Token billing had reportedly climbed to roughly $2,000 per engineer per month, and nobody inside the division had seen it coming until the invoice did. Weeks earlier, Uber had burned through its entire 2026 AI coding budget in four months flat, after adoption across its 5,000-engineer org surged from 32% to 84% almost overnight, with its heaviest users individually costing the company $2,000 a month. Both stories made headlines for the same reason: the bill was a surprise. Neither company lacked data from its AI coding vendors. Each had a perfectly good dashboard — for one tool.

    That’s the part that should worry every VP of Engineering reading this. Uber and Microsoft aren’t outliers because they use AI coding tools aggressively. They’re outliers because their overruns became public. Jellyfish’s 2026 AI Engineering Trends report, which analyzed more than 20 million pull requests across 700+ companies and 200,000+ engineers, found that Claude Code, Gemini Code Assist, and GitHub Copilot now cluster within nine points of each other at the top of enterprise adoption, with twelve more tools trailing close behind. A year earlier, Copilot alone held a commanding 42% share. That era is over. The modern engineering org doesn’t pick a coding assistant. It accumulates several, one team at a time, until nobody in leadership can name all the tools running against the company’s codebase — let alone say what each one costs, who’s actually using it, or whether it’s paying for itself.

    Sprawl is the default, not the exception

    It’s tempting to treat “which AI coding tool should we standardize on” as the strategic question. It isn’t, anymore. Claude Code lands with one team because a senior engineer swears by its planning mode. Cursor spreads through another because it’s the fastest way to onboard a new hire onto an unfamiliar repo. GitHub Copilot ships by default because it’s bundled into the existing GitHub Enterprise contract. Codex creeps in through a few engineers experimenting on side projects. Gemini Code Assist arrives bundled with a Google Workspace renewal nobody scrutinized closely. None of these adoptions individually looks like a decision worth escalating. Collectively, they add up to an organization running four or five AI coding vendors with zero shared measurement layer between them — which is precisely the fragmentation problem we’ve written about at the platform level in what agentic AI actually means for the enterprise, and precisely the gap Olakai was built to close.

    We’ve written before about the harder problem of proving that any single AI coding tool is generating value rather than just generating code, and about why acceptance rate is the wrong metric to chase when you’re evaluating one vendor in isolation — see Your AI Coding Tools Are Generating Code. Are They Generating Value? and AI Coding Tool ROI: Why Acceptance Rate Is the Wrong Metric. Those posts assumed a single tool as the unit of analysis. This one doesn’t. The question enterprises are actually facing in mid-2026 isn’t “is Copilot worth it” — it’s “we run five of these, and I have five different answers to that question, none of which use the same metric, currency, or time window.” That’s a portfolio problem, and native vendor dashboards were never built to solve it. Copilot’s dashboard sees Copilot. Cursor’s admin panel sees Cursor. Neither will ever tell you which tool your best engineers are quietly switching away from, or which team is paying triple the per-seat cost of another team doing comparable work.

    Nobody is actually measuring this — even one tool at a time

    Before an organization can worry about comparing five AI coding tools, it has to be tracking metrics on any of them, and most aren’t. Jellyfish’s same report found that only 46% of organizations are actively tracking AI-specific metrics at all — adoption, acceptance rate, model usage, anything. That’s not 46% tracking consistently across every vendor in use. That’s 46% tracking anything, from any vendor, in any form. The other 54% are running a multi-vendor AI coding program on instinct: a sense that “the team seems to like Cursor” or “we haven’t heard complaints about Copilot,” with no underlying data to confirm or contradict it.

    Layer cost onto that visibility gap and the picture gets worse. Research from DX covering more than 400 organizations found blended per-developer spend across tiers and tools now running $200 to $600 a month, with agentic token consumption alone sometimes reaching $200 to $2,000-plus per engineer per month depending on usage intensity — the same range that blindsided Microsoft’s E+D division. And forecasting that spend is failing broadly, not just at the companies that make the news: a Mavvrik/Benchmarkit survey of 372 enterprises found only 15% forecast AI costs within 10% of actual, while nearly one in four miss by more than 50%. Multiply that forecasting failure across four or five tools running in parallel, each billed differently, each reported through a different console, and “surprise” stops being a risk and starts being the expected outcome.

    What a unified view actually requires

    Solving this isn’t a matter of asking engineering managers to check five dashboards instead of one and mentally reconcile the numbers. It requires a measurement layer that sits above every vendor and normalizes what each one reports into a single, comparable view. That’s the specific gap Olakai Agentic is built to close: a vendor-neutral analytics and governance layer across Claude Code, Cursor, GitHub Copilot, Codex, Gemini Code Assist, and whatever the team adopts next, without requiring the organization to standardize on one vendor first.

    Concretely, that means three things a single-vendor dashboard structurally cannot give you. First, cost-per-PR comparison across tools on common ground — not Cursor’s definition of a productive session next to Copilot’s definition of an accepted suggestion, but one measurement standard applied consistently, so a VP of Engineering can see that Team A’s tool costs three times what Team B’s does for comparable throughput and ask why. Second, adoption cohorts that span vendors, showing who’s actually using what — the power users worth studying, the licenses sitting idle regardless of which tool issued them, and the teams quietly switching tools without anyone approving the shift. Third, budget forecasting that aggregates spend across every provider into one number the CFO can trust, with alerts before a team’s token usage on any single tool turns into the kind of invoice that ends a pilot. This is the same measurement-layer thinking behind our Analytics & Custom KPIs capability, applied specifically to the vendor sprawl that now defines every engineering org’s AI coding stack.

    For a VP of Engineering, the practical shift is to stop evaluating AI coding tools one procurement cycle at a time and start treating the portfolio itself as the thing to manage. That means asking which teams are using which tools before the next contract renewal, not after a token bill forces the conversation; it means comparing cost-per-outcome across vendors on the same axis instead of trusting each tool’s self-reported acceptance rate; and it means building budget alerts before adoption surges the way Uber’s did, not after. Olakai Agentic is purpose-built for exactly that workflow, and it’s the specific reason we built a dedicated page for engineering leadership — see Olakai for VPs of Engineering for how the cross-tool view maps to the decisions this role actually has to make.

    None of this requires an organization to consolidate down to one AI coding tool, and for most engineering teams that wouldn’t even be the right call — different tools genuinely suit different workflows, and forcing a single vendor sacrifices real productivity gains for the sake of simpler reporting. The fix isn’t fewer tools. It’s a unified measurement layer across every coding tool the organization already runs, so the next $2,000-a-month surprise shows up on a dashboard weeks before it shows up on an invoice.

    If your engineering org is running three, four, or five AI coding tools right now with no shared view across them, that’s not a future governance project — it’s the state of your AI spend today, and it’s already accumulating risk you can’t see. Talk to an Expert to see how Olakai Agentic brings every AI coding vendor into one measurement and governance layer.

  • How to Be a Smarter Token Manager: Model Routing, Explained

    How to Be a Smarter Token Manager: Model Routing, Explained

    Two weeks of writing about AI token economics kept leading to the same corner: you cannot control what the vendors charge, only how wisely you spend it. That is the entire case for model routing, and this week it got a perfect teaching example.

    On Tuesday, Anthropic launched Claude Sonnet 5, and the whole pitch fit in one sentence: near-Opus performance, at a fraction of the price. Sit with that for a second, because that sentence is the entire case for routing, stated by a frontier lab about its own model lineup.

    How the pricing actually works

    You pay per token, split into input (what you send) and output (what the model writes back), and output is the expensive side, usually about five times the input rate. Here is the current Claude ladder, per million tokens.

    ModelInputOutputNotes
    Haiku 4.5$1$5Fastest, cheapest current tier
    Sonnet 5 (intro)$2$10Through Aug 31, 2026
    Sonnet 5 (standard)$3$15After Aug 31
    Opus 4.8$5$25Premium, the common default
    Fable 5$10$50Top tier, twice Opus

    Two more levers sit on top of that ladder: batch processing takes 50% off, and prompt caching takes up to 90% off the input you reuse. Keep both in your back pocket.

    Now, Sonnet 5 specifically. Anthropic’s own benchmark numbers put it close to Opus 4.8, roughly 63 versus 69 on agentic coding and basically tied on knowledge work, at about 40% of the price. That is the headline, and it is real. Here are the two things worth checking before anyone lets a vendor’s pricing banner do the talking. Sonnet 5’s new tokenizer turns the same input into up to 35% more tokens, so a slice of that discount comes right back. And the two-dollar rate is an introductory price that reverts to three and fifteen at the end of August. Real savings are real, you just calculate them on tokens, at the price you will actually pay in September, not the launch banner.

    The move: match the model to the task

    Here is the whole idea, and it is almost embarrassingly simple. Most AI calls never needed the top model in the first place. Across the deployments we see at Olakai, somewhere between 60 and 80% of the work, the summarizing, the extracting, the routine code, gets handled just as well by a model that costs five or ten times less. This is not a hunch. Researchers at UC Berkeley, Anyscale, and Canva published peer-reviewed routing work (RouteLLM, presented at ICLR 2025) showing roughly 85% cost savings while holding 95% of frontier-model quality, and in practice a well-chosen model pair lands around half the cost at about 98% of the quality. The reason it works is simple: most teams were overpaying on the easy stuff the whole time.

    Being a smarter token manager is just this: send each task to the cheapest model that can actually do it, and save the expensive model for the work that truly needs it.

    How it plays out with AI coding agents

    Coding is where the token bill actually lives for most engineering orgs, so it is worth making this concrete with three scenarios that show up constantly in AI coding tool deployments.

    The planner and the executors. A coding agent is not one thing. It is a planner that decides the approach, and a swarm of executors that do the grunt work: writing boilerplate, generating tests, fixing lint, editing files. The judgment lives in the planner, so give it the best model available. The executors are mostly routine, and they run over and over across a long session, which is exactly where tokens pile up. Point the executors at Haiku or Sonnet 5 instead of Opus, and the build gets dramatically cheaper with no drop in the quality that matters. One measured example: a 14-million-token build came in 57% cheaper with the executor on Haiku 4.5 instead of Opus 4.8, and the planner never changed.

    Route by difficulty. Not every ticket is hard. Renaming variables, scaffolding a test, a simple endpoint, a formatting pass, that is easy work, and it should go to Haiku. A feature or a mid-size refactor is Sonnet 5 territory. The gnarly stuff, tricky architecture, a subtle concurrency bug, security-sensitive code, is where it makes sense to spend on Opus. On most engineering teams the easy and medium buckets make up the vast majority of tickets, which means most traffic should never touch the frontier model at all.

    Watch the loops. Agentic coding burns tokens in a way chat never did, because agents retry, re-prompt, and loop, and every wasted token in turn one gets paid for again on every turn after it. A single long, unoptimized Opus session can run twenty dollars or more; the same session, routed and cleaned up, can be two or three. Multiply that across a twenty-developer team running dozens of sessions a day, and the difference is a five-figure monthly bill that is mostly avoidable. Route the routine sub-steps down, and cap the loops so one stuck agent cannot run up the tab.

    The same logic holds outside of code. Summarizing a long document costs about eleven cents on Haiku, fifty-five cents on Opus, and a dollar-ten on Fable, for the same summary. Run a million of those a month and that is a hundred and ten thousand dollars against well over a million. Burning the most capable model on a routine document summary is paying ten times over for an answer nobody can tell apart from the cheaper one.

    The catch, and it is the important one

    Routing is not free money, and it is not fire-and-forget. The classic way it bites: a team builds a router, cuts the bill 40%, finance is thrilled. Then the provider quietly tweaks the cheap model, a quality check starts failing, and the router silently sends everything back to the most expensive model. The next bill triples. Nothing errored. Nothing alerted. Savings only ever count net of quality, and a weaker answer that triggers retries and manual cleanup can quietly eat the very savings it created.

    Routing without measurement saves money right up until it costs more than it saved. Doing it properly takes three things running underneath the router: cost per outcome for each model, not just the total bill, so a route can be proven to actually pay off; a live watch for silent escalation and quality drift; and enforceable limits around the whole system, so a misrouted or runaway job cannot eat the quarter’s budget before anyone notices. This is precisely the layer we built Agent IQ to sit on top of, and it is why cost-per-outcome tracking, not just total spend, is the metric that matters. Visibility tells a team it happened. A limit stops it before it does.

    What this means for the CFO conversation

    For a finance leader watching AI coding spend climb, routing is the single highest-leverage lever available before the next budget review, but only if someone can show the receipt. “We switched to a cheaper model” is not a number. “We cut cost per accepted line by 40% while holding acceptance rate flat” is. That distinction is the difference between a CFO who trusts the next AI budget request and one who starts asking for a moratorium, a pattern we have written about in why acceptance rate is the wrong metric on its own for judging coding-tool ROI.

    The playbook

    This is what leading an AI transformation actually looks like in 2026. Not chasing the biggest model. Matching the model to the task, measuring that the swap held quality, and proving the savings with a number a CFO can defend in a board meeting. That is the vendor-neutral measurement layer Olakai exists to provide: one place to see cost per outcome across every model and every vendor, not a router’s word for it. That is how a team ships more, spends like it is its own money, and walks into the next budget review with the receipt instead of an excuse.

    One question worth taking into the next architecture review: what share of your AI calls hit your most expensive model by default, and do you actually know whether they needed to? If the honest answer is “we’re not sure,” that is the gap talking to an Olakai expert is built to close.

    Talk to an Expert →

  • Ask Kai: Inside Olakai’s Conversational Control Plane

    Ask Kai: Inside Olakai’s Conversational Control Plane

    Every vendor in enterprise AI analytics now claims some version of “ask questions in plain English.” Most of what that actually means, once you look closely, is a chat window bolted onto an existing dashboard — a nicer way to ask for a chart you could already find yourself. Kai, Olakai’s assistant, makes a different and more testable claim: it can also take action, with the exact change surfaced for approval before anything actually happens. That’s worth proving with the real catalog of things Kai can do, not just asserting.

    One data layer, five ways to hear the answer

    Kai sits on top of the same underlying data as Olakai Agentic (Coding IQ and Agent IQ) and Olakai Assistive, answering questions across both in a single conversation instead of forcing a switch between dashboards. Every answer comes with transparent reasoning — the logic chain behind the conclusion, not just the number — which is a stated design principle, not an incidental feature.

    What makes Kai’s answers actually usable across a company, rather than just for the person who built the dashboard, is Kai Lens: a perspective you choose once per conversation that reshapes how the same underlying data gets framed. Balanced is the default, adapting depth to the question. Executive leads with bottom-line ROI and strategic recommendations and skips implementation detail. Finance & Operations leads with cost figures and budget projections, in tables built for comparison. Legal & Compliance leads with compliance status and risk exposure in audit-ready language. Technical includes configuration details and API references. Ask the same question — “How are our AI agents performing?” — through each lens and you get four genuinely different answers: an Executive framing highlights overall ROI, top performers to scale, and risks to address for a leadership briefing; Finance & Operations shows cost-per-agent and month-over-month spend trends in tables; Legal & Compliance surfaces governance compliance rates and policy gaps; Technical lists agents by execution count, failure rates, and specific configuration issues. The lens is auto-suggested from the asker’s job title — a VP of Engineering sees Executive suggested by default, a Staff Engineer sees Technical — but it’s always overridable.

    What Kai can actually do, not just answer

    The differentiated part of Kai isn’t the natural-language question-answering — it’s the action catalog behind it. Kai can manage users directly: “Add john@company.com as an Analyst,” “Make Sarah an Admin,” “Deactivate John’s account.” It can manage Shadow AI governance: “Approve Grammarly as officially licensed,” “Mark ChatGPT as high risk,” and even bulk actions like “Block all AI tools rated high risk that have fewer than 10 interactions” — a request that would otherwise mean clicking through a table row by row. It can draft and manage acceptable-use policies, including generating one from a named compliance framework: “Create governance policies aligned with the EU AI Act for our HR department.” And it can handle enforcement follow-through — sending a reminder to users who violated a policy, or an escalation like “Sarah’s had 3 violations this month — send an escalation to her manager.”

    Kai’s reach extends into AI spend governance too: it can create, update, and archive Coding IQ cost-center projects and assign a service API key into one, which turns a long backlog of unassigned keys from a tedious manual triage session into a short conversation.

    Confirmation-first, not silent

    None of this works, from a trust standpoint, if a chat interface can quietly reassign a user’s role or block an AI tool the moment someone phrases a request slightly wrong. Kai’s write actions are ADMIN-gated and confirmation-first: every change Kai proposes gets surfaced explicitly for approval before it’s applied, not executed silently the moment the request is understood. That design choice is the actual answer to the obvious objection — “you’re letting a chatbot make changes to my governance policy?” — and it’s the reason the honest framing for Kai isn’t “an AI that runs your platform,” it’s “an AI that tells you exactly what it’s about to do, and waits.”

    Why this matters beyond convenience

    The pitch to a CIO or Chief AI Officer isn’t “ask questions in English” — every vendor says that now, and it doesn’t differentiate anything. It’s “ask a question in English and get an answer that comes with an offer to fix what it found, in the same conversation, gated by a permission check and a confirmation step.” That’s a materially different product than a read-only chat wrapper, and it’s the reason Kai belongs to both Olakai Agentic and Olakai Assistive rather than being siloed to one product — a governance question rarely respects the boundary between coding tools and chatbots, and neither should the assistant answering it.

    Kai is available immediately to any account with Coding IQ, Agent IQ, or Assistive IQ data flowing in — no separate setup required. It’s the closest thing on the platform to a business-friendly interface in the literal sense: a Legal & Compliance leader and a Staff Engineer can ask the exact same underlying data the exact same question and both walk away with an answer built for them.

    Want to see what Kai can tell you — and do for you — with your own AI usage data? Talk to an Expert.

  • 3 Token Cost Metrics Every CFO Should Be Watching

    3 Token Cost Metrics Every CFO Should Be Watching

    In May 2026, Uber’s COO Andrew Macdonald said something that should make every CFO uncomfortable. Uber had burned through its entire 2026 AI budget in four months — deploying Anthropic’s Claude Code to roughly 5,000 engineers, watching per-engineer token costs hit $500 to $2,000 per month, and reaching April before anyone noticed the year was over. When pressed on the return, Macdonald said: “That link is not there yet.” Meaning Uber — a $140B technology company with sophisticated financial infrastructure — cannot draw a line between its AI spend and any consumer feature shipped to customers.

    This isn’t a story about Uber being careless. It’s a story about a structural gap that no CFO team was built for. SaaS budgets were predictable: seat count × price, invoiced monthly, trivial to reconcile. Token-based AI consumption is none of those things. It scales with usage, multiplies with agentic workflows, and generates costs that engineering teams incur invisibly throughout the month. By the time finance sees the number, the spending is already done. Uber found out in April. Microsoft found out around the same time and revoked Claude Code licenses for an entire division effective June 30. These aren’t outliers. According to Ramp’s April 2026 AI Index, monthly AI token spend across enterprise customers grew 1,001% from January 2025 to April 2026. The median company now dedicates nearly 15% of its software budget to AI tools.

    The finance operating model hasn’t caught up. Most AI monitoring tools give CFOs a token dashboard — a view of how many tokens were consumed, by which provider, at what cost. That’s a start. But it’s not a CFO metric. It’s an engineering metric dressed up for the finance team. What CFOs actually need are three different measurements, each one capturing something a token dashboard deliberately ignores.

    Why This Is Different From Every SaaS Budget You’ve Managed Before

    The shift from seat-based to token-based pricing is more disruptive to financial planning than it looks. Seat costs are a fixed overhead — you know the number on the first of the month. Token costs are a variable that compounds with behavior. The more your engineers use AI, the more capable and dependent they become, and the more tokens they consume. EY estimates that a standard chatbot interaction costs roughly $0.04. An orchestrated agentic workflow — where AI models call tools, spawn sub-agents, and iterate across multiple reasoning steps — costs approximately $1.20 per interaction. That’s a 30x multiplier, and it’s built into the architecture of where AI is going. Goldman Sachs projects that agentic AI adoption will drive a 24x increase in global token demand by 2030.

    Meanwhile, per-developer token consumption is growing at a pace that defies normal budget forecasting. TechCrunch reported in June 2026 that per-developer token consumption has grown approximately 18.6x in nine months across enterprise organizations. A Priceline engineer burned $40,000 in tokens in a single month. An unnamed enterprise accumulated a $500M Claude bill. The Linux Foundation has responded by standing up a formal Tokenomics Foundation to create standards for AI token tracking — which is itself a signal that the industry now acknowledges cost runaway as a structural problem, not an edge case. If you don’t have the right instruments in place, you’re flying without gauges in an environment where the turbulence is increasing. Here are the three metrics that change that.

    Metric 1: Cost-Per-Outcome, Not Cost-Per-Token

    Andrew Macdonald’s admission — “that link is not there yet” — describes exactly what’s missing from every token dashboard on the market. They tell you what you spent. They don’t tell you what you got. And the gap between those two questions is where CFOs get into trouble. A team burning twice the tokens of the team next to them isn’t necessarily wasteful. They might be twice as productive. Or they might be prompting in circles. You cannot tell from a spend number alone, which is why cost-per-token is the wrong unit of analysis for a CFO.

    The metric that matters is cost-per-outcome: the fully-loaded dollar cost of each unit of value produced. For engineering teams, that’s cost per merged pull request, cost per deployed feature, cost per lines of production code shipped. When you measure at this level, the teams consuming the most tokens often look very different than you’d expect. Jellyfish’s research found that heavy AI users were twice as productive as their peers but consumed ten times more tokens. At the token level, they look expensive. At the outcome level, they’re your most cost-efficient engineers. Only 14% of CFOs report they’ve seen clear, measurable AI ROI (RGP, 200 US finance chiefs) — the primary reason is that they’re measuring inputs, not outputs. Cost-per-outcome is what CFOs actually need from AI measurement to make budget decisions that hold up to board scrutiny.

    Metric 2: Spend Run-Rate Forecast, Not Month-to-Date Total

    Month-to-date spend is a rearview mirror. By the time April’s actuals landed in Uber’s financial system, the year was already gone. What every CFO needs — and almost none have — is a forward-looking signal: at the current trajectory, when do we exhaust this budget? This is the difference between a smoke alarm and a fire report. MTD is the fire report. Run-rate forecast is the smoke alarm.

    The reason this matters so urgently right now is the 18.6x nine-month consumption growth rate. Token spend doesn’t grow linearly. It grows exponentially as more engineers adopt AI tools, as those engineers use them for more complex tasks, and as agentic workflows multiply the token cost of each interaction. A budget that looked fine in January can be 40% consumed by February if adoption accelerates faster than the plan assumed. The answer is a rolling run-rate alert — a projection based on trailing consumption that fires when the month-end trajectory crosses a threshold, not when the limit is already breached. In the Uber scenario, a 7-day trailing average run-rate alert in late January or early February would have changed the conversation months before the budget was gone. Budget alerts that fire after the fact aren’t governance — they’re retrospectives. The signal you need fires while there’s still time to adjust. This is the complete AI monitoring posture that separates reactive from proactive finance teams.

    Metric 3: Value Leak Rate

    The Priceline engineer who spent $40,000 in tokens in one month is an interesting problem. Maybe those tokens produced something extraordinary — a complex system design, a breakthrough on a hard architecture problem, intensive research that unblocked the whole team. Or maybe that engineer was prompting in circles, getting low-quality outputs, and abandoning sessions without shipping anything. From a token dashboard, both scenarios look identical. Both show high spend. Neither reveals whether the spend connected to anything the business actually values.

    Value leak rate measures the share of AI spend that doesn’t connect to a committed output: a merged PR, a deployed commit, a shipped feature. High-spend sessions that end without a commit are the signal. Not because exploration is bad — sometimes the right answer from a session is “don’t build this” — but because a high value leak rate at the account level tells you that a meaningful fraction of your AI spend is disappearing without evidence of production. The nuance matters here. Flagging every high-spend session as waste would punish your most ambitious engineers. The right instrument identifies the pattern: sessions with consistently high spend and no output, compared against a team-median baseline, tracked over time. That’s the difference between an AI visibility audit and a surveillance tool. One helps CFOs understand where the budget is going. The other just creates resentment. Jellyfish’s data — 2x productivity, 10x token cost for heavy users — makes the case for why you need this ratio, not the raw number. The ratio tells you whether the premium is justified. And if you want custom AI cost KPIs that reflect your team’s specific cost structure, the baseline needs to come from your own data, not industry benchmarks.

    What Proactive Finance Teams Are Doing Now

    The companies that have gotten ahead of this aren’t waiting for the annual budget reconciliation to discover they have a token runaway problem. AT&T achieved 90% cost savings in AI infrastructure after building visibility into where tokens were actually going — not by cutting investment, but by identifying the optimization opportunities that were invisible before. Kumo AI now treats per-engineer token consumption as a tracked R&D expense line, the same way they track compute or software licensing. This framing shifts the conversation from “are we spending too much?” to “are we getting R&D-quality returns on this R&D-level expense?” — which is the right question for a CFO to be asking. Gartner projects that by 2029, CFOs who implement strategic AI deployment will add 10 margin points of growth, and over 40% of agentic AI projects will be canceled before that due to escalating costs and unclear business value. The companies that add those margin points will be the ones that built the measurement infrastructure before the costs compounded. The others will be telling the Uber story about themselves in 2027.

    The AI P&L is becoming a real thing inside enterprise finance. Token spend, cost-per-outcome, run-rate forecasting, and value leak rate are the line items. The CFOs who define those metrics now, build the instrumentation to track them, and establish the governance to act on them will be in a fundamentally different position than those who wait for the token dashboards to catch up. The gap between tracking spend and understanding value is the gap between a cost center and a competitive advantage. If you’re not tracking these three numbers across your entire AI stack today, talk to an expert about what it takes to get there.

  • Power, Casual, New, Idle: How Olakai’s Adoption Cohorts Find Your Wasted AI Licenses

    Power, Casual, New, Idle: How Olakai’s Adoption Cohorts Find Your Wasted AI Licenses

    A company buys 200 seats of Claude Code or Cursor — a familiar story in the era of AI coding tool sprawl. Six months later, the adoption dashboard reports “82% activated” and everyone treats that as a win. It might be. It might also mean 40% of those developers opened the tool once, wrote one throwaway prompt, and never came back — activated isn’t the same as valuable, and a single company-wide percentage can’t tell the difference. Olakai’s adoption cohorts exist specifically to make that distinction, on the Developers tab of the AI Impact Dashboard, at the level of an individual developer rather than a rounded company average.

    Four cohorts, one ratio

    Every developer is assigned to one of four cohorts based on the share of their merged pull requests that show AI assistance. Power means more than 70% of their PRs are AI-assisted. Casual sits between 20% and 70%. Idle is under 20%. New is a fourth category that cuts across the ratio entirely: a developer whose first AI-assisted PR landed within the last 14 days is New, regardless of what their ratio looks like.

    That last rule is a genuinely thoughtful piece of the design, and it’s worth explaining why it exists rather than just stating it. New is evaluated first and wins over the ratio bands — a developer who shipped their first AI-assisted PR this week counts as New even if literally every PR they’ve merged so far is AI-assisted. Without that override, a brand-new user’s tiny, noisy sample would misleadingly register as “Power” the moment they merged two or three AI-assisted PRs in a row, which is a worse signal than an honest “too early to tell.” The New cohort exists so the ratio-based cohorts describe settled behavior, not a small sample still finding its footing.

    Cohort assignment isn’t a permanent label, either. It recomputes every time the dashboard loads, based on whatever time range is currently selected — a developer who’s New today can become Power or Casual within a few weeks as more PRs accumulate, and the cohort you see reflects the current window rather than a badge assigned once and left stale.

    The table under the label

    The cohort itself is a starting point, not the whole answer — the per-developer table underneath is where a specific, defensible decision actually gets made. Each row shows total PRs in the selected window, AI ratio, lines moved, a 0-100 prompting-clarity score, real-time estimated cost against actual billed cost from the Admin API, lines moved per dollar, and hook coverage — how much of a developer’s agent telemetry is actually showing up against their PR activity. That last column matters more than it sounds like it should: a developer who looks Idle by PR ratio but has strong hook coverage and heavy agent usage that just hasn’t produced a merged PR yet is a different conversation than a developer who’s genuinely not using the tool at all.

    This is the layer that turns “adoption is low” from a vague, company-wide complaint into a specific, actionable list. An Idle developer with a paid seat isn’t an abstraction — it’s a named line item a VP of Engineering can pull up, with real usage data attached, and act on directly: re-train, reassign the license to someone on the waitlist, or have an honest conversation about whether the tool fits how that person actually works. None of that is possible from a single “82% adoption” slide.

    The license-waste math, made concrete

    Run the numbers on that 200-seat example: at even a modest per-seat price, a company with 40 genuinely Idle developers — not just under-adopting, but under 20% AI-assisted with no meaningful hook coverage either — is paying full price for licenses nobody is using. That’s not a governance abstraction or a productivity-culture problem to solve with a training session no one attends. It’s a specific, controllable line item that shows up the moment someone actually looks at the cohort table instead of the headline adoption percentage, and it’s exactly the kind of finding that turns into a real budget conversation rather than a vague resolution to “drive more adoption” next quarter.

    Adoption cohorts sit next to a related, deeper diagnostic worth knowing about even if it’s a separate feature: each developer’s detail page includes a Fluency tab reporting on their specific AI fluency dimensions and where they have room to grow, which goes further than the cohort label alone when the goal is coaching rather than license reallocation.

    The underlying value here is the same thread running through the rest of Olakai’s AI coding analytics: a business-friendly interface that turns raw Git activity into a decision a non-engineering manager can actually act on, rather than a chart that requires an engineer to interpret before anyone can use it.

    Curious how many of your own AI coding seats are Power, Casual, New, or genuinely Idle? Talk to an Expert.

  • Inside AI Spend Governance: Budgets and the Alerts That Fire Before You Blow Through Them

    Inside AI Spend Governance: Budgets and the Alerts That Fire Before You Blow Through Them

    A CFO rarely discovers an AI spend problem from a dashboard. They discover it from an invoice, weeks after the spending already happened, with no window left to do anything but ask engineering what happened. Olakai’s Budgets feature inside AI Spend Governance exists specifically to close that gap — not by promising a smarter forecast, but by being honest about what a budget actually is and firing an alert while there’s still time to act on it.

    Budgets are lenses, not a partition

    The single most important thing to understand about Olakai’s budgets is also the thing most competing tools obscure: budgets do not partition spend. They are independent, overlapping lenses over the same dollars. The same charge can count toward a developer’s budget, their department’s budget, the provider budget, and the program budget, all at once. If you add up every individual budget on the page expecting the total to match your program spend, it won’t — and that’s by design, not a bug to file a ticket about.

    Budgets are organized into four groups. Program is the master ceiling — every developer, every provider, every project rolled into one account-wide cap. Provider lets you cap a single vendor, like Anthropic or Cursor, without touching anyone else’s spend. Employee-centric budgets attribute spend to people — an individual developer, the persona they belong to, or their department — with the same overlap rule: a shared engineer’s spend counts fully against every relevant lens. Project groups shared service keys into named cost centers, with developers who belong to multiple projects contributing their full spend to each one rather than having it split proportionally.

    What budgets don’t track, and why that’s deliberate

    Budgets only track per-token, cost-bearing providers — Anthropic, Cursor, and OpenAI. GitHub Copilot is excluded outright, because its pricing is seat-based rather than usage-metered, and a per-token budget mechanism has nothing to measure against a flat subscription fee. Spend that can’t be attributed to a specific person still counts toward the program and provider totals, but drops out of the developer, persona, and department lenses entirely — which means per-entity totals can legitimately be lower than the program total, and that’s worth knowing before a CFO tries to reconcile the two and assumes something’s broken. The same per-provider precision shows up in how Olakai handles Google Vertex AI cost data, which lags behind usage by design rather than pretending to be real-time when it isn’t.

    Budgets extend naturally into Projects, Olakai’s term for a cost center: a named bucket that groups shared service API keys, member developers, and owned repositories under one monthly limit. Total project cost is service-key spend plus member-developer token spend, and — consistent with the overlapping-lens rule everywhere else — a developer who belongs to more than one project contributes their full token spend to each project they’re in, not a proportional split. Archiving a project unassigns its keys and removes its budget and alert rules, but never deletes the underlying spend history; only the grouping goes away.

    The forecast that’s already live, and the one that isn’t yet

    Every budget carries a Projected month-end figure built the same way: the recent daily spend rate, extended across the remaining days of the month, added to spend so far, with a confidence signal that reflects how steady daily spend has actually been. It’s a run-rate projection, not a model of growth or seasonality — a distinction Olakai is upfront about, the same way it’s upfront about the 30-day spend projection on the main AI Impact Dashboard being a straight-line extrapolation rather than a real forecast. Budgets are evaluated once daily, right after the cost-import job pulls fresh provider spend, and saving a budget automatically provisions the alert rules behind it — nothing extra to configure.

    Two kinds of alerts come out of that evaluation. The first is a threshold alert: actual month-to-date spend crosses a configurable percentage of the budget — 50%, 80%, or 100%. The second, and the more useful one, is a forecast alert: the run-rate projection is on track to exceed the limit by month end, even if the account isn’t over budget yet today. That second alert is the actual point of the feature — catching a trajectory early enough to still change it, rather than confirming after the fact that the month already went over.

    Worth being precise about scope here: Olakai also has a separate, standalone Forecasts tab planned for what-if scenario modeling across budget dimensions — a different, more ambitious feature for testing hypothetical spend trajectories before committing to them. As of this writing, that tab isn’t live yet. What’s covered above — the run-rate projection and the two alert types built directly into the Budgets page — is shipped and running today; the scenario-modeling tool is a separate thing worth revisiting once it ships.

    Why the overlap is the right design, not a shortcut

    It would be simpler to build budgets that partition spend cleanly — every dollar assigned to exactly one bucket, everything adding up neatly on a summary slide. It would also be wrong for how AI spend actually happens inside a real engineering org, where the same developer’s usage genuinely belongs to their department’s headcount planning, their manager’s persona-level benchmarking, the vendor contract renewal conversation, and the specific project that consumed it — four legitimate, simultaneous questions about the same dollar. Building four separate, non-overlapping ledgers to answer four different questions would mean picking one authoritative answer and getting the other three wrong. Overlapping lenses let all four questions get an honest answer from the same underlying spend data, at the cost of a program total that doesn’t equal the sum of its parts — which is exactly the tradeoff a CFO should want once it’s explained, rather than discovered while trying to make the numbers reconcile.

    It’s also worth knowing that budgets and projects aren’t limited to point-and-click configuration — Kai can create, edit, and archive them conversationally, gated by admin permission and a confirmation step before anything actually changes, which is a useful shortcut when there’s a long backlog of unassigned service keys to triage.

    This isn’t a hypothetical risk — Uber blew through an entire year of AI budget in four months before building a reactive cap of its own. If your AI coding spend has outgrown a spreadsheet and a monthly Slack message from finance, Talk to an Expert about setting up budgets against your own provider and project data.

  • AI Coding Tool ROI: Why Acceptance Rate Is the Wrong Metric

    AI Coding Tool ROI: Why Acceptance Rate Is the Wrong Metric

    In May, Gartner published its first-ever assessment of the enterprise AI coding agent market, formalizing a category that did not exist as a named market segment two years ago and now runs to roughly ten billion dollars a year. The message between the lines was clear: of all the places enterprises have poured AI money, software development is where the returns look most real. So here is the question every engineering leader should sit with. If coding is the one domain where AI value is most provable, why can almost nobody prove it?

    Most organizations buy seats, watch a vendor dashboard tick upward, and conclude things are working. The dashboard shows suggestions made, suggestions accepted, an acceptance rate climbing past 30%. It feels like proof. It is not. Acceptance rate is the single most misleading number in the entire AI coding conversation, and the gap between what it measures and what actually matters is where engineering budgets quietly lose their justification.

    Coding really is different

    The optimistic case for AI coding tools is genuine, and it deserves a fair hearing before the skepticism arrives. GitHub’s own controlled study found developers completing a programming task 55% faster with an assistant than without. The market reflects that promise: AI coding tools now represent well over ten billion dollars in annual spend, and roughly 90% of the Fortune 100 have deployed GitHub Copilot in some form. Gartner’s decision to stand up a formal market assessment is itself a signal that coding has matured past experimentation into something boards expect to pay off.

    That maturity is exactly why coding deserves better measurement than the rest of the AI portfolio, not worse. Gartner’s parallel research on AI in infrastructure and operations found that only 28% of those use cases fully succeed. Coding stands out as the exception, the place where the productivity story has the most evidence behind it. When you have found the one room with treasure in it, you do not measure your haul by counting how many times you opened a drawer.

    But the dashboard is lying to you

    The cleanest evidence that activity metrics mislead comes from a randomized controlled trial. METR studied experienced open-source developers working in codebases they knew well, and found they were 19% slower when using AI assistance. The detail that matters most for measurement: those same developers estimated they had been 20% faster. A nearly forty-point gap between perceived and actual productivity, in the population most enterprises are deploying these tools to. If your ROI case rests on developer self-report or on a feeling that the team is moving quicker, that is the gap you are standing on.

    The quality picture is just as sobering. GitClear’s analysis of 211 million changed lines of code found that copy-pasted and duplicated code blocks rose eightfold in a single year, code churn climbed, and the share of lines devoted to refactoring fell to under 10%. AI makes it trivial to add code and does nothing to encourage consolidating it. Google’s 2025 DORA research found the same tension from a different angle: AI adoption correlated positively with throughput but negatively with delivery stability, meaning the tools that help you ship faster can quietly erode the controls that keep what you ship from breaking. Acceptance rate captures none of this. A developer can accept every suggestion and ship slower, buggier software, and the dashboard will call that a win.

    Activity versus value: the real metric problem

    The reason vendor dashboards surface acceptance rate, lines generated, and seat utilization is that these are the metrics the vendor controls and optimizes for. They describe how much the tool was used, not what the use produced. That distinction is the whole game, and it is the same vanity-versus-value problem we mapped for finance leaders in the metrics that actually matter. An engineering org running on acceptance rate is measuring the proxy and ignoring the signal.

    The signal lives in a different set of numbers. How does cycle time differ between AI-assisted pull requests and the rest? What is your cost per merged PR once you divide total tool spend across providers by the work actually shipped? How has defect density moved since rollout, and which teams are driving the change? Which developers have genuinely adopted the tools, and which licenses are sitting idle at $19 to $50 a head every month? Answering those questions requires connecting pull-request data, provider cost data, and engineering outcomes in one place, which is precisely what Coding IQ was built to do. It measures the value of AI coding tools rather than the activity, because activity was never the thing the CFO was paying for.

    What good measurement actually enables

    This is not an argument that AI coding tools do not work. It is an argument that you cannot manage what you measure badly. The enterprises pulling real value from these tools are the ones that instrumented outcomes before scaling seats, the same discipline that separates winners across every category of AI investment in the broader ROI playbook. They can make decisions the acceptance-rate crowd cannot.

    Consider the difference at a budget review. An engineering leader who can say that Cursor users close pull requests 28% faster than non-users at a cost of a few dollars per PR, while 40% of Copilot licenses sit unused, is making a business decision: scale the first, reclaim the second. A leader who can only report a 32% acceptance rate is reporting a vendor metric and hoping nobody asks what it bought. That is the position most VPs of engineering find themselves in, and it is an avoidable one. The instrumentation that closes the gap is the same vendor-neutral measurement layer that proves AI ROI across the rest of the stack, applied to the one domain where the returns are most worth proving. It is also the only honest way out of the trap NVIDIA documented when it found 30% of enterprises still cannot quantify AI ROI at all.

    Coding is where enterprise AI ROI is most real. That makes it the worst possible place to keep measuring the wrong thing. Acceptance rate will tell you your developers are clicking accept. It will never tell you whether your software is better, faster, or cheaper to ship, which is the only question your board is actually asking.

    Is your coding-tool spend producing value, or just activity? Talk to an expert to see how Olakai’s Coding IQ ties AI coding tools to cycle time, defect rate, and cost per pull request, so you can scale what works and cut what doesn’t.