Category: AI Strategy

Strategic guidance for enterprise AI adoption and measurement

  • Ask Kai: Inside Olakai’s Conversational Control Plane

    Ask Kai: Inside Olakai’s Conversational Control Plane

    Every vendor in enterprise AI analytics now claims some version of “ask questions in plain English.” Most of what that actually means, once you look closely, is a chat window bolted onto an existing dashboard — a nicer way to ask for a chart you could already find yourself. Kai, Olakai’s assistant, makes a different and more testable claim: it can also take action, with the exact change surfaced for approval before anything actually happens. That’s worth proving with the real catalog of things Kai can do, not just asserting.

    One data layer, five ways to hear the answer

    Kai sits on top of the same underlying data as Olakai Agentic (Coding IQ and Agent IQ) and Olakai Assistive, answering questions across both in a single conversation instead of forcing a switch between dashboards. Every answer comes with transparent reasoning — the logic chain behind the conclusion, not just the number — which is a stated design principle, not an incidental feature.

    What makes Kai’s answers actually usable across a company, rather than just for the person who built the dashboard, is Kai Lens: a perspective you choose once per conversation that reshapes how the same underlying data gets framed. Balanced is the default, adapting depth to the question. Executive leads with bottom-line ROI and strategic recommendations and skips implementation detail. Finance & Operations leads with cost figures and budget projections, in tables built for comparison. Legal & Compliance leads with compliance status and risk exposure in audit-ready language. Technical includes configuration details and API references. Ask the same question — “How are our AI agents performing?” — through each lens and you get four genuinely different answers: an Executive framing highlights overall ROI, top performers to scale, and risks to address for a leadership briefing; Finance & Operations shows cost-per-agent and month-over-month spend trends in tables; Legal & Compliance surfaces governance compliance rates and policy gaps; Technical lists agents by execution count, failure rates, and specific configuration issues. The lens is auto-suggested from the asker’s job title — a VP of Engineering sees Executive suggested by default, a Staff Engineer sees Technical — but it’s always overridable.

    What Kai can actually do, not just answer

    The differentiated part of Kai isn’t the natural-language question-answering — it’s the action catalog behind it. Kai can manage users directly: “Add john@company.com as an Analyst,” “Make Sarah an Admin,” “Deactivate John’s account.” It can manage Shadow AI governance: “Approve Grammarly as officially licensed,” “Mark ChatGPT as high risk,” and even bulk actions like “Block all AI tools rated high risk that have fewer than 10 interactions” — a request that would otherwise mean clicking through a table row by row. It can draft and manage acceptable-use policies, including generating one from a named compliance framework: “Create governance policies aligned with the EU AI Act for our HR department.” And it can handle enforcement follow-through — sending a reminder to users who violated a policy, or an escalation like “Sarah’s had 3 violations this month — send an escalation to her manager.”

    Kai’s reach extends into AI spend governance too: it can create, update, and archive Coding IQ cost-center projects and assign a service API key into one, which turns a long backlog of unassigned keys from a tedious manual triage session into a short conversation.

    Confirmation-first, not silent

    None of this works, from a trust standpoint, if a chat interface can quietly reassign a user’s role or block an AI tool the moment someone phrases a request slightly wrong. Kai’s write actions are ADMIN-gated and confirmation-first: every change Kai proposes gets surfaced explicitly for approval before it’s applied, not executed silently the moment the request is understood. That design choice is the actual answer to the obvious objection — “you’re letting a chatbot make changes to my governance policy?” — and it’s the reason the honest framing for Kai isn’t “an AI that runs your platform,” it’s “an AI that tells you exactly what it’s about to do, and waits.”

    Why this matters beyond convenience

    The pitch to a CIO or Chief AI Officer isn’t “ask questions in English” — every vendor says that now, and it doesn’t differentiate anything. It’s “ask a question in English and get an answer that comes with an offer to fix what it found, in the same conversation, gated by a permission check and a confirmation step.” That’s a materially different product than a read-only chat wrapper, and it’s the reason Kai belongs to both Olakai Agentic and Olakai Assistive rather than being siloed to one product — a governance question rarely respects the boundary between coding tools and chatbots, and neither should the assistant answering it.

    Kai is available immediately to any account with Coding IQ, Agent IQ, or Assistive IQ data flowing in — no separate setup required. It’s the closest thing on the platform to a business-friendly interface in the literal sense: a Legal & Compliance leader and a Staff Engineer can ask the exact same underlying data the exact same question and both walk away with an answer built for them.

    Want to see what Kai can tell you — and do for you — with your own AI usage data? Talk to an Expert.

  • 3 Token Cost Metrics Every CFO Should Be Watching

    3 Token Cost Metrics Every CFO Should Be Watching

    In May 2026, Uber’s COO Andrew Macdonald said something that should make every CFO uncomfortable. Uber had burned through its entire 2026 AI budget in four months — deploying Anthropic’s Claude Code to roughly 5,000 engineers, watching per-engineer token costs hit $500 to $2,000 per month, and reaching April before anyone noticed the year was over. When pressed on the return, Macdonald said: “That link is not there yet.” Meaning Uber — a $140B technology company with sophisticated financial infrastructure — cannot draw a line between its AI spend and any consumer feature shipped to customers.

    This isn’t a story about Uber being careless. It’s a story about a structural gap that no CFO team was built for. SaaS budgets were predictable: seat count × price, invoiced monthly, trivial to reconcile. Token-based AI consumption is none of those things. It scales with usage, multiplies with agentic workflows, and generates costs that engineering teams incur invisibly throughout the month. By the time finance sees the number, the spending is already done. Uber found out in April. Microsoft found out around the same time and revoked Claude Code licenses for an entire division effective June 30. These aren’t outliers. According to Ramp’s April 2026 AI Index, monthly AI token spend across enterprise customers grew 1,001% from January 2025 to April 2026. The median company now dedicates nearly 15% of its software budget to AI tools.

    The finance operating model hasn’t caught up. Most AI monitoring tools give CFOs a token dashboard — a view of how many tokens were consumed, by which provider, at what cost. That’s a start. But it’s not a CFO metric. It’s an engineering metric dressed up for the finance team. What CFOs actually need are three different measurements, each one capturing something a token dashboard deliberately ignores.

    Why This Is Different From Every SaaS Budget You’ve Managed Before

    The shift from seat-based to token-based pricing is more disruptive to financial planning than it looks. Seat costs are a fixed overhead — you know the number on the first of the month. Token costs are a variable that compounds with behavior. The more your engineers use AI, the more capable and dependent they become, and the more tokens they consume. EY estimates that a standard chatbot interaction costs roughly $0.04. An orchestrated agentic workflow — where AI models call tools, spawn sub-agents, and iterate across multiple reasoning steps — costs approximately $1.20 per interaction. That’s a 30x multiplier, and it’s built into the architecture of where AI is going. Goldman Sachs projects that agentic AI adoption will drive a 24x increase in global token demand by 2030.

    Meanwhile, per-developer token consumption is growing at a pace that defies normal budget forecasting. TechCrunch reported in June 2026 that per-developer token consumption has grown approximately 18.6x in nine months across enterprise organizations. A Priceline engineer burned $40,000 in tokens in a single month. An unnamed enterprise accumulated a $500M Claude bill. The Linux Foundation has responded by standing up a formal Tokenomics Foundation to create standards for AI token tracking — which is itself a signal that the industry now acknowledges cost runaway as a structural problem, not an edge case. If you don’t have the right instruments in place, you’re flying without gauges in an environment where the turbulence is increasing. Here are the three metrics that change that.

    Metric 1: Cost-Per-Outcome, Not Cost-Per-Token

    Andrew Macdonald’s admission — “that link is not there yet” — describes exactly what’s missing from every token dashboard on the market. They tell you what you spent. They don’t tell you what you got. And the gap between those two questions is where CFOs get into trouble. A team burning twice the tokens of the team next to them isn’t necessarily wasteful. They might be twice as productive. Or they might be prompting in circles. You cannot tell from a spend number alone, which is why cost-per-token is the wrong unit of analysis for a CFO.

    The metric that matters is cost-per-outcome: the fully-loaded dollar cost of each unit of value produced. For engineering teams, that’s cost per merged pull request, cost per deployed feature, cost per lines of production code shipped. When you measure at this level, the teams consuming the most tokens often look very different than you’d expect. Jellyfish’s research found that heavy AI users were twice as productive as their peers but consumed ten times more tokens. At the token level, they look expensive. At the outcome level, they’re your most cost-efficient engineers. Only 14% of CFOs report they’ve seen clear, measurable AI ROI (RGP, 200 US finance chiefs) — the primary reason is that they’re measuring inputs, not outputs. Cost-per-outcome is what CFOs actually need from AI measurement to make budget decisions that hold up to board scrutiny.

    Metric 2: Spend Run-Rate Forecast, Not Month-to-Date Total

    Month-to-date spend is a rearview mirror. By the time April’s actuals landed in Uber’s financial system, the year was already gone. What every CFO needs — and almost none have — is a forward-looking signal: at the current trajectory, when do we exhaust this budget? This is the difference between a smoke alarm and a fire report. MTD is the fire report. Run-rate forecast is the smoke alarm.

    The reason this matters so urgently right now is the 18.6x nine-month consumption growth rate. Token spend doesn’t grow linearly. It grows exponentially as more engineers adopt AI tools, as those engineers use them for more complex tasks, and as agentic workflows multiply the token cost of each interaction. A budget that looked fine in January can be 40% consumed by February if adoption accelerates faster than the plan assumed. The answer is a rolling run-rate alert — a projection based on trailing consumption that fires when the month-end trajectory crosses a threshold, not when the limit is already breached. In the Uber scenario, a 7-day trailing average run-rate alert in late January or early February would have changed the conversation months before the budget was gone. Budget alerts that fire after the fact aren’t governance — they’re retrospectives. The signal you need fires while there’s still time to adjust. This is the complete AI monitoring posture that separates reactive from proactive finance teams.

    Metric 3: Value Leak Rate

    The Priceline engineer who spent $40,000 in tokens in one month is an interesting problem. Maybe those tokens produced something extraordinary — a complex system design, a breakthrough on a hard architecture problem, intensive research that unblocked the whole team. Or maybe that engineer was prompting in circles, getting low-quality outputs, and abandoning sessions without shipping anything. From a token dashboard, both scenarios look identical. Both show high spend. Neither reveals whether the spend connected to anything the business actually values.

    Value leak rate measures the share of AI spend that doesn’t connect to a committed output: a merged PR, a deployed commit, a shipped feature. High-spend sessions that end without a commit are the signal. Not because exploration is bad — sometimes the right answer from a session is “don’t build this” — but because a high value leak rate at the account level tells you that a meaningful fraction of your AI spend is disappearing without evidence of production. The nuance matters here. Flagging every high-spend session as waste would punish your most ambitious engineers. The right instrument identifies the pattern: sessions with consistently high spend and no output, compared against a team-median baseline, tracked over time. That’s the difference between an AI visibility audit and a surveillance tool. One helps CFOs understand where the budget is going. The other just creates resentment. Jellyfish’s data — 2x productivity, 10x token cost for heavy users — makes the case for why you need this ratio, not the raw number. The ratio tells you whether the premium is justified. And if you want custom AI cost KPIs that reflect your team’s specific cost structure, the baseline needs to come from your own data, not industry benchmarks.

    What Proactive Finance Teams Are Doing Now

    The companies that have gotten ahead of this aren’t waiting for the annual budget reconciliation to discover they have a token runaway problem. AT&T achieved 90% cost savings in AI infrastructure after building visibility into where tokens were actually going — not by cutting investment, but by identifying the optimization opportunities that were invisible before. Kumo AI now treats per-engineer token consumption as a tracked R&D expense line, the same way they track compute or software licensing. This framing shifts the conversation from “are we spending too much?” to “are we getting R&D-quality returns on this R&D-level expense?” — which is the right question for a CFO to be asking. Gartner projects that by 2029, CFOs who implement strategic AI deployment will add 10 margin points of growth, and over 40% of agentic AI projects will be canceled before that due to escalating costs and unclear business value. The companies that add those margin points will be the ones that built the measurement infrastructure before the costs compounded. The others will be telling the Uber story about themselves in 2027.

    The AI P&L is becoming a real thing inside enterprise finance. Token spend, cost-per-outcome, run-rate forecasting, and value leak rate are the line items. The CFOs who define those metrics now, build the instrumentation to track them, and establish the governance to act on them will be in a fundamentally different position than those who wait for the token dashboards to catch up. The gap between tracking spend and understanding value is the gap between a cost center and a competitive advantage. If you’re not tracking these three numbers across your entire AI stack today, talk to an expert about what it takes to get there.

  • Power, Casual, New, Idle: How Olakai’s Adoption Cohorts Find Your Wasted AI Licenses

    Power, Casual, New, Idle: How Olakai’s Adoption Cohorts Find Your Wasted AI Licenses

    A company buys 200 seats of Claude Code or Cursor — a familiar story in the era of AI coding tool sprawl. Six months later, the adoption dashboard reports “82% activated” and everyone treats that as a win. It might be. It might also mean 40% of those developers opened the tool once, wrote one throwaway prompt, and never came back — activated isn’t the same as valuable, and a single company-wide percentage can’t tell the difference. Olakai’s adoption cohorts exist specifically to make that distinction, on the Developers tab of the AI Impact Dashboard, at the level of an individual developer rather than a rounded company average.

    Four cohorts, one ratio

    Every developer is assigned to one of four cohorts based on the share of their merged pull requests that show AI assistance. Power means more than 70% of their PRs are AI-assisted. Casual sits between 20% and 70%. Idle is under 20%. New is a fourth category that cuts across the ratio entirely: a developer whose first AI-assisted PR landed within the last 14 days is New, regardless of what their ratio looks like.

    That last rule is a genuinely thoughtful piece of the design, and it’s worth explaining why it exists rather than just stating it. New is evaluated first and wins over the ratio bands — a developer who shipped their first AI-assisted PR this week counts as New even if literally every PR they’ve merged so far is AI-assisted. Without that override, a brand-new user’s tiny, noisy sample would misleadingly register as “Power” the moment they merged two or three AI-assisted PRs in a row, which is a worse signal than an honest “too early to tell.” The New cohort exists so the ratio-based cohorts describe settled behavior, not a small sample still finding its footing.

    Cohort assignment isn’t a permanent label, either. It recomputes every time the dashboard loads, based on whatever time range is currently selected — a developer who’s New today can become Power or Casual within a few weeks as more PRs accumulate, and the cohort you see reflects the current window rather than a badge assigned once and left stale.

    The table under the label

    The cohort itself is a starting point, not the whole answer — the per-developer table underneath is where a specific, defensible decision actually gets made. Each row shows total PRs in the selected window, AI ratio, lines moved, a 0-100 prompting-clarity score, real-time estimated cost against actual billed cost from the Admin API, lines moved per dollar, and hook coverage — how much of a developer’s agent telemetry is actually showing up against their PR activity. That last column matters more than it sounds like it should: a developer who looks Idle by PR ratio but has strong hook coverage and heavy agent usage that just hasn’t produced a merged PR yet is a different conversation than a developer who’s genuinely not using the tool at all.

    This is the layer that turns “adoption is low” from a vague, company-wide complaint into a specific, actionable list. An Idle developer with a paid seat isn’t an abstraction — it’s a named line item a VP of Engineering can pull up, with real usage data attached, and act on directly: re-train, reassign the license to someone on the waitlist, or have an honest conversation about whether the tool fits how that person actually works. None of that is possible from a single “82% adoption” slide.

    The license-waste math, made concrete

    Run the numbers on that 200-seat example: at even a modest per-seat price, a company with 40 genuinely Idle developers — not just under-adopting, but under 20% AI-assisted with no meaningful hook coverage either — is paying full price for licenses nobody is using. That’s not a governance abstraction or a productivity-culture problem to solve with a training session no one attends. It’s a specific, controllable line item that shows up the moment someone actually looks at the cohort table instead of the headline adoption percentage, and it’s exactly the kind of finding that turns into a real budget conversation rather than a vague resolution to “drive more adoption” next quarter.

    Adoption cohorts sit next to a related, deeper diagnostic worth knowing about even if it’s a separate feature: each developer’s detail page includes a Fluency tab reporting on their specific AI fluency dimensions and where they have room to grow, which goes further than the cohort label alone when the goal is coaching rather than license reallocation.

    The underlying value here is the same thread running through the rest of Olakai’s AI coding analytics: a business-friendly interface that turns raw Git activity into a decision a non-engineering manager can actually act on, rather than a chart that requires an engineer to interpret before anyone can use it.

    Curious how many of your own AI coding seats are Power, Casual, New, or genuinely Idle? Talk to an Expert.

  • Inside AI Spend Governance: Budgets and the Alerts That Fire Before You Blow Through Them

    Inside AI Spend Governance: Budgets and the Alerts That Fire Before You Blow Through Them

    A CFO rarely discovers an AI spend problem from a dashboard. They discover it from an invoice, weeks after the spending already happened, with no window left to do anything but ask engineering what happened. Olakai’s Budgets feature inside AI Spend Governance exists specifically to close that gap — not by promising a smarter forecast, but by being honest about what a budget actually is and firing an alert while there’s still time to act on it.

    Budgets are lenses, not a partition

    The single most important thing to understand about Olakai’s budgets is also the thing most competing tools obscure: budgets do not partition spend. They are independent, overlapping lenses over the same dollars. The same charge can count toward a developer’s budget, their department’s budget, the provider budget, and the program budget, all at once. If you add up every individual budget on the page expecting the total to match your program spend, it won’t — and that’s by design, not a bug to file a ticket about.

    Budgets are organized into four groups. Program is the master ceiling — every developer, every provider, every project rolled into one account-wide cap. Provider lets you cap a single vendor, like Anthropic or Cursor, without touching anyone else’s spend. Employee-centric budgets attribute spend to people — an individual developer, the persona they belong to, or their department — with the same overlap rule: a shared engineer’s spend counts fully against every relevant lens. Project groups shared service keys into named cost centers, with developers who belong to multiple projects contributing their full spend to each one rather than having it split proportionally.

    What budgets don’t track, and why that’s deliberate

    Budgets only track per-token, cost-bearing providers — Anthropic, Cursor, and OpenAI. GitHub Copilot is excluded outright, because its pricing is seat-based rather than usage-metered, and a per-token budget mechanism has nothing to measure against a flat subscription fee. Spend that can’t be attributed to a specific person still counts toward the program and provider totals, but drops out of the developer, persona, and department lenses entirely — which means per-entity totals can legitimately be lower than the program total, and that’s worth knowing before a CFO tries to reconcile the two and assumes something’s broken. The same per-provider precision shows up in how Olakai handles Google Vertex AI cost data, which lags behind usage by design rather than pretending to be real-time when it isn’t.

    Budgets extend naturally into Projects, Olakai’s term for a cost center: a named bucket that groups shared service API keys, member developers, and owned repositories under one monthly limit. Total project cost is service-key spend plus member-developer token spend, and — consistent with the overlapping-lens rule everywhere else — a developer who belongs to more than one project contributes their full token spend to each project they’re in, not a proportional split. Archiving a project unassigns its keys and removes its budget and alert rules, but never deletes the underlying spend history; only the grouping goes away.

    The forecast that’s already live, and the one that isn’t yet

    Every budget carries a Projected month-end figure built the same way: the recent daily spend rate, extended across the remaining days of the month, added to spend so far, with a confidence signal that reflects how steady daily spend has actually been. It’s a run-rate projection, not a model of growth or seasonality — a distinction Olakai is upfront about, the same way it’s upfront about the 30-day spend projection on the main AI Impact Dashboard being a straight-line extrapolation rather than a real forecast. Budgets are evaluated once daily, right after the cost-import job pulls fresh provider spend, and saving a budget automatically provisions the alert rules behind it — nothing extra to configure.

    Two kinds of alerts come out of that evaluation. The first is a threshold alert: actual month-to-date spend crosses a configurable percentage of the budget — 50%, 80%, or 100%. The second, and the more useful one, is a forecast alert: the run-rate projection is on track to exceed the limit by month end, even if the account isn’t over budget yet today. That second alert is the actual point of the feature — catching a trajectory early enough to still change it, rather than confirming after the fact that the month already went over.

    Worth being precise about scope here: Olakai also has a separate, standalone Forecasts tab planned for what-if scenario modeling across budget dimensions — a different, more ambitious feature for testing hypothetical spend trajectories before committing to them. As of this writing, that tab isn’t live yet. What’s covered above — the run-rate projection and the two alert types built directly into the Budgets page — is shipped and running today; the scenario-modeling tool is a separate thing worth revisiting once it ships.

    Why the overlap is the right design, not a shortcut

    It would be simpler to build budgets that partition spend cleanly — every dollar assigned to exactly one bucket, everything adding up neatly on a summary slide. It would also be wrong for how AI spend actually happens inside a real engineering org, where the same developer’s usage genuinely belongs to their department’s headcount planning, their manager’s persona-level benchmarking, the vendor contract renewal conversation, and the specific project that consumed it — four legitimate, simultaneous questions about the same dollar. Building four separate, non-overlapping ledgers to answer four different questions would mean picking one authoritative answer and getting the other three wrong. Overlapping lenses let all four questions get an honest answer from the same underlying spend data, at the cost of a program total that doesn’t equal the sum of its parts — which is exactly the tradeoff a CFO should want once it’s explained, rather than discovered while trying to make the numbers reconcile.

    It’s also worth knowing that budgets and projects aren’t limited to point-and-click configuration — Kai can create, edit, and archive them conversationally, gated by admin permission and a confirmation step before anything actually changes, which is a useful shortcut when there’s a long backlog of unassigned service keys to triage.

    This isn’t a hypothetical risk — Uber blew through an entire year of AI budget in four months before building a reactive cap of its own. If your AI coding spend has outgrown a spreadsheet and a monthly Slack message from finance, Talk to an Expert about setting up budgets against your own provider and project data.

  • AI Coding Tool ROI: Why Acceptance Rate Is the Wrong Metric

    AI Coding Tool ROI: Why Acceptance Rate Is the Wrong Metric

    In May, Gartner published its first-ever assessment of the enterprise AI coding agent market, formalizing a category that did not exist as a named market segment two years ago and now runs to roughly ten billion dollars a year. The message between the lines was clear: of all the places enterprises have poured AI money, software development is where the returns look most real. So here is the question every engineering leader should sit with. If coding is the one domain where AI value is most provable, why can almost nobody prove it?

    Most organizations buy seats, watch a vendor dashboard tick upward, and conclude things are working. The dashboard shows suggestions made, suggestions accepted, an acceptance rate climbing past 30%. It feels like proof. It is not. Acceptance rate is the single most misleading number in the entire AI coding conversation, and the gap between what it measures and what actually matters is where engineering budgets quietly lose their justification.

    Coding really is different

    The optimistic case for AI coding tools is genuine, and it deserves a fair hearing before the skepticism arrives. GitHub’s own controlled study found developers completing a programming task 55% faster with an assistant than without. The market reflects that promise: AI coding tools now represent well over ten billion dollars in annual spend, and roughly 90% of the Fortune 100 have deployed GitHub Copilot in some form. Gartner’s decision to stand up a formal market assessment is itself a signal that coding has matured past experimentation into something boards expect to pay off.

    That maturity is exactly why coding deserves better measurement than the rest of the AI portfolio, not worse. Gartner’s parallel research on AI in infrastructure and operations found that only 28% of those use cases fully succeed. Coding stands out as the exception, the place where the productivity story has the most evidence behind it. When you have found the one room with treasure in it, you do not measure your haul by counting how many times you opened a drawer.

    But the dashboard is lying to you

    The cleanest evidence that activity metrics mislead comes from a randomized controlled trial. METR studied experienced open-source developers working in codebases they knew well, and found they were 19% slower when using AI assistance. The detail that matters most for measurement: those same developers estimated they had been 20% faster. A nearly forty-point gap between perceived and actual productivity, in the population most enterprises are deploying these tools to. If your ROI case rests on developer self-report or on a feeling that the team is moving quicker, that is the gap you are standing on.

    The quality picture is just as sobering. GitClear’s analysis of 211 million changed lines of code found that copy-pasted and duplicated code blocks rose eightfold in a single year, code churn climbed, and the share of lines devoted to refactoring fell to under 10%. AI makes it trivial to add code and does nothing to encourage consolidating it. Google’s 2025 DORA research found the same tension from a different angle: AI adoption correlated positively with throughput but negatively with delivery stability, meaning the tools that help you ship faster can quietly erode the controls that keep what you ship from breaking. Acceptance rate captures none of this. A developer can accept every suggestion and ship slower, buggier software, and the dashboard will call that a win.

    Activity versus value: the real metric problem

    The reason vendor dashboards surface acceptance rate, lines generated, and seat utilization is that these are the metrics the vendor controls and optimizes for. They describe how much the tool was used, not what the use produced. That distinction is the whole game, and it is the same vanity-versus-value problem we mapped for finance leaders in the metrics that actually matter. An engineering org running on acceptance rate is measuring the proxy and ignoring the signal.

    The signal lives in a different set of numbers. How does cycle time differ between AI-assisted pull requests and the rest? What is your cost per merged PR once you divide total tool spend across providers by the work actually shipped? How has defect density moved since rollout, and which teams are driving the change? Which developers have genuinely adopted the tools, and which licenses are sitting idle at $19 to $50 a head every month? Answering those questions requires connecting pull-request data, provider cost data, and engineering outcomes in one place, which is precisely what Coding IQ was built to do. It measures the value of AI coding tools rather than the activity, because activity was never the thing the CFO was paying for.

    What good measurement actually enables

    This is not an argument that AI coding tools do not work. It is an argument that you cannot manage what you measure badly. The enterprises pulling real value from these tools are the ones that instrumented outcomes before scaling seats, the same discipline that separates winners across every category of AI investment in the broader ROI playbook. They can make decisions the acceptance-rate crowd cannot.

    Consider the difference at a budget review. An engineering leader who can say that Cursor users close pull requests 28% faster than non-users at a cost of a few dollars per PR, while 40% of Copilot licenses sit unused, is making a business decision: scale the first, reclaim the second. A leader who can only report a 32% acceptance rate is reporting a vendor metric and hoping nobody asks what it bought. That is the position most VPs of engineering find themselves in, and it is an avoidable one. The instrumentation that closes the gap is the same vendor-neutral measurement layer that proves AI ROI across the rest of the stack, applied to the one domain where the returns are most worth proving. It is also the only honest way out of the trap NVIDIA documented when it found 30% of enterprises still cannot quantify AI ROI at all.

    Coding is where enterprise AI ROI is most real. That makes it the worst possible place to keep measuring the wrong thing. Acceptance rate will tell you your developers are clicking accept. It will never tell you whether your software is better, faster, or cheaper to ship, which is the only question your board is actually asking.

    Is your coding-tool spend producing value, or just activity? Talk to an expert to see how Olakai’s Coding IQ ties AI coding tools to cycle time, defect rate, and cost per pull request, so you can scale what works and cut what doesn’t.

  • The Real Bill, Not a Guess: How Olakai Reconciles Google Vertex AI Costs to BigQuery

    The Real Bill, Not a Guess: How Olakai Reconciles Google Vertex AI Costs to BigQuery

    Most AI coding-cost dashboards show a number and let you assume it’s a bill. Sometimes it is. Sometimes it’s a model built on token counts and a price sheet, dressed up to look exactly as confident as a real invoice. For a CFO signing off on a growing AI coding spend line, that distinction is the whole ballgame — and it’s one most AI analytics platforms don’t bother making. Olakai does, and Google Vertex AI is the clearest example of why it matters.

    Usage is real-time. Cost isn’t — and the product says so.

    Vertex AI token usage syncs into Olakai in real time — every request, every token, visible almost as soon as it happens. Cost is a different story. Google’s Cloud Billing export into BigQuery, which is what Olakai reconciles against for the real, billed dollar figure, typically takes 24 to 48 hours to catch up. Until it does, the Total Cost stat card, the Cost by Model breakdown, and the Cost by Developer table on the Vertex page all read $0 or an explicit “No cost data yet” — and the page carries an inline notice above the summary cards saying exactly that, not a footnote buried in documentation. Recent spend reads low for a day or two, on purpose, because the alternative is guessing.

    That’s a real tradeoff, not a cosmetic one. A dashboard that always shows a confident-looking cost number, updated in real time, is easier to build and easier to demo. It’s also wrong for the most recent day or two of every reporting period, every time, without telling you. Olakai’s Vertex integration chose the less flattering, more accurate option: token usage updates immediately because it’s genuinely known immediately, and cost updates only once it’s genuinely known — reconciled to Google’s actual bill via the Cloud Billing → BigQuery export, not modeled from a public price sheet.

    Three providers, three different honesty postures

    Vertex isn’t the only provider with a gap between what’s easy to show and what’s actually true — it’s just the most transparent about it. Anthropic and Cursor’s cost data comes straight from their Admin APIs in real time, with no lag at all; what you see is what’s billed, as soon as it happens. OpenAI’s Admin API has a different limitation entirely: it doesn’t expose a user dimension on usage or cost data, so every cost figure on Olakai’s OpenAI page is tracked per API key, not per developer, until someone manually maps a key to a person through Developer Identities. Olakai surfaces that as an inline notice on the OpenAI page too, the same way it surfaces the Vertex cost lag.

    Line those three up and you get three genuinely different honesty postures across five supported providers: Anthropic and Cursor (real-time, no gap), OpenAI (real-time, but per-key rather than per-person until mapped), and Google Vertex (accurate, but lagged 24-48 hours, reconciled to an actual invoice rather than estimated). Most cost dashboards flatten this into one undifferentiated “estimated cost” figure across every vendor. Treating five providers as five different measurement problems, and disclosing the difference in the product itself, is a small thing that adds up to a much more defensible number when it lands in front of finance.

    Unattributed spend, by design

    Per-developer cost allocation for Vertex works by taking each developer’s share of tokens consumed and reconciling that share against the actual billed total once it lands — an allocated figure, not a metered one, and Olakai is specific about which of the two it is. There’s a real limitation baked into that method worth stating plainly: per-developer Vertex cost specifically covers Gemini CLI usage. Antigravity and other Vertex activity that Olakai can’t track at the token level carries no token count to allocate by, so that spend lands in a distinct “Unattributed” row instead of getting force-divided across developers who may not have generated it.

    That’s the same instinct as the cost-lag notice, applied to a different problem: when the system doesn’t have a reliable basis for attributing a dollar to a specific person, it says “Unattributed” instead of guessing. For a finance team trying to reconcile AI coding spend against headcount and productivity, an honest “we don’t know whose this is” line is more useful than a falsely precise number that quietly absorbs error into every developer’s total.

    Why this matters more than it sounds like it should

    None of this is exotic engineering — it’s disclosure. But disclosure is exactly what’s missing from most AI cost tools, and it’s exactly what a CFO needs before treating a number as board-ready. “This $4,200 is Google’s actual invoice, reconciled through BigQuery” and “this other number is a real-time token estimate that hasn’t been billed yet” are two different claims, and conflating them into one undifferentiated “AI spend” figure is how companies end up surprised by their own AI bill months after the fact — the same surprise Olakai’s budget alerts are built to prevent going forward.

    The broader point of building this level of provider-specific honesty into the product is the same one that runs through the rest of the AI Impact Dashboard: a vendor-neutral platform that treats every provider’s data the way that provider’s data actually behaves, rather than smoothing five different billing realities into one uniform-looking chart. It’s a less impressive-sounding pitch than “real-time cost visibility across every AI vendor.” It’s also the one that survives an audit.

    Want to see exactly which of your AI coding costs are real invoices and which are estimates waiting to reconcile? Talk to an Expert.

  • How Olakai Detects AI Coding Tool Usage Without Installing a Single Agent

    How Olakai Detects AI Coding Tool Usage Without Installing a Single Agent

    The first objection a CISO raises to almost any AI analytics pitch is some version of the same question: “so now I need to put another agent on every developer’s laptop?” It’s a fair question — security teams have spent years fighting endpoint sprawl, and a new mandatory install is a real cost even when the tool behind it is useful. For Olakai’s AI coding tool detection, the answer is no. Nothing runs on a developer’s machine. The entire signal comes from pull request data that already exists in GitHub, Bitbucket, or GitLab, read through the same organization-level credential used to power the AI Impact Dashboard.

    That’s a meaningfully different trust posture than most AI monitoring tools, and it’s worth walking through exactly how the detection actually works — not as a black box, but as a stack of specific, checkable signals, with an honest fallback for the cases none of them catch.

    Three signals, applied in order

    Every merged pull request is checked against three detection methods. First, bot author detection: was the PR opened by a known bot account — dependabot, renovate, devin-ai, copilot-swe-agent, claude-code, sweep-ai, snyk-bot, and others recognized by GitHub’s own `user.type === “Bot”` flag? Second, commit co-author trailers: do any commits in the PR carry a Co-authored-by: line matching a known AI tool pattern, covering GitHub Copilot, Cursor, Claude Code, Devin, Amazon Q, and Gemini? Third, PR title and body markers: does the PR’s own text contain a recognizable phrase — “Generated by Cursor,” “claude-code,” “Copilot Workspace,” “Created by Devin,” and similar strings that AI tools leave behind by convention?

    If any one of those three signals fires, the PR is classified AI-assisted — and a single PR can be attributed to more than one tool if different commits carry different signals, which happens more than you’d expect on PRs where a human picks up and finishes AI-started work. A PR earns the stronger label of fully agentic specifically when a bot account opened it: the AI created the branch, wrote the code, and opened the PR itself, without a human author in the loop at all. That’s a real distinction, not a marketing one — a fully agentic PR is a different governance conversation than one where a human developer used AI as a fast collaborator.

    What happens when none of the three signals fire

    Not every AI-assisted PR leaves a clean marker. A developer might paste AI-generated code into a normal commit with no trailer, no bot account, and no mention in the PR body. For that gray area, Olakai runs an LLM classifier — Claude Haiku 4.5 — that reads the PR title, body, and a sample of the diff, and assigns a confidence score. Critically, the PR is only marked AI-assisted through this path at 60% confidence or higher; anything below that threshold is treated as “not detected,” specifically so ambiguous cases don’t get counted and inflate the AI-assisted total.

    That’s a deliberate design decision worth sitting with for a second: the system is built to under-count in ambiguous cases rather than over-claim. Most vendors pitching AI-coding-tool ROI have every incentive to inflate the “AI is helping” number — a bigger adoption percentage is a better story. Olakai’s classifier does the opposite by default, which is exactly the kind of choice that should show up in an honest AI coding tool ROI metric instead of a vanity one.

    Multi-provider integrity, and where the fidelity differs

    Three source-control providers feed the same underlying pull-request table: GitHub, the production integration with the most mature detection signal set; Bitbucket Cloud, currently in beta; and GitLab, covering both SaaS and self-managed instances, using the same signal types plus GitLab-specific Duo markers. Every row is uniquely keyed by account, provider, repository, and PR number, so an organization running both GitHub and Bitbucket at once never gets a collided or double-counted PR — the two providers’ data sits cleanly side by side in the same analytics.

    Fidelity isn’t identical across providers, and Olakai says so rather than presenting every provider as equivalent. Bitbucket and GitLab have no first-class “review” object the way GitHub does, so review rounds and approvals are derived from each platform’s activity stream instead of native review submissions — directionally correct, but lower-fidelity than GitHub’s numbers. Any cross-provider cycle-time comparison should carry that caveat rather than treating a GitHub PR and a Bitbucket PR as measured on perfectly identical instruments.

    Whose PR is this, exactly?

    Detecting AI usage is only half the problem; attributing it to the right developer is the other half, and it’s less glamorous but just as important. Olakai resolves author identity through a four-step priority order: first, the commit email from a commit whose author matches the PR opener, excluding GitHub’s own noreply addresses; if that’s unavailable, the email from the PR’s first commit; failing that, the email on the GitHub user profile; and as a last resort, a constructed noreply fallback built from the GitHub login. That chain exists because real organizations have messy Git configuration — different emails on different machines, corporate SSO aliases, personal accounts used for a first commit — and per-developer ROI attribution is worthless if it silently drops or misattributes a chunk of activity because someone’s `git config` didn’t match their Olakai account exactly.

    The same pull-request pipeline also captures signals beyond raw detection — PR size classification from XS to XL by lines added (used to flag AI PRs that are suspiciously always large, which the documentation itself calls a potential sign of rubber-stamping), issue linkage via Fixes #N or Closes #N references, first-pass approval rate, and a test-file ratio that the product is explicit about not being perfect: it’s based on file count, not line count, so a PR with one large test file and many small source files reads as low-coverage even when it isn’t. Publishing that limitation alongside the metric is the same pattern as the LLM classifier’s 60% confidence floor — under-claim rather than over-claim.

    Put together, the detection layer is the foundation everything else in Olakai’s AI monitoring is built on — the PR mix, the cycle-time comparisons, the productivity score, all of it depends on getting “was this PR actually AI-assisted, and by which tool” right at the source, without asking a single developer to install anything. For a CISO evaluating the tool, that’s the actual pitch: the analytics run on data your organization already generates and already controls access to, not on a new agent asking for a new set of permissions.

    Want to see what this detection stack finds in your own repositories before rolling it out further? Talk to an Expert.

  • The Return of the Desktop App: And the AI Measurement Gap It Creates

    The Return of the Desktop App: And the AI Measurement Gap It Creates

    For 25 years, the entire direction of travel in enterprise software was the same: everything moved to the browser. Salesforce on CDs gave way to Salesforce in a tab, Office gave way to Google Docs, Sketch gave way to Figma, and every installer eventually got replaced by a URL. The logic behind that shift was airtight. Zero friction to distribute, one codebase across every operating system, native multiplayer, continuous deployment, and a subscription revenue model that buyers actually preferred. The web won so decisively that even Adobe capitulated to subscription pricing in 2013, and Microsoft declared itself “cloud-first” within 52 days of Satya Nadella taking over in 2014. If you were building software in 2020 and told a VC you were shipping a desktop app, you were laughed out of the room.

    And then, somewhere in the last 18 months, every AI-native company that could have stayed browser-only started shipping desktop apps instead.

    OpenAI released a ChatGPT Mac app in May 2024, before they had reached feature parity on mobile. Anthropic followed with Claude desktop in November, alongside the Model Context Protocol, which went from around 2 million to 97 million monthly SDK downloads in 16 months. The entire point of MCP is giving AI access to the local filesystem that browsers cannot reach. Perplexity shipped a native Mac app. Cursor, a desktop IDE you download and install the old-fashioned way, is reportedly in talks to raise at a $50 billion valuation, which is roughly the last price tag attached to a desktop-first software company when that company was Microsoft.

    Meanwhile Ollama, which exists purely to run AI models locally on your laptop with no API call involved, went from around 100,000 monthly downloads in early 2023 to over 52 million in early 2026. That is a 520x increase in three years for a product whose defining feature is that it does not touch the cloud. And Microsoft, the same Microsoft that was cloud-first, now mandates that every Copilot+ PC ship with 40 trillion operations per second of on-device AI silicon. The company that spent a decade telling its customers to move everything to Azure is now re-engineering consumer PCs around local inference.

    The Web Won on Five Things, and AI Wants All Five Reversed

    Every piece of enterprise software that moved to the browser did so because the browser offered five structural advantages: zero-friction distribution, SaaS economics, native multiplayer collaboration, cross-device access, and continuous deployment. For most categories, those advantages were decisive. Desktop software only held on in the handful of places where GPU access, filesystem access, offline reliability, or sub-10-millisecond latency were in the critical path. Video editing stayed on the desktop. So did CAD, IDEs, gaming, and anything that needed to push pixels or bits in real time. Those constraints were not ideological. They were physics.

    And every serious AI workload happens to sit squarely inside them. AI agents need to read your actual codebase, not whatever you remembered to paste into a chat window. They run for minutes or hours, not the lifetime of a browser tab that the operating system feels free to suspend the moment you switch windows. Meeting copilots need raw screen and audio access that browsers wall off by design, for good security reasons. Voice AI and autocomplete UX fall apart the moment you introduce a network round-trip, which is why Cursor feels instant and most browser-based AI tools feel laggy. The same constraints that kept Premiere on the desktop in 2005 are now shaping the entire AI application layer in 2026, which means for the first time in a generation the list of software categories that have to live on your machine is growing rather than shrinking.

    And That Creates a Measurement Problem

    Here is where this gets interesting for anyone trying to run an enterprise AI program.

    When AI lived in the browser, you could measure it. Your employees logged into ChatGPT through a centralized account, or they used a SaaS tool whose admin console told you exactly who used what and when. Single sign-on, audit logs, API gateway usage reports, the entire governance stack that evolved for SaaS could be pointed at AI with a few configuration tweaks. The web’s centralization was a pain point for vendors in 2000 and a gift to CIOs in 2020. Everything flowed through a known endpoint, and everything left a trace.

    The desktop renaissance is dismantling that model, category by category, in a matter of months.

    A developer using Cursor is running AI inference against your codebase on their local machine, and your IT team cannot see what they are doing through any centralized log. A knowledge worker using Claude desktop is having conversations with a foundation model that may or may not touch your network. A sales leader using Granola is recording every meeting on their device, with no browser session to inspect. A product team experimenting with Ollama is pulling seventy-billion-parameter models down from Hugging Face and running inference entirely offline, with no API call that your network observability tools can capture. The shadow AI problem that was already keeping CISOs up at night is about to get qualitatively worse, because the new generation of AI tools is specifically engineered to bypass the centralized chokepoints that corporate governance depends on.

    You cannot measure what you cannot see, and you cannot govern what you cannot measure. The measurement gap that enterprises are already struggling to close in their AI ROI programs is about to widen significantly, precisely at the moment when boards and CFOs are starting to demand proof of value.

    What the Old Playbook Got Wrong

    For years, the default AI governance playbook at most enterprises has been some version of: restrict access to sanctioned tools, route traffic through an approved gateway, and generate usage reports from the gateway logs. That playbook works reasonably well when the AI tool in question is a cloud-hosted chatbot that an employee reaches through a browser. It falls apart the moment the AI tool is a desktop app that talks directly to a foundation model provider, or worse, runs inference on the laptop itself.

    The uncomfortable truth is that measurement and governance in an AI-first enterprise cannot be built from the network layer or from SaaS admin consoles alone. You need a visibility layer that works across cloud, browser, desktop, and local-inference environments, and that treats each AI interaction as an observable event regardless of where the compute happened. You need metrics that CFOs actually want to see rather than vanity counts of API calls. And you need a governance model that assumes AI usage is heterogeneous and distributed by default, not centralized and inspectable by default.

    What Leaders Should Do Now

    The shift to AI-native desktop is not a reason to panic, and it is not a reason to try to block desktop AI apps. Every serious study on enterprise AI adoption points to the same conclusion: knowledge workers will use the tools that make them productive, and the companies that lean into that rather than fighting it capture disproportionate value. The question is not whether to allow your teams to use Cursor and Claude and Ollama. The question is whether you can see enough of what is happening across all of them to understand the true ROI of agentic AI, catch governance failures before they become incidents, and make informed decisions about where to invest next.

    That starts with accepting that your AI measurement layer needs to extend into the desktop, the IDE, and the on-device inference runtime, not just the browser. It continues with building unified AI analytics that aggregate events from across environments into a single view. And it ends with a governance model that is resilient to heterogeneity, because the direction of travel for the next five years is more AI, in more places, running on more devices, against more models, not less.

    The desktop is back. The browser is not going away. Most enterprises will run both, permanently. The organizations that win will be the ones that can see across both, measure across both, and make decisions grounded in that visibility.

    If you are thinking about how to build that measurement layer inside your organization, we would love to talk.

  • Tokenmaxxing Is the New Lines of Code: Why Token Leaderboards Won’t Prove AI Value

    Tokenmaxxing Is the New Lines of Code: Why Token Leaderboards Won’t Prove AI Value

    Someone at Meta built a leaderboard called Claudeonomics. It ranked employees by the number of tokens their AI models processed and generated. Top spenders got rewards. Then it leaked to the press. Then Meta quietly shut it down.

    That was earlier this month. This week, Reid Hoffman came out in measured defense of the practice at Semafor’s World Economy Summit. On the same day, an inference-infrastructure startup called Parasail raised $32 million on the thesis that “tokenmaxxing” will create the next compute giant. The company already generates 500 billion tokens a day.

    If you run an AI program and you haven’t yet been asked by your CEO or your board why your engineers aren’t in the top quartile of token consumption, you will be soon. And when that conversation arrives, you need a better answer than a bigger number.

    What tokenmaxxing actually measures

    A token is a small chunk of text an AI model processes. Every prompt consumed, every response generated, every line of code auto-completed — they all add up to a token count. “Maxxing” is Gen Z slang for optimizing something to the extreme. Put them together and you get the idea: rank employees by how many tokens they burn, and call the top of the list your best AI adopters.

    Meta built the internal dashboard. Shopify folded AI usage into performance reviews. Venture capital is now funding the picks and shovels. Hoffman’s defense, notable because he’s one of the more careful voices in the debate, was a cautious endorsement: “You should be getting people at all different kinds of functions actually engaging and experimenting [with AI].” He then immediately added that token tracking “doesn’t mean it’s a perfect example of productivity.”

    Read that second sentence again. The strongest public defender of tokenmaxxing concedes, in the same breath, that it doesn’t measure productivity. Which raises the question everyone at Meta was too polite to ask before the leaderboard leaked: what exactly are we measuring, and why?

    The new lines of code

    If this pattern feels familiar, it should. For decades, engineering organizations tried measuring developer productivity by lines of code written. The metric was easy to count, easy to rank, and spectacularly broken. Engineers who wrote terse, elegant code scored poorly. Engineers who produced verbose, repetitive code scored well. Every competent engineering leader learned the lesson the hard way: when you turn an input metric into a target, people optimize for the metric, not the work.

    Economists call this Goodhart’s Law. When a measure becomes a target, it ceases to be a good measure. Token consumption is lines of code with a fresh coat of paint. It’s an input. It’s easy to count. And it tells you almost nothing about whether the work the AI produced was useful, correct, or worth the compute bill that came with it.

    The cynical version of tokenmaxxing plays out predictably. Employees pad their AI usage with throwaway prompts. Managers celebrate the chart going up. Finance sees the OpenAI and Anthropic invoices climbing and asks what changed. Nobody can tell them, because the leaderboard only measures spend. We covered this exact anti-pattern in AI Metrics That Matter — the gap between what’s easy to count and what a CFO actually wants to see.

    Why it’s seductive anyway

    Tokenmaxxing isn’t popular because executives are naive. It’s popular because real AI measurement is hard and token counts are sitting right there in the API billing dashboard. When a CEO asks the head of AI whether the organization is actually using its new tools, “we processed 4.2 billion tokens last quarter, up 340%” is a satisfying answer to give. It’s specific. It’s directional. It trends up and to the right.

    It’s also, as NVIDIA’s recent survey of 3,200 enterprise leaders revealed, roughly the level of measurement most organizations have settled for. As we covered in our analysis of the NVIDIA State of AI report, 30% of enterprises still cannot measure the ROI of their AI investments at all. Token counts are what you reach for when you’ve given up on measuring the thing you actually care about.

    The other reason tokenmaxxing spreads is that it pushes a real problem — AI adoption — through an easy pipe. In most enterprises, the gap between AI tool licenses purchased and AI tools actually used by employees is enormous. Licenses go unclaimed. Copilots go idle. Shadow AI proliferates in the gap. Counting tokens at least tells you who’s trying something. But “trying something” is a foundation for measurement, not its destination.

    What outcome-based measurement looks like

    The measurement you want isn’t on the API invoice. It’s in the business system the AI was supposed to change. If your developers are using AI coding tools, the question isn’t how many tokens they generated — it’s whether cycle time dropped, whether pull request quality held, whether production incidents stayed flat. If your sales team is using an AI assistant, the question is whether deal velocity improved, not whether reps sent more prompts.

    This is the measurement layer missing from almost every tokenmaxxing dashboard we’ve seen. It’s also the layer that Coding IQ and the rest of the Olakai platform exist to provide. The question we ask our customers to answer isn’t “how much AI did you use?” It’s “what did your AI produce, for whom, and at what business outcome?” Those three questions are the ones a CFO will ask when the bill arrives, and the ones a CISO will ask when governance gets challenged.

    We built an entire framework around this. We call it SEE → MEASURE → DECIDE → ACT. SEE surfaces every AI tool in use, not just the sanctioned ones. MEASURE ties usage to business KPIs the executive team already cares about. DECIDE gives you the evidence to scale, fix, or kill each pilot. ACT turns the answers into an operating rhythm instead of a once-a-quarter scramble. None of those steps begin with token counts. All of them produce numbers your board will actually recognize as value.

    The governance blind spot

    There’s a second problem with tokenmaxxing that rarely gets discussed. A leaderboard that rewards token spend creates an incentive to bypass governance controls to get more of it. Employees who find a sanctioned tool too slow, too throttled, or too narrow in capability will reach for something unsanctioned. Shadow AI already grew fast in the absence of measurement. Adding a scoreboard that rewards consumption accelerates it.

    This is the worry that haunts every CISO we talk to, and it’s why the CFO view and the CISO view of AI can’t live in separate dashboards. You cannot measure AI ROI without measuring AI risk, because the risk is the other half of the cost. Tokenmaxxing, by design, only counts one side.

    Getting started: audit what your AI produces, not what it consumes

    If your organization is under pressure to show AI adoption and you’re being nudged toward tokenmaxxing, there’s a better first step. Pick the three most visible AI deployments in your organization — coding assistants, a customer support copilot, a sales enablement tool — and, for each one, write down the business outcome it was supposed to change. Cycle time. First-contact resolution. Win rate. Whatever it is, write it down. Then measure whether the outcome moved. Our AI ROI framework walks through this end to end.

    Do that for three deployments and you’ll know more about the real state of AI in your organization than any token leaderboard will tell you. You’ll also have the beginnings of a measurement system that survives the next wave of AI hype, whatever it gets called. Lines of code didn’t survive the last one. Tokenmaxxing won’t survive this one. Outcomes always do.

    Olakai helps enterprises measure what their AI is actually producing — across every tool, every user, every workflow — and tie it back to the business KPIs executives already track. If tokenmaxxing is the conversation your board is having, we can help you lead a better one. Talk to an expert.

  • Inside the AI Impact Dashboard: How Olakai Turns PR Data Into Proof of AI Value

    Inside the AI Impact Dashboard: How Olakai Turns PR Data Into Proof of AI Value

    Ask most VPs of Engineering how AI coding tools are doing on their team, and you’ll get an adoption number: “80% of developers used Claude Code or Cursor last month.” That number answers a real question, but not the one the CFO is actually asking. Adoption tells you who opened the tool. It says nothing about whether the team is shipping more, shipping faster, or shipping the same amount of code with an extra subscription line item attached.

    That gap is wider than most engineering leaders assume. A Black Duck survey of over 800 enterprise software engineers and DevOps professionals, published in June 2026, found AI coding assistant adoption had hit 97% — functionally universal — while the same research pointed to governance and measurement, not adoption, as the actual multiplier on ROI. Everyone has the tool. Not everyone can prove what it’s doing. That’s the exact problem the Olakai Agentic AI Impact Dashboard was built to close, and it’s worth walking through how it actually does that, tab by tab, rather than taking the “proof of AI value” claim on faith.

    Overview: what you’re actually spending, and what came back

    The Overview tab starts with money, and it’s careful about which money is real. The Spend Summary section adds two genuinely different billing streams together: Admin API costs, which are token-billed usage pulled straight from Anthropic, Cursor, and OpenAI’s Admin APIs — the same figure that lands on the actual invoice for usage-priced plans — and licensing costs, seat subscriptions like Cursor Business or GitHub Copilot seats, prorated to the selected window. Most companies pay both at once: token bills for power users on usage plans, seat licenses for everyone else on subscription plans. Adding them together is the number finance actually writes the check for.

    What the dashboard refuses to call the resulting monthly figure is instructive. The “Projected 30 day” number is explicitly labeled a straight-line extrapolation, not a forecast — actual spend times 30 divided by 7, answering “if the next 23 days look like the last 7, what does a full month cost?” It doesn’t model growth, seasonality, or seat changes, and Olakai says so in the product rather than letting a rough projection masquerade as a confident prediction. That distinction matters more than it sounds like it should: a lot of AI analytics tools show a single “forecasted spend” number with no indication of how much confidence to put in it.

    The Codebase Outcomes section is where spend turns into a claim about output — AI Code Ratio (the percentage of lines from AI-assisted pull requests, daily) and PR Volume (AI versus non-AI PRs per day), framed explicitly as “the shipped result of the AI coding activity above.” It’s a deliberate causal chain: spend, then activity, then shipped outcome — not three unrelated charts sitting next to each other.

    PR Analysis: a three-way split, not a binary one

    The PR Analysis tab is the place most “is AI making us ship more?” conversations should start. Instead of a simple AI-versus-human split, every pull request in the window lands in one of three buckets: Fully Agentic (an AI agent drove it end to end — Claude Code, Cursor Agent, and similar), Human + AI Assisted (AI helped, but a person drove the work), and Non-AI. Each bucket shows both a count and its share of all PRs, and the distinction between fully agentic and assisted work is the kind of nuance that a single “AI adoption %” figure erases entirely — a team where AI opens and merges PRs unsupervised is a fundamentally different governance conversation than one where AI is a fast autocomplete for human-driven work.

    Underneath the headline split sits a searchable, sortable table with per-PR granularity: repository, author, percentage of AI-attributed code, which specific AI tools touched the PR, lines added and removed, cycle time, and merge date. That level of detail is what turns “our AI adoption looks healthy” into something a VP of Engineering can actually defend in a planning meeting — a specific repo, a specific tool, a specific number, not an aggregate percentage nobody can trace back to real work, and it’s the same granularity that separates a real AI coding tool ROI metric from acceptance-rate vanity numbers.

    Cycle Time: the tail matters as much as the average

    The Cycle Time tab compares AI-assisted and non-AI pull requests at three percentiles — p50, p75, and p90 — each with a “N% faster” or “N% slower” delta column. That’s a deliberate methodological choice, not an arbitrary one: p50 tells you what a typical PR looks like, while p90 tells you whether AI is helping — or actively hurting — the slow tail of your worst cases. An average alone can hide a tool that makes routine work faster while making the hard, unusual PRs meaningfully worse; percentiles don’t let that hide.

    The tab also tracks issue linkage — the average number of tracked issues linked per PR via Fixes #N or Closes #N references — alongside the overall first-pass approval rate, as a check on whether AI-assisted PRs are solving planned, tracked work or generating ad-hoc changes nobody asked for. And it’s explicit about its own limits: breakdowns are available per-repository and per-AI-tool, but there is no per-team view on this tab. If you need a team-level cut, that’s a different report, not a filter you’re missing here.

    The number that goes in the board deck

    All of this rolls up into two composite figures designed for a leadership audience rather than an engineering one. The AI Productivity Score is a 0-100 composite built from four weighted components — Adoption (25 points), Speed (30 points), Quality (25 points), and Efficiency (20 points) — giving a single number that moves as the underlying PR data moves, with an eight-week trend sparkline. Next to it sits AI Equivalent Engineers: the productivity gain expressed as “how many additional full-time engineers’ worth of output your AI tools are producing,” which converts into a quarterly dollar figure at a configurable fully-loaded engineer cost (the default is $200,000 a year).

    The honest part is what happens when the sample is small. Olakai attaches an explicit confidence tier to the Equivalent Engineers figure based on how many developers qualify for the underlying before/after comparison: High confidence needs 15 or more qualifying developers and is described as suitable for executive reporting; Medium is 5 to 14, useful for planning; Low is 3 to 4, a preliminary signal only; and below 3 qualifying developers, the guidance is blunt — do not use this for decisions. That’s an unusual thing for an analytics vendor to put in its own product: a built-in instruction not to trust its own headline metric below a stated threshold. For a platform built on the premise of vendor-neutral, board-ready proof rather than vanity dashboards, that kind of restraint is the point, not an afterthought.

    Two things the Impact Dashboard covers in more depth than this post has room for: the Developers tab, which breaks adoption down into cohorts covered in depth in Power, Casual, New, Idle, and the Productivity tab’s before/after methodology, which compares each developer against their own historical baseline rather than against peers. Both are real, separately documented mechanisms worth their own explanation.

    None of this requires installing anything new on a developer’s machine — it runs on pull request data Olakai already has access to through your connected GitHub, Bitbucket, or GitLab organization, which is also why it works the same way regardless of which AI coding tools your teams actually use. That vendor-neutral posture is what makes the dashboard useful for a mixed fleet — Claude Code here, Cursor there, Copilot somewhere else — instead of a single-vendor usage report dressed up as an ROI tool, and it’s the same reason generating code isn’t the same as generating value across a mixed toolset.

    If your organization already has a mixed toolset and a growing AI coding bill, the harder question isn’t whether to measure impact — it’s whether the number you’re currently reporting up would survive this level of scrutiny. Talk to an Expert to see the AI Impact Dashboard against your own repositories.

    Sources: Black Duck, “AI Coding Hits 97% Enterprise Adoption,” June 2026.