Almost every enterprise finance team is now in the AI cost business, and almost none of them can see the whole bill. The FinOps Foundation’s State of FinOps 2026, drawn from 1,192 practitioners representing more than $83 billion in annual cloud spend, found that 98% now manage AI spend, up from 31% two years earlier. Their number one challenge was not negotiating rates or forecasting growth. It was visibility: simply seeing what AI costs, where it goes, and who incurred it. One practitioner in that research put it more bluntly than any vendor would: “Is your AI providing value? No one can answer that question yet.”
Corroborating evidence is easy to find, though it deserves a careful read. In a survey of 500 finance and technology leaders commissioned by a cloud cost vendor, 79% reported AI-related cost overruns in the past twelve months, and the highest overrun rate in the dataset, 89%, belonged to organizations rating their own FinOps practice most mature. That looks like a paradox, and it is worth resisting the obvious interpretation. The likeliest explanation is not that discipline makes things worse but that mature programs detect overruns the rest never notice. Either way the conclusion holds: more tooling and more process did not produce a reliable picture of AI cost.
The discipline now has a name, AI FinOps, and a growing tool market behind it. AI spend management tools sit at three layers: what the bill says, arriving days after the money is gone; what the application says, inside code someone deliberately instrumented; and the layer budgets are actually decided at, which is who spent it and whether the output justified the cost. Below are the six categories enterprises genuinely deploy. Each solves a real slice of the problem, each has a blind spot that is structural rather than a missing feature, and the last one is ours, held to the same standard.
1. Cloud FinOps Platforms
Vantage, CloudZero, Finout, and nOps come at AI spend from the cloud bill. They ingest AWS, GCP, and Azure billing exports and, increasingly, native usage data from OpenAI, Anthropic, Databricks, and Cursor, then apply allocation rules mapping spend onto teams, models, or customers. Several now offer token-level granularity rather than a monthly invoice line, with anomaly detection that flags a spike before the billing cycle closes. For anyone already running a mature cloud cost practice, extending it to AI is the obvious move.
The structural limit is what they attribute spend to. They allocate against accounts, API keys, and tags, never against people, so attribution quality depends entirely on tagging discipline established before the money was spent. In practice that discipline decays: a shared service key, an untagged project, or a developer who provisioned access outside the standard path all land in an unallocated bucket. Because these platforms read billing records rather than the interactions underneath them, they can tell you a team is expensive but never why. That matters when the fix for one costly team is coaching and the fix for another is a smaller license count.
2. LLM Observability Platforms
Langfuse, Helicone, and Datadog’s LLM observability product start at the application instead. They trace individual calls, record prompts and responses, and attach token counts and cost to each one, producing genuinely precise per-model analytics and, in some cases, budget alerts. Helicone routes through a proxy, so setup is close to a base URL change; Langfuse and Datadog rely on SDK instrumentation inside the application. Engineering teams like these tools because the output is specific enough to act on, in that an expensive prompt template shows up as an expensive prompt template.
The blind spot follows from how the data arrives. These platforms see exactly what someone remembered to instrument and nothing else. A one-off script, an analyst’s notebook, a vendor tool calling a provider API directly, or a newly shipped service whose owner skipped the wiring are all invisible, and invisible in the worst way, producing no error and no gap in any dashboard. They are engineering-owned by design, so finance has no native path to the data. None cover spend that never touches instrumented code, which in most companies means ChatGPT Enterprise seats, Copilot licenses, and the tools employees adopt on their own. That last category is large enough that we treat shadow AI discovery as a cost problem rather than only a security one.
3. Native Provider Dashboards
Every major provider now ships usage reporting, and it is better than it was a year ago. OpenAI breaks cost down by project, model, and API key with daily granularity and unrestricted CSV export. Anthropic‘s Console reports cost by model, date, and workspace. AWS supports per-user attribution for Bedrock by carrying the IAM principal through into Cost Explorer. These dashboards are accurate, free, and require no integration work, which makes them the honest starting point.
Two limits appear quickly. Granularity is the first: Anthropic’s reporting stops at the API key and workspace rather than the individual, with no documented endpoint for pulling that data programmatically, so a monthly CSV download becomes a manual step in an otherwise automated process. Latency is the second, and it varies in ways that matter. Google’s Vertex AI billing export to BigQuery runs 24 to 48 hours behind actual usage by design, and AWS Cost and Usage Report data carries a comparable lag, so neither can stop an overspend while it is happening. We documented how Vertex cost reconciliation actually works precisely because that delay is a property of the system rather than a bug to wait out.
The larger problem is arithmetic. Each dashboard shows one vendor, and the typical enterprise runs several. Answering “what did AI cost us last month” means opening three to seven consoles, each with its own lag window and attribution granularity, sharing no common definition of a team or a user, then reconciling them by hand into a spreadsheet that is stale before the meeting starts.
4. LLM Gateways and Proxies
Gateways are the only category here that can stop spending rather than report on it. LiteLLM tracks spend by key, user, team, and organization, with soft thresholds that notify before a hard limit bites. Portkey rejects requests outright once a budget is exhausted, which is the difference between learning about a runaway retry loop on Monday and never paying for it at all. Cloudflare added cost-based spend limits to its AI Gateway in June 2026, scoped by model, provider, or custom metadata. For teams building AI products, putting a gateway in front of provider APIs is among the higher-leverage decisions available.
A gateway governs the traffic that passes through it, and only that traffic. The moment a developer calls a provider SDK directly with a key they hold, or a business user signs up with a corporate card, or a SaaS vendor bills AI features through an existing contract, the gateway has neither visibility nor authority. This solves the problem of code you control overspending, which is real. It does not touch the harder problem, which is the AI spend you have not yet enumerated. Gateway coverage is also self-reporting in an unhelpful way, because the spend it cannot see produces no record of its own absence.
5. SaaS Management and AI Governance Platforms
This category splits into two groups that get mentioned together and should not be. SaaS management platforms such as Zylo and Productiv track AI cost as an extension of subscription management, covering seat-based tools, expensed applications, and contracted vendors. Zylo’s own 2026 index reported AI-native spend up 108% year over year, rising to 393% at organizations above 10,000 employees. Separately, AI governance platforms such as Credo AI and Holistic AI address regulatory risk under the EU AI Act, ISO 42001, and the NIST AI Risk Management Framework.
The governance platforms serve a compliance officer rather than a finance owner, and cost is not what they are built for; there are no budgets, forecasts, or spend alerts to evaluate. The SaaS platforms do real cost work but see procurement rather than consumption. Spend arriving through a purchase order is well covered, while a developer’s direct provider API key, or Vertex and Bedrock consumption buried inside a cloud bill, sits outside what a subscription-oriented system is designed to ingest. They answer what the company subscribes to. They do not answer what it consumed.
6. Olakai
We built AI Spend Governance for the attribution layer the other five leave open, so it is only fair to describe it the same way. Budgets come in four groups: a program ceiling across every developer, provider, and project; per-provider caps that limit one vendor without touching others; employee-centric lenses attributing spend to an individual, a persona, or a department; and projects grouping shared service keys into named cost centers. They are deliberately overlapping lenses rather than a partition, so the same dollar can count toward a developer’s budget, their department’s, the provider’s, and the program ceiling at once. Adding every budget on the page will not reconcile to program spend, and is not meant to.
What makes that arithmetic decision-grade is the denominator beside it. Cost and value live in the same system across both products: Olakai Agentic expresses coding tool return as AI Equivalent Engineers against a fully loaded engineer cost, and agent return as value created over execution cost, while Olakai Assistive converts estimated time saved into dollars against subscription spend. Those definitions are configurable through custom KPIs rather than fixed, so finance holds the assumptions rather than inheriting ours.
The structural limit is enforcement timing. Olakai is not a gateway, so it cannot reject a call mid-overspend the way Portkey can. Budgets are evaluated once daily after provider data lands, which makes this an early-warning system rather than a hard stop. They also track per-token providers only, so GitHub Copilot is excluded outright, its pricing being seat-based with nothing for a per-token mechanism to measure. Spend that cannot be attributed to a person still counts toward program and provider totals but drops out of the per-entity lenses, which is why those totals can legitimately come in below the program figure. The month-end projection is a run-rate extrapolation, not a model of growth. Teams needing request-level blocking should run a gateway underneath this, not instead of it.
The Pattern: Cost Without a Denominator

Line the blind spots up and they describe one shape. FinOps platforms see accounts rather than people. Observability platforms see instrumented code rather than actual usage. Provider dashboards see one vendor at a time. Gateways see routed traffic only. SaaS and governance tools see contracts rather than consumption. Each fragments along a different axis, and no single category was designed to cover assistive tools, coding agents, and autonomous agents together, attributed to a person, in one place.
That fragmentation is the real finding behind the FinOps Foundation’s visibility result. It is not that these teams lack tools. It is that every tool they own reports a different slice, and slices do not add up.
The deeper issue is why a better cost dashboard does not fix it. Every tool here answers how much. None answers how much relative to what it produced. A team spending $40,000 a month on coding agents is either the best or the worst investment in the company, and the number alone cannot tell you which, which is why spend control and AI ROI measurement are one problem rather than two adjacent ones. Against Gartner’s forecast of $2.67 trillion in worldwide AI spending in 2026, a 49.5% increase with another 36.2% projected for 2027, the cost of answering “how much” without “compared to what” compounds every quarter.
Uber is the version of this that made the news, having encouraged AI coding tool usage without a ceiling and then discovering four months in that the annual budget was already gone. The guardrail it built afterward is real infrastructure. It simply arrived after the crash, which is the outcome every tool in this roundup is trying to prevent and none of them prevents alone.
If you are assembling AI spend management from three of these categories and a spreadsheet, it is worth seeing the consolidated version. Talk to an expert and we will walk through your actual providers, your actual attribution gaps, and what a single view of spend and value would show a CFO reviewing AI spend at your company.
