What is prompt caching, and how much can it actually cut an enterprise AI bill?

Prompt caching applies to the input a model reads again and again: long system prompts, the same context re-read every turn, retrieval, and tool loops. It takes up to 90% off the input you reuse, and Cursor's own data says that without smart caching its cost would be roughly ten times higher.

How Prompt Caching Works in Agentic Workloads

Cached input is what agents run on: long system prompts, the same context re-read every turn, retrieval, and tool loops. That makes it the one line item agentic workloads consume most of, so an agent budget built on a cheap cache hit changes whenever a provider reprices it.

Source: The Cache Tax: Where DeepSeek’s Price Increase Is Concentrated

The Magnitude of Bill Reductions from Caching

In an AI coding bill, the money goes on having the model read your codebase and documentation over and over, on every turn. Output is 0.6% of usage once cache reads are counted, and Cursor's own data says that without smart caching the cost would be roughly ten times higher.

Source: The AI Bill Is Eating Everything Else

Pricing Reductions on Reused Input

On top of choosing the right model, two more levers apply: batch processing takes 50% off, and prompt caching takes up to 90% off the input you reuse.

Source: How to Be a Smarter Token Manager: Model Routing, Explained

Also asked as

  • How much money can prompt caching save on enterprise LLM token costs?
  • What is cached input in agentic AI workloads and how does it reduce bills?
  • How does prompt caching affect token spend for coding tools and AI agents?

Related questions

← All answers