Skip to content
Arrow left icon View all blogs 10 min read

How to Cut LLM Costs with Prompt Caching

Arrow down icon
How to Cut LLM Costs with Prompt Caching
How to Cut LLM Costs with Prompt Caching

TL;DR

  • Prompt caching reduces LLM costs by reusing stable prompt prefixes instead of processing the same context from scratch on every request.
  • The article shows what to cache, system instructions, tool definitions, reference documents, examples, and reusable context, and what should stay dynamic, such as the current user query and live tool results.
  • For agentic workflows, the savings can compound because the same system instructions and tool definitions are repeatedly sent across multiple steps. But caching only works when the prompt is structured correctly: dynamic content inside the stable prefix, short prefixes, expired TTLs, and frequent prompt changes can all destroy cache hits. 

Every request to an LLM API reprocesses your entire prompt from scratch. That prompt includes system instructions, tool definitions, conversation history, reference documents, and policy rules. A lot of this is identical across multiple requests of even sessions. Without caching, you pay for the same tokens again and again.

Prompt caching removes that waste. When a repeated prefix arrives, the provider skips recomputing it again from scratch. You are charged a fraction of the standard input rate.

But just enabling the cache doesn’t instantly get you the benefits. A misconfigured prompt structure can produce near-zero cache hits even when caching is technically active. A workload with short or highly variable prompts may save almost nothing. And ignoring cache-write costs can erode savings you didn’t account for.

The real opportunity in prompt caching comes from understanding when it pays off, how to structure prompts around stable prefixes, and how to measure how much you save from it.

What is Prompt Caching?

When a prompt arrives at an LLM API, the model runs a prefill stage: it processes every token in the input and computes Key (K) and Value (V) vectors at each transformer layer. These K-V pairs use the attention mechanism to generate responses. Computing them is expensive and is where most input processing time is spent.

Prompt caching saves those K-V computations across separate API calls. The first time a prompt prefix arrives, the provider computes and stores the K-V tensors. The second time the same prefix appears, it loads the stored tensors and skips the recomputation, charging a reduced rate for those tokens.

This article focuses on prefix caching. It’s safe and natively supported by all major providers. Two metrics matter for every cached request: 

  • Cache writes (tokens stored on first use, which cost slightly more than standard input) 
  • Cache reads (tokens retrieved from cache, which cost significantly less)

Prompt Caching vs. Semantic Caching

Prompt caching and semantic caching reduce LLM costs in different ways. Prompt caching reuses an exact, repeated portion of a prompt, while semantic caching can reuse a previous response for a sufficiently similar query. The distinction matters when deciding which parts of an AI workload are worth caching.

Prompt CachingSemantic Caching
How it matchesExact prompt prefixSemantic similarity
What it storesReusable prompt contextPrevious query-response pairs
Best forLong, stable context reused across requestsRepeated questions with different wording
SavingsReduces the cost of processing repeated inputCan eliminate the model call entirely
Main riskLow: a mismatch simply results in a cache missHigher: stale or incorrect responses can be returned
FreshnessUsually provider-managed TTLRequires application-level invalidation

How Much Can Prompt Caching Save?

The economics depend on five variables: what percentage of your prompt is stable, how often requests reuse that prefix, what discount the provider applies to cache reads, what premium the provider charges for cache writes, and what your output token costs are.

If your prompt is 1,000 tokens and 200 of them are dynamic, you can only cache 800 tokens per request. If you’re on Anthropic Claude and achieve a 90% cache hit rate on those 800 tokens, you pay 0.1x for 720 tokens per request and 1.25x for 80 tokens that get written on cache misses. Output tokens are unaffected entirely.

This means the absolute savings depend heavily on the size of the stable prefix. A 90% hit rate on a 500-token system prompt saves a modest amount. A 70% hit rate on a 20,000-token enterprise system prompt with tool definitions and reference documents can save tens of thousands of dollars per month.

When Does Prompt Caching Pay Off?

WorkloadCaching valueWhy
Long system prompts (2,000+ tokens)HighLarge stable prefix repeated every request
Tool-heavy agentsHighTool definitions rarely change, often thousands of tokens
RAG applications with repeated reference docsHighSame documents retrieved for similar queries
Multi-turn chat applicationsHighEarlier turns are stable, only new turn varies
Multi-step agent loopsHighSystem instructions repeat across every reasoning step
Few-shot prompting with fixed examplesHighExamples are static, only query changes
Enterprise policy or compliance documentsHighLarge, stable reference material
Short prompts under the minimum thresholdLowProvider threshold not met, no cache activated
Highly dynamic prompts with little shared contentLowToo little stable prefix to cache
Low-frequency workloads with long gaps between requestsLowCache likely expires before next reuse
One-off requestsNoneCache write is never read back

The provider minimum token threshold is worth checking before you assume caching will apply. If your cached prefix is shorter than the provider’s floor, requests still succeed but no caching occurs.

What Should You Cache?

The core principle is simple: stable content belongs in the cached prefix. Dynamic content belongs after it.

Cache these:

  • System instructions and agent personas
  • Tool definitions and API schemas
  • Static few-shot examples
  • Reference documents, policy documents, compliance rules
  • Reusable domain context
  • Fixed knowledge base chunks

Keep these outside the cached prefix:

  • The current user message or query
  • Timestamps, request IDs, session identifiers
  • Live tool results and real-time data
  • User-specific context injected per request
  • Anything that varies between requests

The most common mistake is placing a dynamic value somewhere inside what should be a stable prefix. A timestamp or session ID sitting four paragraphs into an otherwise static system prompt breaks the cache for every token that follows it.

The principle depends on the provider. Anthropic uses explicit cache_control markers on content blocks, OpenAI applies caching automatically based on prefix hashing (with explicit mode available on GPT-5.6 and newer), and Google offers both implicit and explicit cache objects.

Prompt Caching in Agentic AI Workflows

Agents repeat everything on every step of every loop.

Consider what happens in a multi-step agent: the model calls a tool, receives a result, adds it to the conversation, and calls the model again with the full accumulated context. Step 2 re-sends everything from step 1 plus new tool results. Step 3 re-sends everything from step 1 and 2, plus more. The stable portions (system prompt, tool definitions) get reprocessed on every single step, even though they haven’t changed.

A research agent running eight steps with a 5,000-token system prompt reprocesses those same 5,000 tokens eight times. With caching, you pay for them once and read them at 90% discount on the remaining seven steps. Across thousands of agent sessions per day, the compounding effect is substantial.

The key is structuring prompts according to a stability hierarchy:

Layer 1: The core instructions and personas almost never change between requests. Cache aggressively.

Layer 2: Tool definitions and API schemas changes only on deployments. Cache.

Layer 3: Conversation and task history grows over time, but older turns stabilize. Cache the earlier portion of history conditionally. Put the cache boundary after the second-to-last turn and leave the most recent turn uncached.

Layer 4: Current user message and live tool results changes on every step. Never cache.

Multi-agent architectures get an additional benefit. When an orchestrator spawns multiple sub-agents that share the same base system prompt, the first sub-agent’s request writes the cache. 

Every subsequent parallel sub-agent reads from it. A shared foundation of 4,000 tokens across twenty parallel sub-agents means you pay the write premium once and the read discount nineteen times.

Fixed few-shot examples also benefit. A set of ten worked examples at 3,000 tokens, reused across every request in a high-volume text-to-SQL or classification service, is cached after the first call and read cheaply on every subsequent one.

Why Your Prompt Cache Isn’t Hitting

A team might enable caching, check that the API field is present, and assume it’s working. Then they look at their bill and nothing changes. Here are the causes and fixes:

1. Dynamic content inside the stable prefix

A timestamp, request ID, or user name anywhere before the cache boundary invalidates the cache from that point down. Everything after it misses. 

Audit your prompt template character by character. Move every dynamic value below the cache boundary.

2. Stable prefix is below the minimum threshold

If your cacheable content is under the provider’s minimum token count for the model you’re using, no caching occurs. No error is raised. 

Verify the threshold for your specific model, consolidate static content, or expand the cached prefix.

3. Requests arrive too infrequently for the cache to stay warm

Anthropic’s default TTL is 5 minutes, refreshing on each read. If requests arrive every 10 minutes, the cache expires between them. 

Match workload scheduling to cache lifetime, or use the 1-hour TTL option for workloads with longer gaps. OpenAI offers up to 24 hours of extended retention on supported models.

4. Seemingly stable content is actually varying

JSON serialized in a different key order between requests, trailing whitespace differences, or Unicode normalization issues can break the hash even when the content looks identical to a human reviewer. 

Normalize your prompt templates, canonicalize JSON key ordering, and strip unexpected whitespace before sending.

5. Prompt template changes invalidated the cache

Modifying the system prompt, even a single word, invalidates any existing cache. 

Treat your cached prompt prefix like a versioned artifact. Don’t change it casually, and plan for a cold-cache period when you do.

6. Model version changes

Caches are model-specific. A cache written for one model version cannot be read by another. 

After any model update, expect a cold cache period and plan capacity accordingly.

To diagnose the process, log the cache hit and miss counts that every provider exposes in the response metadata. On Anthropic, check cache_read_input_tokens and cache_creation_input_tokens. On OpenAI, check cached_tokens. When hit rate drops unexpectedly, diff your prompt structure between a hit and a miss to locate the change.

How to Implement Prompt Caching

Here is a step-by-step roadmap enterprises can follow to implement prompt caching for their enterprise systems.

Step 1: Identify your repeated context

List everything that appears in every request: system instructions, tool definitions, reference documents, examples. Measure token counts.

Step 2: Separate stable from dynamic content

Sort your prompt elements by how often they change. Anything that varies per user, per session, or per request is dynamic.

Step 3: Structure the prompt with stable content first

System instructions first, then tool definitions, then any reference material, then the cache boundary, then dynamic content, then the user message.

Step 4: Configure provider-specific cache controls

On Anthropic, add cache_control: {“type”: “ephemeral”} to the content blocks you want cached. On OpenAI, caching is automatic for eligible models; use prompt_cache_key if you need routing hints for high-volume workloads. On Gemini, choose implicit (automatic) or explicit (create a cache object with an ID). On Bedrock with Claude, use cache checkpoints at the end of cacheable sections.

A minimal Python example for Anthropic:

import anthropic
client = anthropic.Anthropic()
SYSTEM_PROMPT = """
[Your static instructions, tool definitions, and reference content here]
"""
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system=[
{
"type": "text",
"text": SYSTEM_PROMPT,
"cache_control": {"type": "ephemeral"}
}
],
messages=[
{"role": "user", "content": user_message} # dynamic, never cached
]
)
# Check cache performance
usage = response.usage
print(f"Cache write tokens: {usage.cache_creation_input_tokens}")
print(f"Cache read tokens: {usage.cache_read_input_tokens}")
print(f"Fresh tokens: {usage.input_tokens}")

Step 5: Send repeated requests

The cache is cold until the first request. In high-concurrency systems, consider a deliberate warm-up request before opening full traffic.

Step 6: Measure and iterate

Track hit rate, cost per request, and total input cost. Compare before and after.

Prompt Caching Across OpenAI, Anthropic, Bedrock, and Vertex

The underlying mechanics are similar across providers. The implementation details and pricing differ meaningfully. All figures below reflect current documentation as of mid-2026; verify against the provider’s pricing page before budgeting.

ProviderCaching modeCached input pricingCache writesMinimum cacheable prefixTTL
Anthropic ClaudeAutomatic or explicit0.1× input rate1.25× (5 min) or 2× (1 hour)512–4,096, model-dependent5 min default; 1 hour available
OpenAIAutomatic on GPT-5.5 and earlier; automatic/explicit on GPT-5.6+0.1× on GPT-5.6+; model-dependent elsewhereNo additional write fee on GPT-5.5 and earlier; 1.25× on GPT-5.6+1,024+30 min minimum on GPT-5.6+; model-dependent
Google GeminiImplicit or explicitUp to 0.1×, model-dependentExplicit caching includes token storage charges2,048–4,096, model-dependentProvider-managed for implicit; 1 hour default for explicit
AWS BedrockImplicit or explicit, model/API-dependentUp to 0.1×, model-dependentModel-dependent512–4,096, model-dependent5 min default on many models; 1 hour on supported models

Also note that Anthropic’s automatic caching mode, available on the direct API and some managed platforms, is not available on AWS Bedrock. Bedrock’s checkpoint system achieves similar results but uses a different mechanism.

How to Measure Whether Prompt Caching Is Actually Saving Money

Cache hit rate’s main purpose is cost reduction for your business. Track these metrics in production:

  • Total input tokens per request
  • Cache-write tokens per request
  • Cache-read tokens per request
  • Hit rate (cache-read requests / total eligible requests)
  • Average cost per request
  • Total daily input cost, before and after

A simple cost comparison calculation: take your total cache-eligible input tokens per day, multiply by the standard input price, compare to your actual daily bill accounting for write and read pricing at current provider rates. The gap is your realized savings.

A sudden drop in hit rate usually signals either a prompt structure change or a dynamic value that crept into the stable prefix. Log both cache writes and reads on every request, and set an alert if the hit rate drops below your baseline by more than a few percentage points.

Common Prompt Caching Mistakes

Some of the common prompt caching mistakes can be having dynamic content inside cache, or not measuring cache’s economic value properly. Another reason could be ignoring provider and model level adjustments.

1. Putting dynamic content inside the cached prefix

Timestamps, session IDs, user data, and live tool results can break cache reuse when they appear before the cache boundary. Keep stable instructions, tools, and reusable context first, then append request-specific content.

2. Assuming caching is working without measuring it

Enabling caching does not guarantee cache hits. Monitor cached-token usage and hit rates in production. A workload that performs well in testing can degrade as prompts, traffic patterns, or routing change.

3. Ignoring cache economics and TTLs

Cache writes can cost more than standard input tokens, while cached reads are cheaper. If requests are too far apart for the cache lifetime, you may repeatedly pay write costs without enough discounted reads to offset them. 

4. Ignoring provider and model differences

Minimum cacheable length, TTL, cache controls, and pricing vary across providers and models. Don’t assume an implementation or threshold that works for one model will work for another.

5. Changing cached prompts without accounting for cold starts

Changes to a cached prefix can invalidate existing cache entries and temporarily increase costs while the cache rebuilds. Treat major prompt-template changes like other production changes and monitor their cost impact.

Save LLM Token Costs with Prompt Caching

The problem prompt caching solves is simple: teams pay full price to reprocess context that hasn’t changed since the last request. System prompts, few-shot examples, long reference documents, and tool definitions get sent and reprocessed on every call, even when nothing in them has moved.

Getting caching right takes a stable, well-ordered prompt structure, visibility into actual cache performance in production, and a realistic read on how your traffic patterns interact with provider TTLs. None of that is complicated, but all of it requires attention past the initial setup.

FAQs

What cache hit rate should I target?

It depends on how much of your prompt is static and how often that static portion changes. A large, stable prefix (long system instructions, reference documents) paired with only the user’s input as the variable part will naturally produce a high hit rate. If your prompt changes substantially on nearly every request, a low hit rate is expected.

Does prompt caching change output quality?

No. Caching is purely an infrastructure optimization for cost and latency. The cached tokens are the same tokens, processed the same way, not a summarized or compressed version of the prompt. If output quality shifts after enabling caching, look for another cause rather than the caching itself.

Why did my costs go up after enabling caching?

Usually one of a few things: the prefix changes too often to accumulate cheap reads before the next write, the prompt is too short to qualify for caching at all, gaps between requests exceed the TTL so the cache keeps expiring before reuse, or the cached portion is small relative to the full prompt so the discount barely moves the total. In most cases where costs rise, the cache is being written repeatedly without being read.