AI infrastructure · Anthropic API

Claude prompt caching: stable prefixes that actually hit

A production guide to Claude prompt caching: prefix design, breakpoints, TTLs, invalidation, diagnostics and the metrics that prove a real cache hit.

TOPIC HUBAI, RAG & Vector Search
Conceptual illustration of stable prompt-prefix blocks entering a cache plane while a changing request tail continues to Claude.
An editorial interpretation of the topic, followed by a practical execution diagram.

Claude prompt caching reduces repeated input processing only when consecutive requests share an identical, sufficiently long prefix. The production pattern is simple in principle: order the request as tools, system instructions and stable context; place the cache breakpoint at the last byte-stable block; keep timestamps, user input and request-specific data after it; then prove the result with cache-read tokens and diagnostics. It is best for long, reused instructions or documents—not short prompts, low-reuse traffic or content that changes every request.

Model the request as a stable prefix and a dynamic tail

Anthropic's prompt caching documentation states that the cached prefix follows the request order: tools, system, then messages. A breakpoint hashes everything from the beginning through the marked block. A change anywhere before that point creates a different prefix. The right design question is therefore not “which paragraph is expensive?” but “which ordered prefix stays exactly the same across requests that should share work?”

For a support assistant, the stable prefix might contain deterministic tool schemas, the operating policy, product documentation and approved examples. The dynamic tail contains tenant-specific authorization, the latest customer message and fresh retrieval results. For a document-analysis workflow, the document can be the stable prefix and each question the tail. For a coding agent, the stable portion may be the tool catalog and repository instructions, while the changing task and tool results remain later.

Byte stability matters. A timestamp interpolated into the system prompt, random JSON key ordering, a reordered tool list or a regenerated description invalidates the prefix even when the meaning appears unchanged. Serialize tool schemas deterministically, version stable context explicitly, and keep request IDs, clocks and tracing metadata outside the prompt unless the model truly needs them. A semantic cache is a different architecture; Anthropic prompt caching reuses an identical prefix rather than searching for similar meaning.

Choose automatic caching or explicit breakpoints deliberately

Automatic caching adds top-level cache_control and moves its breakpoint with a growing conversation. It is a strong default for append-only chat where each request replays an unchanged history and adds a small number of blocks. Anthropic documents that the service checks backward for a previously written prefix within a 20-block lookback window. Automatic mode therefore works well when new turns are appended without rewriting earlier content.

Explicit breakpoints fit requests with a stable prefix followed by a variable suffix. Put cache_control on the last stable block, not on a block containing the current time, user query or fresh retrieval result. Anthropic allows up to four breakpoints. Multiple points can separate tool definitions that change rarely from a knowledge bundle that changes daily, or preserve a reachable write point when a long conversation grows beyond the lookback window. Breakpoints themselves do not add token charges; written, read and uncached tokens determine the bill.

A common anti-pattern is placing one breakpoint at the end of the entire request while the final block changes on every call. The service writes an entry for that exact prefix, but the next call cannot match it. Looking backward discovers only entries that earlier calls actually wrote; it does not invent a cache entry at an unmarked stable boundary. Place an explicit breakpoint before the dynamic data instead.

Build one deterministic request contract

The following TypeScript sketch keeps the reusable policy and reference bundle before the breakpoint. The tenant authorization and question remain after it. Treat the model ID, tool order, schemas and system blocks as versioned configuration. Replace the illustrative strings with validated data and enforce authorization before constructing the request.

const stableSystem = [
  { type: 'text', text: POLICY_V3 },
  {
    type: 'text',
    text: PRODUCT_REFERENCE_2026_10,
    cache_control: { type: 'ephemeral', ttl: '1h' }
  }
];

const response = await anthropic.messages.create({
  model: 'claude-sonnet-5-5',
  max_tokens: 1200,
  tools: stableToolsSortedByName,
  system: stableSystem,
  messages: [{
    role: 'user',
    content: [
      { type: 'text', text: `Tenant scope: ${authorizedScope}` },
      { type: 'text', text: userQuestion }
    ]
  }]
});

observe({
  input: response.usage.input_tokens,
  cacheWrite: response.usage.cache_creation_input_tokens,
  cacheRead: response.usage.cache_read_input_tokens
});

The one-hour TTL is not automatically better. Use it when the same large prefix is likely to be reused after gaps longer than five minutes and the additional write price is justified. The default five-minute entry refreshes when read, so a busy workload can keep it warm. Measure the inter-arrival distribution for each prefix version instead of choosing a TTL from intuition. When mixing TTLs, follow Anthropic's ordering rules and keep longer-lived content before shorter-lived content.

The cache key covers an ordered prefix: tools, system instructions and stable context. Dynamic request data stays after the breakpoint.
The cache key covers an ordered prefix: tools, system instructions and stable context. Dynamic request data stays after the breakpoint. Open for a larger view

Calculate break-even with traffic, not a headline percentage

The official pricing page and caching guide currently describe five-minute writes at 1.25 times base input price, one-hour writes at 2 times, and standard cache reads at 0.1 times, with documented model-specific exceptions. Those multipliers can change, so calculate with the active model's current price rather than embedding a permanent budget assumption. Output tokens and the dynamic uncached tail are still billed normally.

For a prefix containing P tokens, compare a no-cache baseline with one write plus expected reads during the TTL. With standard multipliers, a five-minute cache becomes cheaper after reuse because the first write costs more but each hit costs far less. A one-hour write needs more reuse to repay its higher creation cost. The meaningful unit is not requests per day; it is requests per identical prefix version inside the TTL window. Ten tenants with ten different policy prefixes may produce ten cold caches even if aggregate traffic is high.

Use the token counting endpoint before sending the request to estimate tools, documents and messages. Cache thresholds depend on model. Anthropic's current documentation lists 512 tokens for several current models, 1,024 or more for others, and explains that an undersized marked prefix can be processed without caching and without an error. Check the active model's current minimum and verify usage fields rather than assuming cache_control guarantees a write.

Design invalidation as a hierarchy

The prefix hierarchy creates an invalidation hierarchy. Changing tools invalidates the tool prefix and everything after it. Changing the system portion preserves a valid tools prefix but invalidates system and message caches. Changes in message content affect the message layer. Model changes also prevent reuse because cache identity is model-specific. This makes request construction an architecture concern, not a decorator added around an API call.

Keep tool definitions stable across ordinary requests. Do not inject tenant-specific permissions into tool descriptions; authorize tool execution in application code and place only necessary scope data in the dynamic tail. Separate knowledge releases by content hash, such as product-reference:v42. Warm a new prefix with one request, wait until that response begins because the entry is not available earlier, then allow parallel traffic to use it. Retain the old version during a controlled rollout so a rollback does not force every caller onto an untested prefix.

For multi-tenant systems, caching is not an authorization boundary. Never broaden retrieved data or tool permissions to improve hit rate. Reuse only content that is safe and identical for the sharing scope. Place tenant-private context after a public product prefix, or create tenant-specific stable prefixes and include the tenant identity in your own observability dimensions. Security and isolation outrank a cache ratio.

Observe hits, misses and prefix versions

A successful first request normally reports cache_creation_input_tokens. A later hit reports cache_read_input_tokens. Track both alongside regular input tokens, model, prefix version, TTL class and request route. Useful service indicators include cache-read tokens divided by eligible prefix tokens, creation tokens per prefix version, latency distributions for hits versus misses, and cost per completed product operation. Do not treat a high hit ratio as success if stale context or incorrect tenant scope harms the answer.

Anthropic's cache diagnostics can compare consecutive requests when the feature is enabled. Pass the previous response ID; the returned diagnostic can identify the first divergence category, such as model_changed, system_changed, tools_changed or messages_changed. The documentation notes that fingerprints contain hashes and token-count estimates rather than raw prompt text, are scoped to the organization and workspace, and are retained for a limited period. Diagnostics are currently Claude API only, so build basic prefix-version telemetry even when using another platform.

Interpret diagnostics and usage together. No divergence plus zero cache-read tokens can mean the entry expired. A reported system change with zero reads points to request instability. High cache reads plus a late messages change can mean an earlier breakpoint still hit. Preserve a redacted canonical representation or hash of each request layer in your own logs, never raw secrets, so a deployment that changes serialization can be correlated with sudden misses.

Know when prompt caching is the wrong tool

Use prompt caching for long shared policies, large repeated documents, stable tool catalogs, recurring examples and append-only conversations. It can lower model input work; it does not improve retrieval relevance, factual accuracy or authorization. If each request needs different evidence, build a retrieval pipeline. If the answer is deterministic structured data, query SQL or an API. If a short prompt rarely repeats, the added write and operational complexity may not repay itself.

Do not cache stale business facts merely because they are expensive to send. Define an owner and refresh policy for every stable context bundle. Do not move sensitive data into a wider shared prefix to increase reuse. Do not pad prompts with irrelevant text merely to cross a model threshold. Do not test only sequential happy paths: the first parallel wave can all miss because a new entry becomes available only after the first response begins.

Prompt caching also does not replace RAG evaluation or durable agent execution. It composes with them: cache stable instructions and schemas, retrieve fresh authorized evidence into the tail, then evaluate the grounded answer. For model migration decisions, pair the cache contract with the Claude Sonnet 5.5 rollout guide because a model change creates a different cache identity.

Roll out with an explicit acceptance test

Start with one high-volume route and identify the exact stable prefix. Make its serialization deterministic, measure tokens, select a TTL from real inter-arrival times and add a named prefix version. Send one warm-up request, then verify a later request reports cache-read tokens. Change one layer at a time—user query, message suffix, system block and tool schema—to prove which edits preserve or invalidate earlier prefixes.

Before production, test expiry, a cold concurrent burst, a content release and a model fallback. Set alerts on unexpected creation-token growth and read-token collapse, but connect them to product latency and cost rather than an arbitrary universal threshold. The operating goal is a controlled reusable prefix with observable economics and safe invalidation. If the team cannot say what is cached, who may share it, when it changes and how a hit is proven, the cache is not production-ready. For help implementing this as part of an AI product, see AI integrations and retrieval systems.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceAI integrations & retrieval systemsRelevant projectAI Action Studio