Skip to content

Guides

Prompt caching

Reuse a long prompt prefix at the cache read price.

When the start of a prompt repeats across requests, providers can serve it from cache. Cached input is read at a lower price than fresh input.

Check the price

A model supports cache pricing when pricing.input_cache_read in GET /v1/models is above "0" and lower than pricing.prompt. A "0" in input_cache_read or input_cache_write means no cache price is published for the model. Cache writes can cost more than fresh input, so caching pays off when a prefix is reused.

Mark a cache point on Messages

On POST /v1/messages, add cache_control to the last block of the prefix you want cached. Deference forwards system and cache_control untouched.

{
  "model": "anthropic/claude-sonnet-5.5",
  "max_tokens": 1024,
  "system": [
    {
      "type": "text",
      "text": "You are a support agent for Acme. Policies follow...",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [{ "role": "user", "content": "How do refunds work?" }]
}

Keep the cached prefix identical between requests. Changing one character before the cache point is a miss.

Chat Completions and Responses

Some providers cache automatically once a prefix is long enough. The usage block reports it as prompt_tokens_details.cached_tokens on chat completions.

Confirm it works

Read usage.cache_read_input_tokens on Messages or usage.prompt_tokens_details.cached_tokens on Chat completions. Non-zero values confirm a cache read. Activity separates cached input and cache writes, so compare the final cost as well as the token counts. A cache write can increase the first request's cost.

Why a cache misses

CauseFix
The prefix changedMove variable content after the cache point
The prefix is below the provider's minimum lengthCache a longer prefix
The cache expired between requestsSend requests closer together
The system field was flattened to a stringSend it as an array of blocks