Guides
Prompt caching
Reuse a long prompt prefix at the cache read price.
When the start of a prompt repeats across requests, providers can serve it from cache. Cached input is read at a lower price than fresh input.
Check the price
A model supports cache pricing when pricing.input_cache_read in GET /v1/models is above "0" and lower than pricing.prompt. A "0" in input_cache_read or input_cache_write means no cache price is published for the model. Cache writes can cost more than fresh input, so caching pays off when a prefix is reused.
Mark a cache point on Messages
On POST /v1/messages, add cache_control to the last block of the prefix you want cached. Deference forwards system and cache_control untouched.
{
"model": "anthropic/claude-sonnet-5.5",
"max_tokens": 1024,
"system": [
{
"type": "text",
"text": "You are a support agent for Acme. Policies follow...",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [{ "role": "user", "content": "How do refunds work?" }]
}Keep the cached prefix identical between requests. Changing one character before the cache point is a miss.
Chat Completions and Responses
Some providers cache automatically once a prefix is long enough. The usage block reports it as prompt_tokens_details.cached_tokens on chat completions.
Confirm it works
Read usage.cache_read_input_tokens on Messages or usage.prompt_tokens_details.cached_tokens on Chat completions. Non-zero values confirm a cache read. Activity separates cached input and cache writes, so compare the final cost as well as the token counts. A cache write can increase the first request's cost.
Why a cache misses
| Cause | Fix |
|---|---|
| The prefix changed | Move variable content after the cache point |
| The prefix is below the provider's minimum length | Cache a longer prefix |
| The cache expired between requests | Send requests closer together |
The system field was flattened to a string | Send it as an array of blocks |