Endpoints
Messages
Create a message with the Anthropic Messages API. This is the endpoint Claude Code uses.
Follows Anthropic's Messages format. The request and the response pass through unchanged. Use https://deference.si as the SDK base URL, without /v1.
Example
curl https://deference.si/v1/messages \
-H "x-api-key: $DEFERENCE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "Say hello in one sentence." }]
}'import os
import anthropic
client = anthropic.Anthropic(base_url="https://deference.si", api_key=os.environ["DEFERENCE_API_KEY"])
message = client.messages.create(
model="anthropic/claude-sonnet-5.5",
max_tokens=1024,
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(message.content[0].text)import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://deference.si",
apiKey: process.env.DEFERENCE_API_KEY,
});
const message = await client.messages.create({
model: "anthropic/claude-sonnet-5.5",
max_tokens: 1024,
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(message.content);Headers
- x-api-keystringRequired
- Your key.
Authorization: Bearer sk-df-...works too. - anthropic-versionstring
- Forwarded as sent. Anthropic's SDKs set it.
- anthropic-betastring
- Forwarded as sent. Any beta value is accepted.
Request body
- modelstringRequired
- The model id, such as
anthropic/claude-sonnet-5.5. - max_tokensintegerRequired
- The output limit. Deference holds credit for all of it while the request runs, so a high value needs more available credit.
- messagesarrayRequired
- Alternating
userandassistantturns. Content can be a string or an array of blocks. - systemstring | array
- The system prompt. Send blocks to use
cache_control. - streamboolean
- Return server-sent events.
- toolsarray
- Tools the model may call.
- tool_choiceobject
- How the model picks a tool.
- temperaturenumber
- Sampling temperature.
- thinkingobject
- Extended thinking settings, for models that support them.
A query string such as ?beta=true is ignored.
Credit held
Deference reserves credit for the prompt, the full max_tokens and allowed web-search steps. It never lowers max_tokens on Messages. A web_search tool's max_uses is capped to the request's tool-step allowance, which defaults to 30 and can be lowered with max_tool_calls. Omitted or invalid max_uses uses that allowance.
A request whose hold does not fit in your available credit, or in its key's remaining limit, returns 402 with billing_error: lower max_tokens or add credit. A request that free credit pays for runs on what it can hold. See Pricing.
Files and URLs
- A model reads a PDF or other file itself only if it lists
fileinput. For other models, the request returns400 unsupported_parameterunless it asks for the free parser:"plugins": [{ "id": "file-parser", "pdf": { "engine": "cloudflare-ai" } }]. - An image or document given as a URL is held at the model's whole context window, because the provider fetches it.
- A call that gets no answer within 15 minutes returns
504 upstream_timeout, and its hold is charged because the provider may still have run it. Stream long requests.
Response
{
"id": "msg_01Ab3xQz9k",
"type": "message",
"role": "assistant",
"model": "anthropic/claude-sonnet-5.5",
"content": [{ "type": "text", "text": "Hello, nice to meet you." }],
"stop_reason": "end_turn",
"usage": { "input_tokens": 14, "output_tokens": 9 }
}usage also reports cache_creation_input_tokens and cache_read_input_tokens when caching applies.
Errors
Errors use Anthropic's envelope with type: "error" and a request_id. See Errors.