Endpoints
Chat completions
Create a model response from a list of messages, with optional streaming, images and tools.
Follows OpenAI's Chat Completions format. Send the key as Authorization: Bearer sk-df-... or x-api-key.
Example
curl https://deference.si/v1/chat/completions \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"messages": [{ "role": "user", "content": "Say hello in one sentence." }]
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://deference.si/v1",
apiKey: process.env.DEFERENCE_API_KEY,
});
const response = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5.5",
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(response.choices[0].message.content);Request body
- modelstringRequired
- The model id, such as
anthropic/claude-sonnet-5.5. - messagesarrayRequired
- The conversation. Each item has a
role(system,user,assistantortool) andcontent. Content can be a string or an array of parts, includingimage_urlparts for vision. - streamboolean
- Return server-sent events. Defaults to
false. - stream_optionsobject
{ "include_usage": true }makes the final usage chunk follow OpenAI's shape, with an emptychoicesarray.- max_tokensinteger
- The output limit. Also sets the credit held while the request runs.
max_completion_tokensworks the same way. - temperaturenumber
- Sampling temperature.
- top_pnumber
- Nucleus sampling.
- stopstring | array
- Sequences where the model stops.
- toolsarray
- Functions the model may call. See Tool calling.
- tool_choicestring | object
auto,none,requiredor a specific function.- response_formatobject
json_objectorjson_schema. See Structured outputs.- reasoningobject
- How much a reasoning model thinks:
{ "effort": "low" | "medium" | "high" }or{ "max_tokens": 2000 }. - modalitiesarray
["image", "text"]on models that generate images in a chat reply.- seedinteger
- Requests repeatable sampling where the model supports it.
Fields pass through to the provider, which ignores parameters the model does not support. models, route, provider and plugins that add a fee return 400 unsupported_parameter. Deference replaces user with an account tag.
Set max_tokens or max_completion_tokens to control the output limit and credit held for the request. On requests without server tools, Deference can lower the limit to fit available credit or the key's remaining limit. Requests with server tools must fit their full hold. See Credit.
Response
{
"id": "gen-1760000000-Ab3xQz9k",
"object": "chat.completion",
"created": 1760000000,
"model": "anthropic/claude-sonnet-5.5",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello, nice to meet you." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 14, "completion_tokens": 7, "total_tokens": 21 }
}finish_reason is stop, length, tool_calls, content_filter or error.
usage carries the provider's own fields, such as cached and reasoning token counts and a cost. What Deference charges, including the platform fee, is in x-deference-cost and in Activity.
Response headers
| Header | Meaning |
|---|---|
x-request-id | The request id, shown in Activity |
x-deference-cost | Provisional cost in micro-dollars. Set on non-streamed responses |
Streaming
With stream: true the response is a stream of data: chunks that ends with data: [DONE]. The last chunk before it carries usage. See Streaming and Streaming events.