Skip to content

Endpoints

Chat completions

Create a model response from a list of messages, with optional streaming, images and tools.

POST/v1/chat/completions

Follows OpenAI's Chat Completions format. Send the key as Authorization: Bearer sk-df-... or x-api-key.

Example

curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "messages": [{ "role": "user", "content": "Say hello in one sentence." }]
  }'

Request body

modelstringRequired
The model id, such as anthropic/claude-sonnet-5.5.
messagesarrayRequired
The conversation. Each item has a role (system, user, assistant or tool) and content. Content can be a string or an array of parts, including image_url parts for vision.
streamboolean
Return server-sent events. Defaults to false.
stream_optionsobject
{ "include_usage": true } makes the final usage chunk follow OpenAI's shape, with an empty choices array.
max_tokensinteger
The output limit. Also sets the credit held while the request runs. max_completion_tokens works the same way.
temperaturenumber
Sampling temperature.
top_pnumber
Nucleus sampling.
stopstring | array
Sequences where the model stops.
toolsarray
Functions the model may call. See Tool calling.
tool_choicestring | object
auto, none, required or a specific function.
response_formatobject
json_object or json_schema. See Structured outputs.
reasoningobject
How much a reasoning model thinks: { "effort": "low" | "medium" | "high" } or { "max_tokens": 2000 }.
modalitiesarray
["image", "text"] on models that generate images in a chat reply.
seedinteger
Requests repeatable sampling where the model supports it.

Fields pass through to the provider, which ignores parameters the model does not support. models, route, provider and plugins that add a fee return 400 unsupported_parameter. Deference replaces user with an account tag.

Set max_tokens or max_completion_tokens to control the output limit and credit held for the request. On requests without server tools, Deference can lower the limit to fit available credit or the key's remaining limit. Requests with server tools must fit their full hold. See Credit.

Response

{
  "id": "gen-1760000000-Ab3xQz9k",
  "object": "chat.completion",
  "created": 1760000000,
  "model": "anthropic/claude-sonnet-5.5",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hello, nice to meet you." },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 14, "completion_tokens": 7, "total_tokens": 21 }
}

finish_reason is stop, length, tool_calls, content_filter or error.

usage carries the provider's own fields, such as cached and reasoning token counts and a cost. What Deference charges, including the platform fee, is in x-deference-cost and in Activity.

Response headers

HeaderMeaning
x-request-idThe request id, shown in Activity
x-deference-costProvisional cost in micro-dollars. Set on non-streamed responses

Streaming

With stream: true the response is a stream of data: chunks that ends with data: [DONE]. The last chunk before it carries usage. See Streaming and Streaming events.