Skip to content

Guides

Streaming

Receive tokens as they are generated on chat completions, messages and responses.

Set stream to true and the response arrives as server-sent events. Nothing is buffered, so the first token reaches you as soon as the model produces it.

Chat Completions

curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "stream": true,
    "stream_options": { "include_usage": true },
    "messages": [{ "role": "user", "content": "Write a haiku about ledgers." }]
  }'

The last chunk before [DONE] carries usage. With include_usage, it has an empty choices array, as in OpenAI's API. Without it, the chunk keeps one choice with an empty delta, so read choices[0] only when choices is not empty.

Messages and Responses

Messages streams message_start, content block events, message_delta and message_stop. Responses streams events such as response.output_text.delta and ends with response.completed, response.incomplete or response.failed. Both formats are relayed unchanged. See Streaming events for each sequence.

Keepalives

Streams can contain comment lines that begin with : and ping events during long pauses. SSE parsers skip comments. If you parse lines by hand, ignore lines that do not start with data:.

Cancelling

Closing the connection stops your client from receiving tokens. Deference still reads the response to the end, and you are charged for the tokens the model produced.

Errors in a stream

A failure after the stream starts arrives as an event inside an HTTP 200 response. Check for an error event or finish_reason: "error" before trusting the output.

If the connection to the provider drops, no error event is sent: the stream ends. A stream that ends without its final event ([DONE], message_stop, or a final response.* event) is incomplete.