Guides
Streaming
Receive tokens as they are generated on chat completions, messages and responses.
Set stream to true and the response arrives as server-sent events. Nothing is buffered, so the first token reaches you as soon as the model produces it.
Chat Completions
curl https://deference.si/v1/chat/completions \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"stream": true,
"stream_options": { "include_usage": true },
"messages": [{ "role": "user", "content": "Write a haiku about ledgers." }]
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
stream = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "Write a haiku about ledgers."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://deference.si/v1",
apiKey: process.env.DEFERENCE_API_KEY,
});
const stream = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5.5",
messages: [{ role: "user", content: "Write a haiku about ledgers." }],
stream: true,
stream_options: { include_usage: true },
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}The last chunk before [DONE] carries usage. With include_usage, it has an empty choices array, as in OpenAI's API. Without it, the chunk keeps one choice with an empty delta, so read choices[0] only when choices is not empty.
Messages and Responses
Messages streams message_start, content block events, message_delta and message_stop. Responses streams events such as response.output_text.delta and ends with response.completed, response.incomplete or response.failed. Both formats are relayed unchanged. See Streaming events for each sequence.
Keepalives
Streams can contain comment lines that begin with : and ping events during long pauses. SSE parsers skip comments. If you parse lines by hand, ignore lines that do not start with data:.
Cancelling
Closing the connection stops your client from receiving tokens. Deference still reads the response to the end, and you are charged for the tokens the model produced.
Errors in a stream
A failure after the stream starts arrives as an event inside an HTTP 200 response. Check for an error event or finish_reason: "error" before trusting the output.
If the connection to the provider drops, no error event is sent: the stream ends. A stream that ends without its final event ([DONE], message_stop, or a final response.* event) is incomplete.