# Streaming

> Receive tokens as they are generated on chat completions, messages and responses.

Set `stream` to `true` and the response arrives as server-sent events. Nothing is buffered, so the first token reaches you as soon as the model produces it.

## Chat Completions

**curl**

```bash
curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "stream": true,
    "stream_options": { "include_usage": true },
    "messages": [{ "role": "user", "content": "Write a haiku about ledgers." }]
  }'
```

**Python**

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])

stream = client.chat.completions.create(
    model="anthropic/claude-sonnet-5.5",
    messages=[{"role": "user", "content": "Write a haiku about ledgers."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
```

**TypeScript**

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://deference.si/v1",
  apiKey: process.env.DEFERENCE_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5.5",
  messages: [{ role: "user", content: "Write a haiku about ledgers." }],
  stream: true,
  stream_options: { include_usage: true },
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
```

The last chunk before `[DONE]` carries `usage`. With `include_usage`, it has an empty `choices` array, as in OpenAI's API. Without it, the chunk keeps one choice with an empty delta, so read `choices[0]` only when `choices` is not empty.

## Messages and Responses

Messages streams `message_start`, content block events, `message_delta` and `message_stop`. Responses streams events such as `response.output_text.delta` and ends with `response.completed`, `response.incomplete` or `response.failed`. Both formats are relayed unchanged. See [Streaming events](https://deference.si/docs/api-reference/streaming-events) for each sequence.

## Keepalives

Streams can contain comment lines that begin with `:` and `ping` events during long pauses. SSE parsers skip comments. If you parse lines by hand, ignore lines that do not start with `data:`.

## Cancelling

Closing the connection stops your client from receiving tokens. Deference still reads the response to the end, and you are charged for the tokens the model produced.

## Errors in a stream

A failure after the stream starts arrives as an event inside an HTTP `200` response. Check for an `error` event or `finish_reason: "error"` before trusting the output.

If the connection to the provider drops, no error event is sent: the stream ends. A stream that ends without its final event (`[DONE]`, `message_stop`, or a final `response.*` event) is incomplete.
