Reference
Streaming events
The event sequence each endpoint sends when stream is true.
Streams are server-sent events with content-type: text/event-stream. Events are relayed as they arrive.
Sequences
data: {"id":"gen-...","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
data: {"id":"gen-...","choices":[{"index":0,"delta":{"content":"Hello"}}]}
data: {"id":"gen-...","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"gen-...","choices":[],"usage":{"prompt_tokens":14,"completion_tokens":7,"total_tokens":21}}
data: [DONE]event: response.created
event: response.in_progress
event: response.output_item.added
event: response.content_part.added
event: response.output_text.delta
event: response.output_text.done
event: response.content_part.done
event: response.output_item.done
event: response.completedevent: message_start
event: content_block_start
event: content_block_delta
event: content_block_stop
event: message_delta
event: message_stopChat Completions
Each chunk has choices[].delta. The last content chunk carries finish_reason. The last chunk before data: [DONE] carries usage. With stream_options.include_usage, that chunk has an empty choices array, as in OpenAI's API. Without it, the chunk keeps one choice with an empty delta.
Comment lines that begin with : may appear between events. Parsers skip them.
Responses
The event types are response.created, response.in_progress, response.output_item.added, response.content_part.added, response.output_text.delta, response.output_text.done, response.content_part.done, response.function_call_arguments.delta, response.function_call_arguments.done, response.output_item.done, response.completed, response.incomplete, response.failed and error. The stream ends with one of three final events, each carrying the final response object, with its usage when the provider reports it: response.completed when the answer finished, response.incomplete when it stopped early (the reason is in incomplete_details), or response.failed.
Messages
message_start opens with the input token count. Each content block sends content_block_start, one or more content_block_delta events, and content_block_stop. Delta types are text_delta, input_json_delta, thinking_delta and signature_delta. message_delta carries stop_reason and cumulative usage.output_tokens, then message_stop ends the stream.
ping events can appear anywhere, and an error event can appear mid-stream. Ignore event types you do not know.
Images
POST /v1/images streams only on models that support it. The events are image_generation.partial_image as the image forms, then image_generation.completed with the final image and usage, then data: [DONE]. An error event can replace either.
Errors
A failure after the stream starts keeps HTTP 200 and arrives as an event: a top-level error with finish_reason: "error" on Chat Completions, response.failed or error on Responses, and an error event on Messages.