Skip to content

Endpoints

Responses

Create a response with the OpenAI Responses API. This is the endpoint Codex uses.

POST/v1/responses

Follows OpenAI's Responses format and passes through unchanged. Responses runs statelessly: send the whole conversation in input on every request. Send the key as Authorization: Bearer sk-df-....

Example

curl https://deference.si/v1/responses \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "input": "Say hello in one sentence."
  }'

Request body

modelstringRequired
The model id.
inputstring | arrayRequired
A prompt, or a list of input items for a conversation.
instructionsstring
System-level instructions.
streamboolean
Return server-sent events. Codex sets it.
toolsarray
Tools the model may call.
max_output_tokensinteger
The output limit. Also sets the credit held while the request runs.
reasoningobject
Reasoning settings, for models that support them.
storeboolean
Leave unset or false. Stored responses are not supported and return 400.
previous_response_idstring
Not supported. Send the full conversation in input.

Streaming

A stream ends with response.completed, response.incomplete (the reason is in incomplete_details) or response.failed. Each carries the final response object, with its usage when the provider reports it. Codex waits for response.completed. See Streaming events.

Models

Each model reads Responses requests in its own way. If a model fails with a Codex-style request, try another. See Codex.