# Responses

> Create a response with the OpenAI Responses API. This is the endpoint Codex uses.

**POST** `https://deference.si/v1/responses`

Follows OpenAI's Responses format and passes through unchanged. Responses runs statelessly: send the whole conversation in `input` on every request. Send the key as `Authorization: Bearer sk-df-...`.

## Example

**curl**

```bash
curl https://deference.si/v1/responses \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "input": "Say hello in one sentence."
  }'
```

**Python**

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])

response = client.responses.create(
    model="anthropic/claude-sonnet-5.5",
    input="Say hello in one sentence.",
)
print(response.output_text)
```

**TypeScript**

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://deference.si/v1",
  apiKey: process.env.DEFERENCE_API_KEY,
});

const response = await client.responses.create({
  model: "anthropic/claude-sonnet-5.5",
  input: "Say hello in one sentence.",
});
console.log(response.output_text);
```

## Request body

* `model` (string, required): The model id.
* `input` (string | array, required): A prompt, or a list of input items for a conversation.
* `instructions` (string): System-level instructions.
* `stream` (boolean): Return server-sent events. Codex sets it.
* `tools` (array): Tools the model may call.
* `max_output_tokens` (integer): The output limit. Also sets the credit held while the request runs.
* `reasoning` (object): Reasoning settings, for models that support them.
* `store` (boolean): Leave unset or `false`. Stored responses are not supported and return `400`.
* `previous_response_id` (string): Not supported. Send the full conversation in `input`.

## Streaming

A stream ends with `response.completed`, `response.incomplete` (the reason is in `incomplete_details`) or `response.failed`. Each carries the final response object, with its usage when the provider reports it. Codex waits for `response.completed`. See [Streaming events](https://deference.si/docs/api-reference/streaming-events).

## Models

Each model reads Responses requests in its own way. If a model fails with a Codex-style request, try another. See [Codex](https://deference.si/docs/coding-tools/codex-cli).
