# OpenAI compatibility

> Use any OpenAI SDK or tool by changing the base URL and the key.

Deference accepts OpenAI Chat Completions, Responses and Embeddings requests. Set the base URL and key, and keep the rest of your code.

```python
from openai import OpenAI

client = OpenAI(
    base_url="https://deference.si/v1",
    api_key="sk-df-...",
)
```

## Supported endpoints

| Endpoint                    | Notes                                                                                                            |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| `POST /v1/chat/completions` | Streaming, tools, structured outputs and images where the model supports them                                    |
| `POST /v1/responses`        | Stateless: send the whole conversation in `input` each time                                                      |
| `POST /v1/embeddings`       | Embedding models only                                                                                            |
| `POST /v1/images`           | Deference's image endpoint. The OpenAI SDK's `images.generate` calls a different path, so call this one directly |
| `GET /v1/models`            | Public, no key needed                                                                                            |
| `GET /v1/models/{id}`       | One model                                                                                                        |
| `GET /v1/key`               | The calling key's limit, spend and available credit                                                              |

## How requests are handled

* Request fields pass through to the provider, which ignores parameters a model does not support. Chat completions can lower the output limit to fit available credit when no server tools are requested.
* Model ids are `author/name`, for example `anthropic/claude-sonnet-5.5`. Choose an endpoint that matches the model's output: chat or Responses for text, Embeddings for embeddings, and Images for image generation. See [Choose a model](https://deference.si/docs/models).
* Each request runs on the one model it names, at that model's price. `models`, `route`, `provider` and `plugins` that add a fee, such as web search and the paid PDF parsers, return 400 `unsupported_parameter`. `response-healing` and the free PDF parser work.
* On chat completions and embeddings, `user` is replaced with a stable account tag that cannot be reversed. Responses passes through unchanged.
* Non-streamed responses carry `x-request-id` and `x-deference-cost`, the cost in micro-dollars before final settlement.

## Files and server tools

* Models that list `file` input read files themselves. For other models, Chat completions parses files with the free `cloudflare-ai` parser. Responses and Messages return 400 `unsupported_parameter` unless the request sets `"plugins": [{ "id": "file-parser", "pdf": { "engine": "cloudflare-ai" } }]`.
* A file, image or video given as a URL is held at the model's whole context window, because the provider fetches it.
* Tools the provider runs, such as web search, are held step by step. `max_tool_calls` lowers the number of steps, up to 30.
* Search through X (`x_search`) and audio input (`input_audio`) return 400 `unsupported_parameter`. Models that bill per search or per song, such as `perplexity/sonar-deep-research` and the Lyria music models, return 400 `unsupported_model`.

## Streaming

A chat stream ends with a chunk that carries `usage`. With `stream_options.include_usage`, that chunk has an empty `choices` array, as in OpenAI's API. Without it, the chunk keeps one choice with an empty delta.

## Not available

`/v1/completions`, audio, files, batches, assistants, fine-tuning and moderation are not part of Deference, and neither are OpenAI's image paths (`/v1/images/generations`, edits and variations). Calls to them return `404`.

## Errors

Errors use OpenAI's envelope: `error.message`, `error.type`, `error.code` and `error.param`. See [Errors](https://deference.si/docs/api-reference/errors).
