# API overview

> Call models from your code with the OpenAI and Anthropic formats you already use.

Deference serves models through the formats your SDKs already speak. Change the base URL and the key, and keep the rest of your code. To set up a coding tool instead, see [Coding tools](https://deference.si/docs/coding-tools).

* OpenAI format (Most SDKs and tools): `https://deference.si/v1`
* Anthropic format (Claude Code and the Anthropic SDK): `https://deference.si`

Anthropic clients add `/v1/messages` themselves, so their base URL has no `/v1`. A doubled or missing `/v1` returns 404.

Send your key as `Authorization: Bearer sk-df-...` or as `x-api-key: sk-df-...`. See [Authentication](https://deference.si/docs/api-reference/authentication).

## Which endpoint

| Endpoint                                                            | Format                  | Use it for                                                              |
| ------------------------------------------------------------------- | ----------------------- | ----------------------------------------------------------------------- |
| [`POST /v1/chat/completions`](https://deference.si/docs/api-reference/chat-completions) | OpenAI Chat Completions | Most SDKs and tools. Text, vision, tools, structured outputs, streaming |
| [`POST /v1/messages`](https://deference.si/docs/api-reference/messages)                 | Anthropic Messages      | Claude Code and the Anthropic SDK                                       |
| [`POST /v1/responses`](https://deference.si/docs/api-reference/responses)               | OpenAI Responses        | Codex and the OpenAI Responses API                                      |
| [`POST /v1/embeddings`](https://deference.si/docs/api-reference/embeddings)             | OpenAI Embeddings       | Search and retrieval                                                    |
| [`POST /v1/images`](https://deference.si/docs/api-reference/images)                     | OpenRouter Images       | Image generation                                                        |
| [`GET /v1/models`](https://deference.si/docs/api-reference/models)                      | OpenAI list, extended   | Model ids, prices and limits. No key needed                             |
| [`GET /v1/key`](https://deference.si/docs/api-reference/key)                            | Deference               | The calling key's limit, spend and available credit                     |

## Every response

* `x-request-id` is on every response from the five `POST` endpoints, errors included. A request that reaches a model appears in [Activity](https://deference.si/docs/dashboard/activity) under that id. Requests rejected before that, and `GET` requests, are not logged.
* `x-deference-cost` is the provisional cost in micro-dollars, on non-streamed responses that reach a model.
* Errors use the OpenAI envelope on every endpoint except Messages, which uses Anthropic's. See [Errors](https://deference.si/docs/api-reference/errors).

## What Deference changes

* Fields pass through to the provider, which ignores parameters a model does not support. `models`, `route`, `provider` and `plugins` that add a fee return `400 unsupported_parameter`.
* On Chat completions without server tools, the output limit can be lowered to fit available credit. Messages and Responses keep their output limits and refuse a request whose hold does not fit.
* On Chat completions, Embeddings and Images, `user` is replaced with a stable account tag. Messages preserves its body except for capping a web-search tool's `max_uses`; Responses preserves its body.
* Prompts and responses are not stored.

## Limits

300 requests per minute and 8 in flight per key. Accounts on free credit only get 20 requests per minute. See [Rate limits](https://deference.si/docs/api-reference/rate-limits).

## Also in this reference

* [SDKs and clients](https://deference.si/docs/api-reference/sdks-and-clients): the OpenAI and Anthropic SDKs, the Vercel AI SDK, LiteLLM and LangChain
* Compatibility: [OpenAI](https://deference.si/docs/api-reference/openai-compatibility) and [Anthropic](https://deference.si/docs/api-reference/anthropic-compatibility)
* Guides: [Streaming](https://deference.si/docs/api-reference/streaming), [Tool calling](https://deference.si/docs/api-reference/tool-calling), [Structured outputs](https://deference.si/docs/api-reference/structured-outputs) and [Prompt caching](https://deference.si/docs/api-reference/prompt-caching)
* [Docs for agents](https://deference.si/docs/api-reference/docs-for-agents): every page as Markdown
