Start
API overview
Call models from your code with the OpenAI and Anthropic formats you already use.
Deference serves models through the formats your SDKs already speak. Change the base URL and the key, and keep the rest of your code. To set up a coding tool instead, see Coding tools.
- OpenAI formatMost SDKs and toolshttps://deference.si/v1
- Anthropic formatClaude Code and the Anthropic SDKhttps://deference.si
Send your key as Authorization: Bearer sk-df-... or as x-api-key: sk-df-.... See Authentication.
Which endpoint
| Endpoint | Format | Use it for |
|---|---|---|
POST /v1/chat/completions | OpenAI Chat Completions | Most SDKs and tools. Text, vision, tools, structured outputs, streaming |
POST /v1/messages | Anthropic Messages | Claude Code and the Anthropic SDK |
POST /v1/responses | OpenAI Responses | Codex and the OpenAI Responses API |
POST /v1/embeddings | OpenAI Embeddings | Search and retrieval |
POST /v1/images | OpenRouter Images | Image generation |
GET /v1/models | OpenAI list, extended | Model ids, prices and limits. No key needed |
GET /v1/key | Deference | The calling key's limit, spend and available credit |
Every response
x-request-idis on every response from the fivePOSTendpoints, errors included. A request that reaches a model appears in Activity under that id. Requests rejected before that, andGETrequests, are not logged.x-deference-costis the provisional cost in micro-dollars, on non-streamed responses that reach a model.- Errors use the OpenAI envelope on every endpoint except Messages, which uses Anthropic's. See Errors.
What Deference changes
- Fields pass through to the provider, which ignores parameters a model does not support.
models,route,providerandpluginsthat add a fee return400 unsupported_parameter. - On Chat completions without server tools, the output limit can be lowered to fit available credit. Messages and Responses keep their output limits and refuse a request whose hold does not fit.
- On Chat completions, Embeddings and Images,
useris replaced with a stable account tag. Messages preserves its body except for capping a web-search tool'smax_uses; Responses preserves its body. - Prompts and responses are not stored.
Limits
300 requests per minute and 8 in flight per key. Accounts on free credit only get 20 requests per minute. See Rate limits.
Also in this reference
- SDKs and clients: the OpenAI and Anthropic SDKs, the Vercel AI SDK, LiteLLM and LangChain
- Compatibility: OpenAI and Anthropic
- Guides: Streaming, Tool calling, Structured outputs and Prompt caching
- Docs for agents: every page as Markdown