Skip to content

Start

API overview

Call models from your code with the OpenAI and Anthropic formats you already use.

Deference serves models through the formats your SDKs already speak. Change the base URL and the key, and keep the rest of your code. To set up a coding tool instead, see Coding tools.

  • OpenAI formatMost SDKs and toolshttps://deference.si/v1
  • Anthropic formatClaude Code and the Anthropic SDKhttps://deference.si
Anthropic clients add /v1/messages themselves, so their base URL has no /v1. A doubled or missing /v1 returns 404.

Send your key as Authorization: Bearer sk-df-... or as x-api-key: sk-df-.... See Authentication.

Which endpoint

EndpointFormatUse it for
POST /v1/chat/completionsOpenAI Chat CompletionsMost SDKs and tools. Text, vision, tools, structured outputs, streaming
POST /v1/messagesAnthropic MessagesClaude Code and the Anthropic SDK
POST /v1/responsesOpenAI ResponsesCodex and the OpenAI Responses API
POST /v1/embeddingsOpenAI EmbeddingsSearch and retrieval
POST /v1/imagesOpenRouter ImagesImage generation
GET /v1/modelsOpenAI list, extendedModel ids, prices and limits. No key needed
GET /v1/keyDeferenceThe calling key's limit, spend and available credit

Every response

  • x-request-id is on every response from the five POST endpoints, errors included. A request that reaches a model appears in Activity under that id. Requests rejected before that, and GET requests, are not logged.
  • x-deference-cost is the provisional cost in micro-dollars, on non-streamed responses that reach a model.
  • Errors use the OpenAI envelope on every endpoint except Messages, which uses Anthropic's. See Errors.

What Deference changes

  • Fields pass through to the provider, which ignores parameters a model does not support. models, route, provider and plugins that add a fee return 400 unsupported_parameter.
  • On Chat completions without server tools, the output limit can be lowered to fit available credit. Messages and Responses keep their output limits and refuse a request whose hold does not fit.
  • On Chat completions, Embeddings and Images, user is replaced with a stable account tag. Messages preserves its body except for capping a web-search tool's max_uses; Responses preserves its body.
  • Prompts and responses are not stored.

Limits

300 requests per minute and 8 in flight per key. Accounts on free credit only get 20 requests per minute. See Rate limits.

Also in this reference