Skip to content

Start

OpenAI compatibility

Use any OpenAI SDK or tool by changing the base URL and the key.

Deference accepts OpenAI Chat Completions, Responses and Embeddings requests. Set the base URL and key, and keep the rest of your code.

from openai import OpenAI

client = OpenAI(
    base_url="https://deference.si/v1",
    api_key="sk-df-...",
)

Supported endpoints

EndpointNotes
POST /v1/chat/completionsStreaming, tools, structured outputs and images where the model supports them
POST /v1/responsesStateless: send the whole conversation in input each time
POST /v1/embeddingsEmbedding models only
POST /v1/imagesDeference's image endpoint. The OpenAI SDK's images.generate calls a different path, so call this one directly
GET /v1/modelsPublic, no key needed
GET /v1/models/{id}One model
GET /v1/keyThe calling key's limit, spend and available credit

How requests are handled

  • Request fields pass through to the provider, which ignores parameters a model does not support. Chat completions can lower the output limit to fit available credit when no server tools are requested.
  • Model ids are author/name, for example anthropic/claude-sonnet-5.5. Choose an endpoint that matches the model's output: chat or Responses for text, Embeddings for embeddings, and Images for image generation. See Choose a model.
  • Each request runs on the one model it names, at that model's price. models, route, provider and plugins that add a fee, such as web search and the paid PDF parsers, return 400 unsupported_parameter. response-healing and the free PDF parser work.
  • On chat completions and embeddings, user is replaced with a stable account tag that cannot be reversed. Responses passes through unchanged.
  • Non-streamed responses carry x-request-id and x-deference-cost, the cost in micro-dollars before final settlement.

Files and server tools

  • Models that list file input read files themselves. For other models, Chat completions parses files with the free cloudflare-ai parser. Responses and Messages return 400 unsupported_parameter unless the request sets "plugins": [{ "id": "file-parser", "pdf": { "engine": "cloudflare-ai" } }].
  • A file, image or video given as a URL is held at the model's whole context window, because the provider fetches it.
  • Tools the provider runs, such as web search, are held step by step. max_tool_calls lowers the number of steps, up to 30.
  • Search through X (x_search) and audio input (input_audio) return 400 unsupported_parameter. Models that bill per search or per song, such as perplexity/sonar-deep-research and the Lyria music models, return 400 unsupported_model.

Streaming

A chat stream ends with a chunk that carries usage. With stream_options.include_usage, that chunk has an empty choices array, as in OpenAI's API. Without it, the chunk keeps one choice with an empty delta.

Not available

/v1/completions, audio, files, batches, assistants, fine-tuning and moderation are not part of Deference, and neither are OpenAI's image paths (/v1/images/generations, edits and variations). Calls to them return 404.

Errors

Errors use OpenAI's envelope: error.message, error.type, error.code and error.param. See Errors.