Start
OpenAI compatibility
Use any OpenAI SDK or tool by changing the base URL and the key.
Deference accepts OpenAI Chat Completions, Responses and Embeddings requests. Set the base URL and key, and keep the rest of your code.
from openai import OpenAI
client = OpenAI(
base_url="https://deference.si/v1",
api_key="sk-df-...",
)Supported endpoints
| Endpoint | Notes |
|---|---|
POST /v1/chat/completions | Streaming, tools, structured outputs and images where the model supports them |
POST /v1/responses | Stateless: send the whole conversation in input each time |
POST /v1/embeddings | Embedding models only |
POST /v1/images | Deference's image endpoint. The OpenAI SDK's images.generate calls a different path, so call this one directly |
GET /v1/models | Public, no key needed |
GET /v1/models/{id} | One model |
GET /v1/key | The calling key's limit, spend and available credit |
How requests are handled
- Request fields pass through to the provider, which ignores parameters a model does not support. Chat completions can lower the output limit to fit available credit when no server tools are requested.
- Model ids are
author/name, for exampleanthropic/claude-sonnet-5.5. Choose an endpoint that matches the model's output: chat or Responses for text, Embeddings for embeddings, and Images for image generation. See Choose a model. - Each request runs on the one model it names, at that model's price.
models,route,providerandpluginsthat add a fee, such as web search and the paid PDF parsers, return 400unsupported_parameter.response-healingand the free PDF parser work. - On chat completions and embeddings,
useris replaced with a stable account tag that cannot be reversed. Responses passes through unchanged. - Non-streamed responses carry
x-request-idandx-deference-cost, the cost in micro-dollars before final settlement.
Files and server tools
- Models that list
fileinput read files themselves. For other models, Chat completions parses files with the freecloudflare-aiparser. Responses and Messages return 400unsupported_parameterunless the request sets"plugins": [{ "id": "file-parser", "pdf": { "engine": "cloudflare-ai" } }]. - A file, image or video given as a URL is held at the model's whole context window, because the provider fetches it.
- Tools the provider runs, such as web search, are held step by step.
max_tool_callslowers the number of steps, up to 30. - Search through X (
x_search) and audio input (input_audio) return 400unsupported_parameter. Models that bill per search or per song, such asperplexity/sonar-deep-researchand the Lyria music models, return 400unsupported_model.
Streaming
A chat stream ends with a chunk that carries usage. With stream_options.include_usage, that chunk has an empty choices array, as in OpenAI's API. Without it, the chunk keeps one choice with an empty delta.
Not available
/v1/completions, audio, files, batches, assistants, fine-tuning and moderation are not part of Deference, and neither are OpenAI's image paths (/v1/images/generations, edits and variations). Calls to them return 404.
Errors
Errors use OpenAI's envelope: error.message, error.type, error.code and error.param. See Errors.