Skip to content

Get started

What your key can do

Text, vision, image generation, embeddings, tool calling, structured outputs and reasoning, each with a request you can paste.

One key works on every model in the catalog. Each section is a request you can paste, with the models that support it and the full reference.

The examples read your key from DEFERENCE_API_KEY. Quickstart shows how to set it.

Text

Send messages and get a reply. Add "stream": true to stream it.

curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v3.2",
    "messages": [{ "role": "user", "content": "Explain a ledger in one sentence." }],
    "max_tokens": 200
  }'

Browse text models. Details: Chat completions and Streaming.

Vision

Add an image_url part to a message, after the text part. The URL can be public, or a data: URL with base64 content.

curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "messages": [{
      "role": "user",
      "content": [
        { "type": "text", "text": "What is in this image?" },
        { "type": "image_url", "image_url": {
          "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/960px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
        } }
      ]
    }],
    "max_tokens": 300
  }'

PNG, JPEG, WebP and GIF work, and images are billed as input tokens. Browse vision models.

Image generation

POST /v1/images takes a model and a prompt and returns base64 images. Generation can take over a minute.

curl https://deference.si/v1/images \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "recraft/recraft-v4.1-flash",
    "prompt": "A lighthouse on a cliff at dawn, flat illustration"
  }' | jq -r '.data[0].b64_json' | base64 --decode > lighthouse.png

Image generation needs credit: free credit does not pay for it. Browse image models. Details: Images.

Embeddings

Turn text into vectors for search and retrieval.

curl https://deference.si/v1/embeddings \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/text-embedding-3-small",
    "input": "A ledger is a list of entries."
  }'

Browse embedding models. Details: Embeddings.

Tool calling

Describe functions in tools. When the model wants one, the reply carries tool_calls with the function name and its arguments.

curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "messages": [{ "role": "user", "content": "What is the weather in Lisbon?" }],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": { "city": { "type": "string" } },
          "required": ["city"]
        }
      }
    }]
  }'

Browse models with tools. The full loop, with results sent back, is in Tool calling.

Structured outputs

Set response_format to a JSON schema and the reply is JSON that matches it.

curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "messages": [{ "role": "user", "content": "Extract: Ada paid $12.40 on Oct 8." }],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "payment",
        "strict": true,
        "schema": {
          "type": "object",
          "properties": {
            "name": { "type": "string" },
            "amount_usd": { "type": "number" }
          },
          "required": ["name", "amount_usd"],
          "additionalProperties": false
        }
      }
    }
  }'

Check that the model lists response_format or structured_outputs under supported parameters on its page. Details: Structured outputs.

Reasoning

Reasoning models think before they answer. Set reasoning to control how much. The reply carries the thinking in message.reasoning, and reasoning tokens are billed as output.

curl https://deference.si/v1/chat/completions \
  -H "Authorization: Bearer $DEFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5.5",
    "messages": [{ "role": "user", "content": "Is 9.11 larger than 9.9?" }],
    "reasoning": { "effort": "high" },
    "max_tokens": 4000
  }'

effort is low, medium or high. For a token budget instead, send reasoning.max_tokens. On Messages, send Anthropic's thinking object. Browse reasoning models.