Get started
What your key can do
Text, vision, image generation, embeddings, tool calling, structured outputs and reasoning, each with a request you can paste.
One key works on every model in the catalog. Each section is a request you can paste, with the models that support it and the full reference.
- TextChat completions
- VisionChat completions
- Image generationImages
- EmbeddingsEmbeddings
- Tool callingChat completions
- Structured outputsChat completions
- ReasoningChat completions
The examples read your key from DEFERENCE_API_KEY. Quickstart shows how to set it.
Text
Send messages and get a reply. Add "stream": true to stream it.
curl https://deference.si/v1/chat/completions \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v3.2",
"messages": [{ "role": "user", "content": "Explain a ledger in one sentence." }],
"max_tokens": 200
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
response = client.chat.completions.create(
model="deepseek/deepseek-v3.2",
messages=[{"role": "user", "content": "Explain a ledger in one sentence."}],
max_tokens=200,
)
print(response.choices[0].message.content)Browse text models. Details: Chat completions and Streaming.
Vision
Add an image_url part to a message, after the text part. The URL can be public, or a data: URL with base64 content.
curl https://deference.si/v1/chat/completions \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"messages": [{
"role": "user",
"content": [
{ "type": "text", "text": "What is in this image?" },
{ "type": "image_url", "image_url": {
"url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/960px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
} }
]
}],
"max_tokens": 300
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
url = (
"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/"
"Gfp-wisconsin-madison-the-nature-boardwalk.jpg/"
"960px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
)
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": url}},
],
}],
max_tokens=300,
)
print(response.choices[0].message.content)PNG, JPEG, WebP and GIF work, and images are billed as input tokens. Browse vision models.
Image generation
POST /v1/images takes a model and a prompt and returns base64 images. Generation can take over a minute.
curl https://deference.si/v1/images \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "recraft/recraft-v4.1-flash",
"prompt": "A lighthouse on a cliff at dawn, flat illustration"
}' | jq -r '.data[0].b64_json' | base64 --decode > lighthouse.pngimport base64
import os
import requests
response = requests.post(
"https://deference.si/v1/images",
headers={"Authorization": f"Bearer {os.environ['DEFERENCE_API_KEY']}"},
json={
"model": "recraft/recraft-v4.1-flash",
"prompt": "A lighthouse on a cliff at dawn, flat illustration",
},
timeout=180,
)
response.raise_for_status()
image = response.json()["data"][0]["b64_json"]
with open("lighthouse.png", "wb") as file:
file.write(base64.b64decode(image))Image generation needs credit: free credit does not pay for it. Browse image models. Details: Images.
Embeddings
Turn text into vectors for search and retrieval.
curl https://deference.si/v1/embeddings \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/text-embedding-3-small",
"input": "A ledger is a list of entries."
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
result = client.embeddings.create(
model="openai/text-embedding-3-small",
input="A ledger is a list of entries.",
)
print(len(result.data[0].embedding))Browse embedding models. Details: Embeddings.
Tool calling
Describe functions in tools. When the model wants one, the reply carries tool_calls with the function name and its arguments.
curl https://deference.si/v1/chat/completions \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"messages": [{ "role": "user", "content": "What is the weather in Lisbon?" }],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
}]
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "What is the weather in Lisbon?"}],
tools=tools,
)
print(response.choices[0].message.tool_calls)Browse models with tools. The full loop, with results sent back, is in Tool calling.
Structured outputs
Set response_format to a JSON schema and the reply is JSON that matches it.
curl https://deference.si/v1/chat/completions \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"messages": [{ "role": "user", "content": "Extract: Ada paid $12.40 on Oct 8." }],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "payment",
"strict": true,
"schema": {
"type": "object",
"properties": {
"name": { "type": "string" },
"amount_usd": { "type": "number" }
},
"required": ["name", "amount_usd"],
"additionalProperties": false
}
}
}
}'import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "Extract: Ada paid $12.40 on Oct 8."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "payment",
"strict": True,
"schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
"amount_usd": {"type": "number"},
},
"required": ["name", "amount_usd"],
"additionalProperties": False,
},
},
},
)
print(json.loads(response.choices[0].message.content))Check that the model lists response_format or structured_outputs under supported parameters on its page. Details: Structured outputs.
Reasoning
Reasoning models think before they answer. Set reasoning to control how much. The reply carries the thinking in message.reasoning, and reasoning tokens are billed as output.
curl https://deference.si/v1/chat/completions \
-H "Authorization: Bearer $DEFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"messages": [{ "role": "user", "content": "Is 9.11 larger than 9.9?" }],
"reasoning": { "effort": "high" },
"max_tokens": 4000
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://deference.si/v1", api_key=os.environ["DEFERENCE_API_KEY"])
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "Is 9.11 larger than 9.9?"}],
max_tokens=4000,
extra_body={"reasoning": {"effort": "high"}},
)
print(response.choices[0].message.content)effort is low, medium or high. For a token budget instead, send reasoning.max_tokens. On Messages, send Anthropic's thinking object. Browse reasoning models.