Start
SDKs and clients
Use the OpenAI and Anthropic SDKs, the Vercel AI SDK, LiteLLM and LangChain by changing the base URL.
Deference has no SDK of its own. The SDKs you already use work as they are: set the base URL and the key.
| Client | Base URL |
|---|---|
| OpenAI SDK (Python, TypeScript) | https://deference.si/v1 |
| Anthropic SDK (Python, TypeScript) | https://deference.si |
| Vercel AI SDK, LiteLLM, LangChain | https://deference.si/v1 |
The examples read the key from DEFERENCE_API_KEY.
OpenAI SDK
import os
from openai import OpenAI
client = OpenAI(
base_url="https://deference.si/v1",
api_key=os.environ["DEFERENCE_API_KEY"],
)
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://deference.si/v1",
apiKey: process.env.DEFERENCE_API_KEY,
});
const response = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5.5",
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(response.choices[0].message.content);client.responses and client.embeddings work the same way. The SDK's image methods call OpenAI's own image paths, which Deference does not serve. Call POST /v1/images directly.
Anthropic SDK
The Anthropic SDK adds /v1/messages itself, so its base URL has no /v1. It sends the key as x-api-key.
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://deference.si",
api_key=os.environ["DEFERENCE_API_KEY"],
)
message = client.messages.create(
model="anthropic/claude-sonnet-5.5",
max_tokens=1024,
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(message.content[0].text)import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://deference.si",
apiKey: process.env.DEFERENCE_API_KEY,
});
const message = await client.messages.create({
model: "anthropic/claude-sonnet-5.5",
max_tokens: 1024,
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(message.content);messages.count_tokens is not available. The Claude Agent SDK reads ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN, as Claude Code does.
Vercel AI SDK
Use the OpenAI-compatible provider.
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";
const deference = createOpenAICompatible({
name: "deference",
baseURL: "https://deference.si/v1",
apiKey: process.env.DEFERENCE_API_KEY,
});
const { text } = await generateText({
model: deference("anthropic/claude-sonnet-5.5"),
prompt: "Say hello in one sentence.",
});
console.log(text);LiteLLM
Prefix the model with openai/ so LiteLLM uses the OpenAI-compatible path, then add Deference's id.
import os
import litellm
response = litellm.completion(
model="openai/anthropic/claude-sonnet-5.5",
api_base="https://deference.si/v1",
api_key=os.environ["DEFERENCE_API_KEY"],
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)LangChain
import os
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="anthropic/claude-sonnet-5.5",
base_url="https://deference.si/v1",
api_key=os.environ["DEFERENCE_API_KEY"],
)
print(llm.invoke("Say hello in one sentence.").content)Plain HTTP
No SDK is needed. Every endpoint takes JSON and a bearer key. See Chat completions for a curl example.
Setting up a coding tool instead? See Coding tools.