Skip to content

Start

SDKs and clients

Use the OpenAI and Anthropic SDKs, the Vercel AI SDK, LiteLLM and LangChain by changing the base URL.

Deference has no SDK of its own. The SDKs you already use work as they are: set the base URL and the key.

ClientBase URL
OpenAI SDK (Python, TypeScript)https://deference.si/v1
Anthropic SDK (Python, TypeScript)https://deference.si
Vercel AI SDK, LiteLLM, LangChainhttps://deference.si/v1

The examples read the key from DEFERENCE_API_KEY.

OpenAI SDK

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://deference.si/v1",
    api_key=os.environ["DEFERENCE_API_KEY"],
)

response = client.chat.completions.create(
    model="anthropic/claude-sonnet-5.5",
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)

client.responses and client.embeddings work the same way. The SDK's image methods call OpenAI's own image paths, which Deference does not serve. Call POST /v1/images directly.

Anthropic SDK

The Anthropic SDK adds /v1/messages itself, so its base URL has no /v1. It sends the key as x-api-key.

import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://deference.si",
    api_key=os.environ["DEFERENCE_API_KEY"],
)

message = client.messages.create(
    model="anthropic/claude-sonnet-5.5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(message.content[0].text)

messages.count_tokens is not available. The Claude Agent SDK reads ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN, as Claude Code does.

Vercel AI SDK

Use the OpenAI-compatible provider.

import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";

const deference = createOpenAICompatible({
  name: "deference",
  baseURL: "https://deference.si/v1",
  apiKey: process.env.DEFERENCE_API_KEY,
});

const { text } = await generateText({
  model: deference("anthropic/claude-sonnet-5.5"),
  prompt: "Say hello in one sentence.",
});
console.log(text);

LiteLLM

Prefix the model with openai/ so LiteLLM uses the OpenAI-compatible path, then add Deference's id.

import os
import litellm

response = litellm.completion(
    model="openai/anthropic/claude-sonnet-5.5",
    api_base="https://deference.si/v1",
    api_key=os.environ["DEFERENCE_API_KEY"],
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)

LangChain

import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="anthropic/claude-sonnet-5.5",
    base_url="https://deference.si/v1",
    api_key=os.environ["DEFERENCE_API_KEY"],
)
print(llm.invoke("Say hello in one sentence.").content)

Plain HTTP

No SDK is needed. Every endpoint takes JSON and a bearer key. See Chat completions for a curl example.

Setting up a coding tool instead? See Coding tools.