Endpoints
Models
List the catalog or fetch one model, with prices and capabilities.
GET/v1/models
Returns every active model, of every type. Read architecture.output_modalities to tell text models from image and embeddings models. No key is needed. Claude Code reads this endpoint for model discovery, with ?limit=1000.
GET/v1/models/{id}
Returns one model. Ids contain a slash, so write it as part of the path: /v1/models/anthropic/claude-haiku-4.5. An unknown id returns 404 model_not_found.
Example
curl https://deference.si/v1/modelsimport requests
models = requests.get("https://deference.si/v1/models").json()["data"]
print([model["id"] for model in models][:5])const response = await fetch("https://deference.si/v1/models");
const { data } = await response.json();
console.log(data.slice(0, 5).map((model: { id: string }) => model.id));Response
{
"object": "list",
"data": [
{
"id": "anthropic/claude-sonnet-5.5",
"object": "model",
"created": 1760000000,
"owned_by": "anthropic",
"display_name": "Claude Sonnet 5.5",
"description": "A fast model for coding and agents.",
"context_length": 200000,
"max_output_tokens": 64000,
"architecture": { "input_modalities": ["text", "image"], "output_modalities": ["text"] },
"pricing": {
"prompt": "0.000003",
"completion": "0.000015",
"input_cache_read": "0.0000003",
"input_cache_write": "0.00000375",
"request": "0",
"image": "0"
},
"supported_parameters": ["tools", "tool_choice", "max_tokens", "temperature"]
}
],
"has_more": false,
"first_id": "anthropic/claude-sonnet-5.5",
"last_id": "anthropic/claude-sonnet-5.5"
}Fields
- idstring
- The id to send as
model. - display_namestring
- A readable name.
- descriptionstring
- What the model is for.
- context_lengthinteger
- Context window in tokens.
- max_output_tokensinteger
- The largest output the model allows.
- architectureobject
input_modalitiesandoutput_modalities.- pricingobject
- USD per token as strings:
prompt,completion,input_cache_read,input_cache_write.requestandimageare USD per request and per image. - supported_parametersarray
- Request parameters the model accepts.