# Kilo Code

> Add Deference to Kilo Code as an OpenAI-compatible custom provider.

Kilo Code reads one config file in the CLI, VS Code and JetBrains. A custom provider takes a base URL and a key. This guide uses Chat Completions with `https://deference.si/v1`.

## Set up

1. Create a key in [API keys](https://deference.si/keys) and export it as `DEFERENCE_API_KEY`.
2. Add the provider to your global config, `~/.config/kilo/kilo.jsonc` (on Windows, `C:\Users\<you>\.config\kilo\kilo.jsonc`).

```jsonc
{
  "$schema": "https://app.kilo.ai/config.json",
  "model": "deference/claude-sonnet-5.5",
  "provider": {
    "deference": {
      "npm": "@ai-sdk/openai-compatible",
      "options": {
        "baseURL": "https://deference.si/v1",
        "apiKey": "{env:DEFERENCE_API_KEY}"
      },
      "models": {
        "claude-sonnet-5.5": {
          "id": "anthropic/claude-sonnet-5.5",
          "name": "Claude Sonnet 5.5",
          "tool_call": true,
          "limit": { "context": 200000, "output": 32000 }
        }
      }
    }
  }
}
```

3. Restart Kilo Code.

`{env:DEFERENCE_API_KEY}` works only in the global config. In a project file that you commit, Kilo ignores it. Keep the key out of project files.

The key under `models` is the name Kilo shows and uses. `id` is what Kilo sends to Deference. Put the slash-free name in `model` and in `-m`.

To add the provider in VS Code instead, open Settings, then **Providers**, scroll to **Custom provider**, and set **Provider API** to **OpenAI Compatible** with the base URL and key. Token limits and tool calling still need the file.

## Verify

1. Run `kilo models deference`. The model you declared is listed.
2. Run `kilo roll-call "^deference/"`. It sends a short prompt and prints the latency per model.
3. Run `kilo run -m deference/claude-sonnet-5.5 "Reply with exactly: gateway-ok"`.
4. Open [Activity](https://deference.si/activity). The request is the top row, with the model you chose.

## Pick a model

* `limit.context` decides when Kilo compacts the conversation. With no limit, it never does. Set both `context` and `output` yourself.
* `limit.output` is sent as `max_tokens`. Kilo caps it at 32,000 by default.
* `tool_call: true` turns on file and terminal tools.
* Use `@ai-sdk/openai` as `npm` for models served through `POST /v1/responses`, with the same base URL.

## Troubleshooting

| You see                         | Cause                                                                        | Fix                                                                                       |
| ------------------------------- | ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- |
| `401 invalid_api_key`           | `{env:DEFERENCE_API_KEY}` sits in a project file, or the variable is not set | Move the provider to the global config and export the variable                            |
| The model does not appear       | The file is not valid JSONC, or `model` does not match a key                 | Check commas, and use `deference/` plus the key under `models`                            |
| The agent cannot edit files     | `tool_call` is off                                                           | Set `"tool_call": true`                                                                   |
| The conversation never compacts | `limit.context` is missing                                                   | Set `limit.context`                                                                       |
| `404 model_not_found`           | `id` differs from the catalog                                                | Copy the id from the Models page                                                          |
| `402 insufficient_credit`       | No available credit                                                          | [Add credit](https://deference.si/wallet?tab=add)                                                             |
| `402 free_credit_not_eligible`  | The account has only free credit, and this model or request does not qualify | Use an open-weight text model, set a lower `max_tokens`, or [add credit](https://deference.si/wallet?tab=add) |
| `402 key_limit_reached`         | The key reached the credit limit set on it                                   | Raise the key's limit in API keys, or use another key                                     |
