Skip to content

Coding tools

Kilo Code

Add Deference to Kilo Code as an OpenAI-compatible custom provider.

Kilo Code reads one config file in the CLI, VS Code and JetBrains. A custom provider takes a base URL and a key. This guide uses Chat Completions with https://deference.si/v1.

Set up

  1. Create a key in API keys and export it as DEFERENCE_API_KEY.
  2. Add the provider to your global config, ~/.config/kilo/kilo.jsonc (on Windows, C:\Users\<you>\.config\kilo\kilo.jsonc).
{
  "$schema": "https://app.kilo.ai/config.json",
  "model": "deference/claude-sonnet-5.5",
  "provider": {
    "deference": {
      "npm": "@ai-sdk/openai-compatible",
      "options": {
        "baseURL": "https://deference.si/v1",
        "apiKey": "{env:DEFERENCE_API_KEY}"
      },
      "models": {
        "claude-sonnet-5.5": {
          "id": "anthropic/claude-sonnet-5.5",
          "name": "Claude Sonnet 5.5",
          "tool_call": true,
          "limit": { "context": 200000, "output": 32000 }
        }
      }
    }
  }
}
  1. Restart Kilo Code.

{env:DEFERENCE_API_KEY} works only in the global config. In a project file that you commit, Kilo ignores it. Keep the key out of project files.

The key under models is the name Kilo shows and uses. id is what Kilo sends to Deference. Put the slash-free name in model and in -m.

To add the provider in VS Code instead, open Settings, then Providers, scroll to Custom provider, and set Provider API to OpenAI Compatible with the base URL and key. Token limits and tool calling still need the file.

Verify

  1. Run kilo models deference. The model you declared is listed.
  2. Run kilo roll-call "^deference/". It sends a short prompt and prints the latency per model.
  3. Run kilo run -m deference/claude-sonnet-5.5 "Reply with exactly: gateway-ok".
  4. Open Activity. The request is the top row, with the model you chose.

Pick a model

  • limit.context decides when Kilo compacts the conversation. With no limit, it never does. Set both context and output yourself.
  • limit.output is sent as max_tokens. Kilo caps it at 32,000 by default.
  • tool_call: true turns on file and terminal tools.
  • Use @ai-sdk/openai as npm for models served through POST /v1/responses, with the same base URL.

Troubleshooting

You seeCauseFix
401 invalid_api_key{env:DEFERENCE_API_KEY} sits in a project file, or the variable is not setMove the provider to the global config and export the variable
The model does not appearThe file is not valid JSONC, or model does not match a keyCheck commas, and use deference/ plus the key under models
The agent cannot edit filestool_call is offSet "tool_call": true
The conversation never compactslimit.context is missingSet limit.context
404 model_not_foundid differs from the catalogCopy the id from the Models page
402 insufficient_creditNo available creditAdd credit
402 free_credit_not_eligibleThe account has only free credit, and this model or request does not qualifyUse an open-weight text model, set a lower max_tokens, or add credit
402 key_limit_reachedThe key reached the credit limit set on itRaise the key's limit in API keys, or use another key