Coding tools
Codex
Add Deference as a model provider in the Codex config file.
Codex speaks the OpenAI Responses API. Add Deference as a provider in ~/.codex/config.toml. The base URL ends in /v1, and Codex adds /responses itself.
Set up
- Create a key in API keys and export it as
DEFERENCE_API_KEYin the shell that starts Codex. - Add the provider to
~/.codex/config.toml(on Windows,%USERPROFILE%\.codex\config.toml). Keep the top-level keys above the first table.
model = "openai/gpt-5.3-codex"
model_provider = "deference"
web_search = "disabled"
[model_providers.deference]
name = "Deference"
base_url = "https://deference.si/v1"
wire_api = "responses"
env_key = "DEFERENCE_API_KEY"- Restart Codex.
Codex's built-in web search is a tool hosted by OpenAI, so this guide turns it off.
Rules from Codex itself:
- Provider settings are read from the user-level file only. A
model_providerormodel_providersentry in a project.codex/config.tomlis ignored, with a warning at startup. - The ids
openai,ollamaandlmstudioare reserved. Name the provider anything else. wire_api = "responses"is the only accepted value.modelmust be an exact id from Models.- The key must be in the environment of the process that starts Codex. An app launched from the Dock or Start menu does not see variables you export in a terminal.
Verify
- Run
codex doctor. It checks the installation, the config and the authentication. - Start
codexand run/status. It shows the active model and provider. - Send
Reply with exactly: gateway-ok. - Open Activity. The request is the top row, with the endpoint Responses, and its client reads Codex. The reply alone does not prove the route, so check this row.
Pick a model
- Any catalog model can serve
POST /v1/responses, but each model handles Codex's message layout differently. Start with a model made for coding, such asopenai/gpt-5.3-codex, and switch if a run fails. - Choose a model per run with
codex -m <id>, or in a session with/model. - Codex has no catalog entry for ids it does not know, so it uses default limits. To set them, add
model_context_windowandmodel_auto_compact_token_limitat the top level of the file.
Let the agent manage keys
Connect the MCP server and Codex creates, rotates and revokes its own capped keys and checks your credit. You sign in once.
Troubleshooting
| You see | Cause | Fix |
|---|---|---|
stream closed before response.completed | The model rejected Codex's request layout | Try another model |
Unknown model, fallback metadata | Codex has no catalog entry for the id | Safe to ignore. Codex uses default limits |
The config fails to load, naming wire_api | wire_api is set to "chat", which is no longer accepted | Set wire_api = "responses" |
401 missing_api_key | DEFERENCE_API_KEY is not set in the shell that runs Codex | Export it in that shell, then restart Codex |
| The provider is ignored | The block is in a project config, or sits below the first table | Move it to ~/.codex/config.toml and keep top-level keys first |
400 mentioning store or previous_response_id | Responses runs statelessly | Leave store off and send the full conversation |
402 insufficient_credit | No available credit | Add credit |