# Credit

> The dollar balance that pays for requests. Add it with USDC or INFERENCE on Wallet.

Credit is a dollar balance on your Deference account. Every request spends it. [Free credit](https://deference.si/docs/concepts/free-credit) is a separate balance for open-weight text models, spent before your credit.

## Add credit

Open [Wallet](https://deference.si/wallet?tab=add) and choose **Add credit**, then pick what to pay with.

| Pay with  | What happens                                                                         |
| --------- | ------------------------------------------------------------------------------------ |
| USDC      | The USDC goes to the reserve and the same amount of credit is added, in one step     |
| INFERENCE | The INFERENCE is burned and each one adds $1.00 of credit. This is called activation |

DEF holders can also claim their [rewards](https://deference.si/docs/concepts/rewards) straight into credit.

Credit stays on your account. It cannot be sent, sold, withdrawn or turned back into INFERENCE. If you may want to sell later, keep INFERENCE and activate it when you need credit.

## How a request spends it

1. When a request starts, Deference holds the most it could cost: the prompt plus its whole `max_tokens`. Held credit shows beside your balance.
2. When the response ends, the hold is released and the real cost is charged.
3. A few seconds later the cost is checked against the provider's record, and [Activity](https://deference.si/activity) shows the final amount.

Your available credit is your balance minus what running requests hold. A lower `max_tokens` holds less. [Pricing](https://deference.si/docs/concepts/pricing-and-metering) shows how each cost is worked out.

## When a request is refused

A request runs only when its whole hold fits in your available credit and in its key's remaining limit.

| Error                          | Cause                                                                | Fix                                                              |
| ------------------------------ | -------------------------------------------------------------------- | ---------------------------------------------------------------- |
| `402 insufficient_credit`      | Not enough available credit for the hold                             | Add credit, or lower `max_tokens`                                |
| `402 free_credit_not_eligible` | Only free credit is left, and this model or request does not qualify | Use an open-weight text model, lower `max_tokens`, or add credit |
| `402 key_limit_reached`        | The key has less left than the hold                                  | Raise the key's limit in [API keys](https://deference.si/keys), or use another key   |

Requests already running finish. On Chat completions without server tools, a request that does not fit runs with a lower `max_tokens` instead, as long as that leaves room for the prompt and one output token.

## Limit what a key can spend

Give each key its own credit limit and expiry, so one tool cannot spend your whole balance. See [Keys and security](https://deference.si/docs/dashboard/keys-and-security).
