# Pricing

> OpenRouter's price per token with no Deference markup, and how each request is held, charged and settled.

A model's price is OpenRouter's price plus OpenRouter's 5% crypto platform fee, which Deference pays when it buys OpenRouter credit. Deference adds no markup. You pay per request, with no subscription.

## Model prices

Each model has a price in dollars per million tokens for input and for output, and separate prices for cached input read and written. Some models also charge per request or per image. [Models](https://deference.si/models) and `GET /v1/models` show the prices you pay, fee included.

## What a request costs

| Tokens                     | Charged at            |
| -------------------------- | --------------------- |
| Input                      | The input price       |
| Input read from cache      | The cache read price  |
| Input written to cache     | The cache write price |
| Output, reasoning included | The output price      |

When the provider reports the cost itself, as it does for chat completions, embeddings and images, you pay that cost plus the 5% fee. Costs round up to the nearest millionth of a dollar.

## How a request is charged

| Step   | What happens                                                                                                                |
| ------ | --------------------------------------------------------------------------------------------------------------------------- |
| Hold   | Before the request runs, Deference holds the most it could cost: the prompt, the whole output cap and any per-request price |
| Charge | When the response ends, the hold is released and the measured cost is charged                                               |
| Final  | A few seconds later the cost is checked against the provider's record. Activity reads "estimating" until then               |

The output cap is `max_tokens`. Without it, the hold uses the model's largest output. Set `max_tokens` to hold less.

Some requests hold more:

* Each web search step a request allows is held again, with room for the results it carries forward.
* A file, image or video given as a URL is held at the model's whole context window, because the provider fetches it.
* Image generation holds an amount for each requested image, and its charge is final at once.

A request can cost more than its hold, for example when the prompt tokenizes more densely than estimated. A request paid with credit is then charged in full, and your balance can go below zero until you add credit. Free credit never goes below zero.

## Failed and cancelled requests

* A request that fails before the model runs costs nothing.
* A stream that fails partway is charged what the provider reports.
* If you disconnect during a stream, the tokens the model produced are charged.
* A call that is not streamed has 15 minutes. After that it returns `504 upstream_timeout` and its hold is charged, because the provider may still have run it. Stream long requests instead.
* Requests whose cost has no ceiling are refused with a 400 and cost nothing: search through X, audio input, and models that bill per search or per song.

## See what you spent

Non-streamed responses carry `x-deference-cost`, the cost in micro-dollars (1,000,000 is $1). [Activity](https://deference.si/activity) lists each request with its tokens and cost, and [Usage](https://deference.si/usage) totals them by model and key.
