Skip to content

Concepts

Pricing

OpenRouter's price per token with no Deference markup, and how each request is held, charged and settled.

A model's price is OpenRouter's price plus OpenRouter's 5% crypto platform fee, which Deference pays when it buys OpenRouter credit. Deference adds no markup. You pay per request, with no subscription.

Model prices

Each model has a price in dollars per million tokens for input and for output, and separate prices for cached input read and written. Some models also charge per request or per image. Models and GET /v1/models show the prices you pay, fee included.

What a request costs

TokensCharged at
InputThe input price
Input read from cacheThe cache read price
Input written to cacheThe cache write price
Output, reasoning includedThe output price

When the provider reports the cost itself, as it does for chat completions, embeddings and images, you pay that cost plus the 5% fee. Costs round up to the nearest millionth of a dollar.

How a request is charged

StepWhat happens
HoldBefore the request runs, Deference holds the most it could cost: the prompt, the whole output cap and any per-request price
ChargeWhen the response ends, the hold is released and the measured cost is charged
FinalA few seconds later the cost is checked against the provider's record. Activity reads "estimating" until then

The output cap is max_tokens. Without it, the hold uses the model's largest output. Set max_tokens to hold less.

Some requests hold more:

  • Each web search step a request allows is held again, with room for the results it carries forward.
  • A file, image or video given as a URL is held at the model's whole context window, because the provider fetches it.
  • Image generation holds an amount for each requested image, and its charge is final at once.

A request can cost more than its hold, for example when the prompt tokenizes more densely than estimated. A request paid with credit is then charged in full, and your balance can go below zero until you add credit. Free credit never goes below zero.

Failed and cancelled requests

  • A request that fails before the model runs costs nothing.
  • A stream that fails partway is charged what the provider reports.
  • If you disconnect during a stream, the tokens the model produced are charged.
  • A call that is not streamed has 15 minutes. After that it returns 504 upstream_timeout and its hold is charged, because the provider may still have run it. Stream long requests instead.
  • Requests whose cost has no ceiling are refused with a 400 and cost nothing: search through X, audio input, and models that bill per search or per song.

See what you spent

Non-streamed responses carry x-deference-cost, the cost in micro-dollars (1,000,000 is $1). Activity lists each request with its tokens and cost, and Usage totals them by model and key.