Skip to content

Reference

Rate limits

Requests per minute, requests in flight, credit and the limits that apply to images.

Limits

LimitValueWhen exceeded
Requests per minute300 per key429 rate_limited with retry-after
Requests in flight8 per key429 too_many_concurrent_requests with retry-after: 1
Requests per minute on free credit only20 per account, across all its keys429 rate_limited with retry-after
CreditThe account's available credit must cover the request's hold402 insufficient_credit
Key credit limitThe limit you set on the key, less what its running requests hold402 key_limit_reached
Request body32 MB, or 64 MB on Images413 request_too_large
Time to answer15 minutes for a call that is not streamed, 120 seconds for image generation. Streamed text has no limit504 upstream_timeout, and the credit held for the call is charged. A stream of images is cut off

Limits are counted per key, except the free credit limit, which is counted per account. The limits apply to every endpoint that reaches a model.

The 20 per minute limit

An account that has never had credit added is limited to 20 requests per minute in total, whichever key sends them, and the error message says so. Adding credit for the first time lifts it, and the 300 per minute key limit remains. See Free credit.

Retry

A 429 carries retry-after, a whole number of seconds. Wait that long, then retry. If you get too_many_concurrent_requests, a request is still streaming: retry when one finishes, or keep fewer than 8 running.

import random
import time

import requests


def post_with_retry(url, headers, body, attempts=5):
    for attempt in range(attempts):
        response = requests.post(url, headers=headers, json=body, timeout=120)
        retryable = response.status_code in (429, 502, 503, 504)
        if not retryable or attempt == attempts - 1:
            return response
        wait = float(response.headers.get("retry-after", 2**attempt))
        time.sleep(wait + random.random())

Do not retry 400, 401, 402, 404 or 413: the same request fails again. See Errors.

Stay under the limits

  • Reuse connections and cap your own concurrency at 8 per key.
  • Create a key for each tool or job, so one busy job does not use another's budget.
  • Set max_tokens. It also lowers the credit held while a request runs.

Credit that a running request holds is released when it ends. See Pricing.