Reference
Rate limits
Requests per minute, requests in flight, credit and the limits that apply to images.
Limits
| Limit | Value | When exceeded |
|---|---|---|
| Requests per minute | 300 per key | 429 rate_limited with retry-after |
| Requests in flight | 8 per key | 429 too_many_concurrent_requests with retry-after: 1 |
| Requests per minute on free credit only | 20 per account, across all its keys | 429 rate_limited with retry-after |
| Credit | The account's available credit must cover the request's hold | 402 insufficient_credit |
| Key credit limit | The limit you set on the key, less what its running requests hold | 402 key_limit_reached |
| Request body | 32 MB, or 64 MB on Images | 413 request_too_large |
| Time to answer | 15 minutes for a call that is not streamed, 120 seconds for image generation. Streamed text has no limit | 504 upstream_timeout, and the credit held for the call is charged. A stream of images is cut off |
Limits are counted per key, except the free credit limit, which is counted per account. The limits apply to every endpoint that reaches a model.
The 20 per minute limit
An account that has never had credit added is limited to 20 requests per minute in total, whichever key sends them, and the error message says so. Adding credit for the first time lifts it, and the 300 per minute key limit remains. See Free credit.
Retry
A 429 carries retry-after, a whole number of seconds. Wait that long, then retry. If you get too_many_concurrent_requests, a request is still streaming: retry when one finishes, or keep fewer than 8 running.
import random
import time
import requests
def post_with_retry(url, headers, body, attempts=5):
for attempt in range(attempts):
response = requests.post(url, headers=headers, json=body, timeout=120)
retryable = response.status_code in (429, 502, 503, 504)
if not retryable or attempt == attempts - 1:
return response
wait = float(response.headers.get("retry-after", 2**attempt))
time.sleep(wait + random.random())Do not retry 400, 401, 402, 404 or 413: the same request fails again. See Errors.
Stay under the limits
- Reuse connections and cap your own concurrency at 8 per key.
- Create a key for each tool or job, so one busy job does not use another's budget.
- Set
max_tokens. It also lowers the credit held while a request runs.
Credit that a running request holds is released when it ends. See Pricing.