API reference
Rate limits
Per-tier RPM, response headers, retry strategy.
Last updated: 2026-05-19
Rate limits
Limits are per API key. The tier sets the priority your request gets in the queue, not a guaranteed quota.
Read the limit from the response, not from a table
This page used to publish a fixed number of requests per minute for each tier — 60, 120, 240, 1000. Those numbers described nothing: a single limit was applied to every key regardless of tier, so anyone paying for the top tier received a sixteenth of what was written here.
We have not replaced them with four new numbers, because a fixed promise is the wrong shape for capacity that is not fixed. Compute is allocated dynamically by GigaKube according to the load in flight. What your key is allowed right now is in the response headers of every throttled endpoint — chat completions, embeddings, rerank, audio:
X-RateLimit-Limit: 120
X-RateLimit-Remaining: 119
GET /v1/models is not throttled and carries no such headers.
Two different things, often confused
Segnalato dal team di bagai il 18 agosto 2026: questa pagina dice «leggete gli header» mentre Audio pubblica dei numeri fissi. Sembra una contraddizione e non lo è, perché parlano di due grandezze diverse.
Le richieste al minuto sono fisse e per indirizzo. Sono scritte accanto a ogni indirizzo nella sua pagina, ed è giusto pubblicarle perché non cambiano:
| indirizzo | al minuto |
|---|---|
POST /v1/chat/completions |
60 |
POST /v1/responses |
60 |
POST /v1/audio/transcriptions |
30 |
POST /v1/audio/speech |
120 |
POST /v1/embeddings, POST /v1/rerank |
120 |
POST /api/v1/audio/transcribe (mobile) |
30 |
POST /api/v1/audio/synthesize (mobile) |
60 |
La capacità di calcolo dietro quelle richieste non è fissa. Una richiesta accettata dal contatore può comunque mettersi in coda, e quanto aspetta dipende dal carico in volo e dal livello di servizio della chiave. È questa la parte che non promettiamo con un numero.
Quindi: il contatore dice quante richieste potete fare, gli header dicono quante ve ne restano adesso, e nessuno dei due promette quanto aspetterete.
Read those headers and back off on 429. They are authoritative; a table in the
documentation is not.
If you need a contractually guaranteed floor — a known workload that must
not depend on the queue — that is what the company tier puts in writing. Send
us the volume and we will quote it.
How rate limits work
We use a sliding 1-minute window backed by Redis. Each request increments a
counter; if it exceeds the limit currently applied to your key, we return 429.
429 response
HTTP/1.1 429 Too Many Requests
Retry-After: 18
Content-Type: application/json
{
"error": {
"message": "rate limit exceeded (60/min)",
"type": "rate_limit_exceeded"
}
}
Retry-After is in seconds and is the safe time to wait before retrying. The window resets on the next minute boundary.
Recommended retry strategy
Exponential backoff with jitter, capped at 60s:
import time, random
def with_retry(fn, max_attempts=5):
for attempt in range(max_attempts):
try:
return fn()
except RateLimitError as e:
wait = min(60, (2 ** attempt) + random.uniform(0, 0.5))
if e.retry_after:
wait = max(wait, e.retry_after)
time.sleep(wait)
raise RuntimeError("max retries exceeded")
The official OpenAI Python SDK has built-in retry on 429; it works as-is against siati.
Per-organisation limits
For enterprise customers we can set an aggregate org-wide RPM and split between keys. Talk to us if your scale needs that.
Concurrent requests
Independent of RPM. Ogni sistema ha un limite di richieste in parallelo; l'instradamento sceglie quello con posti liberi. If all backends serving your model are saturated, a request queues (with priority based on tier).
The queue depth is exposed in Infrastructure (fleet) for transparency.