siati.ai docs

API reference

Rate limits

Per-tier RPM, response headers, retry strategy.

Last updated: 2026-05-19

Rate limits

Limits are per API key. The tier sets the priority your request gets in the queue, not a guaranteed quota.

Read the limit from the response, not from a table

This page used to publish a fixed number of requests per minute for each tier — 60, 120, 240, 1000. Those numbers described nothing: a single limit was applied to every key regardless of tier, so anyone paying for the top tier received a sixteenth of what was written here.

We have not replaced them with four new numbers, because a fixed promise is the wrong shape for capacity that is not fixed. Compute is allocated dynamically by GigaKube according to the load in flight. What your key is allowed right now is in the response headers of every throttled endpoint — chat completions, embeddings, rerank, audio:

http
X-RateLimit-Limit: 120
X-RateLimit-Remaining: 119

GET /v1/models is not throttled and carries no such headers.

Two different things, often confused

Segnalato dal team di bagai il 18 agosto 2026: questa pagina dice «leggete gli header» mentre Audio pubblica dei numeri fissi. Sembra una contraddizione e non lo è, perché parlano di due grandezze diverse.

Le richieste al minuto sono fisse e per indirizzo. Sono scritte accanto a ogni indirizzo nella sua pagina, ed è giusto pubblicarle perché non cambiano:

indirizzo al minuto
POST /v1/chat/completions 60
POST /v1/responses 60
POST /v1/audio/transcriptions 30
POST /v1/audio/speech 120
POST /v1/embeddings, POST /v1/rerank 120
POST /api/v1/audio/transcribe (mobile) 30
POST /api/v1/audio/synthesize (mobile) 60

La capacità di calcolo dietro quelle richieste non è fissa. Una richiesta accettata dal contatore può comunque mettersi in coda, e quanto aspetta dipende dal carico in volo e dal livello di servizio della chiave. È questa la parte che non promettiamo con un numero.

Quindi: il contatore dice quante richieste potete fare, gli header dicono quante ve ne restano adesso, e nessuno dei due promette quanto aspetterete.

Read those headers and back off on 429. They are authoritative; a table in the documentation is not.

If you need a contractually guaranteed floor — a known workload that must not depend on the queue — that is what the company tier puts in writing. Send us the volume and we will quote it.

How rate limits work

We use a sliding 1-minute window backed by Redis. Each request increments a counter; if it exceeds the limit currently applied to your key, we return 429.

429 response

http
HTTP/1.1 429 Too Many Requests
Retry-After: 18
Content-Type: application/json

{
  "error": {
    "message": "rate limit exceeded (60/min)",
    "type": "rate_limit_exceeded"
  }
}

Retry-After is in seconds and is the safe time to wait before retrying. The window resets on the next minute boundary.

Recommended retry strategy

Exponential backoff with jitter, capped at 60s:

python
import time, random

def with_retry(fn, max_attempts=5):
    for attempt in range(max_attempts):
        try:
            return fn()
        except RateLimitError as e:
            wait = min(60, (2 ** attempt) + random.uniform(0, 0.5))
            if e.retry_after:
                wait = max(wait, e.retry_after)
            time.sleep(wait)
    raise RuntimeError("max retries exceeded")

The official OpenAI Python SDK has built-in retry on 429; it works as-is against siati.

Per-organisation limits

For enterprise customers we can set an aggregate org-wide RPM and split between keys. Talk to us if your scale needs that.

Concurrent requests

Independent of RPM. Ogni sistema ha un limite di richieste in parallelo; l'instradamento sceglie quello con posti liberi. If all backends serving your model are saturated, a request queues (with priority based on tier).

The queue depth is exposed in Infrastructure (fleet) for transparency.