API reference
Chat completions
POST /v1/chat/completions — OpenAI-compatible, with extras for tier routing and priority.
Last updated: 2026-08-18
Chat completions
POST https://api.siati.ai/v1/chat/completions
Drop-in compatible with the OpenAI chat completions endpoint. Same request schema, same response schema, plus our additions for sovereignty (tier, priority).
Request
curl https://api.siati.ai/v1/chat/completions \
-H "Authorization: Bearer $SIATI_API_KEY" \
-H "Content-Type: application/json" \
-H "X-Siati-Tier: medium" \
-d '{
"model": "apertus-70b-instruct",
"messages": [
{"role": "system", "content": "You are a helpful Swiss assistant."},
{"role": "user", "content": "Spiegami il principio di sovranità in 3 frasi."}
],
"temperature": 0.7,
"max_tokens": 400,
"top_p": 0.9,
"stream": false
}'
Parameters
The complete list — what reaches the engine, what is rejected with a 400, and why — is on its own page: Ogni parametro: onorato o rifiutato.
Read it before assuming a parameter is applied. Until 18 August 2026 this table
listed top_p, stop, presence_penalty and frequency_penalty as supported
while the gateway silently dropped them; the table is gone rather than corrected,
because one list in one place is the only way it stays true.
Required: model, messages. Everything else is optional.
Headers (siati-specific)
| Header | Description |
|---|---|
X-Siati-Tier: slow|medium|fast|ludicrous |
Override the default tier of your key for this request. |
X-Request-Id: <uuid> |
Idempotency key for billing. Retry safe. |
Response
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1779203245,
"model": "apertus-70b-instruct",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "..." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 91,
"completion_tokens": 67,
"total_tokens": 158
}
}
Immagini
gemma-4-26b accetta immagini, fino a 2 per richiesta. Il campo content diventa un array di parti:
{
"model": "gemma-4-26b",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Che documento è questo? Estrai data e importo."},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
]
}]
}
image_url.url accetta due forme:
- data URI base64 — consigliata: nessuna latenza di rete aggiuntiva.
- URL pubblico http/https — lo scarichiamo noi e lo passiamo al modello. L'URL deve essere raggiungibile da internet: gli indirizzi di rete privata vengono rifiutati con
400. Non seguiamo i redirect, quindi indicate l'URL finale.
| Limite | Valore |
|---|---|
| Formati | PNG, JPEG, WebP, GIF |
| Dimensione massima | 8 MB per immagine |
| Pixel massimi | 40 megapixel |
| Immagini per richiesta | 2 su gemma-4-26b |
Il tipo viene determinato dai byte, non dall'estensione dichiarata: un base64 valido che non contiene un'immagine ritorna 400.
Costo: l'immagine viene convertita in token e conteggiata nei prompt_tokens, quindi si paga al prezzo di input del modello. Per riferimento misurato, un PNG di 640×420 pixel costa 305 token di prompt.
Modelli senza supporto: mandare un'immagine a un modello che non le accetta ritorna 400 con l'elenco di quelli che le accettano. Il badge vision nel catalogo segnala quali sono.
Models
Pick any model from the live catalog. Highlights:
gemma-4-26b— Google Gemma 4: multimodal (text + images), strong tool calling and robust JSON output. Fast MoE architecture — a great default for agents and structured extraction.apertus-70b-instruct— large Swiss-aligned model when you need depth.qwen2.5:32b/qwen2.5:14b/qwen2.5:1.5b— balanced sizes, low latency.
Multimodal request (Gemma 4)
Gemma 4 accepts up to 2 images per request via the standard OpenAI vision format:
curl https://api.siati.ai/v1/chat/completions \
-H "Authorization: Bearer $SIATI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-26b",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Cosa mostra questa immagine? Rispondi in JSON."},
{"type": "image_url", "image_url": {"url": "https://example.com/foto.jpg"}}
]
}]
}'
Errors
Standard codes from Errors. The most common:
400 invalid_request_error— bad shape, unknown role, etc.401 invalid_api_key— see Authentication.429 rate_limit_exceeded— see Rate limits. IncludesRetry-After.