siati.ai docs

API reference

Chat completions

POST /v1/chat/completions — OpenAI-compatible, with extras for tier routing and priority.

Last updated: 2026-08-18

Chat completions

POST https://api.siati.ai/v1/chat/completions

Drop-in compatible with the OpenAI chat completions endpoint. Same request schema, same response schema, plus our additions for sovereignty (tier, priority).

Request

bash
curl https://api.siati.ai/v1/chat/completions \
  -H "Authorization: Bearer $SIATI_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Siati-Tier: medium" \
  -d '{
    "model": "apertus-70b-instruct",
    "messages": [
      {"role": "system", "content": "You are a helpful Swiss assistant."},
      {"role": "user",   "content": "Spiegami il principio di sovranità in 3 frasi."}
    ],
    "temperature": 0.7,
    "max_tokens": 400,
    "top_p": 0.9,
    "stream": false
  }'

Parameters

The complete list — what reaches the engine, what is rejected with a 400, and why — is on its own page: Ogni parametro: onorato o rifiutato.

Read it before assuming a parameter is applied. Until 18 August 2026 this table listed top_p, stop, presence_penalty and frequency_penalty as supported while the gateway silently dropped them; the table is gone rather than corrected, because one list in one place is the only way it stays true.

Required: model, messages. Everything else is optional.

Headers (siati-specific)

Header Description
X-Siati-Tier: slow|medium|fast|ludicrous Override the default tier of your key for this request.
X-Request-Id: <uuid> Idempotency key for billing. Retry safe.

Response

json
{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1779203245,
  "model": "apertus-70b-instruct",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "..." },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 91,
    "completion_tokens": 67,
    "total_tokens": 158
  }
}

Immagini

gemma-4-26b accetta immagini, fino a 2 per richiesta. Il campo content diventa un array di parti:

json
{
  "model": "gemma-4-26b",
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "Che documento è questo? Estrai data e importo."},
      {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
    ]
  }]
}

image_url.url accetta due forme:

  • data URI base64 — consigliata: nessuna latenza di rete aggiuntiva.
  • URL pubblico http/https — lo scarichiamo noi e lo passiamo al modello. L'URL deve essere raggiungibile da internet: gli indirizzi di rete privata vengono rifiutati con 400. Non seguiamo i redirect, quindi indicate l'URL finale.
Limite Valore
Formati PNG, JPEG, WebP, GIF
Dimensione massima 8 MB per immagine
Pixel massimi 40 megapixel
Immagini per richiesta 2 su gemma-4-26b

Il tipo viene determinato dai byte, non dall'estensione dichiarata: un base64 valido che non contiene un'immagine ritorna 400.

Costo: l'immagine viene convertita in token e conteggiata nei prompt_tokens, quindi si paga al prezzo di input del modello. Per riferimento misurato, un PNG di 640×420 pixel costa 305 token di prompt.

Modelli senza supporto: mandare un'immagine a un modello che non le accetta ritorna 400 con l'elenco di quelli che le accettano. Il badge vision nel catalogo segnala quali sono.

Models

Pick any model from the live catalog. Highlights:

  • gemma-4-26b — Google Gemma 4: multimodal (text + images), strong tool calling and robust JSON output. Fast MoE architecture — a great default for agents and structured extraction.
  • apertus-70b-instruct — large Swiss-aligned model when you need depth.
  • qwen2.5:32b / qwen2.5:14b / qwen2.5:1.5b — balanced sizes, low latency.

Multimodal request (Gemma 4)

Gemma 4 accepts up to 2 images per request via the standard OpenAI vision format:

bash
curl https://api.siati.ai/v1/chat/completions \
  -H "Authorization: Bearer $SIATI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26b",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Cosa mostra questa immagine? Rispondi in JSON."},
        {"type": "image_url", "image_url": {"url": "https://example.com/foto.jpg"}}
      ]
    }]
  }'

Errors

Standard codes from Errors. The most common:

  • 400 invalid_request_error — bad shape, unknown role, etc.
  • 401 invalid_api_key — see Authentication.
  • 429 rate_limit_exceeded — see Rate limits. Includes Retry-After.

Tips