siati.ai docs

API reference

What we change in responses

The only two things the gateway changes in what the model produces, and why. Declared, not hidden.

Last updated: 2026-10-03

The gateway changes two things in what the model produces. Nothing else. We write them down here because a service whose point is saying exactly what it does cannot rewrite answers silently.

1. The model's internal delimiters

After a tool call, some models leave in the text the markers they use to separate their internal channels:

text
<|channel>thought
<channel|>OK. I created the task for the customer Rossi.

Those symbols are protocol, not content: anyone showing that field on a screen would see them. We remove them, together with the word thought when it stays alone at the start of a line.

The list of delimiters is closed: <|channel>, <channel|>, <|start|>, <|end|>, <|message|>, <|im_start|>, <|im_end|> and the role markers. A generic rule on <|…|> could eat legitimate text — for example an answer that talks about syntax — and we do not use one.

2. The word "null" instead of the null value

In the arguments of a tool call, on a field that allows null, the model sometimes writes the four-letter string. A client that checks !== null saves it, and the database ends up with a record whose description literally says "null".

We turn into null only these values, and only when the field contains them entirely, after trimming spaces:

text
null    none    nil    nessuno    n/a    undefined    (empty string)

The comparison is case-insensitive. A sentence that contains the word is not touched: "the field was null and must be fixed" stays as it is.

It applies to the arguments of tool calls, where the field has a declared type. We do not touch the free text of the answer.

3. Silence on input

This is not a change to the answer, but it belongs to the same family and it should be said.

On audio without sound, /v1/audio/transcriptions returns empty text without asking the model. The reason: on digital silence Whisper answers "Grazie." or "Thank you." with no_speech_prob at 0.0000, that is with full confidence. On a dictation path an accidental tap on the microphone would turn into a command.

The check measures the energy of the waveform, not the model's judgement, and triggers below about -50 dBFS — below are digital silence and a closed microphone; above is any speech, even whispered into a distant microphone. It applies only to 16-bit WAV files: compressed formats always go through, because we cannot measure them without decoding and when in doubt, we transcribe. Good audio thrown away is much worse than a hallucination.

What we do not touch

The text of the answer, the numbers, the language, the punctuation, the order of JSON keys. If the model gets a total wrong, that total reaches you as it was written: add up the lines and compare, as you should with any provider.