Live transcription is Solaria-1 only; Solaria-3 is pre-recorded and single language. Match your declared audio metadata to the stream exactly, and keep the callback endpoint public.

Context: Official docs (Gladia SDK docs): live transcription works with Solaria-1 only. If you want Solaria-3, that is pre-recorded audio and a single language per request. Agents that default to the newest model hit this wall fast. Also the audio metadata you declare has to match the stream exactly, and the callback endpoint has to be public or you never get results back.