Gemini Transcribe 3.5
Released on August 26, 2026
Release Summary
Google's most precise speech-to-text model yet, succeeding Chirp 3 for live captioning, voice agents, and file transcription. File API gemini-3.5-transcribe adds speaker attribution and word-level timestamps; Live API gemini-3.5-transcribe-live streams with sub-second latency. Published results show 4.0% WER streaming and 2.6% non-streaming; FLEURS top-locale streaming WER is 5.50% versus 7.32% for Chirp 3. Supports 85+ languages, custom vocabulary, and filler-word cleanup. Public preview in Google AI Studio, Antigravity, and Gemini Enterprise. File pricing $2/M audio input and $12/M text output (about $0.005/min); live is $3.50/$21 (about $0.009/min).
Timeline
Gemini 3.5 Transcribe released
Google opens Gemini 3.5 Transcribe in public preview on the Gemini API, with Rambler on Android and the Gemini app on macOS already using the model.