Alphabet Inc Class CGoogle unveils Gemini 3.5 Transcribe, a new speech-to-text model with superior accuracy, enhancing its AI product lineup.

Google has unveiled Gemini 3.5 Transcribe, its most precise speech-to-text model to date, which converts raw audio directly into accurate, polished, formatted text. The model is designed for developer workflows, enabling voice agents, real-time captioning, and post-call analytics, and is available for real-time streaming via the Live API and for pre-recorded audio via the Interactions API. It supports over 85 languages, can identify up to three speakers, and features a word error rate of 4% for streaming and 2.6% for non-streaming use cases, outperforming OpenAI's GPT Live Transcribe on the FLEURS benchmark. The model is now available in public preview for developers and enterprises, and in the Gemini app on macOS and Android in select countries, with Chrome support coming soon. This release follows Google's earlier launch of Gemini 3.7 Flash and the Gemini app surpassing 1 billion users, while the next frontier model, Gemini 3.5 Pro, remains unreleased.
Alphabet Inc Class CGoogle unveils Gemini 3.5 Transcribe, a new speech-to-text model with superior accuracy, enhancing its AI product lineup.