Google has introduced Gemini 3.5 Transcribe, a new speech-to-text model built to turn natural speech into accurate, polished text. It is designed to cope with background noise, specialized terminology and verbal corrections while preserving the speaker’s intent.

The model supports two APIs: the Live API provides bidirectional, sub-second streaming for real-time voice applications, while the Interactions API handles recorded audio with speaker attribution and word-level timestamps. Gemini 3.5 Transcribe can recognize more than 85 languages and adapt to custom vocabulary.

According to Artificial Analysis, the model reaches a 4.0% Word Error Rate for streaming and 2.6% for non-streaming transcription. It also cuts time to final transcription by 70% compared with Google’s Chirp 3, while supporting multilingual speech and identifying up to three speakers in recorded audio.

Gemini 3.5 Transcribe is available in public preview through Google AI Studio and Google Antigravity, with enterprise access through the Gemini Enterprise Agent Platform. It already powers voice features in the Gemini app on macOS and Rambler on Android, while voice typing in Chrome is planned for a future release.