In the present day, we’re introducing Gemini 3.5 Transcribe, our most exact speech-to-text mannequin but, designed for clever voice interactions. In contrast to typical speech recognition fashions that battle with background noise, advanced jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts uncooked audio instantly into correct, polished, formatted textual content.
Throughout our merchandise just like the Gemini app and on Android, we’ve seen customers already benefiting from this transcription mannequin with new voice capabilities like Rambler on Android and within the Gemini app on macOS. Now, builders can construct related capabilities with Gemini 3.5 Transcribe within the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
We have constructed 3.5 Transcribe to plug seamlessly into your developer workflows, whether or not you’re constructing voice brokers, real-time captioning instruments, or post-call analytics pipelines. The mannequin is obtainable throughout two separate APIs:
- Actual-time streaming: Delivers steady, bidirectional streaming with sub-second latency for interactive voice apps through the Live API utilizing
gemini-3.5-transcribe-live. - Pre-recorded audio processing: Transcribes recorded audio, conferences, name logs, and extra with speaker attribution and word-level timestamps through the Interactions API utilizing
gemini-3.5-transcribe.
Get extra exact and clever transcription
Gemini 3.5 Transcribe is designed to seize your pure talking model to raised perceive your intent and acknowledge customized vocabulary, so you possibly can execute duties together with your voice.
- Good transcription: Seamlessly handles self-corrections (like “let’s meet Tuesday—no, Wednesday”), removes filler phrases (“ums” and ‘“ahs”), auto-formats your textual content.
- Perform calling: The mannequin can delegate advanced duties (reminiscent of picture era and file evaluation) to different Gemini fashions through perform calls. At the moment accessible within the Gemini macOS app.
- Extra exact transcription: As measured by Synthetic Evaluation, achieves a median Phrase Error Charge (WER) of 4.0% for streaming and a couple of.6% for non-streaming use-cases. It exhibits sturdy efficiency throughout noisy, real-world environments, precisely capturing alphanumeric entities like postal codes and order IDs.
- Customized vocabulary: Acknowledges specialised jargon and distinctive spellings by seamlessly adapting transcriptions to your supplied customized vocabulary.
- International language assist: Robotically detects and transcribes over 85 languages, seamlessly dealing with regional accents and numerous dialects.
- Multi-speaker identification: Precisely attributes speech in pre-recorded audio with timestamps for as much as three audio system (assist for 3+ audio system is experimental).
