Gemini 3.5 Transcribe: Convert streamed speech to textual content
Actual-time speech understanding is essential for voice-first interfaces. Final month, we launched Gemini 3.5 Transcribe for low-latency transcription with excessive precision, reaching a 4.0% WER, and helpful options:
- Automated code-switching: Deal with intra-sentence and inter-sentential code- and language-switching with out handbook configuration
- Customized vocabulary biasing: Steer speech recognition towards domain-specific phrases, unusual jargon, firm names, and correct nouns by passing a custom_vocabulary record of as much as 1,000 phrases
- Good transcription mode: Ship polished, reader-ready transcripts with structured formatting, self-corrections, and disfluency removing that eliminates filler phrases
3.5 Transcribe helps 85+ languages and gives a powerful listening engine for voice experiences and stateless duties like sub-second captioning, name heart brokers, and real-time audio analytics. You too can entry the mannequin by way of the Interactions API to transcribe audio recordsdata as much as 1 hour lengthy with structured timestamps and speaker labeling. Learn our developer guide to study extra.
