< Back to all clusters
[TECHNOLOGY] · United States · 44 sources

started · updated

Google launches Gemini 3. 5 Transcribe speech-to-text model

Google has introduced Gemini 3. 5 Transcribe, its most advanced speech-to-text model to date, designed to convert natural, unpolished speech into structured, formatted text. The model is available in two versions: a Live API for real-time, bidirectional streaming with sub-second latency, and an Interactions API for processing pre-recorded audio files.

Key features include the ability to automatically detect over 85 languages, recognize specialized jargon through custom vocabularies, and remove verbal fillers like 'um' and 'ah.' The model also handles self-corrections and can identify up to three speakers in recorded audio. According to Google, the model achieves a word error rate (WER) of 4.0% for streaming and 2.6% for non-streaming audio, representing a 70% improvement in transcription speed compared to the previous Chirp 3 model.

Gemini 3. 5 Transcribe is being integrated across various platforms, including the Gemini app on macOS, Android's Gboard via the Rambler feature, and is planned for future deployment in Chrome. Developers can access the model through Google AI Studio and the Gemini Enterprise Agent Platform in public preview.

Entities

Android · Android · Gemini · Gemini 3.5 Transcribe · Gemini Audio · Gemini Live · Google · Google DeepMind · OpenAI · macOS

Claims

What the coverage asserts, and how many sources carry each claim.

Sources