started · updated
Google launches Gemini 3. 5 Transcribe speech-to-text model
Google has introduced Gemini 3. 5 Transcribe, its most advanced speech-to-text model to date, designed to convert natural, unpolished speech into structured, formatted text. The model is available in two versions: a Live API for real-time, bidirectional streaming with sub-second latency, and an Interactions API for processing pre-recorded audio files.
Key features include the ability to automatically detect over 85 languages, recognize specialized jargon through custom vocabularies, and remove verbal fillers like 'um' and 'ah.' The model also handles self-corrections and can identify up to three speakers in recorded audio. According to Google, the model achieves a word error rate (WER) of 4.0% for streaming and 2.6% for non-streaming audio, representing a 70% improvement in transcription speed compared to the previous Chirp 3 model.
Gemini 3. 5 Transcribe is being integrated across various platforms, including the Gemini app on macOS, Android's Gboard via the Rambler feature, and is planned for future deployment in Chrome. Developers can access the model through Google AI Studio and the Gemini Enterprise Agent Platform in public preview.
Entities
Android · Android · Gemini · Gemini 3.5 Transcribe · Gemini Audio · Gemini Live · Google · Google DeepMind · OpenAI · macOS
Claims
What the coverage asserts, and how many sources carry each claim.
- [● 9 SOURCES] The model can automatically remove filler words such as 'um' and 'ah' and handle self-corrections. www.tokenpost.kr · dailyguardian.ae · news.nextapple.com · dailyguardian.ca · 4sysops.com · +4 more
- [● 6 SOURCES] The model is available in public preview via Google AI Studio and the Gemini Enterprise Agent Platform. www.tokenpost.kr · news.nextapple.com · tech-noisy.com · 4sysops.com · deepmind.google · +1 more
- [● 5 SOURCES] The technology is being integrated into Android's Rambler dictation feature. www.tokenpost.kr · dailyguardian.ca · news.nextapple.com · tech-noisy.com · weel.co.jp
- [● 6 SOURCES] The model can identify up to three speakers in recorded audio, with support for more being experimental. www.tokenpost.kr · news.nextapple.com · tech-noisy.com · 4sysops.com · www.tabletowo.pl · +1 more
- [● 5 SOURCES] The model shows a 70% improvement in time to final transcription compared to the previous Chirp 3 model. www.tokenpost.kr · news.nextapple.com · tech-noisy.com · 4sysops.com · www.tabletowo.pl
- [● 11 SOURCES] Google has released Gemini 3. 5 Transcribe, a new speech-to-text AI model. www.tokenpost.kr · dailyguardian.ae · news.nextapple.com · dailyguardian.ca · tech-noisy.com · +6 more
- [● 8 SOURCES] The model has a word error rate (WER) of 4.0% for real-time streaming and 2.6% for non-streaming audio. www.tokenpost.kr · news.nextapple.com · dailyguardian.ca · tech-noisy.com · 4sysops.com · +3 more
- [● 10 SOURCES] The model automatically detects and supports over 85 languages. www.tokenpost.kr · dailyguardian.ae · news.nextapple.com · dailyguardian.ca · tech-noisy.com · +5 more