started · updated
Meta launches Muse Voice Transcribe real-time audio AI
Meta has launched Muse Voice Transcribe, its first real-time audio perception model developed by Meta Superintelligence Labs. The system integrates streaming speech-to-text, speaker diarization, and endpointing into a single model, allowing it to identify over 20 distinct speakers in recordings lasting more than an hour.
The model is trained on over 70 languages, with 25 validated at launch, including native support for five major Indian languages: Hindi, Tamil, Telugu, Kannada, and Malayalam. It features a unique ‘adaptive delay’ mechanism that uses reinforcement learning to balance speed and accuracy, deciding whether to emit a text token immediately or wait for more context.
Technically, the model processes audio in 80-millisecond chunks. It is designed to handle ‘code-switching,’ where speakers alternate between languages mid-sentence. Muse Voice Transcribe is currently available via the Meta Model API, Meta AI for macOS, and Muse Code. According to Artificial Analysis benchmarks, the model achieved a 3.1% word error rate in English, ranking it first in streaming speech-to-text performance.
Entities
Alexandr Wang · Artificial Analysis · Mark Zuckerberg · Meta · Meta Superintelligence Labs · Muse Voice Transcribe