meta
3.1% word error rate leads the streaming speech chart
Promtime
Meta Superintelligence Labs has introduced Muse Voice Transcribe, a streaming speech-recognition model with a 3.1% word error rate on Artificial Analysis’ English AA-WER Streaming benchmark, according to The New Stack.
At a glance
- The model processes incoming audio in 80-millisecond chunks, can follow conversations lasting more than an hour, and was trained on over 70 languages, including language switching within one conversation.
- Meta says the system separates more than 20 speakers and has extensively verified 25 languages, a narrower group within the more than 70 languages used for training.
- Developers can access the service through the Meta Model API, while Meta also offers it in Meta AI for Mac and Muse Code at a listed price of $3 per 1,000 audio minutes.
Streaming transcription is now a product battleground rather than a narrow model feature. The small differences among reported English results suggest that leadership may shift quickly as general-purpose AI vendors and speech specialists ship new systems. For Meta, the release appears to fit products where ongoing audio understanding is central, including its glasses and Mac app, while the closed model keeps deployment under Meta’s control.
Five English streaming models record word error rates from 3.1% to 4%
Artificial Analysis’ AA-WER Streaming test is limited to English speech. In the comparison reported by The New Stack, Muse Voice Transcribe recorded a 3.1% word error rate, compared with 3.4% for Cartesia Ink-2, 3.6% for ElevenLabs’ Scribe v2 Real-time, 3.9% for GPT Live Transcribe, and 4% for Gemini 3.5 Transcribe Live.
Across several standard speaker-recognition benchmarks, Meta says Muse Voice Transcribe recorded a 17.5% error rate. The source describes speaker identification as a harder task than transcription in real-time settings. That measurement covers recognition of distinct speakers, rather than the English word-error score used by AA-WER Streaming.
Audio arrives in 80-millisecond chunks that govern the model’s timing
Meta describes Muse Voice Transcribe as an autoregressive multimodal model in the Muse Spark family. It receives audio as 80-millisecond segments, or 12.5 each second, compressing each segment into one soft token before deciding whether to output text or request the next piece of audio. Meta recently also released Muse Spark, which can transcribe speech but does not specialize in this use case.
The special next-audio placeholder is replaced by the next audio chunk. When input ends, an empty-audio token tells the model to flush retained text. Meta calls this control over context and response time adaptive delay: difficult words can collect more context, while simpler words can be transcribed sooner. During reinforcement learning, Meta multiplies word-error and delay rewards rather than adding them.
More than 20 speakers can be separated in long multilingual conversations
Meta says the model distinguishes more than 20 speakers and supports conversations longer than an hour. It was trained on more than 70 languages, with 25 extensively verified, and is intended to handle a conversation in which multilingual speakers change languages partway through. Meta calls Muse Voice Transcribe its first real-time audio perception model.
Meta says a similar decision mechanism is used for detecting speakers, tying speaker recognition to the model’s timing decisions. At each point in the incoming stream, Muse Voice Transcribe decides whether it has enough audio context to commit a text token, rather than applying a fixed waiting interval before every word. The system makes that decision for every audio chunk.
Access without open weights Meta has made Muse Voice Transcribe available through the Meta Model API, Meta AI for Mac, and Muse Code. The API costs $3.00 per 1,000 audio minutes, or $0.18 an hour. Unlike Meta’s Muse Glimmer models, the new model’s weights will not be released openly, a Meta spokesperson told The New Stack. The source says OpenAI, Google, xAI, and Alibaba have also released streaming models within weeks of one another this summer.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
