Models

Muse Voice Transcribe Just Raised the Bar for Real-Time Audio AI

CRAZE CRAZE Summary 3 things to know
  • Meta's Muse Voice Transcribe bundles transcription, 20+ speaker separation, and endpoint detection into one streaming pass, eliminating multi-model latency.
  • Meta claims 3.1% WER and 17.5% diarization error, but these are unverified launch-day figures—no outside reproduction yet.
  • The model undercuts Google's live rate at $0.18/hour and supports speaker ID, unlike Gemini's live endpoint; voice API market now has three players.
Jeff Editorial | · 5 min read
Muse Voice Transcribe Just Raised the Bar for Real-Time Audio AI

On September 1, Meta released Muse Voice Transcribe through Meta Model API, the Meta AI Mac app, and Muse Code. It's a single streaming model that handles three tasks in one pass: real-time transcription, speaker separation for 20+ voices, and endpoint detection that knows when someone has actually finished speaking. It costs $3 per 1,000 audio minutes—about $0.18 per hour.

One Model. Three Jobs. Zero Relay Race.

Traditional transcription systems are a relay race. One model does speech-to-text. Another guesses who is speaking. A third detects when a sentence ends. Each leg adds delay and its own mistakes. Muse Voice Transcribe collapses them into a single streaming pass.

The model processes audio in 80-millisecond chunks. It decides whether to output text now or wait for more context on harder words—a technique Meta calls "adaptive delay," trained through reinforcement learning. Simple words get transcribed almost instantly; complex ones get more audio before the model commits.

The result is a model that knows what it's hearing and who is saying it—without asking you to wait.

3.1% WER, 17.5% Diarization Error—but No Independent Verifier Yet

Meta claims Muse Voice Transcribe tops Artificial Analysis's streaming leaderboard with a final-transcript WER of 3.1%, arriving 0.16 seconds after speech ends. That compares favorably to Gemini 3.5 Transcribe Live's independently measured 4.0% WER. For speaker separation, Meta reports average 17.5% diarization error across standard benchmarks, compared to competitors' 21.1–28.6% range.

The honest label: The 3.1% number is Meta's launch-day claim, referencing a leaderboard where the model had been active for less than a day. Outside Meta, no one has independently reproduced it yet. The figure may hold—but it's a claim, not a verified result.

Google and Meta Just Split the Voice API Market in Two

Google shipped Gemini 3.5 Transcribe on August 26. Meta answered six days later with Muse Voice Transcribe. The structural difference between them is the real story.

Gemini 3.5 Transcribe splits into two endpoints: a cheaper, more accurate offline version (2.6% WER) and a more expensive, slightly less accurate live version (4.0% WER). Muse Voice Transcribe is one streaming model that does everything in a single pass. It doesn't offer an offline endpoint, but it does speaker identification for 20+ speakers. Gemini's live endpoint doesn't support speaker identification at all, and its offline version caps at 3 speakers (8 in experimental mode).

Meta's $0.18/hour pricing also undercuts Google's live rates. It's a direct appeal to developers building voice agents, live captioning, and real-time meeting transcripts.

Muse Voice Transcribe Just Raised the Bar for Real-Time Audio AI
Muse Voice Transcribe

OpenAI and Google have been racing on text and reasoning. Meta just entered the voice layer—not with a developer experiment, but with a production-ready model priced to capture the API market. The model is proprietary, not open-weight.

The voice API market now has three serious players: OpenAI, Google, and Meta. The immediate winner may be the one that can ship the lowest-latency, most accurate streaming model at scale. But the strategic winner will be the one that turns real-time audio into a platform—and that is exactly what Meta is trying to build.


P.S. Meta's edge is not just technical. The company already has built-in distribution: the Meta AI Mac app, Muse Code, and whatever comes next for smart glasses. A speech model that powers dictation across its own products can improve faster than one sold purely as an API. The model is the product now. The real market is the layer above it.


Frequently Asked Questions

Q: What is Meta's Muse Voice Transcribe?

A: Muse Voice Transcribe is Meta's first real-time audio perception model, released on September 1, 2026. It is a streaming speech-to-text model that transcribes audio in real time, separates speakers (up to 20+ voices), and detects natural speech endpoints in a single model pass.

Q: What are the key performance metrics?

A: Meta claims Muse Voice Transcribe tops Artificial Analysis's streaming leaderboard with a final-transcript word error rate of 3.1% and latency of 0.16 seconds after speech ends. For speaker separation, Meta reports a 17.5% diarization error rate, compared to competitors' 21.1–28.6% range. Both figures are Meta's launch-day claims and have not been independently verified.

Q: How does it compare to Google's Gemini 3.5 Transcribe?

A: Google shipped Gemini 3.5 Transcribe on August 26—six days before Meta's release. Google's model splits into two endpoints: a cheaper offline version (2.6% WER) and a more expensive live version (4.0% WER). Meta's model is a single streaming endpoint that does transcription, speaker separation (20+ speakers), and endpoint detection in one pass. Google's live endpoint does not support speaker separation, and its offline version caps at 3 speakers (8 in experimental mode).

Q: How much does Muse Voice Transcribe cost?

A: The API costs $3 per 1,000 audio minutes, or approximately $0.18 per hour.

Q: How does the "adaptive delay" work?

A: The model processes audio in 80-millisecond chunks and uses a technique called "adaptive delay" to decide whether to output text immediately or wait for more context on harder words. Simple words are transcribed almost instantly; complex words get more audio before the model commits. The mechanism was trained using reinforcement learning.

Q: What languages does it support?

A: The model was trained on over 70 languages, with 25 languages fully validated at launch.

Q: Where is Muse Voice Transcribe available?

A: The model is available through Meta Model API, the Meta AI Mac app (system-wide dictation by holding the Fn key), and Muse Code.

Q: Is Muse Voice Transcribe open-source?

A: No. The model is proprietary and available only through Meta's API.

Q: Does it work for long audio streams?

A: Yes. Meta says the model can run stably on audio streams exceeding one hour.

Q: What is the significance of this release for the voice AI market?

A: The voice API market now has three serious players: OpenAI, Google, and Meta. With Gemini 3.5 Transcribe shipping six days earlier, Meta and Google are now competing directly in the real-time audio layer. Meta's $0.18/hour pricing is a direct appeal to developers building voice agents, live captioning, and meeting transcripts.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article