Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?
Meta Platforms has entered the speech-to-text market with the launch of Muse Voice Transcribe, an advanced audio perception model designed by Meta Superintelligence Labs. The new model integrates real-time streaming transcription, endpoint detection, and speaker diarization capable of identifying more than 20 speakers. Meta has positioned the API aggressively on pricing at $0.18 per hour of processed audio, matching zero-data-retention processing at the same rate and applying down-to-the-second billing. The model features an autoregressive multimodal architecture processing audio in 80-millisecond chunks with adaptive delay driven by reinforcement learning. In benchmark evaluations, Muse achieved a leading 3.1% final-transcription word error rate (WER) on the Artificial Analysis AA-WER Streaming Index, outperforming competitors like Cartesia Ink-2, ElevenLabs Scribe v2, and OpenAI's GPT Live Transcribe. It also posted a 17.5% average diarization error rate across standard test datasets. Meta's aggressive pricing and built-in diarization directly undercut competitors such as Amazon Transcribe, Speechmatics, and Deepgram, intensifying price and performance competition in the enterprise voice AI infrastructure market.