[{"data":1,"prerenderedAt":28},["ShallowReactive",2],{"nr-en-meta-muse-voice-transcribe-echtzeit-sprachmodell":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":10,"trackerLabel":10,"headlineStat":22,"image":23,"ogImage":24,"imageAlt":5,"csv":10,"minutes":25,"words":26,"html":27},"meta-muse-voice-transcribe-echtzeit-sprachmodell","Meta Launches Muse Voice Transcribe: Real-Time Speech Model for Always-Listening AI Assistants","Meta's Superintelligence Labs have released a streaming transcription model that processes speech in real time, separates up to 20 speakers, and costs just $0.18 per hour. The technology is positioned as the foundation for personal AI agents—including on camera glasses.","2026-09-06","13:03","2026-09-06T13:03:00+02:00","","September 6, 2026","news","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"Speech Recognition","AI Models","Meta","Real-Time Processing","Data Privacy","$0.18 per hour","\u002Fnewsroom\u002Fimg\u002Fmeta-muse-voice-transcribe-echtzeit-sprachmodell.webp","\u002Fog-nr\u002Fmeta-muse-voice-transcribe-echtzeit-sprachmodell.en.png",2,475,"\u003Cp>Meta has unveiled \u003Cstrong>Muse Voice Transcribe\u003C\u002Fstrong>, a real-time speech model that transcribes audio, detects speaker changes, and marks sentence endings—all during live conversations and without separate systems for each task. The model breaks incoming audio into \u003Cstrong>80-millisecond chunks\u003C\u002Fstrong> and decides after each fragment whether to output the next word or listen longer for more context.\u003C\u002Fp>\n\u003Ch2>Quick Facts\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Price:\u003C\u002Fstrong> $0.18 per hour—significantly below competitors like OpenAI or ElevenLabs\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Language Support:\u003C\u002Fstrong> Trained on over 70 languages, with 25 thoroughly validated; supports code-switching between languages within a single sentence\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Speaker Separation:\u003C\u002Fstrong> Distinguishes up to \u003Cstrong>20 speakers simultaneously\u003C\u002Fstrong> and assigns utterances to individual speakers in real time\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Availability:\u003C\u002Fstrong> Now available via Meta AI and its own API\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>The Trick: Adaptive Latency Per Word\u003C\u002Fh2>\n\u003Cp>The core technical innovation lies in \u003Cstrong>dynamic wait time per word\u003C\u002Fstrong>. Rather than accepting fixed latency, the model adjusts for each word individually how long to listen. Simple words are output faster; difficult ones get more time. This behavior was trained using \u003Cstrong>Reinforcement Learning\u003C\u002Fstrong>—the model is rewarded for achieving both low error rates and short delays simultaneously.\u003C\u002Fp>\n\u003Cp>This solves a classic trade-off: the longer a speech model waits, the more accurate the transcription becomes. But latency increases. Muse Voice Transcribe shifts this compromise by deciding per word.\u003C\u002Fp>\n\u003Ch2>Everything From One Model\u003C\u002Fh2>\n\u003Cp>Meta builds additional functions on this foundation—without training separate systems. \u003Cstrong>Speaker attribution\u003C\u002Fstrong> marks where switches occur in the running text and tags each passage with an identifier (A–Z). \u003Cstrong>Sentence detection\u003C\u002Fstrong> marks the beginning and end of utterances. Both tasks are trained jointly with speech recognition.\u003C\u002Fp>\n\u003Cp>According to Meta, the model processes audio recordings of \u003Cstrong>over one hour without post-processing\u003C\u002Fstrong>. In a demo with eight people in a room, the system assigned utterances to individual speakers in real time.\u003C\u002Fp>\n\u003Ch2>Positioning: The Foundation for Always-Listening Assistants\u003C\u002Fh2>\n\u003Cp>Meta explicitly positions Muse Voice Transcribe as the foundation for \u003Cstrong>personal AI agents\u003C\u002Fstrong> that listen in on real conversations via controversial camera glasses. With 80-millisecond latency and the ability to separate multiple speakers in real time, a technical basis emerges for AI assistants that continuously listen in the background—without requiring constant user commands.\u003C\u002Fp>\n\u003Cp>In the \u003Cstrong>Artificial Analysis benchmark from September 1, 2026\u003C\u002Fstrong>, Muse Voice Transcribe achieved the highest accuracy at the lowest price among tested models.\u003C\u002Fp>\n\u003Ch2>What This Means for European Enterprises\u003C\u002Fh2>\n\u003Cp>The release signals that \u003Cstrong>real-time speech recognition with low latency is becoming a commodity\u003C\u002Fstrong>. European companies relying on voice interfaces, call center automation, or real-time transcription now face aggressive pricing competition. At the same time, the model raises \u003Cstrong>data privacy questions\u003C\u002Fstrong>: if AI assistants are to listen continuously, new requirements emerge around transparency, storage, and user control—issues the EU AI Act and GDPR must address.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.de\u002Fmetas-neues-echtzeit-sprachmodell-ist-der-unterbau-fuer-dauerhaft-mithoerende-ki-assistenten\u002F\">The Decoder (DE)\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Fmetas-new-real-time-audio-model-is-the-foundation-for-ai-assistants-that-never-stop-listening\u002F\">The Decoder\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1788692972418]