Meta has launched Muse Voice Transcribe, a new real-time AI speech model designed to transcribe conversations as they happen, identify different speakers and detect when someone has finished speaking. For Indian users, the launch is particularly notable because the model supports Hindi, Tamil, Telugu, Kannada and Malayalam as part of its broader multilingual capabilities.

Developed by Meta Superintelligence Labs, Muse Voice Transcribe is built to handle not only live speech but also multilingual conversations where speakers switch languages mid-sentence. Meta says the model has been trained across more than 70 languages, with 25 languages extensively validated for the initial release.

What is Meta Muse Voice Transcribe?

Muse Voice Transcribe is Meta's first real-time audio perception model. Rather than simply converting a completed audio recording into text, it processes incoming speech continuously and produces a transcript while the conversation is taking place.

The model combines three capabilities in one system: streaming automatic speech recognition (ASR), speaker diarization and endpointing. In simple terms, it can understand what is being said, work out who is speaking and identify when a person has stopped talking.

That combination could make the technology useful for voice assistants, dictation, meetings, coding tools and other applications where waiting for an entire recording to finish is inconvenient.

Which Indian languages does Muse Voice Transcribe support?

Muse Voice Transcribe supports five major Indian languages in addition to its wider global language coverage:

Indian languageMuse Voice Transcribe support
HindiYes
TamilYes
TeluguYes
KannadaYes
MalayalamYes

Meta says the model has been trained on 70+ languages, while 25 have been extensively verified for the initial release. The five Indian languages above are among the languages reported as supported at launch.

For India, the more interesting feature may not simply be the number of languages. It is the model's ability to deal with code-switching, such as when someone moves between English and an Indian language during the same conversation.

Can Muse Voice Transcribe understand Hinglish and other mixed-language conversations?

Yes, code-switching is one of the model's key capabilities.

Meta says Muse Voice Transcribe can natively handle code-switching both within a sentence and between sentences. That means users do not necessarily have to manually change the selected language every time they switch between languages.

This is particularly relevant in India, where conversations frequently mix English with Hindi or regional languages. Meta's own research demonstration shows the model handling mixed-language speech and using context to improve recognition.

How does Muse Voice Transcribe work in real time?

Muse Voice Transcribe is built as an autoregressive multimodal model from Meta's Muse Spark family.

According to Meta's technical explanation, audio is processed in 80-millisecond chunks. The model decides whether it needs to listen to more audio or whether it has enough context to generate the next text token.

This is important because real-time transcription has a basic trade-off: waiting longer can improve accuracy, but it also creates more delay.

Meta addresses this using what it calls adaptive delay. The model can wait longer for a difficult word while committing easier words more quickly, with reinforcement learning used to balance transcription accuracy against latency.

Can Muse Voice Transcribe identify different speakers?

Yes. Speaker diarization is built directly into the model.

Meta says Muse Voice Transcribe can handle 20+ speakers and long audio exceeding one hour without requiring separate post-processing. Its demonstrations show the system identifying different speakers during an extended conversation.

That could be useful for meetings, interviews, podcasts, lectures and group discussions where a simple block of transcribed text is not enough.

Instead of producing one continuous transcript, the system can associate sections of speech with different speakers.

How long can Muse Voice Transcribe audio recordings be?

Meta says the model natively supports audio longer than one hour and can handle 20+ speakers without requiring additional post-processing.

This makes the technology potentially useful beyond short voice commands. Long meetings, interviews and conversations are examples where speaker identification and continuous transcription could be particularly useful.

Where is Meta Muse Voice Transcribe available?

Muse Voice Transcribe is available through the Meta Model API, allowing developers to integrate its speech capabilities into applications. Meta also says it is already being used for dictation in Meta AI for Mac and Muse Code.

For Mac users, Meta's implementation allows voice dictation across applications. Developers can instead access the model through Meta's API and build their own speech-based experiences.

How much does Muse Voice Transcribe cost?

According to current reporting on Meta's Model API, Muse Voice Transcribe is priced at $3 per 1,000 audio minutes, which works out to approximately $0.18 per hour of audio.

DetailMuse Voice Transcribe
API price$3 per 1,000 audio minutes
Approx. hourly cost$0.18
AvailabilityMeta Model API
Meta products using itMeta AI for Mac, Muse Code

Pricing and API availability can change, so developers should verify the current rate before building a production workflow around it.

How accurate is Meta Muse Voice Transcribe?

Meta says Muse Voice Transcribe ranked first on Artificial Analysis' streaming speech-to-text leaderboard based on data dated September 1, 2026. Meta also says it ranks first on public diarization benchmarks.

Reporting on the benchmark puts the model's final-transcription word error rate at 3.1% in that particular evaluation. However, benchmark results should not be treated as a guarantee of the same performance across every Indian accent, microphone, background-noise condition or conversational setting.

For Indian-language users, independent testing across regional accents and code-switched conversations will be particularly useful.

What can Muse Voice Transcribe be used for?

The technology has applications well beyond basic transcription.

  1. Live meeting transcription: Convert conversations into text while participants are speaking.
  2. Interviews and podcasts: Separate speakers automatically during long recordings.
  3. Voice assistants: Give AI systems a faster way to understand ongoing conversations.
  4. Dictation: Convert speech into text across desktop applications.
  5. Coding: Allow developers to communicate with coding tools through voice.
  6. Multilingual communication: Handle conversations that switch between languages.
  7. Long recordings: Process conversations lasting more than an hour.

Meta's current integrations with Meta AI for Mac and Muse Code already demonstrate some of these use cases.

Why is the launch important for India?

India is one of the world's most linguistically diverse technology markets, so speech AI that can handle regional languages and code-switching has an obvious use case.

The five-language launch does not mean Muse Voice Transcribe supports every major Indian language. Bengali, Marathi, Gujarati, Punjabi and other Indian languages should not be presented as confirmed support unless Meta adds them to the validated list.

What makes this launch more interesting is the combination of Indian-language recognition with real-time processing and multilingual code-switching. For users who naturally mix English with Hindi or another regional language, that could be more useful than a transcription system that expects speakers to remain in one language.

Is Muse Voice Transcribe available to everyone?

Not in the same way.

Developers can access the model through Meta's Model API, while Meta says Muse Voice Transcribe is already powering dictation in Meta AI for Mac and Muse Code. Availability of individual features can depend on the product, platform and account.

So, Muse Voice Transcribe is not simply a new standalone consumer transcription app from Meta. It is primarily a model that is being integrated into Meta products and made available to developers.

Meta Muse Voice Transcribe vs traditional transcription tools: What's different?

The main difference is that Meta is combining several audio tasks inside one real-time model.

FeatureMuse Voice Transcribe
Real-time transcriptionYes
Speaker identificationYes
Endpoint detectionYes
Code-switchingYes
70+ languages trainedYes
25 languages extensively validatedYes
20+ speakersYes
1+ hour audioYes
API access                                                Yes

Meta's architecture is designed so that streaming ASR, diarization and endpointing work together rather than requiring separate post-processing stages.

What does Meta Muse Voice Transcribe mean for Indian users?

The immediate impact will likely be seen in voice-based AI tools, dictation and transcription services, rather than as a new standalone chatbot.

Its support for Hindi, Tamil, Telugu, Kannada and Malayalam gives developers more options for building voice experiences aimed at Indian users. The ability to switch languages during a conversation could be especially important because real-world speech does not always follow the clean language boundaries used by many AI systems.

For Meta, the launch is also another step toward its broader vision of AI systems that can understand natural conversations rather than responding only to carefully structured voice commands.

FAQ

What is Meta Muse Voice Transcribe?

Muse Voice Transcribe is Meta's first real-time audio perception model. It combines live speech transcription, speaker identification and endpoint detection in one system.

Which five Indian languages does Muse Voice Transcribe support?

The five Indian languages reported for the launch are Hindi, Tamil, Telugu, Kannada and Malayalam.

Does Muse Voice Transcribe support code-switching?

Yes. Meta says the model can handle multilingual code-switching within and between sentences, allowing speakers to switch languages without manually changing the language setting.

How much does Muse Voice Transcribe cost?

The Meta Model API is currently reported at $3 per 1,000 audio minutes, or approximately $0.18 per hour.

Is Muse Voice Transcribe available in India?

The model is available through Meta's Model API, while Meta says it is already used in Meta AI for Mac and Muse Code. Product-level availability can vary by platform and account.

Our Take

Muse Voice Transcribe is more interesting than a simple speech-to-text upgrade. Its combination of real-time transcription, speaker identification and code-switching addresses some of the messiest parts of real-world conversations.

For India, the five-language support is the headline, but the bigger opportunity is multilingual voice AI that understands how people actually speak. The next test will be how reliably it handles regional accents, noisy environments and everyday English-plus-Indian-language conversations outside controlled benchmarks.