- By Prateek Levi
- Wed, 02 Sep 2026 07:59 PM (IST)
- Source:JND
Meta has introduced Muse Voice Transcribe, and this comes with support for five major Indian languages, which include Hindi, Telugu, Tamil, Kannada, and Malayalam; globally, the platform supports more than 70 languages. This has been developed by Meta Superintelligence Labs and is the first real-time audio perception model.
This model supports multiple tools in one, like endpointing, speech transcription and speaker separation. Apart from that, this model also supports conversations and code-switching, boasting more than a score of 20 speakers at a time.
Real-Time Transcription With Meta Muse Voice Transcript
In a blog post, the company explained that this model comes with the ability to turn voice/speech into text in real time as the person speaks, and not processing it afterwards. Another feature here is that it can also detect individual speakers and also has a sense of things like when someone has started or stopped talking.
Meta’s new speech AI model is built to handle transcription as the conversation happens, rather than waiting for the recording to end and processing everything afterwards. Transcription, speaker separation and endpointing all happen in real time as the audio comes in.
It can also keep up with fairly long recordings. Meta says the model can identify more than 20 speakers in a single recording and work with audio that is over an hour long. That should make it useful for things like long meetings, interviews and group discussions, where cutting the recording into smaller sections can be a hassle.
The model processes audio in 80-millisecond chunks and decides when it has enough information to write down each word. For words that are easy to recognise, it can respond quickly. When something is less clear, it can spend a little more time before committing to the transcription. Meta says reinforcement learning helps the model make these decisions without letting transcription accuracy take a hit.
Language support is another area Meta has focused on. The model was trained on more than 70 languages, while 25 have been extensively validated for the initial release. Among them are five Indian languages: Hindi, Tamil, Telugu, Kannada and Malayalam.
It can also handle conversations where people move between languages. So, if someone switches languages halfway through a sentence, there is no need to manually change the language setting each time. The model can follow the switch on its own.
Meta has also included language, keyword and context biasing to help the system make better sense of what it is hearing. In simple terms, it can use the surrounding conversation and relevant keywords to work out which words are most likely being spoken.
That combination could make the model particularly useful for real-world conversations, where people interrupt each other, switch languages and speak for much longer than a typical short transcription clip.
Pricing And Availability of the Meta Muse Voice Transcribe
This new model is available through the Meta Model API. This is for ease of access for developers so that they can use this model for speech transcription in their own applications. This API will cost $3, which roughly means ₹300 for 1,000 audio minutes, which means that the costing per hour is $0.18 which is almost ₹17.
