- Speech To Text
- English Mandarin Malay Tamil
English-mandarin-malay-tamil speech to text transcription API
Convert English-mandarin-malay-tamil voice into accurate text in seconds. Whether you need English-mandarin-malay-tamil speech to text for real-time applications, voice recordings, or multilingual content, our transcription API delivers fast, secure, and accurate results. Trusted for English-mandarin-malay-tamil voice to text and transcription use cases, integrate high-quality English-mandarin-malay-tamil ASR into your product.
- •High-accuracy transcription of standard English-mandarin-malay-tamil and dialects
- •Supports real-time and batch processing
- •Easy to integrate with our developer-friendly API
- •Built for global enterprise scale, with secure and private processing.
- High-accuracy transcription of standard English-mandarin-malay-tamil and dialects
- Supports real-time and batch processing
- Easy to integrate with our developer-friendly API
- Built for global enterprise scale, with secure and private processing.
English-mandarin-malay-tamil transcription accuracy
Understands every accent We’re trained for variations of dialects and accents. Get accurate transcriptions, no matter the region. Ready for real-time scale High-volume? No problem. Our API handles live recorded and live audio at scale – with secure cloud, on-prem or on-device deployment options. Built for the real world Noisy calls, fast speakers, crosstalk – our tech thrives in messy audio so you get clarity, not compromise. Experience English-mandarin-malay-tamil transcription that works
Try our live English-mandarin-malay-tamil transcription for yourself
Speak into your mic and watch real-time English-mandarin-malay-tamil transcription in action. Fast, accurate, and built for natural conversations.
Four languages, one pack
Combined, the four official languages of Singapore cover the everyday speech of over 50 million people in Singapore, Malaysia, Brunei, and the diaspora across Southeast Asia and beyond.
English (en). Includes Singaporean English and Singlish.
Mandarin (cmn). Accents from China, Taiwan, Singapore, and Malaysia.
Malay (ms). Includes Bahasa Melayu spoken in Singapore and Malaysia.
Tamil (ta). Includes Singapore Tamil and Tamil across the diaspora.

Everything you need for accurate, scalable English-mandarin-malay-tamil speech to text.
Built for real-world use cases and global applications.Everything you need for accurate, scalable English-mandarin-malay-tamil speech to text.
AI speech to text transcription in 53+ languages
Frequently Asked Questions — English, Mandarin, Malay & Tamil
What is English, Mandarin, Malay and Tamil speech to text?
What is English, Mandarin, Malay and Tamil speech to text?
The English, Mandarin, Malay and Tamil multilingual pack, language code cmn_en_ms_ta, is a single speech recognition model that transcribes audio containing any mix of the four languages. It handles code-switching mid-sentence as part of natural speech, unlike general-purpose multilingual ASR, which tends to treat each language as separate.
The pack was built for Southeast Asia, where the four languages are the official languages of Singapore and are spoken across Malaysia, Brunei, and the diaspora. It powers contact center transcription, voice agents, broadcast captioning, and government workflows across the region.
Read more about how the models were built in The real language of business.
Does it handle Singlish and code-switching?
Does it handle Singlish and code-switching?
Yes. Singlish and code-switching are the default case, not an exception.
Singlish, the English-based creole spoken across Singapore, blends English grammar with Malay, Hokkien, Cantonese, and Tamil vocabulary. Code-switching, the practice of moving between languages within one sentence, is how millions of people in Singapore and Malaysia actually speak. The multilingual pack was trained on conversational audio that reflects this. It transcribes what was said, without breaking at the switch.
How does it compare to general-purpose multilingual ASR?
How does it compare to general-purpose multilingual ASR?
General-purpose multilingual ASR supports a long list of languages individually. Each one is handled by a separate model. When a speaker switches languages mid-sentence, most of these systems drop words, produce broken output, or force the developer to make multiple API calls per audio file.
The Speechmatics multilingual pack uses one model and one API call for all four languages, including the switch points. Fewer errors on Singaporean English. Fewer errors on code-switched audio. One transcript, in order, per request.
What's the difference between the Enhanced and Standard operating points?
What's the difference between the Enhanced and Standard operating points?
The English, Mandarin, Malay and Tamil pack runs on two operating points. Enhanced returns the lowest word error rate. It's the default choice for compliance, quality monitoring, and other accuracy-critical work. Standard trades a small amount of accuracy for faster throughput, useful for high-volume batch jobs or cost-sensitive workloads.
Both operating points handle the same four-language code-switching. Set operating_point to enhanced or standard in your transcription config.
How does real-time multilingual transcription work?
How does real-time multilingual transcription work?
Real-time transcription streams audio to the API over a WebSocket connection and returns transcript results as the audio arrives. Partial transcripts land in under a second. Final transcripts follow within two seconds.
Speechmatics supports real-time transcription for the full 53+ language range, including the English, Mandarin, Malay and Tamil pack. The system is built for spontaneous speech, interruptions, background noise, and language switching. For pre-recorded audio and video, batch transcription runs the same model at higher throughput.
Read the real-time overview or the real-time quickstart in docs.
What can the English, Mandarin, Malay and Tamil speech to text API do?
What can the English, Mandarin, Malay and Tamil speech to text API do?
The API integrates multilingual transcription into applications, platforms, and internal systems.
You can:
Transcribe audio and video files programmatically across all four languages.
Stream live audio for real-time transcription with code-switching support.
Return structured transcripts with timestamps and speaker identification.
Prepare text for analytics, subtitling, and downstream workflows.
Add custom dictionary entries for local terms, place names, and technical vocabulary.
The API is production-grade and supports cloud, hybrid, on-premise, and on-device deployment.
What are the main use cases?
What are the main use cases?
The English, Mandarin, Malay and Tamil pack is used across:
Customer interaction analysis and quality monitoring in contact center solutions.
Voice automation in AI voice agents.
Subtitle creation and accessibility in media distribution and captioning.
Collaboration and discussion capture in meeting platforms.
Brand and news monitoring in media monitoring.
Organizations with data residency or scale requirements deploy through enterprise speech recognition.
How do I transcribe a multilingual video to text?
How do I transcribe a multilingual video to text?
Upload the file to the Speechmatics portal or send it through the Batch API. Set the language to cmn_en_ms_ta. The system returns a transcript with timestamps and speaker labels. Export as text, JSON, or SRT.
Do you offer a free multilingual speech to text trial?
Do you offer a free multilingual speech to text trial?
Yes. Create an account and you get eight hours of free transcription every month, across every supported language, including the English, Mandarin, Malay and Tamil pack. See pricing for volume rates.
Can I deploy it on-premise?
Can I deploy it on-premise?
Yes. The pack runs on CPU and GPU containers, Kubernetes, and air-gapped environments. Same model, same accuracy, your infrastructure. Common for organizations with strict data residency requirements across Singapore, Malaysia, and the wider region.
How accurate is the model?
How accurate is the model?
Can speech to text handle noisy audio?
Can speech to text handle noisy audio?
Yes. The model is trained on real-world audio, including phone-quality audio at 8 kHz, background noise, and cross-talk typical of contact center recordings.
What audio formats are supported?
What audio formats are supported?
WAV, MP3, AAC, OGG, MPEG, AMR, M4A, MP4, and FLAC.
What other multilingual options does Speechmatics offer?
What other multilingual options does Speechmatics offer?
Speechmatics supports multilingual audio in two ways.
Specialist bilingual and multilingual language packs — These cover a fixed set of languages that you select in advance. In addition to the four-language cmn_en_ms_ta pack, the pack range includes:
Mandarin and English (cmn_en)
Malay and English (en_ms)
Tamil and English (en_ta)
Arabic and English (ar_en)
Spanish and English
Melia — Melia is Speechmatics' broader multilingual model. It handles code-switching across all 53+ supported languages in a single model, without requiring you to select or manage individual packs. Melia is in production preview for Batch, with real-time support on the roadmap.
See the full language coverage in docs or the Melia model docs.
What industries commonly use multilingual speech to text in Southeast Asia?
What industries commonly use multilingual speech to text in Southeast Asia?
Contact centers, voice agent platforms, government and public-sector workflows, media and broadcasting, meeting and collaboration tools, and accessibility workflows.
![[alt: Industry-leading transcription accuracy in 55+ languages]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2F1dGuTnCrsPeC1XuiYZHdJx%2F854dfedb68eee0749d5b5f2521030fd6%2F9e3ae9aeb3cd6c9da26f9068fe1a29ce1098b1f9.png&w=3840&q=75)