Malaysian speech to text benchmark
Compare speech to text models on Malaysian audio
13 speech to text models, run on the same 1,899 Malaysian utterances: phone calls, podcasts, parliament, street interviews, singing and more. Same audio, same scoring. Lower word error rate is better.
Revolab built this benchmark and two of the models in it. Results as of 12 August 2026. How we tested
Word error rate on 1,899 Malaysian clips.
Ranked by word error rate across all 1,899 clips: mistakes per 100 words spoken, so lower is better. Use the compare button on any row to see that model against Aisyah, kind of audio by kind of audio.
A second AssemblyAI run, universal-3-5-pro, is not listed: its outputs were identical to Universal 2 on every clip.
Aisyah 1.0 Pro against each vendor
Category by category, including where the other model scores lower, with real clips you can play.
- Aisyah vs ElevenLabs ScribeScribe v25.79% vs 6.67%
- Aisyah vs Google GeminiGemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.6 Flash5.79% vs 7.00%
- Aisyah vs YTL ILMUILMU ASR v4.25.79% vs 9.48%
- Aisyah vs OpenAI WhisperWhisper large-v35.79% vs 17.30%
- Aisyah vs Alibaba QwenQwen-Audio 3.0 ASR Flash, Qwen3-ASR-1.7B, Qwen3-ASR-0.6B5.79% vs 9.17%
- Aisyah vs AssemblyAI UniversalUniversal 25.79% vs 18.80%
- Aisyah vs Deepgram NovaNova 35.79% vs 31.99%
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Malaysian speech to text benchmark, answered.
Which speech to text model is most accurate on Malaysian audio?
On this benchmark, Aisyah 1.0 Pro has the lowest word error rate at 5.79% across 1,899 clips, followed by ElevenLabs Scribe v2 at 6.67% and Google Gemini 2.5 Pro at 7.00%. Revolab built both the benchmark and Aisyah.
Is Aisyah 1.0 Pro the most accurate in every category?
No. Aisyah 1.0 Pro has the lowest word error rate in 6 of 12 categories. In News, Parliament, Street interviews, Singing, Common Voice and FLEURS, another model scores lower, and each comparison page shows which.
How is word error rate calculated in this benchmark?
Every wrong, invented and dropped word is counted and divided by the number of words actually said, across all clips. Transcripts are normalised first, so casing, punctuation and digits versus words do not count as errors. Lower is better.
Can I check these results myself?
Yes. The 820 published clips, with the audio and every model's output, are in the benchmark explorer. The other 1,079 clips are held back.
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.