Malaysian speech to text benchmark
Aisyah 1.0 Pro vs Alibaba Qwen for Malaysian speech to text
Word error rates for Aisyah 1.0 Pro and three Alibaba Qwen models on Revolab's benchmark of 1,899 Malaysian speech clips in 12 categories, with sample audio and each model's transcript.
- Aisyah 1.0 Pro
- Fewer mistakes overall, and on 11 of 12 kinds of audio. Biggest leads: short replies and phone calls.
- Alibaba Qwen-Audio 3.0 ASR Flash
- Fewer mistakes on 1 of 12: FLEURS.
Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested
- Aisyah 1.0 Pro5.79%
- Alibaba Qwen-Audio 3.0 ASR Flash9.17%
- Alibaba Qwen3-ASR-1.7B18.43%
- Alibaba Qwen3-ASR-0.6B24.19%
Where Aisyah and Qwen-Audio 3.0 ASR Flash each make fewer mistakes
Aisyah 1.0 ProAlibaba Qwen-Audio 3.0 ASR FlashWord error rate. Further left, fewer mistakes.
Aisyah 1.0 Pro makes fewer mistakes on 11 of 12
- Short repliesShort inputs, 155 clipsAisyah 1.0 Pro 3.28%Alibaba Qwen-Audio 3.0 ASR Flash 26.10%
- Phone callsTelephony, 229 clipsAisyah 1.0 Pro 9.55%Alibaba Qwen-Audio 3.0 ASR Flash 17.43%
- PodcastsPodcast, 155 clipsAisyah 1.0 Pro 4.57%Alibaba Qwen-Audio 3.0 ASR Flash 9.84%
- Drama156 clipsAisyah 1.0 Pro 5.91%Alibaba Qwen-Audio 3.0 ASR Flash 10.80%
- Street interviews143 clipsAisyah 1.0 Pro 14.24%Alibaba Qwen-Audio 3.0 ASR Flash 19.00%
- Singing152 clipsAisyah 1.0 Pro 11.01%Alibaba Qwen-Audio 3.0 ASR Flash 14.39%
- Animation150 clipsAisyah 1.0 Pro 5.66%Alibaba Qwen-Audio 3.0 ASR Flash 8.91%
- News156 clipsAisyah 1.0 Pro 3.48%Alibaba Qwen-Audio 3.0 ASR Flash 5.39%
- Parliament155 clipsAisyah 1.0 Pro 3.79%Alibaba Qwen-Audio 3.0 ASR Flash 5.29%
- Scripted readingRead speech, 154 clipsAisyah 1.0 Pro 1.30%Alibaba Qwen-Audio 3.0 ASR Flash 2.69%
- Common Voicevolunteer read-aloud, 141 clipsAisyah 1.0 Pro 3.74%Alibaba Qwen-Audio 3.0 ASR Flash 4.78%
Qwen-Audio 3.0 ASR Flash makes fewer mistakes on 1 of 12
- FLEURSread-aloud research set, 153 clipsAisyah 1.0 Pro 3.42%Alibaba Qwen-Audio 3.0 ASR Flash 3.29%
By background noise and clip set
The same 1,899 clips, split another way. The lower figure is in bold.
- Clean audio1,364 clipsAisyah 1.0 Pro 4.70%Alibaba Qwen-Audio 3.0 ASR Flash 8.13%
- Some background noise277 clipsAisyah 1.0 Pro 6.12%Alibaba Qwen-Audio 3.0 ASR Flash 8.68%
- Noisy audio249 clipsAisyah 1.0 Pro 12.16%Alibaba Qwen-Audio 3.0 ASR Flash 16.39%
- Published clips820 clips anyone can checkAisyah 1.0 Pro 4.89%Alibaba Qwen-Audio 3.0 ASR Flash 7.22%
- Held-back clips1,079 clips no model has seenAisyah 1.0 Pro 6.58%Alibaba Qwen-Audio 3.0 ASR Flash 10.87%
The other Alibaba models on this page
- Alibaba Qwen3-ASR-1.7B (18.43% overall) does not make fewer mistakes than Aisyah 1.0 Pro on any kind of audio, noise level or clip set.
- Alibaba Qwen3-ASR-0.6B (24.19% overall) does not make fewer mistakes than Aisyah 1.0 Pro on any kind of audio, noise level or clip set.
Every figure for every model is in the full results below.
Hear the difference
Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.
Credit card eligibility question
Phone calls, 1.9 s- What was said
- Credit card am I eligible?
- Aisyah 1.0 Pro
- Credit card am I eligible?
- Alibaba Qwen3-ASR-1.7B
- Credit, am I eligible?
- Alibaba Qwen3-ASR-0.6B
- Credit, am I eligible
- Alibaba Qwen-Audio 3.0 ASR Flash
- Credit card am I eligible.
What to listen for
The reference is "Credit card am I eligible?" Aisyah 1.0 Pro returns "Credit card am I eligible?" and Alibaba Qwen-Audio 3.0 ASR Flash returns "Credit card am I eligible." with the same words. Alibaba Qwen3-ASR-1.7B returns "Credit, am I eligible?" and Alibaba Qwen3-ASR-0.6B returns "Credit, am I eligible", both without the word "card".
Paying RM950 in Manglish
Phone calls, 2.9 s- What was said
- I dah bayar lah, RM950 ni.
- Aisyah 1.0 Pro
- I dah bayar lah, sembilan ratus lima puluh ringgit ni.
- Alibaba Qwen3-ASR-1.7B
- Air bayar lah sebanyak tu sembilan ringgit ni.
- Alibaba Qwen-Audio 3.0 ASR Flash
- I dah bayar lah sembilan ratus lima puluh ringgit ni.
What to listen for
The reference is "I dah bayar lah, RM950 ni." Aisyah 1.0 Pro returns "I dah bayar lah, sembilan ratus lima puluh ringgit ni." with the amount in words, and Qwen-Audio 3.0 ASR Flash returns the same words. Alibaba Qwen3-ASR-1.7B returns "Air bayar lah sebanyak tu sembilan ringgit ni.", where the amount appears as "sembilan ringgit".
Unit number 2105
Phone calls, 3.4 s- What was said
- Yes, my unit number is 2105.
- Aisyah 1.0 Pro
- Yes, my unit number is two one zero five.
- Alibaba Qwen3-ASR-1.7B
- Yes, my unit number is two one zero five.
- Alibaba Qwen3-ASR-0.6B
- Yes, my unit number is two one zero five.
What to listen for
The reference is "Yes, my unit number is 2105." Aisyah 1.0 Pro, Alibaba Qwen3-ASR-1.7B and Alibaba Qwen3-ASR-0.6B each return "Yes, my unit number is two one zero five." The benchmark scores "two one zero five" and "2105" as the same, so each transcript here matches the reference words.
Aisyah vs Alibaba Qwen, answered.
Is Aisyah more accurate than Alibaba Qwen for Malaysian speech?
On Revolab's Malaysian speech benchmark of 1,899 clips, Aisyah 1.0 Pro has a word error rate of 5.79%, against 9.17% for Alibaba Qwen-Audio 3.0 ASR Flash, 18.43% for Qwen3-ASR-1.7B and 24.19% for Qwen3-ASR-0.6B. Revolab built both the benchmark and Aisyah.
How accurate is Alibaba Qwen on Malaysian phone calls?
In the Telephony category of Revolab's Malaysian speech benchmark, Alibaba Qwen-Audio 3.0 ASR Flash has a word error rate of 17.43%, Qwen3-ASR-1.7B 30.69% and Qwen3-ASR-0.6B 30.23%. Aisyah 1.0 Pro has 9.55%.
Is Alibaba Qwen better than Aisyah for any type of Malaysian audio?
In one category on Revolab's Malaysian speech benchmark: Alibaba Qwen-Audio 3.0 ASR Flash has the lower word error rate in FLEURS, 3.29% against 3.42%. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B have a higher rate than Aisyah 1.0 Pro in all 12 categories.
Who built the benchmark behind the Aisyah vs Qwen comparison?
Revolab built both the Malaysian speech benchmark and Aisyah 1.0 Pro. Of its 1,899 clips, 820 are published and 1,079 are held back. On the published clips, Aisyah 1.0 Pro has a word error rate of 4.89% and Alibaba Qwen-Audio 3.0 ASR Flash has 7.22%.
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
All the numbersEvery figure for all 4 models on this page, as tables
Overall
| Model | All clips1,899 clips | Published820 clips | Held back1,079 clips |
|---|---|---|---|
| Aisyah 1.0 Pro | 5.79% (fewer mistakes) | 4.89% (fewer mistakes) | 6.58% (fewer mistakes) |
| Alibaba Qwen-Audio 3.0 ASR Flash | 9.17% | 7.22% | 10.87% |
| Alibaba Qwen3-ASR-1.7B | 18.43% | 15.26% | 21.20% |
| Alibaba Qwen3-ASR-0.6B | 24.19% | 20.73% | 27.21% |
By kind of audio
| Kind of audio | Aisyah 1.0 Pro | Qwen-Audio 3.0 ASR Flash | Qwen3-ASR-1.7B | Qwen3-ASR-0.6B |
|---|---|---|---|---|
| Phone callsTelephony, 229 clips | 9.55% (fewer mistakes) | 17.43% | 30.69% | 30.23% |
| Short repliesShort inputs, 155 clips | 3.28% (fewer mistakes) | 26.10% | 28.10% | 23.08% |
| Scripted readingRead speech, 154 clips | 1.30% (fewer mistakes) | 2.69% | 6.76% | 12.31% |
| PodcastsPodcast, 155 clips | 4.57% (fewer mistakes) | 9.84% | 17.14% | 24.10% |
| Drama156 clips | 5.91% (fewer mistakes) | 10.80% | 27.69% | 37.87% |
| Animation150 clips | 5.66% (fewer mistakes) | 8.91% | 19.58% | 24.73% |
| News156 clips | 3.48% (fewer mistakes) | 5.39% | 16.49% | 18.56% |
| Parliament155 clips | 3.79% (fewer mistakes) | 5.29% | 13.89% | 19.57% |
| Street interviews143 clips | 14.24% (fewer mistakes) | 19.00% | 34.23% | 43.13% |
| Singing152 clips | 11.01% (fewer mistakes) | 14.39% | 18.96% | 31.94% |
| Common Voicevolunteer read-aloud, 141 clips | 3.74% (fewer mistakes) | 4.78% | 8.42% | 12.96% |
| FLEURSread-aloud research set, 153 clips | 3.42% | 3.29% (fewer mistakes) | 7.65% | 14.88% |
By background noise
| Background | Aisyah 1.0 Pro | Qwen-Audio 3.0 ASR Flash | Qwen3-ASR-1.7B | Qwen3-ASR-0.6B |
|---|---|---|---|---|
| Clean audio1,364 clips | 4.70% (fewer mistakes) | 8.13% | 17.18% | 21.91% |
| Some background noise277 clips | 6.12% (fewer mistakes) | 8.68% | 18.23% | 25.37% |
| Noisy audio249 clips | 12.16% (fewer mistakes) | 16.39% | 26.76% | 36.75% |
Types of mistake
| Model | Wrong wordsper 100 words | Invented wordsper 100 words | Dropped wordsper 100 words |
|---|---|---|---|
| Aisyah 1.0 Pro | 3.48% (fewer mistakes) | 0.62% (fewer mistakes) | 1.69% (fewer mistakes) |
| Alibaba Qwen-Audio 3.0 ASR Flash | 5.51% | 1.11% | 2.55% |
| Alibaba Qwen3-ASR-1.7B | 13.22% | 1.09% | 4.13% |
| Alibaba Qwen3-ASR-0.6B | 17.86% | 1.25% | 5.09% |
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Compare Aisyah with other models
See every model in one table- Aisyah vs ElevenLabs ScribeScribe v25.79% vs 6.67%
- Aisyah vs Google GeminiGemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.6 Flash5.79% vs 7.00%
- Aisyah vs YTL ILMUILMU ASR v4.25.79% vs 9.48%
- Aisyah vs OpenAI WhisperWhisper large-v35.79% vs 17.30%
- Aisyah vs AssemblyAI UniversalUniversal 25.79% vs 18.80%
- Aisyah vs Deepgram NovaNova 35.79% vs 31.99%
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.