Malaysian speech to text benchmark

Aisyah 1.0 Pro vs Alibaba Qwen for Malaysian speech to text

Word error rates for Aisyah 1.0 Pro and three Alibaba Qwen models on Revolab's benchmark of 1,899 Malaysian speech clips in 12 categories, with sample audio and each model's transcript.

Aisyah 1.0 Pro
Fewer mistakes overall, and on 11 of 12 kinds of audio. Biggest leads: short replies and phone calls.
Alibaba Qwen-Audio 3.0 ASR Flash
Fewer mistakes on 1 of 12: FLEURS.

Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested

Word error rate on all 1,899 clipsMistakes per 100 words spoken. Shorter bar, fewer mistakes.
  • Aisyah 1.0 Pro5.79%
  • Alibaba Qwen-Audio 3.0 ASR Flash9.17%
  • Alibaba Qwen3-ASR-1.7B18.43%
  • Alibaba Qwen3-ASR-0.6B24.19%

Where Aisyah and Qwen-Audio 3.0 ASR Flash each make fewer mistakes

Aisyah 1.0 ProAlibaba Qwen-Audio 3.0 ASR FlashWord error rate. Further left, fewer mistakes.

Aisyah 1.0 Pro makes fewer mistakes on 11 of 12

  • Short repliesShort inputs, 155 clips
    Aisyah 1.0 Pro 3.28%Alibaba Qwen-Audio 3.0 ASR Flash 26.10%
  • Phone callsTelephony, 229 clips
    Aisyah 1.0 Pro 9.55%Alibaba Qwen-Audio 3.0 ASR Flash 17.43%
  • PodcastsPodcast, 155 clips
    Aisyah 1.0 Pro 4.57%Alibaba Qwen-Audio 3.0 ASR Flash 9.84%
  • Drama156 clips
    Aisyah 1.0 Pro 5.91%Alibaba Qwen-Audio 3.0 ASR Flash 10.80%
  • Street interviews143 clips
    Aisyah 1.0 Pro 14.24%Alibaba Qwen-Audio 3.0 ASR Flash 19.00%
  • Singing152 clips
    Aisyah 1.0 Pro 11.01%Alibaba Qwen-Audio 3.0 ASR Flash 14.39%
  • Animation150 clips
    Aisyah 1.0 Pro 5.66%Alibaba Qwen-Audio 3.0 ASR Flash 8.91%
  • News156 clips
    Aisyah 1.0 Pro 3.48%Alibaba Qwen-Audio 3.0 ASR Flash 5.39%
  • Parliament155 clips
    Aisyah 1.0 Pro 3.79%Alibaba Qwen-Audio 3.0 ASR Flash 5.29%
  • Scripted readingRead speech, 154 clips
    Aisyah 1.0 Pro 1.30%Alibaba Qwen-Audio 3.0 ASR Flash 2.69%
  • Common Voicevolunteer read-aloud, 141 clips
    Aisyah 1.0 Pro 3.74%Alibaba Qwen-Audio 3.0 ASR Flash 4.78%

Qwen-Audio 3.0 ASR Flash makes fewer mistakes on 1 of 12

  • FLEURSread-aloud research set, 153 clips
    Aisyah 1.0 Pro 3.42%Alibaba Qwen-Audio 3.0 ASR Flash 3.29%

By background noise and clip set

The same 1,899 clips, split another way. The lower figure is in bold.

  • Clean audio1,364 clips
    Aisyah 1.0 Pro 4.70%Alibaba Qwen-Audio 3.0 ASR Flash 8.13%
  • Some background noise277 clips
    Aisyah 1.0 Pro 6.12%Alibaba Qwen-Audio 3.0 ASR Flash 8.68%
  • Noisy audio249 clips
    Aisyah 1.0 Pro 12.16%Alibaba Qwen-Audio 3.0 ASR Flash 16.39%
  • Published clips820 clips anyone can check
    Aisyah 1.0 Pro 4.89%Alibaba Qwen-Audio 3.0 ASR Flash 7.22%
  • Held-back clips1,079 clips no model has seen
    Aisyah 1.0 Pro 6.58%Alibaba Qwen-Audio 3.0 ASR Flash 10.87%

The other Alibaba models on this page

  • Alibaba Qwen3-ASR-1.7B (18.43% overall) does not make fewer mistakes than Aisyah 1.0 Pro on any kind of audio, noise level or clip set.
  • Alibaba Qwen3-ASR-0.6B (24.19% overall) does not make fewer mistakes than Aisyah 1.0 Pro on any kind of audio, noise level or clip set.

Every figure for every model is in the full results below.

Hear the difference

Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.

Credit card eligibility question

Phone calls, 1.9 s
What was said
Credit card am I eligible?
Aisyah 1.0 Pro
Credit card am I eligible?
Alibaba Qwen3-ASR-1.7B
Credit, am I eligible?
Alibaba Qwen3-ASR-0.6B
Credit, am I eligible
Alibaba Qwen-Audio 3.0 ASR Flash
Credit card am I eligible.
What to listen for

The reference is "Credit card am I eligible?" Aisyah 1.0 Pro returns "Credit card am I eligible?" and Alibaba Qwen-Audio 3.0 ASR Flash returns "Credit card am I eligible." with the same words. Alibaba Qwen3-ASR-1.7B returns "Credit, am I eligible?" and Alibaba Qwen3-ASR-0.6B returns "Credit, am I eligible", both without the word "card".

Paying RM950 in Manglish

Phone calls, 2.9 s
What was said
I dah bayar lah, RM950 ni.
Aisyah 1.0 Pro
I dah bayar lah, sembilan ratus lima puluh ringgit ni.
Alibaba Qwen3-ASR-1.7B
Air bayar lah sebanyak tu sembilan ringgit ni.
Alibaba Qwen-Audio 3.0 ASR Flash
I dah bayar lah sembilan ratus lima puluh ringgit ni.
What to listen for

The reference is "I dah bayar lah, RM950 ni." Aisyah 1.0 Pro returns "I dah bayar lah, sembilan ratus lima puluh ringgit ni." with the amount in words, and Qwen-Audio 3.0 ASR Flash returns the same words. Alibaba Qwen3-ASR-1.7B returns "Air bayar lah sebanyak tu sembilan ringgit ni.", where the amount appears as "sembilan ringgit".

Unit number 2105

Phone calls, 3.4 s
What was said
Yes, my unit number is 2105.
Aisyah 1.0 Pro
Yes, my unit number is two one zero five.
Alibaba Qwen3-ASR-1.7B
Yes, my unit number is two one zero five.
Alibaba Qwen3-ASR-0.6B
Yes, my unit number is two one zero five.
What to listen for

The reference is "Yes, my unit number is 2105." Aisyah 1.0 Pro, Alibaba Qwen3-ASR-1.7B and Alibaba Qwen3-ASR-0.6B each return "Yes, my unit number is two one zero five." The benchmark scores "two one zero five" and "2105" as the same, so each transcript here matches the reference words.

Aisyah vs Alibaba Qwen, answered.

Is Aisyah more accurate than Alibaba Qwen for Malaysian speech?

On Revolab's Malaysian speech benchmark of 1,899 clips, Aisyah 1.0 Pro has a word error rate of 5.79%, against 9.17% for Alibaba Qwen-Audio 3.0 ASR Flash, 18.43% for Qwen3-ASR-1.7B and 24.19% for Qwen3-ASR-0.6B. Revolab built both the benchmark and Aisyah.

How accurate is Alibaba Qwen on Malaysian phone calls?

In the Telephony category of Revolab's Malaysian speech benchmark, Alibaba Qwen-Audio 3.0 ASR Flash has a word error rate of 17.43%, Qwen3-ASR-1.7B 30.69% and Qwen3-ASR-0.6B 30.23%. Aisyah 1.0 Pro has 9.55%.

Is Alibaba Qwen better than Aisyah for any type of Malaysian audio?

In one category on Revolab's Malaysian speech benchmark: Alibaba Qwen-Audio 3.0 ASR Flash has the lower word error rate in FLEURS, 3.29% against 3.42%. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B have a higher rate than Aisyah 1.0 Pro in all 12 categories.

Who built the benchmark behind the Aisyah vs Qwen comparison?

Revolab built both the Malaysian speech benchmark and Aisyah 1.0 Pro. Of its 1,899 clips, 820 are published and 1,079 are held back. On the published clips, Aisyah 1.0 Pro has a word error rate of 4.89% and Alibaba Qwen-Audio 3.0 ASR Flash has 7.22%.

Check the numbers yourself.

Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.

All the numbersEvery figure for all 4 models on this page, as tables

Overall

Word error rate for each model. Lower is better; the lowest in each column is highlighted.
ModelAll clips1,899 clipsPublished820 clipsHeld back1,079 clips
Aisyah 1.0 Pro5.79% (fewer mistakes)4.89% (fewer mistakes)6.58% (fewer mistakes)
Alibaba Qwen-Audio 3.0 ASR Flash9.17%7.22%10.87%
Alibaba Qwen3-ASR-1.7B18.43%15.26%21.20%
Alibaba Qwen3-ASR-0.6B24.19%20.73%27.21%

By kind of audio

Word error rate by kind of audio. Lower is better; the lowest in each row is highlighted.
Kind of audioAisyah 1.0 ProQwen-Audio 3.0 ASR FlashQwen3-ASR-1.7BQwen3-ASR-0.6B
Phone callsTelephony, 229 clips9.55% (fewer mistakes)17.43%30.69%30.23%
Short repliesShort inputs, 155 clips3.28% (fewer mistakes)26.10%28.10%23.08%
Scripted readingRead speech, 154 clips1.30% (fewer mistakes)2.69%6.76%12.31%
PodcastsPodcast, 155 clips4.57% (fewer mistakes)9.84%17.14%24.10%
Drama156 clips5.91% (fewer mistakes)10.80%27.69%37.87%
Animation150 clips5.66% (fewer mistakes)8.91%19.58%24.73%
News156 clips3.48% (fewer mistakes)5.39%16.49%18.56%
Parliament155 clips3.79% (fewer mistakes)5.29%13.89%19.57%
Street interviews143 clips14.24% (fewer mistakes)19.00%34.23%43.13%
Singing152 clips11.01% (fewer mistakes)14.39%18.96%31.94%
Common Voicevolunteer read-aloud, 141 clips3.74% (fewer mistakes)4.78%8.42%12.96%
FLEURSread-aloud research set, 153 clips3.42%3.29% (fewer mistakes)7.65%14.88%

By background noise

Word error rate by background noise. Lower is better; the lowest in each row is highlighted.
BackgroundAisyah 1.0 ProQwen-Audio 3.0 ASR FlashQwen3-ASR-1.7BQwen3-ASR-0.6B
Clean audio1,364 clips4.70% (fewer mistakes)8.13%17.18%21.91%
Some background noise277 clips6.12% (fewer mistakes)8.68%18.23%25.37%
Noisy audio249 clips12.16% (fewer mistakes)16.39%26.76%36.75%

Types of mistake

Types of mistake per 100 words spoken. Lower is better.
ModelWrong wordsper 100 wordsInvented wordsper 100 wordsDropped wordsper 100 words
Aisyah 1.0 Pro3.48% (fewer mistakes)0.62% (fewer mistakes)1.69% (fewer mistakes)
Alibaba Qwen-Audio 3.0 ASR Flash5.51%1.11%2.55%
Alibaba Qwen3-ASR-1.7B13.22%1.09%4.13%
Alibaba Qwen3-ASR-0.6B17.86%1.25%5.09%
How the test worksClips, scoring, noise tags and what is not compared
  • 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
  • Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
  • Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
  • Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
  • Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
  • Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.

Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.

Compare Aisyah with other models

See every model in one table

Test it on your own audio.

Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.