Malaysian speech to text benchmark
Aisyah 1.0 Pro vs YTL ILMU ASR v4.2: Malaysian speech to text compared
We ran Aisyah 1.0 Pro and YTL ILMU ASR v4.2 on the same 1,899 Malaysian speech clips, from telephony and street interviews to parliament and singing, and compared their word error rates.
- Aisyah 1.0 Pro
- Fewer mistakes overall, and on all 12 kinds of audio. Biggest leads: singing and short replies.
- YTL ILMU ASR v4.2
- Not ahead of Aisyah 1.0 Pro on any kind of audio, noise level or clip set. Closest on podcasts: 6.01% against 4.57%.
Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested
- Aisyah 1.0 Pro5.79%
- YTL ILMU ASR v4.29.48%
Where Aisyah and ILMU ASR v4.2 each make fewer mistakes
Aisyah 1.0 ProYTL ILMU ASR v4.2Word error rate. Further left, fewer mistakes.
Aisyah 1.0 Pro makes fewer mistakes on 12 of 12
- Singing152 clipsAisyah 1.0 Pro 11.01%YTL ILMU ASR v4.2 28.80%
- Short repliesShort inputs, 155 clipsAisyah 1.0 Pro 3.28%YTL ILMU ASR v4.2 20.17%
- Phone callsTelephony, 229 clipsAisyah 1.0 Pro 9.55%YTL ILMU ASR v4.2 20.40%
- Drama156 clipsAisyah 1.0 Pro 5.91%YTL ILMU ASR v4.2 9.25%
- Animation150 clipsAisyah 1.0 Pro 5.66%YTL ILMU ASR v4.2 8.84%
- Street interviews143 clipsAisyah 1.0 Pro 14.24%YTL ILMU ASR v4.2 17.02%
- Scripted readingRead speech, 154 clipsAisyah 1.0 Pro 1.30%YTL ILMU ASR v4.2 3.80%
- Common Voicevolunteer read-aloud, 141 clipsAisyah 1.0 Pro 3.74%YTL ILMU ASR v4.2 6.15%
- FLEURSread-aloud research set, 153 clipsAisyah 1.0 Pro 3.42%YTL ILMU ASR v4.2 5.69%
- Parliament155 clipsAisyah 1.0 Pro 3.79%YTL ILMU ASR v4.2 5.36%
- News156 clipsAisyah 1.0 Pro 3.48%YTL ILMU ASR v4.2 5.01%
- PodcastsPodcast, 155 clipsAisyah 1.0 Pro 4.57%YTL ILMU ASR v4.2 6.01%
ILMU ASR v4.2 makes fewer mistakes on 0 of 12
No kind of audio in this benchmark. It comes closest on podcasts: 6.01% against 4.57%.
By background noise and clip set
The same 1,899 clips, split another way. The lower figure is in bold.
- Clean audio1,364 clipsAisyah 1.0 Pro 4.70%YTL ILMU ASR v4.2 7.79%
- Some background noise277 clipsAisyah 1.0 Pro 6.12%YTL ILMU ASR v4.2 8.93%
- Noisy audio249 clipsAisyah 1.0 Pro 12.16%YTL ILMU ASR v4.2 20.91%
- Published clips820 clips anyone can checkAisyah 1.0 Pro 4.89%YTL ILMU ASR v4.2 7.78%
- Held-back clips1,079 clips no model has seenAisyah 1.0 Pro 6.58%YTL ILMU ASR v4.2 10.97%
Hear the difference
Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.
A ringgit amount in words
Phone calls, 2.9 s- What was said
- I dah bayar lah, RM950 ni.
- Aisyah 1.0 Pro
- I dah bayar lah, sembilan ratus lima puluh ringgit ni.
- YTL ILMU ASR v4.2
- i dah bayar lah sembilan ratus lima puluh ringgit ni
What to listen for
Aisyah 1.0 Pro and YTL ILMU ASR v4.2 returned the same words here, both writing "RM950" as "sembilan ratus lima puluh ringgit". Only capitals and punctuation differ between the two outputs.
Short eligibility question
Phone calls, 1.9 s- What was said
- Credit card am I eligible?
- Aisyah 1.0 Pro
- Credit card am I eligible?
- YTL ILMU ASR v4.2
- Credit card emi eligible
What to listen for
Aisyah 1.0 Pro returned "Credit card am I eligible?", the same as the reference. YTL ILMU ASR v4.2 returned "Credit card emi eligible", with "emi" where the reference has "am I".
Street interview in Manglish
Street interviews, 4.9 s- What was said
- Apa bagus dia ada yang tak bagus dia. Depends lah. Depends pada individu. Haa, aa ah, individu lah.
- Aisyah 1.0 Pro
- Apa bagus dia, ada yang tak bagus dia. depends lah. depends pada individu lah.
- YTL ILMU ASR v4.2
- Apa bagus dia ada yang tak bagus dia. Depends lah. Haah individu lah
What to listen for
For the stretch the reference gives as "Depends pada individu. Haa, aa ah, individu lah.", Aisyah 1.0 Pro returned "depends pada individu lah." and YTL ILMU ASR v4.2 returned "Haah individu lah".
Aisyah vs YTL ILMU, answered.
Is Aisyah more accurate than YTL ILMU?
On Revolab's Malaysian speech benchmark of 1,899 clips, Aisyah 1.0 Pro has a word error rate of 5.79% and YTL ILMU ASR v4.2 has 9.48%. Aisyah 1.0 Pro has the lower WER in all 12 categories and on clean, moderate and noisy audio. Revolab built both the benchmark and Aisyah.
How accurate is YTL ILMU ASR on telephony audio?
On the 229 Telephony clips in Revolab's Malaysian speech benchmark, YTL ILMU ASR v4.2 has a word error rate of 20.40% and Aisyah 1.0 Pro has 9.55%, a gap of 10.85 percentage points. Lower is better.
Is YTL ILMU ASR better than Aisyah in any category?
No. On Revolab's Malaysian speech benchmark, YTL ILMU ASR v4.2 does not have a lower WER than Aisyah 1.0 Pro in any of the 12 categories, at any noise level, or on the 820 published or 1,079 held-back clips. The narrowest category gap is Podcast, at 1.44 percentage points.
Which is more accurate on noisy audio, Aisyah or YTL ILMU?
On the 249 clips tagged noisy in Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has a word error rate of 12.16% and YTL ILMU ASR v4.2 has 20.91%, a gap of 8.75 percentage points. Aisyah 1.0 Pro also has the lower WER on clean and moderate clips.
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
All the numbersEvery figure for all 2 models on this page, as tables
Overall
| Model | All clips1,899 clips | Published820 clips | Held back1,079 clips |
|---|---|---|---|
| Aisyah 1.0 Pro | 5.79% (fewer mistakes) | 4.89% (fewer mistakes) | 6.58% (fewer mistakes) |
| YTL ILMU ASR v4.2 | 9.48% | 7.78% | 10.97% |
By kind of audio
| Kind of audio | Aisyah 1.0 Pro | ILMU ASR v4.2 |
|---|---|---|
| Phone callsTelephony, 229 clips | 9.55% (fewer mistakes) | 20.40% |
| Short repliesShort inputs, 155 clips | 3.28% (fewer mistakes) | 20.17% |
| Scripted readingRead speech, 154 clips | 1.30% (fewer mistakes) | 3.80% |
| PodcastsPodcast, 155 clips | 4.57% (fewer mistakes) | 6.01% |
| Drama156 clips | 5.91% (fewer mistakes) | 9.25% |
| Animation150 clips | 5.66% (fewer mistakes) | 8.84% |
| News156 clips | 3.48% (fewer mistakes) | 5.01% |
| Parliament155 clips | 3.79% (fewer mistakes) | 5.36% |
| Street interviews143 clips | 14.24% (fewer mistakes) | 17.02% |
| Singing152 clips | 11.01% (fewer mistakes) | 28.80% |
| Common Voicevolunteer read-aloud, 141 clips | 3.74% (fewer mistakes) | 6.15% |
| FLEURSread-aloud research set, 153 clips | 3.42% (fewer mistakes) | 5.69% |
By background noise
| Background | Aisyah 1.0 Pro | ILMU ASR v4.2 |
|---|---|---|
| Clean audio1,364 clips | 4.70% (fewer mistakes) | 7.79% |
| Some background noise277 clips | 6.12% (fewer mistakes) | 8.93% |
| Noisy audio249 clips | 12.16% (fewer mistakes) | 20.91% |
Types of mistake
| Model | Wrong wordsper 100 words | Invented wordsper 100 words | Dropped wordsper 100 words |
|---|---|---|---|
| Aisyah 1.0 Pro | 3.48% (fewer mistakes) | 0.62% (fewer mistakes) | 1.69% (fewer mistakes) |
| YTL ILMU ASR v4.2 | 6.15% | 0.87% | 2.46% |
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Compare Aisyah with other models
See every model in one table- Aisyah vs ElevenLabs ScribeScribe v25.79% vs 6.67%
- Aisyah vs Google GeminiGemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.6 Flash5.79% vs 7.00%
- Aisyah vs OpenAI WhisperWhisper large-v35.79% vs 17.30%
- Aisyah vs Alibaba QwenQwen-Audio 3.0 ASR Flash, Qwen3-ASR-1.7B, Qwen3-ASR-0.6B5.79% vs 9.17%
- Aisyah vs AssemblyAI UniversalUniversal 25.79% vs 18.80%
- Aisyah vs Deepgram NovaNova 35.79% vs 31.99%
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.