Malaysian speech to text benchmark
Aisyah 1.0 Pro vs AssemblyAI Universal 2: Malaysian speech to text
Word error rates for Aisyah 1.0 Pro and AssemblyAI Universal 2 on Revolab's Malaysian speech benchmark, across 12 categories from parliament and news to telephony, singing and street interviews.
- Aisyah 1.0 Pro
- Fewer mistakes overall, and on all 12 kinds of audio. Biggest leads: short replies and phone calls.
- AssemblyAI Universal 2
- Not ahead of Aisyah 1.0 Pro on any kind of audio, noise level or clip set. Closest on Common Voice: 5.78% against 3.74%.
Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested
- Aisyah 1.0 Pro5.79%
- AssemblyAI Universal 218.80%
Where Aisyah and Universal 2 each make fewer mistakes
Aisyah 1.0 ProAssemblyAI Universal 2Word error rate. Further left, fewer mistakes.
Aisyah 1.0 Pro makes fewer mistakes on 12 of 12
- Short repliesShort inputs, 155 clipsAisyah 1.0 Pro 3.28%AssemblyAI Universal 2 46.15%
- Phone callsTelephony, 229 clipsAisyah 1.0 Pro 9.55%AssemblyAI Universal 2 51.22%
- News156 clipsAisyah 1.0 Pro 3.48%AssemblyAI Universal 2 30.19%
- Singing152 clipsAisyah 1.0 Pro 11.01%AssemblyAI Universal 2 37.45%
- Drama156 clipsAisyah 1.0 Pro 5.91%AssemblyAI Universal 2 15.50%
- Street interviews143 clipsAisyah 1.0 Pro 14.24%AssemblyAI Universal 2 23.35%
- PodcastsPodcast, 155 clipsAisyah 1.0 Pro 4.57%AssemblyAI Universal 2 13.38%
- Animation150 clipsAisyah 1.0 Pro 5.66%AssemblyAI Universal 2 13.54%
- Parliament155 clipsAisyah 1.0 Pro 3.79%AssemblyAI Universal 2 10.81%
- FLEURSread-aloud research set, 153 clipsAisyah 1.0 Pro 3.42%AssemblyAI Universal 2 6.52%
- Scripted readingRead speech, 154 clipsAisyah 1.0 Pro 1.30%AssemblyAI Universal 2 3.61%
- Common Voicevolunteer read-aloud, 141 clipsAisyah 1.0 Pro 3.74%AssemblyAI Universal 2 5.78%
Universal 2 makes fewer mistakes on 0 of 12
No kind of audio in this benchmark. It comes closest on Common Voice: 5.78% against 3.74%.
By background noise and clip set
The same 1,899 clips, split another way. The lower figure is in bold.
- Clean audio1,364 clipsAisyah 1.0 Pro 4.70%AssemblyAI Universal 2 18.89%
- Some background noise277 clipsAisyah 1.0 Pro 6.12%AssemblyAI Universal 2 13.15%
- Noisy audio249 clipsAisyah 1.0 Pro 12.16%AssemblyAI Universal 2 27.61%
- Published clips820 clips anyone can checkAisyah 1.0 Pro 4.89%AssemblyAI Universal 2 14.86%
- Held-back clips1,079 clips no model has seenAisyah 1.0 Pro 6.58%AssemblyAI Universal 2 22.25%
Hear the difference
Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.
Credit card eligibility question
Phone calls, 1.9 s- What was said
- Credit card am I eligible?
- Aisyah 1.0 Pro
- Credit card am I eligible?
- AssemblyAI Universal 2
- Perdana Menteri, adakah saya wajib?
What to listen for
Aisyah 1.0 Pro returned "Credit card am I eligible?", matching the reference word for word. AssemblyAI Universal 2 returned "Perdana Menteri, adakah saya wajib?", which shares no words with the reference.
Account number on a call
Phone calls, 2.5 s- What was said
- 102105
- Aisyah 1.0 Pro
- one, zero, two, one, zero, five
- AssemblyAI Universal 2
- 202105
What to listen for
The reference is the number 102105. Aisyah 1.0 Pro wrote it as words, "one, zero, two, one, zero, five", with every digit in order. AssemblyAI Universal 2 returned "202105", which differs from the reference in the first digit.
Street interview about films
Street interviews, 10.1 s- What was said
- Filem Malaysia kurang sikit, tapi aa I suka ah macam cerita-cerita gangster semua. Filem apa you suka dan kenapa you suka? Macam KL Gangster, I tengok semua. Juvana.
- Aisyah 1.0 Pro
- Filem Malaysia kurang sikit tapi saya suka macam cerita-cerita gangster semua. Kenapa you suka dan kenapa? Macam cat gangster. Saya tengok semua Jovana.
- AssemblyAI Universal 2
- Film Asia kurang sikit tapi I suka macam cerita-cerita gangster semua Kenapa you suka dan kenapa you suka? Macam Kedang Gangster, I tengok semua Giovanna
What to listen for
Both outputs differ from the reference in places. For the question "Filem apa you suka dan kenapa you suka?", AssemblyAI Universal 2 wrote "Kenapa you suka dan kenapa you suka?" and Aisyah 1.0 Pro wrote "Kenapa you suka dan kenapa?". Universal 2 also has "I tengok semua" as in the reference, where Aisyah 1.0 Pro wrote "Saya tengok semua". Aisyah 1.0 Pro has the opening "Filem Malaysia" as in the reference, where Universal 2 wrote "Film Asia".
Aisyah vs AssemblyAI Universal, answered.
Is Aisyah more accurate than AssemblyAI Universal 2 for Malaysian speech?
On Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has a word error rate of 5.79% across 1,899 clips and AssemblyAI Universal 2 has 18.80%. Aisyah 1.0 Pro has the lower rate in all 12 categories. Revolab built both the benchmark and Aisyah, so test on your own audio as well.
AssemblyAI Universal 2 accuracy on Malaysian phone calls
On the 229 Telephony clips in Revolab's Malaysian speech benchmark, AssemblyAI Universal 2 has a word error rate of 51.22% and Aisyah 1.0 Pro has 9.55%, a gap of 41.67 percentage points. Only Short inputs shows a wider gap between the two models.
Does AssemblyAI Universal 2 have a lower error rate than Aisyah in any category?
No. On Revolab's Malaysian speech benchmark, AssemblyAI Universal 2 does not have a lower word error rate than Aisyah 1.0 Pro in any of the 12 categories. The gaps are narrowest in Common Voice, 5.78% against 3.74% for Aisyah 1.0 Pro, and Read speech, 3.61% against 1.30%.
Which is more accurate on noisy Malaysian audio, Aisyah or AssemblyAI Universal 2?
On the 249 clips tagged noisy in Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has a word error rate of 12.16% and AssemblyAI Universal 2 has 27.61%. Aisyah 1.0 Pro also has the lower word error rate on the clips tagged clean and moderate.
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
All the numbersEvery figure for all 2 models on this page, as tables
Overall
| Model | All clips1,899 clips | Published820 clips | Held back1,079 clips |
|---|---|---|---|
| Aisyah 1.0 Pro | 5.79% (fewer mistakes) | 4.89% (fewer mistakes) | 6.58% (fewer mistakes) |
| AssemblyAI Universal 2 | 18.80% | 14.86% | 22.25% |
By kind of audio
| Kind of audio | Aisyah 1.0 Pro | Universal 2 |
|---|---|---|
| Phone callsTelephony, 229 clips | 9.55% (fewer mistakes) | 51.22% |
| Short repliesShort inputs, 155 clips | 3.28% (fewer mistakes) | 46.15% |
| Scripted readingRead speech, 154 clips | 1.30% (fewer mistakes) | 3.61% |
| PodcastsPodcast, 155 clips | 4.57% (fewer mistakes) | 13.38% |
| Drama156 clips | 5.91% (fewer mistakes) | 15.50% |
| Animation150 clips | 5.66% (fewer mistakes) | 13.54% |
| News156 clips | 3.48% (fewer mistakes) | 30.19% |
| Parliament155 clips | 3.79% (fewer mistakes) | 10.81% |
| Street interviews143 clips | 14.24% (fewer mistakes) | 23.35% |
| Singing152 clips | 11.01% (fewer mistakes) | 37.45% |
| Common Voicevolunteer read-aloud, 141 clips | 3.74% (fewer mistakes) | 5.78% |
| FLEURSread-aloud research set, 153 clips | 3.42% (fewer mistakes) | 6.52% |
By background noise
| Background | Aisyah 1.0 Pro | Universal 2 |
|---|---|---|
| Clean audio1,364 clips | 4.70% (fewer mistakes) | 18.89% |
| Some background noise277 clips | 6.12% (fewer mistakes) | 13.15% |
| Noisy audio249 clips | 12.16% (fewer mistakes) | 27.61% |
Types of mistake
| Model | Wrong wordsper 100 words | Invented wordsper 100 words | Dropped wordsper 100 words |
|---|---|---|---|
| Aisyah 1.0 Pro | 3.48% (fewer mistakes) | 0.62% (fewer mistakes) | 1.69% (fewer mistakes) |
| AssemblyAI Universal 2 | 11.07% | 0.70% | 7.02% |
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
- A second AssemblyAI run, universal-3-5-pro, is not listed separately: its outputs were identical to Universal 2 on every clip.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Compare Aisyah with other models
See every model in one table- Aisyah vs ElevenLabs ScribeScribe v25.79% vs 6.67%
- Aisyah vs Google GeminiGemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.6 Flash5.79% vs 7.00%
- Aisyah vs YTL ILMUILMU ASR v4.25.79% vs 9.48%
- Aisyah vs OpenAI WhisperWhisper large-v35.79% vs 17.30%
- Aisyah vs Alibaba QwenQwen-Audio 3.0 ASR Flash, Qwen3-ASR-1.7B, Qwen3-ASR-0.6B5.79% vs 9.17%
- Aisyah vs Deepgram NovaNova 35.79% vs 31.99%
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.