Malaysian speech to text benchmark
Aisyah 1.0 Pro vs ElevenLabs Scribe v2 for Malaysian speech to text
Revolab ran Aisyah 1.0 Pro and ElevenLabs Scribe v2 on all 1,899 clips of its own Malaysian speech benchmark, from Parliament and News to Telephony and Street interviews. See where each has the lower word error rate.
- Aisyah 1.0 Pro
- Fewer mistakes overall, and on 8 of 12 kinds of audio. Biggest leads: short replies and phone calls.
- ElevenLabs Scribe v2
- Fewer mistakes on 4 of 12: street interviews, FLEURS, Common Voice and singing. Also on noisy audio and the 820 published clips.
Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested
- Aisyah 1.0 Pro5.79%
- ElevenLabs Scribe v26.67%
Where Aisyah and Scribe v2 each make fewer mistakes
Aisyah 1.0 ProElevenLabs Scribe v2Word error rate. Further left, fewer mistakes.
Aisyah 1.0 Pro makes fewer mistakes on 8 of 12
- Short repliesShort inputs, 155 clipsAisyah 1.0 Pro 3.28%ElevenLabs Scribe v2 15.04%
- Phone callsTelephony, 229 clipsAisyah 1.0 Pro 9.55%ElevenLabs Scribe v2 14.61%
- Animation150 clipsAisyah 1.0 Pro 5.66%ElevenLabs Scribe v2 7.53%
- Drama156 clipsAisyah 1.0 Pro 5.91%ElevenLabs Scribe v2 7.41%
- PodcastsPodcast, 155 clipsAisyah 1.0 Pro 4.57%ElevenLabs Scribe v2 5.73%
- Scripted readingRead speech, 154 clipsAisyah 1.0 Pro 1.30%ElevenLabs Scribe v2 2.22%
- Parliament155 clipsAisyah 1.0 Pro 3.79%ElevenLabs Scribe v2 4.65%
- News156 clipsAisyah 1.0 Pro 3.48%ElevenLabs Scribe v2 4.28%
Scribe v2 makes fewer mistakes on 4 of 12
- Street interviews143 clipsAisyah 1.0 Pro 14.24%ElevenLabs Scribe v2 13.05%
- FLEURSread-aloud research set, 153 clipsAisyah 1.0 Pro 3.42%ElevenLabs Scribe v2 2.38%
- Common Voicevolunteer read-aloud, 141 clipsAisyah 1.0 Pro 3.74%ElevenLabs Scribe v2 3.16%
- Singing152 clipsAisyah 1.0 Pro 11.01%ElevenLabs Scribe v2 10.94%
By background noise and clip set
The same 1,899 clips, split another way. The lower figure is in bold.
- Clean audio1,364 clipsAisyah 1.0 Pro 4.70%ElevenLabs Scribe v2 5.88%
- Some background noise277 clipsAisyah 1.0 Pro 6.12%ElevenLabs Scribe v2 6.95%
- Noisy audio249 clipsAisyah 1.0 Pro 12.16%ElevenLabs Scribe v2 11.19%
- Published clips820 clips anyone can checkAisyah 1.0 Pro 4.89%ElevenLabs Scribe v2 4.79%
- Held-back clips1,079 clips no model has seenAisyah 1.0 Pro 6.58%ElevenLabs Scribe v2 8.29%
Hear the difference
Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.
A ringgit amount in words
Phone calls, 2.9 s- What was said
- I dah bayar lah, RM950 ni.
- Aisyah 1.0 Pro
- I dah bayar lah, sembilan ratus lima puluh ringgit ni.
- ElevenLabs Scribe v2
- I dah bayar lah sembilan ratus sembilan puluh ringgit ni.
What to listen for
Both outputs keep "I dah bayar lah" from the reference. Aisyah 1.0 Pro writes the amount as "sembilan ratus lima puluh ringgit", which is RM950 in words, while ElevenLabs Scribe v2 writes "sembilan ratus sembilan puluh ringgit", a different amount from the reference.
Street interview about Malaysian films
Street interviews, 10.1 s- What was said
- Filem Malaysia kurang sikit, tapi aa I suka ah macam cerita-cerita gangster semua. Filem apa you suka dan kenapa you suka? Macam KL Gangster, I tengok semua. Juvana.
- Aisyah 1.0 Pro
- Filem Malaysia kurang sikit tapi saya suka macam cerita-cerita gangster semua. Kenapa you suka dan kenapa? Macam cat gangster. Saya tengok semua Jovana.
- ElevenLabs Scribe v2
- Filem Malaysia kurang sikit tapi aa I suka macam cerita-cerita gangster semua. Filem apa yang you suka? Sebab kenapa suka? Macam Kapten Gangster, I tengok semua. Juvana
What to listen for
ElevenLabs Scribe v2 stays closer to the reference on this clip, keeping "aa I suka" and "I tengok semua. Juvana", where Aisyah 1.0 Pro writes "saya suka" and "Saya tengok semua Jovana." Neither output matches "KL Gangster": Aisyah 1.0 Pro has "cat gangster" and Scribe v2 has "Kapten Gangster".
Account number in words
Phone calls, 2.5 s- What was said
- 102105
- Aisyah 1.0 Pro
- one, zero, two, one, zero, five
- ElevenLabs Scribe v2
- One zero-- two one zero five
What to listen for
The reference is "102105". Aisyah 1.0 Pro returns "one, zero, two, one, zero, five" and ElevenLabs Scribe v2 returns "One zero-- two one zero five", so both give every digit in order as words, which the benchmark scores the same as digits.
Aisyah vs ElevenLabs Scribe, answered.
Is Aisyah more accurate than ElevenLabs Scribe for Malaysian speech?
On Revolab's benchmark of 1,899 Malaysian speech clips, Aisyah 1.0 Pro has a word error rate of 5.79% and ElevenLabs Scribe v2 has 6.67%. Aisyah 1.0 Pro is lower in 8 of the 12 categories. Scribe v2 is lower in Common Voice, FLEURS, Singing and Street interviews, on the 820 published clips and on clips tagged noisy.
Aisyah or ElevenLabs Scribe for Malaysian call centre transcription?
The closest category in Revolab's benchmark is Telephony, 229 clips, where Aisyah 1.0 Pro has a word error rate of 9.55% and ElevenLabs Scribe v2 has 14.61%, a gap of 5.06 percentage points. Aisyah 1.0 Pro also has the lower word error rate across all 1,899 clips.
Is ElevenLabs Scribe better than Aisyah on noisy audio?
On clips tagged noisy in Revolab's benchmark, ElevenLabs Scribe v2 has the lower word error rate: 11.19% against 12.16% for Aisyah 1.0 Pro. On clean clips Aisyah 1.0 Pro is lower, 4.70% against 5.88%, and on moderate noise clips it is lower as well, 6.12% against 6.95%.
Aisyah vs ElevenLabs Scribe accuracy for short voice inputs
In the Short inputs category of Revolab's benchmark, 155 clips, Aisyah 1.0 Pro has a word error rate of 3.28% and ElevenLabs Scribe v2 has 15.04%, a difference of 11.76 percentage points. It is the widest category gap between the two models.
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
All the numbersEvery figure for all 2 models on this page, as tables
Overall
| Model | All clips1,899 clips | Published820 clips | Held back1,079 clips |
|---|---|---|---|
| Aisyah 1.0 Pro | 5.79% (fewer mistakes) | 4.89% | 6.58% (fewer mistakes) |
| ElevenLabs Scribe v2 | 6.67% | 4.79% (fewer mistakes) | 8.29% |
By kind of audio
| Kind of audio | Aisyah 1.0 Pro | Scribe v2 |
|---|---|---|
| Phone callsTelephony, 229 clips | 9.55% (fewer mistakes) | 14.61% |
| Short repliesShort inputs, 155 clips | 3.28% (fewer mistakes) | 15.04% |
| Scripted readingRead speech, 154 clips | 1.30% (fewer mistakes) | 2.22% |
| PodcastsPodcast, 155 clips | 4.57% (fewer mistakes) | 5.73% |
| Drama156 clips | 5.91% (fewer mistakes) | 7.41% |
| Animation150 clips | 5.66% (fewer mistakes) | 7.53% |
| News156 clips | 3.48% (fewer mistakes) | 4.28% |
| Parliament155 clips | 3.79% (fewer mistakes) | 4.65% |
| Street interviews143 clips | 14.24% | 13.05% (fewer mistakes) |
| Singing152 clips | 11.01% | 10.94% (fewer mistakes) |
| Common Voicevolunteer read-aloud, 141 clips | 3.74% | 3.16% (fewer mistakes) |
| FLEURSread-aloud research set, 153 clips | 3.42% | 2.38% (fewer mistakes) |
By background noise
| Background | Aisyah 1.0 Pro | Scribe v2 |
|---|---|---|
| Clean audio1,364 clips | 4.70% (fewer mistakes) | 5.88% |
| Some background noise277 clips | 6.12% (fewer mistakes) | 6.95% |
| Noisy audio249 clips | 12.16% | 11.19% (fewer mistakes) |
Types of mistake
| Model | Wrong wordsper 100 words | Invented wordsper 100 words | Dropped wordsper 100 words |
|---|---|---|---|
| Aisyah 1.0 Pro | 3.48% (fewer mistakes) | 0.62% (fewer mistakes) | 1.69% (fewer mistakes) |
| ElevenLabs Scribe v2 | 4.11% | 0.70% | 1.85% |
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Compare Aisyah with other models
See every model in one table- Aisyah vs Google GeminiGemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.6 Flash5.79% vs 7.00%
- Aisyah vs YTL ILMUILMU ASR v4.25.79% vs 9.48%
- Aisyah vs OpenAI WhisperWhisper large-v35.79% vs 17.30%
- Aisyah vs Alibaba QwenQwen-Audio 3.0 ASR Flash, Qwen3-ASR-1.7B, Qwen3-ASR-0.6B5.79% vs 9.17%
- Aisyah vs AssemblyAI UniversalUniversal 25.79% vs 18.80%
- Aisyah vs Deepgram NovaNova 35.79% vs 31.99%
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.