Malaysian speech to text benchmark

Aisyah 1.0 Pro vs ElevenLabs Scribe v2 for Malaysian speech to text

Revolab ran Aisyah 1.0 Pro and ElevenLabs Scribe v2 on all 1,899 clips of its own Malaysian speech benchmark, from Parliament and News to Telephony and Street interviews. See where each has the lower word error rate.

Aisyah 1.0 Pro
Fewer mistakes overall, and on 8 of 12 kinds of audio. Biggest leads: short replies and phone calls.
ElevenLabs Scribe v2
Fewer mistakes on 4 of 12: street interviews, FLEURS, Common Voice and singing. Also on noisy audio and the 820 published clips.

Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested

Word error rate on all 1,899 clipsMistakes per 100 words spoken. Shorter bar, fewer mistakes.
  • Aisyah 1.0 Pro5.79%
  • ElevenLabs Scribe v26.67%

Where Aisyah and Scribe v2 each make fewer mistakes

Aisyah 1.0 ProElevenLabs Scribe v2Word error rate. Further left, fewer mistakes.

Aisyah 1.0 Pro makes fewer mistakes on 8 of 12

  • Short repliesShort inputs, 155 clips
    Aisyah 1.0 Pro 3.28%ElevenLabs Scribe v2 15.04%
  • Phone callsTelephony, 229 clips
    Aisyah 1.0 Pro 9.55%ElevenLabs Scribe v2 14.61%
  • Animation150 clips
    Aisyah 1.0 Pro 5.66%ElevenLabs Scribe v2 7.53%
  • Drama156 clips
    Aisyah 1.0 Pro 5.91%ElevenLabs Scribe v2 7.41%
  • PodcastsPodcast, 155 clips
    Aisyah 1.0 Pro 4.57%ElevenLabs Scribe v2 5.73%
  • Scripted readingRead speech, 154 clips
    Aisyah 1.0 Pro 1.30%ElevenLabs Scribe v2 2.22%
  • Parliament155 clips
    Aisyah 1.0 Pro 3.79%ElevenLabs Scribe v2 4.65%
  • News156 clips
    Aisyah 1.0 Pro 3.48%ElevenLabs Scribe v2 4.28%

Scribe v2 makes fewer mistakes on 4 of 12

  • Street interviews143 clips
    Aisyah 1.0 Pro 14.24%ElevenLabs Scribe v2 13.05%
  • FLEURSread-aloud research set, 153 clips
    Aisyah 1.0 Pro 3.42%ElevenLabs Scribe v2 2.38%
  • Common Voicevolunteer read-aloud, 141 clips
    Aisyah 1.0 Pro 3.74%ElevenLabs Scribe v2 3.16%
  • Singing152 clips
    Aisyah 1.0 Pro 11.01%ElevenLabs Scribe v2 10.94%

By background noise and clip set

The same 1,899 clips, split another way. The lower figure is in bold.

  • Clean audio1,364 clips
    Aisyah 1.0 Pro 4.70%ElevenLabs Scribe v2 5.88%
  • Some background noise277 clips
    Aisyah 1.0 Pro 6.12%ElevenLabs Scribe v2 6.95%
  • Noisy audio249 clips
    Aisyah 1.0 Pro 12.16%ElevenLabs Scribe v2 11.19%
  • Published clips820 clips anyone can check
    Aisyah 1.0 Pro 4.89%ElevenLabs Scribe v2 4.79%
  • Held-back clips1,079 clips no model has seen
    Aisyah 1.0 Pro 6.58%ElevenLabs Scribe v2 8.29%

Hear the difference

Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.

A ringgit amount in words

Phone calls, 2.9 s
What was said
I dah bayar lah, RM950 ni.
Aisyah 1.0 Pro
I dah bayar lah, sembilan ratus lima puluh ringgit ni.
ElevenLabs Scribe v2
I dah bayar lah sembilan ratus sembilan puluh ringgit ni.
What to listen for

Both outputs keep "I dah bayar lah" from the reference. Aisyah 1.0 Pro writes the amount as "sembilan ratus lima puluh ringgit", which is RM950 in words, while ElevenLabs Scribe v2 writes "sembilan ratus sembilan puluh ringgit", a different amount from the reference.

Street interview about Malaysian films

Street interviews, 10.1 s
What was said
Filem Malaysia kurang sikit, tapi aa I suka ah macam cerita-cerita gangster semua. Filem apa you suka dan kenapa you suka? Macam KL Gangster, I tengok semua. Juvana.
Aisyah 1.0 Pro
Filem Malaysia kurang sikit tapi saya suka macam cerita-cerita gangster semua. Kenapa you suka dan kenapa? Macam cat gangster. Saya tengok semua Jovana.
ElevenLabs Scribe v2
Filem Malaysia kurang sikit tapi aa I suka macam cerita-cerita gangster semua. Filem apa yang you suka? Sebab kenapa suka? Macam Kapten Gangster, I tengok semua. Juvana
What to listen for

ElevenLabs Scribe v2 stays closer to the reference on this clip, keeping "aa I suka" and "I tengok semua. Juvana", where Aisyah 1.0 Pro writes "saya suka" and "Saya tengok semua Jovana." Neither output matches "KL Gangster": Aisyah 1.0 Pro has "cat gangster" and Scribe v2 has "Kapten Gangster".

Account number in words

Phone calls, 2.5 s
What was said
102105
Aisyah 1.0 Pro
one, zero, two, one, zero, five
ElevenLabs Scribe v2
One zero-- two one zero five
What to listen for

The reference is "102105". Aisyah 1.0 Pro returns "one, zero, two, one, zero, five" and ElevenLabs Scribe v2 returns "One zero-- two one zero five", so both give every digit in order as words, which the benchmark scores the same as digits.

Aisyah vs ElevenLabs Scribe, answered.

Is Aisyah more accurate than ElevenLabs Scribe for Malaysian speech?

On Revolab's benchmark of 1,899 Malaysian speech clips, Aisyah 1.0 Pro has a word error rate of 5.79% and ElevenLabs Scribe v2 has 6.67%. Aisyah 1.0 Pro is lower in 8 of the 12 categories. Scribe v2 is lower in Common Voice, FLEURS, Singing and Street interviews, on the 820 published clips and on clips tagged noisy.

Aisyah or ElevenLabs Scribe for Malaysian call centre transcription?

The closest category in Revolab's benchmark is Telephony, 229 clips, where Aisyah 1.0 Pro has a word error rate of 9.55% and ElevenLabs Scribe v2 has 14.61%, a gap of 5.06 percentage points. Aisyah 1.0 Pro also has the lower word error rate across all 1,899 clips.

Is ElevenLabs Scribe better than Aisyah on noisy audio?

On clips tagged noisy in Revolab's benchmark, ElevenLabs Scribe v2 has the lower word error rate: 11.19% against 12.16% for Aisyah 1.0 Pro. On clean clips Aisyah 1.0 Pro is lower, 4.70% against 5.88%, and on moderate noise clips it is lower as well, 6.12% against 6.95%.

Aisyah vs ElevenLabs Scribe accuracy for short voice inputs

In the Short inputs category of Revolab's benchmark, 155 clips, Aisyah 1.0 Pro has a word error rate of 3.28% and ElevenLabs Scribe v2 has 15.04%, a difference of 11.76 percentage points. It is the widest category gap between the two models.

Check the numbers yourself.

Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.

All the numbersEvery figure for all 2 models on this page, as tables

Overall

Word error rate for each model. Lower is better; the lowest in each column is highlighted.
ModelAll clips1,899 clipsPublished820 clipsHeld back1,079 clips
Aisyah 1.0 Pro5.79% (fewer mistakes)4.89%6.58% (fewer mistakes)
ElevenLabs Scribe v26.67%4.79% (fewer mistakes)8.29%

By kind of audio

Word error rate by kind of audio. Lower is better; the lowest in each row is highlighted.
Kind of audioAisyah 1.0 ProScribe v2
Phone callsTelephony, 229 clips9.55% (fewer mistakes)14.61%
Short repliesShort inputs, 155 clips3.28% (fewer mistakes)15.04%
Scripted readingRead speech, 154 clips1.30% (fewer mistakes)2.22%
PodcastsPodcast, 155 clips4.57% (fewer mistakes)5.73%
Drama156 clips5.91% (fewer mistakes)7.41%
Animation150 clips5.66% (fewer mistakes)7.53%
News156 clips3.48% (fewer mistakes)4.28%
Parliament155 clips3.79% (fewer mistakes)4.65%
Street interviews143 clips14.24%13.05% (fewer mistakes)
Singing152 clips11.01%10.94% (fewer mistakes)
Common Voicevolunteer read-aloud, 141 clips3.74%3.16% (fewer mistakes)
FLEURSread-aloud research set, 153 clips3.42%2.38% (fewer mistakes)

By background noise

Word error rate by background noise. Lower is better; the lowest in each row is highlighted.
BackgroundAisyah 1.0 ProScribe v2
Clean audio1,364 clips4.70% (fewer mistakes)5.88%
Some background noise277 clips6.12% (fewer mistakes)6.95%
Noisy audio249 clips12.16%11.19% (fewer mistakes)

Types of mistake

Types of mistake per 100 words spoken. Lower is better.
ModelWrong wordsper 100 wordsInvented wordsper 100 wordsDropped wordsper 100 words
Aisyah 1.0 Pro3.48% (fewer mistakes)0.62% (fewer mistakes)1.69% (fewer mistakes)
ElevenLabs Scribe v24.11%0.70%1.85%
How the test worksClips, scoring, noise tags and what is not compared
  • 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
  • Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
  • Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
  • Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
  • Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
  • Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.

Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.

Compare Aisyah with other models

See every model in one table

Test it on your own audio.

Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.