Malaysian speech to text benchmark

Aisyah 1.0 Pro vs AssemblyAI Universal 2: Malaysian speech to text

Word error rates for Aisyah 1.0 Pro and AssemblyAI Universal 2 on Revolab's Malaysian speech benchmark, across 12 categories from parliament and news to telephony, singing and street interviews.

Aisyah 1.0 Pro
Fewer mistakes overall, and on all 12 kinds of audio. Biggest leads: short replies and phone calls.
AssemblyAI Universal 2
Not ahead of Aisyah 1.0 Pro on any kind of audio, noise level or clip set. Closest on Common Voice: 5.78% against 3.74%.

Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested

Word error rate on all 1,899 clipsMistakes per 100 words spoken. Shorter bar, fewer mistakes.
  • Aisyah 1.0 Pro5.79%
  • AssemblyAI Universal 218.80%

Where Aisyah and Universal 2 each make fewer mistakes

Aisyah 1.0 ProAssemblyAI Universal 2Word error rate. Further left, fewer mistakes.

Aisyah 1.0 Pro makes fewer mistakes on 12 of 12

  • Short repliesShort inputs, 155 clips
    Aisyah 1.0 Pro 3.28%AssemblyAI Universal 2 46.15%
  • Phone callsTelephony, 229 clips
    Aisyah 1.0 Pro 9.55%AssemblyAI Universal 2 51.22%
  • News156 clips
    Aisyah 1.0 Pro 3.48%AssemblyAI Universal 2 30.19%
  • Singing152 clips
    Aisyah 1.0 Pro 11.01%AssemblyAI Universal 2 37.45%
  • Drama156 clips
    Aisyah 1.0 Pro 5.91%AssemblyAI Universal 2 15.50%
  • Street interviews143 clips
    Aisyah 1.0 Pro 14.24%AssemblyAI Universal 2 23.35%
  • PodcastsPodcast, 155 clips
    Aisyah 1.0 Pro 4.57%AssemblyAI Universal 2 13.38%
  • Animation150 clips
    Aisyah 1.0 Pro 5.66%AssemblyAI Universal 2 13.54%
  • Parliament155 clips
    Aisyah 1.0 Pro 3.79%AssemblyAI Universal 2 10.81%
  • FLEURSread-aloud research set, 153 clips
    Aisyah 1.0 Pro 3.42%AssemblyAI Universal 2 6.52%
  • Scripted readingRead speech, 154 clips
    Aisyah 1.0 Pro 1.30%AssemblyAI Universal 2 3.61%
  • Common Voicevolunteer read-aloud, 141 clips
    Aisyah 1.0 Pro 3.74%AssemblyAI Universal 2 5.78%

Universal 2 makes fewer mistakes on 0 of 12

No kind of audio in this benchmark. It comes closest on Common Voice: 5.78% against 3.74%.

By background noise and clip set

The same 1,899 clips, split another way. The lower figure is in bold.

  • Clean audio1,364 clips
    Aisyah 1.0 Pro 4.70%AssemblyAI Universal 2 18.89%
  • Some background noise277 clips
    Aisyah 1.0 Pro 6.12%AssemblyAI Universal 2 13.15%
  • Noisy audio249 clips
    Aisyah 1.0 Pro 12.16%AssemblyAI Universal 2 27.61%
  • Published clips820 clips anyone can check
    Aisyah 1.0 Pro 4.89%AssemblyAI Universal 2 14.86%
  • Held-back clips1,079 clips no model has seen
    Aisyah 1.0 Pro 6.58%AssemblyAI Universal 2 22.25%

Hear the difference

Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.

Credit card eligibility question

Phone calls, 1.9 s
What was said
Credit card am I eligible?
Aisyah 1.0 Pro
Credit card am I eligible?
AssemblyAI Universal 2
Perdana Menteri, adakah saya wajib?
What to listen for

Aisyah 1.0 Pro returned "Credit card am I eligible?", matching the reference word for word. AssemblyAI Universal 2 returned "Perdana Menteri, adakah saya wajib?", which shares no words with the reference.

Account number on a call

Phone calls, 2.5 s
What was said
102105
Aisyah 1.0 Pro
one, zero, two, one, zero, five
AssemblyAI Universal 2
202105
What to listen for

The reference is the number 102105. Aisyah 1.0 Pro wrote it as words, "one, zero, two, one, zero, five", with every digit in order. AssemblyAI Universal 2 returned "202105", which differs from the reference in the first digit.

Street interview about films

Street interviews, 10.1 s
What was said
Filem Malaysia kurang sikit, tapi aa I suka ah macam cerita-cerita gangster semua. Filem apa you suka dan kenapa you suka? Macam KL Gangster, I tengok semua. Juvana.
Aisyah 1.0 Pro
Filem Malaysia kurang sikit tapi saya suka macam cerita-cerita gangster semua. Kenapa you suka dan kenapa? Macam cat gangster. Saya tengok semua Jovana.
AssemblyAI Universal 2
Film Asia kurang sikit tapi I suka macam cerita-cerita gangster semua Kenapa you suka dan kenapa you suka? Macam Kedang Gangster, I tengok semua Giovanna
What to listen for

Both outputs differ from the reference in places. For the question "Filem apa you suka dan kenapa you suka?", AssemblyAI Universal 2 wrote "Kenapa you suka dan kenapa you suka?" and Aisyah 1.0 Pro wrote "Kenapa you suka dan kenapa?". Universal 2 also has "I tengok semua" as in the reference, where Aisyah 1.0 Pro wrote "Saya tengok semua". Aisyah 1.0 Pro has the opening "Filem Malaysia" as in the reference, where Universal 2 wrote "Film Asia".

Aisyah vs AssemblyAI Universal, answered.

Is Aisyah more accurate than AssemblyAI Universal 2 for Malaysian speech?

On Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has a word error rate of 5.79% across 1,899 clips and AssemblyAI Universal 2 has 18.80%. Aisyah 1.0 Pro has the lower rate in all 12 categories. Revolab built both the benchmark and Aisyah, so test on your own audio as well.

AssemblyAI Universal 2 accuracy on Malaysian phone calls

On the 229 Telephony clips in Revolab's Malaysian speech benchmark, AssemblyAI Universal 2 has a word error rate of 51.22% and Aisyah 1.0 Pro has 9.55%, a gap of 41.67 percentage points. Only Short inputs shows a wider gap between the two models.

Does AssemblyAI Universal 2 have a lower error rate than Aisyah in any category?

No. On Revolab's Malaysian speech benchmark, AssemblyAI Universal 2 does not have a lower word error rate than Aisyah 1.0 Pro in any of the 12 categories. The gaps are narrowest in Common Voice, 5.78% against 3.74% for Aisyah 1.0 Pro, and Read speech, 3.61% against 1.30%.

Which is more accurate on noisy Malaysian audio, Aisyah or AssemblyAI Universal 2?

On the 249 clips tagged noisy in Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has a word error rate of 12.16% and AssemblyAI Universal 2 has 27.61%. Aisyah 1.0 Pro also has the lower word error rate on the clips tagged clean and moderate.

Check the numbers yourself.

Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.

All the numbersEvery figure for all 2 models on this page, as tables

Overall

Word error rate for each model. Lower is better; the lowest in each column is highlighted.
ModelAll clips1,899 clipsPublished820 clipsHeld back1,079 clips
Aisyah 1.0 Pro5.79% (fewer mistakes)4.89% (fewer mistakes)6.58% (fewer mistakes)
AssemblyAI Universal 218.80%14.86%22.25%

By kind of audio

Word error rate by kind of audio. Lower is better; the lowest in each row is highlighted.
Kind of audioAisyah 1.0 ProUniversal 2
Phone callsTelephony, 229 clips9.55% (fewer mistakes)51.22%
Short repliesShort inputs, 155 clips3.28% (fewer mistakes)46.15%
Scripted readingRead speech, 154 clips1.30% (fewer mistakes)3.61%
PodcastsPodcast, 155 clips4.57% (fewer mistakes)13.38%
Drama156 clips5.91% (fewer mistakes)15.50%
Animation150 clips5.66% (fewer mistakes)13.54%
News156 clips3.48% (fewer mistakes)30.19%
Parliament155 clips3.79% (fewer mistakes)10.81%
Street interviews143 clips14.24% (fewer mistakes)23.35%
Singing152 clips11.01% (fewer mistakes)37.45%
Common Voicevolunteer read-aloud, 141 clips3.74% (fewer mistakes)5.78%
FLEURSread-aloud research set, 153 clips3.42% (fewer mistakes)6.52%

By background noise

Word error rate by background noise. Lower is better; the lowest in each row is highlighted.
BackgroundAisyah 1.0 ProUniversal 2
Clean audio1,364 clips4.70% (fewer mistakes)18.89%
Some background noise277 clips6.12% (fewer mistakes)13.15%
Noisy audio249 clips12.16% (fewer mistakes)27.61%

Types of mistake

Types of mistake per 100 words spoken. Lower is better.
ModelWrong wordsper 100 wordsInvented wordsper 100 wordsDropped wordsper 100 words
Aisyah 1.0 Pro3.48% (fewer mistakes)0.62% (fewer mistakes)1.69% (fewer mistakes)
AssemblyAI Universal 211.07%0.70%7.02%
How the test worksClips, scoring, noise tags and what is not compared
  • 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
  • Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
  • Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
  • Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
  • Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
  • Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
  • A second AssemblyAI run, universal-3-5-pro, is not listed separately: its outputs were identical to Universal 2 on every clip.

Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.

Compare Aisyah with other models

See every model in one table

Test it on your own audio.

Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.