Malaysian speech to text benchmark
Aisyah 1.0 Pro vs OpenAI Whisper large-v3 for Malaysian speech to text
How Aisyah 1.0 Pro and OpenAI Whisper large-v3 compare on 1,899 clips of Malaysian speech, from parliament and news to phone calls, singing and street interviews.
- Aisyah 1.0 Pro
- Fewer mistakes overall, and on all 12 kinds of audio. Biggest leads: phone calls and short replies.
- OpenAI Whisper large-v3
- Not ahead of Aisyah 1.0 Pro on any kind of audio, noise level or clip set. Closest on scripted reading: 2.78% against 1.30%.
Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested
- Aisyah 1.0 Pro5.79%
- OpenAI Whisper large-v317.30%
Where Aisyah and Whisper large-v3 each make fewer mistakes
Aisyah 1.0 ProOpenAI Whisper large-v3Word error rate. Further left, fewer mistakes.
Aisyah 1.0 Pro makes fewer mistakes on 12 of 12
- Phone callsTelephony, 229 clipsAisyah 1.0 Pro 9.55%OpenAI Whisper large-v3 62.89%
- Short repliesShort inputs, 155 clipsAisyah 1.0 Pro 3.28%OpenAI Whisper large-v3 33.99%
- News156 clipsAisyah 1.0 Pro 3.48%OpenAI Whisper large-v3 22.77%
- Animation150 clipsAisyah 1.0 Pro 5.66%OpenAI Whisper large-v3 20.42%
- Singing152 clipsAisyah 1.0 Pro 11.01%OpenAI Whisper large-v3 25.10%
- PodcastsPodcast, 155 clipsAisyah 1.0 Pro 4.57%OpenAI Whisper large-v3 13.01%
- Parliament155 clipsAisyah 1.0 Pro 3.79%OpenAI Whisper large-v3 9.35%
- Drama156 clipsAisyah 1.0 Pro 5.91%OpenAI Whisper large-v3 11.35%
- Street interviews143 clipsAisyah 1.0 Pro 14.24%OpenAI Whisper large-v3 18.43%
- FLEURSread-aloud research set, 153 clipsAisyah 1.0 Pro 3.42%OpenAI Whisper large-v3 5.49%
- Common Voicevolunteer read-aloud, 141 clipsAisyah 1.0 Pro 3.74%OpenAI Whisper large-v3 5.47%
- Scripted readingRead speech, 154 clipsAisyah 1.0 Pro 1.30%OpenAI Whisper large-v3 2.78%
Whisper large-v3 makes fewer mistakes on 0 of 12
No kind of audio in this benchmark. It comes closest on scripted reading: 2.78% against 1.30%.
By background noise and clip set
The same 1,899 clips, split another way. The lower figure is in bold.
- Clean audio1,364 clipsAisyah 1.0 Pro 4.70%OpenAI Whisper large-v3 17.83%
- Some background noise277 clipsAisyah 1.0 Pro 6.12%OpenAI Whisper large-v3 13.76%
- Noisy audio249 clipsAisyah 1.0 Pro 12.16%OpenAI Whisper large-v3 19.82%
- Published clips820 clips anyone can checkAisyah 1.0 Pro 4.89%OpenAI Whisper large-v3 15.62%
- Held-back clips1,079 clips no model has seenAisyah 1.0 Pro 6.58%OpenAI Whisper large-v3 18.76%
Hear the difference
Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.
Credit card eligibility question
Phone calls, 1.9 s- What was said
- Credit card am I eligible?
- Aisyah 1.0 Pro
- Credit card am I eligible?
- OpenAI Whisper large-v3
- Radhika, am I eligible?
What to listen for
The reference is "Credit card am I eligible?" and Aisyah 1.0 Pro returned exactly that. OpenAI Whisper large-v3 returned "Radhika, am I eligible?", which keeps the end of the question but has the single word Radhika where the reference has the two words Credit card.
Manglish street interview
Street interviews, 4.9 s- What was said
- Apa bagus dia ada yang tak bagus dia. Depends lah. Depends pada individu. Haa, aa ah, individu lah.
- Aisyah 1.0 Pro
- Apa bagus dia, ada yang tak bagus dia. depends lah. depends pada individu lah.
- OpenAI Whisper large-v3
- Bagus dia ada yang tak bagus dia Depends lah Individu lah
What to listen for
The reference mixes Malay and English. Aisyah 1.0 Pro returned "Apa bagus dia, ada yang tak bagus dia. depends lah. depends pada individu lah.", leaving out "Haa, aa ah" and the repeated individu. Whisper large-v3 returned "Bagus dia ada yang tak bagus dia Depends lah Individu lah", leaving out the opening Apa, the phrase Depends pada individu and "Haa, aa ah".
Both match on Okay
Phone calls, 1.2 s- What was said
- Okay.
- Aisyah 1.0 Pro
- Okay
- OpenAI Whisper large-v3
- Okey.
What to listen for
Against the reference "Okay.", Aisyah 1.0 Pro returned "Okay" and Whisper large-v3 returned "Okey.", and the benchmark accepts Okay and Okey as the same word. Both outputs count as a match, so neither model is ahead on this clip.
Aisyah vs OpenAI Whisper, answered.
How accurate is OpenAI Whisper for Malaysian speech?
On Revolab's Malaysian speech benchmark of 1,899 clips, OpenAI Whisper large-v3 has a word error rate of 17.30%, against 5.79% for Aisyah 1.0 Pro. Whisper large-v3's category results range from 2.78% in Read speech to 62.89% in Telephony.
Is Whisper or Aisyah better for call centre transcription?
In the Telephony category of Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has a word error rate of 9.55% and OpenAI Whisper large-v3 has 62.89%, a difference of 53.34 percentage points. This is the widest gap between the two models in any category.
Is Aisyah or Whisper better for Manglish?
Revolab's Malaysian speech benchmark has no separate Manglish category, but its Street interviews category includes clips that mix Malay and English. There, Aisyah 1.0 Pro has a word error rate of 14.24% and OpenAI Whisper large-v3 has 18.43%, a difference of 4.19 percentage points.
Does Whisper beat Aisyah in any category?
No. On Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has a lower word error rate than OpenAI Whisper large-v3 in all 12 categories, on both the published and held-back clips, and at every noise level. The closest category is Read speech, where Aisyah 1.0 Pro has 1.30% and Whisper large-v3 has 2.78%.
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
All the numbersEvery figure for all 2 models on this page, as tables
Overall
| Model | All clips1,899 clips | Published820 clips | Held back1,079 clips |
|---|---|---|---|
| Aisyah 1.0 Pro | 5.79% (fewer mistakes) | 4.89% (fewer mistakes) | 6.58% (fewer mistakes) |
| OpenAI Whisper large-v3 | 17.30% | 15.62% | 18.76% |
By kind of audio
| Kind of audio | Aisyah 1.0 Pro | Whisper large-v3 |
|---|---|---|
| Phone callsTelephony, 229 clips | 9.55% (fewer mistakes) | 62.89% |
| Short repliesShort inputs, 155 clips | 3.28% (fewer mistakes) | 33.99% |
| Scripted readingRead speech, 154 clips | 1.30% (fewer mistakes) | 2.78% |
| PodcastsPodcast, 155 clips | 4.57% (fewer mistakes) | 13.01% |
| Drama156 clips | 5.91% (fewer mistakes) | 11.35% |
| Animation150 clips | 5.66% (fewer mistakes) | 20.42% |
| News156 clips | 3.48% (fewer mistakes) | 22.77% |
| Parliament155 clips | 3.79% (fewer mistakes) | 9.35% |
| Street interviews143 clips | 14.24% (fewer mistakes) | 18.43% |
| Singing152 clips | 11.01% (fewer mistakes) | 25.10% |
| Common Voicevolunteer read-aloud, 141 clips | 3.74% (fewer mistakes) | 5.47% |
| FLEURSread-aloud research set, 153 clips | 3.42% (fewer mistakes) | 5.49% |
By background noise
| Background | Aisyah 1.0 Pro | Whisper large-v3 |
|---|---|---|
| Clean audio1,364 clips | 4.70% (fewer mistakes) | 17.83% |
| Some background noise277 clips | 6.12% (fewer mistakes) | 13.76% |
| Noisy audio249 clips | 12.16% (fewer mistakes) | 19.82% |
Types of mistake
| Model | Wrong wordsper 100 words | Invented wordsper 100 words | Dropped wordsper 100 words |
|---|---|---|---|
| Aisyah 1.0 Pro | 3.48% (fewer mistakes) | 0.62% (fewer mistakes) | 1.69% (fewer mistakes) |
| OpenAI Whisper large-v3 | 9.17% | 2.73% | 5.40% |
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Compare Aisyah with other models
See every model in one table- Aisyah vs ElevenLabs ScribeScribe v25.79% vs 6.67%
- Aisyah vs Google GeminiGemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.6 Flash5.79% vs 7.00%
- Aisyah vs YTL ILMUILMU ASR v4.25.79% vs 9.48%
- Aisyah vs Alibaba QwenQwen-Audio 3.0 ASR Flash, Qwen3-ASR-1.7B, Qwen3-ASR-0.6B5.79% vs 9.17%
- Aisyah vs AssemblyAI UniversalUniversal 25.79% vs 18.80%
- Aisyah vs Deepgram NovaNova 35.79% vs 31.99%
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.