Malaysian speech to text benchmark
Aisyah 1.0 Pro vs Google Gemini for Malaysian speech to text
Word error rates, where lower is better, for Aisyah 1.0 Pro, Google Gemini 2.5 Pro, Google Gemini 2.5 Flash and Google Gemini 3.6 Flash on Revolab's Malaysian speech benchmark: 1,899 clips across 12 categories.
- Aisyah 1.0 Pro
- Fewer mistakes overall, and on 6 of 12 kinds of audio. Biggest leads: short replies and phone calls.
- Google Gemini 2.5 Pro
- Fewer mistakes on 6 of 12: street interviews, FLEURS, singing, Common Voice, parliament and news. Also on noisy audio.
Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested
- Aisyah 1.0 Pro5.79%
- Google Gemini 2.5 Pro7.00%
- Google Gemini 2.5 Flash10.76%
- Google Gemini 3.6 Flash24.55%
Where Aisyah and Gemini 2.5 Pro each make fewer mistakes
Aisyah 1.0 ProGoogle Gemini 2.5 ProWord error rate. Further left, fewer mistakes.
Aisyah 1.0 Pro makes fewer mistakes on 6 of 12
- Short repliesShort inputs, 155 clipsAisyah 1.0 Pro 3.28%Google Gemini 2.5 Pro 35.29%
- Phone callsTelephony, 229 clipsAisyah 1.0 Pro 9.55%Google Gemini 2.5 Pro 19.80%
- Animation150 clipsAisyah 1.0 Pro 5.66%Google Gemini 2.5 Pro 7.82%
- Scripted readingRead speech, 154 clipsAisyah 1.0 Pro 1.30%Google Gemini 2.5 Pro 3.34%
- PodcastsPodcast, 155 clipsAisyah 1.0 Pro 4.57%Google Gemini 2.5 Pro 6.09%
- Drama156 clipsAisyah 1.0 Pro 5.91%Google Gemini 2.5 Pro 6.94%
Gemini 2.5 Pro makes fewer mistakes on 6 of 12
- Street interviews143 clipsAisyah 1.0 Pro 14.24%Google Gemini 2.5 Pro 12.79%
- FLEURSread-aloud research set, 153 clipsAisyah 1.0 Pro 3.42%Google Gemini 2.5 Pro 2.26%
- Singing152 clipsAisyah 1.0 Pro 11.01%Google Gemini 2.5 Pro 10.15%
- Common Voicevolunteer read-aloud, 141 clipsAisyah 1.0 Pro 3.74%Google Gemini 2.5 Pro 3.16%
- Parliament155 clipsAisyah 1.0 Pro 3.79%Google Gemini 2.5 Pro 3.28%
- News156 clipsAisyah 1.0 Pro 3.48%Google Gemini 2.5 Pro 3.24%
By background noise and clip set
The same 1,899 clips, split another way. The lower figure is in bold.
- Clean audio1,364 clipsAisyah 1.0 Pro 4.70%Google Gemini 2.5 Pro 6.43%
- Some background noise277 clipsAisyah 1.0 Pro 6.12%Google Gemini 2.5 Pro 6.65%
- Noisy audio249 clipsAisyah 1.0 Pro 12.16%Google Gemini 2.5 Pro 11.13%
- Published clips820 clips anyone can checkAisyah 1.0 Pro 4.89%Google Gemini 2.5 Pro 5.30%
- Held-back clips1,079 clips no model has seenAisyah 1.0 Pro 6.58%Google Gemini 2.5 Pro 8.49%
The other Google models on this page
- Google Gemini 2.5 Flash (10.76% overall) makes fewer mistakes than Aisyah 1.0 Pro only on FLEURS (3.29% against 3.42%).
- Google Gemini 3.6 Flash (24.55% overall) makes fewer mistakes than Aisyah 1.0 Pro only on FLEURS (2.79% against 3.42%).
Every figure for every model is in the full results below.
Hear the difference
Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.
Street interview about Malaysian films
Street interviews, 10.1 s- What was said
- Filem Malaysia kurang sikit, tapi aa I suka ah macam cerita-cerita gangster semua. Filem apa you suka dan kenapa you suka? Macam KL Gangster, I tengok semua. Juvana.
- Aisyah 1.0 Pro
- Filem Malaysia kurang sikit tapi saya suka macam cerita-cerita gangster semua. Kenapa you suka dan kenapa? Macam cat gangster. Saya tengok semua Jovana.
- Google Gemini 2.5 Pro
- Filem Malaysia kurang sikit tapi aa I suka macam cerita-cerita gangster semua. Macam KL Gangster, I tengok semua, Juvana.
- Google Gemini 2.5 Flash
- filem Malaysia kurang sikit tapi ah I suka macam cerita-cerita gangster semua. Filem apa yang you suka dan kenapa suka? Macam Korean gangster. I tengok semua. Jovanna.
What to listen for
Google Gemini 2.5 Pro wrote "KL Gangster" and "Juvana" as in the reference, but its output has no text for the question "Filem apa you suka dan kenapa you suka?". Google Gemini 2.5 Flash wrote that question as "Filem apa yang you suka dan kenapa suka?", and Aisyah 1.0 Pro wrote "Kenapa you suka dan kenapa?". Aisyah 1.0 Pro also wrote "saya" where the reference has "I", and "cat gangster" for "KL Gangster". On this clip, Gemini 2.5 Flash's text is closer to the reference than Aisyah 1.0 Pro's.
Account number on a call
Phone calls, 2.5 s- What was said
- 102105
- Aisyah 1.0 Pro
- one, zero, two, one, zero, five
- Google Gemini 2.5 Pro
- 10210
- Google Gemini 2.5 Flash
- one zero two one zero
What to listen for
The reference is "102105". Aisyah 1.0 Pro wrote the number as words, "one, zero, two, one, zero, five", which the benchmark scores the same as digits. Google Gemini 2.5 Pro returned "10210" and Google Gemini 2.5 Flash returned "one zero two one zero", both without the last digit of the reference.
Single-word Telephony clip
Phone calls, 1.2 s- What was said
- Okay.
- Aisyah 1.0 Pro
- Okay
- Google Gemini 2.5 Pro
- [ 0m0s - 0m1s ] Okay.
- Google Gemini 3.6 Flash
- It looks like you didn't attach or link the audio, video, or text for the speech! Please provide the audio/video file, a link, or specify which famous speech you would like transcribed, and I will be happy to generate it for you.
What to listen for
Aisyah 1.0 Pro returned "Okay", matching the single-word reference. Google Gemini 2.5 Pro returned "Okay." with bracketed text in front of it, and Google Gemini 3.6 Flash returned a message asking for an audio or video file instead of a transcript.
Aisyah vs Google Gemini, answered.
Is Aisyah more accurate than Gemini 2.5 Pro for Malaysian speech?
On Revolab's Malaysian speech benchmark, Aisyah 1.0 Pro has the lower word error rate across all 1,899 clips: 5.79% against 7.00% for Google Gemini 2.5 Pro. Each has the lower WER in 6 of the 12 categories. Revolab built both the benchmark and Aisyah.
Which Gemini model is most accurate for Malaysian speech to text?
Of the Gemini models tested on Revolab's Malaysian speech benchmark, Google Gemini 2.5 Pro has the lowest word error rate across all 1,899 clips at 7.00%, followed by Google Gemini 2.5 Flash at 10.76% and Google Gemini 3.6 Flash at 24.55%. Aisyah 1.0 Pro is at 5.79%.
How accurate is Gemini at transcribing Malaysian phone calls?
In the Telephony category of Revolab's Malaysian speech benchmark, Google Gemini 2.5 Pro has a word error rate of 19.80%, Google Gemini 2.5 Flash 24.87% and Google Gemini 3.6 Flash 84.42%, where lower is better. Aisyah 1.0 Pro has 9.55%.
Is Gemini 2.5 Pro better than Aisyah on noisy audio?
On clips tagged noisy in Revolab's Malaysian speech benchmark, yes: Google Gemini 2.5 Pro has a word error rate of 11.13% and Aisyah 1.0 Pro has 12.16%. On clean clips, Aisyah 1.0 Pro has 4.70% and Gemini 2.5 Pro has 6.43%.
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
All the numbersEvery figure for all 4 models on this page, as tables
Overall
| Model | All clips1,899 clips | Published820 clips | Held back1,079 clips |
|---|---|---|---|
| Aisyah 1.0 Pro | 5.79% (fewer mistakes) | 4.89% (fewer mistakes) | 6.58% (fewer mistakes) |
| Google Gemini 2.5 Pro | 7.00% | 5.30% | 8.49% |
| Google Gemini 2.5 Flash | 10.76% | 9.05% | 12.25% |
| Google Gemini 3.6 Flash | 24.55% | 15.01% | 32.87% |
By kind of audio
| Kind of audio | Aisyah 1.0 Pro | Gemini 2.5 Pro | Gemini 2.5 Flash | Gemini 3.6 Flash |
|---|---|---|---|---|
| Phone callsTelephony, 229 clips | 9.55% (fewer mistakes) | 19.80% | 24.87% | 84.42% |
| Short repliesShort inputs, 155 clips | 3.28% (fewer mistakes) | 35.29% | 81.79% | 753.63% |
| Scripted readingRead speech, 154 clips | 1.30% (fewer mistakes) | 3.34% | 7.04% | 10.75% |
| PodcastsPodcast, 155 clips | 4.57% (fewer mistakes) | 6.09% | 10.56% | 10.82% |
| Drama156 clips | 5.91% (fewer mistakes) | 6.94% | 9.52% | 15.82% |
| Animation150 clips | 5.66% (fewer mistakes) | 7.82% | 10.02% | 12.27% |
| News156 clips | 3.48% | 3.24% (fewer mistakes) | 3.93% | 5.08% |
| Parliament155 clips | 3.79% | 3.28% (fewer mistakes) | 6.47% | 7.24% |
| Street interviews143 clips | 14.24% | 12.79% (fewer mistakes) | 19.85% | 19.81% |
| Singing152 clips | 11.01% | 10.15% (fewer mistakes) | 20.30% | 28.14% |
| Common Voicevolunteer read-aloud, 141 clips | 3.74% | 3.16% (fewer mistakes) | 4.09% | 4.37% |
| FLEURSread-aloud research set, 153 clips | 3.42% | 2.26% (fewer mistakes) | 3.29% | 2.79% |
By background noise
| Background | Aisyah 1.0 Pro | Gemini 2.5 Pro | Gemini 2.5 Flash | Gemini 3.6 Flash |
|---|---|---|---|---|
| Clean audio1,364 clips | 4.70% (fewer mistakes) | 6.43% | 10.07% | 27.78% |
| Some background noise277 clips | 6.12% (fewer mistakes) | 6.65% | 9.20% | 13.39% |
| Noisy audio249 clips | 12.16% | 11.13% (fewer mistakes) | 17.42% | 20.70% |
Types of mistake
| Model | Wrong wordsper 100 words | Invented wordsper 100 words | Dropped wordsper 100 words |
|---|---|---|---|
| Aisyah 1.0 Pro | 3.48% (fewer mistakes) | 0.62% (fewer mistakes) | 1.69% (fewer mistakes) |
| Google Gemini 2.5 Pro | 3.91% | 1.00% | 2.10% |
| Google Gemini 2.5 Flash | 5.50% | 2.48% | 2.78% |
| Google Gemini 3.6 Flash | 7.44% | 14.35% | 2.75% |
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Compare Aisyah with other models
See every model in one table- Aisyah vs ElevenLabs ScribeScribe v25.79% vs 6.67%
- Aisyah vs YTL ILMUILMU ASR v4.25.79% vs 9.48%
- Aisyah vs OpenAI WhisperWhisper large-v35.79% vs 17.30%
- Aisyah vs Alibaba QwenQwen-Audio 3.0 ASR Flash, Qwen3-ASR-1.7B, Qwen3-ASR-0.6B5.79% vs 9.17%
- Aisyah vs AssemblyAI UniversalUniversal 25.79% vs 18.80%
- Aisyah vs Deepgram NovaNova 35.79% vs 31.99%
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.