Malaysian speech to text benchmark
Aisyah vs Deepgram Nova: Malaysian speech to text compared
How Aisyah 1.0 Pro and Deepgram Nova 3 transcribe Malaysian speech, from phone calls and street interviews to parliament and podcasts, measured on the same 1,899 clips of Revolab's benchmark.
- Aisyah 1.0 Pro
- Fewer mistakes overall, and on all 12 kinds of audio. Biggest leads: phone calls and short replies.
- Deepgram Nova 3
- Not ahead of Aisyah 1.0 Pro on any kind of audio, noise level or clip set. Closest on FLEURS: 7.55% against 3.42%.
Revolab built this benchmark and Aisyah. Results as of 12 August 2026. How we tested
- Aisyah 1.0 Pro5.79%
- Deepgram Nova 331.99%
Where Aisyah and Nova 3 each make fewer mistakes
Aisyah 1.0 ProDeepgram Nova 3Word error rate. Further left, fewer mistakes.
Aisyah 1.0 Pro makes fewer mistakes on 12 of 12
- Phone callsTelephony, 229 clipsAisyah 1.0 Pro 9.55%Deepgram Nova 3 65.39%
- Short repliesShort inputs, 155 clipsAisyah 1.0 Pro 3.28%Deepgram Nova 3 39.71%
- Parliament155 clipsAisyah 1.0 Pro 3.79%Deepgram Nova 3 38.36%
- PodcastsPodcast, 155 clipsAisyah 1.0 Pro 4.57%Deepgram Nova 3 37.00%
- Singing152 clipsAisyah 1.0 Pro 11.01%Deepgram Nova 3 41.23%
- Street interviews143 clipsAisyah 1.0 Pro 14.24%Deepgram Nova 3 43.74%
- Drama156 clipsAisyah 1.0 Pro 5.91%Deepgram Nova 3 35.21%
- News156 clipsAisyah 1.0 Pro 3.48%Deepgram Nova 3 29.75%
- Animation150 clipsAisyah 1.0 Pro 5.66%Deepgram Nova 3 31.91%
- Common Voicevolunteer read-aloud, 141 clipsAisyah 1.0 Pro 3.74%Deepgram Nova 3 14.01%
- Scripted readingRead speech, 154 clipsAisyah 1.0 Pro 1.30%Deepgram Nova 3 9.18%
- FLEURSread-aloud research set, 153 clipsAisyah 1.0 Pro 3.42%Deepgram Nova 3 7.55%
Nova 3 makes fewer mistakes on 0 of 12
No kind of audio in this benchmark. It comes closest on FLEURS: 7.55% against 3.42%.
By background noise and clip set
The same 1,899 clips, split another way. The lower figure is in bold.
- Clean audio1,364 clipsAisyah 1.0 Pro 4.70%Deepgram Nova 3 31.86%
- Some background noise277 clipsAisyah 1.0 Pro 6.12%Deepgram Nova 3 27.04%
- Noisy audio249 clipsAisyah 1.0 Pro 12.16%Deepgram Nova 3 41.02%
- Published clips820 clips anyone can checkAisyah 1.0 Pro 4.89%Deepgram Nova 3 25.63%
- Held-back clips1,079 clips no model has seenAisyah 1.0 Pro 6.58%Deepgram Nova 3 37.55%
Hear the difference
Real clips from the published half of the benchmark. Play the audio, then read what each model returned, unedited.
Unit number, both match
Phone calls, 3.4 s- What was said
- Yes, my unit number is 2105.
- Aisyah 1.0 Pro
- Yes, my unit number is two one zero five.
- Deepgram Nova 3
- yes my unit number is 2105
What to listen for
Both outputs match the reference. Aisyah 1.0 Pro writes the number as words and Deepgram Nova 3 writes it as digits, and the benchmark scores the two forms as the same.
Credit card eligibility question
Phone calls, 1.9 s- What was said
- Credit card am I eligible?
- Aisyah 1.0 Pro
- Credit card am I eligible?
- Deepgram Nova 3
- credit card ml juga
What to listen for
Aisyah 1.0 Pro's output is the same as the reference, word for word. Deepgram Nova 3 has "credit card" as in the reference, then "ml juga" where the reference has "am I eligible".
Malay and English street interview
Street interviews, 4.9 s- What was said
- Apa bagus dia ada yang tak bagus dia. Depends lah. Depends pada individu. Haa, aa ah, individu lah.
- Aisyah 1.0 Pro
- Apa bagus dia, ada yang tak bagus dia. depends lah. depends pada individu lah.
- Deepgram Nova 3
- Returned no text
What to listen for
The reference mixes Malay with the English word "Depends". Deepgram Nova 3 returned no text. Aisyah 1.0 Pro's output matches the reference apart from leaving out "Haa, aa ah" and one "individu".
Aisyah vs Deepgram Nova, answered.
Is Deepgram Nova accurate for Malaysian speech?
On Revolab's Malaysian speech benchmark of 1,899 clips, Deepgram Nova 3 has a word error rate of 31.99% and Aisyah 1.0 Pro has 5.79%, where lower is better. Nova 3 does not have the lower word error rate in any of the 12 categories. Revolab built both the benchmark and Aisyah.
Best speech to text for Malaysian phone calls: Aisyah or Deepgram Nova?
In the Telephony category of Revolab's Malaysian speech benchmark, which has 229 clips, Aisyah 1.0 Pro has a word error rate of 9.55% and Deepgram Nova 3 has 65.39%, a gap of 55.84 percentage points. Two clips on this page, a unit number and a credit card question, come from this category.
How does Deepgram Nova 3 handle Manglish?
Revolab's benchmark has no separate Manglish category, but its Street interviews category has clips that mix Malay and English, such as "Depends lah. Depends pada individu." In that category, Deepgram Nova 3 has a word error rate of 43.74% and Aisyah 1.0 Pro has 14.24%.
Who built the benchmark comparing Aisyah and Deepgram Nova?
Revolab built both the benchmark and Aisyah 1.0 Pro. Of the 1,899 clips, 820 are published and 1,079 are held back. Aisyah 1.0 Pro has a word error rate of 4.89% on the published clips and 6.58% on the held-back clips; Deepgram Nova 3 has 25.63% and 37.55%.
Check the numbers yourself.
Revolab built this benchmark and trained Aisyah, so read the results with that in mind. Every published clip and every model's output is open to check.
All the numbersEvery figure for all 2 models on this page, as tables
Overall
| Model | All clips1,899 clips | Published820 clips | Held back1,079 clips |
|---|---|---|---|
| Aisyah 1.0 Pro | 5.79% (fewer mistakes) | 4.89% (fewer mistakes) | 6.58% (fewer mistakes) |
| Deepgram Nova 3 | 31.99% | 25.63% | 37.55% |
By kind of audio
| Kind of audio | Aisyah 1.0 Pro | Nova 3 |
|---|---|---|
| Phone callsTelephony, 229 clips | 9.55% (fewer mistakes) | 65.39% |
| Short repliesShort inputs, 155 clips | 3.28% (fewer mistakes) | 39.71% |
| Scripted readingRead speech, 154 clips | 1.30% (fewer mistakes) | 9.18% |
| PodcastsPodcast, 155 clips | 4.57% (fewer mistakes) | 37.00% |
| Drama156 clips | 5.91% (fewer mistakes) | 35.21% |
| Animation150 clips | 5.66% (fewer mistakes) | 31.91% |
| News156 clips | 3.48% (fewer mistakes) | 29.75% |
| Parliament155 clips | 3.79% (fewer mistakes) | 38.36% |
| Street interviews143 clips | 14.24% (fewer mistakes) | 43.74% |
| Singing152 clips | 11.01% (fewer mistakes) | 41.23% |
| Common Voicevolunteer read-aloud, 141 clips | 3.74% (fewer mistakes) | 14.01% |
| FLEURSread-aloud research set, 153 clips | 3.42% (fewer mistakes) | 7.55% |
By background noise
| Background | Aisyah 1.0 Pro | Nova 3 |
|---|---|---|
| Clean audio1,364 clips | 4.70% (fewer mistakes) | 31.86% |
| Some background noise277 clips | 6.12% (fewer mistakes) | 27.04% |
| Noisy audio249 clips | 12.16% (fewer mistakes) | 41.02% |
Types of mistake
| Model | Wrong wordsper 100 words | Invented wordsper 100 words | Dropped wordsper 100 words |
|---|---|---|---|
| Aisyah 1.0 Pro | 3.48% (fewer mistakes) | 0.62% (fewer mistakes) | 1.69% (fewer mistakes) |
| Deepgram Nova 3 | 10.97% | 0.94% | 20.09% |
How the test worksClips, scoring, noise tags and what is not compared
- 1,899 Malaysian utterances in 12 kinds of audio, from phone calls and street interviews to parliament and singing. 820 are published and 1,079 are held back.
- Word error rate counts every wrong, invented and dropped word and divides by the number of words actually said, across all clips. Lower is better.
- Transcripts are normalised before scoring, so casing, punctuation and writing a number as digits or as words do not count as errors.
- Clips are tagged for background noise: 1,364 clean, 277 moderate and 249 noisy.
- Speed is not compared. For models served over an API, the time measured is mostly network latency rather than the model itself.
- Results as of 12 August 2026, under the model names shown. Vendors update their models, so results can change.
Product and company names are trademarks of their respective owners and are used only to identify the models tested. Revolab is not affiliated with or endorsed by any of them.
Compare Aisyah with other models
See every model in one table- Aisyah vs ElevenLabs ScribeScribe v25.79% vs 6.67%
- Aisyah vs Google GeminiGemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.6 Flash5.79% vs 7.00%
- Aisyah vs YTL ILMUILMU ASR v4.25.79% vs 9.48%
- Aisyah vs OpenAI WhisperWhisper large-v35.79% vs 17.30%
- Aisyah vs Alibaba QwenQwen-Audio 3.0 ASR Flash, Qwen3-ASR-1.7B, Qwen3-ASR-0.6B5.79% vs 9.17%
- Aisyah vs AssemblyAI UniversalUniversal 25.79% vs 18.80%
Test it on your own audio.
Every new account gets 50 free minutes of speech to text and 50 of text to speech. No card required.