Skip to contentMasud Rabbani
CV
Menu
← All articles

AI & everyday life

Why does voice AI understand some accents better than others?

[2]

A closer look at training voices, fair tests, and what a Bangla speech review can teach us.[2][1]

Imagine asking your phone to set a timer. It gets your friend's words right. When you ask, it writes something strange. You repeat yourself. You slow down. You may even try to sound like someone else. In this imagined scene, a small daily task becomes a test you never agreed to take.

Here, voice AI means the part of a system that turns recorded speech into written words. This is called speech recognition [software that predicts words from sound]. We are looking at the words it writes, rather than whether a voice assistant gives a useful reply.[3]

The voices in its lessons matter[2]

Think of a student who practices with just one small group of people. That is a rough way to picture our question: what happens when a new voice does not sound like the voices in the lessons? The analogy is a starting point, not a description of a human brain.

In a 2025 study, Zachary Hopton and Eleanor Chodroff changed the mix of Catalan dialects used to train a speech model further. A dialect is a form of a language used by a particular group. A broader mix helped this model handle dialect differences. The researchers tested a specific model family and language, so this is evidence from one setting—not a promise to fix every voice tool.[2]

A conceptual flow diagram. Training voices shape the speech model. New speech passes through recording conditions into the model, which produces written words. Review both the training examples and the recording.
A simplified picture of speech recognition: training examples and recording conditions can both affect the result.[2][3]Original conceptual diagram for this article. No measured data or publisher artwork reproduced.

Our research on Bangla speech recognition

Research spotlight · Coauthored work

Natural Language Processing for recognizing Bangla speech with regular and regional dialects: A survey of algorithms and approaches

2024 · 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC)

In 2024, I coauthored a survey of methods, datasets, and challenges in Bangla speech recognition, including regional dialects. Our review brings together evidence on limited data, dialect variation, and computing needs. It also discusses dialect adaptation [adjusting a model for different speech varieties] and methods for speech varieties with little recorded data.[1]

Why it matters: I see this review as a starting map for researchers planning more representative speech tools. It can help frame questions about methods and missing data. Use it as background for Bangla speech recognition, regional dialects, and research gaps—not as proof of a current product's accuracy.[1]

View citation & all coauthors

Paramita Basak Upama, Parama Sridevi, Masud Rabbani, Kazi Shafiul Alam, Munirul Haque, Sheikh Iqbal Ahamed. “Natural Language Processing for recognizing Bangla speech with regular and regional dialects: A survey of algorithms and approaches”. 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), 312–319. (2024). https://doi.org/10.1109/COMPSAC61105.2024.00051.

If this work supports your research, cite the original paper rather than this blog explanation.

The question I would bring to a design meeting is simple: whose voices are in these examples, and whose voices are missing? I would want the team to answer that before showing a polished demo.

A fair test needs more than one score

One common measure is word error rate [the share of word mistakes compared with the correct transcript]. It counts words a system changes, leaves out, or adds, then divides the total by the number of words in the correct transcript. Lower is better on this measure; it does not measure every part of a voice tool's usefulness.[3]

Sound quality needs attention too. A 2024 preprint [research shared before formal peer review] found that recording practices affected recognition results across locations in its data. That is a reason to check recording conditions when comparing groups. It does not show that recording quality explains every difference.[3]

A quick thought experiment

Imagine a team adds many recordings from the same small group, then reports one score for everyone. What would you ask next?

More ways to gather real speech

In May 2025, Mozilla described a Common Voice mode in which people answer prompts in their own words, alongside read-aloud recordings. A June 2026 release added 10 datasets and seven languages. Those announcements describe growth in available speech data; they do not demonstrate equal accuracy for every accent.[4][5]

Sources & further reading

  1. Natural Language Processing for recognizing Bangla speech with regular and regional dialects: A survey of algorithms and approachesParamita Basak Upama, Parama Sridevi, Masud Rabbani, Kazi Shafiul Alam, Munirul Haque, Sheikh Iqbal Ahamed · 2024
  2. The Impact of Dialect Variation on Robust Automatic Speech Recognition for CatalanZachary Hopton and Eleanor Chodroff · May 2025
  3. Reexamining Racial Disparities in Automatic Speech Recognition Performance: The Role of Confounding by ProvenanceChangye Li, Trevor Cohen, and Serguei Pakhomov · July 19, 2024 · preprint
  4. Spontaneous Speech Mode is Coming to Common VoiceJess Rose · Mozilla Common Voice · May 14, 2025
  5. Release LIVE: MCV Scripted Speech v26.0 and Spontaneous Speech v4.0Alexandra Fort · Mozilla Common Voice · June 19, 2026

Disclaimer: This post is for education, not medical or other professional advice; information may contain errors or become outdated, no outcome is guaranteed, views are my own, and third-party materials remain subject to their owners' rights.