Development of a Novel Automatic Speech Recognition to Reduce Racial and Gender Bias

CSEF · 2023 Cognitive Science Third Award

Overview

Problem Statement The growth of automatic speech recognition in the past decade, including Siri, Alexa, and Cortana, has unveiled a serious problem: the implicit racial bias. Studies have shown that all five programs from leading technology companies showed significant racial disparities. Training systems exclude accents and other ways of speaking that have unique linguistic features, effectively censoring nonstandard voices from utilizing ASR. An ASR transcribing program should utilize a diverse database in order to be more inclusive and efficient for all. Research Question § To develop a novel ASR transcribing program that can discern different ethnic accents effectively and evaluate its efficacy compared to Apple, Google, Cortana, and Microsoft. Introduction Automated speech recognition (ASR) systems, which use sophisticated machine-learning algorithms to convert spoken language to text, have become increasingly widespread, powering popular virtual assistants, facilitating automated closed captioning, and enabling digital dictation platforms for health care. The quality of ASR systems has dramatically improved, due to advances in deep learning and the collection of large-scale datasets used to train the systems. There is concern, however, that these tools do not work equally well for all subgroups of the population. By quantifying the disparity and possibly creating a novel ASR, these tools can be made more equitable. Materials and Methods • Use randomly select samples of voices from a variety of ethnicities, genders, and states from available voice and accent databases. The audio samples were 30-60 seconds long. • Transcribe speech samples for each race, and for men and women. • The races included were Hispanic, African American, and Caucasian. • A novel Automated Speech Recognition Program was written in Python using IBM Watson Cloud based on more diverse data samples. • The word error rate was calculated using a code written in Python. • Compare the accuracy and word error rate for race and gender and perform statistical analyses. Results • Based on WER, Siri (38.07%) and Google (35.63%) had the least accuracy. Cortana and the novel ASR performed equally well at 29.04%, and the novel ASR performed better than Siri and Google. Microsoft Word performed the best overall. (Figure 3) • Within the African American group, there was a statistically significant difference between the various ASR systems using ANOVA (p-value of 0.0003). The novel ASR program (AV ASR) performed better than Siri (WER of 39.74% vs 29.48%), and the novel ASR was non-inferior to Google, Cortana, and Word. (Figure 4) • Within the Caucasian group, there was a statistically significant difference between the various ASR systems using ANOVA (p-value =0.008). My ASR program (WER = 21.38%) performed better than Siri (WER= 32.13%) and Google (WER = 26.33%). Siri performed the worst overall, and Word had the lowest WER at 14.55%. Also, Google performed significantly worse than Word (11.78% vs 9.55%). (Figure 5) • Within the Hispanic group, there was a statistically significant difference between the WER in the various ASR systems using ANOVA (p-value 0.001). Google had the highest WER of 46.04%, followed by Siri, and the novel ASR. Word and Cortana performed the best (WER = 34.45 % and WER = 28.09%). Google had a statistically significant difference compared to Cortana and Word. Siri performed significantly worse than word. However, my novel ASR program (AV ASR) was non-inferior to existing systems. (Figure 6) • When comparing based on gender, there was a statistically significant difference between groups when comparing pairs of systems. My ASR program (AV ASR) performed better than Siri (10.61% vs. 8.16%). Siri performed worse than Cortana and Word with a p-value of 0.001. Google performed worse than Word with a p-value < 0.05. • For gender within the African American group, the WER for Microsoft Word performed slightly better for women compared to men. The novel ASR was non-inferior for both genders within the Hispanic and African American groups. • Regardless of race and gender the novel ASR performed significantly better than Siri and Google ( p-value =0.004). Conclusions • Siri performed the worst overall and, in each group,, with the highest WER. • Despite advances in technology, the WER for African Americans and Hispanics was significantly higher than that of Caucasians. • Overall, the novel ASR performed significantly better than Siri and Google. • For African Americans, the novel ASR performed significantly better than Siri and had less accuracy than Word. • There were no statistically significant differences for the Hispanic population. However, the novel ASR performed slightly better than Google and Siri. Limitation of study • The racial disparities present in current ASR systems have implications for the accessibility of technology. By utilizing diverse databases, ASR systems can better serve a wider population. A representative database is vital in these modern times, which is why investing in better data collection on nonstandard varieties of English is necessary. Next year, I will focus on testing a broader range of accents and dialects as well as take my own live samples from study participants. Future Directions • The limitations include the use of recorded samples compared to live samples. There are 8 samples for each subgroup of gender and race, which could make it more difficult to generalize the results to the general population. The word error rate significance tests could easily be over- or underestimated for subgroup analysis. Next year, I aim to increase the power of the study with a larger sample size. References/Bibliography • Chen, X., Li, Z., Setlur, S. et al. Exploring racial and gender disparities in voice biometrics. Sci Rep 12, 3723 (2022). https://doi.org/10.1038/s41598-022-06673-y • Allison Koenecke, Racial disparities in automated speech recognition, https://doi.org/10.1073/pnas.1915768117 • https://blog.usa.gov/press-or-say-1-usagov-expands-its-use-of-interactive-voice-response • Anastasia Atabekova, Communication with Non-native Speakers Through the Service of Speech-To-Speech Interpreting Systems: Testing Technology Capacity and Exploring Specialists’ Views, Services – SERVICES 2021, 10.1007/978-3-030-96585-3_1, (1-17), (2022) • https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8664002/#:~:text=In%20their%20groundbreaking%20study%20of,White%2C%20standard%20American%20English%20speakers • Mengesha, Zion et al. “"I don't Think These Devices are Very Culturally Sensitive."-Impact of Automated Speech Recognition Errors on African Americans.” Frontiers in artificial intelligence vol. 4 725911. 26 Nov. 2021, doi:10.3389/frai.2021.725911

Source coverage

This record comes from a published award list, not a complete project archive. Its abstract comes from CSEF's public project showcase as archived by the Internet Archive before judging (https://web.archive.org/web/20230401224130/https://ca-csef.zfairs.com/showcase/ShowcaseInfo?f=838e60b7-ea75-46e8-865c-fde4864244b3); the version presented may differ.

Awards (2)

Competition history

  • CSEF 2023 Cognitive Science · Entry J0703

Resources

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google