A Novel End-to-End Automatic Speech Recognition (ASR) Model to Improve Tonal Pronunciation in Mandarin for Non-Native Speakers
CSEF · 2023 Computational Systems & Analysis Second Award
Overview
Tonal languages are estimated to comprise 60–70% of the world’s languages. However, tonal language pronunciation is difficult for non-native speakers, as Mandarin has 5 tones, resulting in each tone having a certain pitch. This differentiates thousands of Chinese characters, which are characterized by their phonemes, resulting in 413 distinct tonal combinations compared to 44 tonal sounds in English. As non-native speakers ourselves, we aimed to focus on providing consistent feedback for tonal languages, which is crucial to improving the accuracy of learners’ speaking. Current applications and ML models lack inputted feedback, ignoring linguistic properties of a character-by-character approach. This project focuses on the development of a novel Automatic Speech Recognition (ASR) model that takes in microphone input, transcribes audio into Chinese characters, separates input into individual phonemes, searches for desired tone pronunciation, and produces pitch frequency graphs to compare desired pronunciation with inputted audio. Speech-to-text transcription aimed to optimize results found using NeMo, a toolkit for ASR models that has various features of speech-to-text synthesis, speech classification, and voice activity detection. In addition to utilizing SPICE to create an unsupervised model, visual tonal feedback was implemented to better optimize consistent feedback. Using datasets containing over 10,000 audio files, along with audio files from our Chinese teacher, we achieved 84.5% accuracy in transcription, 87% accuracy in identifying the desired tone, and 89% accuracy in determining desired tone accuracy, creating an end-to-end solution effective in improving tone. Future applications of this project are to implement live pronunciation feedback and further expand our data for improving pronunciation through ASR for other tonal languages. More information can be found in our engineering notebook: https://tinyurl.com/CSEF-S-08-52
Source coverage
This record comes from a published award list, not a complete project archive. Its abstract comes from CSEF's public project showcase as archived by the Internet Archive before judging (https://web.archive.org/web/20230401224130/https://ca-csef.zfairs.com/showcase/ShowcaseInfo?f=838e60b7-ea75-46e8-865c-fde4864244b3); the version presented may differ.
Awards (1)
Competition history
- CSEF 2023
Resources
Related projects
ISEF · 2021
Partially Speaker-Dependent Automatic Speech Recognition Using Deep Neural Networks
ISEF · 2017
Language Identification Based on the Variations in Intonation Using Multi-Classifier Systems
ISEF · 2025
Evaluating Accuracy of Open-Source Automatic Speech Recognition (ASR) Models With Acoustic Analysis
ISEF · 2025
Coding AI To Enhance Speech Therapy
CSEF · 2023
Development of a Novel Automatic Speech Recognition to Reduce Racial and Gender Bias
ISEF · 2026
ATC-Copilot: Automatic Speech Recognition and Natural Language Processing for Air Traffic Control Communications
ISEF · 2026
Accent or Impediment? Using Machine Learning to Prevent Speech Misdiagnosis in Children
ISEF · 2026
ArticuRace: Closing the Global Speech Therapy Gap Through a Low-Resource and Interpretable AI Framework
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: California Science & Engineering Fair public projects