A Novel End-to-End Automatic Speech Recognition (ASR) Model to Improve Tonal Pronunciation in Mandarin for Non-Native Speakers

CSEF · 2023 Computational Systems & Analysis Second Award

Overview

Tonal languages are estimated to comprise 60–70% of the world’s languages. However, tonal language pronunciation is difficult for non-native speakers, as Mandarin has 5 tones, resulting in each tone having a certain pitch. This differentiates thousands of Chinese characters, which are characterized by their phonemes, resulting in 413 distinct tonal combinations compared to 44 tonal sounds in English. As non-native speakers ourselves, we aimed to focus on providing consistent feedback for tonal languages, which is crucial to improving the accuracy of learners’ speaking. Current applications and ML models lack inputted feedback, ignoring linguistic properties of a character-by-character approach. This project focuses on the development of a novel Automatic Speech Recognition (ASR) model that takes in microphone input, transcribes audio into Chinese characters, separates input into individual phonemes, searches for desired tone pronunciation, and produces pitch frequency graphs to compare desired pronunciation with inputted audio. Speech-to-text transcription aimed to optimize results found using NeMo, a toolkit for ASR models that has various features of speech-to-text synthesis, speech classification, and voice activity detection. In addition to utilizing SPICE to create an unsupervised model, visual tonal feedback was implemented to better optimize consistent feedback. Using datasets containing over 10,000 audio files, along with audio files from our Chinese teacher, we achieved 84.5% accuracy in transcription, 87% accuracy in identifying the desired tone, and 89% accuracy in determining desired tone accuracy, creating an end-to-end solution effective in improving tone. Future applications of this project are to implement live pronunciation feedback and further expand our data for improving pronunciation through ASR for other tonal languages. More information can be found in our engineering notebook: https://tinyurl.com/CSEF-S-08-52

Source coverage

This record comes from a published award list, not a complete project archive. Its abstract comes from CSEF's public project showcase as archived by the Internet Archive before judging (https://web.archive.org/web/20230401224130/https://ca-csef.zfairs.com/showcase/ShowcaseInfo?f=838e60b7-ea75-46e8-865c-fde4864244b3); the version presented may differ.

Awards (1)

Competition history

  • CSEF 2023 Computational Systems & Analysis · Entry S0852

Resources

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google