Viseme Analysis: Implementing Contour Sequence Classification to Augment Speech Recognition
CSEF · 2016 Mathematics & Computer Science
Overview
Objectives/Goals To test the efficiency of various lip contour detection methods as well as machine learning algorithms in using a sequence of images of a speaker to determine what the speaker was saying. Methods/Materials Laptop computer with Python development kit, open source computer vision and scientific computing libraries, and dataset of videos obtained from public LiLiR dataset. Implemented feature extraction algorithms and sequence classification algorithms by merging existing models and tested accuracy by running on a separate dataset. Results The final design, which linked adaptive thresholding, a support vector machine, and a hidden Markov model, predicted a spoken letter based on solely video data. It attained a 43% average recognition rate on the test data. Conclusions/Discussion I created a segmented model that took raw image data from a video and predicted the letter that was spoken. This method can be extended to cover dictionaries larger than the English alphabet, and due to its segmented nature, the extracted features can also be added to an audio-based speech recognition system. Though viseme analysis using lip contour detection and hidden Markov models cannot function professionally as a standalone program, it can be effective in filtering out noise in existing speech recognition systems.
Summary statement
I devised an algorithm to implement speech recognition through lip reading.
Help received
All code was written by me using open source libraries. Publicly available datasets were downloaded from the University of Surrey website.
Competition history
- CSEF 2016
Resources
Related projects
ISEF · 2020
A Novel Lip-Reading Method Based on Transfer Learning from MobileNet and LSTM Architecture
ISEF · 2015
Automated Lip-Reading Technique for Speech Disabilities: A Novel Machine Learning Algorithm Using Microsoft Kinect Sensor and Hybrid Approach for Feature Extraction
CSEF · 2019
A Novel Program for the Detection and Translation of the ASL Alphabet through the Use of Deep Learning
ISEF · 2016
iSight: A New Method of Gesture Recognition and Analysis
ISEF · 2020
American Sign Language Classifier
ISEF · 2026
Beyond Audio: A Multimodal EMG-Visual Speech System for Reconstructing Voice From Silence
ISEF · 2021
Visual Sign Language Translator
ISEF · 2025
Coding AI To Enhance Speech Therapy
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: California Science & Engineering Fair public projects