Real-Time Mobile Speech Emotion Recognition: Optimized Transformer VAD Mapping for Low-Latency Offline Performance
CSEF · 2026 Behavioral & Social Sciences (Senior Division)
Overview
This project developed an on-device application for Speech Emotion Recognition (SER) using self-supervised transformer models. The hypothesis stated that audEERING’s wav2vec2-based, dimensional Valence-Arousal-Dominance (VAD) model would be the most suitable of the three evaluated architectures for on-device SER, due to its low-latency inference and externally controllable emotion-mapping process that allows iterative post-deployment accuracy refinement. The procedure consisted of evaluating three fine-tuned architectures – Whisper, SpeechBrain, and audEERING – on the MELD dataset to determine the most suitable model. Unlike Whisper and SpeechBrain, which are classification models, the evaluation of audEERING was done through iterative refinement. The custom mapping system created achieved 63.90% accuracy on MELD, but it achieved the best trade-offs between accuracy and latency. This program improves upon existing solutions by offering offline processing to ensure privacy and eliminate cloud latency. The application features a real-time mode and a results history log. In practice, short audio clips undergo offline inference to display predicted emotions and VAD values, with the option for user confirmation. Program success was evaluated through independent testing with 210 clips, yielding 60.48% accuracy, indicating that while categorical emotion prediction remains challenging, the dimensional VAD-based pipeline functions reliably in a mobile setting. Final performance reached ~420ms latency and ~185MB memory usage, outperforming the 1-3 second delay of cloud-based APIs. Minor emotion classification inaccuracies likely originated from background noise or natural user speech variations. Thus, these results support the hypothesis that a custom-mapped VAD model is suitable for mobile SER. Future applications include emotional assistance for neurodivergent individuals and private mood tracking.
Competition history
- CSEF 2026
Related projects
ISEF · 2021
An Audiovisual Emotion Classification Platform Powered By Deep Neural Networks and Optimized for Individuals with Autism Spectrum Disorder
ISEF · 2020
A Machine Learning Approach to Help Autistic Individuals Recognize Emotions in Vocal Conversation
ISEF · 2021
Voice Emotion Recognition with Audio Data Analysis and Machine Learning Algorithms
ISEF · 2023
Sign2Speak! Synthesizing Emotional Speech From Sign Language Through Deep Learning
ISEF · 2024
Development of Emotion Recognition Device for Visually Impaired
ISEF · 2025
emoTune: New AI-Based Approaches to Assistive Technology for Individuals With Autism Spectrum Disorder
CSEF · 2017
EmNet: Emotion Recognition from Human Voice Using Machine Learning for Affective Computing
CSEF · 2018
Emotion Recognition from Human Speech Using Temporal Information and Deep Learning
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: California Science & Engineering Fair public projects