VoiceFormer: A Deep Learning Approach to Non-Invasive Screening of Speech Related Disorders
CWSF · 2026 Disease & Illness Bronze Medal
Overview
Speech and voice changes can serve as noninvasive indicators of disease, particularly in respiratory, neurological, and psychiatric conditions, where subtle changes in phonation, tone, or articulation can reflect underlying pathology. However, voice-based clinical assessment is not typically integrated in diagnostic pathways, and is difficult to scale. This project investigates the utility of voice-derived features for diagnosis of diverse pathologies with known vocal changes. Using the Bridge2AI Voice Dataset, mel spectrogram representations were paired with MFCC and static acoustic features to train a dual-encoder multimodal deep learning model (VoiceFormer) for the diagnosis of 9 unique conditions. Model performance ranged between 80-98% AUROC for diseases, matching or outperforming previous state-of-the-art models. An interpretability analysis outlined the most important features for classification. This project shows the potential of using multimodal speech features and deep learning to enable real-time, accessible screening for diverse medical conditions.
Awards (2)
- Bronze Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
ISEF · 2018
VoiceEDx: A Novel, Voice-Based End-to-End Multi-Disease Diagnostic Platform Using a Highly Accurate and Expandable Artificial Intelligence Engine for an Early, Secure and Reliable Diagnosis of Disease
CWSF · 2026
NeuroSight: Bridging Oculomic and Acoustic Features for Multi-Disease Prediction via Deep Learning
ISEF · 2026
Machine Learning Model for Early Identification of Vocal Patterns Associated With Glottic Carcinoma Risk by Analysis of Voice Biomarkers and Life Habits
ISEF · 2024
VoiceAD: A Crosslingual Speech-Based Classifier for Early Detection of Alzheimer's Disease
Closest projects by meaning, across every fair and year in the corpus.