VoiceFormer: A Deep Learning Approach to Non-Invasive Screening of Speech Related Disorders

CWSF · 2026 Disease & Illness Bronze Medal

Thumbnail supplied by the source for VoiceFormer: A Deep Learning Approach to Non-Invasive Screening of Speech Related Disorders

Overview

Speech and voice changes can serve as noninvasive indicators of disease, particularly in respiratory, neurological, and psychiatric conditions, where subtle changes in phonation, tone, or articulation can reflect underlying pathology. However, voice-based clinical assessment is not typically integrated in diagnostic pathways, and is difficult to scale. This project investigates the utility of voice-derived features for diagnosis of diverse pathologies with known vocal changes. Using the Bridge2AI Voice Dataset, mel spectrogram representations were paired with MFCC and static acoustic features to train a dual-encoder multimodal deep learning model (VoiceFormer) for the diagnosis of 9 unique conditions. Model performance ranged between 80-98% AUROC for diseases, matching or outperforming previous state-of-the-art models. An interpretability analysis outlined the most important features for classification. This project shows the potential of using multimodal speech features and deep learning to enable real-time, accessible screening for diverse medical conditions.

Video

VoiceFormer achieved an overall AUROC of 0.888 across 9 disease-classification tasks on the Bridge2AI-Voice v3.0 dataset. Performance was strong across all 3 disorder categories: respiratory disorders reached 0.882, voice disorders 0.876, and neurological/mood disorders 0.906. At the individual-task level, the model produced a clear performance hierarchy reflecting the underlying strength of each condition's acoustic signature. Cognitive impairment achieved the highest AUROC at 0.978, followed by airway stenosis (0.947), vocal fold paralysis (0.934), Parkinson's disease (0.898), laryngeal dystonia (0.867), COPD/asthma (0.861), depression (0.843), chronic cough (0.839), and benign lesions (0.827). Notably, all 9 tasks exceeded 0.82 AUROC, indicating that the model achieved clinically meaningful discrimination on every condition tested.

Secondary metrics confirmed consistent performance. F1-scores ranged from 0.718 for benign lesions to 0.943 for cognitive impairment. Precision exceeded recall on every task, indicating that the model operates conservatively, producing fewer false positives at the cost of missing some true cases. For example, depression achieved precision of 0.787 versus recall of 0.699, and Parkinson's achieved 0.841 versus 0.769.

Why?

Despite growing interest in voice-based artificial intelligence, the field still lacks large, clinically linked datasets collected in a standardized, privacy-conscious way. The human voice is produced through the coordination of respiratory, laryngeal, articulatory, and neurological systems, so disease-related changes in any of these systems leave measurable acoustic traces. For example, Parkinson's disease reduces pitch variability and articulatory precision, airway stenosis introduces high-frequency turbulence from narrowed airways, vocal fold paralysis causes breathiness from incomplete glottal closure, and depression flattens prosody and reduces vocal effort.

Because voice can be collected quickly, repeatedly, and at low cost, it has emerged as a promising digital biomarker. However, earlier studies typically relied on small single-site cohorts, inconsistent recording protocols, or models targeting only one disease. Two recent studies on Bridge2AI-Voice v2.0 (442 participants) established initial benchmarks — Liu et al. (2026) achieved AUROC 0.81 using spectrograms alone across 4 broad categories, and Piao et al. (2025) achieved AUROC 0.78 using dual-modal multi-task learning across 9 diseases — but both used limited modalities, simple concatenation fusion, and left conditions like depression unaddressed.

This project tests whether VoiceFormer, a 4-branch cross-attention fusion model, can improve disease detection on Bridge2AI-Voice v3.0, which contains approximately 29,020 recordings from 833 participants including 142 healthy controls. Nine binary tasks are investigated across respiratory (airway stenosis, COPD/asthma, chronic cough), voice (vocal fold paralysis, laryngeal dystonia, benign lesions), and neurological/mood disorders (Parkinson's, cognitive impairment, depression).

How?

Each recording was represented using 4 complementary feature types derived during dataset preparation, since raw audio is not publicly released. Mel spectrograms (60 Mel bins × T frames) captured time-frequency energy, MFCCs (60 × T) summarized the spectral envelope via discrete cosine transform, articulatory features (15 × T) from the SPARC package provided estimated electromagnetic articulography traces for 6 vocal tract articulators plus loudness, periodicity, and pitch, and 132 handcrafted static features from OpenSMILE and Praat captured clinically relevant biomarkers including jitter, shimmer, HNR, formants, and pitch statistics.

VoiceFormer processed these through 4 specialized encoders: a Swin-Tiny Transformer for mel spectrograms, a 1D ResNet-18 for MFCCs, a 1D Temporal Convolutional Network for articulatory features, and a 3-layer MLP for static features. All embeddings were projected to 256 dimensions and fused using a 4-head cross-attention module, which learned task-adaptive modality weighting before passing the fused representation to 9 binary classification heads.

Training used a patient-level 80/20 split with no participant overlap, following a multi-task classification protocol similar to the previous studies in this field. Results were reported over 5 independent runs using AUROC as the primary metric, with secondary outcomes including F1-score, precision, recall, and specificity. Comparators included reproductions of MARVEL and SSL-AST retrained on v3.0, fine-tuned Whisper and Wav2Vec2-BERT, 3 single-modality baselines, and ablated VoiceFormer variants. Interpretability analysis involved cross-task analysis and GradCAM.

What?

VoiceFormer outperformed all comparator models retrained on v3.0 data. MARVEL* achieved 0.840 overall, SSL-AST* 0.824, and single-modality baselines ranged from 0.748 (ResNet18 on MFCCs) to 0.771 (EffNet-B0 on mel spectrograms). DeLong tests confirmed significant improvements on 8 of 9 tasks (all p < 0.01 except laryngeal dystonia at p = 0.006); cognitive impairment was the only nonsignificant comparison (p = 0.137) due to near-ceiling baseline performance at 0.968. The largest gains appeared on benign lesions (+0.096 over SSL-AST*), Parkinson's (+0.056), and chronic cough (+0.055), demonstrating that fusion matters most when individual modality signals are weakest. Training dynamics showed stable convergence with cognitive impairment plateauing by epoch 14 and harder tasks like depression converging around epoch 30, while a mild validation loss uptick after epoch 46 triggered early stopping at epoch 47 with a train-validation gap of 0.094 indicating controlled generalization.

Ablation analysis quantified each component's contribution. Removing the mel-spectrogram branch caused the largest drop (−0.047 overall), followed by MFCC (−0.039), articulatory (−0.014), and static features (−0.007). Replacing cross-attention with concatenation lowered AUROC from 0.888 to 0.867, and the 3-branch ablation without articulatory features scored 0.864, confirming that SPARC features provide a modest but meaningful signal, especially for Parkinson's where motor coordination of the tongue and lips is directly impaired. Single-modality baselines ranged from 0.724 (articulatory only) to 0.788 (mel-spectrogram only), all substantially below the full model.

Interpretability analysis revealed clinically coherent patterns across multiple methods. CKA embedding similarity showed that within-category representations clustered together — respiratory at 0.59–0.66, voice at 0.56–0.73, and neurological/mood at 0.57–0.59 — while cross-category similarity remained low (0.27–0.44), confirming that the model maintains distinct representational spaces for different physiological systems. Laryngeal dystonia and vocal fold paralysis showed the highest within-category pair (0.73), consistent with both being motor-laryngeal conditions. Grad-CAM visualizations showed that airway stenosis activated a narrow high-frequency band corresponding to turbulent airflow, Parkinson's activated low-frequency regions tied to fundamental frequency consistent with hypophonic speech, and depression showed diffuse low-intensity activation spread across mid-frequency bands and temporal transitions, reinforcing that depression's acoustic signature is subtle and distributed. Cross-attention weights further confirmed task-specific modality specialization: airway stenosis relied most on mel-spectrogram information (0.418), Parkinson's on MFCC-based prosodic features (0.474), and depression showed the most balanced distribution across all 4 modalities (0.278–0.308). Cross-site AUROC variation remained below 0.06 for most tasks, supporting generalization, although depression showed the highest variability (0.824–0.861). Together, these findings confirm that VoiceFormer achieves its improvements through physiologically coherent attention to disease-relevant acoustic patterns rather than spurious correlations.

So What?

VoiceFormer demonstrated that a single multi-task model can classify respiratory, voice, neurological, and mood disorders using only derived acoustic features, achieving AUROC above 0.82 on all 9 tasks and 0.888 overall. Most importantly, it provided the first voice-based depression and chronic cough detection on this dataset (0.843 and 0.839 respectively) by fusing individually weak signals across 4 complementary modalities. Cross-attention fusion consistently outperformed concatenation (+0.021 overall), and the 4-modality architecture outperformed all single-modality and 3-modality variants, confirming that each feature type contributes unique diagnostic information. These results support the potential for short voice recordings to serve as a scalable, low-cost, noninvasive screening tool for multiple conditions simultaneously, with privacy-preserving on-device feature extraction eliminating the need to transmit raw audio. However, several limitations must be acknowledged. The dataset contains 833 participants with all tasks classified as borderline-ready, and the binary-versus-control formulation using 142 healthy controls is a different classification target than the disease-versus-other-disorder approach in prior v2.0 studies. Depression labels are self-reported rather than clinically adjudicated, SPARC articulatory features are model-estimated rather than sensor-measured, and all results come from a single dataset without external validation. While the findings strongly support the feasibility of multi-modal voice-based screening, independent replication is essential before clinical deployment.

What's Next?

Future priorities include external validation on independent datasets, extension to additional borderline-ready conditions such as anxiety (107 participants) and obstructive sleep apnea (89 participants), and incorporation of the currently unused phonetic posteriorgrams and Whisper transcriptions as additional model branches. Longer-term directions include combining VoiceFormer with self-supervised pre-training on all 29,020 recordings, adding uncertainty estimation for calibrated confidence scores, developing longitudinal monitoring for progressive conditions like Parkinson's and cognitive decline, and extending the framework to the Bridge2AI-Voice Pediatric cohort.

Thanks

I would like to thank PhysioNet and the Bridge2AI Voice consortium (NIH project 3OT2OD032720-01S1) for providing the Bridge2AI-Voice v3.0 dataset, as well as the developers of OpenSMILE, Praat, Parselmouth, the SPARC articulatory coding package, PyTorch, and torchaudio.

References

Arun, S., Grosheva, M., Kosenko, M., Robertus, J. L., Blyuss, O., Gabe, R., Munblit, D., & Offman, J. (2025). Systematic scoping review of external validation studies of AI pathology models for lung cancer diagnosis. *npj Precision Oncology, 9*(1), 166. https://doi.org/10.1038/s41698-025-00940-7

Baevski, A., Zhou, H., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. *Advances in Neural Information Processing Systems, 33*, 12449–12460.

Bateman, E. D., Hurd, S. S., Barnes, P. J., Bousquet, J., Drazen, J. M., FitzGerald, J. M., Gibson, P., Ohta, K., O’Byrne, P., Pedersen, S. E., Pizzichini, E., Sullivan, S. D., Wenzel, S. E., & Zar, H. J. (2008). Global strategy for asthma management and prevention: GINA executive summary. *European Respiratory Journal, 31*(1), 143–178. https://doi.org/10.1183/09031936.00138707

Bensoussan, Y., Sigaras, A., Rameau, A., Elemento, O., Powell, M., Dorr, D., Payne, P., Ravitsky, V., Bélisle-Pipon, J.-C., Johnson, A., Bahr, R., Watts, S., Bolser, D., Siu, J., Lerner-Ellis, J., Rudzicz, F., Boyer, M., Salvi Cruz, S., Abdel-Aty, Y., ... Ghosh, S. (2025). *Bridge2AI-Voice: An ethically-sourced, diverse voice dataset linked to health information* (Version 2.0.0). PhysioNet. https://doi.org/10.13026/3xt6-rf05

BRIDGE2AI. (2025). *Data generation project—Bridge2AI Voice*. https://bridge2ai.org/data-voice/

Cao, F., Vogel, A. P., Gharahkhani, P., & Renteria, M. E. (2025). Speech and language biomarkers for Parkinson’s disease prediction, early diagnosis and progression. *npj Parkinson’s Disease, 11*(1), 57. https://doi.org/10.1038/s41531-025-00913-4

Chung, Y.-A., Zhang, Y., Han, W., Chiu, C.-C., Qin, J., Pang, R., & Wu, Y. (2021). W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. In *2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)* (pp. 244–250).

Cohen, A. S., Divers, R., & Calamia, M. (2026). Speech pause and speech rate for evaluating Alzheimer’s and mild cognitive impairment: A meta-analysis. *Journal of the International Neuropsychological Society, 32*(1), 24–31. https://doi.org/10.1017/S1355617725101677

Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., et al. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. *BMJ, 385*, e078378.

Davis, S. B., & Mermelstein, P. (1980). Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. *IEEE Transactions on Acoustics, Speech, and Signal Processing, 28*(4), 357–366. https://doi.org/10.1109/TASSP.1980.1163420

Döllinger, M., Kunduk, M., Kaltenbacher, M., Vondenhoff, S., Ziethe, A., Eysholdt, U., & Bohr, C. (2012). Analysis of vocal fold function from acoustic data simultaneously recorded with high-speed endoscopy. *Journal of Voice, 26*(6), 726–733. https://doi.org/10.1016/j.jvoice.2012.02.001

Dubey, P., Fernandes, J. B., & Bhat, M. (2022). Acoustic analysis of voice in laryngopharyngeal cancers pre and post radiotherapy. *Indian Journal of Otolaryngology and Head & Neck Surgery, 74*(Suppl 2), 1973–1978. https://doi.org/10.1007/s12070-020-01934-6

Elendu, C., Amaechi, D. C., Elendu, T. C., Amaechi, E. C., Elendu, I. D., Joseph, M. C., Ajakaye, A. A., Ansong, S. O., Tyagi, V., Anukam, L. I., & Oguoma, C. O. (2025). Guideline-driven and dependable management of asthma: An evidence-based systematic review. *Annals of Medicine and Surgery, 87*(8), 5153–5164. https://doi.org/10.1097/MS9.0000000000003491

Fagherazzi, G., Fischer, A., Ismael, M., & Despotovic, V. (2021). Voice for health: The use of vocal biomarkers from research to clinical practice. *Digital Biomarkers, 5*(1), 78–88. https://doi.org/10.1159/000515346

Fumel, J., Bahuaud, D., Weed, E., Fusaroli, R., & Basirat, A. (2024). A systematic review and Bayesian meta-analysis of prosodic impairment in Parkinson’s disease. *Journal of Speech, Language, and Hearing Research, 67*, 2548–2564. https://doi.org/10.1044/2024_JSLHR-23-00588

Gong, Y., Lai, C.-I. J., Chung, Y.-A., & Glass, J. (2021). SSAST: Self-supervised audio spectrogram transformer [Preprint]. *arXiv*. https://doi.org/10.48550/arXiv.2110.09784

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In *Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition* (pp. 770–778). https://doi.org/10.1109/CVPR.2016.90

Jafari, Z., Andrew, M. K., & Rockwood, K. (2025). Diagnostic utility of speech-based biomarkers in mild cognitive impairment: A systematic review and meta-analysis. *Age and Ageing, 54*(10), afaf316. https://doi.org/10.1093/ageing/afaf316

Jiang, J.-Y., Hsu, P.-M., Pan, Y.-A., Yu, Y.-H., Chen, C.-K., & Hsieh, L.-C. (2025). Cepstral peak prominence: A valuable measure of voice outcome severity in patients with unilateral vocal fold paralysis. *Journal of Voice*. Advance online publication. https://doi.org/10.1016/j.jvoice.2024.11.031

Kalia, A., Boyer, M., Fagherazzi, G., Bélisle-Pipon, J.-C., Bensoussan, Y., & Bridge2AI-Voice Consortium. (2025). Master protocols in vocal biomarker development to reduce variability and advance clinical precision: A narrative review. *Frontiers in Digital Health, 7*, Article 1619183. https://doi.org/10.3389/fdgth.2025.1619183

Li, S., & Tang, H. (2026). Multimodal alignment and fusion: A survey. *International Journal of Computer Vision, 134*, 103. https://doi.org/10.1007/s11263-025-02667-1

Lin, T.-Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017). Focal loss for dense object detection. In *Proceedings of the IEEE International Conference on Computer Vision* (pp. 2980–2988). https://doi.org/10.1109/ICCV.2017.324

Liu, W., Qu, B., Pontell, M., Powell, M., Malin, B., & Yin, Z. (2026). Optimizing domain-adaptive self-supervised learning for clinical voice-based disease classification [Preprint]. *arXiv*. https://doi.org/10.48550/arXiv.2601.22319

Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. In *Proceedings of the IEEE/CVF International Conference on Computer Vision* (pp. 10012–10022). https://doi.org/10.1109/ICCV48922.2021.00986

Loshchilov, I., & Hutter, F. (2017). SGDR: Stochastic gradient descent with warm restarts. In *International Conference on Learning Representations*.

Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. In *International Conference on Learning Representations*.

McLoughlin, I., Pham, L., Song, Y., Miao, X., Phan, H., Cai, P., Gu, Q., Nan, J., Song, H., & Soh, D. (2026). Spectrogram features for audio and speech analysis. *Applied Sciences, 16*(2), 572. https://doi.org/10.3390/app16020572

Moothedan, E., Boyer, M., Watts, S., Abdel-Aty, Y., Ghosh, S., Rameau, A., Sigaras, A., Elemento, O., Bridge2AI-Voice Consortium, & Bensoussan, Y. (2025). The Bridge2AI-voice application: Initial feasibility study of voice data acquisition through mobile health. *Frontiers in Digital Health, 7*, 1514971. https://doi.org/10.3389/fdgth.2025.1514971

Park, D. S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E. D., & Le, Q. V. (2019). SpecAugment: A simple data augmentation method for automatic speech recognition. In *Proceedings of Interspeech 2019* (pp. 2613–2617). https://doi.org/10.21437/Interspeech.2019-2680

Piao, R., Lu, Y., Kemps, H., Xia, T., & Saeed, A. (2025). Unified acoustic representations for screening neurological and respiratory pathologies from voice [Preprint]. *arXiv*. https://doi.org/10.48550/arXiv.2508.20717

Praveen, R. G., & Alam, J. (2024). Recursive joint cross-modal attention for multimodal fusion in dimensional emotion recognition. In *2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)* (pp. 4803–4813). https://doi.org/10.1109/CVPRW63382.2024.00483

Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. In *Proceedings of the 40th International Conference on Machine Learning* (Vol. 202, pp. 28492–28518). PMLR.

Sara, J. D. S., Orbelo, D., Maor, E., Lerman, L. O., & Lerman, A. (2023). Guess what we can hear—Novel voice biomarkers for the remote detection of disease. *Mayo Clinic Proceedings*. Advance online publication. https://doi.org/10.1016/j.mayocp.2023.03.007

Wieczorek, K., Ananth, S., & Valazquez-Pimentel, D. (2024). Acoustic biomarkers in asthma: A systematic review. *Journal of Asthma, 61*(10), 1165–1180. https://doi.org/10.1080/02770903.2024.2344156

Zhang, H., Cisse, M., Dauphin, Y. N., & Lopez-Paz, D. (2018). mixup: Beyond empirical risk minimization. In *International Conference on Learning Representations*.

Images (30)

Awards (2)

  • Bronze Medal
  • Selected for CWSF 2026

Competition history

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: ProjectBoard / Youth Science Canada

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google