NEXUS: A Multimodal Device for Deafblind Communication using Hybrid Dynamic Speaker Diarization
CWSF · 2026 Digital Technology Silver Medal
Overview
Deafblindness is a combined disability of vision and hearing impairments that create significant challenges to communication and information access. Many deafblind individuals rely on intervenors to access spoken information, but access to these services is highly limited due to shortages, uneven availability, and funding restrictions. To address this gap, NEXUS was developed as a wearable assistive device that enables real-time, independent communication. NEXUS captures surrounding speech, identifies who is speaking, and converts conversations into accessible formats like braille and enlarged text, allowing users to follow group conversations without relying on human intermediaries. Notably, the device uses speaker diarization that combines spatial awareness with voice pattern analysis, making speaker identification robust even in complex, multi-speaker environments. This approach outperforms current state-of-the-art models for real-time, far-field diarization. Overall, NEXUS provides an accessible and scalable solution that enhances communication and supports greater autonomy among deafblind communities.
Video
This video could not be played here. Watch it on the original project page.
Video
Video Transcript
Hi! My name is Tyler and my name is Leo, and we developed NEXUS, a wearable device that enables real-time, independent communication for individuals with deafblindness.
NEXUS works by capturing surrounding speech, identifying speakers, and transcribing the spoken content into a labelled transcript, clearly showing who is speaking and what they are saying. This information is delivered through a form of accessible output, such as braille or enlarged text, which can be worn on the user’s arm, allowing them to follow both one-on-one and group conversations in real time.
Today, over 160 million people worldwide live with deafblindness, and communication often depends on intervenors, services that face training shortages and uneven availability, preventing many individuals from receiving sufficient aid. Our device removes that barrier by providing deafblind communities with direct, 24 hour access to conversations.
NEXUS has the potential to transform how deafblind communities engage with the world, bringing independent, real-time communication to everyday life, and restoring access to the conversations that connect us all.
Thank you, and see you at the fair!
Why?
Background Information
Deafblindness affects roughly 160 million individuals worldwide, with over 600,000 people living with deafblindness in Canada alone, creating significant barriers to communication (World Federation of the Deafblind [WFDB], 2018). As a condition that impedes both vision and hearing, many everyday interactions, such as conversations, public services, and social gatherings, become difficult without assistance. Consequently, deafblind individuals rely on intervenors, trained professionals who act as a communication bridge between the person and their environment by conveying spoken information through tactile cues (Deafblind Network of Ontario, 2025).
Problem
However, access to intervenor services is often limited due to shortages of trained professionals, uneven service availability, and funding or policy restrictions on support hours (Figure 4). As a result, deafblind services are frequently unavailable for routine interactions, making education, employment, and social interaction challenging. Nevertheless, deafblind communities continue to depend on interpreters for communication, which can limit autonomy and independence for individuals who wish to communicate and navigate everyday environments on their own (UK Government, 2025).
Research Question
How can an accessible and low-cost solution be designed to enable real-time, independent communication for deafblind individuals in everyday settings?
Research Objectives
Analyze the communication needs and challenges of deafblind individuals in real-world environments to support the design of an accessible, portable assistive system.
Design a system that replicates key functions of intervenors by enabling both one-on-one and group conversations accessibly.
Evaluate the system’s effectiveness in conversational settings through the analysis of speaker diarization accuracy across benchmark datasets and real-world testing.
How?
System Overview
NEXUS works by capturing speech, identifying speakers, transcribing speech content, and converting the information into an accessible format for deafblind users in real time. Prototype I focuses on validating the system’s overall flow using a microphone array to capture spatial audio, connected to STM32 and Raspberry Pi 5 microcontrollers for speaker diarization, which means to separate and label different speakers in a conversation (Figures 7 and 8). Transcribed text is paired with its corresponding speakers to clearly convey who is speaking and what they are saying, which is then delivered to the deafblind user through a refreshable braille display using micro stepper motors and octagonal braille drums (Figure 10).
Diarization Pipeline
NEXUS integrates a dynamic hybrid diarization pipeline designed for two key requirements: real-time processing and far-field operation (Figure 19). The pipeline fuses spatial information via GCC-PHAT localization with vocal characteristics from embedding extraction (Pyannote), weighted through a recency-based exponential decay function (Figures 15 and 18). This reduces the influence of inactive speakers over time, introducing a continuity bias that improves assignment accuracy. The pipeline also adapts to varying conversation scenarios, such as short utterances, by dynamically adjusting speaker weighting. Performance was evaluated against existing methods using both a standardized dataset and in-person testing. All results were benchmarked against the current state-of-the-art for real-time, far-field speaker diarization, Pyannote, which served as the baseline.
Wearable Design
Prototype II refines the physical design into a more compact, wearable system optimized for real-world use (Figures 12 and 13). A hot-swappable interface was proposed to support multiple accessible outputs, accommodating the diverse communication preferences of deafblind individuals, largely attributed to variation in disease severity. The following output modules were designed: a 3D printed electromechanical braille display, enlarged text screen, and haptic feedback module (Figure 14).
What?
Testing Types
Validation was conducted using in-person trials and the AMI Corpus Multiple Distant Microphone (MDM) dataset (University of Edinburgh, n.d.). The in-person trials assessed the end-to-end system performance, including hardware integration, while the AMI evaluation focuses purely on the diarization algorithm outside of hardware limitations (Figures 20 and 21). Thus, real world testing allows for greater evaluation of practical effectiveness, while the standardized evaluation allows for fair benchmarking against existing approaches.
In-Person Testing
Two types of trials were conducted through in-person testing: realistic conversations with some speaker movement and overlapping speech and more difficult conversation scenarios with rapid turn-taking (Table 1).
Under the realistic trials, the baseline achieved a diarization error rate (DER) of 20.5%, whereas the proposed pipeline achieved a DER of 9.9%, representing a substantial reduction in error (Figure 23a). The improvement suggests that the integration of spatial information enables more accurate speaker tracking. Under difficult trials, the baseline achieved an average DER of 81.4%, while the proposed pipeline achieved 11.5%, once again demonstrating a significant improvement in accuracy (Figure 23b). Moreover, the proposed pipeline highlights a significant limitation of traditional embedding-only diarization and static fusion pipelines in that they are unable to adapt to different conversation conditions, making them less robust to real-world environments. In particular, these methods over-rely on speaker embeddings, which become unreliable for diarizing short utterances in real-time environments.
Standardized Dataset
An ablation study was performed across 20 conversation scenarios to validate the pipeline’s improvements over the baseline and optimize the weighting mechanisms (Figure 24). Specifically, the mechanism was applied separately to embeddings and spatial information, allowing for a direct comparison of their impact on performance and determining which configuration yields the greatest accuracy. Results demonstrated that the recency-embedding mechanism achieved the highest reduction in DER from a baseline of 39.4% to 30.6%, while spatial recency weighting achieved a DER of 38.6%, reflecting recency-embedding weighting as the most effective approach (Table 1).
Comparing to Existing Approaches
Existing approaches for real-time diarization face several limitations or lack testing against standardized far-field datasets (Table 3). LS-EEND is a live streaming approach to speaker diarization that relies solely on speaker features (Liang & Li, 2025). However, it was primarily trained on the AMI Individual Headset Microphone (IHM) dataset, which consists of near-field audio recordings captured close to each speaker’s mouth, making it unsuitable for far-field environments. SDSS is another approach that utilizes both speaker embeddings and location data (Zheng et al., 2021). However, its evaluation is limited to a custom dataset and integrates hardware setups that are unfeasible for devices like NEXUS, so it is difficult to directly compare its performance against the proposed diarization system. SpeechCompass is another far-field approach, but like SDSS, it was only evaluated on a custom dataset (Dementyev et al., 2025).
The proposed pipeline is therefore the first study of a real-time, far-field diarization system to be validated under a standardized dataset like AMI Corpus, establishing a new benchmark for future research.
So What?
Impact on Deafblind Communication
Testing reveals NEXUS addresses a major accessibility gap in deafblind communities by enabling real-time communication effectively without intervenors, making everyday communication much more accessible. As a fully wearable and portable system, it allows users to communicate anywhere, at any time, without relying on external support. Designed to be low-cost, NEXUS is significantly more accessible than many existing assistive technologies, reducing financial barriers to adoption (Table 4). At the same time, it offers a broader and more advanced feature set, most notably its support for multi-speaker environments. Being low-cost, NEXUS can be easily scaled and distributed, helping alleviate intervenor shortages and the uneven availability of these services. NEXUS’ advancements will contribute to substantial improvements to the quality of life of deafblind communities by enabling greater independence, making education, work, and daily interactions all the more accessible.
Research Significance of Diarization
Beyond the impact on the deafblind community, NEXUS represents a significant advancement in the research field of speaker diarization by outperforming state-of-the-art baselines. Reliable real-time, far-field diarization systems can also be applied outside of assistive technologies, such as voice assistants and meeting transcription (Figure 26).
Limitations
NEXUS is currently unable to convey nuanced conversational cues such as facial expressions or body language, elements that human intervenors can interpret. Additionally, the system has only been validated on indoor conversations and has not yet been tested in diverse environments, such as outdoors, which is an essential next step for a portable device intended to function reliably in varied settings.
What's Next?
Future work will focus on the construction of prototype II and algorithmic refinement, which aims to be complete for the CWSF in-person (Figure 27). Pilot testing with deafblind individuals will be conducted to further evaluate the system’s usability in real-world environments, while gathering detailed user feedback to optimize user-device interactions and accessibility. Moreover, future research will explore methods to facilitate two-way conversations for deafblind individuals who are unable to speak, such as through the integration of braille keyboards and text-to-speech. Finally, the device aims to be commercialized at a low cost and accessible means to enable global distribution.
Thanks
Ms. Ciobanu and Ms. Pulla
We’d like to extend our gratitude to our teacher advisors, Ms. Ciobanu and Ms. Pulla, who not only supported us throughout the project’s development, but were also the supervisors of our school’s science fair club. Under their guidance, we had the opportunity to mentor dozens of students in developing strong projects and contribute to a record-high number of competitors at our school science fair.
Michelle James - DeafBlind Ontario Services
Michelle provided us with invaluable insights into deafblind accessibility that helped ensure the solution we were building could reach those it aims to serve. We are incredibly grateful for her time and feedback.
Olivia
Thank you to our friend Olivia for letting us use her 3D printer!
Our Parents
Without them, our project would simply not have been possible. Our parents have supported us continuously and words cannot begin to express our gratitude.
References
References
[1] Anthropic. (2024). Claude Sonnet [Large language model]. https://claude.ai
[2] Bredin, H., Yin, R., Coria, J. M., Gelly, G., Korshunov, P., Lavechin, M., Fustes, D., Titeux, H., Bouaziz, W., & Gill, M.-P. (2019, November 4). pyannote.audio: Neural building blocks for speaker diarization. arXiv.org. https://arxiv.org/abs/1911.01255
[3] An intervenor communicating with a person with deafblindness [Photograph]. (2020). London Community Foundation. https://www.lcf.on.ca/stories-backend/2020/8/6/covid-19-grants-deafblind-ontario-services
[4] Cord-Landwehr, T., Gburrek, T., Deegen, M., & Haeb-Umbach, R. (2025, August 29). Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering. arXiv.org. https://arxiv.org/abs/2506.16228
[5] Dawalatabad, N., Ravanelli, M., Grondin, F., Thienpondt, J., Desplanques, B., & Na, H. (2021). ECAPA-TDNN embeddings for speaker diarization. Interspeech 2021, 3560–3564. https://doi.org/10.21437/interspeech.2021-941
[6] Deafblind Information Australia. (2023). Deafblind manual alphabet [Image]. Deafblind Information Australia. https://www.deafblindinformation.org.au/living-with-deafblindness/deafblind-communication/deafblind-manual-alphabet/
[7] DeafBlind Ontario Services. (2018). Open your eyes and ears: To estimates of Canadian individuals with deafblindness and age-related dual sensory loss. https://deafblindontario.com/wp-content/uploads/2025/11/Open-Your-Eyes-and-Ears.pdf
[8] Deafblindness: Challenges, Contributions, and Insights. Hearview. (2025, May 12). https://www.hearview.ai/blogs/news/deafblindness-challenges-contributions-and-insights?srsltid=AfmBOopM2wRilnKXYo0Qv1AgJZgyfQ36iI9WYJJ6WkgsVy5jg91rxdMI
[9] Dementyev, A., Kanevsky, D., Yang, S., Parvaix, M., Lai, C., & Olwal, A. (2025). SpeechCompass: Enhancing mobile captioning with diarization and directional guidance via multi-microphone localization. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–17. https://doi.org/10.1145/3706598.3713631
[10] FireRed Team. (n.d.). FireRedVAD [Computer software]. Hugging Face. https://huggingface.co/FireRedTeam/FireRedVAD/tree/main/Stream-VAD
[11] Hersh, M. (2013). Deafblind people, communication, independence, and isolation. Journal of Deaf Studies and Deaf Education, 18(4), 446–463. https://doi.org/10.1093/deafed/ent022
[12] HumanWare. (n.d.). Brailliant BI 40X braille display. https://store.humanware.com/hus/brailliant-bi-40x-braille-display.html
[13] Landini, F., Profant, J., Diez, M., & Burget, L. (2020, December 29). Bayesian HMM clustering of X-vector sequences (VBX) in speaker diarization: Theory, implementation and analysis on standard tasks. arXiv.org. https://arxiv.org/abs/2012.14952
[14] Learn Deafblind Manual. Deafblind UK. (n.d.). https://deafblind.org.uk/learn-deafblind-manual/
[15] Liang, D., & Li, X. (2025, September 8). LS-EEND: Long-form streaming end-to-end neural diarization with online attractor extraction. arXiv.org. https://arxiv.org/abs/2410.06670
[16] Making a wave from coast to coast since 2015. Deafblind Network of Ontario. (2025, June 2). https://www.deafblindnetworkontario.com/news-events/making-a-wave-from-coast-to-coast-since-2015/
[17] MagniPros. (2025). Image of MagniPros 5X Rechargeable LED page magnifier [Product image]. Amazon Canada. https://www.amazon.ca/MAGNIPROS-Rechargeable-Magnifier-Detachable-HandsFree/dp/B0CZH8PX26/
[18] New England College of Optometry. (2025). Image from "Braille as modern digital assistive technology" [Image]. New England College of Optometry. https://www.neco.edu/news/braille-as-modern-digital-assistive-technology/
[19] Orbit Reader 20 Plus – Braille display, book reader and note-taker. Special Needs Computers. (n.d.). https://specialneedscomputers.ca/products/orbit-reader-20-plus
[20] Pálka, P., Han, J., Delcroix, M., Tawara, N., & Burget, L. (2025, October 22). VBX for End-to-End Neural and Clustering-based Diarization. arXiv.org. https://arxiv.org/abs/2510.19572
[21] SYSTRAN. (n.d.). Faster-whisper [Computer software]. GitHub. https://github.com/SYSTRAN/faster-whisper
[22] UK Government. (2025, November 27). Locked out: Exclusion of deaf and deafblind BSL users from health and social care in the UK (full report – BSL and English versions). https://www.gov.uk/government/publications/bsl-user-experience-of-health-and-social-care-in-uk/locked-out-exclusion-of-deaf-and-deafblind-bsl-users-from-health-and-social-care-in-the-uk-full-report-bsl-and-english-versions
[23] University of Edinburgh. (n.d.). AMI Corpus - Data Problems. AMI Corpus. https://groups.inf.ed.ac.uk/ami/corpus/dataproblems.shtml
[24] Van Den Broeck, B., Bertrand, A., Karsmakers, P., Vanrumste, B., Van hamme, H., & Moonen, M. (2012). Time-domain generalized cross correlation phase transform sound source localization for small microphone arrays. 2012 5th European DSP Education and Research Conference (EDERC), 76–80. https://doi.org/10.1109/ederc.2012.6532229
[25] Varada, V. R. (2023). Electromechanical refreshable braille module [CAD files]. Hackaday.io. https://hackaday.io/project/191181/files
[26] Vispero. (n.d.). Focus 14 Blue 5th generation braille display. https://shop.vispero.com/products/focus-14-blue-5th-generation
[27] Wang, J., Liu, Y., Wang, B., Zhi, Y., Li, S., Xia, S., Zhang, J., Tong, F., Li, L., & Hong, Q. (2022, September 24). Spatial-aware speaker diarization for multi-channel multi-party meeting. arXiv.org. https://arxiv.org/abs/2209.12002
[28] Watters, C., Owen, M., & Munroe, S. (2004). A study of deaf-blind demographics and services in Canada. Canadian National Society of the DeafBlind. https://www.cdbanational.com/wp-content/uploads/2016/03/demographic_study_eng.pdf
[29] World Federation of the Deafblind. (2018). Deafblindness in the world. https://wfdb.eu/deafblindness-in-the-world/
[30] XR ERA. (2022). Image from "The role of wearable haptic devices in XR" [Image]. XR ERA. https://xrera.eu/the-role-of-wearable-haptic-devices-in-xr-meetup-16-recap/
[31] Yong, E. (2016, January 4). The incredible thing we do during conversations. The Atlantic. https://www.theatlantic.com/science/archive/2016/01/the-incredible-thing-we-do-during-conversations/422439/
[32] Zheng, S., Huang, W., Wang, X., Suo, H., Feng, J., & Yan, Z. (2021, July 20). A real-time speaker diarization system based on spatial spectrum. arXiv.org. https://arxiv.org/abs/2107.09321
[33] Zuandi, M. F., Maharani, M. P., & Lim, W. (2018). Performance comparison between steered response power and generalized cross correlation in microphone arrays for sound source localization. ARPN Journal of Engineering and Applied Sciences, 13(9), 3093-3100. https://www.arpnjournals.org/jeas/research_papers/rp_2018/jeas_0518_7036.pdf
Images (27)
Awards (5)
- Young Scientist Award
- Challenge Award
- Special Award
- Silver Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
ISEF · 2024
HearBridge: AR Headset for Interactive Communication via Conversation Textualization and Sign Language Translation for the Hearing Impaired
ISEF · 2024
Dual Sensory
ISEF · 2025
Wearable Translator for Bidirectional Communication Between ASL and Speech
ISEF · 2026
Echo-Glove: Assistive Communication Device
ISEF · 2026
NEXUS: Neurophysiological Emergency eXternal Understanding System- An AI-Integrated Wearable Device Using Multi-Mode Communication and Emotional State Detection in Autism Spectrum Disorder Populations for Elopement Prevention
ISEF · 2021
SoundScape: Real-Time 3D Sound Localization and Classification with Sensory Substitution for the Deaf and Hard of Hearing
ISEF · 2025
HandTalk: A Two-Way Translation System for American Sign Language
ISEF · 2026
A.W.A.R.E.:Development of a Real-Time Wearable Auditory Emergency Detection System for Deaf and Hard of Hearing Individuals Using AI-Based Acoustic Analysis
Closest projects by meaning, across every fair and year in the corpus.