NEXUS: A Multimodal Device for Deafblind Communication using Hybrid Dynamic Speaker Diarization

CWSF · 2026 Digital Technology Silver Medal

Thumbnail supplied by the source for NEXUS: A Multimodal Device for Deafblind Communication using Hybrid Dynamic Speaker Diarization

Overview

Deafblindness is a combined disability of vision and hearing impairments that create significant challenges to communication and information access. Many deafblind individuals rely on intervenors to access spoken information, but access to these services is highly limited due to shortages, uneven availability, and funding restrictions. To address this gap, NEXUS was developed as a wearable assistive device that enables real-time, independent communication. NEXUS captures surrounding speech, identifies who is speaking, and converts conversations into accessible formats like braille and enlarged text, allowing users to follow group conversations without relying on human intermediaries. Notably, the device uses speaker diarization that combines spatial awareness with voice pattern analysis, making speaker identification robust even in complex, multi-speaker environments. This approach outperforms current state-of-the-art models for real-time, far-field diarization. Overall, NEXUS provides an accessible and scalable solution that enhances communication and supports greater autonomy among deafblind communities.

Video

Video

Video Transcript

Hi! My name is Tyler and my name is Leo, and we developed NEXUS, a wearable device that enables real-time, independent communication for individuals with deafblindness.

NEXUS works by capturing surrounding speech, identifying speakers, and transcribing the spoken content into a labelled transcript, clearly showing who is speaking and what they are saying. This information is delivered through a form of accessible output, such as braille or enlarged text, which can be worn on the user’s arm, allowing them to follow both one-on-one and group conversations in real time.

Today, over 160 million people worldwide live with deafblindness, and communication often depends on intervenors, services that face training shortages and uneven availability, preventing many individuals from receiving sufficient aid. Our device removes that barrier by providing deafblind communities with direct, 24 hour access to conversations.

NEXUS has the potential to transform how deafblind communities engage with the world, bringing independent, real-time communication to everyday life, and restoring access to the conversations that connect us all.

Thank you, and see you at the fair!

Why?

Background Information

Deafblindness affects roughly 160 million individuals worldwide, with over 600,000 people living with deafblindness in Canada alone, creating significant barriers to communication (World Federation of the Deafblind [WFDB], 2018). As a condition that impedes both vision and hearing, many everyday interactions, such as conversations, public services, and social gatherings, become difficult without assistance. Consequently, deafblind individuals rely on intervenors, trained professionals who act as a communication bridge between the person and their environment by conveying spoken information through tactile cues (Deafblind Network of Ontario, 2025).

Problem

However, access to intervenor services is often limited due to shortages of trained professionals, uneven service availability, and funding or policy restrictions on support hours (Figure 4). As a result, deafblind services are frequently unavailable for routine interactions, making education, employment, and social interaction challenging. Nevertheless, deafblind communities continue to depend on interpreters for communication, which can limit autonomy and independence for individuals who wish to communicate and navigate everyday environments on their own (UK Government, 2025).

Research Question

How can an accessible and low-cost solution be designed to enable real-time, independent communication for deafblind individuals in everyday settings?

Research Objectives

Analyze the communication needs and challenges of deafblind individuals in real-world environments to support the design of an accessible, portable assistive system.

Design a system that replicates key functions of intervenors by enabling both one-on-one and group conversations accessibly.

Evaluate the system’s effectiveness in conversational settings through the analysis of speaker diarization accuracy across benchmark datasets and real-world testing.

How?

System Overview

NEXUS works by capturing speech, identifying speakers, transcribing speech content, and converting the information into an accessible format for deafblind users in real time. Prototype I focuses on validating the system’s overall flow using a microphone array to capture spatial audio, connected to STM32 and Raspberry Pi 5 microcontrollers for speaker diarization, which means to separate and label different speakers in a conversation (Figures 7 and 8). Transcribed text is paired with its corresponding speakers to clearly convey who is speaking and what they are saying, which is then delivered to the deafblind user through a refreshable braille display using micro stepper motors and octagonal braille drums (Figure 10).

Diarization Pipeline

NEXUS integrates a dynamic hybrid diarization pipeline designed for two key requirements: real-time processing and far-field operation (Figure 19). The pipeline fuses spatial information via GCC-PHAT localization with vocal characteristics from embedding extraction (Pyannote), weighted through a recency-based exponential decay function (Figures 15 and 18). This reduces the influence of inactive speakers over time, introducing a continuity bias that improves assignment accuracy. The pipeline also adapts to varying conversation scenarios, such as short utterances, by dynamically adjusting speaker weighting. Performance was evaluated against existing methods using both a standardized dataset and in-person testing. All results were benchmarked against the current state-of-the-art for real-time, far-field speaker diarization, Pyannote, which served as the baseline.

Wearable Design

Prototype II refines the physical design into a more compact, wearable system optimized for real-world use (Figures 12 and 13). A hot-swappable interface was proposed to support multiple accessible outputs, accommodating the diverse communication preferences of deafblind individuals, largely attributed to variation in disease severity. The following output modules were designed: a 3D printed electromechanical braille display, enlarged text screen, and haptic feedback module (Figure 14).

What?

Testing Types

Validation was conducted using in-person trials and the AMI Corpus Multiple Distant Microphone (MDM) dataset (University of Edinburgh, n.d.). The in-person trials assessed the end-to-end system performance, including hardware integration, while the AMI evaluation focuses purely on the diarization algorithm outside of hardware limitations (Figures 20 and 21). Thus, real world testing allows for greater evaluation of practical effectiveness, while the standardized evaluation allows for fair benchmarking against existing approaches.

In-Person Testing

Two types of trials were conducted through in-person testing: realistic conversations with some speaker movement and overlapping speech and more difficult conversation scenarios with rapid turn-taking (Table 1).

Under the realistic trials, the baseline achieved a diarization error rate (DER) of 20.5%, whereas the proposed pipeline achieved a DER of 9.9%, representing a substantial reduction in error (Figure 23a). The improvement suggests that the integration of spatial information enables more accurate speaker tracking. Under difficult trials, the baseline achieved an average DER of 81.4%, while the proposed pipeline achieved 11.5%, once again demonstrating a significant improvement in accuracy (Figure 23b). Moreover, the proposed pipeline highlights a significant limitation of traditional embedding-only diarization and static fusion pipelines in that they are unable to adapt to different conversation conditions, making them less robust to real-world environments. In particular, these methods over-rely on speaker embeddings, which become unreliable for diarizing short utterances in real-time environments.

Standardized Dataset

An ablation study was performed across 20 conversation scenarios to validate the pipeline’s improvements over the baseline and optimize the weighting mechanisms (Figure 24). Specifically, the mechanism was applied separately to embeddings and spatial information, allowing for a direct comparison of their impact on performance and determining which configuration yields the greatest accuracy. Results demonstrated that the recency-embedding mechanism achieved the highest reduction in DER from a baseline of 39.4% to 30.6%, while spatial recency weighting achieved a DER of 38.6%, reflecting recency-embedding weighting as the most effective approach (Table 1).

Comparing to Existing Approaches

Existing approaches for real-time diarization face several limitations or lack testing against standardized far-field datasets (Table 3). LS-EEND is a live streaming approach to speaker diarization that relies solely on speaker features (Liang & Li, 2025). However, it was primarily trained on the AMI Individual Headset Microphone (IHM) dataset, which consists of near-field audio recordings captured close to each speaker’s mouth, making it unsuitable for far-field environments. SDSS is another approach that utilizes both speaker embeddings and location data (Zheng et al., 2021). However, its evaluation is limited to a custom dataset and integrates hardware setups that are unfeasible for devices like NEXUS, so it is difficult to directly compare its performance against the proposed diarization system. SpeechCompass is another far-field approach, but like SDSS, it was only evaluated on a custom dataset (Dementyev et al., 2025).

The proposed pipeline is therefore the first study of a real-time, far-field diarization system to be validated under a standardized dataset like AMI Corpus, establishing a new benchmark for future research.

So What?

Impact on Deafblind Communication

Testing reveals NEXUS addresses a major accessibility gap in deafblind communities by enabling real-time communication effectively without intervenors, making everyday communication much more accessible. As a fully wearable and portable system, it allows users to communicate anywhere, at any time, without relying on external support. Designed to be low-cost, NEXUS is significantly more accessible than many existing assistive technologies, reducing financial barriers to adoption (Table 4). At the same time, it offers a broader and more advanced feature set, most notably its support for multi-speaker environments. Being low-cost, NEXUS can be easily scaled and distributed, helping alleviate intervenor shortages and the uneven availability of these services. NEXUS’ advancements will contribute to substantial improvements to the quality of life of deafblind communities by enabling greater independence, making education, work, and daily interactions all the more accessible.

Research Significance of Diarization

Beyond the impact on the deafblind community, NEXUS represents a significant advancement in the research field of speaker diarization by outperforming state-of-the-art baselines. Reliable real-time, far-field diarization systems can also be applied outside of assistive technologies, such as voice assistants and meeting transcription (Figure 26).

Limitations

NEXUS is currently unable to convey nuanced conversational cues such as facial expressions or body language, elements that human intervenors can interpret. Additionally, the system has only been validated on indoor conversations and has not yet been tested in diverse environments, such as outdoors, which is an essential next step for a portable device intended to function reliably in varied settings.

What's Next?

Future work will focus on the construction of prototype II and algorithmic refinement, which aims to be complete for the CWSF in-person (Figure 27). Pilot testing with deafblind individuals will be conducted to further evaluate the system’s usability in real-world environments, while gathering detailed user feedback to optimize user-device interactions and accessibility. Moreover, future research will explore methods to facilitate two-way conversations for deafblind individuals who are unable to speak, such as through the integration of braille keyboards and text-to-speech. Finally, the device aims to be commercialized at a low cost and accessible means to enable global distribution.

Thanks

Ms. Ciobanu and Ms. Pulla

We’d like to extend our gratitude to our teacher advisors, Ms. Ciobanu and Ms. Pulla, who not only supported us throughout the project’s development, but were also the supervisors of our school’s science fair club. Under their guidance, we had the opportunity to mentor dozens of students in developing strong projects and contribute to a record-high number of competitors at our school science fair.

Michelle James - DeafBlind Ontario Services

Michelle provided us with invaluable insights into deafblind accessibility that helped ensure the solution we were building could reach those it aims to serve. We are incredibly grateful for her time and feedback.

Olivia

Thank you to our friend Olivia for letting us use her 3D printer!

Our Parents

Without them, our project would simply not have been possible. Our parents have supported us continuously and words cannot begin to express our gratitude.

References

References

[1] Anthropic. (2024). Claude Sonnet [Large language model]. https://claude.ai

[2] Bredin, H., Yin, R., Coria, J. M., Gelly, G., Korshunov, P., Lavechin, M., Fustes, D., Titeux, H., Bouaziz, W., & Gill, M.-P. (2019, November 4). pyannote.audio: Neural building blocks for speaker diarization. arXiv.org. https://arxiv.org/abs/1911.01255

[3] An intervenor communicating with a person with deafblindness [Photograph]. (2020). London Community Foundation. https://www.lcf.on.ca/stories-backend/2020/8/6/covid-19-grants-deafblind-ontario-services

[4] Cord-Landwehr, T., Gburrek, T., Deegen, M., & Haeb-Umbach, R. (2025, August 29). Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering. arXiv.org. https://arxiv.org/abs/2506.16228

[5] Dawalatabad, N., Ravanelli, M., Grondin, F., Thienpondt, J., Desplanques, B., & Na, H. (2021). ECAPA-TDNN embeddings for speaker diarization. Interspeech 2021, 3560–3564. https://doi.org/10.21437/interspeech.2021-941

[6] Deafblind Information Australia. (2023). Deafblind manual alphabet [Image]. Deafblind Information Australia. https://www.deafblindinformation.org.au/living-with-deafblindness/deafblind-communication/deafblind-manual-alphabet/

[7] DeafBlind Ontario Services. (2018). Open your eyes and ears: To estimates of Canadian individuals with deafblindness and age-related dual sensory loss. https://deafblindontario.com/wp-content/uploads/2025/11/Open-Your-Eyes-and-Ears.pdf

[8] Deafblindness: Challenges, Contributions, and Insights. Hearview. (2025, May 12). https://www.hearview.ai/blogs/news/deafblindness-challenges-contributions-and-insights?srsltid=AfmBOopM2wRilnKXYo0Qv1AgJZgyfQ36iI9WYJJ6WkgsVy5jg91rxdMI

[9] Dementyev, A., Kanevsky, D., Yang, S., Parvaix, M., Lai, C., & Olwal, A. (2025). SpeechCompass: Enhancing mobile captioning with diarization and directional guidance via multi-microphone localization. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–17. https://doi.org/10.1145/3706598.3713631

[10] FireRed Team. (n.d.). FireRedVAD [Computer software]. Hugging Face. https://huggingface.co/FireRedTeam/FireRedVAD/tree/main/Stream-VAD

[11] Hersh, M. (2013). Deafblind people, communication, independence, and isolation. Journal of Deaf Studies and Deaf Education, 18(4), 446–463. https://doi.org/10.1093/deafed/ent022

[12] HumanWare. (n.d.). Brailliant BI 40X braille display. https://store.humanware.com/hus/brailliant-bi-40x-braille-display.html

[13] Landini, F., Profant, J., Diez, M., & Burget, L. (2020, December 29). Bayesian HMM clustering of X-vector sequences (VBX) in speaker diarization: Theory, implementation and analysis on standard tasks. arXiv.org. https://arxiv.org/abs/2012.14952

[14] Learn Deafblind Manual. Deafblind UK. (n.d.). https://deafblind.org.uk/learn-deafblind-manual/

[15] Liang, D., & Li, X. (2025, September 8). LS-EEND: Long-form streaming end-to-end neural diarization with online attractor extraction. arXiv.org. https://arxiv.org/abs/2410.06670

[16] Making a wave from coast to coast since 2015. Deafblind Network of Ontario. (2025, June 2). https://www.deafblindnetworkontario.com/news-events/making-a-wave-from-coast-to-coast-since-2015/

[17] MagniPros. (2025). Image of MagniPros 5X Rechargeable LED page magnifier [Product image]. Amazon Canada. https://www.amazon.ca/MAGNIPROS-Rechargeable-Magnifier-Detachable-HandsFree/dp/B0CZH8PX26/

[18] New England College of Optometry. (2025). Image from "Braille as modern digital assistive technology" [Image]. New England College of Optometry. https://www.neco.edu/news/braille-as-modern-digital-assistive-technology/

[19] Orbit Reader 20 Plus – Braille display, book reader and note-taker. Special Needs Computers. (n.d.). https://specialneedscomputers.ca/products/orbit-reader-20-plus

[20] Pálka, P., Han, J., Delcroix, M., Tawara, N., & Burget, L. (2025, October 22). VBX for End-to-End Neural and Clustering-based Diarization. arXiv.org. https://arxiv.org/abs/2510.19572

[21] SYSTRAN. (n.d.). Faster-whisper [Computer software]. GitHub. https://github.com/SYSTRAN/faster-whisper

[22] UK Government. (2025, November 27). Locked out: Exclusion of deaf and deafblind BSL users from health and social care in the UK (full report – BSL and English versions). https://www.gov.uk/government/publications/bsl-user-experience-of-health-and-social-care-in-uk/locked-out-exclusion-of-deaf-and-deafblind-bsl-users-from-health-and-social-care-in-the-uk-full-report-bsl-and-english-versions

[23] University of Edinburgh. (n.d.). AMI Corpus - Data Problems. AMI Corpus. https://groups.inf.ed.ac.uk/ami/corpus/dataproblems.shtml

[24] Van Den Broeck, B., Bertrand, A., Karsmakers, P., Vanrumste, B., Van hamme, H., & Moonen, M. (2012). Time-domain generalized cross correlation phase transform sound source localization for small microphone arrays. 2012 5th European DSP Education and Research Conference (EDERC), 76–80. https://doi.org/10.1109/ederc.2012.6532229

[25] Varada, V. R. (2023). Electromechanical refreshable braille module [CAD files]. Hackaday.io. https://hackaday.io/project/191181/files

[26] Vispero. (n.d.). Focus 14 Blue 5th generation braille display. https://shop.vispero.com/products/focus-14-blue-5th-generation

[27] Wang, J., Liu, Y., Wang, B., Zhi, Y., Li, S., Xia, S., Zhang, J., Tong, F., Li, L., & Hong, Q. (2022, September 24). Spatial-aware speaker diarization for multi-channel multi-party meeting. arXiv.org. https://arxiv.org/abs/2209.12002

[28] Watters, C., Owen, M., & Munroe, S. (2004). A study of deaf­-blind demographics and services in Canada. Canadian National Society of the Deaf­Blind. https://www.cdbanational.com/wp-content/uploads/2016/03/demographic_study_eng.pdf

[29] World Federation of the Deafblind. (2018). Deafblindness in the world. https://wfdb.eu/deafblindness-in-the-world/

[30] XR ERA. (2022). Image from "The role of wearable haptic devices in XR" [Image]. XR ERA. https://xrera.eu/the-role-of-wearable-haptic-devices-in-xr-meetup-16-recap/

[31] Yong, E. (2016, January 4). The incredible thing we do during conversations. The Atlantic. https://www.theatlantic.com/science/archive/2016/01/the-incredible-thing-we-do-during-conversations/422439/

[32] Zheng, S., Huang, W., Wang, X., Suo, H., Feng, J., & Yan, Z. (2021, July 20). A real-time speaker diarization system based on spatial spectrum. arXiv.org. https://arxiv.org/abs/2107.09321

[33] Zuandi, M. F., Maharani, M. P., & Lim, W. (2018). Performance comparison between steered response power and generalized cross correlation in microphone arrays for sound source localization. ARPN Journal of Engineering and Applied Sciences, 13(9), 3093-3100. https://www.arpnjournals.org/jeas/research_papers/rp_2018/jeas_0518_7036.pdf

Images (27)

Awards (5)

  • Young Scientist Award
  • Challenge Award
  • Special Award
  • Silver Medal
  • Selected for CWSF 2026

Competition history

  • CWSF 2026 Digital Technology Qualified through York, ON

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: ProjectBoard / Youth Science Canada

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google