VisionIQ - A Novel Multi-Model AI Approach for Visual Impairment
CWSF · 2026 Digital Technology
Overview
2.2 billion people worldwide live with vision impairment, yet no existing assistive device offers a true level of independence in real-world conditions. VisionIQ is an AI-powered wearable built for $210 CAD, running entirely offline with no phone, screen, or internet required. VisionIQ operates through three modes: Active Mode continuously monitors surroundings, proactively warning about nearby obstacles unprompted. Read Mode reads any printed text aloud. Chat Mode enables open-ended conversation and scene description entirely on device. The system was designed end-to-end, the electrical schematics in KiCad, custom 3D-printed housing enclosure in Fusion 360, and a novel priority-queue multiprocessing architecture coordinating 7 specialized neural networks across all 4 CPU cores. Every model was individually benchmarked. All navigation-critical modes complete <1s, with wake detection as fast as 211ms. VisionIQ was tested with visually impaired users, who confirmed it addresses gaps no existing assistive device has closed. All-in-one. Real-time. Offline. Secure. Hands-free. Affordable. VisionIQ.
Video
This video could not be played here. Watch it on the original project page.
This video could not be played here. Watch it on the original project page.
Video
My project is VisionIQ consist of hardware and software helping the visually impaired to navigate independently with my Innovative solution offering the 3 modes - Active, Read and Chat with a very convenient wearable glass, which is All-in-one, Real-time, Affordable, Hands-free, Edge device, Highly secure and extremely user friendly.
Why?
The inspiration for VisionIQ came from meeting Sabrina, a visually impaired community volunteer with 6% vision due to Stargardt's Disease, struggling daily to navigate, read, and process visual information - even with a traditional cane and an assistive device(OrCam MyEye) costing ~$6000 that she described as unreliable, offline-incapable, and ultimately useless. Her community members faced identical failures with alternatives like Meta smart glasses. The gap in assistive technology is not awareness, it is the actual product itself.
According to the World Health Organization, 2.2 billion people worldwide, and 1.5 million in Canada (CNIB) live with vision impairment or blindness. Traditional assistive devices like white canes cannot provide real-time environmental context or object identification. Current technology alternatives require constant internet connectivity, a paired smartphone, or compromise user privacy entirely. No existing solution combines proactive obstacle detection, OCR, scene understanding, and conversational AI in a single offline edge device. VisionIQ addresses this significant gap.
VisionIQ is an AI-powered wearable that gives visually impaired users a safer, more independent experience through an original priority-queue multiprocessing architecture coordinating 7 specialized neural networks across all 4 CPU cores - running entirely offline, with no phone, screen, or internet required.
VisionIQ continuously monitors surroundings, warning about nearby obstacles with prioritization, distance, and direction - unprompted. It reads printed text aloud, describes scenes in natural language, and holds open-ended conversations with contextual memory, all voice-activated. Everything runs locally, keeping user data completely private and sovereign.
All-in-one, Real-time, Affordable, Hands-free, Edge device, Novel, Private and Highly secure
How?
VisionIQ operates through three core modes: Active, Read, and Chat:
Active Mode continuously monitors the user's surroundings at 9 FPS, proactively warning about nearby obstacles with prioritization, distance, and direction all unprompted.
Read Mode reads printed text aloud from the camera in real time, signs, labels, menus, all just activated by a simple spoken command.
Chat Mode enables the user to describe scenes in natural language and hold open-ended conversations, activated by voice. All modes are voice-activated, with responses delivered through bone conduction audio, keeping the ear canal fully open for natural spatial awareness and safety.
Electrical
VisionIQ's electrical system was designed using KiCad for schematic capture. The compute core is a Raspberry Pi CM5 (8GB) on a Waveshare CM5-NANO-A carrier board. Vision input is provided by a Raspberry Pi Camera Module 3 Wide. Voice commands are captured by an INMP441 MEMS microphone at the nose bridge, feeding into two MAX98357A I2S amplifiers driving Dayton Audio BCE-1 bone conduction transducers for stereo audio output. Power is supplied via an 18650 UPS battery through a coiled USB-C cable.
Mechanical
The enclosure was designed in Autodesk Fusion 360 as a custom housing enclosure for the CM5. A complete CAD assembly verified component fitting and wearability before printing. Sliced using Orca Slicer with gyroid infill for structural efficiency and thermal resistance, printed in PETG material.
Pipeline & Software
VisionIQ is written entirely in Python. OpenCV handles camera frame capture across every mode. SciPy and NumPy manage audio resampling and numerical operations. Ollama runs local LLM and VLM inference. Ultralytics manages the YOLO pipeline. Vosk, Whisper, EasyOCR, Tesseract, and Piper TTS each run as dedicated processes in the multiprocessing architecture, controlled by a main coordinator/process across all 4 CPU cores with 6 priority queues.
Every component runs entirely offline - no cloud, no internet required.
What?
Data Collection & Model Training
YOLO was trained on the COCO Dataset , which has 130,000 images across 80 common object classes,. Augmentations such as flip, mosaic, HSV shift, and scale jitter were applied. Additional images of navigation-critical objects including stairs, curbs, and drop-offs were independently collected and added on top of COCO to prioritize hazards most relevant to visually impaired navigation.
Training was conducted on Google Colab Pro using an NVIDIA A100 GPU, with all pipelines written in Python and monitored within a Jupyter notebook environment for interactive tracking of loss curves and metrics.
5 YOLO26 configs were benchmarked across their resolution, mAP, precision, recall, and inference speed directly on CM5.
YOLO26n at 416px delivers the best overall balance, getting mAP50 of 0.461 and mAP50-95 of 0.318 at ~110ms inference (~9 FPS).
Higher resolution models at 640px gain only 0.475 mAP50 but 3x inference time.
YOLO26s at 460px achieves mAP50 of 0.490 but at ~400ms, nearly 4x slower.
YOLO26n at 416px was selected for final deployment.
NCNN + INT8 Edge Optimization
The trained PyTorch model(.pt) of weights was then exported to ONNX, then converted to NCNN, a high-performance inference framework built specifically for ARM processors like the CM5, eliminating general-purpose framework overhead.
INT8 quantization compresses each 32-bit float weight to an 8-bit int using a scale factor and zero-point offset, reducing memory by 4x and cutting inference time significantly.
Result: 9 FPS continuous object detection on CPU-only hardware with 4x memory reduction, making real-time navigation practical entirely on edge hardware
Frame Analysis - Signals Fallback & Temporal Filter
YOLO cannot detect walls, blank doors, or unmarked surfaces, which is a critical gap for real-world navigation. My own CV fallback was engineered using three seperate signals that run in parallel on raw frames.
Signal 1 (Edge Density) uses Canny detection on the centre 50% of the frame, firing if more than 6% of pixels are edges, which indicates an obstacle ahead.
Signal 2 (Blank Wall) checks for low pixel standard deviation combined with low edge density, confirming a flat uniform surface ahead.
Signal 3 (Floor Hidden) detects when centre brightness matches the floor, showing a wall directly in front. An obstacle confirmed only when 2/3 signals present, balancing sensitivity against false-positive rates.
A temporal filter is implemented, that requires any detected object to appear in at least 2/3 last consecutive frames before VisionIQ acts, eliminating single-frame hallucinations while maximizing detection responsiveness.
Custom Pinhole Distance Estimation & Lateral Positioning
Focal length was mathematically derived from the Camera Module 3's FOV datasheet:
Distance is computed using per-class real-world heights for all 80+ COCO object classes, with a calibration factor of 0.35 that was determined empirically through thorough iterative real-world testing until estimates matched ground truth.
Lateral position is divided into Left (<40%), Centre (40-60%), and Right (>60%), physically verified by walking past objects in both directions to confirm correct directional output.
This enables precise navigation guidance: "Chair to your right, under 1 metre."
So What?
Meeting Sabrina - diagnosed with Stargardt's Disease patient with 6% vision who spent her life savings on an assistive device that now sits unused - is probably the most powerful validation of VisionIQ's purpose. Her emphasis on offline reliability, privacy, and an all-in-one form factor confirmed that the gap in assistive technology is there.
The all-in-one ideology and prototyping is very encouraging - this has the great potential to reach many visually impaired individuals
I spent $6000 of my savings on my OrCam. It sits in the drawer - useless, offline, unreliable, no good interface. VisionIQ works well at all the modes.
We just want to be independent. That's what assistive tech should do.
Individual Model Performance Each of VisionIQ's 7 AI models was benchmarked independently across real-world conditions. EasyOCR achieved 96.7% accuracy in clear conditions. Vosk maintained 100% wake word detection in quiet and 76.7% in crowd noise. Whisper maintained under 8% word error rate across all noise environments. Piper TTS achieved a real-time factor of ~0.11, synthesizing speech 9x faster than real-time.
System Latency All navigation-critical modes complete under 1 second. Wake detection at 211ms ensures a near-instant response with no perceptible delay. Every component runs entirely offline, no cloud, no network latency contribution.
Real-World Navigation Trial Three participants navigated a 10-obstacle indoor course blindfolded, once with cane only, once with cane + VisionIQ. Obstacle avoidance improved from an average of 43% to 85% - a consistent +40 percentage point gain across every single participant, with zero participants failing to improve.
What's Next?
VisionIQ has proven itself as a technically feasible and practically effective assistive system. The next phase focuses on four key areas:
Continued hardware miniaturization to reduce weight and improve wearability
Expanding the YOLO training dataset with even more additional object classes and environments for broader real-world coverage
Scaling user trials through RESNA(Rehabilitation Engineering and Assistive Technology Society of North America) to collect structured feedback from a wider visually impaired community
Implementing hand gesture activation and facial recognition of familiar faces - features directly requested by a legally blind volunteer(Sabrina) and identified as critical gaps absent in most current assistive technology.
Thanks
There are multiple people who I would like to acknowledge for their support over the past year,
Thank you to Youth Science Canada and the Canada-Wide Science Fair for the opportunity to present VisionIQ on a national stage.
A special thank you to the Quinte Regional Science and Technology Fair for selecting VisionIQ to represent our region at CWSF 2026 and Mr.Scott Berry, Mr.Christopher Spencer, and Ms.Natasha Mathieu for making the CWSF journey towards Edmonton smoothly.
My special thanks to Dr.Arjun Yogeswaran for the encouragement and guidance.
To my parents - thank you for your continued encouragement, support, and patience throughout every stage of this project.
To my teachers at Eastside Secondary School - The guidance and support from all the teachers throughout this science fair journey.
My special thanks to Mrs.Sabrina Maracle for testing my Vision IQ and providing her valuable feedback.
References
Statistics and Background
Canadian National Institute for the Blind. (2024). Blindness in Canada. https://www.cnib.ca/en/sight-loss-info/blindness/blindness-canada
Delpero, W., Trope, G., Buys, Y., Yan, P., Brent, M., Liu, S., & Jin, Y. (2023). Prevalence of self-reported visual impairment among people in Canada with and without diabetes. CMAJ Open, 11(6), E1125–E1133. https://doi.org/10.9778/cmajo.20220116
Maulik, P. K., Mascarenhas, M. N., Mathers, C. D., Dua, T., & Saxena, S. (2018). Prevalence and determinants of visual impairment in Canada: Cross-sectional data from the Canadian Longitudinal Study on Aging. Canadian Journal of Ophthalmology, 53(3), 291–297. https://doi.org/10.1016/j.jcjo.2017.11.013
Ngo, C., Man, R., & Fenwick, E. (2024). Visual impairment, employment status, and reduction in income: The Canadian Longitudinal Study on Aging. Canadian Journal of Ophthalmology, 59(3), 201–208. https://doi.org/10.1016/j.jcjo.2024.01.005
Wittenborn, J., & Rein, D. (2016). The impact of vision loss. In A. Welp, R. B. Woodbury, M. A. McCoy, & T. A. Teutsch (Eds.), Making eye health a population health imperative: Vision for tomorrow. National Academies Press. https://www.ncbi.nlm.nih.gov/books/NBK402367/
World Health Organization. (2019). World report on vision. WHO Press. https://www.who.int/publications/i/item/9789241516570
World Health Organization. (2023). Blindness and vision impairment [Fact sheet]. https://www.who.int/news-room/fact-sheets/detail/blindness-and-visual-impairment
Limitations of Commercial AI Assistive Devices
Guide Dogs UK. (2024). Reviewing the Ray-Ban Meta smart glasses for vision impaired users. https://www.guidedogs.org.uk/blog/reviewing-ray-ban-meta-smart-glasses
Perkins School for the Blind. (2024). How I use Be My Eyes with low vision. https://www.perkins.org/resource/be-my-eyes-app-review/
Waisberg, E., Ong, J., Masalkhi, M., Zaman, N., Sarker, P., Lee, A. G., & Tavakkoli, A. (2023). Meta smart glasses - large language models and the future for assistive glasses for individuals with vision impairments. Eye, 38, 1–3. https://doi.org/10.1038/s41433-023-02842-z
Limitations of Traditional Assistive Technology
Pittet, C. E., & Murray, M. M. (2026). Efficacy of electronic travel aids for the blind and visually impaired during wayfinding. Scientific Reports, 16, Article 6423. https://doi.org/10.1038/s41598-026-37578-9
Software and Frameworks
Alpha Cephei. (2024). Vosk offline speech recognition API. https://alphacephei.com/vosk/
Google. (2024). Google Colaboratory. https://colab.research.google.com/
Jaided AI. (2024). EasyOCR: Ready-to-use OCR with 80+ supported languages [Computer software]. GitHub. https://github.com/JaidedAI/EasyOCR
Jocher, G., Chaurasia, A., & Qiu, J. (2023). Ultralytics YOLO (Version 8.0.0) [Computer software]. https://github.com/ultralytics/ultralytics
NumPy Developers. (2024). NumPy: The fundamental package for scientific computing with Python. https://numpy.org/
Ollama. (2024). Ollama: Run large language models locally. https://ollama.com/
Ong, V. (2024). Moondream: A tiny vision language model [Computer software]. GitHub. https://github.com/vikhyat/moondream
OpenAI. (2022). Introducing Whisper. https://openai.com/research/whisper
OpenCV. (2024). OpenCV: Open source computer vision library. https://opencv.org/
Python Software Foundation. (2024). Python programming language. https://www.python.org/
Qwen Team, Alibaba Cloud. (2024). Qwen 2.5: Instruction-tuned language models. https://qwenlm.github.io/blog/qwen2.5/
Rhasspy. (2024). Piper: A fast, local neural text-to-speech system [Computer software]. GitHub. https://github.com/rhasspy/piper
SciPy Developers. (2024). SciPy: Fundamental algorithms for scientific computing in Python. https://scipy.org/
Tencent. (2024). NCNN: High-performance neural network inference framework [Computer software]. GitHub. https://github.com/Tencent/ncnn
Ultralytics. (2024). Ultralytics YOLO documentation. https://docs.ultralytics.com/
Hardware Documentation
Dayton Audio. (2024). BCE-1 bone conduction exciter. Parts Express. https://www.parts-express.com/Dayton-Audio-BCE-1-Bone-Conduction-Exciter-295-214
InvenSense (TDK). (2015). INMP441 omnidirectional microphone with I²S digital output - datasheet (Rev. 1.3). https://invensense.tdk.com/wp-content/uploads/2015/02/INMP441.pdf
Maxim Integrated (Analog Devices). (2020). MAX98357A PCM input class D audio power amplifier - datasheet. https://www.analog.com/media/en/technical-documentation/data-sheets/MAX98357A-MAX98357B.pdf
Raspberry Pi Ltd. (2024). Camera Module 3 - product brief. https://datasheets.raspberrypi.com/camera/camera-module-3-product-brief.pdf
Raspberry Pi Ltd. (2024). Compute Module 5 - datasheet. https://datasheets.raspberrypi.com/cm5/cm5-datasheet.pdf
Raspberry Pi Ltd. (2024). Compute Module 5. https://www.raspberrypi.com/products/compute-module-5/
Waveshare Electronics. (2024). CM5-NANO-A carrier board - wiki. https://www.waveshare.com/wiki/CM5-NANO-A
Datasets
COCO Consortium. (2024). COCO: Common objects in context [Dataset]. https://cocodataset.org/
Images (28)
Awards (2)
- Special Award
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
ISEF · 2024
ViABL: Visual Assistant for the Blind With VLMs
ISEF · 2026
E-Vision: Helping the Blind Experience the World Again
ISEF · 2024
GIVS: A Novel, Generative Artificial Intelligence Vision System for the Visually Impaired
ISEF · 2025
Development of Assistive Technologies for the Visually Impaired Using AI
ISEF · 2022
IVY - Intelligent Vision System for the Visually Impaired
ISEF · 2023
IVY: Intelligent Vision System for the Visually Impaired
ISEF · 2025
A Novel Approach To Using Artificial Intelligence to Aid the Hearing and Vision Impaired
ISEF · 2024
AI-Powered Vision for Enhanced Spatial Navigation of the Visually Impaired
Closest projects by meaning, across every fair and year in the corpus.