VisionSpeak: A Wearable Edge-AI System That Describes the World Through Synthesized Speech for the Visually Impaired

CSEF · 2026 Mathematical Sciences (Junior Division)

Overview

VisionSpeak: A Wearable Edge-AI System That Describes the World Through Synthesized Speech for the Visually Impaired Shreyan Biswas Our project aims to enhance independent mobility and safety for the 2.2 billion individuals globally with visual impairments. While current market solutions exist, they are usually unaffordable and often rely on cloud connectivity, introducing high latency and privacy risks. To address this, we developed "VisionSpeak" a wearable system that provides holistic scene understanding through describing surroundings and issuing safety alerts in real-time without internet dependency. The system is powered by an NVIDIA Jetson Orin Nano, shifting computer vision and language processing from the cloud to edge hardware. A camera continuously captures the user’s surroundings; the onboard GPU runs a vision transformer to generate descriptive captions, which are then analyzed by a locally running Llama large language model to determine whether the situation contains caution or danger (e.g., obstacles, vehicles, drop-offs, or other hazards). Both the scene description and any safety alert are delivered through audio, enabling immediate, hands-free feedback in a portable form factor suitable for daily use. We stress-tested our architecture across diverse environmental conditions, including low-light and dynamic outdoor scenarios, to simulate real-world navigation. Data analysis revealed that the Florence-2 model significantly outperformed other architectures. It achieved a caption relevance score (SPICE/CIDEr) exceeding 0.80, successfully identifying potential hazards and contextual details. It maintained an inference latency under 1s. In conclusion, we successfully created a high quality, privacy-preserving assistive device that runs offline. The final prototype integrates the optimized Florence-2 pipeline on the Orin Nano to convert visual input into near real-time synthesized speech and hazard alerts suitable for daily navigation.

Competition history

  • CSEF 2026 Mathematical Sciences (Junior Division) · Entry J-14-09

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google