MERIT: Mechanistic Explainability of Reasoning Integrity and Transparency
ISEF · 2026 Robotics and Intelligent Machines
Overview
Large language models (LLMs) are increasingly trusted in domains such as education, research, and decision making. Despite their widespread adoption, fundamental questions remain about alignment and logic in LLM reasoning. This assurance gap poses significant risks in high-stakes domains. This study develops a mechanistic framework for reasoning verification by examining internal model computations rather than relying solely on behavioral outputs. I analyzed DeepSeek-R1 Distill Llama-8B latent feature activations across 2,000 mathematical problems of varying difficulty and domain. Employing Sparse Autoencoders for feature extraction, I identify specific internal features directly tied to reasoning behavior. I develop six novel metrics to quantify reasoning quality. Through causal intervention analysis, I demonstrate that general features fracture into domain-specific specialists under increased task complexity. Experimental manipulation of identified features significantly boosts reasoning behavior, establishing that internal features functionally control aspects of reasoning. At higher mathematical difficulties, domain-experts emerge for geometry, number theory, and other subdomains. Additionally, I novelly demonstrate the existence of distinct feature-governed reasoning modalities: a concise calculation-oriented mode and a verbose explanation-oriented mode. Finally, I create an LLM optimization harness with my mechanistic findings, and successfully increase LLM reasoning accuracy by 22% on a test set of 1,000 advanced mathematical problems. This framework enables assurance and optimization of LLMs with implications for all reasoning-enabled AI applications.
Awards (2)
- Third Award of $1,200 $1,200
- Midwest Microelectronics Consortium: Four cash awards of $3,000 each. $3,000
Competition history
- ISEF 2026
Resources
Related projects
ISEF · 2026
Shadow: A Cross Domain, Mathematically Validated, Meta-Cognitive Reasoning Infrastructure for AI Models
ISEF · 2024
Bias in Large Language Models (LLMs); Paving the Way for an Equitable Artificial General Intelligence (AGI)
ISEF · 2026
LLMs Know When We Are Watching: A Lightweight Framework to Quantify Evaluation Awareness
ISEF · 2026
Keep Your Data Close, but Your Failures Closer: Failure-Driven Adversarial Self-Evolution of Language Models
ISEF · 2017
Providing a Method for Neural Networks to Justify Their Conclusions in Both Prediction and Classification Problems
ISEF · 2026
DOTBot: Diffusion-Aligned Chain-of-Thought Bipedal Humanoid Robot for Interpretable Real-Time Autonomous Navigation, Hazard Mitigation, and Recovery in High-Risk Built Environments
ISEF · 2026
The Urgency Axis: Mapping Crisis Severity Representations in Large Language Models
ISEF · 2023
BrainTrain: Aligning Deep Neural Networks to Human Behavior to Improve Robustness and Generalization
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: Regeneron International Science and Engineering Fair