A Novel Mechanistic Framework for Identifying and Neutralizing Latent Deception Signatures in Large Language Models
CSEF · 2026 Computational Science (Senior Division)
Overview
Though AI is embedded in critical infrastructure worldwide, the underlying components remain vulnerable and easily manipulated. Public repositories like Hugging Face host millions of model components downloaded billions of times, with 50K+ models identified as unsafe. Studies show that over 86% of organizations have already reported AI security incidents. Among the most dangerous attacks are backdoors that induce deceptive behavior, causing models to behave normally under safety evaluations but maliciously when triggered. Existing defenses require knowing the trigger in advance or rely on ineffective monitoring, rendering backdoor attacks nearly undetectable. This project introduces the first trigger-agnostic pipeline to detect and remove backdoors in Low Rank Adaptation, the dominant LLM fine-tuning method. The mechanistic framework examines adapter weights to uncover latent deception signatures (structural anomalies invisible to standard evaluation), retaining efficacy where all previous defenses fail. The multi-phase pipeline applies Singular Value Decomposition across all internal weight modules, flagging anomalous spectral norms indicative of backdoors. Causal verification confirms the backdoor’s source using isolated module injection against three statistical controls. Finally, the minimal set of weight directions carrying >95% of the backdoor signal is identified and removed via Rank-K Deflation, a targeted modification requiring no retraining and leaving other parameters untouched. The framework was validated against clean and backdoored adapters, including Anthropic’s Sleeper Agent attack (engineered to bypass all state-of-the-art defenses). The pipeline restored refusal rates of malicious requests from 12% to 88%, matching clean baseline performance, while producing near-zero utility loss and zero false positives. This framework can effectively screen and clean uploaded adapters, establishing a new technical standard for AI supply chain security.
Competition history
- CSEF 2026
Related projects
CSEF · 2026
ClearVision: A Novel Mitigation System Against Unknown Adversarial Attacks for Image Recognition AI Models
ISEF · 2026
Shadow: A Cross Domain, Mathematically Validated, Meta-Cognitive Reasoning Infrastructure for AI Models
ISEF · 2025
Authorship Verification for Academic Dishonesty in the Era of AI
ISEF · 2024
A Novel Approach to Detecting Academic Dishonesty Involving Artificial Intelligence
ISEF · 2025
Integrity: Generalized Artificial Image Classification With Noise Domain Localization
ISEF · 2025
SplitSafe: A Novel Adversarial Attack Detection and Mitigation Technique for Artificial Intelligence Image Recognition Systems
CSEF · 2026
A Deep-Learning Based Cascading Framework for Intrusion Detection and Attack Classification on V2X-Based Autonomous Vehi
CSEF · 2026
Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: California Science & Engineering Fair public projects