A Novel Mechanistic Framework for Identifying and Neutralizing Latent Deception Signatures in Large Language Models

CSEF · 2026 Computational Science (Senior Division)

Overview

Though AI is embedded in critical infrastructure worldwide, the underlying components remain vulnerable and easily manipulated. Public repositories like Hugging Face host millions of model components downloaded billions of times, with 50K+ models identified as unsafe. Studies show that over 86% of organizations have already reported AI security incidents. Among the most dangerous attacks are backdoors that induce deceptive behavior, causing models to behave normally under safety evaluations but maliciously when triggered. Existing defenses require knowing the trigger in advance or rely on ineffective monitoring, rendering backdoor attacks nearly undetectable. This project introduces the first trigger-agnostic pipeline to detect and remove backdoors in Low Rank Adaptation, the dominant LLM fine-tuning method. The mechanistic framework examines adapter weights to uncover latent deception signatures (structural anomalies invisible to standard evaluation), retaining efficacy where all previous defenses fail. The multi-phase pipeline applies Singular Value Decomposition across all internal weight modules, flagging anomalous spectral norms indicative of backdoors. Causal verification confirms the backdoor’s source using isolated module injection against three statistical controls. Finally, the minimal set of weight directions carrying >95% of the backdoor signal is identified and removed via Rank-K Deflation, a targeted modification requiring no retraining and leaving other parameters untouched. The framework was validated against clean and backdoored adapters, including Anthropic’s Sleeper Agent attack (engineered to bypass all state-of-the-art defenses). The pipeline restored refusal rates of malicious requests from 12% to 88%, matching clean baseline performance, while producing near-zero utility loss and zero false positives. This framework can effectively screen and clean uploaded adapters, establishing a new technical standard for AI supply chain security.

Competition history

  • CSEF 2026 Computational Science (Senior Division) · Entry S-07-19

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google