Increasing Computer Vision Models Interpretability
ISEF · 2023 Systems Software
Overview
With the perpetual increase of complexity of the state-of-the-art deep neural networks, it becomes a more and more challenging task to maintain their interpretability. Our work aims to evaluate the effects of adversarial training utilized to produce robust models - less vulnerable to adversarial attacks. It has been shown to make computer vision models more interpretable. Interpretability is as essential as robustness when we deploy the models to the real world. To prove there is a correlation between these two problems, we extensively examine the models using local feature-importance methods (SHAP, Integrated Gradients) and feature visualization techniques (Representation Inversion, Class Specific Image Generation). Standard models, compared to robust ones are less secure, and their learned representations are less meaningful to humans. Conversely, robust models focus on distinctive regions of the images which support their predictions. Moreover, the features learned by the robust model are closer to the real ones.
Competition history
- ISEF 2023
Resources
Related projects
ISEF · 2021
Towards Malware Classifiers Robust to Adversarial Malware
ISEF · 2021
A Biologically Inspired Game Theoretic Adversarial Training Method
ISEF · 2025
SplitSafe: A Novel Adversarial Attack Detection and Mitigation Technique for Artificial Intelligence Image Recognition Systems
ISEF · 2024
AdvMed: Detecting Adversarial Attacks in Medical Deep Learning Systems
Closest projects by meaning, across every fair and year in the corpus.
Source: Regeneron International Science and Engineering Fair