Applying Fractal Structure to Vision Transformer (ViT)
ISEF · 2023 Robotics and Intelligent Machines
Overview
Current ResNet structure uses convolution kernels to preserve local information of an image. In contrast, ViT divides an image into patches to calculate their global attention. With enough amount of training data, ViT can outperform CNN in image classification. However, the lack of local information considered in ViT results in ViT’s lower performance with a small set of data. In this research, to increase the local information considered by ViT, an input image will be patched into tokens of different levels that are designed to form a fractal pattern. The research improves the accuracy of original ViT while decreasing the amount of parameter being used. Using this, Fractal ViT and ViT are trained with satellite images.
Awards (1)
- King Abdulaziz & his Companions Foundation for Giftedness and Creativity: Full Scholarship from King Fahd University of Petroleum and Minerals(KFUPM) (and a $400 cash prize) $400
Competition history
- ISEF 2023
Resources
Related projects
ISEF · 2024
Revolutionizing Non-Invasive Skin Cancer Detection Through a Novel Vision Transformer Application
ISEF · 2024
Multi-Scale Knowledge Transfer Convolutional Transformer: A Novel Deep Learning Framework for 3D Brain Vessel Segmentation
ISEF · 2021
CET-CNN: Modular Hierarchical Image Classification Using Conditional-Execution Tree CNNs
ISEF · 2025
WeCAViT: Weighted Ensemble of CNN, Attention, and a Visual Transformer - A Novel Ensemble Neural Network Architecture for Accurate Pneumonia Diagnosis From Chest X-Rays
Closest projects by meaning, across every fair and year in the corpus.
Source: Regeneron International Science and Engineering Fair