← Back to Explore

Applying Fractal Structure to Vision Transformer (ViT)

ISEF · 2023 Robotics and Intelligent Machines

Overview

Current ResNet structure uses convolution kernels to preserve local information of an image. In contrast, ViT divides an image into patches to calculate their global attention. With enough amount of training data, ViT can outperform CNN in image classification. However, the lack of local information considered in ViT results in ViT’s lower performance with a small set of data. In this research, to increase the local information considered by ViT, an input image will be patched into tokens of different levels that are designed to form a fractal pattern. The research improves the accuracy of original ViT while decreasing the amount of parameter being used. Using this, Fractal ViT and ViT are trained with satellite images.

Awards (1)

  • King Abdulaziz & his Companions Foundation for Giftedness and Creativity: Full Scholarship from King Fahd University of Petroleum and Minerals(KFUPM) (and a $400 cash prize) $400

Competition history

  • ISEF 2023 Robotics and Intelligent Machines · Entry ROBO015 · Dallas, Texas, United States

Resources

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Source: Regeneron International Science and Engineering Fair

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. Browsing stays public.

Continue with Google