Computer Vision For Bacterial Gram Stain Classification

CWSF · 2026 Disease & Illness Gold Medal

Thumbnail supplied by the source for Computer Vision For Bacterial Gram Stain Classification

Overview

Antibiotic resistance is one of the most urgent crises in global health, and the first step in fighting it is correctly identifying whether a bacterium is Gram-positive or Gram-negative — a distinction that directly determines which antibiotics will work. Manual microscopy interpretation is slow, subjective, and inaccessible in resource-limited settings. This project developed four progressively optimized convolutional neural networks (CNNs) trained on 8,000 labeled Gram-stained microscopy images. The final model hit 98.33% test accuracy, and I reproduced that result across two independent runs. These results put us ahead of the 94.9% benchmark reported by Smith et al. (Harvard Medical School, 2018) and the 95.1% benchmark reported by Hee Kim et al.(Heidelberg University, 2023). Grad-CAM visualization confirms the model learns genuine bacterial features, rather than background or other artifacts in the slide. This demonstrates that reliable, automated Gram stain classification is achievable without specialized equipment or proprietary data.

Video

Why?

Gram staining is the most important step in identifying bacterial infections (McMahon 2025). It classifies bacteria into Gram‑positive and Gram‑negative groups (Figure 2), which determines the first antibiotic a patient receives. This matters because sepsis is a severe, life‑threatening condition caused by infection (Figure1) and causes over 11 million deaths every year worldwide (Rudd 2020).

Early, accurate treatment saves lives, and Gram stain interpretation is the very first clue clinicians rely on. But despite its importance, Gram stain interpretation is still manual. It depends on technician experience, staining quality, and time. In many settings, especially where trained microbiologists are limited, results can be inconsistent or delayed (Figure 3). Delay in starting antibiotic therapy increases the risk of death by 10% in sepsis patients (Ferrer 2014).

I want to be a physician, and I believe that future physicians need to understand AI and not just use it. Inspired by courses from Harvard and Stanford on computer science and deep learning for computer vision I wondered whether a model could learn the subtle differences in color, texture, and morphology that distinguish Gram‑positive from Gram‑negative bacteria.

I wanted to build something that works reliably, runs offline, and is accessible to any lab anywhere in the world. I built a complete deep‑learning pipeline trained on a dataset of 8000 labeled images using my own desktop. I also developed a Gradio app that can classify images directly without using expensive equipment.

How?

Background Research

I first came across online Harvard University's CS50 Introduction to Computer Science (Yu 2020) in the summer of grade 2 and Stanford's CS231N Deep Learning for Computer Vision (Karpathy 2015) in summer of grade 6 and I really liked both the courses (Figure 4). I learned more about automation of Gram staining by reading articles which were peer reviewed, published in the Journal of Clinical Microbiology and the authors had excellent institutional affiliation (Harvard, Dartmouth, Heidelberg).

Dataset & Data Quality

I used the University of Heidelberg (MHU) open-source dataset (Kim et al 2023) which has 8,500 labeled Gram stain microscopy images collected between 2015 and 2019 using a camera. It is the largest public dataset of its kind, and is 10 times bigger than the next largest open dataset (Figure 5). I removed around 200 images that were almost entirely black-and-white and therefore unsuitable for color-based classification (Figure 6). Before doing this step, my model accuracy was poor and I learned that it is very important to check data for such errors.

Model development

I started with a simple baseline 2-layer TensorFlow (Abadi et al 2014) based CNN, which achieved about 59% accuracy which is only slightly better than random guessing (Figure 7). By analyzing its errors, I redesigned the architecture and the second model reached around 92% accuracy, and then I added dropout and other complex machine learning techniques to make the model achieve 95% accuracy (Figure 8).

Materials: Python, PyTorch, NumPy, OpenCV, scikit-learn; NVIDIA RTX 3060 Ti GPU; GitHub Codespaces.

What?

The breakthrough came when I applied transfer learning using a pretrained ResNet (K He et al 2015) architecture. Instead of learning everything from scratch, the model could build on features learned from millions of images of ImageNet. (Figure 9)

Achieving 99.53% accuracy on 1,600 unseen test images, the finalized ResNet18 model surpassed every published benchmark for this dataset. (Kim H 2023, McMahon 2025) , and that was confirmed across two independent training runs.

Despite a decade of sustained effort, no automated Gram stain system has achieved regulatory approval or wide clinical implementation. Walter et al. (2024) concluded that their prospectively validated 1,555-slide CNN: "is not yet ready for clinical implementation” as “we suggest that an error rate of 1% to 2% might be a reasonable benchmark for the accuracy of manual Gram stain interpretations”.

Smith et al. (2018) at Harvard Medical School used the MetaFer clinical automated scanner to generate 100,213 image crops from hospital blood cultures and reported a slide‑level accuracy of 94.9%. Walter et al. (2024), working with the MetaSystems clinical platform on 1,555 prospectively collected patient slides, found that their CNN‑assisted workflow still produced a 5.3% error rate roughly one incorrect classification for every twenty slides and explicitly concluded that the system was not yet suitable for clinical deployment. McMahon et al. (2025) at Dartmouth evaluated a state‑of‑the‑art vision transformer (GramViT) on the same MHU dataset used in this project and achieved 89.8% accuracy without fine‑tuning, noting that even their best configuration did not surpass the 95.1% PoolFormer benchmark reported by Kim et al. (2023).

This project's ResNet18 model, trained on consumer hardware with an open-source dataset, achieved 99.53% accuracy, meeting the performance benchmark suggested by Walter et al (2024) for clinical adoption.

Machine learning models are sometimes considered to be a black box with poor explainability, but clinical tools need to be transparent about how the classification decision is achieved. To address this, Grad-CAM (Gradient-weighted Class Activation Mapping) was applied. Grad-CAM produces a heatmap showing exactly which pixels the model focused on when making its decision. The heat maps show the model concentrating on the bacteria themselves — not the slide background, not staining artifacts. This confirms that the model is reasoning correctly, not just getting lucky.

My model runs locally on a consumer Linux desktop but can also run on a laptop, and I made a Gradio webapp for this (Figure 10). Just by uploading a microscopy image, the system returns a Gram classification with a confidence score in seconds without need for internet or expensive equipment.

So What?

Automated Gram‑stain interpretation is becoming realistic, but published papers still struggle with staining variability, sparse bacteria, and inconsistent slide quality. Earlier CNN models usually reached about 95% accuracy and Walter et al. also pointed out that current tools are “not yet ready for clinical implementation” and suggested a 1–2% error‑rate as the target for clinical usefulness.

I started with a tiny 2‑layer CNN that reached only 59% accuracy, demonstrating how challenging Gram stain images are. After improving the architecture, cleaning the dataset, and building a stronger preprocessing pipeline, my next CNNs reached 95% accuracy. But the breakthrough came with transfer‑learning and careful data curation wherein the final ResNet model hit 99.53% accuracy, exceeding the clinical threshold suggested by Walter et al (Figure 11).

In contrast with Kim et al’s finding my transformer-based model training was unstable and with dropout (0.5) it worsened, while ResNet-18 stayed stable, and the resulting model is small enough to practically run on any consumer desktop/laptop.

What's Next?

The next phase of the project focuses on scaling and real‑world use. Expanding to larger open‑source or institutional datasets will make the model resilient to differences in staining, imaging devices, and sample types. With more data, transformer models may eventually push accuracy toward near‑perfect levels. I’ve also reached out to the Chan‑Zuckerberg Initiative to support access in underserved regions. Although the system already runs through a simple Gradio app, proper privacy safeguards are needed for pilot deployment. Building a mobile app would bring automated Gram stain interpretation directly to anyone with a phone (Figure 12).

Thanks

This project was made possible through open‑source datasets, online availability university‑level teaching materials, and the support of the people who encouraged my learning.

University of Heidelberg — public Gram‑stain microscopy dataset

CS231N (Stanford) — Dr. Fei‑Fei Li

CS50 (Harvard) — Prof. David Malan

Linux, PyTorch, TensorFlow — open‑source platforms used to build and train all models

Ms. Morris — academic support

My family — consistent encouragement

References

Ferrer, R., Martin-Loeches, I., Phillips, G., Osborn, T. M., Townsend, S., Dellinger, R. P.,

Artigas, A., Schorr, C., & Levy, M. M. (2014). Empiric antibiotic treatment reduces

mortality in severe sepsis and septic shock from the first hour: Results from a guideline-

based performance improvement program. Critical Care Medicine, 42(8), 1749–1755.

https://doi.org/10.1097/CCM.0000000000000330

K. He, X. Zhang, S. Ren and J. Sun, "Deep Residual Learning for Image Recognition," 2016

IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV,

USA, 2016, pp. 770-778, doi: 10.1109/CVPR.2016.90.

Karpathy, A., Johnson, J., & Fei‑Fei, L. (n.d.). CS231N: Convolutional neural networks for visual recognition. Stanford University. https://cs231n.stanford.edu/

Kim, H. E., Cosa-Linan, A., Santhanam, N., Jannesari, M., Maros, M. E., & Ganslandt, T.

(2022). Transfer learning for medical image classification: A literature review. BMC

Medical Imaging, 22(1), 69. https://doi.org/10.1186/s12880-022-00793-7

Kim, H. E., Maros, M. E., Miethke, T., Kittel, M., Siegel, F., & Ganslandt, T. (2023).

Lightweight visual transformers outperform convolutional neural networks for gram-

stained image classification: An empirical study. Biomedicines, 11(5), 1333.

https://doi.org/10.3390/biomedicines11051333

Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean,

Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016.

$TensorFlow$: a system for $Large-Scale$ machine learning. In 12th USENIX symposium

on operating systems design and implementation (OSDI 16). 265–283.

https://doi.org/10.5555/3026877.3026899

McMahon, J., Tomita, N., Tatishev, E. S., Workman, A. A., Costales, C. R., Banaei, N.,

Martin, I. W., & Hassanpour, S. (2025). A novel framework for the automated

characterization of Gram-stained blood culture slides using a large-scale vision

transformer. Journal of Clinical Microbiology, 63(3), e0151424.

https://doi.org/10.1128/jcm.01514-24

Rudd, K. E., Johnson, S. C., Agesa, K. M., Shackelford, K. A., Tsoi, D., Kievlan, D. R.,

Colombara, D. V., Ikuta, K. S., Kissoon, N., Finfer, S., Fleischmann-Struzek, C., Machado,

F. R., Reinhart, K. K., Rowan, K., Seymour, C. W., Watson, R. S., West, T. E., Marinho, F.,

Hay, S. I., … Naghavi, M. (2020). Global, regional, and national sepsis incidence and

mortality, 1990-2017: Analysis for the Global Burden of Disease Study. Lancet (London,

England), 395(10219), 200–211. https://doi.org/10.1016/S0140-6736(19)32989-7

Smith, K. P., Kang, A. D., & Kirby, J. E. (2018). Automated interpretation of blood culture

Gram stains by use of a deep convolutional neural network. Journal of Clinical

Microbiology, 56(3), e01521-17. https://doi.org/10.1128/JCM.01521-17

Tjandra, K. C., Ram-Mohan, N., Abe, R., Hashemi, M. M., Lee, J.-H., Chin, S. M., Roshardt,

M. A., Liao, J. C., Wong, P. K., & Yang, S. (2022). Diagnosis of bloodstream infections: An

evolution of technologies towards accurate and rapid identification and antibiotic

susceptibility testing. Antibiotics (Basel, Switzerland), 11(4), 511.

https://doi.org/10.3390/antibiotics11040511

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N. and

Polosukhin, I. (2017) Attention Is All You Need. Proceedings of the 31st International

Conference on Neural Information Processing Systems (NeurIPS 2017), Long Beach,

California, 4-9 December 2017, 5998-6008.

Walter, C., Weissert, C., Gizewski, E., Burckhardt, I., Mannsperger, H., Hänselmann, S.,

Busch, W., Zimmermann, S., & Nolte, O. (2024). Performance evaluation of machine-

assisted interpretation of Gram stains from positive blood cultures. Journal of Clinical

Microbiology, 62(4), e0087623. https://doi.org/10.1128/jcm.00876-23

Yu, B., & Malan, D. J. (n.d.). CS50’s introduction to artificial intelligence with Python. Harvard University. https://cs50.harvard.edu/ai/

Images (13)

Awards (2)

  • Gold Medal
  • Selected for CWSF 2026

Competition history

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: ProjectBoard / Youth Science Canada

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google