Computer Vision For Bacterial Gram Stain Classification
CWSF · 2026 Disease & Illness Gold Medal
Overview
Antibiotic resistance is one of the most urgent crises in global health, and the first step in fighting it is correctly identifying whether a bacterium is Gram-positive or Gram-negative — a distinction that directly determines which antibiotics will work. Manual microscopy interpretation is slow, subjective, and inaccessible in resource-limited settings. This project developed four progressively optimized convolutional neural networks (CNNs) trained on 8,000 labeled Gram-stained microscopy images. The final model hit 98.33% test accuracy, and I reproduced that result across two independent runs. These results put us ahead of the 94.9% benchmark reported by Smith et al. (Harvard Medical School, 2018) and the 95.1% benchmark reported by Hee Kim et al.(Heidelberg University, 2023). Grad-CAM visualization confirms the model learns genuine bacterial features, rather than background or other artifacts in the slide. This demonstrates that reliable, automated Gram stain classification is achievable without specialized equipment or proprietary data.
Video
This video could not be played here. Watch it on the original project page.
Why?
Gram staining is the most important step in identifying bacterial infections (McMahon 2025). It classifies bacteria into Gram‑positive and Gram‑negative groups (Figure 2), which determines the first antibiotic a patient receives. This matters because sepsis is a severe, life‑threatening condition caused by infection (Figure1) and causes over 11 million deaths every year worldwide (Rudd 2020).
Early, accurate treatment saves lives, and Gram stain interpretation is the very first clue clinicians rely on. But despite its importance, Gram stain interpretation is still manual. It depends on technician experience, staining quality, and time. In many settings, especially where trained microbiologists are limited, results can be inconsistent or delayed (Figure 3). Delay in starting antibiotic therapy increases the risk of death by 10% in sepsis patients (Ferrer 2014).
I want to be a physician, and I believe that future physicians need to understand AI and not just use it. Inspired by courses from Harvard and Stanford on computer science and deep learning for computer vision I wondered whether a model could learn the subtle differences in color, texture, and morphology that distinguish Gram‑positive from Gram‑negative bacteria.
I wanted to build something that works reliably, runs offline, and is accessible to any lab anywhere in the world. I built a complete deep‑learning pipeline trained on a dataset of 8000 labeled images using my own desktop. I also developed a Gradio app that can classify images directly without using expensive equipment.
How?
Background Research
I first came across online Harvard University's CS50 Introduction to Computer Science (Yu 2020) in the summer of grade 2 and Stanford's CS231N Deep Learning for Computer Vision (Karpathy 2015) in summer of grade 6 and I really liked both the courses (Figure 4). I learned more about automation of Gram staining by reading articles which were peer reviewed, published in the Journal of Clinical Microbiology and the authors had excellent institutional affiliation (Harvard, Dartmouth, Heidelberg).
Dataset & Data Quality
I used the University of Heidelberg (MHU) open-source dataset (Kim et al 2023) which has 8,500 labeled Gram stain microscopy images collected between 2015 and 2019 using a camera. It is the largest public dataset of its kind, and is 10 times bigger than the next largest open dataset (Figure 5). I removed around 200 images that were almost entirely black-and-white and therefore unsuitable for color-based classification (Figure 6). Before doing this step, my model accuracy was poor and I learned that it is very important to check data for such errors.
Model development
I started with a simple baseline 2-layer TensorFlow (Abadi et al 2014) based CNN, which achieved about 59% accuracy which is only slightly better than random guessing (Figure 7). By analyzing its errors, I redesigned the architecture and the second model reached around 92% accuracy, and then I added dropout and other complex machine learning techniques to make the model achieve 95% accuracy (Figure 8).
Materials: Python, PyTorch, NumPy, OpenCV, scikit-learn; NVIDIA RTX 3060 Ti GPU; GitHub Codespaces.
What?
The breakthrough came when I applied transfer learning using a pretrained ResNet (K He et al 2015) architecture. Instead of learning everything from scratch, the model could build on features learned from millions of images of ImageNet. (Figure 9)
Achieving 99.53% accuracy on 1,600 unseen test images, the finalized ResNet18 model surpassed every published benchmark for this dataset. (Kim H 2023, McMahon 2025) , and that was confirmed across two independent training runs.
Despite a decade of sustained effort, no automated Gram stain system has achieved regulatory approval or wide clinical implementation. Walter et al. (2024) concluded that their prospectively validated 1,555-slide CNN: "is not yet ready for clinical implementation” as “we suggest that an error rate of 1% to 2% might be a reasonable benchmark for the accuracy of manual Gram stain interpretations”.
Smith et al. (2018) at Harvard Medical School used the MetaFer clinical automated scanner to generate 100,213 image crops from hospital blood cultures and reported a slide‑level accuracy of 94.9%. Walter et al. (2024), working with the MetaSystems clinical platform on 1,555 prospectively collected patient slides, found that their CNN‑assisted workflow still produced a 5.3% error rate roughly one incorrect classification for every twenty slides and explicitly concluded that the system was not yet suitable for clinical deployment. McMahon et al. (2025) at Dartmouth evaluated a state‑of‑the‑art vision transformer (GramViT) on the same MHU dataset used in this project and achieved 89.8% accuracy without fine‑tuning, noting that even their best configuration did not surpass the 95.1% PoolFormer benchmark reported by Kim et al. (2023).
This project's ResNet18 model, trained on consumer hardware with an open-source dataset, achieved 99.53% accuracy, meeting the performance benchmark suggested by Walter et al (2024) for clinical adoption.
Machine learning models are sometimes considered to be a black box with poor explainability, but clinical tools need to be transparent about how the classification decision is achieved. To address this, Grad-CAM (Gradient-weighted Class Activation Mapping) was applied. Grad-CAM produces a heatmap showing exactly which pixels the model focused on when making its decision. The heat maps show the model concentrating on the bacteria themselves — not the slide background, not staining artifacts. This confirms that the model is reasoning correctly, not just getting lucky.
My model runs locally on a consumer Linux desktop but can also run on a laptop, and I made a Gradio webapp for this (Figure 10). Just by uploading a microscopy image, the system returns a Gram classification with a confidence score in seconds without need for internet or expensive equipment.
So What?
Automated Gram‑stain interpretation is becoming realistic, but published papers still struggle with staining variability, sparse bacteria, and inconsistent slide quality. Earlier CNN models usually reached about 95% accuracy and Walter et al. also pointed out that current tools are “not yet ready for clinical implementation” and suggested a 1–2% error‑rate as the target for clinical usefulness.
I started with a tiny 2‑layer CNN that reached only 59% accuracy, demonstrating how challenging Gram stain images are. After improving the architecture, cleaning the dataset, and building a stronger preprocessing pipeline, my next CNNs reached 95% accuracy. But the breakthrough came with transfer‑learning and careful data curation wherein the final ResNet model hit 99.53% accuracy, exceeding the clinical threshold suggested by Walter et al (Figure 11).
In contrast with Kim et al’s finding my transformer-based model training was unstable and with dropout (0.5) it worsened, while ResNet-18 stayed stable, and the resulting model is small enough to practically run on any consumer desktop/laptop.
What's Next?
The next phase of the project focuses on scaling and real‑world use. Expanding to larger open‑source or institutional datasets will make the model resilient to differences in staining, imaging devices, and sample types. With more data, transformer models may eventually push accuracy toward near‑perfect levels. I’ve also reached out to the Chan‑Zuckerberg Initiative to support access in underserved regions. Although the system already runs through a simple Gradio app, proper privacy safeguards are needed for pilot deployment. Building a mobile app would bring automated Gram stain interpretation directly to anyone with a phone (Figure 12).
Thanks
This project was made possible through open‑source datasets, online availability university‑level teaching materials, and the support of the people who encouraged my learning.
University of Heidelberg — public Gram‑stain microscopy dataset
CS231N (Stanford) — Dr. Fei‑Fei Li
CS50 (Harvard) — Prof. David Malan
Linux, PyTorch, TensorFlow — open‑source platforms used to build and train all models
Ms. Morris — academic support
My family — consistent encouragement
References
Ferrer, R., Martin-Loeches, I., Phillips, G., Osborn, T. M., Townsend, S., Dellinger, R. P.,
Artigas, A., Schorr, C., & Levy, M. M. (2014). Empiric antibiotic treatment reduces
mortality in severe sepsis and septic shock from the first hour: Results from a guideline-
based performance improvement program. Critical Care Medicine, 42(8), 1749–1755.
https://doi.org/10.1097/CCM.0000000000000330
K. He, X. Zhang, S. Ren and J. Sun, "Deep Residual Learning for Image Recognition," 2016
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV,
USA, 2016, pp. 770-778, doi: 10.1109/CVPR.2016.90.
Karpathy, A., Johnson, J., & Fei‑Fei, L. (n.d.). CS231N: Convolutional neural networks for visual recognition. Stanford University. https://cs231n.stanford.edu/
Kim, H. E., Cosa-Linan, A., Santhanam, N., Jannesari, M., Maros, M. E., & Ganslandt, T.
(2022). Transfer learning for medical image classification: A literature review. BMC
Medical Imaging, 22(1), 69. https://doi.org/10.1186/s12880-022-00793-7
Kim, H. E., Maros, M. E., Miethke, T., Kittel, M., Siegel, F., & Ganslandt, T. (2023).
Lightweight visual transformers outperform convolutional neural networks for gram-
stained image classification: An empirical study. Biomedicines, 11(5), 1333.
https://doi.org/10.3390/biomedicines11051333
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean,
Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016.
$TensorFlow$: a system for $Large-Scale$ machine learning. In 12th USENIX symposium
on operating systems design and implementation (OSDI 16). 265–283.
https://doi.org/10.5555/3026877.3026899
McMahon, J., Tomita, N., Tatishev, E. S., Workman, A. A., Costales, C. R., Banaei, N.,
Martin, I. W., & Hassanpour, S. (2025). A novel framework for the automated
characterization of Gram-stained blood culture slides using a large-scale vision
transformer. Journal of Clinical Microbiology, 63(3), e0151424.
https://doi.org/10.1128/jcm.01514-24
Rudd, K. E., Johnson, S. C., Agesa, K. M., Shackelford, K. A., Tsoi, D., Kievlan, D. R.,
Colombara, D. V., Ikuta, K. S., Kissoon, N., Finfer, S., Fleischmann-Struzek, C., Machado,
F. R., Reinhart, K. K., Rowan, K., Seymour, C. W., Watson, R. S., West, T. E., Marinho, F.,
Hay, S. I., … Naghavi, M. (2020). Global, regional, and national sepsis incidence and
mortality, 1990-2017: Analysis for the Global Burden of Disease Study. Lancet (London,
England), 395(10219), 200–211. https://doi.org/10.1016/S0140-6736(19)32989-7
Smith, K. P., Kang, A. D., & Kirby, J. E. (2018). Automated interpretation of blood culture
Gram stains by use of a deep convolutional neural network. Journal of Clinical
Microbiology, 56(3), e01521-17. https://doi.org/10.1128/JCM.01521-17
Tjandra, K. C., Ram-Mohan, N., Abe, R., Hashemi, M. M., Lee, J.-H., Chin, S. M., Roshardt,
M. A., Liao, J. C., Wong, P. K., & Yang, S. (2022). Diagnosis of bloodstream infections: An
evolution of technologies towards accurate and rapid identification and antibiotic
susceptibility testing. Antibiotics (Basel, Switzerland), 11(4), 511.
https://doi.org/10.3390/antibiotics11040511
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N. and
Polosukhin, I. (2017) Attention Is All You Need. Proceedings of the 31st International
Conference on Neural Information Processing Systems (NeurIPS 2017), Long Beach,
California, 4-9 December 2017, 5998-6008.
Walter, C., Weissert, C., Gizewski, E., Burckhardt, I., Mannsperger, H., Hänselmann, S.,
Busch, W., Zimmermann, S., & Nolte, O. (2024). Performance evaluation of machine-
assisted interpretation of Gram stains from positive blood cultures. Journal of Clinical
Microbiology, 62(4), e0087623. https://doi.org/10.1128/jcm.00876-23
Yu, B., & Malan, D. J. (n.d.). CS50’s introduction to artificial intelligence with Python. Harvard University. https://cs50.harvard.edu/ai/
Images (13)
Awards (2)
- Gold Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
CWSF · 2026
DEEP-GRAM: A Deep Learning Model for Gram Stain Species Prediction in Bloodstream Infections
ISEF · 2022
Automatic Classification of Peripheral Neutrophils on Digital Images Analyzed by Artificial Intelligence
ISEF · 2023
MicroScan: A Computer Vision Tool To Assist Malaria Microscopy Diagnosis
ISEF · 2025
Convolutional Neural Networks Applied to Hematological Analysis
ISEF · 2020
Improving Hazard Characterization in Bacterial Pathogens: Predicting Efficiency of Antibiotics Using Machine Learning
ISEF · 2026
Deep Learning-Based Oral Lesion Classification Leveraging Real-Time 3D Gradient-Weighted Class Activation Mapping (Grad-CAM) Imaging
ISEF · 2020
Hey Computer, Am I Sick? Finding Disease Patterns Using Computer Vision
ISEF · 2025
Colorectal Cancer Imaging and Classification - A Deep Learning Approach to Classify Histopathological Images
Closest projects by meaning, across every fair and year in the corpus.