Regeneration-Likeness Scoring of Human Wound-Healing Genes Using Machine Learning
CWSF · 2026 Digital Technology Bronze Medal
Overview
This project created machine learning models trained from axolotl regeneration data to identify which human genes that show patterns similar to those involved in axolotl regeneration. Two machine learning models were created, with the scores combined to create an overall ‘regeneration-likeness’ score for every gene in the input dataset. The input dataset was a human scarring dataset, in order to provide greater contrast between fibrotic (scarring) and more ‘regeneration-like’ genes, highlighting which genes show the strongest conserved signals/patterns and play the most vital roles in this context. Several of the top 10 highest scoring genes (most regeneration-like) were uncharacterized, meaning that there is very limited information currently known—suggesting that they may play more vital roles in healing biology or may be potential candidates for study in regenerative medicine.
Video
This video could not be played here. Watch it on the original project page.
Why?
PROBLEM
Axolotls have the remarkable ability to flawlessly regenerate almost any part of their body without scarring. Humans, however, do not, and have limited regenerative abilities. The main problems this project addresses is that there is no current or widely acknowledged metric for ‘regeneration-likeness’ to measure a genes similarity to patterns observed in regenerative species such as axolotl, the lack of a computational method to detect ‘regeneration-like’ gene expression patterns, and comparative cross-species gene expression problems such as different species, tissues, time scales, baseline biology, and gene regulatory networks which persist while there is no standard method to compare patterns directly.
This project addresses these problems by creating a machine-learning scoring system that detects pattern similarity from RNA-seq data, as opposed to one-to-one gene matching. It deals with biological complexities through filtering and the use of two complementary models (Logistic Regression + Random Forest) and identifies human wound-healing genes whose expression patterns resemble regeneration-associated signatures.
How?
METHODS & MATERIALS
Datasets: GSE16777 (axolotl limb regeneration RNA-seq), GSE113619 (human keloid vs. normal scarring RNA-seq)
Software/ Technology: R coding environment, R packages such as Pheatmap, DESeq2, ggplot2, etc,...
Data Processing:
Normalization and filtering of axolotl and human datasets
Removal of low-expression genes that could become noise
Standardization to ensure model learns biological patterns and not noise
Orthology Mapping (BLASTN):
Axolotl transcripts matched to human equivalents
Best hit orthologs chosen
Ensures that the model compares biologically equivalent genes/orthologs across species
Model Training:
Models were trained on the axolotl regeneration data
Models learned patterns of regeneration, and not species-specific sequences
Application to human dataset:
Each human gene receives a score based on similarity to regeneration pattern
Genes are ranked from highest to lowest based on score
Downstream Analysis:
Score distribution analysis
Ranking of human genes
GO enrichment of top-scoring genes
Biological interpretation of conserved pathways
PROCEDURE
The overall procedure can be broken up and simplified into 7 core steps:
1.Setup
Finding axolotl dataset
Cleaning metadata
Setting up R and organizing files
2.DE Analysis/Normalization
Normalizing, cleaning & running DESeq2 on axolotl dataset to filter and remove noise
3.Ortholog Mapping
Identifying orthologous genes across species using BLASTN
Creating functional annotation for training data based off gene functions
4.Training the Models
Coding, training, and debugging models
Models learn the statistical patterns for axolotl regeneration (not species-specific sequences)
5.Applying the Models to the Human Dataset
Models score every human gene from the human dataset based on how strongly it resembles axolotl regeneration patterns
6.Model Output
The scores of both models are combined
Outputs the combined scores, ranked from highest to lowest
7.Interpretation
Identifying conserved biological processes shared between axolotl regeneration and humans (GO:BP)
Overall analysis and interpretation
***Note: This is a simplified version for ease of understanding
What?
RESULTS
A machine-learning model scoring system for regeneration-likeness was successfully developed.
Both Logistic Regression & Random Forest models were trained using axolotl regeneration gene expression as the reference pattern and applied to human wound-healing genes. The models produced regeneration-likeness scores with a possible range of 0-1 for each gene of the input dataset.
The regeneration-likeness score was extremely skewed.
Most human wound-healing genes scored very low, which indicated biological realism and accuracy since the input dataset was scarring-related; the opposite of regeneration. However, a distinct higher scoring tail emerged, showing that a small subset of genes displayed stronger regeneration-like expression patterns and behavior.
A focused set of top-scoring genes was identified.
These genes outputted the highest scores (~0.30-0.36), meaning they were some of the highest-scoring across both models and suggesting that they may represent conserved repair of regeneration-related signals. This list provides a narrowed set of candidates for future biological investigation and potential candidates for study in regenerative medicine.
Pattern-based comparison outperformed direct gene matching.
The models successfully detected cross-species expression similarities and patterns despite differences in species, tissues, fundamental biological processes, and regulatory networks. This demonstrates that pattern-based machine learning can overcome traditional comparative biology limitations.
Several of the top-scoring genes were uncharacterized loci.
The most interesting finding was that several of the top-scoring human genes were uncharacterized, meaning there was limited to no information known about their functions and/or locations. Because the model evaluates expression-pattern similarity instead of known gene functions, it can identify regeneration-like patterns in genes that have not yet been biologically studied in-depth. The presence of these genes in the high-scoring tail suggests that regeneration-associated expression signatures may exist in understudied regions of the human genome, providing new candidates for future research, as well as underscoring the potential of this method to reveal hidden or novel biology.
CONCLUSION
This project demonstrates that regeneration-like gene expression patterns can be computationally detected in human wound healing using a pattern-based machine-learning approach. By training models on axolotl regeneration RNA-seq data and applying it to human keloid versus normal scar expression data, the models computed regeneration-likeness scores that quantify how closely each human gene’s expression behavior resembles regeneration-associated patterns. The score distribution was extremely skewed: most human genes scored very low, but a distinct high-scoring subset emerged, indicating regeneration-like signatures. Notably, several of the top-scoring genes were uncharacterized or poorly annotated. Because the model evaluates expression-pattern similarity rather than relying on known gene functions, it can identify patterns and genes that have not yet been biologically studied. This suggests that understudied regions of the human genome may contain repair-related or potentially conserved signals that traditional gene-to-gene comparison methods may overlook.
Overall, this project provides a quantitative framework for assessing regeneration-likeness and demonstrates that machine-learning can uncover subtle cross-species expression similarities. While functional conservation cannot be confirmed without biological validation, the identified genes represent promising potential candidates for future research into regeneration, fibrosis, and tissue repair.
So What?
IMPLICATIONS
Regeneration-like patterns exist in humans, but they are subtle and/or not strong enough to enable true regeneration.
Some regenerative signals may be evolutionarily conserved.
Uncharacterized high-scoring genes may represent or be related to overlooked regulators of tissue repair, and may be potential candidates for regenerative medicine.
Computational approaches may help identify targets for improving human healing or reducing scar formation.
Machine learning can detect cross-species similarities and patterns that traditional methods cannot.
This project demonstrates that computational tools can reveal biological patterns and may help guide future research into improving human wound healing. While this project does not prove functional conservation or clinical diagnostic, it highlights promising genes and pathways for further investigation.
LIMITATIONS
In the event of the use of another dataset, the more evolutionary distant a species is from both axolotls and humans, the less meaningful the results may be. While the model would still run, the results may be less reliable.
The models' output(s) reflect whether a genes’ expression patterns are associated with regenerative healing processes. Although higher model scores may imply positive healing outcomes and lower scores are associated with fibrotic pathways, these scores do not predict clinical healing outcomes.
In the event of the use of another dataset/ if used for another purpose than those outlined specifically in the project, varying changes would need to be made to the code to ensure proper format and metadata alignment in addition to whatever changes would be required to fit the alternate purpose or goal intended.
What's Next?
NEXT STEPS
One of the main goals for the future of this project is to make it more accessible and generalizable before making it open source and available via GitHub. Other possible next steps are creating and adding more training data, testing more input datasets, and giving it an interactable interface instead of just code. Ideally, I would like to continue the research done in this project by improving the code/basis of the project and by potentially validating or building upon the results.
Thanks
ACKNOWLEDGEMENTS & THANKS
Patrick Campeau
Mr. Campeau was my Computer Science teacher, one of the people who helped me organize my thoughts in the earliest stages of outlining my project. Thank you for everything you taught me in semester 1 during computer science class, and for your advice!
Dr. Courtin
A knowledgeable person with science communication and public speaking experience who helped me improve the delivery of my project in terms of presentation. Thank you for your advice!
Mathiew Dykstra
Mr. Dykstra has been a very supportive and helpful figure in helping frame, simplify, and idetermine how to present my project in an accessible, clear, and coherent way to present it effectively without losing the audience. Thank you very much for all the time you spent helping me!
Dr. Moise
A Regional Fair judge who gave me excellent feedback on how to improve my project. Thank you very much!
References
Home - GEO - NCBI. (2019). Nih.gov. https://www.ncbi.nlm.nih.gov/geo/
GEO Accession viewer. (2024). Nih.gov. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE116777
GEO Accession viewer. (n.d.). Www.ncbi.nlm.nih.gov. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi
NCBI. (2006). Nucleotide BLAST. Nih.gov. https://blast.ncbi.nlm.nih.gov/Blast.cgi?PROGRAM=blastn&PAGE_TYPE=BlastSearch&LINK_LOC=blasthome
UniProt. (2023). Uniprot. https://www.uniprot.org/
GeeksforGeeks. (2024, February 22). Random Forest Algorithm in Machine Learning. GeeksforGeeks. https://www.geeksforgeeks.org/machine-learning/random-forest-algorithm-in-machine-learning/
(2026). Geeksforgeeks.org. https://media.geeksforgeeks.org/wp-content/uploads/20251216121929631349/random_forest_algorithm.webp
Linear vs Logistic Regression: How to Choose the Right Regression Model for Your Data. (2024, May 28). FreeCodeCamp.org. https://www.freecodecamp.org/news/linear-regression-vs-logistic-regression/
(2026). Andymath.com. https://andymath.com/wp-content/uploads/2019/08/Logistic-Function.jpg
Zach. (2020, October 28). How to Perform Logistic Regression in R (Step-by-Step). Statology. https://www.statology.org/logistic-regression-in-r/
Kurui, M. (2025, September 6). Logistic Regression in R: Your Complete GLM Tutorial - codepointtech.com. Codepointtech.com. https://codepointtech.com/master-logistic-regression-in-r-your-complete-glm-tutorial/
Zach. (2020, November 24). How to Build Random Forests in R (Step-by-Step). Statology. https://www.statology.org/random-forest-in-r/
Random Forests in R. (2026). Rguides.dev. https://rguides.dev/tutorials/r-random-forest/
Bhalla, D. (2014). A complete guide to Random Forest in R. ListenData. https://www.listendata.com/2014/11/random-forest-with-r.html
R Core Team. (2025). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. https://www.r-project.org/
CRAN Packages By Name. (2017). R-Project.org. https://cran.r-project.org/web/packages/available_packages_by_name.html
Rodrigues, M., Kosaric, N., Bonham, C. A., & Gurtner, G. C. (2018). Wound Healing: A Cellular Perspective. Physiological Reviews, 99(1), 665–706. https://doi.org/10.1152/physrev.00067.2017
Libbrecht, M. W., & Noble, W. S. (2015). Machine learning applications in genetics and genomics. Nature Reviews Genetics, 16(6), 321–332. https://doi.org/10.1038/nrg3920
StatQuest with Josh Starmer. (n.d.). StatQuest: edgeR, part 1, Library Normalization [Video]. YouTube. https://www.youtube.com/watch?v=Wdt6jdi-NQo
StatQuest with Josh Starmer. (2018, April 2). StatQuest: Principal Component Analysis (PCA), Step-by-Step [Video]. YouTube. https://www.youtube.com/watch?v=FgakZw6K1QQ
StatQuest with Josh Starmer. (2016, January 6). Drawing and interpreting heatmaps [Video]. YouTube. https://www.youtube.com/watch?v=oMtDyOn2TCc
MIT OpenCourseWare. (2015, January 20). 8. RNA-sequence analysis: expression, isoforms [Video]. YouTube. https://www.youtube.com/watch?v=MniYgsZSp30
Rafael Irizarry. (2012, June 6). Statistics for Genomics: Introduction to RNASEQ [Video]. YouTube. https://www.youtube.com/watch?v=C8RNvWu7pAw
Capcut. (n.d.). CapCut | All-In-One Video Editing Software. Www.capcut.com. https://capcut.com
Images (18)
Awards (2)
- Bronze Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
CYSF · 2025
Regenerative Mechanisms: Decoding Axolotl Regeneration for Human Healing
ISEF · 2022
Enhancing Human Embryonic Kidney Cell Regeneration Through Transducing Exogenous Ambystoma Mexicanum Pax-7 DNA
ISEF · 2025
WoundView: A Novel Comprehensive Tool Utilizing Machine Learning Models for Remote, Cost-Effective, Real-Time Wound Risk Assessment
ISEF · 2026
Glucotoxicity in Regeneration: Modeling Hyperglycemia-Induced Repair Failure in Planaria
ISEF · 2025
Characterization of Overexpressed Neural-Related Genes Based on Single-Cell mRNA Sequencing Data in Regenerated Intestines of Sea Cucumbers Holothuria glaberrima
ISEF · 2018
Proteomic Evolution in Hair Cell Regeneration
ISEF · 2018
Development of Semi-Supervised Machine Learning Models to Predict Enhancer Regions in Polygenic Developmental Diseases
ISEF · 2017
Developing Novel Gene Candidates (MEF2A, LTA, LGALS2, ALOX5AP, and PDE4D), through an Adaptive Genetic Algorithm, Support Vector Cluster, and Dynamic Bayesian Networks, to Analyze in a Learning Classifier System for a Highly Propitious CRISPR Therapy for Ischemic Heart Disease
Closest projects by meaning, across every fair and year in the corpus.