Forecasting Blazing Wildfires with XGBoost ML Algorithm
CWSF · 2026 Digital Technology Silver Medal
Overview
I chose this project because wildfires are becoming a bigger problem in Canada. they can affect forests, wildlife, air quality, and nearby communities. I wanted to see if a machine learning could help people understand fire risk before a fire starts. To do this, I created a wildfire prediction program for Canada. It works by looking at real information such as past fires, weather, land and soil conditions, snow, roads, and nearby population, then using those patterns to predict how likely a fire is to happen, how large it could become, and how severe it might be. This project matters because it shows how computers can help us study complex natural problems and better prepare for wildfire danger.
Video
This video could not be played here. Watch it on the original project page.
Why?
I chose this project because wildfires are becoming one of the most serious environmental and public safety problems in Canada. They affect forests, wildlife, air quality, homes, and entire communities, yet people often only hear about them after they have already become emergencies. I wanted to connect computer science with a real-world problem and explore whether artificial intelligence could help us understand wildfire risk earlier. I was especially interested in this because wildfire danger depends on many factors working together, such as weather, terrain, fuel, snow, lightning, and human activity. My main question was: can a machine-learning system use real Canadian data to predict where wildfire conditions are more dangerous, how serious a fire might become, and display that risk in a way people can understand? I also wanted to build something practical, not just a model hidden in code, so I created an interactive map-based tool. This project could benefit researchers, educators, planners, and communities by improving wildfire awareness and helping people better understand the conditions that lead to fire risk. It is not a replacement for wildfire experts, but it can be a useful decision-support and education tool.
How?
I started by doing background research on what causes wildfires and how people currently measure wildfire danger. I used sources such as government wildfire records, climate data, maps, and weather information because I wanted my project to be based on real Canadian data. Next, I built a computer system that takes information from many sources, including past wildfire locations, land type, elevation, roads, population, snow, lightning, and weather. I turned all of this into a large set of numeral values that machine learning could learn from. My final dataset included 562,397 samples from 2005 to 2024.
Then I trained 2 models, one to predict how likely a wildfire is to start, and another to estimate how large it could become. I also designed the system so it could use live weather or manual weather inputs, which helped me test how the model reacted under different conditions. After that, I tested the system by comparing its predictions to real wildfire records and by trying different weather conditions at the same location, such as snowy, wet, or hot and dry conditions. I also checked whether the model gave similar results in areas with similar conditions and whether it changed when important variables changed and did a statistical study. I recorded the results in graphs and map displays so I could compare patterns clearly without sharing raw data. I then built an interactive map-based website so users could test one location, many locations from a file, or a whole grid of points.
What?
My main finding was that my system was very successful at predicting wildfire ignition risk, but exact final fire size was much harder to predict. The wildfire ignition model performed strongly, reaching about 98.5% accuracy, an ROC-AUC of about 0.998, and an F1-score of about 0.967. I used these statistics because they do more than show how many predictions were correct. They also show how well the model separates higher-risk wildfire conditions from lower-risk ones and how well it balances false alarms with missed wildfire cases. This matters because a wildfire tool should not only be accurate overall, but also useful when risk is high. These results suggest that real Canadian environmental and geographic data contain strong patterns that can help identify when wildfire conditions are more dangerous.
For fire size, my tiered system worked better at predicting general size range than exact hectares. It correctly placed fires into the right size category about 75.9% of the time, which shows that the model often understands whether a fire is likely to stay small or become much larger. However, exact hectare prediction remained much weaker, with an RMSE of about 19,876 hectares and an R² value close to 0.002. This suggests that predicting whether a fire may start is easier than predicting exactly how large it will become, because final fire size depends on many extra factors after ignition, such as changing weather, spread conditions, fuel continuity, and suppression response. Even though the exact size model is still a weakness, the size-bin system was useful because it gave a more realistic general estimate of fire severity than a single size number alone.
My prototype works by taking a location and collecting environmental information such as terrain, fuel type, fire history, weather, snow, lightning, roads, and population. The system then uses machine learning to predict wildfire probability, fire size, and severity, and shows the results on an interactive map. It can test a single point, a CSV list of coordinates, or a grid of points. This allowed me to move beyond only training a model and create a practical tool that users can interact with directly.
I also tested whether the system responded realistically to changing conditions. At the same location, a snowy and cold case gave a wildfire probability of about 0.010, a cool and wet case gave about 0.097, and a hot, dry, and windy case increased to about 0.470. This showed that the model was not simply giving one fixed answer for a location, but was reacting to weather and seasonal changes in a believable way. Overall, my results show that artificial intelligence can model wildfire ignition risk strongly when it uses real wildfire, terrain, climate, snow, fuel, and infrastructure data, while exact fire-size prediction remains the main challenge and the most important area for future improvement.
So What?
My results show that machine learning can be very useful for understanding wildfire risk when it is trained with real Canadian environmental and geographic data. The strongest conclusion from my project is that wildfire ignition risk can be predicted much more successfully than exact final fire size. This means the system is already valuable for identifying places and conditions where wildfire danger is higher, even though it is not yet as strong at predicting exactly how large a fire will become. I learned that wildfire occurrence and wildfire growth are related, but they are not equally easy to model. A fire may start under dangerous conditions, but its final size depends on many additional factors, such as changing weather, fuel continuity, spread, and suppression.
These results are important because they show that wildfire risk is not random. It can be connected to patterns in terrain, weather, fuel type, snow, lightning, roads, population, and fire history. This means machine learning can help turn complex environmental data into useful predictions that people can understand. My project could help researchers, educators, planners, and communities better explore wildfire danger and support earlier decision-making. It is not meant to replace wildfire experts, but it shows how artificial intelligence can be used as a decision-support and learning tool for a major real-world problem.
What's Next?
The next step for my project would be improving the fire-size side of the system, because exact fire size was harder to predict than wildfire ignition risk. I would like to add more spread-related information, such as better fuel and moisture details and weather changes after a fire starts. I would also test the system on more regions of Canada and more future wildfire seasons to see how well it works under changing conditions. In the future, I would like to make the forecasting stronger and continue developing the project into a better wildfire risk/decision-support tool.
10:18 PM
Thanks
I would like to thank everyone who supported me during this project. I am especially grateful to Dad for their guidance, encouragement, and feedback throughout the development of my system. In addition, I appreciate the organizations and open-data sources that made this project possible by providing access to wildfire, climate, terrain, and geographic datasets. I would also like to say thank you to Shell Canada for sponsoring my trip to CWSF
References
References
Canadian Forest Service. (2025). Canadian National Fire Database (CNFDB): Agency provided fire locations (point data) [Data set]. Natural Resources Canada, Canadian Wildland Fire Information System. https://cwfis.cfs.nrcan.gc.ca/index.php/datamart/metadata/nfdbpnt
Canadian Forest Service. (2025). National Burned Area Composite (NBAC), 1972-2024 [Data set]. Natural Resources Canada, Canadian Wildland Fire Information System. https://cwfis.cfs.nrcan.gc.ca/index.php/datamart/metadata/nbac
Canadian Forest Service. (2022). Canadian Wildland Fire Information System (CWFIS) [Data set]. Natural Resources Canada. https://cwfis.cfs.nrcan.gc.ca/index.php/datamart/metadata/fbp
Natural Resources Canada. (2022). 2020 Land Cover of Canada [Data set]. Government of Canada Open Maps. https://search.open.canada.ca/openmap/ee1580ab-a23d-4f86-a09b-79763677eb47
Bondarenko, M., Priyatikanto, R., Tejedor-Garavito, N., Zhang, W., McKeen, T., Cunningham, A., Woods, T., Hilton, J., Cihan, D., Nosatiuk, B., Brinkhoff, T., Tatem, A., & Sorichetta, A. (2025). Constrained estimates of 2015-2030 total number of people per grid square at a resolution of 3 arc (approximately 100 m at the equator) R2025A version v1 [Data set]. WorldPop, School of Geography and Environmental Science, University of Southampton. https://doi.org/10.5258/SOTON/WP00839
NASA/METI/AIST/Japan Spacesystems, & U.S./Japan ASTER Science Team. (2018). ASTER Global Digital Elevation Model V003 [Data set]. NASA EOSDIS Land Processes Distributed Active Archive Center. https://doi.org/10.5067/ASTER/ASTGTM.003
Abatzoglou, J. T., Dobrowski, S. Z., Parks, S. A., & Hegewisch, K. C. (2018). TerraClimate, a high-resolution global dataset of monthly climate and climatic water balance from 1958-2015. Scientific Data, 5, Article 170191. https://doi.org/10.1038/sdata.2017.191
Geofabrik GmbH, & OpenStreetMap contributors. (2026). Canada provincial shapefile extracts [Data set]. Geofabrik Download Server. https://download.geofabrik.de/north-america/canada.html
LANCEMODIS. (2021). MODIS/Aqua Terra thermal anomalies/fire locations 1 km FIRMS NRT (vector data) [Data set]. MODAPS at NASA/GSFC: The Land, Atmosphere Near real-time Capability for EOS (LANCE). https://www.earthdata.nasa.gov/es/data/catalog/lancemodis-mcd14dl-6.1nrt
Open-Meteo. (n.d.). Weather forecast API documentation. Retrieved March 24, 2026, from https://open-meteo.com/en/docs
FastAPI developers. (n.d.). FastAPI (Version 0.135.1) [Computer software]. https://fastapi.tiangolo.com/
GeoPandas developers. (n.d.). GeoPandas (Version 1.1.2) [Computer software]. https://geopandas.org/
Joblib developers. (n.d.). Joblib (Version 1.5.3) [Computer software]. https://joblib.readthedocs.io/
Leaflet contributors. (n.d.). Leaflet [JavaScript library]. https://leafletjs.com/
Matplotlib developers. (n.d.). Matplotlib (Version 3.10.8) [Computer software]. https://matplotlib.org/
NumPy developers. (n.d.). NumPy (Version 2.4.2) [Computer software]. https://numpy.org/
pandas development team. (n.d.). pandas (Version 3.0.1) [Computer software]. https://pandas.pydata.org/
Pallets. (n.d.). Jinja (Version 3.1.6) [Computer software]. https://jinja.palletsprojects.com/
pyogrio developers. (n.d.). pyogrio (Version 0.12.1) [Computer software]. https://pyogrio.readthedocs.io/
pyproj developers. (n.d.). pyproj (Version 3.7.2) [Computer software]. https://pyproj4.github.io/pyproj/
Python Software Foundation. (n.d.). Python (Version 3.14.3) [Computer software]. https://www.python.org/
Rasterio developers. (n.d.). Rasterio (Version 1.5.0) [Computer software]. https://rasterio.readthedocs.io/
scikit-learn developers. (n.d.). scikit-learn (Version 1.8.0) [Computer software]. https://scikit-learn.org/
SciPy developers. (n.d.). SciPy (Version 1.17.1) [Computer software]. https://scipy.org/
Shapely developers. (n.d.). Shapely (Version 2.1.2) [Computer software]. https://shapely.readthedocs.io/
Uvicorn developers. (n.d.). Uvicorn (Version 0.42.0) [Computer software]. https://uvicorn.dev/
xarray developers. (n.d.). xarray (Version 2026.2.0) [Computer software]. https://docs.xarray.dev/
XGBoost contributors. (n.d.). XGBoost (Version 3.2.0) [Computer software]. https://xgboost.readthedocs.io/en/stable/
Agafonkin, V. (n.d.). Leaflet.heat [JavaScript library]. https://github.com/Leaflet/Leaflet.heat
Images (19)
Awards (2)
- Silver Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
CYSF · 2026
Smarter than a SPARK...Machine Learning and Wildfire Prediction
ISEF · 2023
Predicting Large Wildfires Using Machine Learning Approach Towards Environmental Justice via Remote Sensing
JSHS · 2023
ILLINOIS-CHICAGO Predicting Large Wildfires Using Machine Learning Approach towards Environmental Justice via Remote Sensing
ISEF · 2026
A Machine Learning Model for Predicting Fire Probability in Global Peatlands
ISEF · 2022
It's Flaming Out: Using Artificial Intelligence To Emulate Critical Aspects of Wildfire Growth
JSHS · 2024
Predicting Next-Day Wildfire Spread with Environmental Data and Machine Learning
ISEF · 2024
A Comparative Analysis of Machine Learning Models for Wildfire Prediction
ISEF · 2021
An Accessible, Low Cost Tool for Citizen Scientists: Using Remote Sensing Techniques to Predict Fire Damage Propensity
Closest projects by meaning, across every fair and year in the corpus.