Refining the Forecast: Advanced Data Preprocessing for Accurate Wind Prediction
Overview
Accurate wind forecasting is challenging due to wind patterns being affected by weather, geography and seasons. Traditional methods of predicting wind are often incapable of capturing all the patterns and so are unable to effectively predict future behaviour. AI-driven wind forecasting provides a solution, since machine learning models perform well in scenarios with large amounts of past data and the need for accurate predictions of the future. This project focused on the data preprocessing step, when the initial dataset, containing past measurements such as atmospheric pressure or precipitation amount, is transformed in various ways to make it easier for the model to make accurate predictions. Preprocessing has major effects on the final predictions and accuracy, and so carries immense potential for improving wind forecasting. Such predictions are used, for example, in selecting sites for building wind farms and when scheduling maintenance of the existing ones.
Video
This video could not be played here. Watch it on the original project page.
Why?
In the past five years, the interest towards employing AI models in wind predictions has been rapidly increasing. This push has explored machine learning’s applicability to the field of wind predictions and generally found their superiority over traditional physical and statistical methods of prediction (Yang et al., 2024a) . The choice of AI for the task is justified by the complexity of patterns in wind data.
Applications of the method in real-world wind predictions for site selection and operation of wind farms are still limited due to their novelty. Currently, wind predictions in Canada still rely on statistical methods, which produce average values for wind speed in the future, which limits their reliability, especially in the case of short-term predictions, which are essential for operating a wind farm. Wind farms, being a long-term investment, depend totally on the initial site selection process, which basically determines power output for the wind farm for the decades to come. An increase in accuracy and reliability of wind speed data would allow for more efficient choice of locations, further promoting wind energy as an alternative to traditional power sources.
AI techniques must be adapted for the task through fine-tuning of the multitude of its components, including the crucial step for any machine learning model: data preprocessing. This increase in understanding and efficiency could allow for more usage of the technique in real-world situations, potentially increasing energy production of wind farms around the world and ultimately accelerating the transition to renewable sources of energy.
How?
This research investigated how different data preprocessing techniques affect machine learning model accuracy. By controlling the preprocessing method as the independent variable and minimizing other confounding factors, the study isolated its impact on model performance, measured through standard accuracy metrics.
The study used a public meteorological dataset from Sydney, Nova Scotia, obtained from Environment and Climate Change Canada. The dataset covered January 2017, to January 2026, and contained 80,304 rows with multiple weather-related features, such as pressure, humidity, and wind direction. After initial feature engineering, including cyclical encoding of time-based variables and removal of unnecessary columns, the dataset was reduced to 14 features.
Data preprocessing was treated as a two-step process: fitting and transforming. Preprocessing algorithms were fitted only on training data to prevent information leakage, while transformations were applied to both training and test datasets to ensure consistency. The data was split chronologically into 85% training data and 15% test data without shuffling, preserving temporal structure.
A total of eleven machine learning models and four preprocessing techniques were evaluated. Feature engineering choices depended on model families, with polynomial features applied to linear models and discretization used for others. Multiple preprocessing combinations were generated programmatically using nested loops.
For each model–preprocessing combination, hyperparameters were optimized using Optuna with 40 trials/configuration. Each trial included 6‑fold cross-validation to reduce overfitting and bias. Cross-validation accuracy metrics were recorded for further statistical analysis. After identifying optimal hyperparameters, models were retrained on the full training dataset and evaluated on the held-out test set. Final outputs included wind speed predictions and four performance metrics: mean squared error (MSE), mean absolute error (MAE), root mean squared error (RMSE), and R².
All experiments were implemented in Python using libraries such as Scikit-learn, Optuna, and CatBoost, within Google Colab, enabling efficient and reliable large-scale experimentation.
What?
Among the 504 combinations, 94.05% performed better than the RMSE baseline, and only 5.95% performed worse than the baseline. To investigate the effects of each of the preprocessing families, a heatmap of correlation between a preprocessing being applied and MSE was created and is presented in Figure 9. From the heatmap, it can be observed that Time Lags had the highest, almost complete, negative correlation with RMSE, indicating that their usage improved predictions in almost all cases. Feature Engineering and Scaling displayed negative values of lesser magnitude, indicating a less pronounced improvement of the predictions.
The data also showed that among the top 20% (200) combinations, most used models from Linear family, without preference for a specific one, which can also be seen in Figure 10. The figure clearly demonstrates that most Linear models performed better, than the other families. The clear separation into two parts of each of the parts of the graph is due to the immense impact of Time Lags on the final performance: the bottom segment, with lower error, is the one with Time Lags included. Thus, Time Lags were not merely strongly correlated with reduced error but also resulted in large changes in error when applied. This is shown in more detail in Figure 11, which demonstrates that time lags caused a major increase in accuracy.
Additionally, since time lags were split into categories that sometimes included one time lag multiple times, data regarding the interplay of various time lags was collected. Figure 12 summarizes those findings. Most importantly, the lowest RMSE was achieved with the combination that included only Simple Short Lags. Generally, in setups with multiple time lags combined, the cumulative effect of adding more complex lags (seasonal and rolling) is low.
So What?
The findings demonstrate the impact of preprocessing techniques on machine learning prediction accuracy for wind forecasting. Overall, most preprocessing combinations improved model accuracy compared to baseline results, supporting the established consensus that preprocessing generally enhances predictive performance.
Among all techniques evaluated, Time Lags emerged as the most influential preprocessing method. Results showed a strong negative correlation between Time Lag inclusion and mean squared error, indicating consistent performance gains across nearly all models and preprocessing combinations. Visual analyses confirmed that Time Lags were not only associated with higher accuracy but also produced substantially lower error magnitudes. However, the data also revealed diminishing returns from increasingly complex lag structures. Simple short‑term lags delivered the largest benefits, while additional or more complex lag schemes provided little improvement and occasionally reduced accuracy. This suggests that wind speed is primarily dependent on recent observations and that short‑term lagging is sufficient for datasets of similar size.
Scaling techniques showed mixed effectiveness. Their impact was most notable for Linear and KNN models, which are sensitive to feature scaling, while Tree‑based models remained largely unaffected. Common methods such as StandardScaler and MinMaxScaler produced minimal benefits for most models, whereas more aggressive transformations, including PowerTransformer and QuantileTransformer, often degraded performance.
In conclusion, preprocessing plays a critical role in ML‑driven wind forecasting. Time Lags are essential for achieving high accuracy, while other preprocessing methods offer limited or model‑specific benefits and should be applied selectively.
What's Next?
The sheer number of existing preprocessing techniques and machine learning models far exceeds the small subset tested in this research. Most noticeably, no deep learning models were examined due to the time for training. Computing times for deep learning models differ by at least an order of magnitude from the tested models which would result in unsurmountable data collection times. This problem, of course, could be solved by using more powerful servers that are optimized for specifically training such “heavy” models. This would require additional inquiry to fully optimize the workflow.
References
Ahmad, A., Xiao, X., Mo, H., & Dong, D. (2024d). Tuning data preprocessing techniques for improved wind speed prediction. Energy Reports, 11, 287–303. https://doi.org/10.1016/j.egyr.2023.11.056
Environment and Climate Change Canada, (2026). Hourly Data Report, Sydney A, Nova Scotia [Data set]. Government of Canada. https://climate.weather.gc.ca/climate_data/hourly_data_e.html?hlyRange=2014-08-05%7C2026-04-26&dlyRange=2014-08-07%7C2026-04-26&mlyRange=%7C&climate_id=8205701&Prov=NS&urlExtension=_e.html&searchType=stnName&optLimit=yearRange&StartYear=1840&EndYear=2026&selRowPerPage=25&Line=2&searchMethod=contains&Month=4&Day=26&txtStationName=sydney&timeframe=1&Year=2026
Demir, G., Riaz, M., & Deveci, M. (2024). Wind farm site selection using geographic information system and fuzzy decision making model. Expert Systems with Applications, 255. https://doi.org/10.1016/j.eswa.2024.124772
Ebtehaj, I., Bonakdari, H., Zeynoddin, M., Gharabaghi, B., & Azari, A. (2020). Evaluation of preprocessing techniques for improving the accuracy of stochastic rainfall forecast models. International Journal of Environmental Science and Technology, 17(1), 505–524. https://doi.org/10.1007/s13762-019-02361-z
Eroglu, O., Aktas Potur, E., Kabak, M., & Gencer, C. (2023). A Literature Review: Wind Energy Within The Scope of MCDM Methods. In Gazi University Journal of Science (Vol. 36, Number 4, pp. 1578–1599). Gazi Universitesi. https://doi.org/10.35378/gujs.1090337
Gaikwad, N. H., Mitra, A., Payal, J., Das, S., & Keshri, R. kumar. (2025a). Impact of Data Preprocessing on Wind Forecasting Using Machine Learning Techniques. 2025 4th International Conference on Power, Control and Computing Technologies, ICPC2T 2025, 771–776. https://doi.org/10.1109/ICPC2T63847.2025.10958701
Golazad, S. Z., Mohammadi, A., Rashidi, A., & Ilbeigi, M. (2024). From raw to refined: Data preprocessing for construction machine learning (ML), deep learning (DL), and reinforcement learning (RL) models. In Automation in Construction (Vol. 168). Elsevier B.V. https://doi.org/10.1016/j.autcon.2024.105844
Gonzalez Zelaya, C. V. (2019). Towards explaining the effects of data preprocessing on machine learning. Proceedings - International Conference on Data Engineering, 2019-April, 2086–2090. https://doi.org/10.1109/ICDE.2019.00245
Hong, S., McMorland, J., Zhang, H., Collu, M., & Halse, K. H. (2024). Floating offshore wind farm installation, challenges and opportunities: A comprehensive survey. In Ocean Engineering (Vol. 304). Elsevier Ltd. https://doi.org/10.1016/j.oceaneng.2024.117793
Khalaf, O. F., Uçan, O. N., & Alsamarai, N. A. (2025). Wind farm sites selection using a machine learning approach and geographical information systems in Türkiye. Discover Computing, 28(1). https://doi.org/10.1007/s10791-025-09511-7
Maharana, K., Mondal, S., & Nemade, B. (2022). A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, 3(1), 91–99. https://doi.org/10.1016/j.gltp.2022.04.020
Mahmud Sujon, K., Binti Hassan, R., Tusnia Towshi, Z., Othman, M. A., Abdus Samad, M., & Choi, K. (2024a). When to Use Standardization and Normalization: Empirical Evidence from Machine Learning Models and XAI. IEEE Access, 12, 135300–135314. https://doi.org/10.1109/ACCESS.2024.3462434
Mollick, T., Hashmi, G., & Sabuj, S. R. (2024a). Wind speed prediction for site selection and reliable operation of wind power plants in coastal regions using machine learning algorithm variants. Sustainable Energy Research, 11(1). https://doi.org/10.1186/s40807-024-00098-z
D. U. Ozsahin, M. Taiwo Mustapha, A. S. Mubarak, Z. Said Ameen and B. Uzun, "Impact of feature scaling on machine learning models for the diagnosis of diabetes," 2022 International Conference on Artificial Intelligence in Everything (AIE), Lefkosa, Cyprus, 2022, pp. 87-94, doi: 10.1109/AIE57029.2022.00024.
Qian, Z., Pei, Y., Zareipour, H., & Chen, N. (2019). A review and discussion of decomposition-based hybrid models for wind energy forecasting applications. In Applied Energy (Vol. 235, pp. 939–953). Elsevier Ltd. https://doi.org/10.1016/j.apenergy.2018.10.080
Sari, F., & Yalcin, M. (2024). Investigation of the importance of criteria in potential wind farm sites via machine learning algorithms. Journal of Cleaner Production, 435. https://doi.org/10.1016/j.jclepro.2024.140575
Sun, Y., Li, Y., Wang, R., & Ma, R. (2024). Modelling potential land suitability of large-scale wind energy development using explainable machine learning techniques: Applications for China, USA and EU. Energy Conversion and Management, 302. https://doi.org/10.1016/j.enconman.2024.118131
Wang, J., Niu, X., Zhang, L., Liu, Z., & Huang, X. (2024). A wind speed forecasting system for the construction of a smart grid with two-stage data processing based on improved ELM and deep learning strategies. Expert Systems with Applications, 241. https://doi.org/10.1016/j.eswa.2023.122487
Yaman, A. (2024). A GIS-based multi-criteria decision-making approach (GIS-MCDM) for determination of the most appropriate site selection of onshore wind farm in Adana, Turkey. Clean Technologies and Environmental Policy, 26(12), 4231–4254. https://doi.org/10.1007/s10098-024-02866-3
Yan, C. (2025). A review on spectral data preprocessing techniques for machine learning and quantitative analysis. In iScience (Vol. 28, Number 7). Elsevier Inc. https://doi.org/10.1016/j.isci.2025.112759
Yang, Y., Lou, H., Wu, J., Zhang, S., & Gao, S. (2024a). A survey on wind power forecasting with machine learning approaches. In Neural Computing and Applications (Vol. 36, Number 21, pp. 12753–12773). Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/s00521-024-09923-4
Images (31)
Awards (1)
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
ISEF · 2019
Enhancing Wind Power Predictions by Using Weather Data and Improving LSTMs
ISEF · 2017
Looking into the Past for Insight on the Future: Predictive Analytics and Machine Learning for Time Series Data
ISEF · 2023
Short Range Hourly Temperature Forecasting in Relation to National Weather Prediction Models: A Breakthrough
ISEF · 2015
Large Scale Output Predictions for Small Emplacement Renewable Energy
ISEF · 2021
Using Machine Learning to Improve Numerical Weather Prediction
ISEF · 2026
Optimization of Wind Turbine Performance: Integrating Evolutionary Based AI-Assisted Genetic Algorithms and XGBoost With Blade Element Momentum and Aero-Servo-Hydro-Elastic Simulations
ISEF · 2021
An Investigation in Precipitation Prediction with Machine Learning
ISEF · 2023
Energy Consumption Prediction Using Machine Learning With State-Based Appliance Features Identified by Design of Experiments
Closest projects by meaning, across every fair and year in the corpus.