نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسندگان English
This study investigates the influence of ensemble size on the performance of machine learning models for monthly rainfall forecasting using the CHIRPS satellite precipitation dataset. A diverse set of machine learning algorithms, including tree-based, regression-based, kernel-based, instance-based, and ensemble models, was employed to generate independent rainfall predictions. The outputs of these models were then combined using a weighted ensemble approach. For all possible ensemble combinations ranging from 2 to 15 models, optimal weights were determined through a Brute Force search, and the best ensemble was selected based on the Nash–Sutcliffe Efficiency (NSE) criterion. The main objective was to determine the optimal ensemble size while considering both predictive accuracy and computational cost. Results showed a nonlinear relationship between ensemble size and forecasting performance. The best performance was achieved by a four-model ensemble consisting of SMOreg, AdditiveRegression, DecisionTable, and M5Rules, which produced an NSE of 0.88 and an RMSE of 21.10 mm. This ensemble provided the best balance between prediction accuracy, stability, and computational efficiency. Increasing the number of ensemble members beyond four did not improve performance significantly. Instead, greater redundancy and higher correlation among model errors increased computational cost without meaningful accuracy gains. The findings indicate that structural diversity among base learners is more important than simply increasing ensemble size. The proposed framework can be adapted to other regions with similar climatic conditions after retraining with local datasets.
کلیدواژهها English