تحقیقات آب و خاک ایران

تحقیقات آب و خاک ایران

تحلیل اثر اندازه آنسامبل و وزن‌دهی بهینه مدل‌های یادگیری ماشین بر دقت پیش‌بینی بارش ماهانه بارش

نوع مقاله : مقاله پژوهشی

نویسندگان
گروه مهندسی آبیاری و آبادانی، دانشکدگان کشاورزی و منابع طبیعی، دانشگاه تهران، کرج، ایران
10.22059/ijswr.2026.416388.670156
چکیده
در این پژوهش، اثر تعداد اعضای انسامبل بر عملکرد مدل‌های یادگیری ماشین در پیش‌بینی بارش ماهانه با استفاده از داده‌های ماهواره‌ای CHIRPS بررسی شد. برای این منظور، مجموعه‌ای متنوع از مدل‌های یادگیری ماشین شامل مدل‌های درختی (RandomForest، M5P و M5Rules)، رگرسیونی (LinearRegression و AdditiveRegression)، کرنلی (SMOreg و GaussianProcesses)، مبتنی بر فاصله (IBk و LWL) و مدل‌های ترکیبی (Bagging، RandomSubSpace و RandomCommittee) به‌کار گرفته شد و هر مدل به‌صورت مستقل آموزش دید. سپس خروجی مدل‌ها با استفاده از روش انسامبل وزن‌دار ترکیب شد؛ به‌گونه‌ای که برای تمام ترکیب‌های ممکن از ۲ تا ۱۵ مدل، وزن‌های بهینه با روش جستجوی فراگیر تعیین و بهترین ترکیب بر اساس بیشینه‌سازی شاخص نش–ساتکلیف (NSE) انتخاب شد. هدف اصلی پژوهش، تعیین اندازه بهینه انسامبل و بررسی رابطه میان تعداد مدل‌ها، دقت پیش‌بینی و هزینه محاسباتی بود. نتایج نشان داد که رابطه میان اندازه انسامبل و دقت پیش‌بینی غیرخطی است. انسامبل چهارمدلی شامل SMOreg، AdditiveRegression، DecisionTable و M5Rules با دستیابی به NSE برابر 0.88 و RMSE برابر 21.10 میلی‌متر، بهترین تعادل را میان دقت، پایداری و هزینه محاسباتی ایجاد کرد. افزایش تعداد مدل‌ها فراتر از این مقدار، به دلیل افزایش همبستگی خطاها و هم‌پوشانی اطلاعات، بهبود معناداری در دقت ایجاد نکرد و تنها موجب افزایش زمان محاسباتی شد. نتایج نشان داد که در طراحی سامانه‌های انسامبلی، تنوع ساختاری مدل‌های پایه اهمیت بیشتری نسبت به افزایش تعداد آن‌ها دارد. چارچوب پیشنهادی این پژوهش قابلیت کاربرد در سایر مناطق با شرایط اقلیمی مشابه را نیز داراست.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Analyzing the Effect of Ensemble Size and Optimal Weighting of Machine Learning Models on the Accuracy of Monthly Rainfall Prediction

نویسندگان English

Ali Dalir Gabrabad
Afshin Ashrafzadeh
Department of Irrigation and Reclamation Engineering, College of Agriculture and Natural Resources, University of Tehran, Karaj, Iran
چکیده English

This study investigates the influence of ensemble size on the performance of machine learning models for monthly rainfall forecasting using the CHIRPS satellite precipitation dataset. A diverse set of machine learning algorithms, including tree-based, regression-based, kernel-based, instance-based, and ensemble models, was employed to generate independent rainfall predictions. The outputs of these models were then combined using a weighted ensemble approach. For all possible ensemble combinations ranging from 2 to 15 models, optimal weights were determined through a Brute Force search, and the best ensemble was selected based on the Nash–Sutcliffe Efficiency (NSE) criterion. The main objective was to determine the optimal ensemble size while considering both predictive accuracy and computational cost. Results showed a nonlinear relationship between ensemble size and forecasting performance. The best performance was achieved by a four-model ensemble consisting of SMOreg, AdditiveRegression, DecisionTable, and M5Rules, which produced an NSE of 0.88 and an RMSE of 21.10 mm. This ensemble provided the best balance between prediction accuracy, stability, and computational efficiency. Increasing the number of ensemble members beyond four did not improve performance significantly. Instead, greater redundancy and higher correlation among model errors increased computational cost without meaningful accuracy gains. The findings indicate that structural diversity among base learners is more important than simply increasing ensemble size. The proposed framework can be adapted to other regions with similar climatic conditions after retraining with local datasets.

کلیدواژه‌ها English

Brute Force Search
ensemble
WEKA
CHIRPS

مقالات آماده انتشار، پذیرفته شده
انتشار آنلاین از 06 مهر 1405