Document Type : Research Paper
Groundwater, as the primary source of freshwater for drinking, agriculture, and industry, has a vital role in ensuring sustainable development and water security in many regions of the world. In the Mazandaran coastal plain of northern Iran, increasing population growth, industrialization, and intensive agricultural practices have led to escalating pressures on groundwater quality. The infiltration of urban and industrial wastewater, leakage from septic systems, and excessive use of chemical fertilizers and pesticides have resulted in serious chemical degradation of groundwater, manifesting as elevated nitrate, chloride, and total dissolved solids (TDS) concentrations. Because groundwater self-purification processes are extremely slow, such contamination tends to be chronic and often irreversible.
Groundwater quality risk assessment provides an essential basis for identifying pollution‑prone zones and supporting spatially targeted management decisions. However, traditional index-based methods such as DRASTIC, though widely applied, inherently rely on subjective weighting and linear assumptions, which limit their ability to represent non-linear interactions among hydrogeological, anthropogenic, and environmental variables. Recent advances in artificial intelligence and ensemble machine learning have created opportunities to overcome these constraints by capturing hidden patterns and complex dependencies in large multi-source datasets.
This study aims to model and evaluate the groundwater quality risk in the Mazandaran Plain using the Extreme Gradient Boosting (XGBoost) algorithm integrated with Geographic Information System (GIS) datasets. The objectives are to: (i) identify major natural and human-induced factors controlling water quality risk; (ii) generate spatially explicit risk maps showing vulnerable zones; and (iii) assess the interpretability and performance of the XGBoost–GIS combination as a decision-support framework for integrated groundwater management.
Groundwater quality data were collected from monitoring wells across the Mazandaran Plain, spanning both urban–industrial and agricultural zones. Analytical parameters included electrical conductivity (EC), total dissolved solids (TDS), nitrate (NO₃⁻), and chloride (Cl⁻), serving as proxies of salinity and contamination intensity. Predictor variables were derived from multi-source spatial datasets and categorized as:
(1) Topographic and morphological factors, including elevation, slope, and drainage density (derived from a digital elevation model);
(2) Climatic and hydrological variables, such as rainfall and proximity to rivers;
(3) Land-use and remote sensing indices, including the Normalized Difference Vegetation Index (NDVI), Normalized Difference Built-up Index (NDBI), Normalized Difference Water Index (NDWI), and Land Surface Temperature (LST); and
(4) Anthropogenic proximity indicators, particularly distance from industrial estates, urban centers, and major roads, which collectively represent diffusion pathways and contamination pressure.
All spatial data were resampled to a uniform resolution and integrated into the GIS environment for spatial overlay and extraction of pixel-based attribute values. The XGBoost algorithm was employed as a classification model to categorize groundwater quality risk into three levels (low, moderate, high) based on the measured EC and TDS thresholds. Prior to model training, predictor variables were standardized to minimize scale bias, and Pearson correlation coefficients were calculated to verify independence and remove redundant features.
Model parameters, including learning rate, number of estimators, maximum tree depth, and sub-sampling rate, were tuned using a grid-search approach and cross-validation to prevent overfitting. The model accuracy was quantified using performance indicators such as coefficient of determination (R²), overall accuracy (OA), and the confusion matrix. The SHAP (SHapley Additive exPlanations) technique was further applied to interpret feature importance both globally and locally, identifying which variables contributed most strongly to high or low predicted risk.
Finally, the trained model was applied to all spatial pixels to create groundwater quality risk maps within the GIS framework. The resulting raster layers were validated through visual comparison with measured pollutant concentrations and spatial trends known from hydrogeological and land-use data.
The XGBoost algorithm demonstrated reliable performance in predicting and classifying groundwater quality risk patterns across the Mazandaran Plain. The model achieved a coefficient of determination R² ≈ 0.49 for EC and TDS, which indicates a significant predictive ability given the inherent complexity of hydrochemical variability. Confusion matrix analysis suggested that the classifier correctly identified high- and low-risk zones with acceptable precision and recall values, while moderate-risk areas showed the greatest spatial variability.
The feature importance analysis revealed that elevation exerted the dominant influence on risk prediction, suggesting that low-lying coastal areas are more vulnerable to both natural salinization and anthropogenic contamination. The distance to urban and industrial areas ranked as the second most important variable, highlighting the role of point‑source pollution from untreated wastewater and industrial discharge. Among remote sensing indices, NDBI showed a strong positive correlation with groundwater salinity and contamination levels, while NDVI displayed an inverse relation, indicating that vegetated zones provide partial protective effects through enhanced infiltration and dilution. NDWI and LST also contributed moderately, reflecting the interaction between surface moisture, evapotranspiration, and solute accumulation processes.
Spatial risk maps derived from XGBoost predictions showed that high-risk zones are concentrated in low‑elevation regions along the coastal areas of the plain, especially near industrial clusters and densely urbanized sectors. Conversely, low-risk areas correspond to higher elevations and lands dominated by agricultural and forest uses. The pattern of risk distribution aligns well with known hydrogeological pathways and pollution sources, confirming the model’s spatial realism.
The results corroborate previous studies in Iran and similar coastal environments that report the predominance of anthropogenic drivers—particularly urbanization and industrialization—in shaping groundwater quality degradation. The finding that elevation and land-use intensity are decisive predictors is consistent with the fundamental hydrogeologic principle that groundwater flow converges toward lowlands, where accumulation of salts and contaminants is facilitated.
Unlike conventional index-based models (e.g., modified DRASTIC), the XGBoost–GIS framework captured intricate non-linear interactions without requiring subjective weighting, producing a more data-driven depiction of spatial vulnerability. The relatively strong predictive performance obtained with a limited dataset demonstrates the potential of ensemble boosting models to compensate for missing or uncertain hydrogeological inputs by exploiting high-dimensional relationships in hybrid datasets.
Interpretation of SHAP plots shows how certain features, notably elevation and NDBI, interact to reinforce contamination potential. Low-altitude pixels with high NDBI values (dense built-up zones) consistently obtain the highest predicted risk, while increases in NDVI and distance to industrial areas reduce it substantially. This local interpretability provides transparent reasoning for each prediction, which conventional statistical regressions typically fail to furnish.
Furthermore, the relatively high sensitivity of EC/TDS to both anthropogenic and natural factors implies salinity is a composite indicator influenced by groundwater flow paths, lithological dissolution, and human impacts. Therefore, maps based on this indicator provide a reliable representation of integrated groundwater risk.
Comparative evaluation suggests that models combining machine learning and spatial analysis outperform purely deterministic or heuristic approaches, providing accurate and objective risk delineation. However, as with any data-driven model, performance is contingent on the quality and density of monitoring data. Thus, future applications should integrate periodic sampling and real-time sensors to feed adaptive learning pipelines for improved prediction at temporal scales.
This study demonstrates the effectiveness of integrating the Extreme Gradient Boosting (XGBoost) algorithm with GIS for spatial modeling and evaluation of groundwater quality risk in the Mazandaran Plain. The developed model successfully identified the primary controlling factors of risk, including elevation, proximity to urban and industrial areas, and land-use indices (NDBI, NDVI). The resulting risk maps emphasize that the low-lying and highly urbanized parts of the plain are the most vulnerable segments, necessitating urgent management interventions.
The combination of machine learning’s predictive capacity and GIS’s spatial integration capabilities offers a robust analytical framework that can guide groundwater management in data-limited regions. The study underscores the need for continuous water quality monitoring and the integration of spatial modeling results into regional land-use planning and pollution control policies.
Future improvements could involve coupling the model with hydrogeochemical modeling or time‑series satellite data to capture temporal variability. Nevertheless, the XGBoost–GIS approach provides a powerful decision-support tool for identifying priority zones, optimizing monitoring networks, and formulating sustainable management strategies to safeguard groundwater resources in rapidly urbanizing and industrializing coastal plains.
This research was financially supported by the Vice-Chancellery for Research of the Faculty of Engineering, University of Ardakan, Iran.
All authors contributed equally to the conceptualization, methodology, investigation, data analysis, interpretation of results, and writing of the original and revised versions of the manuscript. All authors have read and agreed to the published version of the manuscript.
During the preparation of this work, the authors used OpenAI to improve language clarity and academic writing. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.
Data available on request from the authors.
The authors would like to thank all individuals and organizations who contributed directly or indirectly to the completion of this research. The authors also thank the anonymous reviewers for their constructive comments and valuable suggestions that helped improve the manuscript.
The authors adhered to ethical principles in conducting and reporting the research and avoided data fabrication, falsification, plagiarism, and any form of misconduct.
The authors declare no conflict of interest.