Эффективность различных ансамблевых алгоритмов машинного обучения в классификации индексов качества воды

Farid Hassanbaki Garabaghi a, Semra Benzer b, Recep Benzer c

a Gazi University, Graduate School of Natural and Applied Sciences, Teknikokullar, 06500 Ankara, Turkey; fgarabaghi@gmail.com

b Gazi University, Gazi Education Faculty, Teknikokullar, 06500 Ankara, Turkey; sbenzer@gazi.edu.tr,

c Ankara Medipol University, Department of Management Information System, Ankara, 06050 Ankara, Turkey; recep.benzer@ankaramedipol.edu.tr


https://doi.org/10.29258/CAJWR/2026-R1.v12-1/68-92.eng

*e-mail: sbenzer@gazi.edu.tr

Thematic cluster: Climate & Environment

Type of paper: Research paper

Abstract

The Aral Sea crisis, a landmark case of human-induced environmental disaster, has prompted extensive international intervention. The World Bank (WB), as a major donor, plays a pivotal role in shaping responses to this crisis. This paper argues that the WB’s interventions are legitimized through a powerful discursive strategy: the partial securitization of the crisis as a threat to human security. Using Critical Discourse Analysis (CDA) and securitization theory, this study examines how the WB frames the Aral Sea’s degradation in two key projects: the Syr Darya Control and Northern Aral Sea Phase I Project (NAS) and the Climate Adaptation and Mitigation Program for the Aral Sea Basin (CAMP4ASB). The analysis reveals that by constructing a narrative of imminent threat, the WB prioritizes technical, top-down solutions and measurable economic outcomes. This WB’s strategy in the Aral Sea basin, however, effectively depoliticizes the crisis, obscuring its historical and socio-political roots and sidelining community participation and structural inequities. This paper identifies a process of partial securitization, whereby the WB frames an issue as a critical threat to human security to justify intervention, but deliberately contains it within the realm of technical, market-based solutions, thereby avoiding the full political ramifications of a traditional securitization “speech act”. The study concludes by discussing the implications of this security framing for development policy and advocates for more inclusive, locally grounded approaches to environmental governance.

Available in English

Download the article

For citation:

Hassanbaki, F., Benzer, S., Benzer, R. (2026). Performance of Different Ensemble Machine Learning Algorithms in Classification of the Water Quality Indices. Central Asian Journal of Water Research, 12(1), 68-92. https://doi.org/10.29258/CAJWR/2026-R1.v12-1/68-92.eng

References

Arabgol, R., Sartaj, M., & Asghari, K. (2016). Predicting Nitrate Concentration and Its Spatial Distribution in Groundwater Resources Using Support Vector Machines (SVMs) Model. Environmental Modeling & Assessment, 21:71–82. https://doi.org/10.1007/s10666-015-9468-0


Arora, N., & Kaur, P. D. (2020). A Bolasso based consistent feature selection enabled random forest classification algorithm: An application to credit risk assessment. Applied Soft Computing Journal. 86:105936. https://doi.org/10.1016/j.asoc.2019.105936


Bhati, B. S., Chugh, G., Al-Turjman, F., & Bhati, N. S. (2021). An improved ensemble based intrusion detection technique using XGBoost. Transactions on Emerging Telecommunications Technologies, 32: e4076. https://doi.org/10.1002/ett.4076


Bissenbayeva, S., Abuduwaili, J., Issanova, G., & Samarkhanov, K. (2020). Characteristics and causes of changes in water quality in the Syr Darya River, Kazakhstan. Water Resources, 47, 904-912. https://doi.org/10.1134/S009780782005019X


Bouamar, M., & Ladjal, M. (2007). Evaluation of the performances of ANN and SVM techniques used in water quality classification. In the 14th IEEE International Conference on Electronics, Circuits and Systems, IEEE, Marrakech, Morocco. https://doi.org/10.1109/ICECS.2007.4511173


Breiman, L. (2001). Random Forests. Machine Learning, 45:5–32. https://doi.org/10.1023/A:1010933404324


Bui, D. T., Khosravi, K., Tiefenbacher, J., Nguyen, H., & Kazakis, N. (2020). Improving prediction of water quality indices using novel hybrid machine-learning algorithms. Science of the Total Environment. 721:137612. https://doi.org/10.1016/j.scitotenv.2020.137612


Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. In the 22nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, San Francisco, California. https://doi.org/10.1145/2939672.2939785


Chen, C. W., Tsai, Y. H., Chang, F. R., & Lin, W. C. (2020). Ensemble feature selection in medical datasets: Combining filter, wrapper, and embedded feature selection results. Expert Systems, 37:e12553. https://doi.org/10.1111/exsy.12553


Cheema, M. A., Hanif, M., Albalawi, O., Mahmoud, E. E., & Nabi, M. (2024). Evaluating water-related health risks in East and Central Asian Islamic Nations using predictive models (2020–2030). Scientific Reports, 14(1), 16837.


Danades, A., Pratama, D., & Anggraini, D. (2016). Comparison of Accuracy Level K-Nearest Neighbor Algorithm and Support Vector Machine Algorithm in Classification Water Quality Status. In the 6th International Conference on System Engineering and Technology (ICSET), IEEE, Bandung, Indonesia. https://doi.org/10.1109/ICSEngT.2016.7849638


Danaei Mehr, H., & Polat, H. (2019). Human Activity Recognition in Smart Home With Deep Learning Approach. In the 7th International Istanbul Smart Grids and Cities Congress and Fair (ICSG), IEEE, Istanbul, Turkey. https://doi.org/10.1109/SGCF.2019.8782290


Dezfooli, D., Moghari, S. M. H., Ebrahimi, K., & Araghinejad, S. (2018). Classification of water quality status based on minimum quality parameters: application of machine learning techniques. Modeling Earth Systems and Environment, 4:311–324. https://doi.org/10.1007/s40808-017-0406-9


Dohare, D., Deshpande, S., & Kotiya, A. (2014). Analysis of Ground Water Quality Parameters: A Review. Research Journal of Engineering Sciences, 3 (5):26-31. ISSN: 2278 – 9472


Dong, X., Yu, Z., Cao, W., Shi, Y., & Ma, Q. (2020). A survey on ensemble learning. Frontiers of Computer Science. 14 (2):241–258. https://doi.org/10.1007/s11704-019-8208-z


Freund, Y., & Schapire, R. E. (1997). A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting. Journal of Computer and System Sciences. 55:119-139. https://doi.org/10.1006/jcss.1997.1504


Friedman, J., Hastie, T., & Tibshirani, R. (2000). Additive logistic regression: a statistical view of boosting (With discussion and a rejoinder by the authors). The Annals of Statistics, 28 (2):337-407. https://doi.org/10.1214/aos/1016218223


General Directorate of Environmental Management (2016) Büyük Menderes Basin Pollution Prevention Action Plan (Turkish). Ministry of Environment and Urbanization, Ankara, Turkey


Ighalo, J. O., Adeniyi, A. G., & Marques, G. (2020). Application of linear regression algorithm and stochastic gradient descent in a machine-learning environment for predicting biomass higher heating value. Biofuels, Bioproducts, and Biorefining, 14:1286–1295, 2020. https://doi.org/10.1002/bbb.2140


Khaire, U. M., & Dhanalakshmi, R. (2019). Stability of feature selection algorithm: A review. Journal of King Saud University – Computer and Information Sciences. 34(4): 1060-1073. https://doi.org/10.1016/j.jksuci.2019.06.012


Kumar, Z. M., & Manjula, R. (2012). Regression model approach to predict missing values in the Excel sheet databases. International Journal of Computer Science & Engineering Technology (IJCSET). 3 (4):130-135. ISSN: 2229-3345


Lap, B. Q., Du Nguyen, H., Hang, P. T., Phi, N. Q., Hoang, V. T., Linh, P. G., & Hang, B. T. T. (2023). Predicting water quality index (WQI) by feature selection and machine learning: a case study of An Kim Hai irrigation system. Ecological Informatics, 74, 101991. https://doi.org/10.1016/j.ecoinf.2023.101991


Liu, Q., Wang, X., Huang, X., & Yin, X. (2020). Prediction model of rock mass class using classification and regression tree integrated AdaBoost algorithm based on TBM driving data. Tunnelling and Underground Space Technology. 106:103595. https://doi.org/10.1016/J.TUST.2020.103595


Leng, P., Zhang, Q., Li, F., Kulmatov, R., Wang, G., Qiao, Y., … & Khasanov, S. (2021). Agricultural impacts drive longitudinal variations of riverine water quality of the Aral Sea basin (Amu Darya and Syr Darya Rivers), Central Asia. Environmental Pollution, 284, 117405. https://doi.org/10.1016/j.envpol.2021.117405


Mădălina, P., & Gabriela, B. I. (2014). Water Quality Index – An Instrument for Water Resources Management. Aerul şi Apa: Componente ale Mediului, 2014:391 – 398.


Modaresi, F., & Araghinejad, S. (2014). A Comparative Assessment of Support Vector Machines, Probabilistic Neural Networks, and K-Nearest Neighbor Algorithms for Water Quality Classification. Water Resources Management, 28:4095–4111. https://doi.org/10.1007/s11269-014-0730-z


Motevalli, A., Naghibi, S. A., Hashemi, H., & Berndtsson, R. (2019). Inverse method using boosted regression tree and k-nearest neighbor to quantify effects of point and non-point source nitrate pollution in groundwater. Journal of Cleaner Production, 228:1248-1263. https://doi.org/10.1016/j.jclepro.2019.04.293


Muhammad, S. Y., Makhtar, M., Rozaimee, A., Aziz, A. A., & Jamal, A. A. (2015). Classification Model for Water Quality using Machine Learning Techniques. International Journal of Software Engineering and Its Applications. 9 (6):45-52


Ostad-Ali-Askari, K., Shayannejad, M., & Ghorbanizadeh-Kharazi, H. (2017). Artificial neural network for modeling nitrate pollution of groundwater in marginal area of Zayandeh-rood River, Isfahan, Iran. KSCE Journal of Civil Engineering, 21:134–140. https://doi.org/10.1007/s12205-016-0572-8


Pan, F., Converse, T., Ahn, D., Salvetti, F., & Donato, G. (2009). Feature Selection for Ranking using Boosted Trees. In Proceedings of the 18th ACM conference on Information and knowledge management, Hong Kong, China. https://doi.org/10.1145/1645953.1646292


Radhakrishnan, N., & Pillai, A. S. (2020). Comparison of Water Quality Classification Models using Machine Learning. In the 5th International Conference on Communication and Electronics Systems (ICCES), IEEE, Coimbatore, India. https://doi.org/10.1109/ICCES48766.2020.9137903


Rozemeijer, J. C., & Broers, H. P. (2007). The groundwater contribution to surface water contamination in a region with intensive agricultural land use (Noord-Brabant, The Netherlands). Environmental Pollution, 148:695-706. https://doi.org/10.1016/j.envpol.2007.01.028


Saghebian, S. M., Sattari, M. T., Mirabbasi, R., & Pal, M. (2014). Ground water quality classification by decision tree method in Ardebil region, Iran. Arabian Journal of Geosciences, 7:4767–4777. https://doi.org/10.1007/s12517-013-1042-y


Satybaldiyev, B., Ismailov, B., Nurpeisov, N., Kenges, K., Snow, D. D., Malakar, A., … & Uralbekov, B. (2023). Downstream hydrochemistry and irrigation water quality of the syr darya, aral sea basin, south kazakhstan. Water Supply, 23(5), 2119-2134. https://doi.org/10.2166/ws.2023.114


Sefidian, A. M., & Daneshpour, N. (2019). Missing value imputation using a novel grey based fuzzy c-means, mutual information based feature selection, and regression model. Expert Systems With Applications, 115:68–94. https://doi.org/10.1016/j.eswa.2018.07.057


Siriwardhana, K. D., Jayaneththi, D. I., Herath, R. D., Makumbura, R. K., Jayasinghe, H., Gunathilake, M. B., Azamathulla, H.M., Tota-Maharaj, K., & Rathnayake, U. (2023). A Simplified Equation for Calculating the Water Quality Index (WQI), Kalu River, Sri Lanka. Sustainability, 15(15), 12012. https://doi.org/10.3390/su151512012


Sim, J., Lee, J. S., & Kwon, O. (2015). Missing Values and Optimal Selection of an Imputation Method and Classification Algorithm to Improve the Accuracy of Ubiquitous Computing Applications. Mathematical Problems in Engineering, 2015:538613. https://doi.org/10.1155/2015/538613


Tehrany, M. S., Jones, S., Shabani, F., Martínez-Álvarez, F., & Bui, D. T. (2019). A novel ensemble modeling approach for the spatial prediction of tropical forest fire susceptibility using LogitBoost machine learning classifier and multi-source geospatial data. Theoretical and Applied Climatology, 137:637–653. https://doi.org/10.1007/s00704-018-2628-9


Turkish Standard Institute (TSE), (2005) TURKISH STANDARD (TS-266): Water Intended for Human Consumption. Turkish Standards Institution, Ankara, Turkey


Tyagi, S., Sharma, B., Singh, P., & Dobhal, R. (2013). Water Quality Assessment in Terms of Water Quality Index. American Journal of Water Resources, 1 (3): 34-38. https://doi.org/10.12691/ajwr-1-3-3


Uddin, M. d. G., Nash, S., & Olbert, A. I. (2021). A review of water quality index models and their use for assessing surface water quality. Ecological Indicators, 122:107218. https://doi.org/10.1016/j.ecolind.2020.107218


Uyun, S., & Sulistyowati, E. (2020). Feature selection for multiple water quality status: integrated bootstrapping and SMOTE approach in imbalance classes. International Journal of Electrical and Computer Engineering (IJECE), 10 (4):4331- 4339. http://doi.org/10.11591/ijece.v10i4.pp4331-4339


Varol, M., & Şen, B. (2012). Assessment of nutrient and heavy metal contamination in surface water and sediments of the upper Tigris River, Turkey. Catena, 92:1–10. https://doi.org/10.1016/j.catena.2011.11.011


World Health Organization (WHO) (2006) Guidelines for Drinking-water Quality: incorporating first addendum. Vol. 1, Recommendations. – 3rd ed. Geneva, Switzerland. ISBN: 92 4 154696 4


Yozgatligil, C., Aslan, S., Iyigun, C., & Batmaz, I. (2013). Comparison of missing value imputation methods in time series: the case of Turkish meteorological data. Theoretical and Applied Climatology, 112:143–167. https://doi.org/10.1007/s00704-012-0723-x


Zebari, R. R., Abdulazeez, A. M., Zeebaree, D. Q., Zebari, D. A., & Saeed, J. N. (2020). A Comprehensive Review of Dimensionality Reduction Techniques for Feature Selection and Feature Extraction. Journal of Applied Science and Technology Trends, 1 (2):56-70. https://doi.org/10.38094/jastt1224


Zhou, Q., Zhou, H., & Li, T. (2016). Cost-sensitive feature selection using random forest: Selecting low-cost subsets of informative features. Knowledge-Based Systems, 95:1–11. https://doi.org/10.1016/j.knosys.2015.11.010

Добавить комментарий