| Issue |
BIO Web Conf.
Volume 238, 2026
VI International Scientific and Practical Conference “Ensuring Sustainable Development in the Context of Agriculture, Energy, Ecology and Earth Science” (ESDCA 2026)
|
|
|---|---|---|
| Article Number | 02010 | |
| Number of page(s) | 10 | |
| Section | Ecology and Conservation of Biological Diversity | |
| DOI | https://doi.org/10.1051/bioconf/202623802010 | |
| Published online | 10 June 2026 | |
Comparative analysis of classical machine learning algorithms for binary classification of drinking water portability
Russian State Agrarian University – Moscow Timiryazev Agricultural Academy, Moscow, Russian Federation
* Corresponding author: This email address is being protected from spambots. You need JavaScript enabled to view it.
Abstract
Today, drinking-quality water is becoming a scarce resource, where water quality serves as an indicator of anthropogenic impact on the environment. This study presents a solution to the problem of binary classification of drinking water based on hydrochemical parameters, utilizing a comparative analysis of classical machine learning algorithms. The relevance of this research stems from the need for rapid, automated methods for the preliminary assessment of water quality within the framework of decision support systems for water treatment and monitoring. The objective of the study was to identify the most balanced algorithm for classifying water as either suitable or unsuitable for drinking, based on standard physicochemical parameters. Evaluating water against these parameters aids in assessing the state of the aquatic environment and its ecological safety. To address this classification task using machine learning algorithms, a complete data analysis pipeline was implemented. Three algorithms were subjected to detailed examination: Decision Tree, Random Forest, and Gradient Boosting. Automated modeling was employed using the PyCaret library. All models were evaluated on a dedicated test dataset using the following metrics: Accuracy, Precision, Recall, F1-score, and AUC-ROC. The most balanced results were obtained with the Random Forest model, which achieved an Accuracy of 0.61 and an F1-score of 0.48. The Gradient Boosting model demonstrated comparable overall accuracy, while the Decision Tree model yielded the highest Precision score, reaching 0.71. A feature importance analysis for the Random Forest model revealed that pH is the most significant predictor, a finding consistent with fundamental principles of drinking water quality assessment. These results establish a foundation for the preliminary identification of potentially substandard water samples and for supporting water quality monitoring efforts.
© The Authors, published by EDP Sciences, 2026
This is an Open Access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.
Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.
Initial download of the metrics may take a while.

