scieee AI-readable full text Open interactive document viewer

Evolutionary and ensemble machine learning predictive models for evaluation of water quality

Aldrees, Ali

Abstract

Study region Bisham Qilla and Doyian stations, Indus River Basin of Pakistan Study focus Water pollution is an international concern that impedes human health, ecological sustainability, and agricultural output. This study focuses on the distinguishing characteristics of an evolutionary and ensemble machine learning (ML) based modeling to provide an in-depth insight of escalating water quality problems. The 360 temporal readings of electric conductivity (EC) and total dissolved solids (TDS) with several input variables are used to establish multi-expression programing (MEP) model and random forest (RF) regression model for the assessment of water quality at Indus River. New hydrological insight for the region The developed models were evaluated using several statistical metrics. The findings reveal that the determination coefficient (R2) in the testing phase (subject to unseen data) for the all the developed models is more than 0.95, indicating the accurateness of the developed models. Furthermore, the error measurements are much lesser with root mean square logarithmic error (RMSLE) nearly equals to zero for each developed model. The mean absolute percent error (MAPE) of MEP models and RF models falls below 10% and 5%, respectively, in all three phases (training, validation and testing). According to the sensitivity study of generated MEP models about the relevance of inputs on the predicted EC and TDS, shows that bi-carbonates and chlorine content have significant influence with a sensitiveness score more than 0.90, whereas the impact of sodium content is less pronounced. All the models (RF and MEP) have lower uncertainty based on the prediction interval coverage probability (PICP) calculated using the quartile regression (QR) approach. The PICP% of each model is greater than 85% in all three stages. Thus, the findings of the study indicate that developing intelligent models for water quality parameter is cost effective and feasible for monitoring and analyzing the Indus River water quality.

Full text

Journal of Hydrology: Regional Studies 46 (2023) 101331 Available online 7 February 2023 2214-5818/© 2023 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/). Evolutionary and ensemble machine learning predictive models for evaluation of water quality Ali Aldrees a , Muhammad Faisal Javed b , Abubakr Taha Bakheit Taha a , Abdeliazim Mustafa Mohamed a , Michał Jasi´ nski c , d , * , Miroslava Gono d a Department of Civil Engineering, College of Engineering, Prince Sattam bin Abdulaziz University, Alkharj 16273, Saudi Arabia b Department of Civil Engineering, COMSATS University Islamabad (CUI), Abbottabad Campus, Islamabad, Pakistan c Faculty of Electrical Engineering, Wroclaw University of Science and Technology, 50-370 Wroclaw, Poland d Faculty of Electrical Engineering and Computer Science, VSB-Technical University of Ostrava, 708-00 Ostrava, Czech Republic ARTICLE INFO Keywords: Water quality assessment Evolutionary algorithm Ensemble learning Random forest regression ABSTRACT Study region: Bisham Qilla and Doyian stations, Indus River Basin of Pakistan Study focus: Water pollution is an international concern that impedes human health, ecological sustainability, and agricultural output. This study focuses on the distinguishing characteristics of an evolutionary and ensemble machine learning (ML) based modeling to provide an in-depth insight of escalating water quality problems. The 360 temporal readings of electric conductivity (EC) and total dissolved solids (TDS) with several input variables are used to establish multiexpression programing (MEP) model and random forest (RF) regression model for the assessment of water quality at Indus River. New hydrological insight for the region: The developed models were evaluated using several statistical metrics. The findings reveal that the determination coefficient (R2) in the testing phase (subject to unseen data) for the all the developed models is more than 0.95, indicating the accurateness of the developed models. Furthermore, the error measurements are much lesser with root mean square logarithmic error (RMSLE) nearly equals to zero for each developed model. The mean absolute percent error (MAPE) of MEP models and RF models falls below 10% and 5%, respectively, in all three phases (training, validation and testing). According to the sensitivity study of generated MEP models about the relevance of inputs on the predicted EC and TDS, shows that bi-carbonates and chlorine content have significant influence with a sensitiveness score more than 0.90, whereas the impact of sodium content is less pronounced. All the models (RF and MEP) have lower uncertainty based on the prediction interval coverage probability (PICP) calculated using the quartile regression (QR) approach. The PICP% of each model is greater than 85% in all three stages. Thus, the findings of the study indicate that developing intelligent models for water quality parameter is cost effective and feasible for monitoring and analyzing the Indus River water quality. * Corresponding author at: Faculty of Electrical Engineering, Wroclaw University of Science and Technology, 50-370 Wroclaw, Poland. E-mail address: [email protected] (M. Jasi´ nski). Contents lists available at ScienceDirect Journal of Hydrology: Regional Studies journal homepage: www.elsevier.com/locate/ejrh https://doi.org/10.1016/j.ejrh.2023.101331 Received 16 June 2022; Received in revised form 30 January 2023; Accepted 31 January 2023 Journal of Hydrology: Regional Studies 46 (2023) 101331 2 1. Introduction The important element for managing and supervising the water cycle, environment, social welfare and economic development is surface water linked with streams and rivers (Pandhiani et al., 2020). There are various factors which affect the quality of river water such as environmental events (like annual-precipitation and erosion) and social components such as industrial and agricultural production processes (Singh et al., 2019). Surface water acts as an utmost resource of fresh water across the globe. Its deficiency can cause severe problems in accessing drinking water, and on environmental and economic development (Shahzad et al., 2019). Water pollution is caused by exchange of river systems and its surroundings, waste products of industries and agriculture (Solangi et al., 2019). Polluted water is an important problem which is creating deadly situation for human welfare, agriculture, and environment (Singh et al., 2019). The accumulated salt due to salinity produce a negative hydrological effect in water, which creates hurdles in its domestic, agricultural, and commercial use. Inspecting water quality and saline water is becoming important thus balance between water supply and demand has reached the required limit (Kim et al., 2016). Electric conductivity (EC) and total dissolved solids (TDS) is an important scale which can be used to determine water quality and its suitability for irrigation and drinking purposes (Jamei et al., 2020). The EC reflects the number of dissolved substances and minerals. Furthermore, TDS is a combination of dissolved inorganic salts like sodium (Na + ), calcium (Ca 2+ ), magnesium (Mg 2+ ), nitrates (NO 3 - ), chloride (Cl - ), sulfate (SO 4 2- ), and other dissolved organic particles (Jagaba et al., 2020). The quality of water can be determined by levels of salt and organic matter, as higher level of salt and organic matter replicates the poor water quality (Jagaba et al., 2020). The assessment, monitoring and regulation of water quality is turn out to be decisive and thus the water supply and demand reached the threshold (Velmurugan et al., 2020). TDS and EC scales have been the topic of laboratory experimental events, but the results of manual laboratory are time consuming and inaccurate and has system errors too (Jamei et al., 2020). Also, the sensors used for calculation of EC and TDS are expansive and need proper maintenance. Since, the sensors are installed on river sites and can be stolen or damaged. So computerized tests can be used to check the water quality (Sattari et al., 2016). By using numerical, statistical, and mechanical methods, numerous scientific publications have analyzed different qualities of water (Bozorg-Haddad et al., 2017; Salami et al., 2016). Such type of traditional sources can still forecast the linear and uniform data sources (Deng et al., 2015). During last two decades, the artificial intelligence (AI) techniques, such as supervised machine learning (ML) has been used widely to overcome different environmental challenges especially water quality index simulation (Aldrees et al., 2022; Alizadeh et al., 2018; Kargar et al., 2020; Singh et al., 2019). ML technique is a step forward for developing the supervision and management of engineering projects (Sihag et al., 2019; Yaseen et al., 2019). Such techniques can give suitable predictions without modern programming officials. ML models gain support from the collected data. The data subsets i.e., training set, validation set and testing techniques are then used to develop and test performance of different predictive models (Najafzadeh et al., 2018). The ML techniques are important and prevalent for predicting the quality of water (Tung and Yaseen, 2020). Study of (Tripathi and Singal, 2019) report the establishment of innovative water-quality-index (WQI) computation method for the Indian-Ganges-River. Principal Components Analysis (PCA) approach was used in which twenty-eight inputs were reduced to only nine, which includes total coliform (TC), dissolved oxygen (DO), chlorine content (Cl), magnesium content (Mg), sulfates content (SO 4 ), electrical conductivity (EC), biochemical oxygen demand (BOD), total dissolved solids (TDS) and power of hydrogen ions (pH). The usage of these nine variables saves a lot of time and give quick analysis medium for water quality assessment. Likewise, (Zali et al., 2011) exercised the most prevalent ML method (ANN: artificial neural network) to check the effect of six inputs (SS: suspended solids, COD: chemical oxygen demand, DO, BOD, pH and nitrates) for the calculation of WQI. By using sensitivity analysis, similarities of each variable were compared in WQI prediction which showed that DO, NO 3 and SS are very important inputs. (Nigam and SM, 2019) assessed fuzzy based models and traditional calculation techniques on the basis of prediction performance for the computation of ground water WQI. It was noted that fuzzy based models have more predictive power. It leaves behind the traditional techniques of calculation and categorizes the quality of water. (Srinivas and Singh, 2018) further extended the study of (Nigam and SM, 2019). They gave the idea of an interactive-fuzzy-model (IFM) which was a remarkable fuzzy decision-making approach for forecasting river WQI. Their research showed that WQI predictive power was more accurate in IFM as compared to traditional fuzzy technique. (Yaseen et al., 2018) did research on the performance of adaptive neuro fuzzy inference system (ANFIS) based hybridized models which includes grid-partitioning (GP), subtractive-clustering (SC), and Fuzzy C-mean data clustering (FCM). It was noted that ANFIS-SC’s performance is persistent, continuous and best. To make a correlation among WQI and biological inputs such as DO, SS, COD, BOD, pH, and nitrates, different algorithms like Radial-basis-function-neural-network (RBFNN) and Back-propagation-neural-network (BPNN) were also used (Hameed et al., 2017). The study reveals that the predictive results of RBFNN model were comparatively good. The functioning of least-square-support-vector-regression (LSSVR) and genetic-programming (GP) was tested for the estimation of potassium, sodium, magnesium, sulphates, pH, TDS and EC in the Sefidrood-River situated in Iran (Bozorg-Haddad et al., 2017). The R 2 is above 0.9 for all established models, which shows good correlation. To check the quality of Abu-Ziriq river in Iran, (Al-Mukhtar and Al-Yaseen, 2019) exercised multiple-linear-regression (MLR), ANFIS, and ANN approach. They tested different models of TDS and EC with most important inputs like magnesium, chloride, nitrate, calcium, sulfate and hardness. During this trial, they reported that ANFIS gave the best results. Similarly, for the estimation of DO in river water, (Sarkar and Pandey, 2015) used ANN technique in different water vicinity with four distinct variables (like pH, DO, BOD and temperature) and concluded that the correlation coefficient (R) is above 0.90 between the predicted and measured data. Furthermore, to estimate the quantity of drinking water produced from chemical treatment plants, (Zhang et al., 2019) exercised the ANN and GP hybridized model and revealed that the performance of these models was better. The further data added to the algorithm improves the efficacy of the output models. Many other different procedures were also used for different hydrological and meteorological challenges (rainfall fall estimation). These methods contain A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 3 tree based algorithmic procedure which include support vector machine (SVM), decision tree (DT) and random forest (RF) model. Different research had revealed that these algorithms have outstanding forecasting abilities (Aghelpour et al., 2019; Hazarika et al., 2020). The authors used SVM and RF models estimate the TSS, BOD, TDS, and COD content (Granata et al., 2017). The predictive power of SVM was excellent during learning stage, however, when this model was exposed to unseen data (fresh data not used while modelling), its result was not efficient. Common and typical algorithms are still used widely but experts are in research of new, flexible and intelligent algorithm (Alizadeh et al., 2018; Kargar et al., 2020). It is worth remembrance that neural networks do not provide the physical aspect of the problem. It acts just like a black box. Models based in ANN are simply called as a relation between input and output (Azimi et al., 2019). Moreover, the relation will either be linear or will depend on the pre-defined functions (Ismael et al., 2021; Kim et al., 2021). Similarly, the results of tree-based algorithms are better in linear relation only (Koranga et al., 2022; Zhang et al., 2021). For the mentioned problems, researchers are using evolutionary algorithm (EA) like GEP for simulation of water quality parameters (WQPs) (Jiang and Chui, 2022; Li et al., 2021; Najafzadeh et al., 2019; Shah et al., 2021). EA works best in a condition where the requirement is precise mathematical equation, and higher generalization capabilities. For creating an ideal predictive model, GEP is unable to work with deviating output values so it must skip in the training period for improving the performance of a model. Moreover, GEP are only suitable in a simple relation between input and output because it encodes one expression at a time and gives its output (Oltean and Grosan, 2003). MEP is opposite to GEP on basis of encoding chromosomes. MEP can code multiple chromosomes in a single program (Oltean and Grosan, 2003). By checking the unknown complications, it has the capability of giving accurate results (Arabshahi et al., 2020). The MEP evolution stage can remove the error complexities from the resultant expression. As compared to other ML methods, decoding process of MEP is quite instinctive. The enormous advantages of MEP algorithm demand its wide scale adoption and recognition. However, it is rarely used by environmentalists. Based on the provided literature review, in current research, water quality indices such as EC and TDC of the Upper Indus Basin were designed at Bisham Qilla monitoring stations using MEP technique and random forest (RF) regression technique based on the most influencing variables. The eight effective parameters considered in this study for the estimation of electric conductivity and totaldissolved-solids are graphically presented in Fig. 1. The total of 360 readings recoded on monthly basis are retrieved from Water-andPower-Development-Authority (WAPDA). These readings are divided into three sets i.e., learning set (training and validation) and testing set. A deep statistical analysis tests and sensitivity study is done on the presented models to check their efficiency and reliability. The established models are beneficial in reducing the difficulty and cost which forecast EC and TDS concentrations properly on a small group of inputs. 2. Research methodology In this section the methods of multi expression programming (MEP) and an ensemble learning approach known as random forest (RF) will be explained. Checking the water quality indicator (TDS and EC) with the selected study area and performance evaluation indicators will be discussed. Fig. 1. Factors affecting the EC and TDS of water. A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 4 2.1. Multi-expression programming Genetic algorithms (GA) are based on stochastic and metaheuristic technique for creating best solutions that depends on the principles of biological selection and genetics (Goldberg, 2006). By using the traditional optimization methods, a series of binary strings is created by GA for describing the results. Genetic programming (GP) was an updated version of GA, which encodes string expressions into executable computer program (Koza and Koza, 1992). For solving any problem, GP implements the Darwin’s biological evolution theory. GP’s main purpose is to use the fitness metric for finding a code that can merge the inputs and outputs. Generally, there are three types of GP, in which linear GP is more productive than graph based and tree based GP (Alavi et al., 2013; Cheng et al., 2020). Reason of its effectiveness is that it does not need active researchers. It also improves the accuracy in real time scale. For forecasting the water quality indices (TDS and EC), a multi expression programming (linear GP approach) was used in current paper. MEP keeps a track of all records by using linear chromosomes. The most suitable answer that defines chromosome is determined by comparing the fitness values of the script. First step in MEP algorithmic process is to create code from a randomly selected population. All the subsequent steps are performed until the required condition is achieved (Oltean and Grosan, 2003): 1. By using simple binary tournament process, parents are selected after which they are established again on the basis of specific crossover probability. 2. The combination of two parents produces two offspring. 3. After that the weakest is changed with the best among them using the mutation process. MEP works same as Pascal and C compilers that interpret mathematical expression to coding (Alavi et al., 2013, 2010). A series of formulae represents the genes of MEP. The length of chromosomes is defined by the set of genes that remains constant throughout the whole calculation. The gene is created by a function set “F” and terminals “T” (1 or 2). For generating the proper code, the first gene of chromosome must match to the randomly chosen terminal inside a terminal set. Having reference is very important for the function gene. Values of the terminal scores produced by the specific gene are less than the gene’s chromosomal location. Simple MEP chromosome is explained by the example provided in Fig. 2. The function set i.e., F= { +, /, Pow}and the terminal set i.e., T= {B1,B2,B3,B4}are uses to encode MEP gene. The MEP gene can be transformed into program by decryption the chromosomes. The decryption process starts from top to bottom. The particular code is offered in the shape of gene in Fig. 2. The genes number 0, 1, 3 and 5 code just single terminal i.e., Y0=B1, Y1=B2, Y3=B3, and Y5=B4 respectively. At positions 0 and 1, the gene 2 is used to represent the ′Pow′function applied on supposed operands and encodes the expression Y2=BB2 1. Similarly, at locations 2 and 3, gene 4 denotes the ′/′function applied on the operands, thus encoding the expression Y4=BB2 1/B3. Ultimately, the gene 6 represent the final expression as Y4=BB2 1/B3+B4. Accordingly, the chromosomal order can be seen in the form of forest of gene trees, entertaining numerous expressions. Following the effectiveness of different expression, the best expression tree is revealed (Oltean and Grosan, 2003). Fig. 2. The typical chromosomes used to denote the genes forming tree in MEP algorithm. A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 5 2.2. Random forest regression approach The RF technique is a popular approach for addressing regression and classification problems. It is a (Breiman, 2001) developed ensemble learning approach that integrates a number of regression trees. Random Forest is basically a set of decision trees whose outcomes are combined in a single outcome. Generally, the correlations and statistical errors achieved in the current RF model are equivalent to those found in previous optimization methods. However, the simplicity of predicting with categorical data utilizing RF makes it a preferred forecasting method. The use of decision trees has the benefit of relatively rapid simulation for huge datasets (Xiong et al., 2019). Furthermore, it can process both numerical and categorical (descriptive) data. The collection of arranged circumstances in compliance with leaf of the regression tree is represented by the root of the regressive trees (Iqbal, M. et al., 2021; Vickers, 2017). Initially, bootstrapping data are generated randomly via substitution of the entire training set, and the prediction trees are assembled to every bootstrap instance. RF then receives the weights of various evidentiary attributes, i.e., the (x) input array. Following that, RF constructs K regression trees and aggregates the findings. The K{T(x)}k 1 are constructed, the RF regression predictor is shown as Eq. (1):  fYf K(x) = ∑ K k=1 T(x) K(1) The bagging technique improves the variability of the forest, causing trees to develop using various training data groups designed to minimize overlapping with competing trees. The data that were not chosen for constructing the k th tree during the bagging procedure are included in additional set known as out-of-bag (OOB) set. These OOB samples are used by the kth tree to evaluate performance (Peters et al., 2007). RF can thus construct a reliable and unbiased estimate of the prediction error not relying on an extrinsic data subset (Xiong et al., 2019). Fig. 3 depicts the flowchart of the RF algorithm. In this research, the results obtained from MEP (metaheuristic approach) are compared with RF (ensemble machine learning method) to estimate the parameters used for water quality assessment, i.e., electric-conductivity and total-dissolved-solids. An attempt is made to make sure the superiority of modern machine learning technique. Ensemble machine technique (RF) used in this study is performed with the help of RapidMiner Studio available online as an open-source tool, while MEP is performed using open-source package MEPX V.2.0. (Chu et al., 2021; Shah et al., 2020). 2.3. Study area and data collection The Indus River is one of the transboundary rivers of the Asia, which flows in three different countries. The length of the river, which is 3180-kilometer, starts from Tibet and enters into India near Ladakh and ends in the Karachi by touching Arabian Sea. The annual flow of Indus River is around 253 km 3 and is one of the 50 largest rivers of the world. The area surrounded by Indus River consists of mountains and glaciers mostly. The water quality assessment in these areas is important for the life of ecosystem established around the river (Shah et al., 2021; Tahir et al., 2011). Water and Power Development Authority (WAPDA) of Pakistan recorded the water quality data on monthly basis at different points of River Indus. The data used in this research was obtained from WAPDA. A total of 360 datapoints were collected from them recorded from 1975 to 2005. The data used in this research was recorded at the Bisham Qilla and Doyian stations. The pictorial representation of the study area is provided in Fig. 4. After vast literature review, eight parameters i.e., water temperature ( 0 C), Calcium (Ca), magnesium (Mg), sulphate (SO 4 ), sodium (Na), pH, bicarbonates (HCO 3 ) and chloride (Cl) were selected as input variables while two parameters i.e., TDS and EC were selected as output. Although the discharge is also among the most important variable, however, the data retrieved from the WAPDA does not include any information about the measurement of discharge at Bisham Qilla and Doyian stations. The statistical details like mean, standard deviation, kurtosis, skewness, and range of all the parameters are shown in Table 1. It can be seen that all values of TDS and EC lie within the acceptable range provided by WHO for drinkable water (TDS: 300–600 mg/L; EC: 88–770 μ S/cm) (Jamei et al., 2020). Also the dispersion statistics lies in the permissible range (kurtosis: [−10, 10]; skewness: [−3, 3] (Khan et al., 2022b; Khan et al., 2022c). As accuracy of machine learning models largely depends on the input data, therefore its necessary to calculate the correlation coefficient amongst the input variables and output to avoid “multicollinearity problem” (Khan et al., 2021a). The correlation coefficients for both outputs are shown in Table 2. All values of correlation coefficient between inputs are less than 0.8, showing good dispersion of data. In addition, the values show a strong connection between selected input variables and output parameter with correlation values greater than 0.8 (Azim et al., 2021; Jalal et al., 2021b). Furthermore, the significance (P-value) of the outputs corresponding to correlation coefficients at significance level of 5% ( α =0.05) are less than 0.00001, showing a significant relationship Fig. 3. Flow chart of Random Forest (RF) algorithm. A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 6 between outputs and input parameters considered in this study. Hence, the data is randomly divided into three different sets (training, testing and validation) as suggested by different researchers (Khan et al., 2022a; Nguyen et al., 2021). 70% of the obtained data is used for training set, while 15% of the data is used for testing set, whereas the remaining 15% is used for validation set. 2.4. Statistical metrics for performance assessment of models To check the accuracy and applicability of the developed modes, various statistical parameters like mean absolute percent error (MAPE), root mean squared logarithmic error (RMSLE), R-square, root mean square error (RMSE) and mean absolute error (MAE) of all three sets are calculated and assessed. Amongst all the statistical parameters, R-square is the most reliable one in assessing the Fig. 4. Graphical description of study area under consideration. Table 1 Statistical analysis of composed data. Parameters Descriptive statistics Range Maximum Minimum Mean Std. Deviation Kurtosis Skewness Concentration of Ca 2+ (mEq/L) Input 103.39 104.00 0.61 1.823 5.420 3.22 1.26 Concentration of Mg 2+ (mEq/L) Input 2.61 2.64 0.03 0.615 0.343 4.42 1.92 Concentration of Na + (mEq/L) Input 8.95 9.00 0.05 0.543 0.672 4.44 1.86 Concentration of HCO 3 - (mEq/L) Input 7.29 7.40 0.11 1.740 0.697 2.66 1.24 Concentration of Cl - (mEq/L) Input 4.20 4.20 0.00 0.297 0.312 5.80 1.25 Concentration of SO 4 2- (mEq/L) Input 3.10 3.20 0.10 0.576 0.383 1.87 2.18 pH Input 8.40 8.40 0.00 7.841 0.621 4.56 -1.86 WT (◦C) Input 20.11 21.11 1.00 12.18 3.812 -0.75 -0.32 Concentration of TDS (ppm) Output 464 524 60 148.35 62.231 37.225 4.584 Concentration of EC ( μ S/cm) Output 690 770 88 250.42 98.691 31.808 3.801 A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 7 applicability of the model. The detail description along with their formulae is provided in Table 3. Where, the ‘P′and ‘M′represents the forecasted and measured reading, respectively and ‘k′is the overall sum of experiments. 2.5. Uncertainty analysis In current research, quantile regression (QR) was applied for evaluation of the overall uncertainty of a predictive model. QR is a simple linear mathematical method used in several studies to estimate the quantiles of a predictor variables in the absence of a hypothesis (Kouadri et al., 2022; Tas¸an et al., 2022). The implementation of QR to analyse error rates based on anticipated water quality indicators has been shown to be a comparatively simple, effective, and reliable approach for evaluating predictive models’ uncertainty (Weerts et al., 2011). For every quartile (q), the mathematical relationship among measured (M) and predicted (P) water quality indicators (EC and TDS) are provided as; M=aqP+bq(2) In Eq. (2), aq and bq are the linear regression coefficients that can be derived and calculated through minimization of the sum of residuals as provided in Eq. (3). min ∑ k j=1 ρ q(Mj−(aqPj+bq)) (3) Wherein Mj and Pj are the jth observations out of a total of J observations and ρ q is the function used for quantile regression: ρ q( ε j)={(q−1) • ε j, ε j≤0 q• ε j, ε j>0}(4) Eq. (4) is used to calculate the error ( ε j), which can be described as the difference between the measured values (Mj) and the estimated linear QR (aqP+bq) for the specified quantile (5–95%). (Shrestha and Solomatine, 2006) proposed uncertainty indicator using the performance evaluation of the QR method i.e., predictive interval coverage probability (PICP) (Tas¸an et al., 2022). The PICP is a useful indicator reflecting that how many readings fit inside the expected area (5–95%). When the PICP is equal to 1, there is no uncertainty. Eq. (5) shows the calculation of PICP. The mean Table 2 Pearson’s correlation-matrix of input and output variables under consideration. Variables Ca Mg Na HCO 3 Cl SO 4 PH WT (◦C) Concentration of Ca 2+ (mEq/L) 1.000 Concentration of Mg 2+ (mEq/L) 0.315 1.000 Concentration of Na + (mEq/L) 0.398 0.332 1.000 Concentration of HCO 3 - (mEq/L) 0.626 0.482 0.594 1.000 Concentration of Cl - (mEq/L) 0.446 0.380 0.490 0.449 1.000 Concentration of SO 4 2- (mEq/L) 0.381 0.444 0.454 0.178 0.198 1.000 pH -0.13 0.116 -0.005 0.023 -0.051 -0.01 1.000 WT (◦C) -0.46 -0.390 -0.249 -0.296 -0.239 -0.31 0.021 1.000 Concentration of TDS (ppm) 0.704 0.619 0.759 0.769 0.587 0.546 -0.630 -0.664 Significance (P-value) Less than 0.00001 Concentration of EC ( μ S/cm) 0.715 0.601 0.728 0.767 0.555 0.508 -0.616 -0.649 Significance (P-value) Less than 0.00001 Table 3 Description of the statistical indicators used for evaluation of developed models. Statistical indicators Formula Description Coefficient of determination (R2) 1−∑k j=1(Pj−Mj)2 ∑k j=1(Pj−Mj)2 The highest value of R-square is 1 for perfect fit. Closer the value to 1, the more accurate and effective is the model will be (Nafees et al., 2022; Shah et al., 2021). Mean absolute error (MAE) ∑k j=1Pj−Mj k MAE is calculated by adding the absolute of all the differences between predicted output and experimental values and then dividing it by total number of observations. MAE can be used to detect the low error values in a set (Iqbal, M.F. et al., 2021). Root mean square logarithmic error (RMSLE)  ∑k j=1[log(Mj+1)−log(Pj+1)]2 k √RMSLE determines the relative variance amongst the forecasted and measured readings. It uses the logarithmic function to incorporate the enlarged error and beneficial for skewed readings (Iqbal, M.F. et al., 2021). Root mean square error (RMSE) ∑k j=1(Mj−Pj)2 k √It is exercised to effectively assess the magnitude of larger errors having substantial divergences. The RMSE provides more weightage to the outliers and error is measured by pondering the root of squared-error (Iqbal, M.F. et al., 2021). Mean absolute percent error (MAPE)∑k j=1Mj−Pj Mj k It is used to illustrates the percentage of either positive or negative error and to classify the machine learning models as Excellent: 0%≤MAPE≤10%; Good: 10%≤MAPE≤20%; Acceptable: 20%≤MAPE<50%; and Inaccurate: 50%≤MAPE (Khan, S. et al., 2022). A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 8 width of the predictive interval (uncertainty band) is the range between the lower predictive limit (PLlower j) and upper predictive limit (PLupper j). PICP =1 k∑ k j=1 C(5) C={1,PLlower j<Pj<PLupper j 0,otherwise }(6) 2.6. Tunning the MEP model algorithmic parameters When beginning modeling, specific MEP algorithmic parameters are turned to generate robust, accurate, and adaptive models. The appropriate settings are established considering literature recommendations and a hit and trial technique (Mousavi et al., 2010). The overall number of programs raised in the population is determined by the size of the population. An increased population take longer to reach convergence and provide more complicated, but realistic and correct code. Yet, increasing the population above certain level may cause overfitting in the simulated model. The optimal algorithmic parameters for the best MEP models (EC and TDS) are provided in Table 4. The modeling process started with the supposition of total ten subpopulations. Basic arithmetic operators (-, +, x, /) were chosen to generate a simple formula and to reach convergence with a better level of accurateness and robustness. Different Iterations are used to define the level of accuracy which the program should achieve until being stopped. Simulations with multiple iterations provide a model with the least error. Also, mutation and crossover rates manage the chance of offspring participating in the genetic process. The crossover rate ranges from 50% to 95%. Following the evaluation of numerous configurations of these parameters, the top potential combinations were selected based on the efficacy of the model, as shown in Table 4. An overfitting of the model performance is a common issue with supervised ML modeling. A model seems to perform acceptably on the presented data set, but its effectiveness degrades substantially when evolved using the unknown data. To tackle the stated difficulty, the research proposed to evaluate the trained model on an unseen data (testing dataset) (Qiu et al., 2020). In a nutshell, the whole readings in the database were randomly divided into 3 consistent and unique sets. Training set comprised of 252 readings (70%), while both the validation and testing set contains 54 readings (15%). It was confirmed that data distribution is consistent across all sets. The optimal model was computed using the learning set (training and validation). The accuracy of the validated models was next assessed with unknown data kept aside in the testing set. The generated models performed well in all 3 levels. The process involves the production of a population for finding an optimal solution. The procedures are recurring, with each cycle getting us nearer to the best solution. Each iteration performance is evaluated over the whole population. The algorithm will continue to evolve till a pre-defined fitness (R 2 or RMSE) value becomes stagnant. When the model results for any of the three datasets were inappropriate, the process was repeated by gradually increasing the number and size sub-populations. The best solution is then picked due to the least RMSE and high R 2 . It is worth remembering that evolving duration and number of generations make an impact on the correctness of the models generated. Because of the incorporation of more parameters into the operation, a model would evolve indefinitely. In the present simulation, however, the modeling was stopped at 1000 generations or when the increase in fitness score dropped to 0.1%. Additionally, as described in Section, an optimal solution should meet several performance measures. 2.7. Tunning the RF model algorithmic parameters The preceding approach was used to create the RF model in the RapidMiner Studio (Version 9.8) framework. The data from 360 temporal readings of the water quality indices was obtained in the processing window during the first stage of model development. A number of operators was used to execute the process. The “select-attribute” operator to select the attributes-subset for the modelling process. The command “set-role” was then employed to modify the position of an attribute as a targeting variable. Among several factors assessed, the EC and TDS were picked as the target attribute. Applying the "split-data" operation, the data were randomly Table 4 The tunned algorithmic parameters of developed EC-MEP and TDS-MEP models. Tunned operators Optimized setting Linking operators Subtraction, addition, division, multiplication, Ln, Sqrt Subpopulation size 250 Num subpopulations 50 Crossover probability 0.9 Code length 50 Mutation probability 0.01 Crossover type uniform Operators probability 0.5 Tournament size 2 Num generations 1000 Variables probability 0.5 A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 9 divided into 70% train and 30% validation set. "Optimize-parameter (grid)" was utilized to determine the optimum solution of the specified parameters. Numerous other criterions were used to analyze the best possible productivity of the presented model, including the sum of trees, maximum depth, minimum leaf size, confidence, apply pre-pruning and pruning, RF minimum gain, guess subset ratio, random splits, minimum size for split, and enable parallel execution. The optimum RF model was obtained by employing the lowest value of RF minimal gain (i.e., zero), while the highest value over the 1000 steps on a linear scale is equals to infinity. The optimum RF model was obtained by employing the lowest value of RF minimal gain (i.e., zero), while the highest value over the 1000 steps on a linear scale is equals to infinity. A greater minimum gain value leads in smaller splits and, consequently, fewer trees. A too high value eliminates splitting, resulting in trees containing only one node. As a result, the frequency of repetitions was altered to get an optimal model in many trails. Ultimately, 1000 trials produced the best model. The operator’s ports were precisely linked to get the finest RF model. The "cross-validation" with automated sampling technique was included inside the "optimize-grid-parameter" using the k-fold cross validation approach. For this reason, the whole data was split into ten different subsets. The nine subsets of data were used for learning by performing the regression analysis, while the one subset is kept for cross-validation. The operation is repeated till all the subsets are used as a cross-validation set. The ultimate cross-validation performance is calculated by taking the average of performance metrics of all cross-validation sample. After expanding the density of trees indefinitely, it was discovered that 200 trees produced the best results. The processing duration of the model development doubled as the values of hyper-parameters increasing, without enhancing the correlations or creating a decrease in error. Table 5 lists the additional training parameters for the learned model. The generated RF model’s performance was evaluated using R 2 , RMSLE, MAE, RMSE and MAPE. 3. Results and discussions 3.1. Formulation of TDS and EC using MEP model The MEP code of EC and TDS is denoted by S-1 and S-2. The code utilizes nine input variables as provided in supplementary information. The experimental EC predicted results is graphically presented in Fig. 5. Best position of regression line is 45 0 which means that the slope must be equal to 1 for strong correlation (Iqbal, M.F. et al., 2021). The slope for all three sets is shown in Fig. 5 (a) as training: 0.9885, validation: 0.9921 and testing: 0.9918. These results show a mutual relationship between experimental and forecasted EC. Moreover, the results are nearly close to each other which is a best fit for sets. Due to this the proper training of model is defined (Iqbal, M.F. et al., 2021; Khan, S. et al., 2022). Consequently, it has great generalization power as it works almost similar on testing set (fresh unseen data). Similar analysis was also done for TDS findings of developed MEP model as shown in Fig. 5 (b). Generated model of TDS was well trained and has the capacity of following experimental TDS. The slope of regression line of all three sets is nearly equal i.e., training: 0.9924, validation: 1.0119, and testing: 0.9989. The functioning of TDS model in testing stage is exceptional just like an EC model which shows that the over-fitting problem was solved in a best manner (Anonna et al., 2021; Iqbal, M.F. et al., 2021; Jalal et al., 2021a). Moreover, the R 2 of each MEP model is higher than 0.9 in training, validation and testing phase. It is clear from Fig. 5 (a) and 5 (b) that the spread of data points is nearer to 45 0 lines, showing the greater prediction accuracy. Like the slope, the R 2 values of both models are also above 0.9. Specifically on the unseen data (testing phase), it is equals to 0.9725 and 0.9674 for TDS and EC, respectively. Performance of both the models is similar in each phase of the developed MEP model with a minor difference between slope or R 2 values (Khan et al., 2021b). Moreover, the efficiency and adaptability of the models also depend on the number of data points use during simulation (Gholampour et al., 2017). The higher number of data points equaling 360, results in an accurate prediction with less errors in both models (Mansouri et al., 2018). 3.1.1. Comprehensive performance analysis of developed MEP models Reliability of model and number of used readings in foundation of model are directly proportional to each other. Ration of instances Table 5 The tunned algorithmic parameters of developed EC-RF and TDS-RF models. Tunned parameters Optimized setting Operator RF Chosen parameter Minimal gain Grid scale Linear Grid range Zero to infinity Error handling Fail on error Log performance Checked Sampling type Automatic Number of cross-validation folds Ten Number of trees 20 Maximum depth 20 Criteria Least square Performance indicators R 2 , RMSLE, MAE, RMSE, MAPE A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 16 Goldberg, D.E., 2006. Genetic algorithms. Pearson Education, India. Granata, F., Papirio, S., Esposito, G., Gargano, R., De Marinis, G., 2017. Machine learning algorithms for the forecasting of wastewater quality indicators. Water 9 (2), 105. Hameed, M., Sharqi, S.S., Yaseen, Z.M., Afan, H.A., Hussain, A., Elshafie, A., 2017. Application of artificial intelligence (AI) techniques in water quality index prediction: a case study in tropical region, Malaysia. Neural Comput. Appl. 28 (1), 893–905. Hazarika, B.B., Gupta, D., Ashu, Berlin, M., 2020. A Comparative Analysis of Artificial Neural Network and Support Vector Regression for River Suspended Sediment Load Prediction. Springer, Singapore, Singapore, pp. 339–349. Ismael, M., Mokhtar, A., Farooq, M., Lü, X., 2021. Assessing drinking water quality based on physical, chemical and microbial parameters in the Red Sea State, Sudan using a combination of water quality index and artificial neural network model. Groundw. Sustain. Dev. 14, 100612. Jagaba, A., Kutty, S., Hayder, G., Baloo, L., Abubakar, S., Ghaleb, A., Lawal, I., Noor, A., Umaru, I., Almahbashi, N., 2020. Water quality hazard assessment for hand dug wells in Rafin Zurfi, Bauchi State, Nigeria. Ain Shams Eng. J. 11 (4), 983–999. Jalal, F.E., Xu, Y., Iqbal, M., Jamhiri, B., Javed, M.F., 2021a. Predicting the compaction characteristics of expansive soils using two genetic programming-based algorithms. Transp. Geotech. 30, 100608. Jalal, F.E., Xu, Y., Iqbal, M., Javed, M.F., Jamhiri, B., 2021b. Predictive modeling of swell-strength of expansive soils using artificial intelligence approaches: ANN. ANFIS GEP 289, 112420. Jamei, M., Ahmadianfar, I., Chu, X., Yaseen, Z.M., 2020. Prediction of surface water total dissolved solids using hybridized wavelet-multigene genetic programming: new approach. J. Hydrol. 589, 125335. Jiang, L., Chui, T.F.M., 2022. A review of the application of constructed wetlands (CWs) and their hydraulic, water quality and biological responses to changing hydrological conditions. Ecol. Eng. 174, 106459. Kargar, K., Samadianfard, S., Parsa, J., Nabipour, N., Shamshirband, S., Mosavi, A., Chau, K.-w, 2020. Estimating longitudinal dispersion coefficient in natural streams using empirical models and machine learning algorithms. Eng. Appl. Comput. Fluid Mech. 14 (1), 311–322. Khan, M.A., Shah, M.I., Javed, M.F., Khan, M.I., Rasheed, S., El-Shorbagy, M., El-Zahar, E.R., Malik, M., 2021a. Application of random forest for modelling of surface water salinity. Ain Shams Engineering Journal. Khan, M.A., Zafar, A., Farooq, F., Javed, M.F., Alyousef, R., Alabduljabbar, H., Khan, M.I., 2021b. Geopolymer concrete compressive strength via artificial neural network, adaptive neuro fuzzy interface system, and gene expression programming with K-fold cross validation. Front. Mater. 8, 621163 https://doi.org/ 10.3389/fmats. Khan, M.A., Farooq, F., Javed, M.F., Zafar, A., Ostrowski, K.A., Aslam, F., Malazdrewicz, S., Ma´ slak, M., 2022b. Simulation of depth of wear of eco-friendly concrete using machine learning based computational approaches. Materials 15 (1), 58. Khan, S., Ali Khan, M., Zafar, A., Javed, M.F., Aslam, F., Musarat, M.A., Vatin, N.I., 2022. Predicting the ultimate axial capacity of uniaxially loaded CFST columns using multiphysics artificial intelligence. Materials 15 (1), 39. Kim, H., Jeong, H., Jeon, J., Bae, S., 2016. Effects of irrigation with saline water on crop growth and yield in greenhouse cultivation. Water 8 (4), 127. Kim, J., Seo, D., Jang, M., Kim, J., 2021. Augmentation of limited input data using an artificial neural network method to improve the accuracy of water quality modeling in a large lake. J. Hydrol. 602, 126817. Koranga, M., Pant, P., Kumar, T., Pant, D., Bhatt, A.K., Pant, R.P., 2022. Efficient water quality prediction models based on machine learning algorithms for Nainital Lake, Uttarakhand. Materials Today: Proceedings. Kouadri, S., Pande, C.B., Panneerselvam, B., Moharir, K.N., Elbeltagi, A., 2022. Prediction of irrigation groundwater quality parameters using ANN, LSTM, and MLR models. Environ. Sci. Pollut. Res. 29 (14), 21067–21091. Koza, J.R., Koza, J.R., 1992. Genetic programming: on the programming of computers by means of natural selection. MIT Press. Li, L., Rong, S., Wang, R., Yu, S., 2021. Recent advances in artificial intelligence and machine learning for nonlinear relationship analysis and process control in drinking water treatment: A review. Chem. Eng. J. 405, 126673. Liu, Q.-f, Iqbal, M.F., Yang, J., Lu, X.-y, Zhang, P., Rauf, M., 2021. Prediction of chloride diffusivity in concrete using artificial neural network: Modelling and performance evaluation. Constr. Build. Mater. 268, 121082. Mansouri, I., Gholampour, A., Kisi, O., Ozbakkaloglu, T., 2018. Evaluation of peak and residual conditions of actively confined concrete using neuro-fuzzy and neural computing techniques. Neural Comput. Appl. 29 (3), 873–888. Mousavi, S., Alavi, A., Gandomi, A., Arab Esmaeili, M., Gandomi, M., 2010. A data mining approach to compressive strength of CFRP-confined concrete cylinders. Struct. Eng. Mech. 36 (6), 759. Nafees, A., Amin, M.N., Khan, K., Nazir, K., Ali, M., Javed, M.F., Aslam, F., Musarat, M.A., Vatin, N.I., 2022. Modeling of Mechanical Properties of Silica Fume-Based Green Concrete Using Machine Learning Techniques. Polymers 14 (1), 30. Najafzadeh, M., Rezaie-Balf, M., Tafarojnoruz, A., 2018. Prediction of riprap stone size under overtopping flow using data-driven models. Int. J. River Basin Manag. 16 (4), 505–512. Najafzadeh, M., Ghaemi, A., Emamgholizadeh, S., 2019. Prediction of water quality parameters using evolutionary computing-based formulations. Int. J. Environ. Sci. Technol. 16 (10), 6377–6396. Nguyen, Q.H., Ly, H.-B., Ho, L.S., Al-Ansari, N., Le, H.V., Tran, V.Q., Prakash, I., Pham, B.T., 2021. Influence of Data Splitting on Performance of Machine Learning Models in Prediction of Shear Strength of Soil. Math. Probl. Eng. 2021, 4832864. Nigam, U., SM, Y., 2019. Development of computational assessment model of fuzzy rule based evaluation of groundwater quality index: comparison and analysis with conventional index, Proceedings of International Conference on Sustainable Computing in Science, Technology and Management (SUSCOM), Amity University Rajasthan, Jaipur-India. Oltean, M., Grosan, C., 2003. A comparison of several linear genetic programming techniques. Complex Syst. 14 (4), 285–314. Pandhiani, S.M., Sihag, P., Shabri, A.B., Singh, B., Pham, Q.B., 2020. Time-series prediction of streamflows of Malaysian rivers using data-driven techniques. J. Irrig. Drain. Eng. 146 (7), 04020013. Peters, J., De Baets, B., Verhoest, N.E., Samson, R., Degroeve, S., De Becker, P., Huybrechts, W., 2007. Random forests as a tool for ecohydrological distribution modelling. Ecol. Model. 207 (2–4), 304–318. Qiu, R., Wang, Y., Wang, D., Qiu, W., Wu, J., Tao, Y., 2020. Water temperature forecasting based on modified artificial neural network methods: two cases of the Yangtze River. Sci. Total Environ. 737, 139729. Salami, E., Salari, M., Ehteshami, M., Bidokhti, N., Ghadimi, H., 2016. Application of artificial neural networks and mathematical modeling for the prediction of water quality variables (case study: southwest of Iran). Desalin. Water Treat. 57 (56), 27073–27084. Sarkar, A., Pandey, P., 2015. River water quality modelling using artificial neural network technique. Aquat. Procedia 4, 1070–1077. Sattari, M.T., Joudi, A.R., Kusiak, A., 2016. EstimatioN Of Water Quality Parameters With Data-driven Model. J. Water Works Assoc. 108 (4), E232–E239. Shah, M.I., Javed, M.F., Abunama, T., 2020. Proposed formulation of surface water quality and modelling using gene expression, machine learning, and regression techniques. Environ. Sci. Pollut. Res. Shah, M.I., Javed, M.F., Abunama, T., 2021. Proposed formulation of surface water quality and modelling using gene expression, machine learning, and regression techniques. Environ. Sci. Pollut. Res. Int. 28 (11), 13202–13220. Shahzad, G., Rehan, R., Fahim, M., 2019. Rapid performance evaluation of water supply services for strategic planning. Civ. Eng. J. 5 (5), 1197–1204. Sharafati, A., Pezeshki, E., 2020. A strategy to assess the uncertainty of a climate change impact on extreme hydrological events in the semi-arid Dehbar catchment in Iran. Theor. Appl. Climatol. 139 (1), 389–402. Shrestha, D.L., Solomatine, D.P., 2006. Machine learning approaches for estimation of prediction interval for the model output. Neural Netw. 19 (2), 225–235. Sihag, P., Tiwari, N., Ranjan, S., 2019. Prediction of unsaturated hydraulic conductivity using adaptive neuro-fuzzy inference system (ANFIS). ISH J. Hydraul. Eng. 25 (2), 132–142. A. Aldrees et al. Journal of Hydrology: Regional Studies 46 (2023) 101331 17 Singh, A.P., Dhadse, K., Ahalawat, J., 2019. Managing water quality of a river using an integrated geographically weighted regression technique with fuzzy decisionmaking model. Environ. Monit. Assess. 191 (6), 1–17. Solangi, G.S., Siyal, A.A., Siyal, P., 2019. Analysis of Indus Delta groundwater and surface water suitability for domestic and irrigation purposes. Civ. Eng. J. 5 (7), 1599–1608. Srinivas, R., Singh, A.P., 2018. Application of fuzzy multi-criteria approach to assess the water quality of river Ganges. Soft Computing: Theories and Applications. Springer,, pp. 513–522. Tahir, A.A., Chevallier, P., Arnaud, Y., Neppel, L., Ahmad, B., 2011. Modeling snowmelt-runoff under climate scenarios in the Hunza River basin, Karakoram Range, Northern Pakistan. J. Hydrol. 409 (1), 104–117. Tas¸an, M., Tas¸an, S., Demir, Y., 2022. Estimation and uncertainty analysis of groundwater quality parameters in a coastal aquifer under seawater intrusion: a comparative study of deep learning and classic machine learning methods. Environ. Sci. Pollut. Res. 1–25. Tripathi, M., Singal, S.K., 2019. Use of principal component analysis for parameter selection for development of a novel water quality index: a case study of river Ganga India. Ecol. Indic. 96, 430–436. Tung, T.M., Yaseen, Z.M., 2020. A survey on river water quality modelling using artificial intelligence models: 2000–2020. J. Hydrol. 585, 124670. Velmurugan, A., Swarnam, P., Subramani, T., Meena, B., Kaledhonkar, M., 2020. Water demand and salinity. Desalin. Chall. Oppor. IntechOpen. Vickers, N.J., 2017. Animal communication: when i’m calling you, will you answer too? Curr. Biol. 27 (14), R713–R715. Wang, H.-L., Yin, Z.-Y., 2020. High performance prediction of soil compaction parameters using multi expression programming. Eng. Geol. 276, 105758. Weerts, A., Winsemius, H., Verkade, J., 2011. Estimation of predictive hydrological uncertainty using quantile regression: examples from the National Flood Forecasting System (England and Wales). Hydrol. Earth Syst. Sci. 15 (1), 255–265. Xiong, H., Wang, L., Wang, Z., 2019. Chalcogenide microlens arrays fabricated using hot embossing with soft PDMS stamps. J. Non-Cryst. Solids 521, 119542. Yaseen, Z.M., Ramal, M.M., Diop, L., Jaafar, O., Demir, V., Kisi, O., 2018. Hybrid adaptive neuro-fuzzy models for water quality index estimation. Water Resour. Manag. 32 (7), 2227–2245. Yaseen, Z.M., Sulaiman, S.O., Deo, R.C., Chau, K.-W., 2019. An enhanced extreme learning machine model for river flow forecasting: State-of-the-art, practical applications in water resource engineering area and future research direction. J. Hydrol. 569, 387–408. Zali, M.A., Retnam, A., Juahir, H., Zain, S.M., Kasim, M.F., Abdullah, B., Saadudin, S.B., 2011. Sensitivity analysis for water quality index (WQI) prediction for Kinta River. Malays. World Appl. Sci. J. 14, 60–65. Zhang, Q., Li, Z., Zhu, L., Zhang, F., Sekerinski, E., Han, J.-C., Zhou, Y., 2021. Real-time prediction of river chloride concentration using ensemble learning. Environ. Pollut. 291, 118116. Zhang, Y., Gao, X., Smith, K., Inial, G., Liu, S., Conil, L.B., Pan, B., 2019. Integrating water quality and operation into prediction of water production in drinking water treatment plants by genetic algorithm enhanced artificial neural network. Water Res. 164, 114888. A. Aldrees et al.