Full text
water Article Possibilities of Using Neuro-Fuzzy Models for Post-Processing of Hydrological Forecasts Tomas Kozel 1,2,*, Tomas Vlasak 3and Petr Janal 1 Citation: Kozel, T.; Vlasak, T.; Janal, P. Possibilities of Using Neuro-Fuzzy Models for Post-Processing of Hydrological Forecasts. Water 2021, 13, 1894. https://doi.org/10.3390/ w13141894 Academic Editor: Marco Franchini Received: 11 June 2021 Accepted: 6 July 2021 Published: 8 July 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). 1Czech Hydrometeorological Institute, Kroftova 43, 616 00 Brno, Czech Republic; petr[email protected] 2Faculty of Civil Engineering, Institute of Landscape Water Management, Brno University of Technology, Veveˇrí331/95, 602 00 Brno, Czech Republic 3Czech Hydrometeorological Institute, A. Staška 1177, Rožnov, 370 07 ˇ CeskéBudˇejovice, Czech Republic; [email protected] *Correspondence: [email protected].cz Abstract: When issuing hydrological forecasts and warnings for individual profiles, the aim is to achieve the best possible results. Hydrological forecasts themselves are burdened by an error (uncertainty) at the inputs (precipitation forecast) as well as on the side of the hydrological model used. The aim of the method described in this article is to reduce the error of the hydrological model using post-processing the model results. Models based on neuro-fuzzy models were selected for the post-processing itself. The whole method was tested on 12 profiles in the Czech Republic. The catchment size of the individual profiles ranged from 90 to 4500 km 2 and the profiles varied in their character, both in terms of elevation as well as land cover. After finding the suitable model architecture and introducing supporting algorithms, there was an improvement in the results for the individual profiles for selected criteria by on average 5–60% (relative culmination error, mean square error) compared to the results of re-simulation of the hydrological model. The results of the application show that the method was able to improve the accuracy of hydrological forecasts and thus could contribute to better management of flood situations. Keywords: hydrological forecast; floods; artificial intelligence methods; post-processing 1. Introduction Short-term hydrological forecasts are an essential part of many early warning systems today. Based on the forecast, it is possible to adjust the management of reservoirs or start preparatory measures against floods. The problem of hydrological forecasts is their accuracy. The forecast itself is affected by uncertainties both on the input side (precipitation, river basin saturation, etc.) and on the hydrological model side (schematization and calibration of the model). The errors caused by the above-mentioned uncertainties are partly systematic (repeated). They can be identified, but often cannot be removed by modifying the hydrological model, because the model cannot capture the whole complexity of the runoff process [ 1 ]. However, these errors can be eliminated by one of the methods of adjusting the resulting forecasts—in so-called post-processing. For the above reason, post-processing is commonly introduced in predictions to reduce prediction error. In practice, commonly used post-processing methods usually only deal with the error associated with the precipitation amount or with hydrological uncertainty, but only with stochastic (ensemble) prediction [ 2 – 5 ]. Other methods that work directly with the hydrological model error work on the basis of daily flows [ 6 ] or longer time periods. In the case of other methods such as the Kalman filter [ 7 ], the methods are usually limited to repeated correction of the first value and subsequent adjustment of the prediction (shift of values by the magnitude of the error, time shift). The problem with their use is that the effect of adjusting the first value decreases with increasing time. In addition, repeated recalculations of the model are very time consuming. In the case of the Czech Republic, Water 2021,13, 1894. https://doi.org/10.3390/w13141894 https://www.mdpi.com/journal/water
Water 2021,13, 1894 2 of 15 the forecast length of 72 h is usually used (Czech Hydrometeorological Institute, CHMI). For this reason, a method was developed that is able to correct the error of the hydrological model during a time period of 72 h or more in advance and does not need recalculation (recalibration). One of the possibilities of post-processing the outputs of the hydrological model are trained neuro-fuzzy models [ 8 ]. Trained neuro-fuzzy models are an adaptive-networkbased fuzzy inference systems (ANFIS). The advantage of ANFIS is the fact that it represents a combination of neural networks and a fuzzy set, which ANFIS gives certain advantage over classic neural network or trained fuzzy models. For further understanding of ANFIS we recommend a scientific paper of the authors Jang and Sun [ 8 ]. The further described model is based on this theory. The advantage of the above models is their calculation speed and the ability to approximate even very complex relationships between the individual factors influencing the resulting course of hydrological forecast. The artificial intelligence methods themselves are very often used for flood flow forecast or reservoir management [9–12] . Usage of ANFIS as the post-processing method is what makes this study unique. The aim of the method described in this paper is to be able to apply a model, based on the ANFIS method, for post-processing and it is therefore assumed that the reader has certain understanding of the ANFIS method. The ANFIS itself could be used for flood flow forecasting, but this application has some restrictions and usually has a high demand on sufficient training episodes. The major issue associated with a direct application is related to extrapolation (when forecasted episodes are outside of the training data matrix). The second problem is that a flood is a continuous curve, values of which usually do not differ from hour to hour by a certain difference (results stability problems). The AI method can give rather fluctuating values, without another algorithm. These problems are solved by using a hydrological model. The ANFIS is used for post-processing of hydrological model results to enhance their accuracy. Moreover, the ANFIS is also intended to be used in flood situations only (when it makes sense to apply the post-processing). The hydrological model forecast is provided on a daily basis even when flow values are low (drought forecast). The uncertainty of hydrological modeling is caused, among other things, by the simplification of reality on a spatial scale (homogenization of river basins into modeled areas) and on a time scale (same set of parameters for different runoff phases). It can be assumed that for a specific river basin, this part of the uncertainty of hydrological modeling will manifest itself by systematic errors, which are repeated in similar runoff situations. The magnitude of these biases depends on many factors such as saturation, initial flow, precipitation, season and others. Therefore, a simple mathematical relationship cannot be established for them. However, trained neuro-fuzzy models can identify (evaluate) these relationships and use them to make the forecast more accurate. The aim of the described post-processing models based on trained neuro-fuzzy models is to reduce the error of the outputs from the hydrological model for events that were caused by rainfall. In the following text, hydrological forecast will be understood as the deterministic forecast obtained by the continuous hydrological model for gauged (forecasting) profile. The presented method focuses on runoff situations in which there is a risk of floods and thus a considerable pressure to provide the greatest possible reliability of flow forecasts. At the same time, however, we have to emphasize that the method does not solve the errors caused by inaccurate predictions of precipitation, which are often the dominant source of errors during floods. Figure 1shows a flowchart for the process and also shows where the post-processing stage comes into play during this process.
Water 2021,13, 1894 3 of 15 Water 2021, 13, x FOR PEER REVIEW 3 of 15 Figure 1. Scheme of important process. 2. Materials and Methods It is, therefore, clear from the previous text that the use of neuro-fuzzy models as a post-processing tool could bring the desired improvement of the existing hydrological forecast methods since the model itself, based on artificial intelligence methods, will be able to mitigate the prediction error. In the model architecture of the trained neuro-fuzzy model NFM, sequential aggregation of inputs was used [13] if the model used 3 or more inputs (Figure 2). Figure 3 depicts the model scheme without sequential aggregation of inputs. The advantage of a gradual aggregation of inputs is easier construction of the rule matrix. First, it is necessary to build a model architecture and a matrix of target behavior containing input-output relationships, because models based on artificial intelligence methods need patterns for their training (learning). Table 1 below lists the individual inputs that were used to find a suitable post-processing model architecture. All flow values are modeled flows by the hydrological model, unless stated otherwise for the flow parameter. Table 1. Parameters used for model construction. Value Marking Unit Instantaneous flow value Qsim m3/s Total precipitation amount (1 h) Hs Mm Previous flow value Qpm m3/s Next flow value Qfm m3/s Topsoil saturation indicator UTWZ Mm Average of the hourly sum of precipitation totals over a period of time Hspi Mm Corrected flow (post-procesing flow) Qpost m3/s Difference between current and previous flow value ΔQmodp m3/s Difference between current and next flow value ΔQmodf m3/s Average of the flow values over time Qpmodi m3/s Figure 2. Model scheme using sequential aggregation of inputs. Figure 3. Model scheme without using sequential aggregation of inputs. Figure 1. Scheme of important process. 2. Materials and Methods It is, therefore, clear from the previous text that the use of neuro-fuzzy models as a post-processing tool could bring the desired improvement of the existing hydrological forecast methods since the model itself, based on artificial intelligence methods, will be able to mitigate the prediction error. In the model architecture of the trained neuro-fuzzy model NFM, sequential aggregation of inputs was used [ 13 ] if the model used 3 or more inputs (Figure 2). Figure 3depicts the model scheme without sequential aggregation of inputs. The advantage of a gradual aggregation of inputs is easier construction of the rule matrix. First, it is necessary to build a model architecture and a matrix of target behavior containing input-output relationships, because models based on artificial intelligence methods need patterns for their training (learning). Table 1below lists the individual inputs that were used to find a suitable post-processing model architecture. All flow values are modeled flows by the hydrological model, unless stated otherwise for the flow parameter. Water 2021, 13, x FOR PEER REVIEW 3 of 15 Figure 1. Scheme of important process. 2. Materials and Methods It is, therefore, clear from the previous text that the use of neuro-fuzzy models as a post-processing tool could bring the desired improvement of the existing hydrological forecast methods since the model itself, based on artificial intelligence methods, will be able to mitigate the prediction error. In the model architecture of the trained neuro-fuzzy model NFM, sequential aggregation of inputs was used [13] if the model used 3 or more inputs (Figure 2). Figure 3 depicts the model scheme without sequential aggregation of inputs. The advantage of a gradual aggregation of inputs is easier construction of the rule matrix. First, it is necessary to build a model architecture and a matrix of target behavior containing input-output relationships, because models based on artificial intelligence methods need patterns for their training (learning). Table 1 below lists the individual inputs that were used to find a suitable post-processing model architecture. All flow values are modeled flows by the hydrological model, unless stated otherwise for the flow parameter. Table 1. Parameters used for model construction. Value Marking Unit Instantaneous flow value Qsim m3/s Total precipitation amount (1 h) Hs Mm Previous flow value Qpm m3/s Next flow value Qfm m3/s Topsoil saturation indicator UTWZ Mm Average of the hourly sum of precipitation totals over a period of time Hspi Mm Corrected flow (post-procesing flow) Qpost m3/s Difference between current and previous flow value ΔQmodp m3/s Difference between current and next flow value ΔQmodf m3/s Average of the flow values over time Qpmodi m3/s Figure 2. Model scheme using sequential aggregation of inputs. Figure 3. Model scheme without using sequential aggregation of inputs. Figure 2. Model scheme using sequential aggregation of inputs. Water 2021, 13, x FOR PEER REVIEW 3 of 15 Figure 1. Scheme of important process. 2. Materials and Methods It is, therefore, clear from the previous text that the use of neuro-fuzzy models as a post-processing tool could bring the desired improvement of the existing hydrological forecast methods since the model itself, based on artificial intelligence methods, will be able to mitigate the prediction error. In the model architecture of the trained neuro-fuzzy model NFM, sequential aggregation of inputs was used [13] if the model used 3 or more inputs (Figure 2). Figure 3 depicts the model scheme without sequential aggregation of inputs. The advantage of a gradual aggregation of inputs is easier construction of the rule matrix. First, it is necessary to build a model architecture and a matrix of target behavior containing input-output relationships, because models based on artificial intelligence methods need patterns for their training (learning). Table 1 below lists the individual inputs that were used to find a suitable post-processing model architecture. All flow values are modeled flows by the hydrological model, unless stated otherwise for the flow parameter. Table 1. Parameters used for model construction. Value Marking Unit Instantaneous flow value Qsim m3/s Total precipitation amount (1 h) Hs Mm Previous flow value Qpm m3/s Next flow value Qfm m3/s Topsoil saturation indicator UTWZ Mm Average of the hourly sum of precipitation totals over a period of time Hspi Mm Corrected flow (post-procesing flow) Qpost m3/s Difference between current and previous flow value ΔQmodp m3/s Difference between current and next flow value ΔQmodf m3/s Average of the flow values over time Qpmodi m3/s Figure 2. Model scheme using sequential aggregation of inputs. Figure 3. Model scheme without using sequential aggregation of inputs. Figure 3. Model scheme without using sequential aggregation of inputs. Table 1. Parameters used for model construction. Value Marking Unit Instantaneous flow value Qsim m3/s Total precipitation amount (1 h) Hs Mm Previous flow value Qpm m3/s Next flow value Qfm m3/s Topsoil saturation indicator UTWZ Mm Average of the hourly sum of precipitation totals over a period of time Hspi Mm Corrected flow (post-procesing flow) Qpost m3/s Difference between current and previous flow value ∆Qmodp m3/s Difference between current and next flow value ∆Qmodf m3/s Average of the flow values over time Qpmodi m3/s 3. Application The model described in the previous chapter was applied to selected prediction profiles to test the possibilities of post-processing under various conditions. The area
Water 2021,13, 1894 4 of 15 of interest was chosen because there are very diverse types of river basins. There are typical mountain catchments and catchments strongly influenced by agricultural activity as well as catchments with a high proportion of water areas. The profiles themselves were selected from a set of forecasting profiles included in the national flood forecasting service performed by the Czech Hydrometeorological Institute (CHMI). All the selected profiles are located within the Vltava river basin in the Czech Republic in the region of south Bohemia. Their positions are shown in Figure 4and their brief description is given in Table 2 . A total of 12 profiles were selected, which were divided into three groups according to their location. At the end of each profile name, a number is given in parentheses, which represents the position of the profile in Figure 3. The first group consists of the catchments of the Lenora (2), Liˇcov (3) and ˇ CeskéBudˇejovice (1) profiles. The profile of ˇ CeskéBudˇejovice is in this application the final profile of the whole first group and the area of its catchment area represents the whole area of the catchment area of the first group. The catchments of the Liˇcov and Lenora profiles are then sub-catchments. The basin of the Liˇcov and Lenora profiles mostly consists of forests and mountains. The profile of Lenor itself lies above the Lipno reservoir. The second group consists of the profiles Bechynˇe (4) and Rodvínov (5). The Bechynˇe profile is the final profile of the whole group. The Bechynˇe profile catchment area and its partial part of the Rodvínov catchment area are relatively flat with a large percentage of agricultural land and include a large number of ponds and pond systems (5.1% of the catchment area). The third group consists of the profiles Bohumilice (11), Katowice (7), Modrava (9), Písek (6), Podedvory (12), Stod˚ulky (10) and Sušice (8). The Písek profile is the final profile of the entire third group. The catchment area of the Bohumilice, Modrava, Stod˚ulky and Sušice profiles is mainly made up of forests. The catchments of the Bohumilice, Modrava and Stod˚ulky profiles themselves have a mountain character. Water 2021, 13, x FOR PEER REVIEW 4 of 15 3. Application The model described in the previous chapter was applied to selected prediction profiles to test the possibilities of post-processing under various conditions. The area of interest was chosen because there are very diverse types of river basins. There are typical mountain catchments and catchments strongly influenced by agricultural activity as well as catchments with a high proportion of water areas. The profiles themselves were selected from a set of forecasting profiles included in the national flood forecasting service performed by the Czech Hydrometeorological Institute (CHMI). All the selected profiles are located within the Vltava river basin in the Czech Republic in the region of south Bohemia. Their positions are shown in Figure 4 and their brief description is given in Table 2. A total of 12 profiles were selected, which were divided into three groups according to their location. At the end of each profile name, a number is given in parentheses, which represents the position of the profile in Figure 3. The first group consists of the catchments of the Lenora (2), Ličov (3) and České Budějovice (1) profiles. The profile of České Budějovice is in this application the final profile of the whole first group and the area of its catchment area represents the whole area of the catchment area of the first group. The catchments of the Ličov and Lenora profiles are then sub-catchments. The basin of the Ličov and Lenora profiles mostly consists of forests and mountains. The profile of Lenor itself lies above the Lipno reservoir. The second group consists of the profiles Bechyně (4) and Rodvínov (5). The Bechyně profile is the final profile of the whole group. The Bechyně profile catchment area and its partial part of the Rodvínov catchment area are relatively flat with a large percentage of agricultural land and include a large number of ponds and pond systems (5.1% of the catchment area). The third group consists of the profiles Bohumilice (11), Katowice (7), Modrava (9), Písek (6), Podedvory (12), Stodůlky (10) and Sušice (8). The Písek profile is the final profile of the entire third group. The catchment area of the Bohumilice, Modrava, Stodůlky and Sušice profiles is mainly made up of forests. The catchments of the Bohumilice, Modrava and Stodůlky profiles themselves have a mountain character. Figure 4. Locations of river profiles and their catchments. Table 2 shows the catchment area and the average long-term flow Q a . Furthermore, for the individual profiles, values of flows are given, at which the individual degrees of flood activity are announced, which are defined by the Water Act of the Czech Republic Figure 4. Locations of river profiles and their catchments.
Water 2021,13, 1894 5 of 15 Table 2. Profile specification. Num. Profile River Catchment Area [km2] Qa [m3/s] 1. FLD [m3/s] 2. FLD [m3/s] 3. FLD [m3/s] N1 [m3/s] N100 [m3/s] FIP [%] 1ˇ CeskéBudejovice Vltava 2850 106 244 361 489 172 908 35 2 Lenora Tepla Vltava 177 3.06 29 53.5 70.8 26 113 78 3 Liˇcov Cerna 127 1.29 12.7 21.4 29.7 21 188 53 4 Bechynˇe Luznice 4057 22.2 87.9 140 187 111 577 33 5 Rodvínov Nezarka 297 2.2 18.7 26.9 43.7 20 91 10 6 Písek Otava 2914 23.4 135 214 297 146 837 27 7 Katovice Otava 1134 13.8 118 169 255 133 510 42 8 Sušice Otava 533 10.5 64.2 94.3 127 101 369 77 9 Modrava Vydra 90 3.01 30.5 41.2 54.6 29 120 96 10 Stod˚ulky Kremelna 135 3.24 24 37 52.4 40 153 89 11 Bohumilice Spulka 105 0.97 24 35.1 47.7 11 84 48 12 Podedvory Blanice 204 2.04 15.3 28.2 38.6 25 165 60 Table 2shows the catchment area and the average long-term flow Q a . Furthermore, for the individual profiles, values of flows are given, at which the individual degrees of flood activity are announced, which are defined by the Water Act of the Czech Republic 254/2001 of the Collection of Laws. The values of the flood level danger (FLD) levels themselves represent the individual limit values for the declaration of a flood danger (FLD 1—low level of danger—yellow color; FLD 2—high level of danger-orange color; FLD 3—extreme level of danger—red color). The particular level of danger usually corresponds with this color set in a global warning system. The individual values of flows in the FLD columns were determined according to the measurement curve valid at the time of validation of the post-processing model. Column N1 shows the one-year flood flow rate value and column N100 shows the 100-year flood flow rate value. The last column of the table shows the percentage of forest area with respect to the total catchment area FIP. The FIP parameter is used for better understanding of the basin. Input data of model flow, soil saturation and average precipitation data for each river basin were obtained by using the Aqualog [ 14 ] hydrological model, which is used in this part of the river basin for operational calculation of the hydrological forecast. The Aqualog is a hydrological modeling system that integrates the rainfall-runoff model Sacramento (SAC-SMA) [ 15 ], the snow accumulation and melting model Snow17 [ 16 ] and the channel routing model TDR [ 17 ]. The spatial structure of modeling techniques is semi-distributed with an average area of the catchments between 30 and 50 km 2 . The required inputs are data on river discharge, precipitation and air temperature in 1-h step. The precipitation data, which were used for resimulation, were merged with rainfall data (combination of adjusted radar observation and values measured by automatic weather station). Aqualog re-simulation data were calculated on the observed data in continuous intervention-free operation. Deviations of the calculated flow from the observed values therefore represent the error of the input observed data and hydrological modeling. The re-simulation period 2005–2016 includes a number of significant flood episodes. The data were grouped into individual months and the tendency of the simulation model compared to the reality in the individual months was determined. Data from a particular month were selected for the training of the model for the that particular month, and if the neighboring months had the same tendency (underestimated, overestimated), the training matrix was extended by these data. Data of distant months were not considered, even if they had the same tendency. The reason for this selection was the time variability of the influence of the model sensitivity on the individual input parameters of the simulation model and the variable boundary conditions of the simulation model. Data from neighboring months that met the condition of the same tendency were included only due to the lack of training data during the individual months.
Water 2021,13, 1894 6 of 15 During the calibration process, optimal settings and architecture of the individual models were sought. For the NFM model, it was the shape number of membership functions further marked mf and sequence of input. The calibration period was chosen from 2005 to 2010 and was gradually extended by new (past) events. Individual events were selected for the calibration process on the basis of their peak flow and total precipitation Hs. The event started when the hourly precipitation Hs exceeded 1 mm and ended 15 h (35 h catchments exceeding 250 km 2 ) from the last value of the precipitation Hs, which was higher than 0. From the events defined above, those events were selected in the target behavior matrix, for which the peak flow exceeded the half the flow value of the FLD 1 of the given profile basin. In the first phase of the calibration process, boundaries were sought for the results obtained by the NFM models, and therefore, the NFM models were trained over the entire data period and then verified on the validation period. Criterion Ewas introduced to assess the success rate of the model: E= n ∑ i=1 (Qreal −Qsim)2(1) which is defined as the average sum of squares of the deviations between the simulated flow Q sim and the actual flow Q real for the observed period. The results showed that the NFM models showed a lower error of criterion Eby 20 to 50% compared to the results of unadjusted (without post-processing) outputs of the hydrological model. These results set the limits that can be achieved in this study when NFM models are used for post-processing. It is possible to achieve very different values of criterion Efor selected episodes. For the very need of the actual operation and validation, an auxiliary algorithm was compiled for scenarios where the modified event was outside the limits of the training area. Commonly used methods for extending the limits of training data in AI methods such as the addition of white noise have not been able to sufficiently capture the range of future events being modified. In many cases, the peak flow was more than two times greater than the peak flows in the training data. This algorithm is referred to in the following text as EEA (extreme events algorithm). Before starting the post-processing itself (model training phase), the EEA algorithm verifies whether the modeled culmination (modeled by the hydrological model) is higher than the historically modeled events. If the first condition is met, the EEA algorithm finds the event with the highest peak in the calibration data. The value of the current simulated culmination is then increased by 25%. Such increase ensures a sufficiently large space for training. It also calculates the difference between the adjusted current peak value and the peak value of the selected calibration event. Subsequently, the hourly gradients for the individual simulated flows for the selected calibration event ∆ Qka are calculated and the culmination value is determined in time. The values of the rising limb (including the peak flow) are obtained according to the Equation (2) and the values of the falling limb of the artificial simulated event according to the Equation (3): Qvvi=Qsim,i+3i×∆Qkai 3c(2) where iis the element order and cis the value of the peak flow order. Qsvi=Qsim,i+2k×∆Qkai 2c(3) where k=c+(c−i)and cis the number of peak members in the time series. (4) The algorithm then calculates the ratio between the simulated and real values for each member of the series of the selected calibration event. The real values of the artificial event are obtained by multiplying the simulated values of the artificial event by the vector described above. According to the procedure described above, the EEA algorithm is able to
Water 2021,13, 1894 7 of 15 prevent misleading of the NFM model. Models based on AI are usually not able to provide sufficient results outside their range of training (they are not appropriate for extrapolation). The best results in the validation stage were achieved using an architecture that used the current value of Q sim as input the gradient between the current and previous value ∆ Qmodp and the average of hourly precipitation total Hspi for the selected time period Γ [h]. The architecture described above will hereafter be referred to as the NFM 1 model. Figure 5shows a diagram of the NFM 1 model and Figure 6shows one of the local model (N-F). Water 2021, 13, x FOR PEER REVIEW 7 of 15 provide sufficient results outside their range of training (they are not appropriate for extrapolation). The best results in the validation stage were achieved using an architecture that used the current value of Qsim as input the gradient between the current and previous value ΔQmodp and the average of hourly precipitation total Hspi for the selected time period Γ [h]. The architecture described above will hereafter be referred to as the NFM 1 model. Figure 5 shows a diagram of the NFM 1 model and Figure 6 shows one of the local model (N-F). The learning process itself took place in two phases. In the first phase, the architecture of the fuzzy model of Sugeno [18,19] type was constructed using the method of fuzzy Cmean clustering [20]. This step determines the value of the input and output, value and shape of membership function (mf) of the model and sets rule matrix. Then, training was performed on selected episodes using the backpropagation method and the use of the built-in function of the MATLAB anfis program [21]. The best results were obtained when the value of mf was set to 5 (Gaussian curve) for all local N-F models. Figure 5. Schema of model NFM. Figure 6. Simplified schema of model local model NFM 1. For the NFM model, an analysis of the influence of the length of the selected period of the Hspi parameter on the results was performed. The results of the analysis are shown in Table 3. The length of the chosen period for the calculation of Hspi is referred to as Γ in the following text. It can be seen from Table 3 that the length values used to generate the Hspi parameter generally reach higher values for larger river basins. This statement corresponds to the length of the river basin's reaction time to previous precipitation. Table 3. Values of Γ, for which the lowest values of the selected criteria were reached. Num. Basin 1 2 3 4 5 6 7 8 9 10 11 12 Γ [h] 6 2 2 4 3 6 6 5 3 2 2 3 Area [km2] 2848 177 127 4057 297 2914 1134 534 90 135 105 203 While testing the effect of length on the Γ parameter, the choice and number of training events were also tested. Based on the testing results, it was found that better results Figure 5. Schema of model NFM. Water 2021, 13, x FOR PEER REVIEW 7 of 15 provide sufficient results outside their range of training (they are not appropriate for extrapolation). The best results in the validation stage were achieved using an architecture that used the current value of Qsim as input the gradient between the current and previous value ΔQmodp and the average of hourly precipitation total Hspi for the selected time period Γ [h]. The architecture described above will hereafter be referred to as the NFM 1 model. Figure 5 shows a diagram of the NFM 1 model and Figure 6 shows one of the local model (N-F). The learning process itself took place in two phases. In the first phase, the architecture of the fuzzy model of Sugeno [18,19] type was constructed using the method of fuzzy Cmean clustering [20]. This step determines the value of the input and output, value and shape of membership function (mf) of the model and sets rule matrix. Then, training was performed on selected episodes using the backpropagation method and the use of the built-in function of the MATLAB anfis program [21]. The best results were obtained when the value of mf was set to 5 (Gaussian curve) for all local N-F models. Figure 5. Schema of model NFM. Figure 6. Simplified schema of model local model NFM 1. For the NFM model, an analysis of the influence of the length of the selected period of the Hspi parameter on the results was performed. The results of the analysis are shown in Table 3. The length of the chosen period for the calculation of Hspi is referred to as Γ in the following text. It can be seen from Table 3 that the length values used to generate the Hspi parameter generally reach higher values for larger river basins. This statement corresponds to the length of the river basin's reaction time to previous precipitation. Table 3. Values of Γ, for which the lowest values of the selected criteria were reached. Num. Basin 1 2 3 4 5 6 7 8 9 10 11 12 Γ [h] 6 2 2 4 3 6 6 5 3 2 2 3 Area [km2] 2848 177 127 4057 297 2914 1134 534 90 135 105 203 While testing the effect of length on the Γ parameter, the choice and number of training events were also tested. Based on the testing results, it was found that better results Figure 6. Simplified schema of model local model NFM 1. The learning process itself took place in two phases. In the first phase, the architecture of the fuzzy model of Sugeno [ 18 , 19 ] type was constructed using the method of fuzzy C-mean clustering [ 20 ]. This step determines the value of the input and output, value and shape of membership function (mf) of the model and sets rule matrix. Then, training was performed on selected episodes using the backpropagation method and the use of the built-in function of the MATLAB anfis program [ 21 ]. The best results were obtained when the value of mf was set to 5 (Gaussian curve) for all local N-F models. For the NFM model, an analysis of the influence of the length of the selected period of the Hspi parameter on the results was performed. The results of the analysis are shown in Table 3. The length of the chosen period for the calculation of Hspi is referred to as Γ in the following text. Table 3. Values of Γ, for which the lowest values of the selected criteria were reached. Num. Basin 1 2 3 4 5 6 7 8 9 10 11 12 Γ[h] 6 2 2 4 3 6 6 5 3 2 2 3 Area [km2]2848 177 127 4057 297 2914 1134 534 90 135 105 203 It can be seen from Table 3that the length values used to generate the Hspi parameter generally reach higher values for larger river basins. This statement corresponds to the length of the river basin’s reaction time to previous precipitation.
Water 2021,13, 1894 8 of 15 While testing the effect of length on the Γ parameter, the choice and number of training events were also tested. Based on the testing results, it was found that better results were achieved by selecting particular training episodes rather than using all the data. When choosing whether to use the training episode, two basic criteria were decisive. The first of these was the simulated peak flow, which should not differ by more than 25% from the value of the simulated peak flow of the adjusted episode. If the total number of training episodes was less than three, the 25% limit was gradually increased until at least one training episode was found. The second criterion was the maximum sum of the moving average of the total precipitation of a length three. If for the selected training episodes according to the first selection criterion, the maximum moving average of the three values differed by more than 60% from the adjusted episode and the total number of episodes was higher than five, episodes that did not meet the criterion were excluded. Figure 7 shows a comparison between an application of post-processing to a selected episode using all training episodes and using a selection of training episodes. Figure 8shows training results for select episode June 2013 Bechynˇe (including artificial episode for training made by EEA). Water 2021, 13, x FOR PEER REVIEW 8 of 15 were achieved by selecting particular training episodes rather than using all the data. When choosing whether to use the training episode, two basic criteria were decisive. The first of these was the simulated peak flow, which should not differ by more than 25% from the value of the simulated peak flow of the adjusted episode. If the total number of training episodes was less than three, the 25% limit was gradually increased until at least one training episode was found. The second criterion was the maximum sum of the moving average of the total precipitation of a length three. If for the selected training episodes according to the first selection criterion, the maximum moving average of the three values differed by more than 60% from the adjusted episode and the total number of episodes was higher than five, episodes that did not meet the criterion were excluded. Figure 7 shows a comparison between an application of post-processing to a selected episode using all training episodes and using a selection of training episodes. Figure 8 shows training results for select episode June 2013 Bechyně (including artificial episode for training made by EEA). Figure 7. Results for the selected episode in the Bechyně profile. Figure 7 shows the course of the selected episode, where the dashed blue line shows the flow value N1 (period of return = 1) and the green value N100 (period of return = 100). The measured data are plotted in orange. The re-simulation of the episode using the Aqualog model is shown in gray. The results are shown in pink with the use of post-processing without the use of the EEA algorithm and the selection of episodes was not used during the training. The results of post-processing using the EEA algorithm during training and without selection of individual episodes are displayed in blue. The yellow color shows the results using post-processing using the EEA algorithm during training and using the selection of episodes. 050 100 150 200 250 0 100 200 300 400 500 600 700 Time [h] Q [m3/s] Bechyně profile (June 2013) Observed data Aqualog model NFM+EEA+filtred data NFM+EEA+all data NFM+all data N1 N100 Figure 7. Results for the selected episode in the Bechynˇe profile. Figure 7shows the course of the selected episode, where the dashed blue line shows the flow value N1 (period of return = 1) and the green value N100 (period of return = 100). The measured data are plotted in orange. The re-simulation of the episode using the Aqualog model is shown in gray. The results are shown in pink with the use of post-processing without the use of the EEA algorithm and the selection of episodes was not used during the training. The results of post-processing using the EEA algorithm during training and without selection of individual episodes are displayed in blue. The yellow color shows the results using post-processing using the EEA algorithm during training and using the selection of episodes.
Water 2021,13, 1894 9 of 15 Water 2021, 13, x FOR PEER REVIEW 9 of 15 Figure 8. Calibration results for the selected episode in the Bechyně profile. Figure 8 plots the training data (blue) and the result of the trained post-processing model (red). An artificial training episode provided by the EEA algorithm is also shown (black circles). 4. Results To better evaluate the results, an Ek criterion was introduced, which is defined as the average relative deviation between the actual value of the peak flow Qp,i and the predicted value Qs,i for the entire validation period, where the individual members are averaged in absolute values. 𝐸𝑘=𝑎𝑏𝑠𝑄,−𝑄, 𝑄, (5) The Nash–Suctliffe criterion (NSE) was introduced to further evaluate the benefits of post-processing: 𝑁𝑆𝐸=1−∑𝑄,−𝑄, ∑𝑄,−𝑄 (6) The root mean square error (RMSE) was chosen as the last criterion: 𝑅𝑀𝑆𝐸=∑𝑄,−𝑄, 𝑛 (7) The individual results were further divided according to the N-year culmination into two groups (N ≤ 1 and N > 1). First, the overall results were evaluated and these were then evaluated separately for each group, because it was necessary to assess how the model performs under situations of different level of seriousness. The Ek criterion was introduced because the issuance of alerts is very often governed by the peak flow value. Table 4 shows the criteria values for all profiles together for the whole validation period for the episodes when the above-mentioned post-processing and the values obtained for the data without post-processing (clean data) were used. The Ek and NSE values are averaged results for the whole validation period. The criterion E is the sum of the individual results of the criteria E for whole validation period. Figure 8. Calibration results for the selected episode in the Bechynˇe profile. Figure 8plots the training data (blue) and the result of the trained post-processing model (red). An artificial training episode provided by the EEA algorithm is also shown (black circles). 4. Results To better evaluate the results, an Ek criterion was introduced, which is defined as the average relative deviation between the actual value of the peak flow Q p,i and the predicted value Q s,i for the entire validation period, where the individual members are averaged in absolute values. Ek =abs Qp,i−Qs,i Qp,i!(5) The Nash–Suctliffe criterion (NSE) was introduced to further evaluate the benefits of post-processing: NSE =1−∑n i=1(Qsim,i−Qreal,i)2 ∑n i=1Qreal,i−Qreal2(6) The root mean square error (RMSE) was chosen as the last criterion: RMSE =s∑n i=1(Qsim,i−Qreal,i)2 n(7) The individual results were further divided according to the N-year culmination into two groups (N ≤ 1 and N > 1). First, the overall results were evaluated and these were then evaluated separately for each group, because it was necessary to assess how the model performs under situations of different level of seriousness. The Ek criterion was introduced because the issuance of alerts is very often governed by the peak flow value. Table 4shows the criteria values for all profiles together for the whole validation period for the episodes when the above-mentioned post-processing and the values obtained for the data without post-processing (clean data) were used. The Ek and NSE values are averaged results for the whole validation period. The criterion Eis the sum of the individual results of the criteria Efor whole validation period.