scieee AI-readable full text Open interactive document viewer

Jucazinho Dam Streamflow Prediction: A Comparative Analysis of Machine Learning Techniques

Silva, Erickson Johny Galindo da; Coutinho, Artur Paiva; Firmino Cardoso, Jean; Marques Bezerra, Saulo de Tarso

Abstract

[ENGLISH]Abstract: The centuries-old history of dam construction, from the Saad el-Kafara Dam to global expansion in the 1950s, highlights the importance of these structures in water resource management. The Jucazinho Dam, built in 1998, emerged as a response to the scarcity of water in the Agreste region of Pernambuco, Brazil. After having less than 1% of its water storage capacity in 2016, the dam recovered in 2020 after interventions by the local water utility. In this context, the reliability of influent flow prediction models for dams becomes crucial for managers. This study proposed hydrological models based on artificial intelligence that aim to generate flow series, and we evaluated the adaptability of these models for the operation of the Jucazinho Dam. Data normalization between 0 and 1 was applied to avoid the predominance of variables with high values. The model was based on machine learning and employed support vector regression (SVM), random forest (RF) and artificial neural networks (ANNs), as provided by the Python Sklearn library. The selection of the monitoring stations took place via the Brazilian National Water and Sanitation Agency's (ANA) HIDROWEB portal, and we used Spearman's correlation to identify the relationship between precipitation and flow. The evaluation of the performance of the model involved graphical analyses and statistical criteria such as the Nash-Sutcliffe model efficiency coefficient (NSE), the percentage of bias (PBIAS), the coefficient of determination (R²) and the root mean standard deviation ratio (RSR). The results of the statistical coefficients for the test data indicated unsatisfactory performance for long-term predictions (8, 16 and 32 days ahead), revealing a downward trend in the quality of the fit with an increase in the forecast horizon. The SVM model stood out by obtaining the best indices of NSE, PBIAS, R² and RSR. The graphical results of the SVM models showed underestimation of the flow values with an increase in the forecast horizon due to the sensitivity of the SVM to complex patterns in the time series. On the other hand, the RF and ANN models showed hyperestimation of the flow values as the number of forecast days increased, which was mainly attributed to overfitting. In summary, this study highlights the relevance of artificial intelligence in flow prediction for the efficient management of dams, especially in water scarcity and data-scarce scenarios. A proper choice of models and the ensuring of reliable input data are crucial for obtaining accurate forecasts and can contribute to water security and the effective operation of dams such as Jucazinho. [PORTUGUESE]Resumo: A história secular da construção de barragens, desde a Barragem de Saad el-Kafara até a expansão global na década de 1950, destaca a importância dessas estruturas na gestão de recursos hídricos. A Barragem de Jucazinho, construída em 1998, surgiu como resposta à escassez de água na região Agreste de Pernambuco, Brasil. Após atingir menos de 1% de sua capacidade de armazenamento em 2016, a barragem se recuperou em 2020 após intervenções da concessionária local. Nesse contexto, a confiabilidade dos modelos de previsão de vazão afluente para barragens torna-se crucial para os gestores. Este estudo propôs modelos hidrológicos baseados em inteligência artificial visando gerar séries de vazão, e avaliou a adaptabilidade desses modelos para a operação da Barragem de Jucazinho. A normalização dos dados entre 0 e 1 foi aplicada para evitar a predominância de variáveis com valores elevados. O modelo baseou-se em aprendizado de máquina e empregou regressão de vetores de suporte (SVM), floresta aleatória (RF) e redes neurais artificiais (ANNs), fornecidas pela biblioteca Python Sklearn. A seleção das estações de monitoramento ocorreu via portal HIDROWEB da Agência Nacional de Águas e Saneamento Básico (ANA), e utilizou-se a correlação de Spearman para identificar a relação entre precipitação e vazão. A avaliação do desempenho do modelo envolveu análises gráficas e critérios estatísticos como o coeficiente de eficiência de Nash-Sutcliffe (NSE), a porcentagem de viés (PBIAS), o coeficiente de determinação (R²) e a razão do desvio padrão médio da raiz (RSR). Os resultados dos coeficientes estatísticos para os dados de teste indicaram desempenho insatisfatório para previsões de longo prazo (8, 16 e 32 dias à frente), revelando uma tendência de queda na qualidade do ajuste com o aumento do horizonte de previsão. O modelo SVM destacou-se obtendo os melhores índices de NSE, PBIAS, R² e RSR. Os resultados gráficos dos modelos SVM mostraram subestimação dos valores de vazão com o aumento do horizonte de previsão devido à sensibilidade do SVM a padrões complexos nas séries temporais. Por outro lado, os modelos RF e ANN mostraram superestimação dos valores de vazão à medida que o número de dias de previsão aumentou, o que foi atribuído principalmente ao overfitting. Em resumo, este estudo destaca a relevância da inteligência artificial na previsão de vazão para a gestão eficiente de barragens, especialmente em cenários de escassez hídrica e dados escassos. A escolha adequada de modelos e a garantia de dados de entrada confiáveis são cruciais para obter previsões precisas e podem contribuir para a segurança hídrica e a operação eficaz de barragens como Jucazinho. [SPANISH]Resumen: La historia centenaria de la construcción de presas, desde la presa Saad el-Kafara hasta la expansión global en la década de 1950, destaca la importancia de estas estructuras en la gestión de recursos hídricos. La presa Jucazinho, construida en 1998, surgió como respuesta a la escasez de agua en la región Agreste de Pernambuco, Brasil. Tras tener menos del 1% de su capacidad de almacenamiento de agua en 2016, la presa se recuperó en 2020 tras intervenciones de la empresa local de aguas. En este contexto, la fiabilidad de los modelos de predicción de caudal afluente para presas se vuelve crucial para los gestores. Este estudio propuso modelos hidrológicos basados en inteligencia artificial que tienen como objetivo generar series de caudales, y evaluamos la adaptabilidad de estos modelos para la operación de la presa Jucazinho. Se aplicó la normalización de datos entre 0 y 1 para evitar el predominio de variables con valores altos. El modelo se basó en el aprendizaje automático y empleó regresión de vectores de soporte (SVM), bosque aleatorio (RF) y redes neuronales artificiales (ANN), proporcionadas por la biblioteca Python Sklearn. La selección de las estaciones de monitoreo se realizó a través del portal HIDROWEB de la Agencia Nacional de Aguas y Saneamiento Básico (ANA) de Brasil, y utilizamos la correlación de Spearman para identificar la relación entre precipitación y caudal. La evaluación del rendimiento del modelo involucró análisis gráficos y criterios estadísticos como el coeficiente de eficiencia del modelo de Nash-Sutcliffe (NSE), el porcentaje de sesgo (PBIAS), el coeficiente de determinación (R²) y la relación de la desviación estándar media de la raíz (RSR). Los resultados de los coeficientes estadísticos para los datos de prueba indicaron un rendimiento insatisfactorio para predicciones a largo plazo (8, 16 y 32 días por delante), revelando una tendencia a la baja en la calidad del ajuste con un aumento en el horizonte de pronóstico. El modelo SVM se destacó obteniendo los mejores índices de NSE, PBIAS, R² y RSR. Los resultados gráficos de los modelos SVM mostraron una subestimación de los valores de caudal con un aumento en el horizonte de pronóstico debido a la sensibilidad del SVM a patrones complejos en las series temporales. Por otro lado, los modelos RF y ANN mostraron una sobreestimación de los valores de caudal a medida que aumentaba el número de días de pronóstico, lo que se atribuyó principalmente al sobreajuste. En resumen, este estudio destaca la relevancia de la inteligencia artificial en la predicción de caudales para la gestión eficiente de presas, especialmente en escenarios de escasez de agua y datos escasos. Una elección adecuada de modelos y asegurar datos de entrada fiables son cruciales para obtener pronósticos precisos y pueden contribuir a la seguridad hídrica y la operación efectiva de presas como Jucazinho. [CHINESE]摘要:大坝建设有着数百年的历史,从萨德埃尔-卡法拉大坝到20世纪50年代的全球扩张,突显了这些结构在水资源管理中的重要性。建于1998年的Jucazinho大坝是为了应对巴西伯南布哥州阿格雷斯特地区的水资源短缺而出现的。在2016年蓄水量不足1%后,该大坝在当地水务部门干预后于2020年恢复。在此背景下,大坝入流预测模型的可靠性对管理者至关重要。本研究提出了基于人工智能的水文模型,旨在生成流量序列,并评估了这些模型对Jucazinho大坝运行的适应性。对数据进行了0到1之间的归一化处理,以避免高值变量占主导地位。该模型基于机器学习,采用了Python Sklearn库提供的支持向量回归(SVM)、随机森林(RF)和人工神经网络(ANN)。监测站的选择通过巴西国家水务和卫生局(ANA)的HIDROWEB门户进行,并使用斯皮尔曼相关性来确定降水与流量之间的关系。模型性能的评估涉及图形分析和统计标准,如纳什-萨特克利夫模型效率系数(NSE)、偏差百分比(PBIAS)、决定系数(R²)和均方根标准差比(RSR)。测试数据的统计系数结果表明,长期预测(提前8、16和32天)的性能不令人满意,显示出随着预测范围的增加,拟合质量呈下降趋势。SVM模型通过获得最佳的NSE、PBIAS、R²和RSR指数脱颖而出。SVM模型的图形结果显示,随着预测范围的增加,流量值被低估,这是由于SVM对时间序列中复杂模式的敏感性。另一方面,RF和ANN模型显示,随着预测天数的增加,流量值被高估,这主要归因于过度拟合。总之,本研究强调了人工智能在流量预测中对大坝高效管理的相关性,特别是在水资源短缺和数据稀缺的情况下。正确选择模型和确保可靠的输入数据对于获得准确的预测至关重要,并有助于水安全和Jucazinho等大坝的有效运行。 [GERMAN]Zusammenfassung: Die jahrhundertealte Geschichte des Talsperrenbaus, vom Saad el-Kafara-Staudamm bis zur globalen Expansion in den 1950er Jahren, unterstreicht die Bedeutung dieser Bauwerke für die Wasserbewirtschaftung. Der 1998 errichtete Jucazinho-Staudamm entstand als Reaktion auf die Wasserknappheit in der Region Agreste in Pernambuco, Brasilien. Nachdem der Damm 2016 weniger als 1 % seiner Wasserspeicherkapazität aufwies, erholte er sich 2020 nach Eingriffen des lokalen Wasserversorgers. In diesem Zusammenhang ist die Zuverlässigkeit von Zuflussvorhersagemodellen für Talsperren für die Betreiber von entscheidender Bedeutung. Diese Studie schlug hydrologische Modelle auf der Grundlage künstlicher Intelligenz vor, die darauf abzielen, Abflussreihen zu generieren, und wir bewerteten die Anpassungsfähigkeit dieser Modelle für den Betrieb des Jucazinho-Staudamms. Eine Datennormalisierung zwischen 0 und 1 wurde angewendet, um die Dominanz von Variablen mit hohen Werten zu vermeiden. Das Modell basierte auf maschinellem Lernen und verwendete Support Vector Regression (SVM), Random Forest (RF) und künstliche neuronale Netze (ANNs), wie sie von der Python-Bibliothek Sklearn bereitgestellt werden. die Auswahl der Überwachungsstationen erfolgte über das HIDROWEB-Portal der brasilianischen Nationalen Wasser- und Sanitäragentur (ANA), und wir verwendeten die Spearman-Korrelation, um die Beziehung zwischen Niederschlag und Abfluss zu identifizieren. Die Bewertung der Leistung des Modells umfasste grafische Analysen und statistische Kriterien wie den Nash-Sutcliffe-Modelleffizienzkoeffizienten (NSE), den prozentualen Bias (PBIAS), das Bestimmtheitsmaß (R²) und das Verhältnis der mittleren Standardabweichung (RSR). Die Ergebnisse der statistischen Koeffizienten für die Testdaten deuteten auf eine unbefriedigende Leistung bei langfristigen Vorhersagen (8, 16 und 32 Tage im Voraus) hin und zeigten einen Abwärtstrend bei der Qualität der Anpassung mit zunehmendem Vorhersagehorizont. Das SVM-Modell stach hervor, indem es die besten Indizes für NSE, PBIAS, R² und RSR erzielte. Die grafischen Ergebnisse der SVM-Modelle zeigten eine Unterschätzung der Abflusswerte mit zunehmendem Vorhersagehorizont aufgrund der Empfindlichkeit der SVM gegenüber komplexen Mustern in den Zeitreihen. Andererseits zeigten die RF- und ANN-Modelle eine Überschätzung der Abflusswerte, wenn die Anzahl der Vorhersagetage zunahm, was hauptsächlich auf Overfitting zurückzuführen war. Zusammenfassend hebt diese Studie die Relevanz künstlicher Intelligenz bei der Abflussvorhersage für das effiziente Management von Talsperren hervor, insbesondere in Szenarien mit Wasserknappheit und Datenmangel. Eine richtige Auswahl von Modellen und die Sicherstellung zuverlässiger Eingabedaten sind entscheidend für die Erzielung genauer Vorhersagen und können zur Wassersicherheit und zum effektiven Betrieb von Talsperren wie Jucazinho beitragen. [FRENCH]Résumé : L'histoire séculaire de la construction de barrages, du barrage de Saad el-Kafara à l'expansion mondiale dans les années 1950, souligne l'importance de ces structures dans la gestion des ressources en eau. Le barrage de Jucazinho, construit en 1998, est apparu comme une réponse à la pénurie d'eau dans la région de l'Agreste à Pernambuco, au Brésil. Après avoir atteint moins de 1 % de sa capacité de stockage d'eau en 2016, le barrage s'est rétabli en 2020 après des interventions du service des eaux local. Dans ce contexte, la fiabilité des modèles de prévision du débit entrant pour les barrages devient cruciale pour les gestionnaires. Cette étude a proposé des modèles hydrologiques basés sur l'intelligence artificielle visant à générer des séries de débit, et nous avons évalué l'adaptabilité de ces modèles pour l'exploitation du barrage de Jucazinho. La normalisation des données entre 0 et 1 a été appliquée pour éviter la prédominance de variables à valeurs élevées. Le modèle était basé sur l'apprentissage automatique et utilisait la régression vectorielle de support (SVM), la forêt aléatoire (RF) et les réseaux de neurones artificiels (ANN), tels que fournis par la bibliothèque Python Sklearn. La sélection des stations de surveillance a eu lieu via le portail HIDROWEB de l'Agence nationale brésilienne de l'eau et de l'assainissement (ANA), et nous avons utilisé la corrélation de Spearman pour identifier la relation entre les précipitations et le débit. L'évaluation de la performance du modèle a impliqué des analyses graphiques et des critères statistiques tels que le coefficient d'efficacité du modèle de Nash-Sutcliffe (NSE), le pourcentage de biais (PBIAS), le coefficient de détermination (R²) et le rapport de l'écart-type moyen quadratique (RSR). Les résultats des coefficients statistiques pour les données de test ont indiqué une performance insatisfaisante pour les prévisions à long terme (8, 16 et 32 jours à l'avance), révélant une tendance à la baisse de la qualité de l'ajustement avec une augmentation de l'horizon de prévision. Le modèle SVM s'est distingué en obtenant les meilleurs indices de NSE, PBIAS, R² et RSR. Les résultats graphiques des modèles SVM ont montré une sous-estimation des valeurs de débit avec une augmentation de l'horizon de prévision en raison de la sensibilité du SVM aux modèles complexes dans les séries temporelles. D'autre part, les modèles RF et ANN ont montré une surestimation des valeurs de débit à mesure que le nombre de jours de prévision augmentait, ce qui a été principalement attribué au surapprentissage. En résumé, cette étude souligne la pertinence de l'intelligence artificielle dans la prévision du débit pour la gestion efficace des barrages, en particulier dans les scénarios de pénurie d'eau et de rareté des données. Un choix approprié de modèles et la garantie de données d'entrée fiables sont cruciaux pour obtenir des prévisions précises et peuvent contribuer à la sécurité de l'eau et à l'exploitation efficace de barrages tels que Jucazinho.

Full text

5.93.2 Jucazinho Dam Streamflow Prediction: A Comparative Analysis of Machine Learning Techniques Erickson Johny Galindo da Silva, Artur Paiva Coutinho, Jean Firmino Cardoso and Saulo de Tarso Marques Bezerra Article https://doi.org/10.3390/hydrology11070097 Citation: Silva, E.J.G.d.; Coutinho, A.P.; Cardoso, J.F.; Bezerra, S.d.T.M. Jucazinho Dam Streamflow Prediction: A Comparative Analysis of Machine Learning Techniques. Hydrology 2024,11, 97. https:// doi.org/10.3390/hydrology11070097 Academic Editor: Andrea Petroselli Received: 3 April 2024 Revised: 11 June 2024 Accepted: 11 June 2024 Published: 4 July 2024 Copyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). hydrology Article Jucazinho Dam Streamflow Prediction: A Comparative Analysis of Machine Learning Techniques Erickson Johny Galindo da Silva , Artur Paiva Coutinho * , Jean Firmino Cardoso and Saulo de Tarso Marques Bezerra Agreste Campus, Federal University of Pernambuco, Av. Marielle Franco, Km 59, Caruaru 55014-900, Brazil; [email protected] (E.J.G.d.S.); [email protected] (J.F.C.); [email protected] (S.d.T.M.B.) *Correspondence: arthur[email protected] Abstract: The centuries-old history of dam construction, from the Saad el-Kafara Dam to global expansion in the 1950s, highlights the importance of these structures in water resource management. The Jucazinho Dam, built in 1998, emerged as a response to the scarcity of water in the Agreste region of Pernambuco, Brazil. After having less than 1% of its water storage capacity in 2016, the dam recovered in 2020 after interventions by the local water utility. In this context, the reliability of influent flow prediction models for dams becomes crucial for managers. This study proposed hydrological models based on artificial intelligence that aim to generate flow series, and we evaluated the adaptability of these models for the operation of the Jucazinho Dam. Data normalization between 0 and 1 was applied to avoid the predominance of variables with high values. The model was based on machine learning and employed support vector regression (SVM), random forest (RF) and artificial neural networks (ANNs), as provided by the Python Sklearn library. The selection of the monitoring stations took place via the Brazilian National Water and Sanitation Agency’s (ANA) HIDROWEB portal, and we used Spearman’s correlation to identify the relationship between precipitation and flow. The evaluation of the performance of the model involved graphical analyses and statistical criteria such as the Nash–Sutcliffe model efficiency coefficient (NSE), the percentage of bias (PBIAS), the coefficient of determination (R 2 ) and the root mean standard deviation ratio (RSR). The results of the statistical coefficients for the test data indicated unsatisfactory performance for long-term predictions (8, 16 and 32 days ahead), revealing a downward trend in the quality of the fit with an increase in the forecast horizon. The SVM model stood out by obtaining the best indices of NSE, PBIAS, R 2 and RSR. The graphical results of the SVM models showed underestimation of the flow values with an increase in the forecast horizon due to the sensitivity of the SVM to complex patterns in the time series. On the other hand, the RF and ANN models showed hyperestimation of the flow values as the number of forecast days increased, which was mainly attributed to overfitting. In summary, this study highlights the relevance of artificial intelligence in flow prediction for the efficient management of dams, especially in water scarcity and data-scarce scenarios. A proper choice of models and the ensuring of reliable input data are crucial for obtaining accurate forecasts and can contribute to water security and the effective operation of dams such as Jucazinho. Keywords: support vector machine; random forest; artificial neural network; hydrological modeling; rainfall; flow; forecasting 1. Introduction Many archaeologists consider the Saad el-Kafara dam to be one of the first in the world. It was probably built during the reign of Khufu, who was the king of Egypt around 2900–2877 B.C. [ 1 ]. Since then, there has been a long tradition of dam construction that spans millennia and with various purposes, such as flood control, providing water for human consumption, irrigation and animal thirst and, more recently in the history of dams, for industrial purposes and electricity generation. Hydrology 2024,11, 97. https://doi.org/10.3390/hydrology11070097 https://www.mdpi.com/journal/hydrology Hydrology 2024,11, 97 2 of 20 As we arrived in the 1950s, which was a period of great expansion of global populations and economies, dams began to be increasingly considered as a solution to meet the growing demands for water and energy. Since then, according to the World Commission on Dams (WCD) [ 2 ], at least 45,000 large dams have been built worldwide, with almost half of the world’s rivers having at least one large dam in their course. In the Brazilian city of Surubim in the state of Pernambuco, the Jucazinho Dam was inaugurated in 1998; it is located in the Capibaribe watershed and bars the river that is also called the Capibaribe. It was built to minimize the water scarcity in the rural region of Pernambuco and to control floods on the Capibaribe River. According to the Pernambuco Water and Climate Agency (APAC) [ 3 ], the dam has an accumulation capacity of 204.82 million cubic meters of water at an elevation of 292 m from the main spillway crest, and, as recorded by the Brazilian National Water and Sanitation Agency (ANA) [ 4 ], it has 92% of the total withdrawal demand for water for human supply, 7% for animal thirst, and 1% for irrigation. This amount of water is relevant due to the large number of municipalities it serves: a total of 15 cities. The Hydroenvironmental Plan for the Capibaribe Hydrographic Basin of the Secretariat of Water Resources of the State of Pernambuco highlights the economic potential of the municipalities served by the dam, nine of which belong to the second largest textile and clothing pole in Brazil, three with relevant agricultural activities, two belonging to the furniture and tourism pole and one with a strong thermal tourism industry. Even with the efforts of the Pernambuco Sanitation Company (COMPESA), which is responsible for the operation of the Jucazinho Dam, to combat the prolonged drought, the dam’s water level dropped to below 0.01% in 2016 and only recovered four years later, in 2020, surpassing 1% [ 5 ]. At the height of the water crisis, the company implemented a rotation in the supply, providing the population with potable water only seven days per month [6]. To address such challenges and improve the management of water resources, researchers have increasingly turned to advanced computational methods. In recent decades, several artificial intelligence models have been used by researchers to predict influentary flows due to the great accuracy and flexibility of considering physical processes with all their characteristics. Among the methods used, artificial neural network (ANN) [ 7 – 9 ], random forest (RF) [10–12] and support vector machine (SVM) [13–15] models stands out. Based on the literature, there are models that are better suited depending on the watershed. Adnan et al. [ 13 ] conducted an evaluation of various models, including the Optimally Pruned Extreme Learning Machine (OP-ELM), Least Square Support Vector Machine (LSSVM), Multivariate Adaptive Regression Splines (MARS) and M5 Model Tree (M5Tree), for modeling monthly streamflows with precipitation and temperature inputs. Their findings highlighted the superiority of LSSVM and MARS-based models for streamflow prediction without the need for local input data, surpassing the OP-ELM and M5Tree models. Parisouj et al. [ 8 ] investigated the predictive accuracy of three renowned machine learning algorithms—Support Vector Regression (SVR), ANN and Extreme Learning Machine—for monthly and daily streamflows across four rivers in the United States. The study identified SVR as the most effective model at both the monthly and daily scales, whereas the ANN model exhibited the least satisfactory performance. Meshram et al. [ 9 ] compared the efficacy of three AI techniques—Adaptive Neuro Fuzzy Inference System (ANFIS), Genetic Programming and ANN—for forecasting streamflow within India’s Shakkar watershed. The findings underscored that models incorporating cyclic terms outperformed those that did not consider periodicity and relied solely on previous streamflow data. Sousa Jr. et al. [ 12 ] assessed K-Nearest Neighbor, SVM, RF and ANN models for daily streamflow prediction in a transitional region between the Savanna and Amazon biomes in Brazil. The results demonstrated that these models achieved promising streamflow predictions for up to three days ahead, even in basins with scarce hydrological data. Hydrology 2024,11, 97 3 of 20 The motivation of this study is the delivery of a hydrological model for reservoir management based on emerging artificial intelligence techniques, adding to the literature machine learning adaptations for the prediction of tributary flows and optimizing the operation of water reservoirs. Therefore, the objective of the present study is to generate synthetic series of tributary flows for the Jucazinho Dam, which is located in the Agreste region of Pernambuco, from stochastic models. By evaluating the adaptability of machine learning models using SVM, RF and ANN to generate synthetic series of tributary flows to the dam, analyzing the influence on the quality of adjustments with the increase in the number of forecast days and verifying the quality of the adjustments based on statistical metrics, we determine the best model for the hydrological variables under study. 2. Materials and Methods 2.1. Study Area Located in the Brazilian city of Surubim in the state of Pernambuco in the Capibaribe watershed, the Jucazinho Dam bars the river that is also called the Capibaribe, as we can see in the map of the situation in Figure 1. Its construction began in 1995 and was completed in 1998, and at the time, there was the expropriation of more than two thousand hectares in the riverside areas, involving about 5000 people [16]. Figure 1. Situation map. One of the reasons for the construction of the dam was the scenario of scarcity of water supply in the rural region of Pernambuco. With the construction of Jucazinho, which has a maximum storage capacity of 245.26 million cubic meters of water at the maximum maximorum elevation of 295 m, 21 municipalities could be served, which impacted the lives of approximately 800 thousand inhabitants. As stated in Figure 2, 92% of the dam’s use is for water for human supply. In addition, the dam is used for fish farming, livestock and agricultural. The other reason for the construction of the dam was for flood control planning in Capibaribe, which involved the construction of several dams in order to protect the metropolitan region of Recife (about 135 km away) from historical floods such as those that occurred between the years 1960 and 1980. The dam has a flood control volume of 106m³. Hydrology 2024,11, 97 4 of 20 Figure 2. Total withdrawal demands. Its construction type is gravity with roller-compacted concrete, and it has a central stepped spillway with a ski-jump-type dissipation basin. There are also two side spillways connected to the dam’s abutment. It contains a gallery with access at the two abutments and which extends over the entire embankment of the dam. Above the central spillway is a bridge that connects the abutments. It has a water intake for supply and another for the release of the ecological flow downstream, both with a pipe of 2.0 m in diameter and a reduction to 1.5 m. In Table 1, we provide data from the dam’s technical file, with information provided by Department of Water Resources of Pernambuco and data measured by Neves et al. [17]. 2.2. Input Data The “garbage in, garbage out” principle refers to the fundamental idea that the quality of the output of a data processing system is directly influenced by the quality of the input data. In other words, if inaccurate, incomplete or inadequate information is fed into a system, it is inevitable that the resulting output will also be inaccurate or of poor quality. This principle highlights the critical importance of reliable and high-quality data entry to ensure accurate and useful results in any computing or decision-making process. The data for this study were obtained from the HIDROWEB portal: an online platform by ANA that offers information on Brazil’s water resources. The portal provides real-time data from a vast network of monitoring stations across the country, covering hydrometeorological, hydrographic and water quality aspects. In the Jucazinho Dam’s basin, 32 rainfall monitoring stations and two river flow monitoring stations were identified, as shown in Figure 3. Table 1. Technical data sheet of the Jucazinho Dam. Name Dimension Unit Embankment Latitude 07°57′59.39′′ S - Longitude 35°44′3.16′′ W - Incremental drainage area 2865.60 km2 Total drainage area 4149.90 km2 Maximum volume 204.82 hm3 Minimum volume 0.29 hm3 Usable volume 326.75 hm3 Maximum operating water level 292.00 m Minimum operating water level 253.00 m Elevation of the bottom of the lake 238.00 m Hydrology 2024,11, 97 5 of 20 Table 1. Cont. Name Dimension Unit Crest Length 442.00 m Width 8.00 m Elevation 299.00 m Main Spillway Length 170.00 m Crest elevation 292.00 m Distance between spillway and embankment crest 7.00 m Maximum flow rate 5446.69 m3·s−1 Side Spillways Length 57.00 m Crest elevation 295.00 m Maximum flow rate 1291.30 m3·s−1 Gallery Length 2.00 m Width 2.00 m Elevation 250.00 m Maximum flow rate 2.72 m3·s−1 To choose stations for the study, those with records within the same time frame were initially selected. Fluviometric stations 39100000 and 39130000 had records matching with pluviometric stations 735159, 736040, 736041, 736042 and 836092 and were the most recently updated and were thus chosen for the study. A key aspect of hydrological modeling is the correlation between rainfall and flow data. Rainfall drives surface runoff and groundwater recharge, directly affecting river flow levels. Spearman’s correlation ( ρ ), a robust non-parametric measure, was used to assess this correlation, as shown in Equation (1). It ranges from − 1 to 1 and indicates negative (ρ< 0), positive (ρ> 0) or no correlation (ρ= 0). ρ=1− 6∑n i=1di2 n(n2−1)!(1) where di and n are, respectively, the difference in ranks between the original series and the series sorted in ascending order for the i-th observation and the total number of observations. The highest correlation was observed between pluviometric station 736042 and fluviometric station 39130000, as demonstrated in Figure 4, with a correlation coefficient of 0.18, leading to their selection for the study. Despite the low correlation, this research addresses a real situation where the case study (Jucazinho Dam) was selected by the study funder (COMPESA). The choice was due to its importance to the region, its history of “collapse”, and its operation based on the technicians’ empiricism. Precisely due to data limitations, this study can contribute to the literature by evaluating artificial intelligence models in real situations with scarce data. The general data of both stations are contained in Table 2. Due to the beginning and end of both series, the records from 1 January 1986 to 1 June 2023 were used to develop the hydrological model. Hydrology 2024,11, 97 6 of 20 Figure 3. Pluviometric and fluviometric stations in the Jucazinho Dam’s catchment area. Figure 4. Correlation matrix graph. Regarding the flow data, to fill the faults of station 39130000, data from station 39100000 were used and were multiplied by a factor of 1.57, referring to the ratio between the drainage area of 2450 km 2 of station 39130000 and the drainage area of 1560 km 2 of station 39100000, which were obtained through data from ANA’s HIDROWEB portal. It is noteworthy that both stations are located on the Capibaribe River, which is the main river of the Capibaribe Hydrographic Basin, which is barred by the Jucazinho Dam, and that the number of faults is insignificant when compared to the total series, allowing, without major damage to the model, this type of fault filling. The rest of the missing data, both flow and precipitation, for stations 736042 and 39130000, were interpolated linearly. The series with the gaps filled is presented in Figure 5 and covers a total of 13,668 days. Initially, stations 736042 and 39130000 presented, respectively, 0.61% and 5.38% of failures in this number of days. After using the data from station 39100000 multiplied by the factor of 1.57, the percentage of failures from station 39130000 dropped to 4.15%. Finally, both series presented 13,668 days of records without failures. It should be noted that linear interpolation is not the appropriate methodology for filling in daily failures of a precipitation or flow series, but specifically because we were filling for a period with low precipitation and flow, with values equal to zero or close to it, it was acceptable to apply the method without the association of significant errors in the results of the filled series. Hydrology 2024,11, 97 7 of 20 Table 2. General data of the pluviometric and fluviometric stations used. Information Pluviometric Station Fluviometric Station Station Name Taquaritinga do Norte Toritama Code 736042 39130000 Basin 3—Atlantic, NW/NE section 3—Atlantic, NW/NE section Sub-basin 39—Capibaribe, Ipojuca, Una, Goiana, Mundaú, Paraíba do Meio, Coruripe, Sirinhaém, São Miguel and Camaragibe Rivers 39—Capibaribe, Ipojuca, Una, Goiana, Mundaú, Paraíba do Meio, Coruripe, Sirinhaém, São Miguel and Camaragibe Rivers City Taquaritinga do Norte Santa Cruz do Capibaribe State Pernambuco Pernambuco Accountable ANA ANA Operator Geological Survey of Brazil (CPRM) CPRM Latitude −7.9039 −8.0128 Longitude 36.0469 −36.0578 Elevation (m) 785 376 Drainage area (km2) - 2450 Distance to Jucazinho Dam (km) 34.31 35.16 Start of the series 1 January 1986 1 January 1973 End of the series 30 June 2023 1 June 2023 Series size (years) 36.5 49.5 It also should be noted that fluviometric station 391300000 is about 35.16 km from the Jucazinho Dam embankment and 14.64 km from the Jucazinho inundation area, and side spillway crest elevations are about 295 m. Therefore, the flow series generated by the artificial intelligence models trained from the data presented in Figure 5should be multiplied by a factor of 1.69 for a practical application of reservoir management. This factor refers to the ratio between the drainage area of 4149.90 km 2 of the Jucazinho Dam and the drainage area of 2450.00 km2of station 39130000. Figure 5. Data series used in this study. Hydrology 2024,11, 97 8 of 20 Figure 6presents the average of the monthly accumulations of all the years of the catchment basin. It is verified that there is little rainfall in the area, with approximately four months of rain and eight months of drought, indicating that the Capibaribe watershed has no hydrological memory. The hydrological memory of a watershed represents its ability to store and release water over time in response to climatic conditions. It is influenced by factors such as geology and land use, and basins with permeable soils tend to have greater hydrological memory. Hydrological models face challenges in basins without hydrological memory because they may have less predictable responses, impairing the model’s ability to capture anomalous climate events, given that temporal variability in water retention and release directly influences the hydrological response. Figure 6. Average of the cumulative monthly index of all the years in the series. 2.3. Model Construction For the development of the model, the input variables are presented in Table 3; the model predicts a sequence of next steps from a sequence of past observations. As Q(t−1) represents the flow rate for the time prior to Q(t) , the delayed flows of 1, 2, 3, ..., 32 days with respect to t are called, respectively, Q(t−1) , Q(t−2) , ..., Q(t−32) ; likewise, the flows with an advance of 1, 2, 3, ..., 32 days in relation to t are called, respectively, Q(t+1) , Q(t+2) , ..., Q(t+32). The same nomenclature logic is used for precipitation. Thus, in the first scenario, the prediction of the flow for the next day was made from the previous data of one day of flow and precipitation (C-1). In the second scenario, previous data from two days of flow and precipitation were considered in order to predict the next two days of flow (C-2). In the other scenarios, the same logic was used but with 4 (C-4), 8 (L-8), 16 (L-16) and 32 days of data (L-32). The models were grouped into: • (Group C): Short-term prediction for 1 (C-1), 2 (C-2) and 4 (C-4) days of prediction; • (Group L): Long-term prediction for 8 (L-8), 16 (L-16) and 32 days (L-32) of prediction. The models were constructed in an orderly manner without mixing the chronological order of the pairs containing the input and output variables and with recursive, repeating values in these pairs. Using model C-1 as an example, we show the pairs in Table 4. Note that P(t−2) , Q(t−2) and Q(t+1) are repeated in both the input and output sets. In this way, historical values are used both to predict future flow values and to provide information about past patterns that influence these predictions. In a short-term context, usually covering periods of up to a week, flow prediction is essential for immediate decision-making. This includes real-time control of water flow, flood prevention, and reservoir water level management. The ability to anticipate intense weather events or sudden changes in hydrological conditions allows for rapid responses, such as the controlled release of water to prevent flooding or the immediate adjustment of the stored volume to meet current demand. Hydrology 2024,11, 97 15 of 20 Table 9. Cont. R2 SVM RF ANN 65–35 70–30 75–25 80–20 65–35 70–30 75–25 80–20 65–35 70–30 75–25 80–20 C-1 0.54 0.83 0.91 0.94 0.70 0.87 0.85 0.82 0.73 0.92 0.92 0.93 C-2 0.40 0.75 0.85 0.91 0.66 0.85 0.80 0.69 0.71 0.87 0.88 0.79 C-4 0.27 0.63 0.78 0.86 0.57 0.63 0.72 0.56 0.64 0.76 0.77 0.71 L-8 0.16 0.47 0.68 0.80 0.50 0.50 0.62 0.48 0.53 0.60 0.70 0.59 L-16 0.08 0.30 0.52 0.70 0.37 0.35 0.50 0.20 0.41 0.37 0.57 0.38 L-32 0.01 0.17 0.41 0.57 0.21 0.05 0.03 –1.81 0.24 0.15 0.41 0.00 RSR SVM RF ANN 65–35 70–30 75–25 80–20 65–35 70–30 75–25 80–20 65–35 70–30 75–25 80–20 C-1 0.67 0.41 0.30 0.25 0.55 0.36 0.38 0.42 0.52 0.28 0.28 0.27 C-2 0.78 0.50 0.38 0.30 0.58 0.38 0.44 0.55 0.54 0.35 0.34 0.46 C-4 0.86 0.61 0.47 0.37 0.66 0.61 0.53 0.66 0.60 0.49 0.48 0.54 L-8 0.92 0.73 0.56 0.44 0.70 0.71 0.61 0.72 0.69 0.63 0.55 0.64 L-16 0.96 0.84 0.69 0.55 0.80 0.80 0.71 0.90 0.77 0.79 0.65 0.79 L-32 0.99 0.91 0.77 0.66 0.89 0.97 0.97 1.68 0.87 0.92 0.76 1.00 Table 10. Better statistical criteria for training. C-N NSE PBIAS R2RSR C-1 SVM 80-20 SVM 75-25 SVM 80-20 SVM 80-20 C-2 SVM 80-20 RF 70-30 SVM 80-20 SVM 80-20 C-4 SVM 80-20 SVM 80-20 SVM 80-20 SVM 80-20 The results of the models are presented in Figures 7–12. The models took 90% of the code execution time to train, which means that the computational cost of training is significantly higher than what is required to use the trained model. In the SVM models, with the increase in the number of days, there was underestimation of the flow forecast values (points approaching the x-axis). Despite being effective at modeling nonlinear relationships, the SVM model is sensitive to complex patterns in time series. For this reason, as the forecast horizon increased, the model was unable to predict the significant changes in flow patterns, as can be seen in Figure 7. The AF and ANN models had the opposite results of the SVM model: with the increase in days, there was an overestimation of the flow prediction values (points approaching the y-axis), as can be seen in Figures 8and 9. This is mainly due to overfitting, where the models have adjusted too much to the training data and incorporated transient noise and patterns that are not representative of the actual flow behavior. Han et al. [ 22 ] applied SVM to the Bird Creek watershed for flood prediction and found that, like ANN models, SVM also suffers from underfitting and overfitting problems, with overfitting being more harmful than underfitting. The study also reveals an interesting result in the response of the SVM to different rainfall inputs, where lighter rains generated very different responses than more intense rainfall, similar to what occurred in the present work. Hydrology 2024,11, 97 16 of 20 Figure 7. Scatter plots of the tests for SVM 80-20. On the graph of the readings over time shown in Figure 10, it is possible to verify the underestimation of the peaks mentioned above for the SVM models. The C-1 model was able to predict peaks better than the L-32 model. It is also verified that there is a delay between the graphs of the observed and measured flow for the L-32 model during the increase in the number of pairs of input and output data points. This delay explains the low statistical coefficients because the accumulated error is summed. According to Figures 11 and 12, the RF models presented similar results as those of the ANN: generating noise that oscillated much above the measured flow. The underlying relationship between the features and the target variable is highly nonlinear and complex; that is why both RF and ANN struggled to capture it effectively, leading to instability and noisy predictions. As with the SVM model, there is also a delay between the graphs of the observed and measured flow for the L-32 model during the increase in the number of pairs of input and output data points. Figure 8. Scatter plots of tests for RF 80-20. The ANN models were able to represent the flow peaks well, as Figure 12 shows. Above all, generated noise oscillated the values that should be zero between values slightly higher or lower than zero, even predicting negative flow values. As with both models presented above, there is also a delay between the graphs of the observed and measured flow for the L-32 model during the increase in the number of pairs of input and output data points. Hydrology 2024,11, 97 17 of 20 Figure 9. Scatter plots of tests for ANN 80-20. Figure 10. Comparison between observed and simulated flows for SVM 80-20. Figure 11. Comparison between observed and simulated flows for RF 80-20. Hydrology 2024,11, 97 18 of 20 The results confirm that the Capibaribe watershed does not have hydrological memory. It is recommended in future works using the application of SVM for different watersheds to verify the adaptability of the model according to the hydrological memory of each basin. Figure 12. Comparison between observed and simulated flows for ANN 80-20. 4. Conclusions The general objective of this research was to generate synthetic series of tributary flows to the Jucazinho Dam, which is located in the Agreste region of Pernambuco, based on stochastic models. The specific objectives were to evaluate the adaptability of SVM, RF and ANN machine learning models to generate the synthetic flow series, to analyze the influence of the quality of the adjustments with the increase in the number of days of the forecast, and to verify the quality of the adjustments made by the models using statistical metrics. Based on the results obtained and the discussions presented, it is concluded that: • All models showed satisfactory performance for short-term prediction, which includes 1, 2 and 4 days, and unsatisfactory for long-term prediction, which includes 8, 16 and 32 days. • The graphical results of the SVM models showed underestimation of the flow values with an increase in the forecast horizon due to the sensitivity of the SVM to complex patterns in the time series. • On the other hand, the RF and ANN models showed hyperestimation of the flow values as the number of forecast days increased, which was mainly attributed to overfitting. • For all models, an increase in the number of prediction days led to a tendency to decrease the quality of the adjustment; this was mainly justified as due to the delay in the predictions, which generated an accumulation of errors. • According to the statistical criteria, the model that best adapted to the series was SVM, which resulted in the best statistical indices of NSE, PBIAS, R2and RSR. • Even in situations where data are scarce, artificial intelligence models SVM, RF and ANN have the potential to be applied for short-term prediction. • The Capibaribe watershed does not have hydrological memory, which impacted model training. It is recommended in future works using the application of ANNs in different watersheds to verify the adaptability of the model according to the hydrological memory of each basin. Hydrology 2024,11, 97 19 of 20 Author Contributions: Conceptualization, E.J.G.d.S. and A.P.C.; Formal analysis, A.P.C. and S.d.T.M.B.; Investigation, E.J.G.d.S.; Methodology, E.J.G.d.S.; Project administration, S.d.T.M.B.; Software, E.J.G.d.S. and J.F.C.; Supervision, A.P.C. and S.d.T.M.B.; Validation, A.P.C.; Writing—original draft, E.J.G.d.S.; Writing—review and editing, E.J.G.d.S. and S.d.T.M.B. All authors have read and agreed to the published version of the manuscript. Funding: This research was funded by the Academic Master’s and Doctoral Program for Innovation (MAI/DAI), which was promoted by the Brazilian National Council for Scientific and Technological Development (Conselho Nacional de Desenvolvimento Científico e Tecnológico—CNPq, Brazil), in a partnership between the Federal University of Pernambuco (Universidade Federal de Pernambuco—UFPE, Brazil) and the Pernambuco Sanitation Company (Companhia Pernambucana de Saneamento—COMPESA, Brazil). Data Availability Statement: The data presented in this study are available on request from the corresponding authors. Acknowledgments: The authors would like to thank Vinnycius Luz and Milton Melo Neto from the Pernambuco Sanitation Company (Companhia Pernambucana de Saneamento—COMPESA, Brazil), the National Council for Scientific and Technological Development (Conselho Nacional de Desenvolvimento Científico e Tecnológico—CNPq, Brazil) for the productivity scholarships for Artur Coutinho [process 315927/2021-6] and Saulo Bezerra [process 308202/2022-8], the Foundation for Support of Science and Technology of the State of Pernambuco (Fundação de Amparo à Ciência e Tecnologia de Pernambuco—FACEPE, Brazil) [process APQ-1767-3.01/22], and the Coordination for the Improvement of Higher Education Personnel (Coordenação de Aperfeiçoamento de Pessoal de Nível Superior—CAPES, Brazil) [Finance Code 001]. Conflicts of Interest: The authors declare no conflicts of interest. References 1. Jansen, R.B. Dams and Public Safety: A Water Resources Technical Publication; United States Printing Office: Denver, CO, USA, 1980. 2. World Commission on Dams. Dams and Development: A New Framework for Decision-Making: The Report of the World Commission on Dams; Earthscan: Oxford, UK, 2000. 3. Agência Pernambucana de Águas e Clima-Apac. Available online: https://acesse.one/hyipu (accessed on 20 November 2022). 4. Agência Nacional das Águas-Ana. Reservatórios do Semiárido Brasileiro: Hidrologia, Balanço Hídrico e Operação; ANA: Brasília, Brazil, 2017; Volume Anexo E; 178 p. 5. Companhia Pernambucana de Saneamento-Compesa. Available online: https://l1nq.com/RCLDM (accessed on 20 November 2022 ). 6. Santana, R.A.; Bezerra, S.T.M.; Santos, S.M.; Coutinho, A.P.; Coelho, I.C.L.; Pessoa, R.S.V. Assessing alternatives for meeting water demand: a case study of water resource management in the Brazilian Semiarid region. Util. Policy 2019,61, 100974. 7. Ali, S.; Shahbaz, M. Streamflow forecasting by modeling the rainfall–streamflow relationship using artificial neural networks. Model. Earth Syst. Environ. 2020, 6, 1645–1656. [CrossRef] 8. Parisouj, P.; Mohebzadeh, H.; Lee, T. Employing machine learning algorithms for streamflow prediction: a case study of four river basins with different climatic zones in the United States. Water Resour. Manag. 2020,34, 4113–4131. [CrossRef] 9. Meshram, S.G.; Meshram, C.; Santos, C.A.G.; Benzougagh, B.; Khedher, K.M. Streamflow prediction based on artificial intelligence techniques. Iran. J. Sci. Technol. Trans. Civ. Eng. 2022,46, 2393–2403. [CrossRef] 10. Sun, N.; Zhang, S.; Peng, T.; Zhang, N.; Zhou, J.; Zhang, H. Multi-variables-driven model based on random forest and Gaussian process regression for monthly streamflow forecasting. Water 2022,14, 1828. [CrossRef] 11. Islam, K.I.; Elias, E.; Carroll, K.C.; Brown, C. Exploring random forest machine learning and remote sensing data for streamflow prediction: An alternative approach to a process-based hydrologic modeling in a snowmelt-driven watershed. Remote Sens. 2023, 15, 3999. [CrossRef] 12. de Sousa, M.F., Jr.; Uliana, E.M.; Aires, R.V.; Rápalo, L.M.; da Silva, D.D.; Moreira, M.C.; Lisboa, L.; da Silva Rondon, D. Streamflow prediction based on machine learning models and rainfall estimated by remote sensing in the Brazilian Savanna and Amazon biomes transition. Model. Earth Syst. Environ. 2024,10, 1191–1202. [CrossRef] 13. Adnan, R.M.; Liang, Z.; Heddam, S.; Kermani, M.; Kisi, O.; Li, B. Least square support vector machine and multivariate adaptive regression splines for streamflow prediction in mountainous basin using hydro-meteorological data as inputs. J. Hydrol. 2020,586, 124371. [CrossRef] 14. Essam, Y.; Huang, Y.F.; Ng, J.L.; Birima, A.H.; Ahmed, A.N.; El-Shafie, A. Predicting streamflow in Peninsular Malaysia using support vector machine and deep learning algorithms. Sci. Rep. 2022,12, 3883. [CrossRef] [PubMed] 15. Ikram, R.M.A.; Hazarika, B.B.; Gupta, D.; Heddam, S.; Kisi, O. Streamflow prediction in mountainous region using new machine learning and data preprocessing methods: a case study. Neural Comput. Appl. 2023,35, 9053–9070. [CrossRef] Hydrology 2024,11, 97 20 of 20 16. Girão, L.C.P. Uma Análise da Contribuição dos Programas Básicos Ambientais Como Instrumento de Gestão Ambiental Para a Barragem de Jucazinho Localizada no Município de Surubim/PE. Master’s Thesis, Universidade Federal de Pernambuco-UFPE, Recife, Brazil, 2004. 17. Neves, Y.T.; Rodrigues, A.; Cabral, J.J.S.P. Modelagem computacional do rompimento hipotético da barragem de Jucazinho no estado de Pernambuco (Brasil). Rev. DAE 2021,69, 167–182. [CrossRef] 18. Nash, J.E.; Sutcliffe, J.V. River flow forecasting through conceptual models part I-A discussion of principles. J. Hydrol. 1970,10, 282–290. [CrossRef] 19. Moriasi, D.N.; Arnold, J.G.; van Liew, M.W.; Bingner, R.L.; Harmel, R.D.; Veith, T.L. Model evaluation guidelines for systematic quantification of accuracy in watershed simulations. Am. Soc. Agric. Biol. Eng. 2007,50, 885–900. 20. Cheng, M.; Fang, F.; Kinouchi, T.; Navon, I.M.; Pain, C.C. Long lead-time daily and monthly streamflow forecasting using machine learning methods. J. Hydrol. 2020,590, 125376. [CrossRef] 21. Al-Mukhtar, M. Random forest, support vector machine, and neural networks to modelling suspended sediment in Tigris River-Baghdad. Environ. Monit. Assess. 2019,191, 673. [CrossRef] [PubMed] 22. Han, D.; Chan, L.; Zhu, N. Flood forecasting using support vector machines. J. Hydroinform. 2007,9, 267–276. [CrossRef] Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.