scieee AI-readable full text Open interactive document viewer

APPLICATION OF BIG DATA TECHNOLOGY TO THE FIELD OF HYDROGEOLOGY AND ENGINEERING GEOLOGY WITH A LOGICAL-STRUCTURAL APPROACH

Djumanov J.X., Sayfullaeva N.A.

Abstract

At present, the core challenges in the fields of geology, hydrogeology, engineering geology, irrigation and land reclamation, ecology, and environmental monitoring center around obtaining objective data about the environment and ensuring its accurate and timely processing. In the study of the Earth's surface layer, particularly the component referred to as the hydrosphere, the integration of advanced computer networks, big data technologies, neural networks, and artificial intelligence methods-including models, algorithms, and software tools-plays a pivotal role. This is especially critical for extensive urban agglomerations characterized by high population density and encompassing major groundwater drinking resources, mineral reserves, primary irrigation and reclamation water systems, as well as large-scale industrial and hydrotechnical complexes. In recent years, there has been a growing interest in the theoretical and scientific foundations for the development of information-analytical systems and methods in these disciplines. This trend is driven, on the one hand, by the reduction of ground-based observation network nodes in hydrogeology, engineering geology, ecology, and environmental monitoring, and on the other hand, by the rapid advancement and increasing application of information and communication technologies and software solutions in these areas.

Full text

INTERNATIONAL SCIENTIFIC JOURNAL SCIENCE AND INNOVATION SPECIAL ISSUE “MODERN PROBLEMS AND PROSPECTS FOR THE DEVELOPMENT OF DIGITAL TRANSFORMATION IN ENERGY” SEPTEMBER 24, 2025 91 APPLICATION OF BIG DATA TECHNOLOGY TO THE FIELD OF HYDROGEOLOGY AND ENGINEERING GEOLOGY WITH A LOGICAL-STRUCTURAL APPROACH Djumanov J.X., Sayfullaeva N.A. Tashkent University of Information Technologies named after Muhammad al-Khwarizmi https://doi.org/10.5281/zenodo.17558355 At present, the core challenges in the fields of geology, hydrogeology, engineering geology, irrigation and land reclamation, ecology, and environmental monitoring center around obtaining objective data about the environment and ensuring its accurate and timely processing. In the study of the Earth's surface layer, particularly the component referred to as the hydrosphere, the integration of advanced computer networks, big data technologies, neural networks, and artificial intelligence methods-including models, algorithms, and software tools-plays a pivotal role. This is especially critical for extensive urban agglomerations characterized by high population density and encompassing major groundwater drinking resources, mineral reserves, primary irrigation and reclamation water systems, as well as large-scale industrial and hydrotechnical complexes. In recent years, there has been a growing interest in the theoretical and scientific foundations for the development of information-analytical systems and methods in these disciplines. This trend is driven, on the one hand, by the reduction of ground-based observation network nodes in hydrogeology, engineering geology, ecology, and environmental monitoring, and on the other hand, by the rapid advancement and increasing application of information and communication technologies and software solutions in these areas. The integration of various topological and cartographic networks within geological disciplines has been examined through the application of invariant and quasi-differential transformation theories, as well as linear and nonlinear filtering techniques. Moreover, diverse methods of intelligent processing of big data-based on both deterministic and statistical approaches-have been explored. Particular attention has been given to the development of algorithms grounded in the theory of syntactic analysis for image-based recognition of cartographic information. Additionally, the operational principles of neurocomputers, their core characteristics, and their applications in image interpretation have been studied in detail. In the development of models, algorithms, and information-analytical systems, the author has drawn upon extensive experience accumulated over many years of scientific and pedagogical activity in the fields of computer networks and systems, as well as the architecture of computing machines and complexes. The primary directions for the application of big data and intelligent data analysis methods include obtaining objective and real-time information on hydrogeology, engineering geology, environmental conditions, and the rational use of natural resources. These approaches are also essential for the monitoring of natural and technogenic hazardous situations and disasters. At the same time, information on natural resources and the results of their intelligent processing not only influence socio-economic relations but also hold significant value for scientific advancement, technological progress, as well as for addressing strategic military and political objectives. INTERNATIONAL SCIENTIFIC JOURNAL SCIENCE AND INNOVATION SPECIAL ISSUE “MODERN PROBLEMS AND PROSPECTS FOR THE DEVELOPMENT OF DIGITAL TRANSFORMATION IN ENERGY” SEPTEMBER 24, 2025 92 As an intangible resource, information possesses certain advantages; compared to other resources-particularly material ones-it requires minimal costs for recording, storage, digital processing, analytical evaluation, the formulation of proposals and recommendations for decisionmaking, and for transportation (i.e., transmission). Furthermore, it is not subject to restrictions on initial copying or use. Natural information resources are measured, recorded, or generated with the direct participation of observers and are disseminated to state cadaster agencies and higher-level institutions according to established regulatory frameworks. The advancement of computer and network technologies, the emergence of the Internet of Things (IoT), and cloud services, along with the development of automated algorithms for data collection and analysis, have fundamentally transformed the processes of information circulation and significantly increased the volume and diversity of data involved. The appearance and accumulation of vast and structurally heterogeneous datasets have led to the emergence of a new phenomenon known as Big Data. The era of the "Big Data" revolution has begun; however, with the advent of digital capabilities for measuring, recording, and storing data, the uncontrolled accumulation of information has led to significant challenges. Due to the vast memory required by unstructured data, compression and long-term archival storage are generally not recommended. These issues have become increasingly acute, particularly affecting the efficiency of working with such data volumes and, consequently, the interests of large corporations and the effective monitoring and management of critical natural resources. Traditional methods of data recording, storage, and processing are constrained by limited memory capacity and outdated management mechanisms, and their modernization requires substantial financial investment. As a response to these challenges, Big Data technologies and their capabilities have attracted considerable interest, not only in the medical field but also in the domains of natural resource management, ecology, and environmental protection. Governmental bodies and international organizations alike are showing growing engagement in exploring and implementing these technologies. However, in the management of natural resources, working with big data necessitates an entirely different approach to the processing of unstructured and unsystematized information. In general, Big Data technology must fulfill the following core functions: − Digitalization and automation of data acquisition and recording processes; − Cleansing of data arrays from irrelevant or redundant information; − Digital processing and systematization of large-scale datasets; − Intelligent analysis of complex and heterogeneous data structures; − Ensuring access to the entire volume of continuously changing data in real time; − Guaranteeing the integrity and protection of information as a unified whole. It is important to note that among all the aforementioned functions, the analysis of continuously updated data holds paramount significance. The operational principles of a Big Data system fundamentally differ from traditional concepts of information collection, storage, or analytical architecture. By its nature, Big Data analysis represents a novel approach to intelligent information management. It involves the construction of a fundamentally new and complex structure, such as the analysis of ambiguous or unstructured datasets, and the distribution of tasksdata collection, storage, digital processing, and analysis-among multiple executable applications, which operate in accordance with algorithms defined by control modules. In the current context, the outcomes of such analysis play a decisive role in the evaluation of natural resources. These results provide the foundation for substantiating and diagnosing new reserves, supporting scientific INTERNATIONAL SCIENTIFIC JOURNAL SCIENCE AND INNOVATION SPECIAL ISSUE “MODERN PROBLEMS AND PROSPECTS FOR THE DEVELOPMENT OF DIGITAL TRANSFORMATION IN ENERGY” SEPTEMBER 24, 2025 93 investigation and exploration efforts, and developing services for organizations and enterprises. Moreover, they enable reliable forecasting of future development trajectories in the sector. Initially, the principles and sources of data collection are defined, and then the services for working with big data based on the Hadoop system are implemented within the MapReduce framework. Unlike the concepts of “Information” and “Data,” which are associated with the technical aspects of processing, the term “Big Data” does not primarily refer to the structuring or specific types of data. Rather, it may include structured, unstructured, and semi-structured data. In scientific literature, it is common to describe the various characteristics of Big Data using the 7 “V”s. All data collected through Big Data technologies can be classified according to their sources as follows. 1) Volume - Data volume, i.e., quantity – big data in the field of natural resources is composed of information originating from millions of automated devices, satellite imagery, electronic network equipment, and the applications in use. Unlike traditional mass data processing, these data volumes are processed in real time. This means that they are collected instantly, and the continuity of the data stream is highly significant. Thus, Big Data not only digitally captures data streams but must also record them in a lossless manner and reanalyze them; 2) Velocity – Speed – rapid data processing – unlike traditional methods, it not only records information streams and processes data packets in real time but also stores them in the required format and processes them digitally. The faster this happens, the better. An example of stream data processing is the Geoservers service, which not only includes complete georeferencing and coordinate projection but also analyzes user data based on materials skipped or deemed irrelevant by users. For the author's purposes, it offers the creation of specialized automated monitoring network channels and, in addition, provides services for collecting information about user interests, geographical characteristics, content preferences, and proposals for the target audience; 3) Variety – Types and diversity of data. Big Data is formed from various sources and data in different formats (video data, satellite images, aerial photographs, operational data recorded by automated devices, textual reports and project documentation, geological cross-section schemes, diagrams and maps, transaction files, well log interpretations), as well as link corrections and web page structures, among others. The main types of "Big Data" in terms of volume are those formed from social networks and social media services, representing partially structured or unstructured data. Thus, in the context of Big Data, the term does not merely refer to "large amounts of data" in a narrow sense – it is much broader, encompassing not only high-speed data processing but also the variety of sources and formats from which the data is obtained; 4) Veracity – Reliability – the authenticity of the initial data. Due to the large volume and variability of incoming data sources, it is difficult to control the reliability of Big Data. The relevance, accuracy, and precision of the received data can only be confirmed through careful analysis and comparative evaluation of the obtained results; 5) Variability – Variability refers to the fluctuation of initial data sets over a certain period of time. During processing and comparison, the original value of the obtained data may change, meaning it depends on contextual operational elements. Primarily, this characteristic appears when working with cartographic and textual data. To understand the precise meaning of individual features, it is necessary to develop models, algorithms, and complex software products that can determine the semantic load based not only on the literal meaning but also on the context; INTERNATIONAL SCIENTIFIC JOURNAL SCIENCE AND INNOVATION SPECIAL ISSUE “MODERN PROBLEMS AND PROSPECTS FOR THE DEVELOPMENT OF DIGITAL TRANSFORMATION IN ENERGY” SEPTEMBER 24, 2025 94 6) Value – "Value" refers to the potential change in the significance of data. The potential value shift of big data can be very high. For instance, fluctuations in groundwater levels, that is, amplitude, are influenced by the above-mentioned characteristics of Big Data: complete and accurate analysis of data, the relevance of the data, and conclusions derived through visualization. The data that can be used to solve the urgent problems of a particular user, as well as the results of analysis that help generate new ideas, represent major management challenges and arouse scientific interest; 7) Visualization – "Visualization is the process of representation, that is, imagining, and it deals with the issue of clear understanding and precise perception." The set of resulting data is sometimes not compatible with human perception or imagination. Therefore, it requires a special visualization process in an appropriate format - such as 2D, 3D, or 4D. A typical example of data visualization is the creation of maps, images, graphs, and diagrams that reflect the results of data analysis, and sometimes illustrate models and algorithms over time. The ability of Big Data visualization to self-adjust is important: in creating final outputs, parameters taken into account can be defined independently by users according to their goals and tasks, using techniques such as machine learning or self-configuration. All data collected by Big Data can be classified based on their sources - whether the data is needed for analytical processing or not. There is data that is not deliberately stored or collected by organizations, but is generated incidentally (in passing) during interaction with monitoring or network services, and remains in archive systems. The analytical mechanism within the framework of Big Data does not fundamentally differ in logic or structure from traditional analysis algorithms: information is recorded and collected - followed by analysis, comparative assessment, and conclusion formulation. However, the emergence of the "7-V" factors in the digital processing domain necessitates a new analytical approach. These factors include large data volumes, high update velocity, and the multiplicity of data sources - none of which can be effectively addressed by traditional analytical tools alone. Consequently, while the physical growth of computing systems allows for compensation of computational demands, it simultaneously leads to a decline in analytical processing speed. The analysis of Big Data differs fundamentally from traditional concepts of storage or monitoring. By its nature, Big Data analysis represents a novel approach to information management: it involves the creation of a fundamentally new and complex analysis architecture. This approach includes the distribution of data collection, storage, and analytical functions among multiple execution programs operating according to algorithms defined by control modules. These processes are implemented using the Hadoop software suite, which enables distributed (parallel) processing of large-scale data and the resolution of distributed tasks across multiple clusters comprising thousands of nodes. This is achieved through the integration of simple programs, auxiliary tools, and a collection of libraries. The relatively simple design of the Hadoop infrastructure ensures its popularity and broad applicability, even when operating across thousands of machines, each with its own processing and storage capabilities. HDFS (Hadoop Distributed File System) is a distributed file system that serves as the primary data storage system used by other components of Hadoop. Its architecture enables highperformance access to data and is designed specifically for working with large volumes of information. Files entered into the system are divided into blocks and distributed across individual nodes (DataNodes) of the computing cluster. INTERNATIONAL SCIENTIFIC JOURNAL SCIENCE AND INNOVATION SPECIAL ISSUE “MODERN PROBLEMS AND PROSPECTS FOR THE DEVELOPMENT OF DIGITAL TRANSFORMATION IN ENERGY” SEPTEMBER 24, 2025 95 In the integration of Hadoop and MapReduce systems, a distributed computing environment is created for processing large datasets. The primary task of MapReduce is to divide the input data set into independently processed blocks, in accordance with a task map (“map” – plan, scheme, chart, diagram), which allows parallel data processing. Thus, the fundamental operation of data handling involves dividing input information into smaller sub-tasks distributed among parallel processes (the “divide and conquer” principle). Each of these sub-tasks is executed in a strictly defined manner, ensuring the consistency of the results across parallel analyses. The outcomes of the sub-tasks are then aggregated in accordance with the main task in the reduce phase (“reduce” – to aggregate, to collapse). Typically, the computational and storage nodes are colocated, meaning that the MapReduce infrastructure and the HDFS operate on the same set of nodes. This approach provides the platform with the ability to optimally schedule tasks on nodes where data is already located, thereby ensuring high processing and data transmission speeds across the cluster. Together, HDFS and MapReduce form the core of the Hadoop ecosystem. The MapReduce system consists of one master module (JobTracker) and one subordinate module (TaskTracker) per cluster node. Together, they form a dedicated control mechanism responsible for scheduling tasks across cluster devices, monitoring their execution, and resubmitting any failed tasks for processing. Libraries such as HBase, Hive, and Mahout contain algorithms for machine learning and data mining, including algorithms for data classification and clustering. These algorithms are developed in separate modules, and Mahout also allows users to create their own algorithms to perform personalized tasks. The resulting algorithms are compatible with MapReduce, making them suitable for processing large volumes of data. Different types of data require different processing applications - Pig determines which one to use and serves as a platform for analyzing large datasets. It is based on a high-level language for expressing data analysis programs, including an evaluation system for these programs. BIBLIOGRAPHY 1. Джуманов Ж.Х. Геоинформационные технологии в гидрогеологии. -Т.: ГП «Институт ГИДРОИНГЕО», 2016. -258 с. 2. Мавлонов А.А., Джуманов Ж.Х. Гидрогеоинформационная модель подземных вод в геоинформационных системах//Геология и минеральные ресурсы. –Т. 2006. № 2. С.5559. 3. Хабибуллаев И., Джуманов Ж.Х. Об информационно-коммуникационной технологии в гидрогеологии. Геология и минеральные ресурсы. –Т.: 2014. № 1. С.48-54. 4. Веретенников А.В. Big Data: анализ больших данных сегодня. -М.: Молодой ученый, 2017. - № 32. С. 9-12. 5. Джуманов Ж.Х. Особенности компьютерной технологии при создании и использовании гидроинформационных моделей подземных вод// Гидрогеологические исследования в Узбекистане. Тр. посвящ. 50-летию гидрогеологической службы Узбекистана. –Т.: 2007. С.134-138.