scieee AI-readable full text Open interactive document viewer

Leveraging spatial data to enrich analytics systems

Galindo González, Diana Rocío

Abstract

Advancements in technology and data analytics are revolutionising the field of data science in business. These developments are increasing the availability of resources and enhancing interoperability capabilities. Big data technologies propel the capacity for information collection, processing, and cutting-edge analysis. As a result, we are seeing a growing demand for advanced analysis and specialised systems. Geospatial data is playing a role in this evolving paradigm. The integration of information and technology plays a vital role in shaping territorial decision-making, particularly in measuring the Sustainable Development Goals (SDGs) linked to the UN's 2030 agenda from a geospatial perspective. At the European level, initiatives such as INSPIRE (Infrastructure for Spatial Information in Europe) highlight the importance of integration in smart city planning and territorial information systems. This master's project proposes a comprehensive framework for data science projects that involve geospatial data. The framework builds upon DataOps concepts and existing data science frameworks, seeking to bridge the gap between conceptual and operative geospatial data science projects. It considers the particularities and technical aspects of managing and analysing geographical information. The project includes a proof of concept with two use cases, demonstrating the practical application of the proposed framework using real geospatial data.

Full text

id184303   LEVERAGING SPATIAL DATA TO ENRICH ANALYTICS SYSTEMS DIANA ROCÍO GALINDO GONZÁLEZ Thesis supervisor: JUANMANUELMIRASLÓPEZ(SISTEMASDEINFORMACIONTERRITORIALY POSICIONAMIENTO,SL) Tutor:PETARJOVANOVIC(DepartmentofServiceandInformationSystemEngineering) Degree:Master'sDegreeinDataScience Master's thesis Facultat d'Informàtica de Barcelona (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech 25/01/2024  Abstract Advancements in technology and data analytics are revolutionising the field of data science in business. These developments are increasing the availability of resources and enhancing interoperability capabilities. Big data technologies propel the capacity for information collection, processing, and cutting-edge analysis. As a result, we are seeing a growing demand for advanced analysis and specialised systems. Geospatial data is playing a role in this evolving paradigm. The integration of information and technology plays a vital role in shaping territorial decisionmaking, particularly in measuring the Sustainable Development Goals (SDGs) linked to the UN’s 2030 agenda from a geospatial perspective. At the European level, initiatives such as INSPIRE (Infrastructure for Spatial Information in Europe) highlight the importance of integration in smart city planning and territorial information systems. This master’s project proposes a comprehensive framework for data science projects that involve geospatial data. The framework builds upon DataOps concepts and existing data science frameworks, seeking to bridge the gap between conceptual and operative geospatial data science projects. It considers the particularities and technical aspects of managing and analysing geographical information. The project includes a proof of concept with two use cases, demonstrating the practical application of the proposed framework using real geospatial data. 1 Acknowledgements I want to thank my supervisor, Juan Manuel Miras, for his direction throughout my internship at SITEP, always bringing ideas and challenges to remind me how much I love what I do. I also want to thank the entire SITEP work team for their support, respect, and willingness to help me during this time. Furthermore, I am grateful to my tutor, Petar, for his orientation and guidance throughout this project and the whole master’s; truly the most pro of the pros. I also want to thank Vice Dean Oscar Romero for his endless patience and advice during the program and all my lecturers and professors for always being available in my learning process, especially to Profe Josefina in the first semester. This final project marks the end of a great effort from me and my loved ones. I could not have faced this challenge and achieve this goal without the love, help, and huge support of mis padres amados, mijn top amoro Raymond, and my friends and colleagues Alejo, Fernando FS, and Adrian. Brothers, Nephew, Schoonzuses, family, friends, colleagues at the distance and classmates and pals, divine treasures adding love, cheer, and joy to life. Thank you all for your support and encouragement. Building bridges takes us further than building walls. DaShanne Stokes 2 Contents 1 Background 11 1.1 Conceptual framework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 1.1.1 DataOps framework and architecture . . . . . . . . . . . . . . . . . . . . . . 13 1.1.2 Spatial key concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 1.2 State of the practice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 2 Proposed framework 29 2.1 Data management of spatial data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 2.1.1 Data governance proposal for data management . . . . . . . . . . . . . . . 31 2.2 Data analysis of spatial data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 2.2.1 Data Governance proposal for data analysis . . . . . . . . . . . . . . . . . . 35 2.3 Framework implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 3 Deployment technologies 42 3.1 Spatial data management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 3.1.1 Spatial data visualisation and deployment . . . . . . . . . . . . . . . . . . . 46 3.2 Spatial data analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 4 Proof of concept 50 4.1 Use case: Vector data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.1.1 Data management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51 4.1.2 Data analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 4.1.3 Solution implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 4.2 Use case: Raster data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 4.2.1 Data management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 4.2.2 Data analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 4.3 Results and Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 5 Conclusions and future work 64 Appendix 66 3 List of Figures 1 Methodology of the project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 1.1 DataOps methodology and architecture used as reference framework. . . . . . . 13 1.2 Data management backbone proposed by [30].. . . . . . . . . . . . . . . . . . . . 14 1.3 Data analysis backbone [35].. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 1.4 Metadata artifacts for governance [30].. . . . . . . . . . . . . . . . . . . . . . . . . 17 1.5 Attributes of metadata artifacts for Data management backbone presented by [30].. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 1.6 Taxonomy of spatial statistics prediction methods as presented by [17].. . . . . 22 1.7 Comparison of spatial statistics analysis methods by [27].. . . . . . . . . . . . . . 24 2.1 Proposed framework. Extended from [30]and [35].. . . . . . . . . . . . . . . . . 29 2.2 Data management backbone with spatial /spatiotemporal component. Extended from [30].. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 2.3 Data governance artifacts in data management for proposed framework. Adapted from [30].. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 2.4 Data analysis backbone with spatial /spatiotemporal component. Extended from [30].. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 2.5 Flowchart to implement the proposed framework . . . . . . . . . . . . . . . . . . . 37 3.1 Deployment technologies illustration for the Data management backbone of the proposed framework. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 3.2 Deployment technologies illustration for the Data analysis backbone of the proposed framework. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48 4.1 Implemented system architecture for vector data case. . . . . . . . . . . . . . . . . 51 4.2 Data sources and landing zone folder structure for vector use case. . . . . . . . . 54 4.3 Data Collector and log application for vector use case. . . . . . . . . . . . . . . . . 54 4.4 SQL Output - Minimum distances from grid population centroids to bus /train station. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 4.5 Exploitation zone Dashboard output for vector use case. . . . . . . . . . . . . . . . 56 4.6 Data analysis deployment visualisation exploratory map. . . . . . . . . . . . . . . 57 4.7 Data analysis deployment visualisation Moran cluster map. . . . . . . . . . . . . . 57 4.8 Data analysis deployment visualisation Moran scatterplot . . . . . . . . . . . . . . 58 4.9 Implemented system architecture for raster data case. . . . . . . . . . . . . . . . . 59 4 List of Tables 1.1 Data science tasks presented by [21].. . . . . . . . . . . . . . . . . . . . . . . . . . . 16 1.2 Errors in databases for spatial data [10].. . . . . . . . . . . . . . . . . . . . . . . . 25 2.1 Spatial based methods for data analysis tasks . . . . . . . . . . . . . . . . . . . . . . 34 2.2 Metadata artifact proposed for data analysis governance . . . . . . . . . . . . . . . 36 3.1 NoSQL databases spatial support according to [1].. . . . . . . . . . . . . . . . . . 44 3.2 Big spatiotemporal Data Processing Infrastructures presented by [1].. . . . . . . 45 3.3 Geospatial Technologies for web map visualisation /dashboarding. . . . . . . . . 46 4.1 Data sources, analysis and task required to calculate KPI of the use case for vector data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 4.2 Data sources, analysis and task required to calculate KPI of the use case for raster data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 5 Listings 4.1 Tessellation centroid to closest station distance calculation . . . . . . . . . . . . . . 54 5.1 Data collector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 5.2 Data persistant loader . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78 5.3 Data Formatter . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 5.4 Data analysis model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 6 Introduction As technology and data analytics continue to advance, so does the availability of resources and interoperability capabilities. With enhanced capacity for collecting and processing information, cutting-edge analysis methods, and big data technologies, new opportunities for data science in business contexts are emerging and results in a greater demand for advanced analysis and utilization of previously specialized systems. This is the case with geographic information and geospatial data [44], as well. The convergence of information and technology plays a crucial role in territorial decisionmaking. A notable illustration of this impact is evident in the measurement of the Sustainable Development Goals (SDGs) linked to the United Nations (UN) 2030 agenda from the geospatial perspective1. Over the past decade, the UN has spearheaded several Expert Group initiatives such as the United Nations Integrated Geospatial Information Framework (UN-IGIF) and the Committee of Experts on Big UN Data and Data Science for Official Statistics, emphasizing the strategic use and exploitation of geospatial data. [39]presents a study of the role of geospatial information in contributing to sustainable development. At the European level, there is a growing need for cities and territorial information systems to be integrated thoughtfully. This need is highlighted by initiatives such as INSPIRE (Infrastructure for Spatial Information in Europe). INSPIRE is a directive from the European Union that aims to create a spatial data infrastructure, making it easier for public sector organizations to share environmental spatial information. Similarly, the paper by [8]proposes smart city planning based on data integration. At the local level, the Generalitat of Catalonia has more than 60 systems that use geospatial resources in territorial information systems and that are available through Hipermapa of the Generalitat that enable the exploitation of data and, in turn, require greater technological capabilities. Model variability based on location evolution demands expanded data management and analytical capabilities, promotes interoperability, and redefines spatial analytics within dynamic data lake and big data environments. This shift in approach extends even to technology services such as conventional data science. 1https://sdgtransformationcenter.org/ 7 Data Management Backbone The data management backbone encompasses collecting, storing, processing, and readying data in alignment with the gathered specifications. As a result of this phase, the required views of the data are expected to perform the analysis[4]. In the context of this framework, it is divided into two phases: data ingestion, where the data sources of interest of the project are mapped and by means of data collectors pass to the data Storage. In the storage phase, through different zones (landing, formatted, and trusted), data is prepared in a pipeline to be engineered until it reaches a level that can be exploited and analised. An illustration of this part of the architecture is presented in Figure 1.2. Figure 1.2: Data management backbone proposed by [30]. The landing zone has two parts: temporal where data is stored in temporary files in the simplest way before passing to the persistent zone, where it is organized based on data source and timestamp, creating versions and setting up elements to facilitate data indexing. The next step involves homogenizing the data according to a standardized data model in the denominates formatted zone, ensuring syntactic homogenization and a proper arrangement of data to be exploited and, according to the requirements based on the business understanding, in a database storage. Subsequently, the trusted zone is set to transform landing data into qualified data to be input in the exploitation zone. It is the space to conduct data quality process and to perform generic and standard data cleaning not related to the specific one of the project [35]. It could include removal of duplicated records and outlier detection. Finally, the exploitation zone can include visualisation displays and storage of relational tables to be used in the data analysis backbone but it is mainly oriented to integrate information from the previous zones to achieve discovery, entity resolution, ad-hoc transformations and data loading into a target schema in a potentially different data model. 14 Data Analysis Backbone The data analysis backbone involves uncovering relevant information and utilizing models that can address the specific questions the system needs to solve using the accessible data assuring quality and thresholds criteria [4]. In real life, a project always may consider several pipelines [35]and through an iterative process achieve an appropriate modelling to discover consistent and robust information. Figure 1.3 depicts the architecture of an analytical pipeline. Figure 1.3: Data analysis backbone [35]. After managing and presenting the necessary data in the exploitation zone of the project, Analytical Sandboxes are utilized to capture specific subsets of relevant elements for the analysis phase. This phase aims to address questions expected to be answered through the system. The pipelines generated in this project stage encompass the phases outlined in methodologies such as CRISP-DM [42]. The analytical sandboxes capture subsets of the elements created in the exploitation zone that are pertinent to the ongoing analysis. This involves applying data preparation rules tailored to the type of analysis. If necessary, labeling is also performed at this stage. Consequently, two sets of datasets are produced: the training and validation datasets. Using features derived from these analytical sandboxes a Feature generation is performed. If necessary, labeling is also done at this stage. Consequently, two data sets are produced: the training and validation data sets. According to the task and algorithm chosen to pursue the business requirements, the model training step is performed. This step is an iterative process where hyper-parameters are identified and adjusted using the training data set. The model is then trained and validated based on quality criteria, such as accuracy and recall. [21]presents a classification of modeling alternatives according to the type of analysis pursued presented in table 1.1. Taking as reference the [21]approach, data science tasks include: Classification: Determ15 Table 1.1: Data science tasks presented by [21]. Tasks Algorithms Example Classification Decision trees, neural networks, bayesian models, induction rules, k-nearest neighbors Assigning voters into known buckets by political parties Regression Linear regression, logistic regression Predicting the unemployment rate for the next year Anomaly detection Distance-based, density-based, LOF Detecting fraudulent credit card transactions Time series forecasting Exponential smoothing, ARIMA, regression Sales forecasting Clustering k-Means, density-based clustering Finding customer segments in a company based on transactions Association analysis Apriori, FP-growth, Eclat Identifying frequent itemsets in a market basket Recommendation engines Collaborative filtering, content-based filtering, hybrid recommenders Finding the top recommended movies for a user ine whether a given data point falls into predefined classes, Regression: Predict the numerical target label of a data point; Anomaly Detection: Predict whether a data point deviates significantly from the rest of the data in the dataset, Time Series Forecasting: Predict the future value of the target variable based on historical values, Clustering: Identify inherent clusters within the dataset based on its properties, Association Analysis: Identify relationships within an item set based on transactional data and Recommendation Engines: Predict a user’s preference for an item. Statistical machine learning, and deep learning methods and techniques cover the modeling stage. Mapping or description of the algorithms to perform model tasks provides such a vast number of possibilities, and variability in pros and cons according to the domain of each analysis is out of the scope of this project. A specific oriented location-centered method is available in the 1.1.2 chapter. Literature references of best practices for statistics, machine learning [43], and artificial intelligence modeling apply in this stage of the project and should be adapted according to the purpose. Data Governance The data governance component is transversal to the two backbones and aims to create a structure that promotes quality, facilitates cohesive management, ensures systematic oversight, and enables efficient reporting, ultimately enhancing the value of data within the corporate /business environment [30]. Figures 1.4 and 1.5 expose a diagram for each corresponding backbone. In the data management backbone aData Governance Process (DGP), explained by [30] is responsible for moving data across the zones. Data collectors in charge of extracting data from specific data sources and load them in a temporal landing zone or data lake, a Data persistence loader (DPL) which permanently and systematically stores the ingested data files 16 and a Data formatter managing the conversion of the data into a unified format and making it available for the exploitation zone. In the temporal landing, the data collector fetches data from the sources and stores them using a hierarchical storage structure, allowing rapid storage and easy searching for the data files. Subsequently, in the persistence landing, the DPL takes charge of organizing and enriching the information. Figure 1.4: Metadata artifacts for governance [30]. This involves adding metadata, such as schema definitions, to facilitate automatic parsing of datasets during further integration. To handle this information effectively, Persistent Landing is deployed using a Key-Value store. This approach allows for easy access and quick querying of stored data through keys, providing flexibility in storing various information from the data source as values. The key is designed to enable efficient querying, with the prefix determining the data’s name, followed by determinators specific to the source (e.g., format, data type, domain), and a timestamp as a suffix. This design allows for swift retrieval of data from a specific source, different data formats within the same source, or consecutive extractions from the same source [30].Figure 1.4 depicts structure of the Metadata artifacts the and Figure 1.5 provides details for attributes per artifact. The Formatted Zone is a structured subset of ingested data, providing a uniform approach for accessing information from different sources. Data Formatter DGPs, specific to each data format, convert and transfer data from the Persistent Landing submodule to the Formatted Zone. The conversion relies on schema information to achieve a canonical data model, assumed to be tabular in this framework. Once in the Formatted Zone, data is ready for use, although original quality issues may persist. The Exploitation Zone serves as the endpoint for analysis-ready data intended for end-users. The framework considers three exploitation models: tables, Dataframes, and tensors, covering a broad range of analytical needs. This zone can host various views over the data previously 17 Figure 1.5: Attributes of metadata artifacts for Data management backbone presented by [30]. stored in the Formatted Zone, potentially created for different data services. The proposal of [30]emphasizes how systematic governance facilitates the final exploitation and analysis of data by diverse end-users and do not go into detail of Analysis Backbones. 1.1.2 Spatial key concepts As explained by [31]Spatial data involves describing the positions of observations through coordinates, which are established within a coordinate system. Various coordinate systems exist, with a key distinction being whether the coordinates are specified in a 2-dimensional or 3-dimensional space with reference to orthogonal axes (Cartesian coordinates), or if they are determined based on distance and directions (polar, spherical, and ellipsoidal coordinates). In addition to pinpointing the observation locations, all observations are linked to the time of their occurrence. A detailed explanation of coordinates systems can be found in [40],[16] and cited by [31],[24]. We utilized two important factors or concepts that are specific to geospatial data as opposed to conventional data in this project. The first factor involves the use of coordinates and their reference system (CRS) and the second is related to the data representation model. 18 Based on [28], a Coordinate Reference System (CRS) is an abstract mathematical concept that defines a sequence of axes with specific units of measurement. A coordinate system is associated with a particular reference object, such as Earth or any other object of interest, through a datum. In terms of geographic coordinates, the Datum refers to a set of reference points on the Earth’s surface that are used to measure positions, along with a corresponding model of the Earth’s shape. There are two types of coordinates: geographic coordinates, which use latitude and longitude, and plane coordinates, which involve projecting the datum model onto a plane. These various coordinate reference systems are identified through the use of EPSG coding. It is necessary to take into account how projected systems cause distortion in measuring quantities. A detailed explanation was presented by [29]. This is strictly related to the purpose of the usage of the data and the scale of the input and output data and, by general rule, projecting the input data in a single CRS. The local reference system established by the geographic national agency should minimize serious mistakes or distortions in the output information. Compared to the conventional data, this is an added factor arising from the deployment and integration of data across various systems going beyond statistical or mathematical assumptions, impacting overall performance. Errors may arise from the improper blending of information in Cartesian coordinates, the use of different geoids, or actions that invalidate or introduce inaccuracies into the final results. Additionally, the inclusion of time in computations can lead to significant variations and affect the precision or accuracy of quantifying measurements and shapes. The second key concept in the geospatial context is the existence of two models for the real world: vector and raster. The vector model describes feature geometries, where features encompass entities with geometry, potentially implicit temporal properties, and additional attributes, including descriptive labels or quantitative values. The primary application of simple feature geometries lies in describing two-dimensional geometries using points, lines, or non-self-intersecting polygons. These geometries can be queried for properties, transformed, combined into new geometries, and further queried for additional properties. This is the case of spatial networks like transportation networks, with nodes representing points of interest or intersections. In the same way, a tessellation involves subdividing space into smaller elements using polygons. Regular tessellations, accomplished with regular polygons like triangles, squares, or hexagons, are employed for spatial data; those using squares are known as raster data [31]. The raster model corresponds to a regularly laid out lattice of usually square pixels. Raster dimensions describe how the rows and columns relate to spatial coordinates. The pixels can represent either discrete or continuous variables. 19 Spatial data can be captured with temporal information, forming spatiotemporal data, which integrates both spatial and temporal aspects. This data type involves geometries changing over time and is represented by various models. spatiotemporal data types combine timestamps with spatial data types, including points, lines, and polygons and also multitemporal rasters [1]. An integral view of spatial information theory can be found in [22], where the author explains in detail and contextually spatial information associated concepts as: location, neighborhood, field, object, network, event, granularity, accuracy, meaning, and value. Even though these concepts are relevant in the analysis of geographic data and derive other factors of importance in the context of geospatial data, this master project considers these two tools in the analysis backbone as the basic structure required to have an extended framework in a geospatial context. In the specific case of data exploitation and analysis, particular ecosystems and methods that respond to different phenomena of interest where space and territory are related to the project’s objective are generated. These methods incorporate and quantify spatial variability or spatial autocorrelation and are framed in the field of spatial analysis. Spatial analysis One of the main differences between the conventional data analysis and those related to the spatial dimension of geoinformation that refers to the joins. An spatial join is a query where two or more geometries are compared based on their locations. The conventional join uses standard keys or identifiers between tables to relate them. In a spatial join, uses one or more spatial functions, as predicates, to determine the spatial relationship between its objects as overlay, intersect, contain, union, difference, and symetric difference. These operations correspond to the basic functions of spatial analysis. Likewise, the concept of Spatial autocorrelation is decisive in analysing geospatial data concerning conventional data. As explained by [13], the concept of spatial autocorrelation originated in the late 1950s, based on the principle of nearness, suggesting that nearby areas exert a stronger influence on each other than distant areas. This perspective aligns with Tobler’s First Law, stating that everything is related, but proximity enhances the degree of relationship. It was referred to by various names in the social science and statistical literature, such as spatial dependence, spatial association, spatial interaction, and spatial interdependence, and their mathematical characteristics were outlined by statisticians such as Moran, Krishna-Iyer, and Geary. Methods associated with spatial autocorrelation allow exploratory, clustering and regression analysis of spatial data. [5]explains the spatial point patterns implemented in R stats. In summary, the exploratory analysis is based on defining intensity, correlation and spacing functions. Statistical inference uses logistic regression models (including Poisson processes), 20 hypothesis testing and simulation as in conventional data methods such as Monte Carlo. The author also explains cases with covariates and the tests necessary for model validations and spatial interaction models. As in statistical prediction, the application of the methods is subject to assumptions, confidence intervals and uncertainty calculation. Adopting the framework proposed by [6], specialized methods in the geospatial context can be categorized into several fields: Spatial Point Pattern Analysis (SPP), Interpolation, Geostatistics, and Modeling Areal Data. Similar to conventional regression methods, these techniques rely on statistical modelling for specific purposes. Within each field, [6]explains various methods. In Spatial Point Pattern Analysis (SPP), which involves the statistical analysis of spatial point processes, there are techniques such as Homogeneous and Inhomogeneous Poisson Processes, Estimation of Intensity, and Likelihood of an Inhomogeneous Poisson Process it is frequently employed in epidemiology. In the realm of Interpolation and Geostatistics, notable methods include models associated with spatial autocorrelation, Spatial Prediction (Universal, Ordinary, and Simple Univariate Kriging), Multivariable Prediction (Cokriging, Model-Based Geostatistics, and Bayesian Approaches). For Modeling Areal Data, the emphasis is on Spatial Statistics, encompassing Simultaneous and Conditional Autoregressive Models, Spatial Regression Models, Mixed-Effects Models, Spatial Econometrics, Generalized Additive Models (GAM), Generalized Estimating Equations (GEE), and Generalized Linear Mixed-Effect Models (GLMM). According to the authors, these techniques find extensive application in modelling various phenomena across diverse fields such as biology, spatial economics, image processing, environmental and earth science, ecology, geography, epidemiology, agronomy, forestry, and mineral prospection. [17]presents a study summarizing algorithms for spatio temporal prediction and states a taxonomy of spatial prediction methods depicted in 1.7. The classification of the authors expands the possibilities of data analysis specifically considering spatial component and simultaneously enclose a new set of techniques from statistical procedures. According to the advances in ML and AI, and to their requirements, nearness and spatial correlation are also derived. In machine learning and artificial intelligence methodologies, [19]introduces hybrid options that combine geostatistics and neural networks, optimizing geoinformation analysis computation. Various techniques, including different architectures of artificial neural networks and statistical learning theory, as well as kernel-based methods like support vector machines and support vector regression, play a key role in these methodologies. The authors also point out that it is crucial to underscore that these data-driven approaches, often depicted as black boxes, heavily depend on the quality and quantity of data. Consequently, incorporating diverse statistical and geostatistical tools becomes imperative to effectively supervise and regulate the quality of data analysis and modelling within the framework of Machine Learning. 21 Figure 1.6: Taxonomy of spatial statistics prediction methods as presented by [17]. Additionally, researchers like [38]highlight the significance of cross-validation samples in models through conventional approaches. Algorithms may not inherently adapt to data science projects that include geospatial information. Moreover, the spatial dimension of the data should be considered during the sampling for model training. This particularization of sampling also extends to statistical modeling. In the case of statistical methods for alphanumeric and geographic data, compliance with a series of assumptions is required to guarantee the particular scope of each method to be applied. In this work, [38]discuss the impact of spatial autocorrelation on modeling ecological phenomena and emphasizes the use of Grouped Cross-Validation strategies to mitigate bias in estimations when employing machine learning methods. The study evaluates various machine learning algorithms, including Boosted Regression Trees (BRT), k-Nearest Neighbors (KNN), Random Forest (RF), and Support Vector Machine (SVM), comparing them with traditional parametric algorithms such as Logistic Regression (GLM) and semi-parametric ones like Generalized Additive Models (GAM) in terms of predictive performance. Different data resampling approaches are also explored. The findings suggest that spatial hyperparameter tuning, particularly in spatial cross-validation, aligns well with spatial estimation of classifier performance, outperforming non-spatial hyperparameter optimization. The study reveals substantial performance differences (up to 47%) between bias-reduced (spatial cross-validation) and overoptimistic (non-spatial crossvalidation) settings, underscoring the importance of accounting for the influence of spatial autocorrelation. While machine learning algorithms have gained popularity for handling high-dimensional and correlated data with fewer model assumptions, their ability to make inferences is still limited compared to parametric models. Hyperparameter optimization is essential for achieving the best model performance, typically conducted through automatic procedures like random 22 search or Bayesian optimization. In contrast, parameters of parametric models are estimated during model fitting. The study also highlights the significance of cross-validation, a resampling-based technique for estimating predictive performance. Spatial cross-validation, inspired by Brenning’s approach, utilizes kmeans clustering to diminish the impact of spatial autocorrelation by dividing the data into spatially disjoint subsets. Regarding model performance, in spatial settings, Random Forest (RF) demonstrates the best predictive performance, followed by Boosted Regression Trees (BRT), k-Nearest Neighbors (KNN), and Logistic Regression (GLM). Generalized Additive Models (GAM) exhibit higher variance in spatial settings compared to other algorithms. Finally, it emphasizes the necessity of spatial cross-validation for accurate performance estimation in spatial predictive modeling, particularly in the context of epidemiological analysis. The study suggests that this approach is crucial for addressing the challenges posed by spatial autocorrelation in ecological modeling. For Spatial clustering [41]explains approaches referencing divided major clustering methods into four categories, which were partitioning methods, hierarchical methods, grid-based methods, and density-based methods. The authors describe several approaches concerning spatial clustering methods: Partitioning methods, exemplified by K-means, use iteration to identify clusters and their centres, proposed CLustering for Large Applications based on Random Search (CLARANS) to detect both points and polygonal objects. Hierarchical methods use distance or density functions to create clusters, such as Balanced Iterative Reduction and Hierarchy Clustering (BIRCH) and Chameleon. Density-based methods, such as Density-Based Spatial Clustering of Applications with Noise (DBSCAN), can discover clusters in various ways, and OPTICS addresses the sensitivity of DBSCAN to input parameters. They also mention grid-based methods use a grid structure to form groups and model-based approaches. The authors emphasize that spatial clustering differs from spatiotemporal (ST) clustering due to the introduction of the "time" element. ST clustering integrates spatial and temporal information for more detailed analysis. spatiotemporal interaction methods, spatiotemporal k nearest neighbour test, and scan statistics are introduced as ST clustering methods. In this sense, mainly, the scan statistics method involves a circular scan window to identify spatial clusters, and the space-time scan statistics extend this to detect clusters in both space and time. Partitional clustering methods, such as ST-DBSCAN, introduce a temporal neighbourhood radius to cluster spatiotemporal data. The appropriate spatial and temporal radii selection is determined using the k-distance graph, ensuring a clear separation between noise and cluster data. They also mention Kernel Density Estimation (KDE), a non-parametric method that uses a Gaussian function to detect clusters in spatial data. [31]states that currently, the 23 Figure 2.2: Data management backbone with spatial /spatiotemporal component. Extended from [30]. Regarding Data Ingestion, geospatial data source formats differ in structure and variety from conventional alphanumeric data sources that can be received in a project. In the landing zone, particularly in the temporal domain, when expanding data sources, it is recommended to categorise them similarly to conventional approach, maintaining hierarchical storage structure but encompassing three types of data: alphanumeric or non-spatial, spatial and spatiotemporal. This division allows a more effective selection of sources, enables the identification of available resources for exploitation and data analysis areas, and can facilitate the identification of strategies for optimisation in information loading operationalisation. In the persistent zone, we propose incorporating the EPSG code of the CRS as part of the key structure. Every data science project involving spatial methods or geoinformation must consider the reference system. To manage geospatial sources, CRS is required for the formatted zone. Additionally, an information model identifier makes data processing more convenient before launching into the formatted zone. It aims to identify projections in spatial data and is part of data governance. In the Persistent zone, we also suggest using composite labels according to the information model. For the vector case, including the geometry as: vector_point,vector_line,vector_polygon,vector_tesselation or vector_grid. For the raster case a label including spatial resolution as raster_15cm. In the vector case, it is useful to know the associated geometry to identify possible transformations and homogenisations required in posterior zones. In the raster case, we consider including spatial resolution as part of the easy-identify transformations when sources with different resolutions should be integrated. This is a common practice in raster data source labelling. 30 In the Formatted zone, a technology that considers spatial data must be chosen according to the model and format of the available data. Another storage technology decision factor depends on the purpose of the exploitation zone; the data format available from the formatted area should ensure ready-to-use data sets. In this zone it is also suggested to use spatial indices, which would require technologies that support spatial data. Also, in the Formatted zone, homogenization through the projection of the data to the same CRS is imperative as some analysis methods or technological tools require a specific planar or geographic coordinate system and to ensure that the visualization is effective according to the technologies chosen. Regarding handling big spatial data, a proper framework such as the exposed in [18]for this purpose should be considered. However, as a basic statement, in the scope of this project it should be pointed out that the raster model inherently corresponds to large volumes of data, demanding more processing capacity due to the information’s spatial pixel resolution and terrain coverage. The storage requirements increase further with high or broad spectral resolution data, scaling further in the Data Analysis stage. Therefore, it is suggested in this framework, for projects including raster data priorisation of scalable and big data oriented storage technologies. In the Exploitation zone, since spatial data frequently involves mapping as a resource for displaying information, we suggest using technologies able to support mapping and consider dashboarding technologies able to provide the user capabilities to download the data in spatial formats. Making the output available in spatial formats facilitates data interoperability and reproducibility of projects. It is also preferable to use technologies that can connect the displayed information on the map and other graphical resources to facilitate their exploitation purpose. In the exploitation zone, it is important to select the appropriate spatial data format for further analysis. Depending on the technology chosen, data frames or spatial data frames must be used. In geospatial data science projects, spatial data formats, databases, data frames, and tensors are used, but it is suggested to avoid multiple transformations of the geospatial data if a data analysis backbone is planned to deployed. A more tool-oriented perspective on available technologies at the time of this project’s elaboration is provided in Chapter 3. 2.1.1 Data governance proposal for data management As presented in the 1,Data Governance Processes (DGP) in this backbone corresponds to: Data collectors,Data Persistance Loader, and Data formatters which include a set of artifacts. In this proposed framework, both, processes and metadata artifacts preserve the same structure, but metadata attributes can be expanded in relation with the geospatial data.Figure 2.3 illustrates the proposed extension to the conventional framework adapted from [30]. 31 Figure 2.3: Data governance artifacts in data management for proposed framework. Adapted from [30]. As exposed, the Data collectors in the temporal landing fetch data from the sources and stores them using, in this case, the hierarchical storage structure according to non-spatial, spatial and spatiotemporal data. Subsequently, the Data Persistance Loader organises in a Key-value storage and adds the required metadata. In this case, at minimum the Key contains the CRS, and data model(vector or raster), geometry for vector sources, and spatial resolution for raster sources and a timestamp as a suffix as in the original approach. Depending on the project, and accuracy and scale or level of detail required, this should include or not additional metadata using references such as the geographic information. For Data formatters it is intended that there be one per type of data (spatial and non-spatial) and data format as a dedicated process for converting and transferring information from the Persistent Landing to the Formatted Zone. Assuming that the selected technology for implementing the data governance application has the capability to execute operations in the persist32 ent zone, it is anticipated to, when necessary, perform a reprojection to a unified Coordinate Reference System (CRS) of spatial data at this stage. This step aims to prevent redundant information in various reference systems unless it is absolutely essential. In the Exploitation Zone, we consider tables, spatial data in files or web services, dataframes and tensors intended for end-users in a ready-to-use format. Likewise, in spite of the exploitation zone outputs, there should be a preview of the spatial data to verify the location consistency of the data. 2.2 Data analysis of spatial data In the data analysis backbone, the foundation of an extended framework relies on addressing specific methods able to model phenomena of interest, particularly where space and territory are an important factor within the project’s objectives. There are methods explicitly designed to quantify spatial variability or spatial autocorrelation or incorporate spatial criteria in this stage. The data analysis backbone of proposed framework is illustrated in 2.2. Figure 2.4: Data analysis backbone with spatial /spatiotemporal component. Extended from [30]. Table 2.1 provides an overview of different techniques and methods for analyzing geospatial data. Each row of the table represents potential inputs of geospatial information, while the columns present methods and expected types of output data for analyzing the task at hand. The primary objective is exploring possibilities aligned with exposed tasks and methods. 33 Table 2.1: Spatial based methods for data analysis tasks SPATIAL ANALYSIS TASK INPUT CLUSTERING/ SEGMENTATION CLASSIFICATION REGRESSION / INTERPOLATION Point LISA* ■ DBSCAN • Spatial point patterns • Geostatistics ■ Spatial regression ■ Spatial point patterns ■ Network SPP over linear networks ⋋ Spatial Interaction Data ⋋ Polygon / Tessellation Local and Global Moran’s I ◊⊞ CLARANS ◊ DBSCAN ◊ BIRCH ◊ Spatial regression ◊ Raster GRF** SPP: Spatial Point Patterns * LISA: Local Indicators Spatial Association ** Geographical Random Forests OUTPUT: Point •Network ⋋Polygon ◊Tessellation /Grid ⊞Raster ■ It is important to emphasize that while this table presents known methods, conducting a thorough literature review during the analysis stage is highly recommended to uncover additional possibilities. The aim is to explore possibilities aligned with exposed tasks and methods. As outlined in Chapter 1, the distinction between conventional clustering and spatial clustering involves grouping objects with specific dimensions so that objects within a group share similar characteristics compared to those in other groups. Particular methods for spatial clustering encompass Ripley’s K Function, Density Kernel Estimation, Exploratory Data for Networks, Moran Local, Geary Local, Getis-Ord, Moran’s I Global, Gearys’ C Global, Moran Local, Geary Local, Getis-Ord. For regression purposes in spatial data analysis, specialized methods cover requirements such as explaining the pattern of event occurrences, predicting outcomes based on multiple inputs, and predicting the number of events starting in one node and ending in another. Notable methods include Kriging interpolation, Poisson Processes, Cox Processes, Regression for Correlated Data, SAR, CAR, Spatial Econometrics, Geographically Weighted Regression, and Gravitational Models. When dealing with spatial data across different timestamps or periods, the analysis is called spatiotemporal analysis, often involving data cubes. However, this type of spatial analysis falls outside this project’s scope. [31]explores the concept of data cubes, highlighting that are generally not required for point patterns and trajectories. Nevertheless, the authors note that the initial computational steps for such data often involve creating data cube representations through aggregation based on either time-fixed spatial or space-fixed temporal discretization. 34 The utilisation of spatial data cubes is discussed by [31], noting that those associated with point patterns and trajectories generally do not requires them. Spatiotemporal point patterns represent sets of coordinates over time for events or objects, such as accidents, disease cases, traffic jams, lightning strikes, etc. Trajectory data, on the other hand, consist of time sequences of spatial locations for moving objects like people, cars, satellites, and animals. For such data, the primary information lies in the coordinates. Although transforming these into a limited set of regularly discretized grid cells covering space may facilitate some analyses, such as a rapid exploration of patterns in areas with higher densities, the loss of exact coordinates also limits certain analysis approaches involving distance, direction, or speed calculations. Nevertheless, the initial computational steps for such data often involve generating data cube representations by aggregating to a time-fixed spatial and/or space-fixed temporal discretisation. While spatial methods are widely implemented in research and development, a notable bottleneck emerges in the industry or business intelligence spectrum due to processing limitations. The intricacies of calculations render these algorithms resource-intensive, hindering their widespread use. However, the current status of spatial data management and analysis requires leveraging big data resources to parallelize processing and overcome these computational limitations, thereby unlocking new possibilities for spatial analysis. 2.2.1 Data Governance proposal for data analysis The use of of geospatial information or spatial-based methods also requires particularisation in the Data Governance Component. The DGP processes considered for this stage are data transformation and data preparation. Following the conventional framework, we adopt the four essential metadata artifacts: Data Views Registry, Datasets Registry, Models, and Performance. Expanding on this structure, in the domain of spatial data analysis, both can be extended to register the particularities exposed previously; we introduce an additional log within the Model artifact. This log comprises a registry, documenting details such as the task performed, the input layers employed, the CRS used for calculations, a mention of the method undertaken, resulting layers and resolution or spatial scale in the output data. Additionally, the log provides insights distinguishing between short and long-term analysis needs and offers a view of the data analysis lifecycle. An explanation of the structure of the log table is supplied in 2.2. By incorporating this detailed log, our proposal enhances the tracking and documentation of the analysis stage, ensuring a more thorough understanding of the spatial data analysis processes. This completes the proposed framework for data science projects with geospatial components object of this project. 35 Table 2.2: Metadata artifact proposed for data analysis governance Registry Detail Analysis task Described as a general purpose (as in Table 2.1 ) EPSG_Code CRS Used / required to perform the analysis Input data Taken from datasets registry Method Algorithm or technique use Output data Taken from datasets registry Output geo level scale / spatial resolution of output Short / long term Analysis or model review requirement 2.3 Framework implementation The flowchart in Figure 2.5 aims to provide a comprehensive and step-by-step diagram to implement the proposed framework. It encompasses the proposal to implement a data science project considering spatial data. 1. KPI Definition and Identification of Expected Visualization The initial step in implementing this extended framework is to collaboratively define or identify, with the client, the expected responses from the system or application under design. Determining if the spatial component is relevant in these responses is crucial, which generally involves where and in what areas. In what places? In other words, if the location is determinative or adds value to the decision-making process. At this point, it is essential to identify if, in the client’s need, the visualization of the information on a map helps to understand or make better use of the data for its purpose. 2. Data Sources Identification The next step involves identifying the sources of information. In addition to user-provided data, this extension considers additional geographical information. This encompasses details from available portals, spatial data infrastructures, national statistical and mapping agencies, earth observation data, and any other relevant information that can enhance data visualization, exploitation, and analysis. 3. Follow conventional framework Determining the client’s needs and the availability of information for data exploitation and analysis is crucial for the viability of the framework. Otherwise, the conventional framework, as exposed by [4],[30], and [18], may be more suitable. 4. Classify Sources While data sources may not inherently possess explicit geospatial attributes, evaluating their potential for spatialisations or representation with geographical data is imperative. In this context, we propose a classification scheme for data sources, categorising them into non-spatial, spatial, or spatiotemporal (ST). 36 Figure 2.5: Flowchart to implement the proposed framework 37 Non-spatial data refers to information that does not have any associated geometry or is not available in a raster format. On the other hand, spatial data includes different types such as geographic information layers in either vector or raster formats, various data formats, web spatial services, and the ones explained in Section 3.1. Furthermore, spatio-temporal data includes datasets that have geographic information or location-based attributes available for different timestamps or periods. During this stage, a verification step examines the feasibility of spatialisation for non-spatially formatted sources. 5. Spatialise data sources Spatialisation of a variable from a non-spatial source becomes viable when the dataset’s attributes incorporate geographic location information. This location data may manifest as coordinates, references to political-administrative entities, delineations of river basins, geological units, areas of influence, or geometries of interest to the client. If the attribute stores address information, it can be transformed into coordinates using geocoding tools available through APIs, scripts or software for such a purpose. Territorial entities often provide geometries through National Geographic Institutes, National Spatial Data Infrastructure Agencies, or Open Data Portals. The same applies to spatial entities like river basins, geological units, or similar entities where incorporating data reconciliation sources is imperative. Grids, tessellations, and raster model information may also be available in nonspatial formats like matrices, arrays, or plain files, requiring associated metadata and reference systems for utilisation. Generally, if a data source contains a set of coordinates latitude and longitude or X and Y, the process to spatialize it is commonly known as Geocoding. In contrast, if a source contains information about a location associated with a geometry, such as administrative boundaries, it is necessary to implement a reconciliation mechanism to join the data source and representative geometry. Likewise, there is the case that an image corresponds to a geographic space with no coordinates but where the user knows the location, where what is called a georeferencing process would apply. In any of the cases, the CRS should be known. 6. Identify model and CRS This step is intended to verify the presence of a Coordinate Reference System (CRS) in geographic information. The effective utilisation of specific analysis methods and the precise deployment and overlaying of information depend on the availability of a CRS. In this case, obtaining a reference to the CRS is expected using its EPSG code website. Additionally, it is crucial to discern the display format of each information source, distinguishing between geometries (vector model) and images (raster model). It is essential to identify the type of geometry employed for vector data, while for raster data, both spatial and spectral resolutions should be determined. This information helps establish the capabilities 38 and possibilities in subsequent management and analysis phases, laying the foundation for informed decision-making in the data processing pipeline. 7. Identify scope of exploitation and data analysis Once data type and CRS per source are determined, the focus shifts to identifying, at a glance, the scope exploitation zone and data analysis tasks. It means to evaluate if the project scope includes spatial quantitative analysis, technologies like GIS tools or data visualisation platforms may be selected as those exposed in sections 1.1.2 and 3.1. If the project involves data exploitation without spatial quantitative analysis, the exploitation zone can display the information directly on Dashboarding Maps using technologies that are not geospatial-focused. Refer to Chapter 3for more details. If spatialisation is not feasible, the next step would be to identify data management technologies that can replicate the spatial data flow. It involves reviewing the means and relationships between KPI and data sources and defining transformations and the geospatial level or scale to display the information. 8. Determine output spatial model and level In this step, it is expected to identify the level of geographic information to be used in the data exploitation area and/or in the analysis stage. Based on consolidated information from sources and expected outputs, output levels and formats need to be confirmed. The levels in the vector case will be the geometries, and in the cases that apply the scale of the information or the size of the grid; In the raster case it would be the spatial resolution. The formats would correspond to ready-to-use elements such as geographic layers, tables, geodataframes or those exposed in the 3.1 section. This may also include defining whether or not to use geographic information standards in the project. 9. Define data management approach After gathering all the necessary information from the previous steps, it is important to evaluate the technology that will be used for data management. This could include a relational database with extensions for geospatial data, a NoSQL database with capabilities for managing geographic data, or a big data infrastructure. The decision should be based on whether the technology can support the project’s needs. To make the right choice, it is important to consider factors such as the technology’s ability to handle input formats, scalability, costs, and systematic deployment requirements. Chapter 3provides a detailed description of some of the most frequently used options. 39 Table 3.3: Geospatial Technologies for web map visualisation /dashboarding. Technology Description OpenLayers An open-source JavaScript library for displaying map data in web browsers. It provides a powerful set of features for building web mapping applications. Leaflet JavaScript library for interactive maps. CesiumJS JavaScript library for 2D maps and 3D globes. D3.js JavaScript library for creating dynamic and interactive data visualizations, including maps. CARTO Offers a cloud-based platform for spatial analysis and visualization, allowing users to create interactive maps and dashboards. Mapbox Offers a variety of mapping tools and services, including Mapbox Studio for designing custom maps and Mapbox GL JS for building interactive web maps. ESRI ArcGIS Online A cloud-based mapping and analysis tool provided by ESRI Inc. for creating and sharing interactive maps and applications. Google Maps Platform Provides APIs for embedding Google Maps into web applications, including features like markers, geocoding, and directions. Tableau A data visualization platform that allows users to create interactive dashboards, including mapping capabilities. Power BI A business analytics tool by Microsoft that includes mapping and dashboarding features for data visualization. Grafana A multi-platform open source analytics and interactive visualization web application. It provides charts, graphs, and alerts for the web when connected to supported data sources, including visualization of spatiotemporal information. 3.1.1 Spatial data visualisation and deployment Spatial data visualisation involves specialised characteristics closely linked to technologies focused on the presentation and interoperability of geographic information. Besides spatial data formats and database storage presented, Web geographic services can be considered part of the implementation technologies specifically to make available spatial data in the Exploitation Zone or analysis within a data science project. Table 3.3 explains most common technologies in spatial data deployment and dashboarding using geospatial information. In a conventional geographic information technology the main way to share geographic information is Geographic viewers or Geovisor. Under these applications it is common to find, as in a conventional approach, services in HTTP (Hypertext Transfer Protocol) protocol, encoding data using key-value pair (KVP) structures or eXtensible Markup Language (XML). The most common geospatial web services are WMS (Web Map Service), WFS (Web Feature Service), WFS allows editing of geographic entities and attributes in vector-type layers, WMTS (Web Map Tile Service) and WCS (Web Coverage Service). WMS implements maps and layers in image format; WMTS provides georeferenced map mosaics, and WCS helps obtain and query coverages oriented to raster layers while preserving the values of each cell. However, data science projects must always enable the exploitation of non-spatial data and link spatial information to graphics or dashboarding components for better analysis, and web services are just a piece of the dashboarding or the exploitation zone. 46 When considering Dashboarding, tools equipped with map visualization functionalities are incorporated, such as Leaflet. In this context, Geographic Information System (GIS) tools provide additional elements to the dashboard, integrating conventional dashboard components, such as graphics. In particular, GIS tools such as those provided by ESRI or other GIS platforms augment the dashboard experience by including spatially oriented graphical representations alongside traditional dashboard elements. This integration ensures a comprehensive, visual presentation of geographic and non-geographic data within a unified dashboard framework. Finally, based on the presented information, Figure 3.1 depicts the deployment technologies associated with the Data management backbone according to the framework proposed. Figure 3.1: Deployment technologies illustration for the Data management backbone of the proposed framework. 3.2 Spatial data analysis Geographic Information Systems are ecosystems where geographic data is the core of data analytics. In any case, these systems can integrate and interoperate with both spatial and alphanumeric data. In some particular cases, making a complete technological deployment of a data science project can rely directly on a Geographic Information System, but given the growth of data science, it is increasingly common to find different tools that allow modular development that respond to specific needs outside the context of GIS. Among the vast number of possibilities, there are commercial and open-source GIS software available, such as ArcGIS, QGIS, GRASS, and SAGA. Most GIS software can interoperate directly with Python or R through libraries, enabling side and side options to perform data analytics. 47 Python is a programming language and stands out as a key language for data modelling, analysis, and visualisation. The extension in the data analysis for geospatial data started with that integration of Python into GIS platforms such as ArcGIS and QGIS. Further, it solidified its role in spatial data processing. Shapely, GeoPandas, SciPy.Spatial and rasterstats are Python libraries that support spatial, temporal, and spatiotemporal data analysis. PySAL addresses statistical modelling and analysis and offers extensive capabilities for spatial analysis of vector and raster data. Figure 3.2: Deployment technologies illustration for the Data analysis backbone of the proposed framework. As stated by [1]spatial libraries and packages in Python were initially developed for singlenode computing environments, requiring parallel and distributed computing platforms such as Hadoop and Spark for large data sets. Several Hadoop APIs, such as Hadoop Streaming, mrjob, Pydoop, and Luigi, allow Python users to access Hadoop MapReduce and HDFS, and PyArrow provides additional client access to HDFS. Overall, the continued growth of Python’s community and capabilities and its integration into GIS platforms position it as a powerful language for spatial data science. Spatial analysis methods implementations presented in python are explained by [2]. Most of these libraries are designed for single-core CPU execution and mainly memory data processing, which limits Python’s scalability for big data. One of the options to expand this capability is PySpark, it introduces overhead when compiling Python code to Java before execution. Additionally, DASK is a native Python library for parallel and distributed computing, which scales seamlessly across distributed nodes and parallelises tasks within a single node. On the other hand, R project, a prominent language in data science, initially conceived for statistical purposes and data analysis, has been evolving in this way since the 90s. The R ecosystem for spatiotemporal data analysis includes sf, raster, terra, stars, and spacetime packages. 48 Tools such as Hadoop Streaming, RHadoop, RHIPE, and ORCH API facilitate MapReduce jobs using Hadoop within the R environment. The Sparklyr and SparkR packages, which provide support for Spark’s machine learning algorithms, are used to interact with Spark. Although Hadoop and Spark lack native support for processing spatial data, custom R packages like Apache Sedona close this gap, allowing spatial analysis. R spatial analysis packages are listed in detail in [34], tools and examples well explained by [6],[31],[23]and [3]. Based on the presented information Figure 3.2 depicts the deployment technologies associated with the Data management backbone according to the framework proposed. 49 Chapter 4 Proof of concept This chapter aims to describe the use of the proposed framework using real data samples from Catalonia. Both of the cases presented aim to illustrate the proposed framework for each spatial data model in the simplest way, understanding that the project’s complexity may vary according to the requirements of SITEP clients. The first use case considers raster model information associated with information processing to obtain unexpected pattern locations, and the second presents an example solution orientation implementation. This exercise intends to provide an end-to-end example of a complete use of the proposed framework and implement an analytical solution representing the process with real data. For each case, each of the proposed stages is explained using the framework: Solution design, Data management, Data analysis, Data governance. Although the objective is to illustrate the process, it does not focus on obtaining the real results in the analytical part of the process as it should in a real client case. In both cases, the data sources are assumed to come from the same date. We develop the solution designed using the framework just for the vector use case according to resources availability in time and infrastructure. Similarly, in spite of the different alternatives available in terms of technologies, both cases use the python library logging for the data governance component application mainly due to restricted access to resources in licenses and hardware. The intended objective is to exemplify the complete implementation of the proposed framework. 4.1 Use case: Vector data In this use case, the main objective is to analise and visualize the walking time required for residents to access the transportation system in the Province of Barcelona. This project leverages geoinformation to reveal insights related to identify areas with the highest and lowest times for accesing bus and train stations and to detect clusters of municipalities with similar conditions. Implementing the proposed framework and following the workflow, the available data sources that allow obtaining information associated with the questions posed are identified to solve this case as depicted in Figure 4.1. 50 Figure 4.1: Implemented system architecture for vector data case. This example framework implementation starts identifying the KPIs of relevance: 1. How large a population is covered by ATM services and which fare zone apply? 2. What is the average walking time to access a bus station? 3. What is the average walking time to access a train station? 4. Are there population clusters where access time is too low or too high? 4.1.1 Data management Based on the availability of data to solve the KPI, we identify the following data sources: • Population Data: Obtained from EUROSTAT JRC-GEOSTAT 2018 is a regular grid map of 1 x 1 km cells reporting the number of residents for the year 2018 for Europe developed in the second half of 2020 by the European Commission Joint Research. • Transport systems stops: Open Street Map Data OpenStreetMap1volunteer geographic information downloaded from Geofabrik web site. • ATM fare zones extracted from the website. • Geospatial entities of Municipality: Available in the Institut Cartogràfic i Geològic de Catalunya administrative boundaries download website (also available in both Open Data portals of the Generalitat de Catalunya and Barcelona) a escala 5000 en formato shapefile. In this case, both, KPI and data sources are within geospatial context. Classifying the data 1https://www.openstreetmap.org/about 51 Table 4.1: Data sources, analysis and task required to calculate KPI of the use case for vector data. KPI Data source related Analysis /Task Exploitation zone Data analysis Population covered by ATM services and fares Population Fare zones Grouping Barplot Map /Layer Train distance average per price zone Population Fare zones Bus stops Distance calculation Barplot Bus distance average per price zone Population Fare zones Train stops Distance calculation Barplot Groups by distance All sources Clustering Map /Layer Moran’s I sources: There is a non-spatial, spatializable source, and three spatial sources. In the case of ATM fare zone, it corresponds to an alphanumerical sources object of spacialisation. The population grid, it is a tessellation of 1km cell size with CRS EPSG: 4326, the stops of the transportation system points with CRS EPSG: 4326 and the administrative boundaries of the municipalities of Catalonia with CRS EPSG: 25831. The list of municipalities covered by ATM comes from a list in txt format that is subject to spatialisation as explained in the 2section, as corresponds to a list of municipalities. In this case, it is necessary to reconcile the data contained in the ATM rate information and the administrative limits provided by the ICGC. A reconciliation source is incorporated in the persistance landing stage that allows the use of spatial methods, which in this case was called LUT referring to a Lookup table. From this information, we can identify that the first three KPIs can be deployed in the data exploitation area and the fourth KPI demands the use of an analysis method. Therefore, the questions expected to be answered by the system and requires a data analysis task. We set out the relationships between data sources and KPIs outlined in Table 4.1. For this example, the distance is calculated using Postgis extension tools. This analysis only illustrates the use case, considering spatial location data. Still, in a more exact case, it should incorporate route and shortest path information using network analysis using the PgRouting extension in Postgres and data of the road network in the area. Identifying that the most granular level of information is found at the population grid level, 1 km2, we decided to use this as Output spatial level. Information display formats in the exploitation area are expected to contain both possibilities, spatial information and alphanumeric information to display statistical graphs. The Data management technology that we decided on in this use case was Postgres with PostGIS extension because it contemplates the operations required for the data exploitation area, it does not require data scaling because they do not vary constantly and it corresponds to a short term analysis. to change scale or extension and through SQL queries that allow 52 responding to the needs described with the PostGIS spatial extension. If recurring analyses are required, the use of a technology more oriented to managing large volumes of data should be reviewed. Regarding Data exploitation technologies, the information is displayed in a Dashboard using the python library Dash that allows you to interact with graphs and maps. 4.1.2 Data analysis As Task and scope of data analysis a spatial quantitative analysis is required for the 4th KPI, we choose the spatial clustering methods approach 4.1. In this case, Moran’s I and Local Moran’s Indexes allow to highlight the spatial clustering based on an statistical text. The clusters are composed of grid elements with similar conditions in high-values or low values. In this case, the method groups. It corresponds to an exploratory data analysis and does not requires regression models or parameters. Based on the roadmap derived from this review Data Exploitation Technologies and taking advantage of Spatial Dataframe in the formatted area, we chose to use Geopandas and PySAL to display Local statistics. The result expected shows the trend of the variable average walking time to access the transportation system in the area covered by ATM. The proposed analysis, given the objective of the proof of concept, has a Life cycle of the analysis method of execution and the result of the implementation of the model corresponding to a Colab Notebook to explain the result of the spatial statistical method. 4.1.3 Solution implementation Data management Below are some of the outputs obtained during the implementation process of the development of the technological solution made from the framework proposal for this use case, for each stage. The development of the tools provided ideas of real cases and challenges that could be face implementing the framework. Figures shown in this section are illustrative, the annex 5contains the implementation scripts. The governance files and outputs of Dashboards and model deployment are available at: this link. In the data management, data sources were stored in a local file system, in this phase, we classify the data as specified, the results are illustrated in Figure 4.4. The development of the governance application, both the data collectors and the data persistance loader, using Python 3 and relying on Geopandas library tools to obtain the necessary records associated with the geospatial component. Also within this application the log 53 Figure 4.2: Data sources and landing zone folder structure for vector use case. files are extracted. Figure 4.4 shows a part of the implemented code and the complete code as mentioned is presented in the Annex 5. Figure 4.3: Data Collector and log application for vector use case. According to the arquitecture we created a Postgres DB with postgis extension using Postgres v.16.0. Likewise, using SQL queries we prepare the ready-to-use data for the exploitation zone„ in this case we only used the ST_Distance method of postgis to illustrate the calculations of distances. Figure 4.1.3 present the SQL query to find the distances between each grid centroid and the closest bus or train station and Figure 4.4 illustrates the output of this intermediate step. 54 1/*KPI 4 Ready to use population grid polygon with minimum distance*/ 2alter table public."POPULATION_25831" add column geom_centroids geometry(Point, 25831) 3update public."POPULATION_25831" set geom_centroids = st_centroid(geom) 4drop index population_centroids_idx 5create index population_centroids_idx on population using gist (geom_centroids) 6 7create table UC_VECTOR_DA_GRID_DISTANCE_STATION_POP as 8select p.fid, 9p."TOT_P_2018" as tot_p_2018, 10 ST_Distance(p.geom_centroids, t.geom) as dist, 11 p.geom_centroids as geom_population, 12 t.geom as geom_transport, 13 st_makeline(p.geom_centroids, t.geom) as geom_line 14 from population as p 15 join lateral ( 16 select tt.geom 17 from transport as tt 18 where fclass in ('railway_station','bus_stop') 19 order by p.geom_centroids <-> tt.geom 20 limit 1 21 )as t 22 on true Listing 4.1: Tessellation centroid to closest station distance calculation Figure 4.4: SQL Output - Minimum distances from grid population centroids to bus /train station. 55 This use case uses the expected metadata artefacts exposed in Chapter 2. Also, It is expected that Data Collectors and source registries do not imply additional or complex applications, given there are only two data sources. The Data Persistence Loader, as mentioned before, do not store the information but links the data source Path to the formatted zone and create the Key-Value document using metadata available in the Data Source. 4.3 Results and Evaluation The main results of implementing the framework and evaluation associated with the instantiation of the technological solution for the use cases presented allowed us to validate whether the approach was feasible and identify possible improvements or particularities associated with each geographic information model in the following aspects: Framework implementation stage: • We were able to implement the framerwork end-to-end, obtaining an architecture design for each geographic data model. • We identified specific considerations for both vector and raster data. • Spatial data, with overlay capabilities requirement, requires advanced capabilities for exploitation zone. • It is required big Data infrastructure development for raster data management to suggests potential solutions. Technical findings: • Geospatial Data in Persistent Landing using Key-Value Store implies a high resource demand, especially for raster data implementation.Duplicity in raster data may be spaceconsuming and not always the best option for persistent fast querying in a new format. • Python and R libraries for reading spatial data and generating required metadata • Additional considerations for geospatial data quality in data governance processes can rely on geospatial deployment technologies tools. Proof of concept using real data: • We had limited exploration of options due to the lack of access to licensed tools and time constraints in the project. • Inclusion of Reference System and Proposed Model facilitated queries in data loaders tasks during solution instantiation. It implies redundancy, but a trace from the data sources identification stage optimizes development in both raster and vector use cases. 62 • Vector data is lighter, but analysis algorithms may be more complex. • Raster data involves simpler calculations, but volume and density exceed conventional capabilities. • In spatial data projects, generating just one layer of geographic information from the model deployment can constitute an entire data science project, addressing client needs. • Python and R libraries for reading spatial data and generating required metadata were easy and practical to use. • Significant difference observed between establishing architecture for vector and raster data. In the project development, identified several lines of further study or aspects outside the temporal scope, detailed in Chapter 5. 63 Chapter 5 Conclusions and future work We have provided a comprehensive overview, including a conceptual background, a proposed framework, a step-by-step implementation guide, and a range of technologies to support the development of spatial data science projects. Additionally, we have conducted a proof of concept by implementing the framework in two use cases and developed test applications to assess its effectiveness. Our research suggests that there are numerous possibilities for creating data science projects using a geospatial approach within an extensive framework. The particular considerations that we have included can help enhance the design process and improve the management and analysis of such projects. As in the conventional approach, there is more than a one-size-fits-all solution when deploying spatial technology. Every spatial project is unique and requires a tailored approach considering various factors. The decision-making process is primarily guided by client preferences, including their motivations and available resources. Critical considerations include balancing economic and technical resources, sticking to project timelines, and ensuring alignment with the solution’s systematic needs. The only differential factor here is the tools’ capacity to manage data and perform queries based on combinations of geographic information layers and analysis using spatial-specific methods. In practice, it’s become clear that client preferences and motivations drive the project’s core from the KPI definition, followed by data sources. The choice of technology plays a secondary role in the decision-making process. Depending on the complexity of the task, it may constrain alternatives. Another finding was the growing significance of explainability and reproducibility tools in decoding algorithmic processes. These tools, exemplified in spatial data analytics by different authors, are in concordance with the governance component and enhance transparency and comprehension, ultimately contributing to the reproducibility of results mainly in machine learning and deep learning approaches. For algorithms that use statistical methods, quantifying uncertainty or error can explain the method’s effectiveness. 64 The ever-evolving nature of data science projects, particularly those involving spatial data, necessitates a continuous adaptation and testing paradigm. Regular benchmarking, preliminary trials with sample data, and iterative testing are recommended to assess the viability of the chosen technologies. This approach allows for real-world testing and feedback, facilitating refinements aligned with project objectives. Future work While this project aimed to make progress in developing a comprehensive framework for spatial data science projects, it inevitably encountered limitations tied to its temporal scope. Notably, topics relevant to data science projects, such as ethical considerations and user mapping, were intentionally left outside the project. One potential area for future exploration involves the ethical considerations of spatial data science projects. An extended framework could delve into ethics by developing codes of conduct at both the business and project levels. This extension would encourage a comprehensive evolution, promoting transparency principles that guide both producers and users of the framework. Incorporating ethical considerations is crucial to ensure responsible and moral practices, recognizing the significance of ethics in the dynamic landscape of data science. By connecting this with the decisions associated with the territory, the framework is proposed as a contribution to transparency processes in the management of the territory and the appropriate use of technologies for the benefit of citizens. Moreover, the identified expansion branches present promising directions for enriching the proposed framework. First and foremost, handling scalability in spatial contexts for both vector and raster models, mainly when used together, emerges as a crucial concern. A future investigation should explore scalable solutions tailored to the unique challenges posed by spatial data, contributing to the framework’s adaptability in diverse scenarios. There is a chance to assess and compare different prediction techniques explicitly used in spatial contexts. Further investigation could be carried out to comprehensively examine these methods, highlighting their effectiveness in spatial scenarios. This in-depth understanding would improve the predictive abilities of the framework, ensuring it remains relevant and valuable for a range of spatial data science projects. Finally, it would be interesting to compare the results obtained from traditional and spatial analytics using the data analysis framework. It will help identify the differences and similarities in the outcomes generated by these approaches. Such analysis can improve the recommendations and guidelines provided by the framework and provide valuable insights to practitioners for making informed decisions in spatial data science projects. 65 Bibliography [1]Md Mahbub Alam, Luis Torgo and Albert Bifet. “A survey on spatio-temporal data analytics systems”. In: ACM Computing Surveys 54.10s (2022), pp. 1–38. [2]Luc Anselin, Xun Li and Julia Koschinsky. “GeoDa, from the desktop to an ecosystem for exploring spatial data”. In: Geographical Analysis 54.3 (2022), pp. 439–466. [3]Luc Anselin, Ibnu Syabri and Youngihn Kho. “GeoDa: an introduction to spatial data analysis”. In: Handbook of applied spatial analysis: Software tools, methods and applications. Springer, 2009, pp. 73–89. [4]Claudia Patricia Ayala Martínez et al. “DOGO4ML: Development, operation and data governance for ML-based software systems”. In: Joint Proceedings of RCIS 2022 Workshops and Research Projects Track: co-located with the 16th International Conference on Research Challenges in Information Science (RCIS 2022): Barcelona, Spain, May 17-20, 2022. CEUR-WS. org. 2022. [5]Adrian Baddeley, Ege Rubak and Rolf Turner. Spatial point patterns: methodology and applications with R. CRC press, 2015. [6]Roger S Bivand et al. Applied spatial data analysis with R. Vol. 747248717. Springer, 2008. [7]Martin Breunig et al. “Geospatial data management research: Progress and future directions”. In: ISPRS International Journal of Geo-Information 9.2 (2020), p. 95. [8]Bénédicte Bucher et al. “EuroGeographics and EuroSDR”. In: (2023). [9]Serena Coetzee et al. “Open geospatial software and data: A review of the current state and a perspective into the future”. In: ISPRS International Journal of Geo-Information 9.2 (2020), p. 90. [10]Rodolphe Devillers, Robert Jeansoulin and Michael F Goodchild. Fundamentals of spatial data quality. ISBN: 1905209568. ISTE London, 2006. [11]Julian Ereth. “DataOps-Towards a Definition.” In: LWDA 2191 (2018), pp. 104–112. [12]Martin Ester, Hans-Peter Kriegel and Jörg Sander. “Knowledge discovery in spatial databases”. In: Annual Conference on Artificial Intelligence. Springer. 1999, pp. 61–74. [13]Manfred M Fischer and Arthur Getis. Handbook of applied spatial analysis: software tools, methods and applications. Springer, 2010. [14]Inan Gür et al. “Requirements for DataOps to foster Dynamic Capabilities in OrganizationsA mixed methods approach”. In: 2022 IEEE 24th Conference on Business Informatics (CBI). Vol. 1. IEEE. 2022, pp. 166–175. [15]Otto Huisman, Rolf A de By et al. “Principles of geographic information systems”. In: ITC Educational Textbook Series 1 (2009), p. 17. 66 [16]Volker Janssen. “Understanding coordinate reference systems, datums and transformations”. In: (2009). [17]Zhe Jiang. “A Survey on Spatial and Spatiotemporal Prediction Methods”. In: arXiv preprint arXiv:2012.13384 (2020). [18]Petar Jovanovic et al. “Quarry: a user-centered big data integration platform”. In: Information Systems Frontiers 23 (2021), pp. 9–33. [19]M Kanevski, A Pozdnukhov and V Timonin. “Machine learning algorithms for geospatial data. Applications and software tools”. In: (2008). [20]Slava Kisilevich et al. Spatio-temporal clustering. Springer, 2010. [21]Vijay Kotu and Bala Deshpande. Data science: concepts and practice. Morgan Kaufmann, 2018. [22]Werner Kuhn. “Core concepts of spatial information for transdisciplinary research”. In: International Journal of Geographical Information Science 26.12 (2012), pp. 2267–2276. [23]Xun Li and Luc Anselin. Rgeoda: R library for spatial data analysis. 2022. [24]Roger Lott. “Geographic information-well-known text representation of coordinate reference systems”. In: (2015). [25]Robin Lovelace, Jakub Nowosad and Jannes Muenchow. Geocomputation with R. CRC Press, 2019. [26]Nikos Mamoulis. Spatial data management. Springer Nature, 2022. [27]Samuel de França Marques and Cira Souza Pitombo. “Transit ridership modeling at the bus stop level: comparison of approaches focusing on count and spatially dependent data”. In: Applied spatial analysis and policy 16.1 (2023), pp. 277–313. [28]Dimitar Misev, Mihaela Rusu and Peter Baumann. “A semantic resolver for coordinate reference systems”. In: Web and Wireless Geographical Information Systems: 11th International Symposium, W2GIS 2012, Naples, Italy, April 12-13, 2012. Proceedings 11. Springer. 2012, pp. 47–56. [29]Karen A Mulcahy and Keith C Clarke. “Symbolization of map projection distortion: a review”. In: Cartography and geographic information science 28.3 (2001), pp. 167–182. [30]Sergi Nadal et al. “Operationalizing and automating Data Governance”. In: Journal of big data 9.1 (2022), pp. 1–31. [31]Edzer Pebesma and Roger Bivand. Spatial data science: With applications in R. CRC Press, 2023. [32]Sergio Rey, Dani Arribas-Bel and Levi John Wolf. Geographic data science with python. CRC Press, 2023. [33]Brian D Ripley. “Spatial statistics in R”. In: R news 1.2 (2001), pp. 14–15. [34]Jakub Nowosad Roger Bivand. Analysis of Spatial Data.https://cran.r-project. org/web/views/Spatial.html.[Online; accessed 10-January-2024]. 2023. [35]Oscar Romero. Class notes: Data Science End-to-End Project. Nov. 2022. [36]Miquel Sànchez-Marrè. “Intelligent Decision Support Systems”. In: Intelligent Decision Support Systems. Springer, 2022, pp. 77–116. 67 [37]Markus Schneider. Spatial data types for database systems: finite resolution geometry for geographic information systems. Springer, 1997. [38]Patrick Schratz et al. “Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data”. In: Ecological Modelling 406 (2019), pp. 109–120. [39]Greg Scott and Abbas Rajabifard. “Sustainable development and geospatial information: a strategic framework for integrating a global policy agenda into national geospatial capabilities”. In: Geo-spatial information science 20.2 (2017), pp. 59–76. [40]H Seeger. “Spatial referencing and coordinate systems”. In: Geographical information systems 1 (1999), pp. 757–766. [41]Zhicheng Shi and Lilian SC Pun-Cheng. “Spatiotemporal data clustering: A survey of methods”. In: ISPRS international journal of geo-information 8.3 (2019), p. 112. [42]Rüdiger Wirth and Jochen Hipp. “CRISP-DM: Towards a standard process model for data mining”. In: Proceedings of the 4th international conference on the practical applications of knowledge discovery and data mining. Vol. 1. Manchester. 2000, pp. 29–39. [43]Brett Wujek, Patrick Hall and Funda Günes. “Best practices for machine learning applications”. In: SAS Institute Inc (2016), pp. 1–23. [44]Yiqun Xie et al. “Transdisciplinary foundations of geospatial data science”. In: ISPRS International Journal of Geo-Information 6.12 (2017), p. 395. 68 Appendix Use case vector code Data management 1## SOURCES --> LANDING 2!pip install pyogrio 3 4import os 5import shutil 6from datetime import datetime 7import pandas as pd 8import geopandas as gpd 9import json 10 import logging 11 import pathlib 12 import fiona; help(fiona.open) 13 import pyogrio; help(pyogrio.read_dataframe) 14 from shapely.geometry import shape 15 import ast 16 17 formatter = logging.Formatter('%(asctime)s,%(msecs)d %(name)s %(levelname)s %(message)s', datefmt='%H:%M:%S') 18 19 def setup_logger(name, log_file, level=logging.INFO): 20 governance_path = '/content/drive/MyDrive/UC_VECTOR/03_DG' 21 if not os.path.exists(governance_path): 22 os.makedirs(governance_path) 23 24 file_handler = logging.FileHandler(log_file) 25 file_handler.setFormatter(formatter) 26 27 logger = logging.getLogger(name) 28 logger.setLevel(level) 29 logger.addHandler(file_handler) 30 31 return logger 32 33 def lSources_lTemp(sources_path, temp_path): 34 35 files = os.listdir(sources_path) 36 txt_files = [file for file in files if file.endswith('.txt')] 69 37 geo_files = [file for file in files if file.endswith('.shp')] 38 39 base_directory = temp_path 40 Non_spatial = "Data" 41 Spatial = "sData" 42 SpatioTemporal = "stData" 43 44 temp_path_non_spatial = os.path.join(base_directory, Non_spatial) 45 temp_path_spatial = os.path.join(base_directory, Spatial) 46 temp_path_ST = os.path.join(base_directory, SpatioTemporal) 47 48 os.makedirs(temp_path_non_spatial, exist_ok=True) 49 os.makedirs(temp_path_spatial, exist_ok=True) 50 os.makedirs(temp_path_ST, exist_ok=True) 51 52 timestamp = datetime.now().strftime("%Y%m%d%H%M%S") 53 Wildcard = datetime.now().strftime("%Y-%m-%d") 54 55 for file in txt_files: 56 input_sources = os.path.join(sources_path, file) 57 logger.info(f"Reading: {input_sources}") 58 59 # Sources to Temporal of "ATM" Source 60 if 'atm'in file.lower(): 61 SourceName = "ATM_Zones" 62 # Data Governance Logs 63 # Input by user 64 DCR_logger.info(f"DestPath:{temp_path_non_spatial}") 65 66 Sources_logger.info(f"SourceName: '{SourceName}'") 67 Sources_logger.info(f"Wildcard: '{input_sources}'") 68 Sources_logger.info(f"ValidFrom: '11/2023'") 69 70 Sources_logger.info(f"MaintainancePolicy: KeepLast") 71 Sources_logger.info(f"SourceURL: https://www.atm.cat/sistema-tarifari-integrat/ sistema-de-transport/mapa-de-la-zonificacio") 72 73 logger.info(f"FilePath: '{input_sources}'") 74 logger.info(f"Timestamp: '{Wildcard}'") 75 76 file_name = f"{file}".replace('.txt','') 77 DCR_logger.info(f"\u007b'KeyAttribute:''{timestamp}'_{SourceName}_'{file_name}'") 78 Key = f"{timestamp}_{SourceName}_{file_name}" 79 json_file_name = f"{Key}.json" 80 81 # Read txt 82 df = pd.read_csv(input_sources, sep=";") 83 print("readingatm") 84 schema = list(df.columns) 85 Sources_logger.info(f"Metadata: '{schema}'") 86 70 87 file_extension = pathlib.Path(file).suffix 88 Sources_logger.info(f"SourceType: '{file_extension}'") 89 90 # Create the new JSON file path 91 json_file_path = os.path.join(temp_path_non_spatial, json_file_name) 92 # Create a dictionary 93 dict_data_temp = {'Key': Key, 'Value': input_sources, 'Metadata': schema} 94 json_string = json.dumps(dict_data_temp) 95 #json_data_temp = f" \u007b'Key':{Key},'Value': {input_sources}, 'Metadata': { schema}\u007d" 96 97 # Write JSON data to the new file 98 with open(json_file_path, 'w') as json_file: 99 json.dump(json_string,json_file) 100 101 #print(f"File '{file}'converted to JSON and saved as '{json_file_name}'") 102 Sources_logger.info(f"File '{file}'read and JSON and saved as '{json_file_name}'") 103 logger.info(f"File '{file}'read and JSON to JSON and saved as '{json_file_name}'") 104 105 elif 'lut'in file.lower(): 106 SourceName = "LUT" 107 # Data Governance Logs 108 # Input by user 109 DCR_logger.info(f"DestPath:{temp_path_non_spatial}") 110 111 Sources_logger.info(f"SourceName: '{SourceName}'") 112 Sources_logger.info(f"Wildcard: '{input_sources}'") 113 Sources_logger.info(f"ValidFrom: '11/2023'") 114 115 Sources_logger.info(f"MaintainancePolicy: KeepLast") 116 117 logger.info(f"FilePath: '{input_sources}'") 118 logger.info(f"Timestamp: '{Wildcard}'") 119 120 file_name = f"{file}".replace('.txt','') 121 DCR_logger.info(f"\u007b'KeyAttribute:''{timestamp}'_{SourceName}_'{file_name}'") 122 Key = f"{timestamp}_{SourceName}_{file_name}" 123 json_file_name = f"{Key}.json" 124 125 # Read txt 126 df = pd.read_csv(input_sources,sep='\t', engine='python') 127 print("readingLUT") 128 schema = list(df.columns) 129 Sources_logger.info(f"Metadata: '{schema}'") 130 131 file_extension = pathlib.Path(file).suffix 132 Sources_logger.info(f"SourceType: '{file_extension}'") 133 134 # Create the new JSON file path 135 json_file_path = os.path.join(temp_path_non_spatial, json_file_name) 136 # Create a dictionary 71 429 elif 'population'in Key_ORG.lower(): 430 Layer = "POPULATION" 431 SourceName = "JRC" 432 else: 433 print("Error in name") 434 435 current_crs = gdf.crs 436 EPSG = target_crs.replace(':','_') 437 438 if current_crs != target_crs: 439 # Reproject the GeoDataFrame to the target CRS 440 gdf = gdf.to_crs(target_crs) 441 Landing_logger.info(f"GeoDataFrame has been reprojected to {target_crs}") 442 else: 443 print("GeoDataFrame already has the target CRS") 444 445 geojson_file_name = f"{timestamp}_{SourceName}_{Layer}_VECTOR_{EPSG}.geojson" 446 geojson_file_path = os.path.join(persistent_path, geojson_file_name) 447 448 gdf.to_file(geojson_file_path, driver='GeoJSON') 449 logger.info(f"GeoDataFrame has been written to {geojson_file_path}") 450 Landing_logger.info(f"GeoDataFrame has been written to {geojson_file_path}") 451 452 if __name__ == "__main__": 453 governance_path = '/content/drive/MyDrive/UC_VECTOR/03_DG' 454 temp_path = '/content/drive/MyDrive/UC_VECTOR/01_DM/01_LANDING/01_TEMPORAL' 455 persistent_path = '/content/drive/MyDrive/UC_VECTOR/01_DM/01_LANDING/02_PERSISTENT' 456 457 print("Setting up logging...") 458 Landing_logger = setup_logger('landing', os.path.join(governance_path, 'LogLanding.txt')) 459 logger = setup_logger('LogDataCollector', os.path.join(governance_path, 'LogDataCollector .log')) 460 DCR_logger = setup_logger('DataCollectorRegistry', os.path.join(governance_path, ' DataCollectorRegistry.log')) 461 462 lTemp_lPers(temp_path, persistent_path) 463 print("Temporal to persistent landing executed successfully.") Listing 5.1: Data collector 1## LANDING ---> PERSISTENT 2 3def lTemp_lPers(temp_path, persistent_path): 4timestamp = datetime.now().strftime("%Y%m%d%H%M%S") 5 6# Process files in 'Data'directory 7temp_Data_path = os.path.join(temp_path, 'Data') 8 9for file_name in os.listdir(temp_Data_path): 10 file_path = os.path.join(temp_Data_path, file_name) 11 78 12 with open(file_path) as f: 13 content = f.read() 14 content = content.strip('"') 15 content = content.replace('\\"','\"') 16 d = json.loads(content) 17 logger.info(f"json loaded: '{d}'") 18 19 try: 20 df = pd.read_csv(d["Value"]) 21 print(df) 22 except pd.errors.ParserError as e: 23 print(f"Error in file {file_name}: {e}") 24 25 Key_ORG = d["Key"] 26 logger.info(f"Reading...{Key_ORG}") 27 28 txt_file_name = f"{timestamp}_{Key_ORG}.txt" 29 txt_file_path = os.path.join(persistent_path, txt_file_name) 30 31 df.to_csv(txt_file_path, index=False, header=True, sep='\t') 32 logger.info(f"DataFrame has been written to {txt_file_path}") 33 34 35 # Process files in 'sData'directory 36 temp_sData_path = os.path.join(temp_path, 'sData') 37 VECTOR_SOURCES = [os.path.join(temp_sData_path, f) for fin os.listdir(temp_sData_path) if 'vector'in f.lower()] 38 39 for file in VECTOR_SOURCES: 40 with open(file) as f: 41 target_crs = 'EPSG:25831' 42 43 content = f.read() 44 content = content.strip('"') 45 content = content.replace('\\"','\"') 46 47 d = json.loads(content) 48 logger.info(f"json loaded: '{d}'") 49 gdf = gpd.read_file(d["Value"]) 50 Key_ORG = d["Key"] 51 logger.info(f"Reading...{Key_ORG}") 52 53 if 'administratives'in Key_ORG.lower(): 54 Layer = "MUNICIPIS" 55 SourceName = "ICGC" 56 elif 'transport'in Key_ORG.lower(): 57 Layer = "STATIONS" 58 SourceName = "OSM" 59 elif 'population'in Key_ORG.lower(): 60 Layer = "POPULATION" 61 SourceName = "JRC" 79 62 else: 63 print("Error in name") 64 65 current_crs = gdf.crs 66 EPSG = target_crs.replace(':','_') 67 68 if current_crs != target_crs: 69 # Reproject the GeoDataFrame to the target CRS 70 gdf = gdf.to_crs(target_crs) 71 Landing_logger.info(f"GeoDataFrame has been reprojected to {target_crs}") 72 else: 73 print("GeoDataFrame already has the target CRS") 74 75 geojson_file_name = f"{timestamp}_{SourceName}_{Layer}_VECTOR_{EPSG}.geojson" 76 geojson_file_path = os.path.join(persistent_path, geojson_file_name) 77 78 gdf.to_file(geojson_file_path, driver='GeoJSON') 79 logger.info(f"GeoDataFrame has been written to {geojson_file_path}") 80 Landing_logger.info(f"GeoDataFrame has been written to {geojson_file_path}") 81 82 if __name__ == "__main__": 83 governance_path = '/content/drive/MyDrive/UC_VECTOR/03_DG' 84 temp_path = '/content/drive/MyDrive/UC_VECTOR/01_DM/01_LANDING/01_TEMPORAL' 85 persistent_path = '/content/drive/MyDrive/UC_VECTOR/01_DM/01_LANDING/02_PERSISTENT' 86 87 print("Setting up logging...") 88 Landing_logger = setup_logger('landing', os.path.join(governance_path, 'LogLanding.txt')) 89 logger = setup_logger('LogDataCollector', os.path.join(governance_path, 'LogDataCollector .log')) 90 DCR_logger = setup_logger('DataCollectorRegistry', os.path.join(governance_path, ' DataCollectorRegistry.log')) 91 92 lTemp_lPers(temp_path, persistent_path) 93 print("Temporal to persistent landing executed successfully.") Listing 5.2: Data persistant loader 1select * 2from LUT 3 4select * 5from LUT l 6 7SELECT * 8FROM public."MUNICIPIS_25831" M 9 10 select 11 lut.id, 12 lut.municipi, 13 lut.nom_muni, 14 lut.nom_original_atm, 80 15 a.Tarifa, 16 pop.grd_id, 17 pop.tot_p_2018, 18 pop.geom as pop_geometry, 19 mb.codimuni 20 from 21 lut 22 join 23 public."ATM" aon lut.nom_original_atm = a.Nombre 24 join 25 public."MUNICIPIS_25831" mb on LUT.municipi = mb.codimuni 26 join 27 public."POPULATION_25831" pop on ST_Intersects(mb.geom, pop.geom); 28 29 30 /*KPI 1 Population per fare zone*/ 31 with TrainDistances as ( 32 select 33 a.Tarifa AS tarifa, 34 pop.grd_id, 35 ST_Distance(osm.geom, ST_Centroid(pop.geom)) AS distance_train 36 from 37 public."OSM_STATIONS_25831" osm 38 cross join 39 public."POPULATION_25831" pop 40 join 41 lut ON ST_Intersects(osm.geom, pop.geom) 42 join 43 public."ATM" aON lut.nom_original_atm = a.Nombre 44 where 45 osm.fclass = 'railway_station' 46 ) 47 select 48 t.tarifa, 49 AVG(t.distance_train) AS avg_min_distance_train, 50 SUM(pop.tot_p_2018) AS sum_tot_p_2018 51 from 52 TrainDistances t 53 join 54 public."POPULATION_25831" pop ON t.grd_id = pop.grd_id 55 join 56 public."MUNICIPIS_25831" mb ON pop.codi_muni = mb.codimuni 57 group by 58 t.tarifa; 59 60 61 /*KPI 2 Train distance average per price zone*/ 62 select 63 tarifa, 64 MIN(min_distance_bus) AS min_distance_train, 65 SUM(sum_tot_p_2018) AS sum_tot_p_2018 81 66 from ( 67 select 68 a.tarifa, 69 MIN(ST_Distance(osm.geom, ST_Centroid(pop.geom))) AS min_distance_train, 70 MIN(ST_Distance(osm.geom, ST_Centroid(pop.geom)) *15 / 1000) AS walking_time_train, 71 SUM(pop.tot_p_2018) AS sum_tot_p_2018 72 from 73 public."OSM_STATIONS_25831" osm 74 cross join 75 public."POPULATION_25831" pop 76 join 77 public."ATM_Municipis" aON ST_Intersects(osm.geom, pop.geom) 78 where 79 osm.fclass = 'railway_station' 80 group by 81 a.tarifa, pop.grd_id 82 )AS subquery 83 group by 84 tarifa; 85 86 87 88 /*KPI 3 Bus distance average per price zone*/ 89 select 90 tarifa, 91 MIN(min_distance_bus) as min_distance_train, 92 SUM(sum_tot_p_2018) as sum_tot_p_2018 93 from ( 94 select 95 a.tarifa, 96 MIN(ST_Distance(osm.geom, ST_Centroid(pop.geom))) AS min_distance_bus, 97 MIN(ST_Distance(osm.geom, ST_Centroid(pop.geom)) *15 / 1000) AS walking_time_bus, 98 SUM(pop.tot_p_2018) AS sum_tot_p_2018 99 from 100 public."OSM_STATIONS_25831" osm 101 cross join 102 public."POPULATION_25831" pop 103 join 104 public."ATM_Municipis" aON ST_Intersects(osm.geom, pop.geom) 105 where 106 osm.fclass = 'bus_stop' 107 group by 108 a.tarifa, pop.grd_id 109 )AS subquery 110 group by 111 tarifa; 112 113 114 /*KPI 4 Ready to use population grid polygon with minimum distance*/ 115 alter table public."POPULATION_25831" add column geom_centroids geometry(Point, 25831) 116 update public."POPULATION_25831" set geom_centroids = st_centroid(geom) 82 117 drop index population_centroids_idx 118 create index population_centroids_idx on population using gist (geom_centroids) 119 120 121 create table UC_VECTOR_DA_GRID_TARIFA_POP as 122 select p.fid, 123 p."TOT_P_2018" as tot_p_2018, 124 ST_Distance(p.geom_centroids, t.geom) as dist, 125 p.geom_centroids as geom_population, 126 t.geom as geom_transport, 127 st_makeline(p.geom_centroids, t.geom) as geom_line 128 from population as p 129 join lateral ( 130 select tt.geom 131 from transport as tt 132 where fclass in ('railway_station','bus_stop') 133 order by p.geom_centroids <-> tt.geom 134 limit 1 135 )as t 136 on true 137 138 139 pg_dump -U postgres db_formatted_use_case_vector > db_formatted_use_case_vectorexport_.pgsql Listing 5.3: Data Formatter Data analysis 1## NOTEBOOK DATA ANALYSIS 2## REFERENCE: 3 4!pip install pysal 5!pip install libpysal 6 7from google.colab import drive 8drive.mount('/content/drive') 9 10 # Commented out IPython magic to ensure Python compatibility. 11 # %matplotlib inline 12 13 import matplotlib.pyplot as plt 14 from libpysal.weights.contiguity import Queen 15 from libpysal import examples 16 import numpy as np 17 import pandas as pd 18 import geopandas as gpd 19 import os 20 import splot 21 22 gdf = gpd.read_file('/content/drive/MyDrive/UC_VECTOR/01_DM/03_EXPLOITATION/ UC_VECTOR_DA_GRID_TARIFA_POP.gpkg') 83 23 gdf['AVG'] = gdf['Walking_time_bus']+gdf['Walking_time_train']/2 24 25 import folium 26 27 m = gdf.explore( 28 column="AVG", 29 scheme="naturalbreaks", 30 legend=True, 31 k=6, 32 tooltip=False, 33 legend_kwds=dict(colorbar=False), 34 cmap = 'cool', 35 name="Average time" 36 ) 37 m 38 39 y = gdf['AVG'].values 40 w = Queen.from_dataframe(gdf) 41 w.transform = 'r' 42 43 from esda.moran import Moran 44 45 w = Queen.from_dataframe(gdf) 46 moran = Moran(y, w) 47 moran.I 48 49 from splot.esda import moran_scatterplot 50 fig, ax = moran_scatterplot(moran, aspect_equal=True) 51 plt.show() 52 53 from splot.esda import plot_moran 54 55 plot_moran(moran, zstandard=True, figsize=(10,4)) 56 plt.show() 57 58 moran.p_sim 59 60 from splot.esda import moran_scatterplot 61 from esda.moran import Moran_Local 62 63 # calculate Moran_Local and plot 64 moran_loc = Moran_Local(y, w) 65 fig, ax = moran_scatterplot(moran_loc) 66 ax.set_xlabel('AVG') 67 ax.set_ylabel('Spatial Lag of AVG') 68 plt.show() 69 70 fig, ax = moran_scatterplot(moran_loc, p=0.05) 71 ax.set_xlabel('AVG') 72 ax.set_ylabel('Spatial Lag of AVG') 73 plt.show() 84