Full text
TESE DE DOUTORAMENTO ACQUISITION AND DECLARATIVE ANALYTICAL PROCESSING OF SPATIO-TEMPORAL OBSERVATION DATA Sebastián Villarroya Fernández ESCOLA DE DOUTORAMENTO INTERNACIONAL PROGRAMA DE DOUTORAMENTO EN INVESTIGACIÓN EN TECNOLOXÍAS DA INFORMACIÓN SANTIAGO DE COMPOSTELA ANO 2018
DECLARACIÓN DO AUTOR DA TESE Acquisition and Declarative Analytical Processing of Spatio-Temporal Observation Data D. Sebastián Villarroya Fernández Presento miña tese, seguindo o procedemento adecuado ao Regulamento, e declaro que: 1) A tese abarca os resultados da elaboración do meu traballo. 2) No seu caso, na tese se fai referencia as colaboracións que tivo este traballo. 3) A tese é a versión definitiva presentada para a súa defensa e coincide ca versión enviada en formato electrónico. 4) Confirmo que a tese non incorre en ningún tipo de plaxio de outros autores nin de traballos presentados por min para a obtención de outros títulos. En Santiago de Compostela, 20 de Xullo de 2018 Asdo. Sebastián Villarroya Fernández
AUTORIZACIÓN DO DIRECTOR / TITOR DA TESE Acquisition and Declarative Analytical Processing of Spatio-Temporal Observation Data D. José Ramón Ríos Viqueira D. José Manuel Cotos Yáñez INFORMA/N: Que a presente tese, correspóndese co traballo realizado por D. Sebastián Villarroya Fernández, baixo a miña dirección, e a utorizo a súa presentación , considerando que reúne os r equisitos esixidos no R egulamento de Estudos de Doutoramento da USC, e que como director desta non incorre nas causas de abstención establecidas na Lei 40/2015. En Santiago de Compostela, 20 de Xullo de 2018 Asdo. José Ramón Ríos Viqueira Asdo. José Manuel Cotos Yáñez
A Sabela, Álex y Sandra
Maybe the paths that you each shall tread are already laid before your feet, though you do not see them. Lady Galadriel
xiv I also want to remember my parents. For all the sacrifice they had to do to give me a good academic training. For the education they gave me. For so many things. To my mother, for being a fighter, for never giving up, for overcoming the unbeatable. To my father, who taught me the most important lesson: to love life above all things. To my grandparents. Although they are not directly related to this thesis, they have always supported everything I have wanted to do. To the colleagues of the COGRADE research group. For making me feel an important part of the group from the beginning. For all the good times at work and, specially, outside of it. And for all the help given to me so many times. This would not be the same without you. To the Research Center on Information Technologies (CiTIUS). For all the administrative help provided. For the technical help in many projects. And, above all, for co-funding my attendance at the 1st Summer School on Data Science, organized by ACM Sigmod. To the Diputación de A Coruña. For granting me the 2012 Research Scholarship. To the Galicia Supercomputing Center (CESGA). For all the infrastructure and help provided during the execution of the experiments of this thesis. And especially I want to thank Javier Cacheiro López. Finally, I want to thank the organizations and institutions that supported, contributed or funded the following research projects related to this thesis: –Patrimonio cultural de la Eurorregión Galicia-Norte de Portugal: Valoración e Innovación. GEOARPAD (0358_GEOARPAD_1_E). INTERREG V-A España-Portugal (POCTEP) Program, 2014-2020. European Regional Development Fund (ERDF). European Union. –FUTURE-HDA: Internet del Futuro en el Hogar Digital Asistencial (ITC-20113075). Center for the Development of Industrial Technology (CDTI) and FEDER-INNTERCONECTA Program. –Desarrollo de un servicio de análisis espacial y su aplicación en la implementación de un sistema de gestión de hábitats humanos (TIN2010-21246-C02-02). National Research Program, Ministry of Science and Innovation. –Proyecto Minieólica: Fomento de la tecnología eólica de pequeña potencia. Subproyecto 3.4: Evaluación y diseño de un proyecto demostrador de energía eólica e hidrógeno (PS-120000-2006-5). Ministry of Science and Innovation.
xv –Proyecto Peixe Verde. Subproyecto 1: Toma de Datos (PSE-370300-2007-1). Ministry of Education and Science. –Rede de Tecnoloxías Cloud e Big Data para HPC (R2014-049). Xunta de Galicia. –Sistema de Información Xeográfica para a xestión e difusión da información meteorolóxica e oceanográfica de Galicia. Subproxecto USC (09MDS034522PR). Xunta de Galicia. July 2018
Contents Abstract 1 Resumen 3 1 Introduction 13 1.1 Background................................... 13 1.2 Problemdescription............................... 17 1.3 Motivation.................................... 18 1.4 Objective and contribution . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 1.5 OutlineoftheThesis .............................. 22 2 Background and related work 25 2.1 Introduction................................... 25 2.2 Data Acquisition Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 2.2.1 CORFUFramework .......................... 26 2.2.2 TOREROProject............................ 26 2.2.3 Chimaris and Papadopoulos, 2007 . . . . . . . . . . . . . . . . . . . 27 2.2.4 Horsburgh et ál., 2011 . . . . . . . . . . . . . . . . . . . . . . . . . 27 2.2.5 GEOSWIFT Infrastructure . . . . . . . . . . . . . . . . . . . . . . . 28 2.2.6 LIFE UNDER YOUR FEET (LUYF) Sensor Network . . . . . . . . 29 2.2.7 SPINEFramework........................... 30 2.3 DataAnalysisSystems ............................. 30 2.3.1 OGCSWEStandards.......................... 33 2.3.2 Observation Data Models . . . . . . . . . . . . . . . . . . . . . . . 34 2.3.3 Geographic Information Systems (GIS) . . . . . . . . . . . . . . . . 35
xviii Contents 2.3.4 Sensor Stream Processing Approaches . . . . . . . . . . . . . . . . . 35 2.3.5 Spatial and Spatio-Temporal DBMSs . . . . . . . . . . . . . . . . . 35 2.3.6 Spatial NoSQL Databases . . . . . . . . . . . . . . . . . . . . . . . 36 2.3.7 Spatial High Performance Data Warehouses Approaches . . . . . . . 36 2.3.8 Array Data Managers . . . . . . . . . . . . . . . . . . . . . . . . . . 36 2.3.9 SciQL.................................. 37 2.3.10 Distributed Processing Frameworks . . . . . . . . . . . . . . . . . . 37 2.3.11 SODA.................................. 38 2.4 Distributed Spatial Data Processing Systems . . . . . . . . . . . . . . . . . 38 2.4.1 HadoopGIS .............................. 39 2.4.2 SpatialHadoop ............................. 42 2.4.3 SpatialSpark .............................. 45 2.4.4 GeoSpark................................ 45 2.4.5 GeoTrellis................................ 48 2.4.6 Magellan ................................ 49 2.4.7 LocationSpark ............................. 49 2.4.8 Simba.................................. 51 2.5 Distributed Spatio-Temporal Data Processing Systems . . . . . . . . . . . . 53 2.5.1 ST-Hadoop ............................... 53 2.5.2 Stark .................................. 54 3 GeoDADIS 59 3.1 Introduction................................... 59 3.2 SystemArchitecture .............................. 60 3.3 MainComponents ............................... 63 3.3.1 DataDissemination .......................... 63 3.3.2 DataAcquisition ............................ 66 3.3.3 ConfigurationManager ........................ 68 3.3.4 DataManager ............................. 68 3.4 Experimental Implementation . . . . . . . . . . . . . . . . . . . . . . . . . 69 4 SODA Design 73 4.1 Introduction................................... 73 4.2 Observation Data Warehouse . . . . . . . . . . . . . . . . . . . . . . . . . . 74
Contents xix 4.2.1 Spatio-temporal Data Model . . . . . . . . . . . . . . . . . . . . . . 74 4.2.2 Observation Data Model . . . . . . . . . . . . . . . . . . . . . . . . 88 4.3 Observation Data Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 96 4.3.1 Mapping Analysis Language (MAPAL) . . . . . . . . . . . . . . . . 97 4.3.2 Analytical Processes . . . . . . . . . . . . . . . . . . . . . . . . . . 108 4.3.3 System Operators . . . . . . . . . . . . . . . . . . . . . . . . . . . 111 4.3.4 Evaluation of MAPAL Expressions . . . . . . . . . . . . . . . . . . 116 5 MAPAL Implementation 119 5.1 Introduction................................... 119 5.2 Data Types Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . 121 5.2.1 Conventional Data Types Implementation . . . . . . . . . . . . . . . 123 5.2.2 Temporal Data Types Implementation . . . . . . . . . . . . . . . . . 123 5.2.3 Point1D Data Type Implementation . . . . . . . . . . . . . . . . . . 125 5.2.4 Point2D Data Type Implementation . . . . . . . . . . . . . . . . . . 125 5.2.5 Geometric Data Type Implementation . . . . . . . . . . . . . . . . . 125 5.3 Data Structures Implementation . . . . . . . . . . . . . . . . . . . . . . . . 127 5.3.1 In-Memory Structures Implementation . . . . . . . . . . . . . . . . 127 5.3.2 Disk Structures Implementation . . . . . . . . . . . . . . . . . . . . 132 5.4 User-Defined Data Types and Functions . . . . . . . . . . . . . . . . . . . . 135 5.4.1 User-Defined Data Types (UDTs) . . . . . . . . . . . . . . . . . . . 135 5.4.2 User-Defined Functions (UDFs) . . . . . . . . . . . . . . . . . . . . 137 5.5 Data Channels Implementation . . . . . . . . . . . . . . . . . . . . . . . . . 138 5.6 Operators Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . 139 5.6.1 Dimension Operators . . . . . . . . . . . . . . . . . . . . . . . . . . 139 5.6.2 Extensional MappingSet Operators . . . . . . . . . . . . . . . . . . 149 5.7 Experimental Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 5.7.1 Clustersetup .............................. 157 5.7.2 Experimentsetup............................ 158 5.7.3 Evaluation Results . . . . . . . . . . . . . . . . . . . . . . . . . . . 161 5.7.4 Scalability ............................... 163 5.8 OptimizationExample ............................. 165 6 Conclusions and Future research 171
xx Contents 6.1 Conclusions................................... 171 6.2 Future lines of research . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 A Primitive mappings 175 B Publications 185 B.1 InternationalJournals.............................. 185 B.2 International Conferences . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185 B.3 NationalConferences.............................. 186 B.4 BookChapters ................................. 186 B.5 OtherPublications ............................... 187 Bibliography 189 List of Figures 203 List of Tables 205
Abstract A myriad of data acquisition devices is observing every day more variables and generating a vast amount of data in almost every application domain. Environmental observation data is an essential portion of such generated data, whose spatio-temporal nature has posed interesting challenges in the area of Environmental Observation Data Management Systems. Two features are common to all these systems: spatio-temporal observations and heterogeneity. In the context of this Thesis, the Observations and Measurements (O&M) conceptual schema was adopted as the theoretical framework for the definition of the concept of observation. Heterogeneity specifically concerns the data acquisition part of the aforementioned systems, which need to access data produced by heterogeneous sensing following different software/hardware specifications that are accessed through several communication protocols. A major challenge is to provide the required flexibility to enable data acquisition from heterogeneous sensing devices and data dissemination through heterogeneous end-user applications. The system must provide simple and straightforward mechanisms for the incorporation of the following components: 1) new in-situ sensing devices, 2) new data dissemination services, and 3) different persistent data storage technologies. Focusing on observation data management, a system must provide the following general functionalities to effectively manage observation data: 1) management of conventional Entity/Relationship data related to non-observed properties of entities, 2) management of sampled data over temporal, spatial (1D and 2D) and spatiotemporal domains, 3) Support for observation data semantics, and 4) efficient implementation for large scale shared-nothing distributed hardware architectures. Moreover, the INSPIRE Directive of the European Union encourages the creation of a Spatial Data Infrastructure (SDI) to ensure the interoperability of spatial information systems in Europe. The application of INSPIRE in the Spanish legislative system forces public administrations to make their geographic data available through SDI services. Therefore, the new
2Abstract enriched geographical knowledge allows for the appearance of many applications in different areas of knowledge that require spatial analysis capabilities. In spite of the above needs, to the best of my knowledge, none of the available technologies and approaches found in data acquisition and data management literature provide support for all the aforementioned functionalities. Therefore, the main objective of this Thesis is the design and implementation of a generic framework for spatio-temporal observation data acquisition and declarative analytical processing. This overall goal can be divided into three independent specific objectives: – Design and implementation of a generic observation data acquisition and dissemination server. – Design of a framework for declarative spatio-temporal analysis in very large spatiotemporal data warehouses. – Efficient implementation of spatio-temporal on-line analytical processing in large scale distributed shared-nothing hardware architectures. The main contributions of this Thesis may be summarized as follows: – Generalization of a data acquisition and dissemination server, with great applicability in many scientific and industrial domains, providing flexibility in the incorporation of different technologies for data acquisition, data persistence and data dissemination. – Definition of a new hybrid logical-functional paradigm to formalize a novel data model for the integrated management of entity and sampled data. – Definition of a novel spatio-temporal declarative data analysis language for the previous data model. – Definition of a data warehouse data model supporting observation data semantics, including application of the above language to the declarative definition of observation processes executed during observation data load. – Column-oriented parallel and distributed implementation of the spatial analysis declarative language. The huge amount of data to be processed forces the exploitation of current multi-core hardware architectures and multi-node cluster infrastructures.
Resumen Una enorme cantidad de dispositivos de adquisición de datos observan cada día más variables y generan ingentes cantidades de datos en la práctica totalidad de dominios de aplicación. Al mismo tiempo, cada día más áreas de investigación centran sus esfuerzos en la adquisición y gestión eficiente de los datos (p. ej., Redes de Sensores, Internet de las Cosas), y en el aprovechamiento inteligente de la información (p. ej., Minería y Análisis de datos). Los datos de observaciones medioambientales constituyen una parte fundamental de dichos datos y, debido a su naturaleza espacio-temporal, presentan algunos desafíos interesantes en el área de la gestión de datos. De hecho, durante las últimas décadas se ha realizado un gran esfuerzo en la investigación de Sistemas de Gestión de Datos de Observaciones Medioambientales. Dichos sistemas presentan dos características comunes: observaciones espaciotemporales y heterogeneidad. Observaciones espacio-temporales La localización de cada observación en un determinado espacio de referencia y el instante temporal en el que el valor observado de una observación se aplica a la propiedad observada son elementos fundamentales de los metadatos, imprescindibles durante la ejecución del análisis. En el contexto de esta Tesis, el esquema conceptual definido por el estándar Observations & Measurements (O&M) del Open Geospatial Consortium (OGC) ha sido adoptado como marco teórico para la definición del concepto de observación y otros conceptos relacionados (p. ej., propiedad observada,valor observado). Una observación contiene un valor observado y los metadatos que proporcionan la semántica de observación necesaria para interpretarlo correctamente. Así, por ejemplo, un valor observado (25) con una unidad de medida específica (ºC) de una propiedad observada (temperatura) está proporcionado por una determinada entidad observada (estación_meteorológica). Una entidad observada puede
10 Resumen va obteniendo. La aproximación cliente/servidor permite a los clientes consultar directamente los datos almacenados en el sistema. La flexibilidad se consigue en GeoDADIS gracias al uso de diferentes patrones de diseño software tanto en la implementación como en el diseño de sus diferentes componentes. El patrón Adapter facilita la incorporación de nuevos servicios de datos, servicios de control remoto y canales de adquisición de datos con cambios mínimos en los componentes internos del sistema. La incorporación de dichos elementos requiere únicamente de actualizaciones de la información de configuración. La flexibilidad, escalabilidad y extensibilidad han sido validadas durante el desarrollo de un prototipo para adquisición y diseminación de datos basado en GeoDADIS que permite la monitorización del estado de salud en entornos educativos. El diseño de SODA se ha dividido en diferentes tareas. En primer lugar, se ha definido un modelo de datos espacio-temporal que incluye nuevos tipos (espaciales y temporales), y estructuras de datos (Dimensiones, Extensional MappingSets, Intensional Mappings) necesarias para la correcta representación de datos de Entidades y datos Raster de forma integrada. Sobre este primer modelo de datos se ha construido un modelo de datos que dota al sistema de la semántica de observación requerida gracias a la definición de un nuevo lenguaje llamado XODDL. A continuación, se ha definido un lenguaje declarativo espacio-temporal para el análisis de datos llamado MAPAL. Además de la especificación de datos y tareas de análisis, este lenguaje permite la definición de procesos analíticos. Estos procesos se ejecutan de forma interna y proporcionan nuevas observaciones a partir de observaciones externas registradas por los diferentes canales de adquisición. Finalmente, se han definido los operadores de sistema que se encargan de ejecutar las tareas definidas por el usuario en MAPAL. Las ventajas principales de SODA se puntualizan a continuación: – Tanto el modelo de datos de observaciones espaciales como la definición declarativa de los procesos analíticos incorporan y dan soporte a la semántica de datos de observaciones. – Se proporciona soporte directo a la representación y análisis integrado tanto de datos E/R convencionales como datos espaciales, temporales y espacio-temporales muestreados. – Los nuevos tipos de datos espaciales y temporales permiten la representación y transformación entre diferentes resoluciones tanto en el dominio espacial como temporal.
Resumen 11 – El concepto matemático de función se utiliza para representar tanto datos (mediante la nueva estructura de datos Extensional MappingSet) como comportamiento (mediante la nueva estructura de datos Intensional Mapping). Así pues, esta solución debería ser de fácil uso para usuarios del ámbito científico. Adicionalmente, esta aproximación funcional facilita la definición y reutilización de resultados intermedios. – La incorporación de los nuevos lenguajes propuestos, MAPAL y XODDL, en servicios web es muy sencilla debido a que están basados en el lenguaje declarativo XML. – La implementación eficiente de SODA se ha visto beneficiada por las estructuras de datos no anidadas que se han definido en el modelo de datos. Se ha propuesto también la implementación de un prototipo de forma que se pueda comparar a SODA con las soluciones existentes en el estado del arte para el análisis de datos espaciales y espacio-temporales. Los beneficios de los modelos de datos y operadores definidos se han demostrado en los resultados de rendimiento obtenidos por el prototipo implementado. Los tiempos de ejecución obtenidos para la operación de join espacial, implementada para esta comparativa, mejoran los obtenidos por las soluciones existentes actualmente (p. ej., GeoSpark, LocationSpark, Stark). Dichos tiempos de ejecución son órdenes de magnitud inferiores a los obtenidos por los competidores. Además, los test de escalabilidad demuestran un rendimiento similar al de las mejores soluciones actuales. El principal inconveniente de SODA viene dado por la adopción de un nuevo paradigma funcional de gestión de datos por parte de los usuarios de bases de datos tradicionales. Sin embargo, el formalismo funcional se ha combinado con el bien conocido formalismo lógico a la hora de definir MAPAL. Por tanto, los constructores de MAPAL son muy similares a los constructores que están presentes en otros lenguajes disponibles actualmente como XQuery. Para finalizar este resumen, se van a comentar las líneas de trabajo futuro que se pueden derivar de esta Tesis. Respecto a GeoDADIS, las futuras líneas de investigación deberían estar relacionadas con la ampliación del sistema de forma que se pueda dar soporte para la adquisición y diseminación de las observaciones complejas que producen los sensores remotos (p. ej., lidar, radar). En cuanto a SODA, se pueden identificar diferentes vías de trabajo futuro. A continuación, se detallan las más relevantes:
12 Resumen – Incorporación de nuevas técnicas de optimización de consultas. – Definición de nuevas estructuras de indexación. – Diseño e implementación de nuevas estrategias de particionamiento para Dimensiones, Extensional MappingSets y datos espaciales. – Incorporación de técnicas de procesamiento aproximado de consultas sobre Extensional MappingSets almacenados.
CHAPTER 1 INTRODUCTION 1.1 Background A myriad of data acquisition devices is observing every day more variables and generating a vast amount of data in almost every application domain, e.g., health care, home automation, clean energy production, weather forecast, natural disaster prediction. Furthermore, an increasing number of research areas are involved in the efficient acquisition and management of data, e.g., Sensor Networks, Data Logging, Internet of Things (IoT), large scale data management, and in the intelligent exploitation of information, e.g., Data Analytics and Mining. Environmental observation data is an essential portion of such generated data, whose spatio-temporal nature has posed interesting data management challenges. More specifically, important research efforts have been devoted to Environmental Observation Data Management Systems for decades. Two features are common to all these systems: spatio-temporal observations and heterogeneity. Spatio-temporal observations The location of each observation in some reference space and the time when the observed value of an observation applies to the observed property are important pieces of metadata that must be used during the analysis. In the context of this Thesis, the Observations and Measurements (O&M) conceptual schema [30] of the Open Geospatial Consortium (OGC) was adopted as the theoretical framework for the definition of the concept of observation and other related concepts (e.g., observed property, observed value). An observation en-
14 Chapter 1. Introduction : OM_Observation +phenomenonTime: TM_Object = 07/02/2018 11:49 : Measure +value = -15 +uom = ºC +observedValue temperature_sensor: OM_Process +serial_number = LXA5506000EM00 +model = TM09-BELL +trigger_type = time-triggered +time_frequency = 10 min +resolution = 0.5 air_temperature: GFI_PropertyType +observedProperty +observationProcess EOAS_weather_station: Weather_Station +name: String = EOAS-Santiago +owner: String = Xunta de Galicia +geometry: GM_Object = Point(536101 , 4747354, 23029) +observedEntity Weather_Station +name: String +owner: String +geometry: GM_Object +air_temperature: Measure +air_humidity: Measure +soil_temperature: Measure +soil_humidity: Measure +wind_speed: Measure +solar_radiation: Measure «instanceOf» Figure 1.1: OGC Observation example. closes both an observed value and the relevant observation metadata that provides observation semantics required to adequately interpret it. An example of an observation is shown in Fig. 1.1. An observed value (-15) with a specific unit of measure (ºC) of an observed property (air_temperature) is provided by an observed entity (EOAS_weather_station). An observed entity may have both conventional properties (name,owner,geometry) and observed properties (air_temperature,air_humidity,soil_temperature,soil_humidity, wind_speed,solar_radiation). Values of conventional properties are usually assigned by some authority whereas values of observed properties are estimated by some observation process (temperature_sensor). It is mandatory to register properties (serial_number, model,trigger_type,time_frequency,resolution) of the specific observation process used to generate the observed value.Observation processes may be of very different nature, including physical devices (e.g., temperature sensors), tasks performed by people (e.g., data registered by an operator) and data processing algorithms (e.g., weather forecast). It is also mandatory to register the phenomenonTime (07/02/2018 11:49), i.e., the time instant when the observed value applies to the observed property. Notice for example that the weather forecast (observation process) might take into account historic data values obtained some time ago. The type of observation data produced by an observation process is determined by two characteristics: i) whether the process is executed periodically (time_triggered) or triggered by specific events (event_triggered); ii) the relative position of the process with
1.1. Background 15 Figure 1.2: Observation data types. respect to the observed entity (in-situ or remote). Different available data types obtained by combining such characteristics are shown in Fig. 1.2. Trigger type Event-triggered processes start at some time instant determined by a specific event. For instance, a Light Detection and Ranging (LIDAR) image taken at some specific time instant. Time-triggered processes are executed at some predefined time frequency producing regular samplings in the temporal domain. As an example, we might register air temperature values obtained by the temperature sensor of a weather station every ten minutes. Sensor location Focusing on sensors (one of the aforementioned observation process types), in-situ sensors are located at the spatial position of the observed entity. They produce a single observa-
16 Chapter 1. Introduction Sodar Unit Acoustic Pulse Echo Spatial resolution Scattering volume (a) SODAR (static platform) VIIRS Suomi NPP Spatial resolution (b) VIIRS (mobile platform) Figure 1.3: Illustration of 1D and 2D spatial samplings. tion value at each time instant. Examples of such sensors are a temperature sensor installed in a meteorological station (static platform) and a GPS device installed in a car (mobile platform). Unlike in-situ sensors, remote sensors are located far away from the observed entity. They provide several observed values (one for each observed entity) at each time instant. An example of static remote sensor is the Sonic Detection And Ranging (SODAR), Fig. 1.3(a), used to register wind speed at different heights above the ground by measuring the scattering of sound waves produced by atmospheric turbulence. SODAR generates a 1D sampling of wind speed along consecutive discrete locations of a vertical line profile. An example of a remote sensor installed in a mobile platform is the Visible Infrared Imaging Radiometer Suite (VIIRS) installed in the Suomi NPP satellite (Fig. 1.3(b)). VIIRS allows high resolution images to be acquired both in visible and infrared spectrum, providing a whole view of the Earth every two days with a spatial resolution of 750 meters. Generated data include 2D
1.2. Problem description 17 regular samplings (called Rasters in the area of Geographic Data Management) of color and temperature of the sea surface. Heterogeneity Heterogeneity specifically concerns the data acquisition part of the aforementioned systems, which need to access data produced by heterogeneous sensing devices (e.g., humidity sensors, GPS devices, radar, lidar, multispectral scanner sensors) following different software/hardware specifications that are accessed through several communication protocols (e.g., WiFi, RS-485, Ethernet). General system architectures of relevant Data Acquisition and Monitoring Applications [62, 102, 76, 78, 63], and Non Real-Time Supervisory Control Systems [26, 32] are usually composed of three main components, namely, End-User Applications,Data Servers and Sensing Devices.End-User applications are in charge of data analysis and visualization. Sensing devices run the observation process and are highly heterogeneous both in functionality and communication capabilities, as already stated. Acting as a gateway between the heterogeneous specific domains of End-User Applications and the heterogeneous collection of Sensing Devices, one or more Data Servers are added to the system in order to provide homogeneous data access. 1.2 Problem description Based on the above, several challenging problems arise during the design and implementation of Environmental Observation Data Acquisition and Management Systems. Specific issues related to both observation data acquisition and observation data management are detailed below. Regarding observation data acquisition, generalization efforts in sensing devices programming and end-user application development tend to be worthless because of the strong conditioning on vendor specifications and the high dependency on specific domain and user preferences, respectively. On the other hand, the functionality and architecture of data servers tend to be very similar in the broad majority of applications. A major challenge however is to provide the required flexibility to enable data acquisition from heterogeneous sensing devices and data dissemination through heterogeneous end-user applications. The system must provide simple and straightforward mechanisms for the incorporation of the following components:
18 Chapter 1. Introduction – New in-situ sensing devices. – New data dissemination services. – Different persistent data storage technologies for different observed properties. Focusing on observation data management, a system must provide the following general functionalities to effectively manage observation data: – Management of conventional Entity/Relationship (ER) data related to non-observed properties of entities. – Management of sampled data over temporal, spatial (1D and 2D) and spatio-temporal domains. – Support for observation data semantics. Relevant observation metadata of observed properties of entities must be provided. – Efficient implementation for large scale shared-nothing distributed hardware architectures. 1.3 Motivation In terms of spatial data software architectures for GIS, recent developments and trends propose the decomposition of systems into simple and well-defined services, which are often web-based and whose interfaces follow international interoperability standards of the OGC and the International Organization for Standardization (ISO). Thus, Spatial Data Infrastructures (SDI) integrated by distributed services through the Internet can be made available to GIS developers. Beyond the previous technological consideration, relevant policies are being adopted to improve the availability of spatial data sets generated by different public administrations. In particular, the INSPIRE Directive of the European Union (2007/2/CE, March 14th 2007) encourages the creation of a SDI to ensure the interoperability of spatial information systems in Europe. The application of INSPIRE in the Spanish legislative system forces public administrations to make their geographic data available through SDI services. Therefore, the new enriched geographical knowledge allows for the appearance of many applications in different areas of knowledge that require spatial analysis capabilities.
1.3. Motivation 19 In spite of the above needs, to the best of my knowledge, none of the available technologies and approaches found in data acquisition and data management literature provide support for all the functionalities required in Section 1.2. More details related to the this assertion are given in the following paragraphs. Most of data acquisition systems provide high flexibility to obtain observation data from heterogeneous sensing devices but lack the required flexibility in data storage and dissemination capabilities. As opposed to [79], in [17, 26, 53, 67, 89, 96] flexible mechanisms to attach new sensing devices are provided. However, [96] lacks flexibility to extend implemented data storage technologies whereas [17, 26, 53, 67, 89] provide limited capabilities. Flexible ways to attach new dissemination services are provided in [79] as stored procedures and userdefined functions accessible through web-form interfaces. Such flexibility is not available in [89, 96] and is very limited in [17, 26, 53, 67]. A huge amount of research effort devoted to observation data management may be found in data management literature. The area of spatial databases [46, 68] is one of the most active research areas providing many research approaches. Even the ISO SQL standard [58], implemented by well known DBMSs [83] has been extended with relevant spatial functionality. These tools currently enable declarative querying over spatial data, including support for 2D rasters. High performance Data Warehouse [54] and NoSQL [77] tools implement spatial extensions although raster data are not supported. Recording and processing of conventional and spatial data, including rasters, are currently supported by available GIS tools [81]. Declarative data analysis, not provided by such GIS solutions, is supported by array data managers [15, 22] for very large collections of array raster data. Even though declarative analysis of relational data through array data structures is not very user friendly, an attempt of integrated management of relational and array data was tried in [111] but the user has to deal with both relational and array semantics. Systems providing declarative analysis of data streams of sensor data have been developed in [40, 71]. However, raster data are not supported. Finally, observation data semantics are only supported by standards of OGC, Sensor Web Enablement (SWE) initiative [30, 84, 23] and specific observation data models and ontologies [20, 29, 72], although declarative analysis of observation data is not supported.
26 Chapter 2. Background and related work 2.2.1 CORFU Framework A Common Object-oriented Real-time Framework for the Unified (CORFU) development of distributed IPMCS (Industrial Process Measurement and Control Systems) applications is defined in [96]. The CORFU framework adopts the function block concept defined by IEC standards [55, 56] and proposes a new network topology for fieldbus interconnection. The core element in the proposed network topology, called interworking unit, is composed of the following building blocks. –Virtual Field Bus (VFB): the main component of the interworking unit abstracts any commercial fieldbus to the IEC 61499 [55] level. This abstraction allows for interoperability in fieldbus level. –Fieldbus Wrapper: allows for wrapping different fieldbus specifications to the VFB. –Industrial Process-Control Protocol (IPCP): each interworking unit implements the IPCP on top of TCP/UDP layers. The IPCP has been defined for the development, distribution, and operation of function block based industrial process measurement and control applications. In this solution the interworking units are located between each fielbus and a backbone network that provides real time interconnection of fieldbus segments. The architecture of the interworking unit adopts the Adapter pattern [41] to ease the incorporation of new wrappers that enable the interconnection of heterogeneous fieldbuses to the selected backbone. Similarly, GeoDADIS implements the Adapter pattern to access data acquisition channels, data services and control clients. 2.2.2 TORERO Project The research project TORERO (total life cycle web-integrated control) specifies a new DCS environment. The main element of the TORERO DCS [89] is a mechatronic component, called torero device, providing intelligent control. In a TORERO environment the required control functionality is realized by all torero devices working in collaboration. The architecture of a torero device is divided into three layers: –Physical layer: sensor/actuator elements.
2.2. Data Acquisition Systems 27 –Hardware layer: processor, storage, RAM, Ethernet interface, connectors to the sensor/actuator elements, etc. –Software layer: Operating System, Java Virtual Machine, FTP, HTTP, the control application, etc. The control application can only access hardware components via so called device functions. A device function is an abstraction of the underlying hardware, i.e., a wrapper that allows heterogeneous hardware to be controlled by the same control application software. As stated in Subsection 2.2.1, GeoDADIS also implements the Adapter pattern. 2.2.3 Chimaris and Papadopoulos, 2007 A generic component-based framework that can be used to build telecontrol applications was defined and implemented in [26]. The main components of this framework, called remote units, are small intelligent subsystems. Each remote unit must perform the following tasks. – Handle every connected device (alarms, lights, heating, etc.). – Monitor the connected devices and transmit data changes and message alerts to the control center. – Change its behavior based on received update and control messages. – Support secure and consistent communication. – Ensure the availability and efficiency of the communication channel. Based on the above, a remote unit may include control functionality that goes beyond the control capabilities of GeoDADIS. However, the flexibility requirements imposed during the design of GeoDADIS for data dissemination, data storage and remote control are not present in [26]. 2.2.4 Horsburgh et ál., 2011 An environmental observatory information system that supports collection, organization, storage, analysis and publication of hydrologic observations is described in [53]. The architectural and procedural components are described as follows.
28 Chapter 2. Background and related work –Data Observation and Communication: sensors and telemetry systems used to collect observations. –Data Storage and Metadata: data models, database systems and software required to create a persistent data repository. –Quality Assurance, Quality Control and Provenance: software and procedures for transforming raw data into publishable data products. –Data Publication and Interoperability: software, protocols, formats and vocabularies used for publishing data in interoperable formats. –Discovery and Presentation: tools provided to data consumers for visualization and analysis purposes. Related to GeoDADIS are the Data Storage and Metadata and Data Publication and Interoperability components. The former enables persistent storage of both sensor data and metadata. An important added-value step in this component involves the mediation across the variety of software supporting sensor and communication systems. Such a mediation is achieved in GeoDADIS by the implementation of wrappers for different data acquisition channels, as already mentioned. The latter provides data dissemination functionality, achieved in GeoDADIS by the implementation of data services. 2.2.5 GEOSWIFT Infrastructure GeoSWIFT is a distributed geospatial infrastructure for the Sensor Web1proposed in [67]. The core component of GeoSWIFT is the open geospatial sensing service, which serve as a single queryable global sensor for Sensor Web users. Each sensing service role and its behavior are explained below. –Sensing Server: provides a web-enabled interface for sensor systems and their geospatial information. The standard for sensor data access exposed by GeoSWIFT is based on the specifications provided by the Sensor Web Enablement (SWE) initiative [19] of the OGC. 1“A Sensor Web is a system of intra-communicating spatially distributed sensor pods that can be deployed to monitor and explore new environments” [61].
2.2. Data Acquisition Systems 29 –Sensing Registry: plays a central role in publishing, finding, and binding to networkaccessible services by providing a common mechanism to classify, register, describe, search, maintain, and access information about Sensor Webs and other Web Services. –Viewer: GeoSWIFT Viewer is based on GeoServNet Viewer2. As in the case of GeoDADIS, the Sensing Server of GeoSWIFT acts as a gateway that hides the different communication protocols, data formats and standards of sensor systems and provides a standard interface for clients to collect and access sensor observations. Despite of the similarities between GeoSWIFT and GeoDADIS, GeoSWIFT does not achieve the flexibility requirements imposed during the design of GeoDADIS. 2.2.6 LIFE UNDER YOUR FEET (LUYF) Sensor Network A data access gateway is implemented in [79] to gather data from a Wireless Sensor Network (WSN) for soil monitoring. Core components of LUYF are detailed below. –Data Collection Subsystem: composed of motes and a base station. Each mote is connected to a data acquisition board providing ambient light, temperature and soil moisture sensors. Motes sample data at some predefined temporal resolution, typically every minute, and store them in local memory. The base station requests stored data from motes once every two weeks and stores the retrieved measurements in the database. –Database: raw measurements arrive from the base station as ASCII files. First, received data are loaded into a temporary table. Next, duplicates are removed and data are stored as raw data. A multi-step pipeline is required for converting raw data to scientifically meaningful values. Such process is automatically performed by a stored procedure for all sensors within the database. Stored procedures and user defined functions, accessible through web-form interfaces, provide access to aggregated data. The major drawback of LUYF, when compared to GeoDADIS, is the lack of flexibility that allows users to add new data acquisition wrappers. 2GeoServNet Viewer is a 2D/3D Web GIService viewer designed for streaming large amount of spatial data via Internet.
30 Chapter 2. Background and related work 2.2.7 SPINE Framework The general architecture of the SPINE framework [17] is composed of a collection of sensor nodes connected to the coordinator node that manages the network, collects and analyzes the retrieved data, and acts as a gateway to connect sensors and wide area networks. The sensor node manages and abstracts sensors (providing a standard interface to diverse sensor drivers), and is responsible for sampling and storing sensor data in properly defined buffers. Two major differences may be found between SPINE and GeoDADIS. First, SPINE enables the incorporation of signal processing functionality that is out of the scope of GeoDADIS. Second, the flexibility in the incorporation of new data dissemination and remote control services of GeoDADIS is not present in SPINE. 2.3 Data Analysis Systems In this section a comparison between different solutions in the area of data analysis is provided. Based on generic functionalities required for all observation management systems, the comparison criteria are specified below. 1. Direct support for observation semantics3: the representation of terms related to an observation is required. Observed entities allow for the representation of entities with conventional and observation properties. Effective analysis and correct interpretation of observed values of some observed property require relevant metadata to be recorded. Specifically, important metadata to be recorded are the observation process and the phenomenon time.Observation process and observation entity instances have to be classified into process types and entity types respectively. Moreover, the recording of observation process properties should be also supported. 2. Support for the management of sampled data: it is not only for classical E/R data that an observation data management system must support efficient processing. Data structures and operations have to be provided to enable effective processing of sampled data. As aforementioned in Section 1.1, time-triggered processes generate temporal sampling data and remote sensors usually produce spatial sampling data (raster data). As we will see below, either highly inefficient approaches or complex nested models arise when applying the classical relational-based models to sampled data. 3We refer here to observation semantics provided by [30] and detailed in Section 1.1
2.3. Data Analysis Systems 31 3. Support for multi-resolution temporal and spatial data:observation data is currently generated with different temporal and spatial resolutions by a huge amount of available sensors. Because of that, evaluation of operations often implies transformations between diverse temporal and spatial resolutions. An appropriate data type system should be provided by observation data management systems in order to simplify these transformations. 4. Simple data modeling approach: for evaluation purposes in the context of this Thesis, we are considering as non-simple data models those that fulfill one or two of the following features: – more than one non-nested data structure. – nested data structures including records and collections. It is clear that simple data models have some advantages over non-simple ones, e.g., an efficient implementation of a simple model is far more straightforward than a nested data model implementation, and implementation of different semantics in diverse data structures often results in not user friendly interfaces. 5. Model based on a well known paradigm: a rapid progression up the learning curve is enabled when data models are defined based on well known paradigms on account of many years of user experience. 6. Stream processing approach: stream processing approaches are required when real time prerequisites are present and there is not a large amount of data to be recorded. These systems implement small stored data structures and efficiently process input data streams in order to produce output data streams. Stream processing approaches are commonly known as Complex Event Processing4(CEP) and they rely on the evaluation of Continuous Query Language (CQL) expressions [13, 59]. 7. On Line Transaction Processing (OLTP) approach: OLTP approaches are required when real time prerequisites are present with simple temporal patterns and there is a large amount of data to be recorded. This approach is traditionally supported by conventional DBMSs for reasonably large data collections and provided by both NoSQL [77, 8] and NewSQL [105] solutions in the new era of Big Data Management. 4Also known as Information Flow Processing Systems [31]
32 Chapter 2. Background and related work 8. On Line Analytical Processing (OLAP) approach: OLAP approaches are required when real time prerequisites are not present and there is a large amount of data to be recorded. We usually identify these systems in Data Warehouse solutions implemented by Bussines Intelligence (BI) applications. Examples of high performance implementations are Hewlett-Packard Vertica [99], which is an evolution of C-Store [93], and the open source MonetDB database [54]. Recent research solutions on column-oriented technologies serve as a basis for efficient implementations of those Big Data solutions. Indeed, the column-based storage of relational data, instead of the traditional row-based storage, is the major contribution of these approaches. Main features are: – efficient compression techniques. – processing over compressed data. – columns not involved in computations are not retrieved from storage. As major drawback we can mention the inefficient performance of insertions, updates and deletions of data. This makes them suitable for data warehouses. 9. Support for declarative processing: taking into account that procedural solutions are dominant in application domains such as environmental applications handling sampled observation data, and that procedural approaches have well known disadvantages compared to declarative languages, it is clear the motivation for applying declarative data management technologies to these environmental applications. 10. Support for aggregated queries: effective observation data analysis in OLAP systems cannot be accomplished without statistical methods providing aggregation functionality. 11. Support for iterative processing: recursive queries are required in only few data management applications. This is the reason why such functionalities were out of the scope of first SQL implementations. Current ISO SQL standard and DBMSs vendors support a kind of limited recursion. Regarding the analysis of observation data in environmental applications, such functionalities are commonly required to perform many simulations. Examples of these are forest fire propagation, oil spills, and flooding. Thus, although it is not a key functionality, support for iterative processing is a desirable feature.
2.3. Data Analysis Systems 33 12. Data processing based on a well known language: as stated for data models, a clear advantage for data management systems is that the definition of query languages is based on well known paradigms. 13. Distributed processing of spatial, temporal and spatio-temporal data: traditional technologies are no longer suitable for processing the large amount of spatial, temporal and spatio-temporal data currently generated. In fact there is a growing demand for solutions that support high performance queries on such data. This makes distributed and parallel processing of spatial, temporal and spatio-temporal data no longer desirable but required. Hence, data management systems currently created have as essential requirement such cluster-based processing. 14. Availability of efficient implementation: a data management approach is really useful if it can be efficiently implemented. A prototype implementation demonstrates the approach feasibility and its use in real application domains shows its maturity. The degree of compliance with the previous evaluation criteria is now detailed for several related research solutions and available technologies, including also the SODA framework. Table 2.1 provides an overview of such evaluation. For each approach, Prepresents that the relevant criterion is partially supported andYrepresents that is completely supported. A more detailed discussion is given below. 2.3.1 OGC SWE Standards The Sensor Web Enablement (SWE) of the OGC provides standards for interfaces of web services that are related to the management of environmental observation data. In particular, the Observations and Measurements (O&M) [30] and Sensor Model Language (SensorML) [84] were already mentioned in Chapter 1. The Sensor Observation Service (SOS) [23] defines a web service interface to query observation data collections, either stored or directly obtained from devices. Data is transferred between client and server in standard XML encodings of O&M and SensorML models. Query capabilities of SOS are limited to just filtering. Regarding data processing, OGC defines the Web Processing Service (WPS)[88] interface that enables the invocation of data processing algorithms through the web. Various implementations of the above standards already exist in the market, both with commercial and open source licenses. In general it is obvious that O&M provides appropriate support for the modeling of
34 Chapter 2. Background and related work Obs. Semantics Sampled Data Multiresolution Simple Model Well Known Model Stream Proc. OLTP OLAP Declarative Proc. Aggregation Iterative Proc. Well Known Lang. Impl. Available OGC SWE Stds. Y Y Y Y Obs. Data Models Y Y Y GIS Y Y Y Y Sensor Stream Y Y Y P P P Y Spat. and ST DBMSs Y Y Y Y P P P P Y Spatial NoSQL Y Y Y Spatial HP DW Y Y Y P P P Y Array Data Managers Y Y Y Y Y Y SciQL Y P Y Y Y P Y Dist. Proc. Systems Y Y Y Y Y Y SODA Y Y Y Y Y Y Y Table 2.1: Comparison of related technologies. observation semantics and sampled data. Different spatial and temporal resolutions are supported but transformations are a user matter. The underlying object oriented data modeling approach with XML encodings is well known. However, nested structures are required to support sampled data. Declarative data processing is not supported at all as WPS just provides means for remote procedure calls. 2.3.2 Observation Data Models Beyond the above O&M OGC standards, several data models and ontologies have been proposed to support observation data semantics [20, 29, 72]. They are based on well known paradigms and provide observation data semantics with simple data modeling approaches.
2.3. Data Analysis Systems 35 However, sampled data and multi-resolution is out of the scope of these models as well as any kind of data processing. 2.3.3 Geographic Information Systems (GIS) Currently, a wide variety of GIS tools, both with commercial and open source licenses, are available. A representative example of them is GRASS [81], which supports the management of any kind of geographic data, including rasters, recorded in many different well known models and formats. Raster data management is usually formalized with relevant raster algebras [24]. Observation semantics are not considered in GIS and although the managed data may have many different spatial resolutions, transformations between them have to be explicitly done by the user to perform operations. Spatial data processing is a strength of tools like GRASS. However, it is performed by the execution of a very large amount of different commands. Therefore, a declarative language is missing. Notice that the user must know which is the functionality of each command and how to combine them, thus only expert users may take real advantage of spatial data analysis with GIS tools. 2.3.4 Sensor Stream Processing Approaches Various Stream Processing approaches have been explicitly proposed for the management of data generated by sensor networks [40, 71]. Although they were defined for sensor data management, observation data semantics are not explicitly incorporated and are delegated to user interpretation. Any kind of spatial data management is out of the scope of these approaches. They support declarative real-time processing of streams with aggregation functionality based on SQL like languages. Real-time requirements of these approaches are clearly in conflict with the support of iterative processing. 2.3.5 Spatial and Spatio-Temporal DBMSs Many temporal extensions have been proposed for the classical relational model [33, 91]. Recently, some characteristics have been incorporated into ISO SQL standard [64]. Various spatial [46, 68, 98] and spatio-temporal [47, 104] extensions to classical models have been proposed in the literature. Spatial functionality has already been added to ISO SQL standard [58], which is currently implemented by most of the available DBMSs (see [83] for an example). Direct support of observation semantics is out of the scope of spatial DBMSs. They
42 Chapter 2. Background and related work Figure 2.2: Architecture of SpatialHadoop. source: Eldawy and Mokbel, in ICDE, 2015 [37]. 2.4.2 SpatialHadoop SpatialHadoop [37] is a MapReduce framework with native support for spatial data. Unlike previous approaches (e.g., Parallel-Secondo [69], M D-HBase [82], Hadoop GIS), Spatial- Hadoop do not rely on Hadoop as a black box. Such a different approach prevents Spatial- Hadoop from suffering the limitations and performance bottlenecks of Hadoop. The main attributes that allow SpatialHadoop to overcome the limitations of previous approaches are detailed below. – Provision of built-in code. SpatialHadoop code is built inside the Hadoop base code to extend Hadoop core with spatial data functionality. This feature allows SpatialHadoop to be more powerful and efficient than previous solutions. – Support for skewed spatial data distributions by implementing a set of spatial index structures. – Users are enabled to develop a huge amount of spatial functions, e.g, spatial join,range queries. As shown in Fig. 2.2, the SpatialHadoop architecture is divided into four main layers. A detailed description of each layer is provided below.
2.4. Distributed Spatial Data Processing Systems 43 Language Layer. A novel high level SQL-based language, called Pigeon [36], is implemented in this layer. Several languages have been recently defined to reduce coding effort when working with MapReduce-based paradigms, e.g., HiveQL [52], Pig Latin [85], SCOPE [112], and YSmart [65]. Pigeon is an extension to Pig Latin providing OGC-compliant spatial data types, functions and operations. Standard spatial data types (e.g., Point,LineString, and Polygon) are supported. User-defined functions (UDFs) are harnessed to define spatial aggregations (e.g., Union), spatial predicates (e.g., Overlaps), and other spatial functions (e.g., Buffer). A new kNN (knearest neighbors) statement has been added to support kNN- queries. In addition, two Pig Latin statements have been overridden. SpatialHadoop overrides the Filter statement to support range queries, and the Join statement to support spatial joins. Storage Layer. As pointed out in [37], several challenges arise when applying traditional spatial indexes (e.g., Grid file, R-tree [48]) in Hadoop. To overcome these limitations, SpatialHadoop follows a two-layer indexing approach. A global index, stored in the master node, allows SpatialHadoop to split data across a set of partitions stored in slave nodes. A local index, stored in each partition, enables local data to be arranged. Regardless of the underlying spatial index structure, SpatialHadoop defines an index building algorithm composed of three main phases. 1. Partitioning. Spatial partitioning of the input file into npartitions performed through the following steps. a) Compute the number of partitions, n. b) Define partition boundaries, i.e., the spatial area covered by each single partition. This process strongly depends on the underlying index being constructed. c) Perform the physical partition of the input file, given the above partition boundaries, through a MapReduce job. 2. Local Indexing. A reduce function is used to build a local index on the stored data of each physical partition. To make this happen, the reduce function stores the records of each partition in a spatial index, written in a local index file. 3. Global Indexing. Once local indexing is performed, the master node builds a global index that indexes all partitions. First, concatenates all local index files into one final
44 Chapter 2. Background and related work Figure 2.3: Map phase in Hadoop and SpatialHadoop. source: Eldawy and Mokbel, in ICDE, 2015 [37]. indexed file. Second, indexes all file blocks using their rectangular boundaries as the index key to build the in-memory global index. MapReduce Layer. This layer is responsible for running the MapReduce jobs that process the required queries. Fig. 2.3 shows the Map phase of the MapReduce plan in both Hadoop and SpatialHadoop, highlighting the differences between them. In Hadoop, the File- Splitter takes the input file and divides the data into nsplits, where nis determined based on the number of available slave nodes. Then, the RecordReader extracts records as key-value pairs and passes them to the Map function. SpatialHadoop enriches traditional Hadoop systems modifying the FileSplitter and RecordReader components. The new SpatialFileSplitter early prunes file blocks not contributing to the answer and generates data splits by exploiting the global spatial index stored on input files. And the new SpatialRecordReader efficiently process the previous splits using local indexes. Operations Layer. The language layer is provided with a myriad of spatial operations (e.g., range queries,kNN-queries,spatial joins). The operations layer is responsible for the efficient implementation of all these spatial operations.
2.4. Distributed Spatial Data Processing Systems 45 2.4.3 SpatialSpark SpatialSpark [107] is a prototype system to process large-scale spatial join queries over Spark, and supports indexed spatial joins based on point-in-polygon test and point-to-polyline distance computation. The following main goals have been defined for SpatialSpark. – Identify limitations and advantages of Spark for spatial data processing in cluster environments from an architectural point of view. – Determine the potential performance of modern hardware for large-scale spatial join query processing. Different indexing techniques for spatial filtering have been implemented in SpatialSpark. For spatial refinement, SpatialSpark relies on the well known Java Topology Suite (JTS) package [60]. To make SpatialSpark compatible with Hadoop-based systems, strings are used to represent geometries. Although higher efficiency could be possible by representing geometries as binary, avoiding string pairing overheads and allowing flexible disk accesses, this option is left for future work in SpatialSpark. As all intermediate data are memory resident in Spark, higher performance is achieved in SpatialSpark by minimizing expensive disk I/Os, and utilizing finer grained data parallelism. 2.4.4 GeoSpark GeoSpark [108] is an in-memory cluster computing framework for processing large-scale spatial data, providing support for spatial data types, indexes, and operations by extending the core of Spark. Specifically, GeoSpark enhances the resilient distributed datasets (RDDs) to support spatial data (SRDDs). The key features of GeoSpark are the following. – Support for loading, processing, and analyzing large-scale spatial data over Spark. – Support for geometrical and distance operations is given by the definition of a set of SRDD types, e.g., PointRDD and PolygonRDD. Moreover, Spark programmers may easily develop spatial analysis applications by using the Application Programming Interface (API) provided by SRDDs. – Support for different global spatial data indexing techniques. Input SRDDs are partitioned using a grid structure. Then, these resulting grids are assigned to computing machines for parallel execution.
46 Chapter 2. Background and related work Figure 2.4: Architecture of GeoSpark. source: Yu et ál., in Proc. SIGSPATIAL, 2015 [108]. The architecture of GeoSpark is composed of three layers, as depicted in Fig. 2.4. Apache Spark Layer serves as the basis where GeoSpark is built on, Spatial RDD Layer provides support for geometrical and spatial objects and operations, and Spatial Query Processing Layer executes efficient spatial query processing algorithms. Apache Spark Layer. Comprises the basic functions natively provided by Spark such as loading/storing data from/to persistent storage and regular RDD operations. Spatial RDD Layer. Efficient partition of spatial data elements across cluster nodes is enabled by the definition of the extended spatial version of the Spark RDD. To write spatial data analytics applications, novel parallelized spatial transformations and actions in SRDDs provide users with an intuitive interface. Main features of this layer are pointed out below. –Spatial Objects. Three new SRDDs (PointRDD,RectangleRDD, and PolygonRDD), implemented in this layer, allow the storage of different spatial objects. Furthermore, GeoSpark provides a Geometrical Operations Library which natively supports geometrical operations such as Overlap(),MinimumBoundingRectangle()and Union(). –SRDD Partitioning. GeoSpark automatically partitions every SRDD using a global grid file. The algorithm for partitioning the SRDDs is as follows. First, a global grid file
2.4. Distributed Spatial Data Processing Systems 47 is created by splitting the spatial space into a number of equal geographical size grid cells. Then, each element in the SRDD is assigned to every overlapping grid cell, i.e., if an element intersects with two or more grid cells, it is duplicated and different grid IDs are assigned to its copies. –SRDD Indexing.Spatial IndexRDDs which inherit from SRDDs are implemented in GeoSpark to provide spatial indexes such as Quad-Tree [39] and R-Tree [48]. Furthermore, a local spatial index may be adaptively created on a SRDD partition to find an optimal trade-off between the run time performance and the memory/cpu usage in the cluster. Spatial Query Processing Layer. Once geometrical objects are pre-processed and stored in the Spatial RDD Layer, users may invoke spatial queries (e.g., Range Query,Join Query) supported by this layer over large-scale spatial datasets. Query execution is parallelized in GeoSpark by using features such as partitioned SRDDs, spatial indexing, and fast in-memory computation. GeoSpark’s algorithms for spatial range, spatial join, and kNN queries are described as follows. –Spatial Range Query. The range query algorithm is executed by GeoSpark following the steps below. 1: Load target dataset 2: Partition data 3: Create a spatial index on each SRDD partition (optional) 4: Broadcast the query window to each SRDD partition 5: Check the spatial predicate in each partition 6: Remove duplicate spatial objects generated in data partitioning phase –Spatial Join Query. The algorithm for processing spatial join queries in GeoSpark is given below.
48 Chapter 2. Background and related work 1: Load two input SRDDs 2: Partition data 3: Create a spatial index on each SRDD partition (optional) 4: Join the two SRDDs by their keys (grid IDs) 5: Calculate spatial relations of spatial objects that have the same grid ID 6: Keep in the final results only the elements satisfying the spatial relation 7: Group results by grid ID 8: Remove duplicate results –Spatial kNN Query. GeoSpark implements the following heap based top-k algorithm [87] to process spatial kNN queries. 1: Load a partitioned SRDD (pSRDD), a point (P), and a number (k) 2: for partition in pSRDD do 3: Calculate distances from the given point Pto every object within partition 4: Maintain a local heap containing the nearest kobjects around the point Pbased on the calculated distances 5: end for 6: Merge results from each partition 2.4.5 GeoTrellis Geotrellis [43] is a high performance geoprocessing engine and programming toolkit. The main objective of GeoTrellis is the incorporation of geospatial analysis functionalities to real time interactive web applications. Focused on raster data processing, the following core problems are behind the development of GeoTrellis. – Create scalable high performance geoprecessing web services. – Parallelize geoprocessing operations to harness multi-core architectures. – Create large-scale distributed geoprocessing services. GeoTrellis helps developers to create simple, standard REST [38] services that return geoprocessing models results. These geoprocessing models are automatically parallelized and optimized. Both creating new operators and composing new operators with existing ones are easy tasks in GeoTrellis.
2.4. Distributed Spatial Data Processing Systems 49 2.4.6 Magellan Magellan [73] is a distributed execution engine for geospatial analytics on big data implemented on top of Spark. Modern database techniques are exploited to optimize geospatial queries. Once the application developer has written standard SQL or dataframe queries to evaluate geometric expressions, the execution engine efficiently lays data out in memory, picks the right query plan, and optimizes the query execution with efficient spatial indexes. Magellan supports multiple spatial data types (e.g., Point,LineString,Polygon, MultiPoint,MultiPolygon) and several spatial predicates (e.g., Intersects,Contains, Within). Spatial indexes in Magellan support the so called Z-Order curves [49] and are mainly used to speed up the spatial join performance. 2.4.7 LocationSpark LocationSpark [95] is a spatial data processing system built as a library on top of Spark, providing spatial query APIs on top of the standard dataflow operators. The main features of LocationSpark are shown next. – Support for spatial querying, spatial data updates, and spatial analytics. A rich set of spatial query operators (e.g., spatial range,spatial kNN,spatial join, and kNN join) is provided. LocationSpark supports data updates and spatio-textual operations. Moreover, spatial data analysis functions such as spatial data clustering, spatial data skyline computation, and spatio-textual topic summarization are provided by LocationSpark. – Support for global and local in-memory spatial data indexes. Furthermore, an efficient spatial Bloom filter has been embedded into LocationSpark’s indexes to avoid unnecessary network communication overhead. – Tracking of frequently accessed spatial data and dynamic flushing of less frequently accessed data into disk. – Storing spatial data as key-value pairs, where the key is a spatial geometric key (e.g., latitude-longitude value, line segment, polyline, rectangle, polygon) and the value type can be specified by the user (e.g., text type). The layered system architecture of LocationSpark is depicted in Fig. 2.5. A detailed discussion of main layers of such architecture is provided below.
50 Chapter 2. Background and related work Figure 2.5: Architecture of LocationSpark. source: Tang et ál., in Proc. VLDB Endow., 2016 [95]. Query Scheduler. This layer is responsible for managing query skew7to mitigate runtime performance degradation of spatial queries. First, LocationSpark dynamically collects statistical information from each partition and detects hotspot data partitions. Then, in order to choose a set of partitions to be further reallocated to optimal workers, a cost model evaluates the overhead of repartitioning the hotspot partitions. Query Executor. Specific query evaluation plans are executed in slave nodes once spatial queries and related data have been scheduled. For various alternative execution plans, LocationSpark evaluates the runtime and memory usage trade-offs. The best execution plan is selected and executed on each slave node. Spatial Indexing. Two layers of spatial indexes (global and local) are provided by LocationSpark. The global index is responsible for partitioning data among worker nodes. Based on the underlying data distribution in space, the global index is built to ensure that each data partition has the same amount of data. A grid index and a region Quad-tree are provided as global indexes in LocationSpark. Furthermore, to match the needs of different scenarios, users can specify the type of the local index (e.g., grid local index, R-tree, a variant of the Quad-tree, or an IR-tree) to be executed on each data partition. 7Similarly to data skew, query skew occurs in a distributed computing environment when some queries are unevenly distributed in space and a number of data partitions are overwhelmed.
2.4. Distributed Spatial Data Processing Systems 51 Memory Management. It is very common for spatial data analysis systems that certain partitions are queried more frequently than others. To deal with this issue, LocationSpark records access frequencies and corresponding time stamps in the spatial index. Then, access frequencies are aggregated to detect the most frequently accessed data. Finally, the most frequently accessed data is cached into memory and the less frequently used data is stored into disk. 2.4.8 Simba Simba [106] is a scalable distributed in-memory analytics engine supporting efficient spatial queries and analytics over big spatial data. The main objectives of Simba are pointed out below: – Simple and expressive programming interfaces. – Low query latency. – High analytics throughput. – Excellent scalability. Next, key features of Simba are highlighted: – Support for rich spatial queries and analytics by extending Spark SQL [14] with core spatial operations. An expressive programming interface for these operations is offered in both SQL and DataFrame API. – Support for spatial indexing to provide low query latency. – Execution of multiple spatial queries in parallel to improve analytical throughput. – Selection of good spatial query plans by using cost-based optimizations (CBO). – Supply of novel algorithms for efficient and scalable execution of spatial operators. The architecture of Simba, depicted in Fig. 2.6, shows the novel components added to the Apache Spark stack. A brief explanation of these components is provided below.
CHAPTER 3 GEODADIS 3.1 Introduction According to the International Energy Agency (IEA), energy efficiency is “a mainstream tool for economic and social development”, with potential “to support economic growth, enhance social development, advance environmental sustainability, ensure energy-system security and help build prosperity” [4]. To reach a significant improvement of energy efficiency in fishing vessels, the Green Fish project1[11] attempted to characterize the generation and consumption of energy during fishing activities. Different data acquisition systems [102] were developed and deployed in several fishing vessels. To leverage the background on designing and deploying the previous data acquisition systems, GeoDADIS [100] enables data acquisition and data dissemination in in-situ sensor platforms. A wide range of technologies must be supported for data dissemination tasks. Data acquisition must fulfill the following requirements. – Any sensor type must be supported. – Any communication channel type must be supported. – Addition of new sensors and new communication channels with a minimum effort. The effort required to add new sensors that measure new parameters through new data acquisition channels must be minimum. Both synchronous and asynchronous data acquisition 1Founded by the Spanish and Galician public administrations.
60 Chapter 3. GeoDADIS channels must be supported. For the former, measures are pulled from the sensors by the framework. For the latter, sensors push measurements to the framework. Two special data acquisition channels used to (1) provide the geographic location of the platform and (2) synchronize the clock, must be specified by the system configuration. Furthermore, additional metadata must be stored in system configuration to enable the specification of the range of historic recorded measurements and the frequency of the sampling process, for each observed parameter. Both client/server (i.e., services query the framework to pull data) and publish/subscriber (i.e., the framework pushes data to the services) data services must be supported. Similarly to data acquisition channels, a minimum effort in the implementation of new data services is required. New remote administration services must be implemented with a minimum effort as well. A minimum effort is also required in the implementation of new data storage technologies for distinct measured parameters. Notice for example the different storage capabilities required by a single temperature value and a complex satellite image. The remainder of this Chapter is organized as follows. Section 3.2 provides a detailed description of the general layered architecture of GeoDADIS, introducing the relevant functionality of each layer and corresponding components. An in-depth description of main components of GeoDADIS architecture is given in Section 3.3. Thus, relevant details about structure and functionality of components DataDissemination, DataAcquisition, Configuration- Manager and DataManager are provided in Section 3.3.1, Section 3.3.2, Section 3.3.3, and Section 3.3.4, respectively. Finally, Section 3.4 introduces an experimental implementation of GeoDADIS developed to monitor people health status in educational environments. 3.2 System Architecture The general component architecture of GeoDADIS, Fig. 3.1, is composed of three main software layers. General purpose functionality related to system control, sensor data management and configuration metadata management is provided by three main components in layer Data and Control Management.DataManager provides persistent storage functionality required by component DataAcquisition and data query functionality demanded by component DataDissemination. Component ConfigurationManager enables the remainder components to access
3.2. System Architecture 61 RemoteAdm DataDissemination DataSubscriber iDataDisQuery DataClient AdmClient iRemoteControl iDataMan DataManager iDataDisPublish ConfigurationManager iSysControl DATA ACQUISITION DataAcquisition SystemControl AsynchChannelManager SynchChannelManager iDataAcqInsert iStRefresh iStamp StampManager DATA AND CONTROL MANAGEMENT EXTERNAL INTERACTION Figure 3.1: GeoDADIS component architecture. configuration settings. Functionality enabling to start up, stop and restart different components is provided by SystemControl. External data acquisition channels enable heterogeneous sensors to provide measures for the sampling processes implemented in layer Data Acquisition at the bottom of GeoDADIS architecture, Fig. 3.1. For each measured parameter of each sensor, administration staff configure both the time range and the frequency of the sampling process. Regarding trigger properties of available sensors, the proposed architecture enables both the time-triggered approach (i.e., sampling frequency is determined by the system) and the event-triggered approach (i.e., sensors deliver measures to be sampled independently of the system). As shown in Fig. 3.1,
62 Chapter 3. GeoDADIS each data acquisition channel must be associated to an external channel manager component. AsynchChannelManagers enable the communication with event-triggered sensors whereas SynchChannelManagers enable to query time-triggered sensor at configured sampling rate. Every new measure generated by an event-triggered sensor is delivered to GeoDADIS by the relevant AsynchChannelManager through the iDataAcqInsert interface. A buffer located in the DataAcquisition component temporarily records the last measure of each sensor. Geo- DADIS queries SynchChannelManagers of time-triggered sensors to sample new measures at relevant sampling rate. A spatio-temporal stamp (i.e,. a time value together with the geographic location of the platform) is requested to StampManager through interface iStamp for each sampled measure. Then, interface iDataMan of component DataManager is used to deliver the sampled measure and relevant time stamp. Component StampManager periodically requests the current value of the spatio-temporal stamp by using the interface iStRefresh. Location and time of stamps are provided by sensors, thus external data acquisition channels must be used to obtain such values. Configuration data provided by ConfigurationManager must store the refresh rate and the data acquisition channels used to get the stamp components. Users and administration staff interaction with GeoDADIS is enabled by the functionality provided by the layer External Interaction. Component DataDissemination is responsible for the proper communication between the DataManager and the available data services. Two different communication approaches are enabled here. The client/server approach enables communication with DataClients whereas publish/subscribe approach is used in communication with DataSubscribers. Notice that both DataClient and DataSubscriber data services are external to GeoDADIS. Interface iDataDisQuery is used by DataClients to deliver queries over stored measures. DataDissemination delegates such queries to DataManager through interface iDataMan. Component DataManager notifies the component DataDissemination through interface iDataDisPublish every time a new measure is recorded. Then, the new measure is delivered to appropriate DataSubscribers which in turn may use interface iDataDis- Query to request queries over stored measures. Configuration and control functionality is provided to administration staff by component RemoteAdmin which delegates actual requests to ConfigurationManager and SystemControl respectively.
3.3. Main Components 63 3.3 Main Components Prior to the detailed description of GeoDADIS’ main components, a brief introduction to three well known design patterns [41] and how they relate to GeoDADIS is provided below. –Singleton pattern: used to enable coordinated access to shared resources by restricting the instantiation of a class to only one object. Configuration data, data acquisition channels, data services and control clients are examples of such resources in GeoDADIS. –Adapter (wrapper) pattern: translates one interface for one class into a compatible interface to enable the cooperation of classes implementing different interfaces. GeoDADIS uses adapters to uniformly access the specific interfaces of control clients, data services and data acquisition channels. –Observer (publish/subscriber) pattern: notifies changes in the state of an object (subject) to a list of objects (observers). The publish/subscribe communication between component DataDissemination and component DataSubscriber implements this pattern in GeoDADIS. 3.3.1 DataDissemination Component DataDissemination combines the three above design patterns to provide flexible external data access. Both data clients, following a client/server approach through the iDataDisQuery interface, and data subscribers, following a publish/subscribe protocol, may access the DataDissemination component. By using the Adapter design pattern, the effort to add a new data service is restricted to the implementation of a new data subscriber or data client adapter class. The Singleton design pattern is used to coordinate the access to the available data services. The main class of the component (DataServicesManager) implements the three interfaces of DataDissemination, Fig. 3.2. Pseudocode describing the implementation of relevant methods of the interfaces that enable the querying (iDataDisQuery)and publishing (iDataDisPublish) of measures is depicted in the figure as well. Methods getParameters and getMeasures of the interface iDataDisQuery enable external data service components to query the recorded data. Implementation of such methods is delegated to relevant methods of components ConfigurationManager and DataManager respectively. The implementation of method publish- Measure of interface iDataDisPublish, which is used by component DataManager to deliver
64 Chapter 3. GeoDADIS iDataServiceAdapter +startDataService() +stopDataService() DataServicesManager +getInstance(): DataServicesManager DataClientAdapterDataSubscriberAdapter DataSubscriber iDataDisPublish +publishMeasure(m: MeasureVO) iDataDisQuery +getParameters(): parameterVO[0..*] +getMeasures(paramId: String, f: STlFilter[0..1]): MeasureVO[0..*] iConfMan iDataMan DataClient iDataSubscriberAdapter +publishMeasure(m: MeasureVO) -instance IDataDisControl -confMan -dataMan -dataSubscribers 1..* -paramId -dataServices getParameters(){ return (confMan.getParameters()); } getMeasures(paramId, f){ return (dataMan.getMeasures(paramId, f)) } publishMeasure(m){ paramId=m.getParam().getParamId(); for each s in dataSubcribers[paramId]{ s.publishMeasure(m); } } «interface» «interface» «interface» «interface» «interface» «interface» «interface» Figure 3.2: UML Class Diagram of component DataDissemination. measures to appropriate external data subscribers, follows a combination of the Observer and Adapter design patterns. When a new measure is received, DataServicesManager (subject class of the Observer pattern) uses the measures’ paramId to notify appropriate DataSubscriberAdapters (both Observer class and Adapter class) by calling method publishMeasure of interface iDataSubscriberAdapter. The list of DataSubscriberAdapter names for each paramId as well as the name of the DataClientAdapter for each data service is part of the component configuration metadata. Adapter classes are also required to enable GeoDADIS to access control functionality (start, restart and stop) of data clients.
3.3. Main Components 65 DataAcqManager +getInstance(): DataAcqManager iDataAcqInsert +insertMeasure(m: MeasureVO) +insertCurrentTime(t: TimeStamp) +insertCurrentLocation(p: Point) AsynchChannelManager «thread» ParameterSampler -paramId: String -samplingInterval: Float +Run() +samplers -paramId iStRefresh +getStamp(): Stamp iDataMan -dataMan iStamp +getCurrentStamp(): Stamp -stampMan iConfMan -confMan iDataAcqControl AsynchDataAcqChannelAdapter -instance AsynchInputBuffer -currentTime: TimeStamp -currentLocation: Point +getInstance(): AsynchDataReciver +getMeasure(param: paramId): MeasureVO +getCurrentTime(): TimeStamp +getCurrentLocation(): Point MeasureComponentVO -componentName: String -componentValue: Object -components MeasureVO -measureTime: TimeStamp -measureLocation: Point -measureSimpleValue: Object -componentName -measures -paramId AsynchDataAcqChannel SynchChannelManager SynchDataAcqChannel -channelAdapter SynchDataAcqChannelAdapter iSynchDataAcqChannelAdapter +getMeasure(paramId: String): MeasureVO +getCurrentTime(): TimeStamp +getCurrentLocation(): Point -instance -buffer iDataAcqChannelAdapter +dataChannelStart() +dataChannelStop() -dacChannels -dacChannelId -dataChannel -timeSource AbstractDataAcqChannel +getMeasure(paramId: String): MeasureVO +getCurrentTime(): TimeStamp +getCurrentLocation(): Point -locationSource getStamp(){ s.setTime(timeSource.getCurrentTime()); s.setLocation(locationSource.getCurrentLocation()): return(s); } Run(){ While(true){ sleep(samplingInterval); m = dataChannel.getMeasure(paramId); s = stampMan.getCurrentStamp(); m.setTime(s.getTime()); m.setPlatformLocation(s.getLocation()); dataMan.insertMeasure(m); } } «interface» «interface» «interface» «interface» «interface» «interface» «interface» «interface» Figure 3.3: UML Class Diagram of component DataAcquisition.
66 Chapter 3. GeoDADIS 3.3.2 DataAcquisition Component DataAcquisition provides general purpose functionality to enable data acquisition over both synchronous and asynchronous data communication channels. Similarly to DataDissemination component, the effort to add a new channel is restricted to the implementation of a new data acquisition channel adapter class due to the use of the Adapter design pattern. The functionality of the DataAcquisition component (see Fig. 3.3 for a graphical representation of its internal structure) is accessed through the DataAcqManager class. Interface iDataAcqControl provides component control functionality whereas interface iStRefresh is used to obtain the current spatio-temporal stamp from the appropriate data acquisition channels. The implementation of the class follows a Singleton design pattern in order to coordinate the concurrent access to both the sampling threads and the data acquisition channels. As it is shown in pseudocode given in the figure for class DataAcqManager, the implementation of method getStamp of the interface iStRefresh access directly the data channels (class Abstract- DataAcqChannel) configured as location and time sources. Regarding the data acquisition process, each sensed parameter is sampled by a different ParameterSampler thread. This implementation is also illustrated with pseudocode in the figure. First, the thread sleeps during a given samplingInterval that is obtained from the configuration data of the specific parameter. After waking up, the thread uses its data acquisition channel, which is also obtained from the configuration data, to obtain the next measure of the parameter. Next, the current stamp obtained from the iStamp interface is used to associate current time and location to the measure. Finally, the stamped measure is delivered to the component DataManager through interface iDataMan. A data acquisition channel (AbstractDataAcqChannel) may access measures of either a synchronous (SynchDataAcqChannel) or an asynchronous (AsynchDataAcqChannel) external channel manager component. Measures of synchronous channel managers are accessed directly through a relevant adapter class (SynchDataAcqChannelAdapter), which is obtained from configuration data and that implements the interface iSynchDataAcqChannelAdapter. On the other hand, measures of asynchronous channel managers are obtained from a buffer (AsynchInputBuffer), that is populated with the last measure of each parameter through the interface iDataAcqInsert. Coordinated access to the buffer is achieved through the use of the Singleton design pattern for its implementation. Notice that an adapter class is still required
3.3. Main Components 67 «interface» iConfMan ConfigurationManagerFacade +getInstance(): ConfigurationManagerFacade -instance SensorVO -sensorId: String -sensorName: Varchar -sensorDesc: String DacqChannelVO -dacqChannelId: String -dacqChannelName: String -dacqChannelDesc: String -dacqChannelType: DacChannelType -dacqChannelAdapterClassName: String -channel 1..* ParameterVO -paramId: String -paramName: String -paramDesc: String -samplingInterval: Float -recordRange: Float -dataAccessClass: String -sensor 1..* DataServiceVO -dataServiceId: String -dataServiceName: String -dataServiceDesc: String -dataServiceType: DataServiceType -dataServiceAdapterClassName: String -params -paramId simpleParameterVO -units: String -dataType: String compoundParameterVO paramComponentVO -componentName: String -units: String -dataType: String -components1..* DataSubscriberVO «interface» iConfManControl -dacqChannels -dacqChannelId -dataServices -dataServiceId -parameters -paramId -timeSource -locSource Figure 3.4: UML Class Diagram of component ConfigurationManager. for asynchronous channel managers in order to access their control functionality (start, stop and restart) from GeoDADIS.
74 Chapter 4. SODA Design stored data. Finally, required operators must be defined to actually perform the analysis tasks defined by users though MAPAL queries. The remainder of this chapter is organized as follows. Section 4.2 is devoted to the definition of the Observation Data Warehouse in SODA. Spatio-temporal and observation data models are described in Section 4.2.1 and Section 4.2.2, respectively. The observation data analysis system proposed in SODA is explained in Section 4.3. First, MAPAL language is fully described in Section 4.3.1. Then, the proposed syntax to define internal processes is explained in Section 4.3.2. Next, Section 4.3.3 defines the required operators to perform observation data analysis. And finally, Section 4.3.4 shows an example of how a MAPAL query is translated into a sequence of SODA operators. 4.2 Observation Data Warehouse Based on general functionalities for an observation data management system specified in Section 1.2, an observation data warehouse should meet the following requirements. – Support for the representation of data coming from the integration of temporal, spatial and spatio-temporal samplings with classical E/R data. – Direct support for observation semantics, through the representation of appropriate required metadata (introduced in Section 1.1), must be provided. – System data types must enable the representation of temporal and spatial data with parametric resolution. Additionally, implicit and explicit castings should be provided to ease the transformation between these resolutions. From the above requirements, an underlying spatio-temporal data model is first defined. Capabilities of this data model go beyond those of an observation data model and enable the processing of any type of spatio-temporal data. To build the definitive data model for the observation data warehouse, structures for observation metadata are then added on top of such underlying data model. 4.2.1 Spatio-temporal Data Model It is well known that relational formalism applied to sampled data processing results in highly inefficient approaches. On the other hand, functional models, which fit well sampled data,
4.2. Observation Data Warehouse 75 have already been used to manage E/R data in the area of Functional Databases [45]. A novel spatio-temporal data model based on the well known mathematical concept of function is defined in this section. Definition and storage of various data types (conventional, temporal and spatial), functions to manipulate data values (Intensional Mappings), functions to record data values (Extensional Mappings and Extensional MappingSets), and data singletons (Constants) are supported. Data types Conventional data types They consist of the data types usually supported by general purpose data management systems, including Boolean,CString (variable size character string), Integer, and Real. In scientific applications, the user knowledge about the precision and scale of numeric data is very important to choose the most appropriate physical representation for real numbers. Hence, a fixed point numeric representation is supported by the parametric data type FixedPrecision(P,S), with conventional semantics for P(precision, i.e., maximum of number of decimal digits) and S(scale, i.e., number of decimal digits in the fractional part). Default and maximum values for Pand Sare system defined: DP (default P), DS (default S), MP (maximum P), and MS (maximum S). Every value Nof type FixedPrecision(P, S) may be written in the form N=n·10−S where nis the integer value in the range (−10P,10P)that actually will be stored, together with Pand S. Thus, Pdetermines the underlying primitive1integer data type used to store n. Additionally, all data types enable the representation of a special undefined value denoted by ⊥. Temporal data types TimeInstant(R) t=n·R|n∈Z;−10MP <n<10MP ∪{⊥} 1Considered primitive data types are: byte,short,int and long.
76 Chapter 4. SODA Design n=0 n=-1 5 seconds n=1 n=10MP-1n=-10MP+1 ··· ··· R 1970/01/01 T 00:00:00.000Z (a) TimeInstant(5). n=0 n=1 n=2 n=17279 5 seconds ··· 00:00:00.000 R (b) Time(5). Figure 4.1: Example of TimeInstant(R) and Time(R) data types where R=5s. Time(R) t=n·R|n∈Z; 0 ≤n·R<24 hours ·3600 seconds hour ∪{⊥} Date(R) TimeInstant(86400) The above temporal data types have been defined to enable the representation of discrete multi-resolution time values, where R(temporal resolution) is a value of data type Double and nis the corresponding index in the defined temporal sampling. Similarly to FixedPrecision values, only Rand nvalues are stored2. Notice that the discrete temporal value t1=n1·R actually represents the continuous time interval defined by the following set of instants {t|n1·R≤t<(n1+1)·R} The semantics of a TimeInstant value is a positive or negative time shift in seconds from an absolute reference time instant, t=0. The most commonly used value for this reference time instant in current DBMS, and also in SODA, is 1970-01-01 T 00:00:00.000000Z. The semantics of a Time value is a positive time shift in seconds from a relative reference time instant, t=0. The value used for this time instant in SODA is the beginning of each day in civil time throughout the world, i.e., 00:00:00.000000. As an example, the available values that can be represented by data types TimeInstant(5) and Time(5) are depicted in Fig. 4.1(a) and Fig. 4.1(b), respectively. 2The primitive integer data type used to store temporal values is long.
4.2. Observation Data Warehouse 77 Spatial data types Point1D(P,R) x=nx·R|nx∈Z;−10P<nx<10P∪{⊥} Point2D(P,R) (x,y) = (nx·R,ny·R)|nx,ny∈Z;−10P<nx,ny<10P∪{⊥} To enable users to represent discrete multi-resolution spatial values, the above spatial data types have been defined, where P(precision) is of type Integer and R(spatial resolution) is of type Double. Notice that the discrete Point1D value x1=n1·Ractually represents the continuous 1D spatial interval x|x∈R;x1−R 2≤x<x1+R 2 and the discrete Point2D value (x1,y1)=(nx1·R,ny1·R)represents the following 2D rectangle (x,y)|(x,y)∈R2;x1−R 2≤x<x1+R 2;y1−R 2≤y<y1+R 2 Similarly to time instant values, a Point1D value is a 1D positive or negative spatial shift in meters from a specific origin point, x=0. In fact, Point1D(P,R) data type defines a 1D spatial sampling in Rand provides a 1D cartesian coordinate system. Fig. 4.2 depicts feasible values, nx, for Point1D(1,1) and Point1D(1,0.5). Likewise, a Point2D value is a positive or negative 2D spatial shift in meters from a specific origin point, x= (0,0).Point2D(P,R) data type defines a 2D spatial sampling in R2and provides a 2D Cartesian coordinate system. Fig. 4.3 depicts all possible values, (nx,ny), for Point2D(1,1) and Point2D(1,0.5). Similarly to previous data types, only P,Rand integer indexes niare stored. As FixedPrecision values, Pdetermines the underlying primitive integer data type used to store the integer index value. Geometric data types Based on Point2D(P,R) data type and on the standard specification defined in [58], the following data types enable the modeling of geometries in 2D euclidean spaces:
78 Chapter 4. SODA Design x 0 -1 1 R2 xxxx xxx xx xx xx x x xx 0 3 2456 7 8 9 -2 -3-4 -5 -6-7-8 -9 -1 -2 -3-4 -5 -6-7-8-9 123 4 56 7 8 9 R1 - Point1D(1,0.5), R2=0.5m- Point1D(1,1), R1=1m x x Figure 4.2: Spatial data types Point1D(1,1) and Point1D(1,0.5). –LineString(P,R): vector polylines defined by sequences of elements of Point2D(P,R). –Polygon(P,R): vector polygons, possibly with holes, whose borders are defined by sequences of elements of Point2D(P,R). –GeometryCollection(P,R): heterogeneous collections of geometries of any of the following data types: Point2D(P,R),Polyline(P,R) and Polygon(P,R). –MultiPoint(P,R): homogeneous collections of Point2D(P,R) geometries. –MultiLineString(P,R): homogeneous collections of LineString(P,R) geometries. –MultiPolygon(P,R): homogeneous collections of Polygon(P,R) geometries. –Geometry(P,R): abstract type that enables the representation of geometries or geometry collections of any of the above 2D data types. Data Structures The following data structures, Dimensions,Extensional MappingSets and Constants enable the modeling and recording of spatio-temporal entity and sampled data. Dimensions ADimension is a finite set of elements of a given data type. More formally, a Dimension dover data type T, denoted d:T, is defined as a non-empty finite subset of T−{⊥}.Dimensions may be defined only over conventional, temporal and spatial data types, and not over geometric data types. Temporal and spatial Samplings are special cases of Dimensions of major interest for the modeling of sampled spatio-temporal data. Thus, if min and max are two values of the same
4.2. Observation Data Warehouse 79 x x x x x x x x x x x - Point2D(1,0.5), R2=0.5m- Point2D(1,1), R1=1m x x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx x xxxx xxx xx xx xx x x xx (0,0) (9,-9)(-9,-9) (-9,9) (9,9) (-9,9) (9,9) (9,-9)(-9,-9) R2 R1 Figure 4.3: Spatial data types Point2D(1,1) and Point2D(1,0.5). Time,TimeInstant,Date or Point1D data type Twhere min <max, then a 1D Sampling S from min to max, denoted S(min,max), is defined as the following Dimension over T S(min,max) = {s∈T|min ≤s≤max } Likewise, if sm= (xm,ym)and sM= (xM,yM)are two values of the same Point2D data type Twhere xm<xMand ym<yM, then a 2D Sampling S from smto sM, denoted S(sm,sM), is defined as the following Dimension over T: S(sm,sM) = {(x,y)∈T|xm≤x≤xM,ym≤y≤yM}
80 Chapter 4. SODA Design Notice that, in general, a Dimension is stored by the explicit recording of each of its elements. However, Samplings are implicitly recorded by the storage of their limit and parametric values. Extensional MappingSets An Extensional MappingSet is a finite set of mappings, Extensional Mappings, with a common domain defined by the Cartesian product of Dimensions. If d1,d2,...,dnis a sequence of not necessarily distinct Dimensions and Tis a data type, then an Extensional Mapping with signature M(d1,d2,...,dn):Tis defined as a function M:d1,d2,...,dn→T. An Extensional MappingSet with signature EM(d1,d2,...,dn|M1:T1,M2:T2,...,Mm: Tm)is defined as the following finite set of Extensional Mappings: EM(d1,d2,...,dn|M1:T1,M2:T2, ... , Mm:Tm) = {M1(d1,d2,...,dn):T1,M2(d1,d2,...,dn):T2, ... , Mm(d1,d2,...,dn):Tm} An Extensional MappingSet EM is extensionally defined by a finite set of nested tuples of the form (d,m), where d∈d1×d2×...×dnand m∈T1×T2×...×Tm. Constants AConstant C of type T, denoted by C:T, is defined as an atomic value of type T. Following a functional database approach [45], Dimensions and Extensional MappingSets enable the modeling of Entities and Relationships between them. Hence for example, Dimensions StationId and MunCode in Fig. 4.4 record identifiers and codes of weather stations and municipalities, respectively. The remainder properties of stations and municipalities are modeled by relevant Extensional MappingSets Station and Municipality. Beyond classical ER data, this model integrates the representation of temporal and spatial 1D and 2D sampled data. For example, Dimensions ObsData and Loc5m in Fig. 4.4 are respectively a temporal Sampling and a 2D spatial Sampling. These Samplings are used to model the phenomenonTime of temperature, humidity and wind speed observations at each weather station in Extensional MappingSet Observation, and the geolocation of elevation observations in Extensional MappingSet Topo.
4.2. Observation Data Warehouse 81 2014/01/01 Start 2015/07/08 End ObsDate P1 Start PN End Loc5m ... 10092 10093 ... StationId ... 15077 15078 ... MunCode StationId Name Station ... 10092 10093 ... ... Punta Candieira Malpica ... Loc ... P3 P4 ... StationId ObsDate Observation ... 10092 10092 ... ... 2014/01/01 2014/01/02 ... Temperature ... 10.93 12.02 ... Humidity ... 90 89 ... Wind Speed ... 19.13 15.66 ... MunCode Municipality ... 15077 15078 ... Name ... Santa Comba Santiago de Compostela ... Geo ... PG1 PG2 ... Dimensions MappingSets P4 P3 Loc5m Elevation Topo 432.13 P1 P5 PN ... ... ... ... P1 PN Figure 4.4: Data Structures. Image in Fig. 4.4 depicts the following geolocated Extensional Mappings:Station.Loc (green starred locations), Municipality.Geo (red line geometries) and Topo.Elevation (gray-scale raster). Intensional Mappings An Intensional Mapping is a function defined over the available data types, either by an algorithm or an analytical expression. If T1,T2,...,Tnis a possibly empty sequence of not necessarily distinct data types and T is also a data type, then an Intensional Mapping with signature M(T1,T2,...,Tn):Tis defined as a function M:T1,T2,...,Tn→T. Primitive intensional mappings. Defined by algorithms, primitive mappings may be already incorporated into the system or provided by the user through user-defined functions.
82 Chapter 4. SODA Design (9,-9)(-9,-9) (-9,9) (9,9) (0,0) Figure 4.5: Point2D Space Filling Curve. They include conventional, temporal and spatial functions like those supported by SQL and relevant extensions [58]. Comparison and arithmetic operators are also supported and defined even for temporal and spatial data types. Fig. 4.5 depicts the specific space filling curve used in SODA for defining a total ordering in Point2D data type. Implicit type castings are automatically applied between data types of the same family during the evaluation of functions and operations. For illustration purposes, primitive intensional mappings defined for all MAPAL data types are shown in Table 4.1. A complete list of intensional mappings defined for specific data types is provided in Appendix A. Argument data types must be compatible for the underlying operation or function to successfully execute each mapping. Thus, overloaded mappings (i.e., different argument versions of each mapping) have been defined for each data type in SODA. For instance, the overloaded mappings defined in SODA for mapping equal(o1,o2)are described below. equal(Boolean b1,Boolean b2): returns the Boolean value true iff (a∧b)∨(¯a∧¯ b). equal(CString s1,CString s2): returns the Boolean value true iff s1is lexicographically equal to s2, i.e., represents the same sequence of char values.
4.2. Observation Data Warehouse 83 Primitive mapping Description distinct(o1,o2)Returns true if o16=o2and returns false otherwise. equal(o1,o2)Returns true if o1=o2and returns false otherwise. getDataType(o)Returns the data type of o. greaterThan(o1,o2)Returns true if o1>o2and returns false otherwise. greaterThanOrEqualTo(o1,o2)Returns true if o1≥o2and returns false otherwise. isDe fined(o)Returns true if ois not null and returns false otherwise. lowerT han(o1,o2)Returns true if o1<o2and returns false otherwise. lowerT hanOrEqualTo(o1,o2)Returns true if o1≤o2and returns false otherwise. setUnde f ined(o)Sets oto null. Table 4.1: Description of primitive common mappings. equal(Numeric n1,Numeric n2): returns the Boolean value true iff n1and n2represent the same Numeric value. If n1and n2have different Numeric data types an implicit casting is applied before the comparison following the next rules. – if one argument (nr) is a Real value and the other argument (ni) is an Integer value, then nris cast to an Integer value yielding trunc(nr). – if one argument (nr) is a Real value and the other argument (nf p) is a FixedPrecision(P,S) value, then nf p is cast to a Real value. – if one argument (ni) is an Integer value and the other argument (nf p) is a Fixed- Precision(P,S) value, then nf p is cast to an Integer value yielding trunc(nf p). equal(Temporal t1,Temporal t2): returns the Boolean value true iff t1and t2represent the same Temporal value. If t1and t2have different Temporal data types or have the same data type but different resolution an implicit casting is first applied. Notice that the casting process may introduce truncation errors. Once both arguments have the same data type and resolution, t1=i1·Rand t2=i2·R, the Boolean value true is returned if and only if i1=i2. Rules for the implicit casting are detailed below. – if both arguments have the same Temporal data type but different resolution (thus, Date data type does not apply here), then the argument with higher resolution is
90 Chapter 4. SODA Design <?xml version="1.0" encoding="UTF-8"?> <xs:schema attributeFormDefault ="unqualified" elementFormDefault ="qualified" targetNamespace="es.usc.citius.de.soda.xoddl" version="1.0.0" xmlns="es.usc.citius.de.soda.xoddl" xmlns:xs="http://www.w3.org/2001/XMLSchema"> <xs:element name="ObservationSchema"> <xs:complexType> <xs:sequence> <xs:element name="ProcessType" type="ProcessType_Type" maxOccurs="unbounded"/> <xs:element name="FeatureType" type="FeatureType_Type" maxOccurs="unbounded"/> </xs:sequence> </xs:complexType> </xs:element> <xs:complexType name="FeatureType_Type" > <xs:sequence> <xs:element name="KeyProperty" type="KeyProperty_Type" maxOccurs="unbounded"/> <xs:element name="Property" type="FeatureProperty_Type" minOccurs="0" maxOccurs="unbounded"/> </xs:sequence> <xs:attribute name="name" type="xs:QName" use="required"/> </xs:complexType> <xs:complexType name="ProcessType_Type" > <xs:sequence> <xs:element name="Property" type="ProcessPropertyType" minOccurs="0" maxOccurs="unbounded"/> </xs:sequence> <xs:attribute name="name" type="xs:QName" use="required"/> <xs:attribute name="type" type="ProcessTypeEnum" use="required"/> <xs:attribute name="triggeredBy" type="TriggeredByType" use="required"/> <xs:attribute name="timeResolution" type="xs:string" use="optional"/> </xs:complexType> <xs:simpleType name="ProcessTypeEnum"> <xs:restriction base="xs:string"> <xs:enumeration value="Internal"/> <xs:enumeration value="External"/> </xs:restriction> </xs:simpleType> <xs:simpleType name="TriggeredByType"> <xs:restriction base="xs:string"> <xs:enumeration value="Time"/> <xs:enumeration value="Event"/> </xs:restriction> </xs:simpleType>
4.2. Observation Data Warehouse 91 <xs:complexType name="KeyPropertyType"> <xs:attribute name="name" type="xs:NCName" use="optional"/> <xs:attribute name="type" type="xs:string" use="optional"/> <xs:attribute name="sampling" type="xs:boolean" use="optional" default="false"/> </xs:complexType> <xs:complexType name="FeaturePropertyType"> <xs:attribute name="name" type="xs:NCName" use="required"/> <xs:attribute name="type" type="xs:string" use="required"/> <xs:attribute name="sourceProcessType" type="xs:QName" use="optional"/> </xs:complexType> <xs:complexType name="ProcessPropertyType"> <xs:attribute name="name" type="xs:NCName" use="required"/> <xs:attribute name="type" type="xs:string" use="required"/> </xs:complexType> </xs:schema> Code 4.1: XML schema definition of XODDL. A more detailed description of main elements in the observation data model, Feature Type and Process Type (and associated Dimensions and ExtensionalMappingSets generated to store their data), are provided below. The UML object diagram of Fig. 4.9 shows a running example used to ease the understanding of the above concepts. Feature Type Feature Type has been defined to enable the integrated modeling of entities and samplings. In the running example, three Feature Types have been defined. Topo: used to model a geographic sampling. The spatial Dimension of Topo is modeled by Key Property Loc5m which defines a sampling Dimension of data type Point2D(9,5). The elevation above the sea level at each point of Loc5m is provided by Feature Property Elevation. Municipality: used to model municipal entities. Each municipality is uniquely identified by non sampling Key Property MunCode.Feature Properties Name and Geo provide the name and geometry of each municipality respectively. Station: models meteorological facilities (entities). Similarly to Municipality, a Key Property StationId uniquely identifies each meteorological station. Non observed
92 Chapter 4. SODA Design : KeyProperty name = MunCode dataType = CString sampling = False : FeatureProperty name = Name dataType = CString : FeatureProperty name = Geo dataType = MultiPolygon(9,0.01) : FeatureProperty name = Location dataType = Point2D(9,0.01) : FeatureProperty name = Name dataType = CString : FeatureProperty name = Elevation dataType = FixedPrecision(7,2) : KeyProperty name = StationId dataType = Integer sampling = False : FeatureType name = Municipality : FeatureType name = Station : FeatureType name = Topo : KeyProperty name = Loc5m dataType = Point2D(9,5) sampling = True : FeatureProperty name = Temperature dataType = Double : ProcessType name = HumidityTempProbe type = External triggeredBy = Time timeResolution = 10 minutes : ProcessProperty name = Description dataType = CString : FeatureProperty name = FrostAlert dataType = CString : ProcessType name = FrostControl type = Internal triggeredBy = Event timeResolution = 10 minutes : ProcessProperty name = Description dataType = CString : FeatureProperty name = Humidity dataType = Integer : FeatureProperty name = WindSpeed dataType = Double : ProcessType name = Anemometer type = Extenal triggeredBy = Time timeResolution = 5 seconds : ProcessProperty name = Description dataType = CString Figure 4.9: Running example (UML object diagram) properties Name and Location provide the name and location of each station. Observed properties Temperature,Humidity,WindSpeed and FrostAlert model observed values provided by relevant observation processes. Temperature and relative humidity observations are provided by an external process of type HumidityTempProbe, wind speed observations are provided by an external process of type Anemometer and the frost risk index for each weather station is generated by an internal process of type FrostControl according to temperature and relative humidity observation values provided by an external process of type HumidityTempProbe.
4.2. Observation Data Warehouse 93 A detailed description of Dimensions and Extensional MappingSets generated within the system for each Feature Type FT is provided below. a) Let KP be a Key Property of data type DT . Different Dimensions will be generated depending on the value of attribute sampling. – if sampling=true, a sampling Dimension FT.KP(lo,hi)is generated, where lo and hi of data type DT define the boundaries of the generated sampling. – if sampling=false, a non sampling Dimension FT.KP :DT is generated. The running example generates the following Dimensions: Topo.Loc5m(lo:Point2D(9,5), hi:Point2D(9,5)) Municipality.MunCode:CString Station.StationId:Integer Notice that duplicate names are not allowed, thus the name of the Feature Type is added as a prefix to the name of the relevant property in order to avoid name conflicts. b) Let KP1,...,KPnbe the Key Properties of FT. The following Extensional MappingSet is also generated FT(FT.KP1,...,FT.KPn|M1:DT1,...,Mm:DTm) where Mi:DTi=FT.FPi(FT.KP1,...,FT.KPn):DTi is the Extensional Mapping generated for the non observed Feature Property FPiof FT . The Extensional MappingSets generated in the running example are the following: Topo( Topo.Loc5m | Elevation:FixedPrecision(7,2) Municipality( Municipality.MunCode | Geo:MultiPolygon(9,0.01)) Station( Station.StationId | Name:CString, Location:Point2D(9,0.01))
94 Chapter 4. SODA Design Similarly to Key Properties, the name of the Feature Type is added as a prefix to the name of the relevant Extensional Mapping to avoid name conflicts. c) Let FP1:DT1,...,FPn:DTnbe the Feature Properties generated by an observation source process of type PT. Let KP1,...,KPnbe the Key Properties of FT. The following Extensional MappingSet is stored to enable the recording of generated observation values. FT.PT(FT.KP1,...,FT.KPn,PT.Time |FP1:DT1,...,FPn:DTn,Process :Integer) where FPi:DTi=FPi(FT.KP1,...,FT.KPn,PT.Time):DTi is the Extensional Mapping recording the observation values of FPi, Process(FT.KP1,...,FT.KPn,PT.Time):Integer records the identifier of the specific observation process used at each time instant to generate observations, and PT.Time is a Dimension generated by PT to store the time instants of observation values. In the running example the following Extensional MappingSets are generated: Station.HumidityTempProbe( Station.StationId, HumidityTempProbe.Time | Temperature:Double, Humidity:Integer, Process:Integer) Station.Anemometer( Station.StationId, Anemometer.Time | WindSpeed:Double, Process:Integer) Station.FrostControl( Station.StationId, FrostControl.Time | FrostAlert:CString, Process:Integer)
4.2. Observation Data Warehouse 95 Process Type Source Processes metadata of Observed Properties are stored by Process Type objects. Each Process Type may be either Time-triggered or Event-triggered at a given timeResolution R.Time-triggered processes generate Temporal samplings at resolution R, i.e., a new observed value is generated every Rseconds since lo to hi (lo and hi are respectively the lowest and highest TimeInstant(R) values defined in the system for observation times generated by the relevant process). The semantics for timeResolution in Event-triggered processes is slightly different, meaning that the time at which the event is fired will be stored as a TimenInstant(R) value. In the running example, process HumidityTempProbe generates temperature and humidity observations every 10 minutes, whereas process Anemometer generates wind speed observations every 5 seconds. Owing to the external nature of these processes, observation values must be provided by external systems. On the contrary, internal FrostControl process computes the frost risk value according to meteorological observations provided by HumidityTempProbe. To generate calculated Feature Properties, the system executes internal processes during ETL tasks. For each Process Type PT , the following Dimensions and Extensional MappingSets are recorded. a) Since the sourceProcess that actually generates observation values for PT may change over time, identifiers of such processes are automatically generated by the system and recorded in a Dimension PT :Integer. Notice that these identifiers are used in Extensional Mappings Process, defined in the above subsection, to identify the sourceProcess that generates each observation. Following Dimensions are generated within the system to record all the required process identifiers in the running example: HumidityTempProbe:Integer Anemometer:Integer FrostControl:Integer b) Let PP1,...,PPnbe the Process Properties of PT. The following Extensional MappingSet is recorded PT.Properties(dPT |M1:PPT1,...,Mn:PPTn) where dPT =PT :Integer
96 Chapter 4. SODA Design is the Dimension of PT identifiers, and Mi:PPTi=PT.PPi(PT):PPTi is the Extensional Mapping that enables the recording of the PPivalues for each PT instance. For the running example, the following Extensional MappingSets are generated: HumidityTempProbe.Properties( HumidityTempProbe | Description:CString) Anemometer.Properties( Anemometer | Description:CString) FrostControl.Properties( FrostControl | Description:CString) c) If PT is a time-triggered Process of resolution Rthen a Sampling PT.Time(lo :TimeInstant(R),hi :TimeInstant(R)) is recorded, where lo and hi are respectively the lowest and highest time instants configured in SODA for observations generated by PT . If PT is an event-triggered Process of resolution Rthen a non-sampling Dimension PT.Time :TimeInstant is stored. These Dimensions store the time instant assigned to each observation value and are, therefore, added to the domain of the relevant Extensional MappingSet that stores such observation values. As shown in the above subsection, Dimensions HumidityTemp- Probe.Time,Anemometer.Time and FrostControl.Time have been added to the domain of Extensional MappingSets Station.HumidityTempProbe,Station.Anemometer and Station.FrostControl, respectively. 4.3 Observation Data Analysis Given the above data models for the representation of both spatio-temporal and observation data, a novel XML based language, called MAPAL (Mapping Analysis Language) is provided
4.3. Observation Data Analysis 97 for the analysis of the proposed data structures. Such a language should fulfill the following requirements based on the generic functionality of an observation data management system. – Follow a declarative paradigm. – Support for OLAP over large data warehouses of spatial observation data. – Support the definition of Internal Processes. – Support the integrated analysis of both entity data and spatial, temporal and spatiotemporal sampled data. – Support for aggregation functionality. 4.3.1 Mapping Analysis Language (MAPAL) Owing to the functional nature of the proposed data models, extensions of well known languages like SQL and XQuery cannot be directly used. However, constructs of these well known languages are used by the hybrid logical-functional paradigm of MAPAL. The combination of such constructs with the XML syntax enables their insertion in currently dominating web services interfaces. Three types of expressions (Functional,Conditional and Aggregate) may be used to define derived Constants,Intensional Mappings and Extensional MappingSets. Additionally, Sampling and Dimension expressions are used to define derived Dimensions. MAPAL syntax and semantics, with illustrative examples, are provided below. The XML Schema definition of MAPAL is shown in Code 4.2. Notice that, for illustration purposes, pieces of code defining Dimensions,Intensional Mappings,Constants and Extensional MappingSets are explained separately and corresponding references have been inserted in Code 4.2. An abstract super type DefinitionType is defined to encapsulate the required common attribute name. As it is shown in following code snippets, all definitions extend DefinitionType. ExternalReferenceType specifies the syntax required to access input and output data channels in order to import and export Constants,Dimensions and Extensional MappingSets. Two required attributes, dataChannel and name, have been defined to specify the data channel and the name of the Constant,Dimension or Extensional MappingSet in the data channel, respectively. Code 4.3 shows the DimensionType schema that enables the definition of sampling and non sampling Dimensions.ConstantType schema that enables the definition of Constants is depicted in Code 4.4. In Code 4.5 the IntensionalMappingType schema that enables the defi-
98 Chapter 4. SODA Design <?xml version="1.0" encoding="UTF-8"?> <xs:schema attributeFormDefault="unqualified" elementFormDefault="qualified" targetNamespace="es.usc.citius.de.mapal" version="1.0.0" xmlns="es.usc.citius.de.mapal" xmlns:xs="http://www.w3.org/2001/XMLSchema" > <!-- ***************************** --> <!-- DEFINITION --> <!-- ***************************** --> <xs:complexType abstract="true" name="DefinitionType"> <xs:attribute name="name" type="xs:NCName" use="required"/> </xs:complexType> <xs:element abstract="true" name="Definition" type="DefinitionType"/> <!-- ***************************** --> <!-- EXTERNAL REFERENCES --> <!-- ***************************** --> <xs:complexType name="ExternalReferenceType"> <xs:attribute name="dataChannel" type="xs:NCName" use="required"/> <xs:attribute name="name" type="xs:QName" use="required"/> </xs:complexType> <!-- DIMENSION DEFINITION: Code 4.3 --> <!-- INTENSIONAL MAPPING DEFINITION: Code 4.5 --> <!-- CONSTANT DEFINITION: Code 4.4 --> <!-- EXTENSIONAL MAPPING DEFINITION: Code 4.6 --> </xs:schema> Code 4.2: XML schema definition of MAPAL. nition of Intensional Mappings is shown. Code 4.6 shows the ExtensionalMappingSetType schema that enables the definition of Extensional MappingSets. Dimensions DimensionType, Code 4.3, specifies the syntax to define Dimensions. Such Dimensions may be defined by using the element <Dimension>.DimensionType extends DefinitionType with an optional attribute storeName. If this attribute is defined, the Dimension is persisted to disk with the name storeName and added to the system catalog in order to be accessible for subsequent operations. Notice that the common attribute name refers to the name of the Dimension in main memory. The new generated Dimension may be exported to a specific number of data channels by adding one optional element <Output> of type ExternalReferenceType for each
4.3. Observation Data Analysis 99 <xs:group name="DimensionSpecification"> <xs:sequence> <xs:element name="ForEach" type="ForEachType" maxOccurs="unbounded"/> <xs:element name="Where" type="xs:string" minOccurs="0"/> <xs:element name="Return" type="xs:string"/> </xs:sequence> </xs:group> <xs:complexType name="ForEachType"> <xs:simpleContent> <xs:extension base="xs:string"> <xs:attribute name="var" type="xs:NCName" use="required"/> </xs:extension> </xs:simpleContent> </xs:complexType> <xs:complexType name="DimensionType"> <xs:complexContent> <xs:extension base="DefinitionType"> <xs:sequence> <xs:choice> <xs:element name="Input" type="ExternalReferenceType"/> <xs:element name="Sampling" type="SamplingSpecificationType"/> <xs:group ref="DimensionSpecification" /> </xs:choice> </sequence> <xs:attribute name="storeName" type="xs:QName" use="optional"/> </xs:extension> </xs:complexContent> </xs:complexType> <xs:element name="Dimension" substitutionGroup="Definition" type="DimensionType"/> <xs:complexType name="SamplingSpecificationType"> <xs:sequence> <xs:element name="Start" type="xs:string"/> <xs:element name="End" type="xs:string"/> </xs:sequence> <xs:attribute name="type" type="xs:string"/> </xs:complexType> Code 4.3: XML schema definition of MAPAL Dimension. output data channel. For the <Output> element, dataChannel is the data channel to which the Dimension is exported and name is the storage name of the Dimension in the data channel. As already stated, a Dimension may be either a sampling Dimension or a non sampling Dimension. Moreover, each non sampling Dimension may be either derived from Dimensions previously loaded in the system or loaded from external data channels.
106 Chapter 4. SODA Design <IntensionalMapping name="meteoProperty" domain="m, s, t"> <When> m="Temperature" </When> <ThenReturn> ffr:Observation.Temperature(s,t) </ThenReturn> <When> m="Humidity" </When> <ThenReturn> ffr:Observation.Humidity(s,t) </ThenReturn> <When> m="WindSpeed" </When> <ThenReturn> ffr:Observation.WindSpeed(s,t) </ThenReturn> </IntensionalMapping> <IntensionalMapping name="IDW" domain="m,p,t"> <ForEach var ="s">StationId </ForEach> <Where> distance(Station.Loc(s), p) < IDWDistance </Where> <Aggregate> sum(meteoProperty(m,s,t)/distance(Station.Loc(s), p)^2) / sum(1/distance(Station.Loc(s), p)^2) </Aggregate> </IntensionalMapping> An Intensional Mapping may be defined by a Conditional Expression that enables the introduction of if-then-else structures. An unbounded number of <When> and <ThenReturn> elements enable the definition of conditions and corresponding results. An optional element <ElseReturn> enables the definition of the default result. In previous code, Intensional Mapping meteoProperty returns temperature, humidity or wind speed observations through aConditional Expression. Extensional MappingSets The schema definition of an Extensional MappingSet of type ExtensionalMappingSetType is shown in Code 4.6. Similarly to Dimensions and Constants, an Extensional MappingSet may be imported from an external data channel and exported to several external data channels. Of course, it may also be persisted into local catalog. The following code shows the definition of Extensional MappingSet Observation imported from Extensional MappingSet Observation in data channel Postgis, persisted to local catalog as ffr:Observation and exported to data channel NetCDF as ObservationFromPostgis. <ExtensionalMappingSet name="Observation" storeName="ffr:Observation"> <Input dataChannel="Postgis" name="Observation"/> <Output dataChannel="NetCDF" name="ObservationFromPostgis"/> </ExtensionalMappingSet> An optional attribute domain may be used to specify a list of Dimensions so that the Cartesian product of such Dimensions is the domain of the resulting Extensional MappingSet.
4.3. Observation Data Analysis 107 <xs:complexType name="ExtensionalMappingType"> <xs:simpleContent> <xs:extension base="xs:string"> <xs:attribute name="name" type="xs:NCName" use="required"/> </xs:extension> </xs:simpleContent> </xs:complexType> <xs:complexType name="ExtensionalMappingSetType"> <xs:complexContent> <xs:extension base="DefinitionType"> <xs:sequence> <xs:choice> <xs:element name="Input" type="ExternalReferenceType"/> <xs:element name="ExtensionalMapping" type="ExtensionalMappingType" minOccurs="0" maxOccurs="unbounded"/> </xs:choice> <xs:element name="Output" type="ExternalReferenceType" minOccurs="0" maxOccurs="unbounded"/> </xs:sequence> <xs:attribute name="storeName" type="xs:QName" use="optional"/> <xs:attribute name="domain" type="xs:string" use="optional"/> </xs:extension> </xs:complexContent> </xs:complexType> <xs:element name="ExtensionalMappingSet" substitutionGroup="Definition" type="ExtensionalMappingSetType"/> Code 4.6: XML schema definition of MAPAL Extensional Mapping. Since an Extensional MappingSet is composed of Extensional Mappings that provide an output value for each domain element, <ExtensionalMapping> enables the definition of an Extensional Mapping through a Functional Expression. In the example below, a forest fire risk index is provided by Extensional MappingSet ForestFire for each combination of time instant (within ffr:ObsDate) and location (within ffr:Loc5m). Resulting Extensional MappingSet is exported to data channel NetCDF as ForestFire. <IntensionalMapping name="normalize" domain="v, min, max"> <When> v < min </When> <ThenReturn>0</ThenReturn> <When> v > max </When> <ThenReturn>1</ThenReturn> <ElseReturn> (v-min)/(max-min) </ElseReturn> </IntensionalMapping> <ExtensionalMappingSet name="ForestFire" domain="p ffr:Loc5m, t ffr:ObsDate"> <ExtensionalMapping name="Risk"> normalize(IDW("Temperature",p,t), minTemperature, maxTemperature)*TemperatureWeight+
108 Chapter 4. SODA Design (IDW("Humidity", p,t)/100)*HumidityWeight+ normalize(IDW("WindSpeed",p, t), minWindSpeed, maxWindSpeed)*WindSpeedWeight+ normalize(slope(p), 0, maxSlope)*SlopeWeight </ExtensionalMapping> <Output dataChannel="NetCDF" name="ras:ForestFire"/> </ExtensionalMappingSet> 4.3.2 Analytical Processes An appropriate syntax to define internal analytical Observation processes executed during ETL tasks is now defined. Similarly to MAPAL and XODDL, a declarative XML-based syntax has been defined to ease the definition of internal processes. Element <Process> enables the definition of such processes. A required attribute <processType> specifies the process data type. An optional element <Description> within <Process> may provide a textual process description. A required element <Definition> comprises the required MAPAL elements to properly define the internal process. – First, a number of optional Dimensions and Intensional Mappings may be defined to be used in remainder elements. – Next, the temporal Dimension of the resulting process is defined by using either an element <TriggeredByTime> or an element <TriggeredByEvent>. Recall that each Process Type PT has a Dimension PT.Time, which is either a sampling Dimension for time-triggered processes or a non sampling Dimension for event-triggered processes. The temporal Dimension of a time-triggered internal Process Type PT with temporal resolution Ris defined with an expression of the following form: <TriggeredByTime> PT1.Time,...,PTn.Time </TriggeredByTime> where each PTi.Time is the temporal Dimension of a Process Type PTi. The semantics are those of the 1D Sampling S(m,M), where m=cast(min{v|v∈PT1.Time ∪PT2.Time ∪...∪PTn.Time}as TimeInstant(R)) M=cast(max{v|v∈PT1.Time ∪PT2.Time ∪...∪PTn.Time}as TimeInstant(R)) An expression of the following form enables the definition of the temporal Dimension of an event-triggered internal Process Type PT of time resolution R:
4.3. Observation Data Analysis 109 <TriggeredByEvent> <Event var="t">PT1.Time, PT2.Time, ... , PTn.Time </Event> <Condition> c(t) </Condition> </TriggeredByEvent> where each PTi.Time is the temporal Dimension of a Process Type PTiand c(t)is a functional expression of Boolean type. The semantics are those of the non sampling Dimension defined by the set {cast(t as TimeInstant(R)) |t∈PT1.Time ∪PT2.Time ∪... ∪PTn.Time ∧c(t)} – Finally, an <ExtensionalMapping> MAPAL element has to be used to define a Feature Property FT.FP observed by an internal Process Type PT. This Extensional Mapping is automatically added to the relevant Extensional MappingSet recording FP observation values. The evaluation of each Extensional Mapping during ETL tasks is restricted to the evaluation of the elements of PT.Time to be imported, avoiding reevaluation of the Extensional Mapping for the whole temporal extension of the data warehouse. Definition of Process Type FrostControl of running example is provided below for illustration purposes. <?xml version="1.0" encoding="utf-8"?> <pd:ProcessDefinitions xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="es.usc.citius.de.soda.ProcessDefinition Soda_ProcessDefinition.xsd" xmlns="es.usc.citius.de.mapal" xmlns:fish="es.usc.citius.de.fish" xmlns:pd="es.usc.citius.de.soda.ProcessDefinition"> <pd:Process processType="FrostControl"> <pd:Description> Calculates the frost risk index for each Station from Temperature and Humidity observation values. </pd:Description> <pd:Definition> <IntensionalMapping name="FrostAlert" domain="t, h"> <When> t < 0 AND h > 95 </When> <ThenReturn> VERY HIGH </ThenReturn> <When> t < 0 AND h > 85 AND h <= 95 </When> <ThenReturn> HIGH </ThenReturn>
110 Chapter 4. SODA Design <When> t < 0 AND h > 75 AND h <= 85 </When> <ThenReturn> MEDIUM </ThenReturn> <When> t < 0 AND h > 65 AND h <= 75 </When> <ThenReturn> LOW </ThenReturn> <When> t < 0 AND h > 55 AND h <= 65 </When> <ThenReturn> VERY LOW </ThenReturn> <ElseReturn> VERY LOW </ElseReturn> </IntensionalMapping> <IntensionalMapping name="StationsInRisk" domain="t"> <ForEach var="s">Station.StationId </ForEach> <Where> Station.HumidityTempProbe.Temperature(s, t) < 0 AND Station.HumidityTempProbe.Humidity(s, t) > 85 </Where> <Aggregate> not EMPTY(s) </Aggregate> </IntensionalMapping> <pd:TriggeredByEvent> <pd:Event var="t">HumidityTempProbe.Time </pd:Event> <pd:Condition> StationsInRisk(t)</pd:Condition> </pd:TriggeredByEvent> <ExtensionalMapping name="FrostAlert" domain="Station.StationId s, FrostControl.Time t"> <Return> FrostAlert(Station.HumidityTempProbe.Temperature(s, t), Station.HumidityTempProbe.Humidity(s, t)) </Return> </ExtensionalMapping> </pd:Definition> </pd:Process> </pd:ProcessDefinitions> An accurate implementation of a frost risk alert system is out of the scope of this Thesis, thus some simplifications are made for an easy understanding. For this example, the frost risk is calculated based on the following rules: – if T<0∧RH >95, then frost risk is VERY HIGH – if T<0∧85 <RH ≤95, then frost risk is HIGH – if T<0∧75 <RH ≤85, then frost risk is MEDIUM – if T<0∧65 <RH ≤75, then frost risk is LOW – if T<0∧55 <RH ≤65, then frost risk is VERY LOW
4.3. Observation Data Analysis 111 where Tis the measured temperature in Celsius degrees and RH is measured relative humidity in percentage units. Every time that FrostControl code is executed a new process identifier is automatically generated by the system. This identifier is stored in both Dimension FrostControl and Extensional Mapping Process of Extensional MappingSet Station.FrostControl. Process FrostControl is defined as an event-triggered process that is fired every time a station is in risk VERY HIGH or HIGH, i.e., measured temperature is below 0ºC and measured relative humidity is above 85%. Such conditions are implemented by Intensional Mapping StationsInRisk.Feature Property Station.FrostAlert generates output observations applying the above rules (implemented by Intensional Mapping FrostAlert) over temperature and relative humidity values. 4.3.3 System Operators Query processing performs the evaluation of the above MAPAL expressions. A many-sorted algebra over three different data structures (Dimensions,Constants and Extensional MappingSets) that enables the evaluation of MAPAL queries is defined next. Operations of this algebra are classified into three different groups according to the result structure that they produce. The general syntax of such operations is the following: operatorName[paramList]...[paramList](argumentList) where paramList is a comma separated list of parameters and argumentList is a comma separated list of arguments. Dimension Operators –ImportDimension[Name][channelName][storageName]. Imports a Dimension from an external data channel. Parameter Name is the name of the new Dimension. Parameter channelName is the name of the external data channel. Parameter storageName is the name, in the external data channel, of the Dimension to import. –ScanDimension[Name]. Reads a Dimension from local catalog. Parameter Name is the name of the Dimension to read. –SamplingDimension[Name](k1,k2). Generates an new in-memory sampling Dimension with all the values of Sampling S(k1,k2). Parameter Name is a CString containing the
112 Chapter 4. SODA Design name of the generated Dimension. Operands k1and k2are Constants with identical temporal or spatial data type, obtained from some Constant operator. –Union(d1,d2). Computes the Union of Dimensions as defined in 4.3.1. Operands d1 and d2are Dimensions produced by some other operator. –Intersection(d1,d2). Computes the Intersection of Dimensions as defined in 4.3.1. Operands d1and d2are Dimensions produced by some other operator. –ProjectDimension[d][c](MS). Generates a result Dimension containing all the distinct elements of parameter dwhere Extensional Mapping c has a true value. Operand MS is an Extensional MappingSet produced by a relevant operation. Parameter dis the name of either a Dimension or an Extensional Mapping of MS. Optional parameter cis the name of a Boolean Extensional Mapping of MS. Notice that if dis a sampling Dimension and cis provided, then the resulting Dimension will not be a sampling Dimension anymore. –StoreDimension[Name](d). Saves Dimension d to disk using Name as its storage name. Parameter Name is a CString and parameter dis a Dimension generated by some operator. If dis a sampling Dimension only the metadata and its limits are written into the local catalog. –ExportDimension[channelName][storageName](d). Exports the Dimension d to an external data channel. Parameter dis a Dimension generated by some operator. Parameter channelName is the name of the external data channel. Parameter storageName is the name, in the external data channel, of the new exported Dimension. Extensional MappingSet Operators –ImportMappingSet[Name][channelName][storageName][domain]. Imports an external Extensional MappingSet from a data channel. Parameter Name is the name of the new Extensional MappingSet. Parameter channelName is the name of the external data channel. Parameter storageName is the name, in the external data channel, of the Extensional MappingSet to import. Parameter domain is a comma separated list of Dimensions, i.e., the domain of the Extensional MappingSet. Notice that all Dimensions within the domain must be already imported before importing the Extensional MappingSet.
4.3. Observation Data Analysis 113 –Product[Name](d1,...,dn). Generates a new Extensional MappingSet, without Extensional Mappings, whose domain is the Cartesian product d1×...×dn. Parameter Name is the name of the new Extensional MappingSet. Each operand diis a Dimension produced by some operator. –Product(MS,d). Generates a result Extensional MappingSet whose domain is the Cartesian product of dwith the domain of MS. All Extensional Mappings of MS are kept in the resulting Extensional MappingSet. –ProjectMappingSet[Name][s1,...,sn](MS). Generates a new Extensional MappingSet whose domain is equal to the domain of MS and with one Extensional Mapping for each si. Operand MS is an Extensional MappingSet produced by some operator. Parameter Name is the name of the new Extensional MappingSet. Each parameter siis either a Dimension or an Extensional Mapping of MS. –EvaluateIntensionalMappings[m1,...,mn](MS). Adds new Extensional Mappings to MS. Operand MS is an Extensional MappingSet produced by some operator. Each mi is an intensional mapping expression of the form newMappingName =pm(s1,...,sm), where pm is the name of a primitive mapping and each siis the name of either a Dimension or an Extensional Mapping of MS. The actual expression of miis generated from the expression of each siobtained from MS. If some Extensional Mapping of MS has been already computed by an expression equivalent to the actual expression of mi, the new name newMappingName is added to the list of names referencing such Extensional Mapping. Otherwise, the primitive mapping is evaluated for each element of the domain of MS to produce a new Extensional Mapping called newMappingName. –EvaluateExtensionalMapping[m](MS). Appends a new Extensional Mapping to MS (an Extensional MappingSet produced by some operator). Parameter mis an extensional mapping expression of the form newMappingName =ems.em(s1,...,sm), where ems references an Extensional MappingSet,em references an Extensional Mapping of ems, and each siis the name of either a Dimension or an Extensional Mapping of MS. The domain of ems must be defined by the Cartesian product of mdimensions d1×d2×...× dmin such a way that the data type of each siis compatible with the data type of each di. If some Extensional Mapping of MS has been computed by an expression equivalent to m, then the name newMapping will be added to the names of such Extensional Mapping, otherwise em will be evaluated.
114 Chapter 4. SODA Design –EvaluateConstant[Name](MS,k). Generates a new Extensional Mapping from the expression of k. Parameter Name is the name of the new Extensional Mapping. Operand MS is an Extensional MappingSet produced by some operator. Operand kis a Constant value. If the expression of kis already present in some Extensional Mapping of MS, then the name newMapping will be added to such Extensional Mapping. Otherwise, a new Extensional Mapping called newMapping is added to MS, which records the value obtained from kfor each element of the domain. The data type and expression of the new mapping will be obtained from k. The domain of the new mapping in MS will be empty. –EvaluateAggregateMappings[Name][groupBy][orderBy][c][ag1,...,agn](MS).MS is an Extensional MappingSet obtained from some operator. Parameter groupBy is a list of names of Dimensions of MS. Parameter ordSpec is an optional ordering specification composed of a list of pairs (s,o), where each sis the name of either a Dimension or an Extensional Mapping of MS, and ois an ordering direction, either ascending or descending. Optional parameter cis the name of an Extensional Mapping of MS of Boolean data type. Each agiis an expression of the form newMappingi= AggMappingi(s1,s2,...,sn), where AggMappingiis the name of a primitive aggregate mapping and each sjis the name of either a Dimension or an Extensional MappingSet of MS. The result Extensional MappingSet will have as domain the list of Dimensions referenced in groupBy and as Extensional Mappings the list of mappings of MS whose domain does not contain Dimensions not present in groupBy, together with the new mappings generated by ag1,ag2,...,agn. To achieve this, first MS is grouped by Dimensions in groupBy and Extensional Mappings whose domain does not contain Dimensions out of groupBy. Next, the sequence of tuples of each group is filtered using c. Then, the result is ordered according to ordSpec. Finally, each aggregate mapping is evaluated in the ordered sequence of tuples to produce just one value for each group. If some Extensional Mapping of MS already has the same expression of one of the aggregates aggMappingito be evaluated, then newMapping will be added to the names of such Extensional Mapping and aggMappingiwill not be evaluated again. –StoreMappingSet[Name](MS). Parameter Name is a CString and operand MS is an Extensional MappingSet obtained from some operation. The metadata of MS and the data of each of its Extensional Mappings is saved to disk. This operation has not effect if MS does not have any Extensional Mapping.
4.3. Observation Data Analysis 115 –ExportMappingSet[channelName][Name](MS). Exports the Extensional MappingSet MS to an external data channel. Parameter MS is an Extensional MappingSet generated by some operator. Parameter channelName is the name of the external data channel. Parameter Name is the name, in the external data channel, of the new exported Extensional MappingSet. Constant Operators –Literal[Name][l]. Transforms a literal into a data value that is recorded in the result Constant. Parameter Name is the name of the new generated Constant. Parameter lis the CString representation of a literal. The expression associated to the Constant will be l. –ImportConstant[Name][channelName][storageName]. Imports a Constant from an external data channel. Parameter Name is the name of the new Constant. Parameter channelName is the name of the external data channel. Parameter storageName is the name, in the external data channel, of the Constant to import. –ScanConstant[Name]. Reads a Constant from local catalog. Parameter Name is the name of the Constant to read. –EvaluateIntensionalMapping[Name][m](k1,...,kn). Each operand kiis a Constant obtained from some operation and mis the name of a primitive mapping. Mapping mis evaluated using the values of kias parameters to obtain the value for the result Constant. The expression of the result will also be generated from mand the expression of each ki. –EvaluateExtensionalMapping[Name][m](k1,...,kn). It is similar to the above operation, however now mreferences an Extensional Mapping of an Extensional MappingSet. The data of both mand the Dimensions of the domain of mhave to be accessed to obtain the result value. The expression of the result will be generated from mand the expression of each ki. –EvaluateAggregateMapping[Name][orderBy][c][agg](MS).MS is an Extensional MappingSet obtained from some operation. Parameter ordSpec is an ordering specification composed of either Dimensions or Extensional Mappings of MS and ordering directions (ascending or descending). Parameter cis an Extensional Mapping of MS of Boolean
122 Chapter 5. MAPAL Implementation MapalValue +distinct(o2): BooleanValue +equal(o2): BooleanValue +getDataType(): DataTypeMetadata +greaterThan(o2): BooleanValue +greaterThanOrEqualTo(o2): BooleanValue +isDefined(): BooleanValue +lowerThan(o2): BooleanValue +lowerThanOrEqualTo(o2): BooleanValue +setUndefined() BooleanValue -value: Boolean[0..1] +BooleanValue(v: Boolean) +and(o: BooleanValue): BooleanValue +getValue(): Boolean +not(): BooleanValue +or(o: BooleanValue): BooleanValue +toCString(): CStringValue CStringValue -value: String[0..1] +CStringValue(v: String) +concat(s: CString): CStringValue +getValue(): String +length(): FixedPrecisionValue +lower(): CStringValue +toBoolean(): BooleanValue +toDate(): DateValue +toFixedPrecision(): FixedPrecisionValue +toFixedPrecision(p: integer, s: integer): FixedPrecisionValue +toInteger(): IntegerValue +toReal(): RealValue +toTime(): TimeValue +toTime(r: FixedPrecisionValue): TimeInstantValue +toTimeInstant(): TimeInstantValue +toTimeInstant(r: FixedPrecisionValue): TimeInstantValue +upper(): CStringValue NumericValue +abs(): NumericValue +acos(): NumericValue +asin(): NumericValue +atan(): NumericValue +atan2(y: NumericValue): NumericValue +ceil(): IntegerValue +cos(): NumericValue +divide(n: NumericValue): NumericValue +floor(): IntegerValue +ln(): NumericValue +log(): NumericValue +mod(a: NumericValue): IntegerValue +multiply(n: NumericValue): NumericValue +power(e: NumericValue): NumericValue +round(): IntegerValue +round(n: IntegerValue): FixedPrecisionValue +sin(): NumericValue +sqrt(): NumericValue +subtract(n: NumericValue): NumericValue +sum(n: NumericValue): NumericValue +tan(): NumericValue RealValue -value: Double[0..1] +RealValue(v: Double[0..1]) +getValue(): Double +toBoolean(): BooleanValue +toCString(): CStringValue +toFixedPrecision(): FixedPrecisionValue +toFixedPrecision(p: IntegerValue, s: IntegerValue): FixedPrecisionValue +toInteger(): IntegerValue +toPoint1D(): Point1DValue +toPoint1D(p: IntegerValue, r: RealValue): Point1DValue FixedPrecisionValue -value: Number[0..1] -precision: integer -scale: integer +FixedPrecisionValue(p: IntegerValue[0..1], s: IntegerValue[0..1], v: Number[0..1]) +getPrecision(): IntegerValue +getScale(): IntegerValue +getStoredValue(): Number +getValue(): RealValue +toBoolean(): BooleanValue +toFixedPrecision(p: IntegerValue, s: IntegerValue): FixedPrecisionValue +toInteger(): IntegerValue +toPoint1D(): Point1DValue +toPoint1D(p: IntegerValue, r: RealValue): Point1DValue +toReal(): RealValue +toCString(): CStringValue IntegerValue -value: Long[0..1] +IntegerValue(v: Long[0..1]) +getValue():IntegerValue +toBoolean(): BooleanValue +toCString(): CStringValue +toReal(): RealValue +toFixedPrecision(): FixedPrecisionValue +toFixedPrecision(p: IntegerValue, s: IntegerValue): FixedPrecisionValue +toPoint1D(): Point1DValue +toPoint1D(p: IntegerValue, r: RealValue): Point1DValue Figure 5.2: Class diagram of Conventional data types in the prototype implementation.
5.2. Data Types Implementation 123 5.2.1 Conventional Data Types Implementation The class diagram of conventional MAPAL data types implemented in the developed prototype is depicted in Fig. 5.2. An abstract class MapalValue encapsulates the common operations (primitive mappings) that every data type must implement, as defined in Table 4.1. Thus, the rest of the classes representing MAPAL data types must inherit from MapalValue. Class BooleanValue has an attribute value of Java type Boolean to store the actual boolean value. Moreover, BooleanValue implements the primitive boolean mappings defined in Table A.1, and provides the required constructor and accessor methods. Similarly to BooleanValue, class CStringValue provides constructor and accessor methods, and implements the string primitive mappings defined in Table A.2. The Java type String is used to store the actual CString value. Primitive mappings common for all numeric (Integer,Real,FixedPrecision) data types, defined in Table A.3, are encapsulated by the abstract class NumericValue from which all the numeric classes inherit. Attribute value of IntegerValue and RealValue are of Long and Double Java types, respectively. Whereas, to store the actual integer value of attribute value in FixedPrecisionValue, the abstract Java type Number is used. Depending on the combination of attributes precision and scale, the most appropriate Java integer type inheriting from Number (Byte,Short,Integer and Long) is actually used. Furthermore, constructors, accessor methods and appropriate casting mappings are defined for all numeric data types. Notice that constructor and accessor methods are defined for internal use within the implementation and are not accessible to MAPAL users through primitive mappings. Conventional data type values are generated automatically by the system from literal values in MAPAL sentences. 5.2.2 Temporal Data Types Implementation A major added value in MAPAL are those data types that enable the definition of temporal Samplings. Abstract class SamplingValue in Fig. 5.3 defines the primitive mapping subtract for all those data types that may be used in Sampling definitions. Inheriting from SamplingValue, abstract class TemporalValue encapsulates attributes and mappings common for all defined temporal data types. Specifically, attribute resolution uses the Java primitive type double to store the temporal resolution. Moreover, the mapping sum and an overloaded version of mapping subtract, defined in Table A.4, are added here together with the accessor method for attribute resolution. Similarly to numeric data types, TimeValue
124 Chapter 5. MAPAL Implementation MapalValue +distinct(o2): BooleanValue +equal(o2): BooleanValue +getDataType(): DataTypeMetadata +greaterThan(o2): BooleanValue +greaterThanOrEqualTo(o2): BooleanValue +isDefined(): BooleanValue +lowerThan(o2): BooleanValue +lowerThanOrEqualTo(o2): BooleanValue +setUndefined() TemporalValue -resolution: double +getResolution(): RealValue +subtract(n: IntegerValue): TemporalValue +sum(n: IntegerValue): TemporalValue TimeInstantValue -value: long[0..1] +TimeInstantValue(r: RealValue, t: IntegerValue) +toCString(): CStringValue +toTime(): TimeValue +toTime(r: RealValue): TimeValue +toTimeInstan(r: RealValue): TimeInstantValue TimeValue -value: intType[0..1] +TimeValue(r: RealValue, t: intType) +toCString(): CString Value +toTime(r: RealValue): TimeValue +toTimeInstant(): TimeInstantValue +toTimeInstan(r: RealValue): TimeInstantValue intType DateValue +DateValue(t: IntegerValue) +toCString(): CStringValue Point1DValue -coord: coordType[0..1] -precision: integer -resolution: double +Point1DValue(p: IntegerValue[0..1], r: RealValue[0..1], c: coordType[0..1]) +Point1DValue(coord: FixedPrecisionValue) +Point1DValue(coord: RealValue) +getCoord(): FixedPrecisionValue +getResolution(): FixedPrecisionValue +getPrecision(): IntegerValue +getValue(): coordType +subtract(n: IntegerValue): Point1DValue +sum(n: IntegerValue): Point1DValue +toCString(): CStringValue +toFixedPrecision(): FixedPrecisionValue +toFixedPrecision(p: IntegerValue, s: IntegerValue): FixedPrecisionValue +toInteger(): IntegerValue +toPoint1D(p: IntegerValue, r: RealValue): Point1DValue +toReal(): RealValue coordType SamplingValue +subtract(v: SamplingValue): IntegerValue Figure 5.3: Class diagram of Temporal and Point1D data types in prototype implementation. and TimeInstantValue implement constructors and the appropriate casting mappings also defined in Table A.4. While Java primitive type long is used to store the integer temporal value in TimeInstantValue, class TimeValue uses the most appropriate Java primitive integer type (intType) to store such integer value. The selection of intType is strongly dependent on the attribute resolution. Since MAPAL data type Date is a shortcut for TimeInstant(86400), class DateValue only defines a constructor and an overloaded version of mapping toCString. Similarly to conventional data types, constructor and accessor methods are only for internal use and are not
5.2. Data Types Implementation 125 accessible by MAPAL users. The system generates them from literal values in MAPAL sentences. 5.2.3 Point1D Data Type Implementation Another important MAPAL data type contributed by this Thesis is Point1D. Definition of spatial 1D Samplings is enabled by this data type. Thus, class Point1DValue in Fig. 5.3, also inheriting from SamplingValue, stores attributes precision and resolution using Java primitive types integer and double, respectively. As TimeValue does, the most appropriate Java primitive integer type (coordType) is used to store the integer spatial attribute coord. In this case, such Java type is selected depending on the value of attributes precision and resolution. Primitive Point1D mappings defined in Table A.5 are also present in Point1DValue together with constructors and accessor methods. As previous data types, such constructor and accessor methods are not accessible by MAPAL users. 5.2.4 Point2D Data Type Implementation Probably the most important contribution of MAPAL, regarding data types, is the definition of data type Point2D that enables, in turn, the definition of 2D spatial Samplings. Class Point2DValue in Fig. 5.4 encapsulates the Point2D mappings defined in Table A.6, together with constructors and accessor methods. Attributes precision and resolution store relevant Point2D parameters using Java primitive types integer and double, respectively. Attribute value uses a JTS2[60] object Point to store the 2D spatial coordinates. Since MAPAL data type Point2D also represents a geometry, Point2DValue implements the primitive mappings defined in Geometry2DInterface, which contains the vast majority of primitive mappings defined for all Geometries in Table A.7. Recall that Point2D enables the definition of 2D spatial Samplings, thus Point2DValue inherits from SamplingValue as well. Point2DValue constructor and accessor methods are also not accessible by MAPAL users. 5.2.5 Geometric Data Type Implementation The developed prototype also implements the geometric data types defined in Section 4.2.1. Thus, class Geometry2DValue in Fig. 5.4 provides constructors and accessor methods to create 2Java Topology Suite.
126 Chapter 5. MAPAL Implementation MapalValue +distinct(o2): BooleanValue +equal(o2): BooleanValue +getDataType(): DataTypeMetadata +greaterThan(o2): BooleanValue +greaterThanOrEqualTo(o2): BooleanValue +isDefined(): BooleanValue +lowerThan(o2): BooleanValue +lowerThanOrEqualTo(o2): BooleanValue +setUndefined() Point2DValue -value: Point -precision: integer -resolution: double +Point2DValue(p: IntegerValue[0..1], r: RealValue[0..1], v: Point[0..1]) +Point2DValue(c1: FixedPrecisionValue, c2: FixedPrecisionValue) +4neigh(p: Point2DValue): BooleanValue +8neigh(p: Point2DValue): BooleanValue +getPosition(): IntegerValue +getValue(): Point +getX(): FixedPrecisionValue +getXint(): IntegerValue +getY(): FixedPrecisionValue +getYint(): IntegerValue +shift(x: IntegerValue, y: IntegerValue): Point2DValue +toPoint2D(p: IntegerValue, r: RealValue): Point2DValue SamplingValue +subtract(v: SamplingValue): IntegerValue Geometry2DValue -precision: integer -resolution: double -value: Geometry +Geometry2DValue(p: IntegerValue[0..1], r: FixedPrecisionValue[0..1], v: LineString[0..1]) +Geometry2DValue(p: IntegerValue[0..1], r: FixedPrecisionValue[0..1], v: Polygon[0..1]) +Geometry2DValue(p: IntegerValue[0..1], r: FixedPrecisionValue[0..1], v: GeometryCollection[0..1]) +Geometry2DValue(p: IntegerValue[0..1], r: FixedPrecisionValue[0..1], v: MultiPoint[0..1]) +Geometry2DValue(p: IntegerValue[0..1], r: FixedPrecisionValue[0..1], v: MultiLineString[0..1]) +Geometry2DValue(p: IntegerValue[0..1], r: FixedPrecisionValue[0..1], v: MultiPolygon[0..1]) +Geometry2DValue(v: Point2DValue[1..*], isLineString: BooleanValue) +Geometry2DValue(e: LineString2DValue, h: LineString2DValue[0..*]) +Geometry2DValue(m: LineString2DValue[1..*]) +Geometry2DValue(m: Polygon2DValue[1..*]) +area(): RealValue +centroid(): Point2DValue +exterior(): Geometry2DValue +endPoint(): Point2DValue +getValue(): Geometry +holes(): Geometry2DValue +isClosed(): BooleanValue +isRing(): BooleanValue +isSimple(): BooleanValue +length(): RealValue +perimeter(): RealValue +startPoint(): Point2DValue +voronoi(): Geometry2DValue «interface» Geometry2DInterface +buffer(n: RealValue): Geometry2DInterface +contains(g: Geometry2DInterface): BooleanValue +convexHull(): Geometry2DInterface +crosses(g: Geometry2DInterface): BooleanValue +difference(g: Geometry2DInterface): Geometry2DInterface +disjoint(g: Geometry2DInterface): BooleanValue +distance(g: Geometry2DInterface): RealValue +envelope(): Geometry2DInterface +equals(g: Geometry2DInterface): BooleanValue +fromGml(s: CString): Geometry2DInterface +fromWkt(s: CString): Geometry2DInterface +getPrecision(): IntegerValue +getResolution(): FixedPrecisionValue +gml(): CStringValue +intersection(g: Geometry2DInterface): Geometry2DInterface +intersects(g: Geometry2DInterface): BooleanValue +overlaps(g: Geometry2DInterface): BooleanValue +symDifference(g: Geometry2DInterface): Geometry2DInterface +touches(g: Geometry2DInterface): BooleanValue +union(g: Geometry2DInterface): Geometry2DInterface +within(g: Geometry2DInterface): BooleanValue +wkt(): CStringValue Figure 5.4: Class diagram of Point2D and Geometric data types in prototype implementation.
5.3. Data Structures Implementation 127 and access geometries, respectively. Furthermore, primitive mappings defined in Table A.8, Table A.9, Table A.10 and Table A.11 are also incorporated into Geometry2DValue. Property value uses a JTS abstract object Geometry to store the geometry value. Current implemented3subclasses of Geometry include LineString,Polygon,MultiLineString,MultiPoint and MultiPolygon. Even though all mentioned primitive mappings are implemented in Geometry2DValue, only appropriate ones may be represented depending on the actual subclass of Geometry. Similarly to Point2DValue,Geometry2DValue stores geometry precision and resolution, and implements the primitive mappings defined in Geometry2DInterface. 5.3 Data Structures Implementation Efficient structures are required for recording data and metadata related to Dimensions,Extensional MappingSets and Constants both in disk and main memory. A detailed description of such structures implemented in the developed prototype is provided in this section. 5.3.1 In-Memory Structures Implementation Each Dimension,Extensional MappingSet and Constant, either obtained from disk or calculated, is recorded in main memory in a structure composed of a header, with appropriate metadata, and a data area. Metadata recorded in a Dimension header include name, size and data type. A Dimension might be obtained from disk or generated in memory as a result of some operation. A boolean IsStored is kept in the header to identify these two types of Dimensions. Attribute Storage- Name is used to identify Dimensions stored in the local catalog (introduced in next Section). Boolean attribute IsMaterialized shows whether a Dimension is materialized in main memory or not. A non-materialized Dimension (Fig. 5.5(b)) only records unique element identifiers in main memory, which reference element positions. Such references are generated in main memory using the size of the Dimension. A materialized Dimension might have been generated in memory, in such case it has only data values (Fig. 5.5(c)), or it may have been read from disk, in such case it has both references and data values (Fig. 5.5(a)). Non-materialized stored Dimensions enable the implementation of late materialization [2], which avoids having to read data values from disk which are not involved in any calculation. 3http://locationtech.github.io/jts/javadoc/.
128 Chapter 5. MAPAL Implementation Name: StationId Size: 80 IsSampling: False IsStored: True StorageName: StationId IsMaterialized: True DataType: Integer DataChannel: PostGis Header Data Refs Values 0 1 2 3 4 5 6 ... 893 894 896 899 1000 1001 1002 ... Name: StationId Size: 80 IsSampling: False IsStored: True StorageName: StationId IsMaterialized: False DataType: Integer DataChannel: PostGis Header Data Refs Values 0 1 2 3 4 5 6 ... (a) Materialized stored Dimension (b) Non-materialized stored Dim- Name: ObsDate Size: 365 IsSampling: True IsStored: False StorageName: IsMaterialized: True DataType: TimeInstant(86400) DataChannel: Start: 16060 End: 16425 Header Data Refs Values 16060 16061 16062 16063 16064 16065 16066 ... (c) Materialized non-stored Dimension ension Figure 5.5: Example of in-memory Dimension structures. An Extensional MappingSet header must record global metadata (Name,IsStored and StorageName) as well as metadata of building Dimensions and Extensional Mappings. Fig. 5.6 illustrates an Extensional MappingSet computed in memory to record the temperature, elevation and location of each meteorological station at each observation date. Extensional Mapping Temperature records the result of the evaluation of the expression: Observation.Temperature(StationId,ObsDate).Extensional Mapping m1 is a temporary one that records the result of the expression Station.Loc(StationId).Extensional Mapping Loc shares both data and metadata with m1, therefore it will share also the same header entry. Extensional Mapping m2 is also temporary and records the result of the expression: cast(m1(StationId,ObsDate)AS Point2D(7,5)). Finally, Extensional Mapping Elevation records the result of the expression Topo.Elevation(m2(StationId,ObsDate)). It is noticed that Extensional Mapping m2 records Point2D(7,5) values that reference elements in Dimension Loc5m. These references are needed to obtain the elevation values from disk. It is therefore noticed that beyond the references and values of the Dimensions and the values of each Extensional Mapping, both Dimensions and Extensional Mappings might record additional columns that contain references to stored Dimensions. Besides, the header of each Extensional Mapping keeps record of the subset of Dimensions from which it is dependent, i.e., its real domain and the expression that was used to compute it. The former is used during aggregate operations, as it will be shown in Section 5.6, whereas the latter helps in avoiding the computation of duplicate Extensional Mappings in the same Extensional MappingSet as it is the case of Extensional Mappings m1 and Loc.
5.3. Data Structures Implementation 129 ReferencedBy m2 {Loc5m} RefDimensions RefDims Dimensions Name StationId ObsDate False True IsSampling True True IsStored False False IsMaterialized Integer TimeInstant(86400) DataType 16060 Start 16424 End Size Header Data Refs Values 0 0 0 ... 1 1 ... StationId Refs Values ObsDate 0 1 2 ... 0 1 ... Temperature Values m1 Values m2 Values Loc5m Elevation Values 1124 1136 1206 ... 1645 1705 ... (51320501, 479952838) (51320501, 479952838) (51320501, 479952838) ... (53881597, 475086604) (53881597, 475086604) ... (102641, 959906) (102641, 959906) (102641, 959906) ... (107763, 950173) (107763, 950173) ... 1388891280 1388891280 1388891280 ... 1001420550 1001420550 ... 40271 40271 40271 ... 332839 332839 ... Name: TemperatureElevationAtStation IsStored: False StorageName: StorageName StationId ObsDate Name {Temperature} {m1, Loc} {m2} {Elevation} FixedPrecision(5,2) Point2D(9, 0.01) Point2D(7,5) FixedPrecision(7,3) DataType Mappings Domain {StationId, ObsDate} {StationId} {StationId} {StationId} Expression Observation.Temperature(StationId, ObsDate) Station.Loc(StationId) Cast(Station.Loc(StationId), "Point2D(7,5)") Topo.Elevation(m2) 80 365 DataChannel PostGIS PostGIS Figure 5.6: Example of in-memory MappingSet structures. The structure of each Constant will record the following attributes: name, storage name, IsStored (whether the Constant has been stored), data value and the expression evaluated to generate the data value (used to avoid duplicate computations). These in-memory structures for the representation of Dimensions,Extensional Mapping- Sets and Constants during the execution of operations are implemented in Java, making use of the DataFrame structure of Spark to record data columns. Fig. 5.7 depicts an UML diagram of the Dimension,Extensional MappingSet and Constant structures designed for the implementation. Each Constant resulting from some MAPAL operation is represented with an object of class Constant, Fig. 5.7. Notice that, together with relevant names, this class has also attributes to represent both its data value (of class MapalValue) and the expression used to compute it. As explained in Section 5.2.1, class MapalValue is the root of a hierarchy of classes that enable the representation of all the system data types. Each subclass of the hierarchy provides an appropriate data structure for the data value and a collection of methods to implement required
130 Chapter 5. MAPAL Implementation Constant -name: CStringValue -isStored: BooleanValue -storageName: CStringValue -expression: CStringValue -data: MapalValue +Literal(name: CStringValue, expression: CStringValue) +Import(name: CStringValue, storageName: CStringValue, channelName: CStringValue) +Scan(name: CStringValue) +EvaluateIntensionalMapping(name: CStringValue, mapping: CStringValue, arguments: Constant[1..*]) +EvaluateExtensionalMapping(name: CStringValue, mapping: CStringValue, arguments: Constant[1..*]) +EvaluateAggregateMapping(name: CStringValue, orderBy: OrderBySpecification, c: CStringValue, agg: AggMappingCall, ms: MappingSet) +Store(storageName: CStringValue) +Export(storageName: CStringValue, channelName: CStringValue) Dimension -data: DataFrame +Import(name: CStringValue, storageName: CStringValue, channelName: CStringValue) +Scan(name: CStringValue) +SamplingDimension(name: CStringValue, start: Constant, end: Constant) +Union(dim: Dimension) +Intersection(dim: Dimension) +Project(dimensionName: CStringValue, booleanMappingName: CStringValue, ms: MappingSet) +Store(storageName: CStringValue) +Export(storageName: CStringValue, channelName: CStringValue) DimensionHeader -name: CStringValue -storageName: CStringValue -dataChannel: CStringValue -size: IntegerValue -dataType: DataTypeMetadata -isSampling: BooleanValue -isStored: BooleanValue -isMaterialized: BooleanValue SamplingDimensionHeader -start: MapalValue -end: MapalValue MappingSet -data: DataFrame +Import(name: CStringValue, storageName: CStringValue, channelName: CStringValue, domain: CStringValue[1..*]) +Product(name: CStringValue, dims: Dimension[1..*]) +Product(dim: Dimension) +Project(name: CStringValue, components: CStringValue[1..*]) +EvaluateIntensionalMappings(mappingNames: CStringValue[1..*], mappings: IntMappingCall[1..*]) +EvaluateExtensionalMapping(mappingName: CStringValue, mapping: ExtMappingCall) +EvaluateConstant(mappingName: CStringValue, k: Constant) +EvaluateAggregateMappings(name: CStringValue, groupBy: CStringValue[1..*], orderBy: OrderBySpecification, c: CStringValue, mappingNames: CStringValue[1..*], aggs: AggMappingCall[1..*]) +Store(storageName: CStringValue) +Export(storageName: CStringValue, channelName: CStringValue) MappingSetHeader -name: CStringValue -storageName: CStringValue -isStored: BooleanValue -dimensions1..* MappingHeader -name: CStringValue[1..*] -dataType: DataTypeMetadata -domain: CStringValue[1..*] -expression: CStringValue -mappings1..* ReferencedDimensions -ReferencedBy: CStringValue -refDimensions: CStringValue[1..*] -refDims 0..* AggMappingCall -aggMapping: CStringValue -arguments: CStringValue[1..*] ExtMappingCall -extMapping: CStringValue -arguments: CStringValue[1..*] IntMappingCall -intMapping: CStringValue -arguments: CStringValue[1..*] OrderBySpecification -component: CStringValue[1..*] -direction: OrderingDirection[1..*] «enumeration» OrderingDirection Ascending Descending DataTypeMetadata -dataTypeName: MapalDataType -precision: IntegerValue +scale: IntegerValue -resolution: RealValue «enumeration» MapalDataType Tuple DataValue BooleanValue CStringValue NumericValue IntegerValue RealValue FixedPrecisionValue SamplingValue TemporalValue TimeValue TimeInstantValue DateValue Point1DValue Point2DValue Geometry2DValue Figure 5.7: In-memory structures implementation with Spark.
5.3. Data Structures Implementation 131 primitive mappings and operators. Each Constant operation described in Subsection 4.3.3 has a relevant method in class Constant. Each Dimension generated by some operation is represented in memory by an instance of class Dimension, Fig. 5.7. The header of class Dimension is an instance of class Dimension- Header. Notice that class SamplingDimensionHeader is provided for the proper representation of SamplingDimensions.Dimension data is stored in a Spark DataFrame, which may have one or two columns depending on whether the Dimension is stored or/and materialized, as shown in the example of Fig. 5.5. Each Dimension operator described in Subsection 4.3.3 is implemented with a relevant method in class Dimension. Each Extensional MappingSet generated by some operation is stored in memory by an instance of class MappingSet. The header is an instance of type MappingSetHeader, which includes one DimensionHeader per Dimension of the Extensional MappingSet and one MappingHeader per Extensional Mapping. Information related to which Extensional Mapping or Dimension references values of stored Dimensions is provided by instances of class ReferencedDimensions. All the data columns of an Extensional MappingSet are represented in a single DataFrame. Such DataFrame includes columns for Dimension references and Dimension values, for Extensional Mapping values and for Dimensions referenced by either Extensional Mappings or Dimensions. Each Extensional MappingSet operation described in Subsection 4.3.3 is implemented by a relevant method in class MappingSet. Additional classes have been defined to ease the proper representation of several class attributes and method arguments. Thus, attribute dataType in DimensionHeader and MappingHeader stores information of recorded data type in a DataTypeMetadata object, which in turn stores the data type name in a MapalDataType object. Optional argument orderBy, in methods for calculating aggregate mappings both in Constant and MappingSet classes, stores the ordering specification in an OrderBySpecification object, which in turn stores the ordering direction in an OrderingDirection object. Classes IntMappingCall,ExtMappingCall and AggMappingCall enable the representation of expressions of the form mapping(a1,··· ,an)used in the evaluation of intensional,extensional and aggregate mappings, respectively. Spark class DataFrame provides methods to properly depict the DataFrame schema (printSchema()) as well as the recorded DataFrame data (show()). For illustration purposes, Fig. 5.8 depicts the output of such methods in the stored DataFrame of Extensional MappingSet Municipality. Notice that Dimensions and Extensional Mappings values are recorded