Full text
A Methodology and a Software Tool for Sensor Data Validation/Reconstruction: Application to the Catalonia Regional Water Network Miquel ` A. Cuguer´ o-Escofeta, Diego Garc´ ıaa, Joseba Quevedoa, Vicenc¸ Puiga, Santiago Espinb, Jaume Roquetb aSupervision, Safety and Automatic Control Research Center (CS2AC), Polytechnic University of Catalonia (UPC), Terrassa Campus, Gaia Research Bldg. Rambla Sant Nebridi, 22. 08222 Terrassa, Barcelona, Spain (e-mail: {miquel.angel.cuguero,diego.garcia,joseba.quevedo,vicenc.puig}@upc.edu). bATLL Concession`aria de la Generalitat de Catalunya, SA. Sant Mart´ı de l’Erm, 30. 08970 Sant Joan Desp´ı, Barcelona, Spain (e-mail: {sespin, jroquet}@atll.cat). Abstract In this paper, a sensor data validation/reconstruction methodology applicable to water networks and its implementation by means of a software tool, are presented. The aim is to guarantee that the sensor data are reliable and complete in case that sensor faults occur. The availability of such dataset is of paramount importance in order to successfully use the sensor data for further tasks e.g. water billing, network efficiency assessment, leak localisation and real-time operational control. The methodology presented here is based on a sequence of tests and on the combined use of spatial models (SM) and time series models (TSM) applied to the sensors used for real-time monitoring and control of the water network. Spatial models take advantage of the physical relations between different system variables (e.g. flow and level sensors in hydraulic systems) while time series models take advantage of the temporal redundancy of the measured variables (here by means of a Holt-Winters (HW) time series model). First, the data validation approach, based on several tests of different complexity, is described to detect potential invalid or missing data. Then, the reconstruction process is based on a set of spatial and time series models used to reconstruct the missing/invalid data with the model estimation providing the best fit. A software tool implementing the proposed data validation and reconstruction methodology is also described. Finally, results obtained applying the proposed methodology to a real case study based on the Catalonia regional water network is used to illustrate its performance. Keywords: Sensor Data Validation/Reconstruction, Fault Isolation, Model-Based Fault Diagnosis, Time Series 1. Introduction Critical Infrastructure Systems (CIS), including water, gas or electricity networks, are complex large-scale systems geographically distributed and decentralized with a hierarchical structure. These systems require highly sophisticated supervisory and real-time control schemes, to ensure high performance achievement and maintenance when conditions are non-favourable [1, 2] due to e.g. sensor and actuator malfunctions (faults). Regarding the measurements in water systems, the commonly measured hydraulic and quality variables include water flow rate in links, pressure in nodes, water level in tanks, pH, conductivity and turbidity, as well as disinfectant and pollutant concentrations. For each measurement obtained from a sensor, the data (signals) are usually represented in the form of one dimensional time 1
2 series. Each sensor measures a physical quantity and converts it into a signal that can be read by the appropriate instrumentation. Then, the measuring system converts the sensor signals into values aiming to represent a certain real physical quantity. These values, known as raw data, need to be validated before further use in order to assure the reliability of the results derived from their usage. In systems like CIS, a telecontrol system is acquiring, storing and validating data gathered from different kind of sensors every given sampling time (e.g. every few minutes) to accurately real-time monitor the whole system. In the data acquisition process, several problems can occur, as those related with the communication system (e.g. between sensors and data loggers or in the telecontrol system itself) or outliers, producing lost or corrupted data which may be of great concern in order to have valid historic records. When this is occurring, lost data should be replaced by a set of estimated data which should be representative of the data lost, since missing data may severely jeopardise further processes needing complete datasets in order to get meaningful conclusions/analysis. Another common problem in CIS is caused by the unreliable sensors, which may be affected by faults e.g. offset, drift, freezing in the measurements [3, 4, 5]. These unreliable data should also be detected and replaced by forecasted data, since it may be used for system management tasks e.g. maintenance, planning, investment plans, billing, security and operational control [6] and system fault detection and isolation (Figure 1). In the case of water network applications, this system fault diagnosis may include e.g. network leaks isolation, as considered in [7, 8, 9]. However, the methodology presented here may well be applied to different applications involving a telemeasured sensor network, such as smart buildings or environmental systems (see, e.g. [10, 11]). In addition to the possible measurement deviations related to the sensor performance itself, the errors can also occur due to heterogeneous reasons, e.g. sensor installation, calibration or electrical problems. Thus, it is important to provide the data system with procedures that can detect these problems and assist the user in the monitoring and the processing of the incoming data. The data validation is an essential step to improve data reliability. Sensor data validation and reconciliation have been intensively addressed using least-squares approaches including Kalman filters (see, e.g. [12, 13]). These have been also used for data forecasting, as pointed out in [14], where a review of techniques for prediction of consumption in water and natural gas grids is presented. The basic idea of Kalman filter based methods in data validation and reconciliation is to allow gross error detection and to provide reconstructed data that is consistent with model/balance equations describing the system operation. The approach presented in this paper aims to assess the validity of each single sensor measurement by means of a set of tests exploiting not only the model equations (spatial redundancy) but also temporal redundancy, using time series models and a bank of low-level tests (non-model based) aiming to label the data with a certain quality index. Traditionally, data validation process has been developed by manual data analysis, performed by experienced users with the only assistance of basic data analysis and visualisation tools [15], which significantly limits the amount of data to be validated [16] and the abnormal situations which may be correctly detected [17]. However, the volume of real data acquired in CIS is dramatically increasing due to the increment of automated measurement systems allowing their monitoring [18]. Also, real-time operation, paramount in many real applications, makes human data validation even harder to pursuit. In order to cope with this situation and increase the reliability of the data diagnosis system, automatic data validation tools have arised e.g. NIKLAS for real and non-real time diagnosis of meteorological data [19]. Also, in [20] a data validation module is considered in the framework of an on-line water quality fault tolerant control system. Over the last 15 years, more and more affordable on-line sensors have become available, leading to ever increasing acceptance of on-line water monitoring [21]. These on-line systems allow to deploy control mechanisms that are optimized for and respond to the actual process conditions. However, on-line systems require a data validation method that is applicable to real-time incoming data. The major difference between on-line and off-line data validation lies in the available information and the required execution time. In contrast with the on-line execution, the off-line operation has no time restrictions because real-time constraints do not apply, and regarding the information, the whole set of data is available. Moreover, on-line data validation is usually required by a further real-time control system and thus the data are used for decision support (or decision making) just after being obtained. Consequently, the on-line data validation process should have low execution time, whereas the off-line data validation does not have this requirement. According to the nature of the available knowledge, different kinds of data validation approaches may be considered, with varying degrees of sophistication. In general, one may distinguish between elementary signal-based (“low-level”) methods and model-based (“high-level”) methods (see, e.g. [6]). Elementary signal based methods use simple heuristics and limited statistical information of a given sensor [22]. Typically, these methods are based on validating either signal values or signal variations. On the one hand, in the signal value-based approach, data are assessed as valid or invalid according to two different thresholds (high and low) so data are assumed to be invalid when lying 2
3 outside these threshold values. On the other hand, methods based on signal variations look for high variations (peaks in the curve) and low variations (flat curve) in the signals. Model-based methods rely on the use of models to check the consistency of sensor data [21]. This consistency check is based on computing the difference between the predicted value from the model and the real value measured by the sensors. Then, this difference (known as residual) is compared with a threshold value (zero in the ideal case). When the residual is bigger than the corresponding threshold, a fault is assumed in the sensor; otherwise, the sensor is assumed to work properly. Moreover, the information of all the available residuals and models allows performing fault isolation in order to discover the faulty sensor. Models are usually derived using either multivariate procedures exploiting the correlation or analytical relations between several variables, sometimes measured at different times (“temporal redundancy”) and/or locations (“spatial redundancy”). In this paper, a methodology is developed for validation and reconstruction of sensors data in a water network, taking into account not only spatial models (SM) but also time series models (TSM) for each flow and level meter here. Also, internal models of every component in the local equipment units (e.g. pumps, valves, flows, levels) are considered. SM take advantage of the relation between different variables in the system (e.g. demand, pump flows and tank levels) while TSM take advantage of the temporal redundancy of the measured variables, by means of Holt-Winters (HW) time series models [23]. Moreover, after the corrupted sensor data are detected, they must be replaced by adequate estimated data using the available temporal/spatial redundancy. The methodology is mainly applied to flow and level meters, since it exploits the temporal redundancy of flow and level data in a water network. In this paper, an operative software tool implementing the presented methodology which is able to properly handle raw sensor data (including storage, querying and visualization) is also presented. The proposed approach and tool are applied to several subsystems in the Catalonia regional water network (Figure 8) using raw data collected from ATLL Concession`aria de la Generalitat de Catalunya, SA (ATLL), the company managing this water network. The structure of the paper is as follows: In Sections 2 and 3, the methodology to validate/reconstruct the sensor data, in order to provide a reliable dataset when faulty situations occur within the sensor set, is proposed. In Section 4, the software tool implementing the proposed methodology is introduced. In Section 5, the application case study is presented, based on the Catalonia regional water network. The sensors in this network measure several real magnitudes of interest such demand and input tank flows and levels, considering real-world scenarios. Also, the corresponding results obtained applying the proposed methodology are detailed in Section 5. Finally, conclusions of this work are outlined in Section 6. 2. Proposed Methodology 2.1. Description In real water networks such as the one considered here, there is usually a telemeasurement system acquiring, recording and validating data gathered from different kind of sensors at each sample time to accurately real-time monitor the whole network [6]. As discussed in the introduction, in this data acquisition process, problems in the communication system (e.g. between sensors and data loggers) or in the telemeasurement system itself (e.g. sensors may be affected by e.g. offset, drift or freezing faults), often arise and produce data loss, which may be of great concern in order to have valid historic records. These unreliable data should be detected and replaced by estimated data before they can be used for system management tasks such as maintenance, planning, billing and operational control, as depicted in the procedure in Figure 1. The input to this procedure is the raw data yraw gathered from the sensors. The process is divided in two different stages: the first stage is related with the data validation, while the second stage addresses the reconstruction of invalid/missing data, before the data are stored in an operational database (DB) for further use. At the first stage (data validation, detailed in Figure 2), if the datum yraw(k) at a certain sample time kis validated, flag vis set to 1 and datum yval(k)=yraw(k) is stored in the aforementioned operational DB as validated data. Conversely, if the datum yraw(k) is invalidated, flag vis set to 0 and the datum reconstruction process (second stage) is started, in order to provide a reconstructed estimation yrec(k) of the invalid/missing data yraw(k) to be stored in the DB. The whole procedure is further detailed in Algorithm 1 for the data validation stage and in Algorithm 2 for the data reconstruction stage. Here, communication and sensor faults are considered as faults affecting the telemeasurement system and the sensors, respectively, and the data detection/reconstruction procedure is used as a prefilter to estimate the invalid/missing data when these type of faults are occurring. As discussed in the introduction, different types of data detection methods with distinct degrees of complexity may be considered according to the available system knowledge. This is the approach that the proposed methodology here 3
4 Data Validation Validated? Data Reconstruction No Operational DataBase Yes Raw sensor data ( ) raw y k ( ) 1v k = ( ) 0v k = ( ) val y k ( ) rec y k Figure 1. Raw data validation/reconstruction procedure will follow. Generally, two types of methods are considered, one for elementary ‘low-level’ signal-based methods and another for ‘high-level’ model-based methods. The first type uses simple heuristics and limited statistical information from the sensors [22] [15] and is typically based on checking either signal values or variations, whilst the second type uses models for consistency-checking of the sensor data [21]. 2.2. Validation Tests The data detection process presented is inspired by the Spanish AENOR-UNE norm 500540 developed for data validation in meteorological stations [6]. The methodology presented here applies a set of consecutive detection tests to a given dataset (Figure 2), to finally assign a certain quality level qdepending on the number of tests passed. Also, the corresponding tests passed are characterized by a validation vector l, as shown in Figure 2. If the datum yraw(k) at a certain sample time kis voided at any validation level, flag vis set to 0 and the datum reconstruction process (second stage) is started. Conversely, if the datum yraw(k) pass all the validation levels, flag vis set to 1 and the data are validated (i.e. yval(k)=yraw(k)). In the latter situation i.e. validated datum yraw(k), q(k)=6 and l(k)=[1,1,1,1,1,1]. The validation tests include a set of ’low-level’ tests (Levels 0 to 3, included) which check elementary signal properties, and a set of ’high-level’ tests (Level 4 and Level 5), which rely on the use of models to check the consistency of the sensor data. The latter models are also used in the reconstruction stage of the potentially invalidated data, as explained in Section 3.3. As introduced in the last paragraph, if any of the validation tests in Figure 2 is not satisfied, v=0. The validation procedure is also detailed in Algorithm 1. Level 0 Communications Level 1 Physical range Raw sensor data Level 2 Trend Level 3 Equipment State Level 4 Spatial Consistency Level 5 Time Series Consistency ( ) ( ) 0 ( ) 1 raw y k q k v k = = ( )v k Passed? Yes No 0( ) 1 () ( ) 1 l k q k q k = = + 0( ) 0 ( ) 0 l k v k = = Passed? Yes No 1( ) 1 () ( ) 1 l k q k q k = = + 1( ) 0 ( ) 0 l k v k = = Passed? Yes No 2( ) 1 () ( ) 1 l k q k q k = = + 2( ) 0 ( ) 0 l k v k = = Passed? Yes No 3( ) 0 ( ) 0 l k v k = = Passed? No 4( ) 0 ( ) 0 l k v k = = Yes 4 ( ) ( ) 1 ( ) 1 q k q k l k = + = Passed? No 5( ) 0 ( ) 0 l k v k = = Yes 5 ( ) ( ) 1 ( ) 1 q k q k l k = + = ( )kl ( )q k 3( ) 1 ( ) ( ) 1 l k q k q k = = + Figure 2. Data Validation Tests An explanation of the test applied to each level is given next: 4
5 •Level 0: This level, also called communications level, checks whether the data are properly recorded at a regular sample rate by the acquisition system. If this is not fulfilled, there is some communication problem involving e.g. the data transmission from the ground sensors to the operational database. Hence, this level allows detecting problems in the data acquisition or communication system, which is one of the most common faults affecting telemeasurement systems, as the one considered here. •Level 1: This level, also called physical range level, checks whether the data are within the physical range of the sensor acquiring the corresponding measurement. The expected range of the measurements may be obtained from e.g. sensor specifications or historical records of the data. •Level 2: This level, also called trend level, checks whether the data derivative i.e. the magnitude change of the data among consecutive sample times, are within their expected rate. This allows detecting unexpected and possibly undesired sudden changes in the data, e.g. in a water network, tank water level sensors measurements cannot change more than several centimeters per minute. •Level 3: This level, also called equipment state level, allows to check the consistency of the variables in a given equipment unit i.e. sensor or actuator. For example, in a water network system, in a pipe with a valve and a flow meter installed, there is a relation between the valve state and the flow meter reading. •Level 4: This level, also called spatial consistency level, checks the consistency of the data collected by a certain sensor with its SM [24], i.e. the correlation between data coming from spatially-related sensors. This SM is obtained from the physical relations among these variables. In hydraulic systems, this relation is generally obtained from the mass balance model of the element relating the different measured variables involved. •Level 5: This level, also called time series consistency level, checks for temporal consistency of a given sensor measurement, by means of a TSM obtained from sensor historical records under faultless assumption ([6]). A common method for time series signal forecasting is the HW approach ([23], [25]) because of its simplicity and low computational and storage requirements. In contrast to spatial consistency level, time series consistency level only uses information of the considered sensor without needing additional information (e.g. network topology or extra measurements from the system) to perform the validation, which makes it convenient when there is no such additional information available or the sensors needed by the corresponding spatial consistency level are unreliable. At this level, the analysis of the historic measurement records of a certain sensor are used to obtain the corresponding HW TSM sensor model and to validate the current data acquired by this element. 3. Model-based Data Validation/Reconstruction Levels 3.1. Models for Data Validation/Reconstruction Model-based data validation/reconstruction relies on using models that exploit the temporal or spatial redundancy existing among the sensors. On the one hand, SM takes advantage of the relation between different variables physically related within the system. In water networks, this relation is generally obtained from the mass balance relating the different measured variables involved in a particular hydraulic element. For example, in a water tank (see Figure 3) the corresponding SM level estimation may be stated ˆxS M(k)=x(k−1) +∆t A(qin(k−1) −qout(k−1)),(1) where ˆxS M is the spatial model tank level estimation, xis the measured tank level, qin is the incoming tank flow, qout is the outcoming tank flow and ∆tis the sampling time, respectively. Estimation of other variables (e.g. ˆqin, ˆqout) may be obtained in a similar manner. Real elements include uncertainty (e.g. due to noise, inaccuracy of the model, etc.) which may lead to the nonsatisfaction of the mass balance in the element considered. Hence, consistency of the data collected by a certain sensor with its SM [24] (i.e. the correlation between data coming from spatially-related sensors) should take this uncertainty into account. 5
6 flow level flow in q out q x Figure 3. Single tank system schematics with single input and single demand Alternatively, TSM takes advantage of the temporal redundancy of the measured variables. A wide used method for time series modelling because of its simplicity, low computational and storage requirements and ease of automation, is the HW approach [25]. This method, which was originally created for sales demand forecasting, has been used in a broad range of applications since its appearance. Exponential methods are first introduced in [26], where decreasing series of exponential weights are used. In [25], the former method is extended to include trend and seasonality terms. In [27, 28] multiple (i.e. double and triple) seasonality is explored, expanding the initial single seasonality expression of the former HW method, designed to cope with the sales demands monthly variations across a year period. Further alternative approaches to exponential smoothing forecasting may be found in [29] and [30]. Some issues of interest regarding its performance include the effect of the outliers in the forecasting, the consideration of the aforementioned different seasonal periods which may characterise the corresponding time series data sequence to be modelled (e.g. sales demands, water demands) or the consideration of prediction intervals which may provide reliability to the forecast. Regarding outliers, which may be produced by unexpected component behaviors (e.g. sensor malfunctions) these may degrade the performance of the HW method if not accommodated. In [31], this problem is considered and a robust version of the HW method against the outliers is presented, by recursively filtering their effect in the data main stream and applying the standard HW approach to the obtained filtered data. The latter approach is also considered here to provide robustness against the outliers. Moreover, there are different versions of the HW method e.g. additive or damped trend, additive or multiplicative seasonality, single or multiple seasonality [23]. Here, good performance has been attained with the additive single seasonality version, which estimated value is obtained for a forecasting horizon ` ˆxTS M(k)=¯ R(k−`)+`¯ G(k−`)+¯ S(k−L),(2) where ¯ Ris the level estimation removing seasonality, ¯ R(k−`)=αx(k−`)−¯ S(k−`−L) +(1−α)¯ R(k−`−1) +¯ G(k−`−1), (3) ¯ Gis the trend estimation, ¯ G(k−`)=β¯ R(k−`)−¯ R(k−`−1) +(1−β)¯ G(k−`−1), (4) 6
7 S¯is the seasonal component estimation, ¯ S(k−`)=γx(k−`)−¯ R(k−`) +(1−γ)¯ S(k−`−L), (5) and Lis the season (daily) periodicity, α,βand γare the HW parameters (level, trend and season smoothing factors, respectively), xis the measured value and ˆxTS M is the TSM estimated value. The parameters α,βand γare in the interval [0,1] and can be estimated from historical data using the least-squares approach. Hence, analysing the historic records of a certain sensor, a HW TSM model can be obtained and used to estimate missing data of this element when a fault is affecting its readings. 3.2. Data Validation On the one hand, the test check for the so-called ’low-level’ tests are straightforward, since they rely on basic signal-based heuristics. On the other hand, the ’high-level’ model-based tests rely on checking for consistency by means of the residuals ri(k), obtained from the difference between the system measurements and the corresponding SM or TSM estimations, expressed in input-output regressor form ri(k)=xi(k)−ˆxi(k)=xi(k)−φT i(k)θi,(6) where θiare the nominal parameters obtained using a training dataset, xiis the sensor measurement, ˆxiis the model prediction and φi(k) is the regressor vector of dimensions nθi×1 including inputs (ui(k),ui(k−1),ui(k−2), ...) and outputs (yi(k),yi(k−1),yi(k−2), ...). The particular models used to compute the prediction ˆxiat instant kdepend on the validation level considered (i.e. model-based level 4 or 5 in Figure 2, respectively), and are introduced in Section 3.3. Considering the uncertainty (e.g. modelling errors, noise), the detection test involves checking the condition |ri(k)|< τi,(7) where τiis the detection threshold. The detection threshold can be determined using statistical methods [32] or setmembership approaches [33]. In the case of statistical methods, the noise is assumed to follow a normal distribution with known mean value µiand standard deviation σi[34]. Then, the threshold of the i-th residual can be determined as follows: τi=µi+3σi, including the 99.7 % of the values of a normal distribution according to the 3-sigma rule. Alternatively, when using a set-membership approach the noise is assumed to be unknown but bounded, with a priori known bound. Then, the threshold can be obtained by propagating the uncertainty to the residual computation [33]. Using either one or the other approach, the threshold in (7) is determined to include the values of the whole residual distribution in the faultless situation and hence, it may be used for fault detection purposes. This threshold is also useful to provide prediction interval bounds for the data forecasting process, so test condition (7) can be equivalently expressed as follows xi(k)∈[ ˆxi(k),¯ ˆxi(k)],(8) where ¯ ˆxi(k)=ˆxi(k)+τiand ˆxi(k)=ˆxi(k)−τi, respectively. Condition (8) applies both to SM (1) and TSM (2) models. These interval bounds (8) consider the corresponding model behavior under faultless conditions including the uncertainty effect, as introduced in the residual bound condition (7). Hence, these bounds could alternatively be used in the data validation process, in order to decide whether a data sample at time instant kis reliable. The whole data validation process is detailed in Algorithm 1. 7
8 3.3. Data Reconstruction As introduced in Section 2, when a fault is detected at the validation stage and the corresponding data are voided, a reconstruction process is started until the sensor data are validated again. The output of the data validation process (Figure 1) is used to identify the invalidated data that should be reconstructed. SM, related with Level 4 in Figure 2, and TSM, related with Level 5 in Figure 2, are used for this purpose, depending on the performance of each model. This data reconstruction process is detailed in Algorithm 2. The performance of each model is measured by the Mean Squared Error (MSE), evaluated in a moving horizon window MS E(k)=1 m k X j=k−m e(j)2,(9) where mis the number of data samples considered in the window, e(j)=x(j)−ˆx(j) is the error at instant j,x(j) is the measured value at instant j, ˆx(j) is the estimated value by the model (SM or TSM, respectively) at instant jand kis the actual time instant. The model having best MSE index before the fault occurs (i.e. when the data validation process is not satisfactory) is used to produce the reconstructed sensor signal. In order to produce the forecasted signal, it is desirable to use measured data instead of estimated data, to avoid model uncertainty effects in the forecasted value. This calls for the computation of ˆxi(k)|`using (2) with `,1 when possible and the usage of the different models obtained in a gain-scheduling fashion when e.g. the data are invalidated for more than a single time instant. HW TSM models may obtain forecasted values for different prediction horizons ` by design, if forecasted value at time kin (2) is rewritten as follows ˆxTS M(k)|`=¯ R(k)+`¯ G(k)+¯ S(k−L+`),(10) Then, measured values may be used to produce the TSM forecasted signal within a complete season (day) L without using old forecasted values. In order to achieve this, the complete set of models (i.e. the models for each step within the complete season L) must be obtained at the calibration stage under faultless assumption, i.e. a HW TSM model [α`,β`,γ`] may be obtained for `=1,· · · ,L. Similar procedure may be used in the same case study for alternative applications not related with data validation/reconstruction, as e.g. consumer demand prediction, in order to forecast water network user behavior beforehand. HW TSM models are specially suited to this end, since they were created in order to predict market product sales evolution according to consumer periodical behaviors [25], and user water consumption in district metered areas have a similar behavior. 4. Software Framework The architecture of the software framework implemented is depicted in Figure 4. There are two main components: the Data Management Web application and the Validation and Reconstruction tool1. On the one hand, the Data Management component is a web application focused on collecting and serving time series data, i.e. observations coming from any kind of sensor. It allows authorized users to upload new data, download historical data and visualize data from anywhere using a device with a browser and Internet connection. Thus, this web-based data repository is highly available and provides a solution to the data-driven users to keep centralized data from different projects and sources. It also avoids typical datasets-usage related drawbacks e.g. data loss, sparse and duplicated data locations and emails with large datasets between project members. On the other hand, the Validation and Reconstruction component allows users to apply the methodologies described in Sections 2 and 3 on data provided by the Data Management web application. 1Both software tools are proprietary software. 8
9 Figure 4. Software architecture diagram 4.1. Data Management Web Application This module provides a user-friendly tool allowing to import and export data so that stored data are available to registered users with read permissions on the dataset. This point is important in order to respect existing confidential agreements: a user must have explicit permission on a dataset to be able to access or visualize it. Only the dataset’s owner and the administrator can grant read permissions. People working with data usually need to collect and prepare the raw data, e.g. remove outliers and fill missing data, before being able to apply further analysis e.g. statistical, exploratory or even to focus on the real objective of working with the corresponding data. These sort of tasks are time-consuming: there are many situations when the first two steps introduced take the 80 % of the whole data treatment process time. Thus, this tool provides three services in order to focus the efforts on the data themselves and not on how to collect, obtain and prepare them. These services are the following: the data import service, the data export service and the visualization service. The data import service handles the data ingestion from different file formats (e.g. CSV, Excel, Access). An import wizard allows the user to specify the input data format, allowing the data to be loaded into the database after being specified. The data export service handles the data extraction. The user can specify the time period to export and the output file format. The current version of the tool allows to download data in CSV, Excel and SAC format2. Finally, the data visualization service provides a tool to visually explore the collected data. Hence, the user can plot multiple signals (e.g. time series) to do some exploratory analysis before downloading and to select only the relevant data. The visualization tool allows zooming and panning the time series. This web application is implemented in two layers, a back-end (server layer) that handles the data storage and access with an underlaying data model, and a front-end (visual layer) to provide a friendly web-based user interface to interact with the three services described before. The back-end is developed with the Django3web framework connected to a database based on PostgreSQL. The front-end is implemented in HTML and JavaScript (see Figure 5). The Import and Export modules handle the operations of saving and querying data against the PostgreSQL Database server. 4.2. Validation and Reconstruction Matlab Tool The Validation and Reconstruction methodologies, detailed in Sections 2 and 3, are summed up in Algorithm 1 and Algorithm 2, respectively. These methodologies are implemented in a software tool developed in Matlab. Matlab 2SAC format is a binary Mat-file containing a defined data structure. 3Django is a free open source web framework. Its primary goal is to facilitate the creation of complex, database-driven websites. 9
16 01/01 01/02 01/03 01/04 01/05 0 10 20 30 40 50 60 70 80 m³/h. E6FT00102_CI 01/01 01/02 01/03 01/04 01/05 0 200 400 600 MSE MSE 01/01 01/02 01/03 01/04 01/05 Valid Invalid Validation measured estimated reconstructed value time series time series test limits test derivative time series Figure 10. Results of the validation and reconstruction methodology on the flow meter E6FT00102 CI 16
17 820 840 860 880 900 920 940 960 980 1000 1020 −400 −200 0 200 400 600 800 time [h] [m³/h.] Sensor D6FT00204_CI Original data Measured data Spatial model TS model Estimated data 820 840 860 880 900 920 940 960 980 1000 1020 0 0.5 1 1.5 2x 104 time [h] MSE TS model Spatial model Figure 11. Results of the validation and reconstruction methodology, flow meter D6FT00204 CI case due to the lack of information. Again, the only available model for reconstruction is the TSM (similarly as in the scenario in Figure 10) which is finally used for the missing data reconstruction in the scenario considered in Figure 12, based on the limited available information in this particular case. Finally, two different scenarios involving the flow meter E6FT00502 CI are considered. On the one hand, in the scenario in Figure 14, a communication fault affecting the corresponding flow meter is presented, which does not transmit data in one day period (from t=1024 h to t=1048 h). In this particular case, the communication fault only affects the latter sensor, hence the corresponding SM is available because the spatially related sensors (e.g. D6FT00201 CI) are not affected by this fault. In this scenario, the SM model is used for missing data reconstruction, since it performs better than the corresponding HW TSM model (bottom subplot in Figure 14). It may be noted that the use of the SM assumes that the model input sensor measurement is faultless when the SM is used for data reconstruction. This may be assured since the input model integrity is checked by the methodology presented here at its corresponding stage and, if not verified, the validation test at this stage is not fulfilled. On the other hand, an offset fault of magnitude 25 % full scale affecting the flow meter E6FT00502 CI, also common in this kind of sensors, is presented in Figure 15, lasting for three days (from t=1024 h to t=1096 h). As in the previous scenario, the SM model performs better than HW TSM before the fault is produced (see Figure 15 bottom subplot) and hence it is used for invalid data estimation. In this particular scenario, it may be noted how the measured signal is out of the SM threshold boundaries (red dotted line) for the whole fault scenario, whilst it remains bounded by the HW TSM threshold boundaries (blue dotted line) for part of the first day after the fault is produced. This behavior is due to the adaptation of the HW TSM to the input signal, as its estimation depends on the historic records of the measurements, as detailed in Section 3.1. Hence, it should be considered that, when used for data validation, the prognosis derived from the application of the time series consistency test will expire after a certain time after the fault is produced, when using measurement historic records. 6. Conclusions In this paper, a data validation and reconstruction methodology is introduced to overcome the sensor problems arising in CIS, such as water networks. The validation strategy is based on a set of data quality tests that allow to detect potentially erroneous data. Then, a reconstruction scheme is defined using SM and TSM to provide an estimation based on the model having the best fit, also providing prediction intervals for the forecasted reconstructed data. In addition, a software tool is described to provide a homogeneous and accessible database by a user-friendly interface, and to apply the methodology presented here. Finally, some results obtained using data from a real network 17
18 820 840 860 880 900 920 940 960 980 1000 1020 −600 −400 −200 0 200 400 600 800 1000 time [h] [m³/h.] Sensor E6FT00502_CI Original data Measured data Spatial model TS model Estimated data 820 840 860 880 900 920 940 960 980 1000 1020 0 1 2 3 4x 104 time [h] MSE TS model Spatial model Figure 12. Results of the validation and reconstruction methodology, flow meter E6FT00502 CI 820 840 860 880 900 920 940 960 980 1000 1020 −400 −200 0 200 400 600 800 time [h] [m³/h.] Sensor D6FT00201_CI Original data Measured data Spatial model TS model Estimated data 820 840 860 880 900 920 940 960 980 1000 1020 0 0.5 1 1.5 2x 104 time [h] MSE TS model Spatial model Figure 13. Results of the validation and reconstruction methodology, flow meter D6FT00201 CI 18
19 1000 1010 1020 1030 1040 1050 −600 −400 −200 0 200 400 600 800 time [h] [m³/h.] Sensor E6FT00502_CI Original data Measured data Spatial model TS model Estimated data 1000 1010 1020 1030 1040 1050 0 1 2 3x 104 time [h] MSE TS model Spatial model Figure 14. Results of the validation and reconstruction methodology on the flow meter E6FT00502 CI 960 980 1000 1020 1040 1060 1080 1100 1120 1140 1160 −600 −400 −200 0 200 400 600 800 1000 1200 time [h] [m³/h.] Sensor E6FT00502_CI Original data Measured data Spatial model TS model Estimated data 960 980 1000 1020 1040 1060 1080 1100 1120 1140 1160 0 1 2 3x 105 time [h] MSE TS model Spatial model Figure 15. Results of the validation and reconstruction methodology on the flow meter E6FT00502 CI 19
20 located in the Catalonia area are presented using the software described, showing the ability of the methodology to detect and reconstruct anomalous data. In future steps of this work, the proposed methodology and tool are going to be applied to the whole Catalonia Regional network, since in the latter network not only the hydraulic sensors considered here are monitored, but also e.g. the water quality sensors. Acknowledgement This work has been partially funded by the Spanish Ministry of Science and Technology through the Project ECOCIS (Ref. DPI2013-48243-C2-1-R) and Project HARCRICS (Ref. DPI2014-58104-R), and by EFFINET grant FP7-ICT-2012-318556 of the European Commission. References [1] M. Schtze, A. Campisano, H. Colas, W. Schilling, P. A. Vanrolleghem, Real time control of urban wastewater systems - where do we stand today?, Journal of Hydrology 299 (2004) 335–348. doi:10.1016/j.jhydrol.2004.08.010. [2] M. Marinaki, M.; Papageorgiou, Optimal Real-time Control of Sewer Networks, Springer, 2005. [3] V.K.Kanakoudis, D.K.Tolikas, The role of leaks and breaks in water networks: technical and economical solutions, Water Supply: Research and Technology-Aqua 50 (5) (2001) 301–311. [4] V.K.Kanakoudis, S. Tsitsifli, Water pipe network reliability assessment using the dac method, Desalination and Water Treatment 33 (1-3) (2011) 97–106. [5] S.Tsitsifli, V.K.Kanakoudis, I. Bakouros, Pipe networks risk assessment based on survival analysis, Water Resources Management 25 (14) (2011) 3729–3746. [6] J. Quevedo, V. Puig, G. Cembrano, J. Blanch, J. Aguilar, D. Saporta, G. Benito, M. Hedo, A. Molina, Validation and reconstruction of flow meter data in the Barcelona water distribution network, Control Engineering Practice 18 (6) (2010) 640–651. doi:10.1016/j.conengprac.2010.03.003. [7] R. P´ erez, G. Sanz, V. Puig, J. Quevedo, M. A. Cuguer´ o-Escofet, F. Nejjari, J. Meseguer, G. Cembrano, J. M. M. Tur, R. Sarrate, Leak Localization in Water Networks. A Model-Based Methodology Using Pressure Sensors Applied to a Real Network in Barcelona, IEEE Control Systems Magazine 34 (2014) 24–36. doi:10.1109/MCS.2014.2320336. [8] J. Quevedo, M. A. Cuguer´ o, R. P´ erez, F. Nejjari, V. Puig, J. M. Mirats, Leakage location in water distribution networks based on correlation measurement of pressure sensors, in: 8th IWA Symposium on System Analysis and Integrated Assessment (WATERMATEX 2011), San Sebastian, 2011. [9] R. P´ erez, V. Puig, J. Pascual, J. Quevedo, E. Landeros, A. Peralta, Methodology for leakage isolation using pressure sensitivity analysis in water distribution networks, Control Engineering Practice 19 (10) (2011) 1157–1167. doi:10.1016/j.conengprac.2011.06.004. [10] M. A. Cuguer´ o-Escofet, M. Christodoulou, J. Quevedo, V. Puig-Cayuela, D. Garc´ ıa, M. Michaelides, Combining Contaminant Event Diagnosis with Data Validation /Reconstruction : Application to Smart Buildings, in: In Proceedings of IEEE 22nd Mediterranean Conference on Control and Automation (MED14), Palermo, Italy, 2014, pp. 293–298. doi:10.1109/MED.2014.6961386. [11] M. A. Cuguer´ o, J. Quevedo, V. Puig, D. Garc´ ıa, Inconsistent sensor data detection/correction: Application to environmental systems, in: Neural Networks (IJCNN), 2014 International Joint Conference on, 2014, pp. 84–90. [12] S. Narasimhan, C. Jordache, Data Reconciliation and Gross Error Detection: an Intelligent Use of Process Data, Gulf Professional Publishing, 1999. doi:10.1016/B978-088415255-2/50018-5. [13] J. A. Romagnoli, M. C. S´ anchez, Data Processing and Reconciliation for Chemical Process Operations, Vol. 2 of Process Systems Engineering, Elsevier, 1999. doi:10.1016/S1874-5970(00)80015-6. [14] M. Fagiani, S. Squartini, L. Gabrielli, S. Spinsante, F. Piazza, A Review of Datasets and Load Forecasting Techniques for Smart Natural Gas and Water Grids: Analysis and Experiments, Neurocomputing 170 (October) (2015) 448–465. doi:10.1016/j.neucom.2015.04.098. [15] H. Jorgensen, S. Rosenrn, H. Madsen, P. Mikkelsen, Quality control of rain data used for urban runoffsystems, Water Science and Technology 37 (11) (1998) 113–120. URL http://www.sciencedirect.com/science/article/pii/S0273122398003230 [16] M. Mourad, J.-L. Bertrand-Krajewski, A method for automatic validation of long time series of data in urban hydrology, Water Science & Technology 45 (4-5) (2002) 263–270. [17] V. Venkatasubramanian, R. Rengaswamy, K. Yin, S. N. Kavuri, A review of process fault detection and diagnosis: Part i: Quantitative model-based methods, Computers & Chemical Engineering 27 (3) (2003) 293–311. doi:10.1016/S0098-1354(02)00160-6. [18] N. Branisavljevi´ c, Z. Kapelan, D. Prodanovi´ c, Improved real-time data anomaly detection using context classification, Journal of Hydroinformatics 13 (2011) 307. doi:10.2166/hydro.2011.042. [19] G. Lempio, C. Podlasly, T. Einfalt, NIKLAS - Automatical quality control of time series data (2000) (2010) 1–6. [20] F. Edthofer, J. Van Den Broeke, J. Ettl, W. Lettl, a. Weingartner, Reliable online water quality monitoring as basis for fault tolerant control, Conference on Control and Fault-Tolerant Systems, SysTol’10 - Final Program and Book of Abstracts (2010) 57– 62doi:10.1109/SYSTOL.2010.5675985. [21] K. Tsang, Sensor data validation using gray models, ISA Transactions 42 (2003) 9–17. [22] D. Burnell, Auto-validation of district meter data, in: CCWI ’03 Advances in Water Supply Management, London, 2003. [23] S. Makridakis, S. Wheelwright, R. Hyndman, Forecasting methods and applications, John Wiley & Sons, 1998. [24] J. Quevedo, J. Blanch, V. Puig, J. Saludes, S. Espin, J. Roquet, Methodology of a data validation and reconstructions tool to improve the reliability of the water network supervision, in: International Conference of IWA Water Loss, Sao Paolo, Brazil, 2010. 20
21 [25] P. R. Winters, Forecasting sales by exponentially weighted moving averages, Management Science 6 (52) (1960) 324–342. [26] R. G. Brown, Statistical Forecasting for Inventory Control, New York: McGraw-Hil, 1959. [27] J. Taylor, Short-term electricity demand forecasting using double seasonal exponential smoothing, The Journal of the Operational Research Society 54 (8) (2003) 799–805. [28] J. W. Taylor, Triple seasonal methods for short-term electricity demand forecasting, European Journal of Operational Research 204 (1) (2010) 139–152. doi:10.1016/j.ejor.2009.10.003. URL http://linkinghub.elsevier.com/retrieve/pii/S037722170900705X [29] C. Pegels, Exponential forecasting: Some new variations, Management Science 15 (1969) 311–315. [30] E. S. Gardner Jr., Exponential Smoothing: The State of the Art - Part II, International Journal of Forecasting 22 (4) (2006) 637–666. URL http://www.sciencedirect.com/science/article/pii/S0169207006000392 [31] S. Gelper, R. Fried, C. Croux, Robust Forecasting with Exponential and Holt-Winters Smoothing, Journal of forecasting 29 (June 2009) (2010) 285–300. doi:10.1002/for. URL http://onlinelibrary.wiley.com/doi/10.1002/for.1125/abstract [32] M. Basseville, I. Nikiforov, Detection of Abrupt Changes: Theory and Application, Prentice-Hall, Inc., 1993. [33] V. Puig, Fault diagnosis and fault tolerant control using set-membership approaches: Application to real case studies, International Journal of Applied Mathematics and Computer Science 20 (4) (2010) 619–635. [34] S. X. Ding, Model-based Fault Diagnosis Techniques, Springer, 2008. 21