scieee AI-readable full text Open interactive document viewer

Mining Geographic Data for Fuel Consumption Estimation

Vítor Daniel Ferreira da Cunha Ribeiro

Full text

FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO Mining Geographic Data for Fuel Consumption Estimation Vitor Daniel Ferreira da Cunha Ribeiro Master in Informatics and Computing Engineering Supervisor: Ana Cristina Costa Aguiar (Prof.) Co-Supervisor: João Rodrigues (PhD Student) 31st January, 2013 Mining Geographic Data for Fuel Consumption Estimation Vitor Daniel Ferreira da Cunha Ribeiro Master in Informatics and Computing Engineering Approved in oral examination by the committee: Chair: Ricardo Morla (Prof.) External Examiner: Pedro Abreu (Prof.) Supervisor: Ana Aguiar (Prof.) 18st February, 2013 Abstract In today’s world, society highly depends on mobility. Modern cities possess large and complex transportation networks, whose operating conditions sometimes preclude the minimization of human mobility associated costs, and contribute to an unsustainable growth of carbon emissions. Even in this technological era people live in, with space and mobile technology, the advantages or consequences of each mobility option, is not always direct or available to a user in real time. In this scenario, sometimes a user is imbued with finding the best option with limited resources. In an effort to reduce both emissions and, mostly important, costs, since money is a global motivator, an innovative solution for fuel consumption estimation is proposed, taking advantage of the opportunities generated by the growing market of mobile devices. The proposed solution involves calculating fuel consumption with a GPS only, present in every modern smartphone, taking advantage of its ubiquitous availability. In order to achieve this goal, a large dataset of GPS and fuel consumption data was gathered by various volunteers, using an Android mobile application as a gathering unit. The application logs the smartphone’s embedded sensor data, and uses an external device known as an On-Board Diagnostics device, to gather vehicle data. In the end, a solution oriented to emission maps is proposed to enhance the current urban monitor. i ii Sumário É possivel afirmar, nos dias correntes, que a sociedade depende largamente da mobilidade. Nas cidades modernas que possuem sistemas de transporte complexos, por vezes, a falta de interoperabilidade entre estes serviços, limita a minimização dos custos associados com a mobilidade, e contribui para um crescimento insustentável de emissões de carbono. Mesmo na era tecnológica em que vivemos, com acesso à tecnologia espacial e móvel, as vantagens ou consequências de cada opção de mobilidade não estão sempre disponíveis em tempo real. Neste cenário, recai por vezes à pessoa a tarefa de encontrar a melhor opção pelos seus próprios recursos. Num esforço para reduzir estas emissões e, principalmente, custos, dado que o dinheiro é um motivador global, é proposta uma solução inovadora para estimar o consumo de combustível, tirando partido das oportunidades geradas pelo crescimento notório do mercado móvel. A solução envolve o cálculo de consumo de combustivel através de um sensor GPS, presente em qualquer smartphone moderno, tirando assim partido da disponibilidade ubíqua desta tecnologia. Para alcançar este objectivo, foi feita uma recolha de dados por parte de vários voluntários, usando uma aplicação móvel para Android. Esta aplicação grava os dados recolhidos pelos sensores embebidos no smartphone, e usa um dispositivo externo, conhecido como OBD, para recolher dados de veículos. A solução encontrada pode ser usada para mapas de emissões. iii iv Acknowledgements First and foremost, i dedicate this work to the kind volunteers that agreed to spend their fuel in the name of this cause. I would like then to thank Professor Carlos Tavares de Pinho and Carlos Valentim for their time and patience. I thank Professor Jaime Villate for his friendship and Professor António Augusto de Sousa for his kindness and availability, for they are both valuable pillars in the MIEIC structure. I also thank their capability of reducing the initial awkwardness towards the teaching class. A special thanks goes to my supervisor Ana Aguiar for all the help and guidance provided, and for putting up with me and my constant confusion, and never loosing the will to give one more push. I also thank my co-supervisor and dear friend João Rodrigues for the constant intellectual bullying, the bags of popcorn for the late hours, for teaching me about sugar turning into caramel, and being a light at times of despair more times than a sane person could (although i suspect he has already lost it). This work would have not been possible without my godly supervisors. Finally, before the music starts, i thank my loved ones, family, friends, my dogs, Google, and all the people who supported me through my academic life, all the workers at Instituto de Telecomunicações for their effort in maintaining a fun and healthy work environment, and all the unnamed people involved in this thesis. Thank you also my imaginary friend for reappearing 15 years later and not letting me talk alone to my bedroom wall. v CONTENTS xii List of Figures 2.1 Approach and Fields of Study . . . . . . . . . . . . . . . . . . . . . . . 5 2.2 EngineEvolution .............................. 6 2.3 GasolineVSDiesel............................. 7 2.4 Garmin Mechanic Interface . . . . . . . . . . . . . . . . . . . . . . . . . 10 2.5 DashCommandInterface.......................... 11 2.6 alOBDInterface............................... 12 2.7 DashDynoInterface............................. 13 2.8 TransportationPyramid........................... 15 2.9 EMIT Model Structure [CCN02]...................... 16 2.10TripStatistics................................ 18 2.11 CMEM Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . 19 2.12MOVESInterface.............................. 21 2.13ecorioinstructions.............................. 24 2.14greenMeterInterface ............................ 25 2.15MDDArchitecture ............................. 26 2.16MDDWebsite................................ 26 4.1 Model.................................... 32 4.2 DetailedModel ............................... 32 5.1 CoordinateSystem ............................. 36 5.2 ELM327clones............................... 38 5.3 Mounting an OBD device to a Peugeot 207 . . . . . . . . . . . . . . . . 38 5.4 OBDMessageFormat ........................... 39 5.5 CAN OBD Message Format . . . . . . . . . . . . . . . . . . . . . . . . 39 5.6 Average Number Of OBD Responses Per Second In Various Vehicles . . 42 5.7 Four-StrokeCycle.............................. 43 5.8 MDDInterface ............................... 50 5.9 MDDArchitecture ............................. 50 5.10ServiceLifecycle .............................. 56 5.11 Number Of Points For Each Vehicle . . . . . . . . . . . . . . . . . . . . 59 5.12 Speed versus Steepness points . . . . . . . . . . . . . . . . . . . . . . . 61 5.13 Delay suffered in System Time minus GPS Time . . . . . . . . . . . . . 62 5.14 GPS Speed versus OBD Speed in metres per second . . . . . . . . . . . . 63 5.15 GPS Speed versus OBD Speed (Mean) in metres per second . . . . . . . 64 5.16 Inclination versus Fuel Consumption . . . . . . . . . . . . . . . . . . . . 65 5.17 Inclination versus Fuel Consumption (Mean) . . . . . . . . . . . . . . . 66 xiii LIST OF FIGURES 5.18InclinationHistogram............................ 66 5.19 Acceleration versus Fuel Consumption . . . . . . . . . . . . . . . . . . . 67 5.20 Acceleration versus Fuel Consumption (Mean) . . . . . . . . . . . . . . 67 5.21 Acceleration Histogram . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 5.22 Speed versus Fuel Consumption . . . . . . . . . . . . . . . . . . . . . . 68 5.23 Speed versus Fuel Consumption (Mean) . . . . . . . . . . . . . . . . . . 69 5.24SpeedHistogram .............................. 69 5.25 Inclination versus Fuel Consumption (Mean) without outliers . . . . . . . 70 5.26 Acceleration versus Fuel Consumption (Mean) without outliers . . . . . . 71 5.27 Inclination versus Fuel Consumption (Mean) for positive (Red) and negative (Black) values and Acceleration versus Fuel Consumption (Mean) for positive (Blue) and negative (Green) values . . . . . . . . . . . . . . 72 5.28 Inclination versus Fuel Consumption (Mean) for different vehicles . . . . 73 5.29 Acceleration versus Fuel Consumption (Mean) for different vehicles . . . 73 5.30 Speed versus Fuel Consumption (Mean) for different vehicles . . . . . . . 74 5.31 Inclination versus Acceleration (Mean) . . . . . . . . . . . . . . . . . . . 74 5.32 Speed versus Fuel Consumption from Vehicle V42 . . . . . . . . . . . . 75 5.33 Fuel Consumption Histogram . . . . . . . . . . . . . . . . . . . . . . . . 75 5.343DHistogram................................ 76 6.1 FuelEfficiency ............................... 78 6.2 Fuel Efficiency with Simple Linear Regression . . . . . . . . . . . . . . 79 6.3 Fuel Efficiency with Polynomial Regression . . . . . . . . . . . . . . . . 79 6.4 Component+Residual Plots, Linear Model Summary And Studentized Residuals for Inclination Versus Fuel Consumption . . . . . . . . . . . . 81 6.5 Component+Residual Plots, Linear Model Summary And Studentized Residuals for Inclination Versus Fuel Consumption (Mean) . . . . . . . . 82 6.6 Component+Residual Plots, Linear Model Summary And Studentized Residuals for Acceleration Versus Fuel Consumption . . . . . . . . . . . 83 6.7 Component+Residual Plots, Linear Model Summary And Studentized Residuals for Acceleration Versus Fuel Consumption (Mean) . . . . . . . 84 6.8 Component+Residual Plots, Linear Model Summary And Studentized Residuals for Speed Versus Fuel Consumption . . . . . . . . . . . . . . . 85 6.9 Component+Residual Plots, Linear Model Summary And Studentized Residuals for Speed Versus Fuel Consumption (Mean) . . . . . . . . . . 86 6.10 Linear Model Summary And Studentized Residuals for All Values Versus FuelConsumption.............................. 87 6.11 Logarithmic Transform on Speed and Fuel Consumption . . . . . . . . . 87 6.12 Speed vs Fuel Consumption (Mean) plot for lower speeds . . . . . . . . . 88 6.13 Studentized Residuals for All Values Versus Fuel Consumption . . . . . . 90 A.1 Speed Vs Fuel Consumption (Mean) on 20 Trips . . . . . . . . . . . . . 97 A.2 Speed Vs Fuel Consumption (Mean) on 40 Trips . . . . . . . . . . . . . 98 A.3 Speed Vs Fuel Consumption (Mean) on 60 Trips . . . . . . . . . . . . . 98 A.4 Speed Vs Fuel Consumption (Mean) on 80 Trips . . . . . . . . . . . . . 99 A.5 Speed Vs Fuel Consumption (Mean) on 100 Trips . . . . . . . . . . . . . 99 A.6 Summary of Final Formula . . . . . . . . . . . . . . . . . . . . . . . . . 100 xiv List of Tables 5.1 Stoichiometric Ratios For Common Fuel Types . . . . . . . . . . . . . . 45 5.2 Tested Smartphones and Vehicles . . . . . . . . . . . . . . . . . . . . . . 57 5.3 VehicleIdentification............................ 59 5.4 Calculated Versus Shown Values Sample . . . . . . . . . . . . . . . . . . 64 5.5 Percentile of Data Points . . . . . . . . . . . . . . . . . . . . . . . . . . 65 5.6 Correlation And Covariance between variables . . . . . . . . . . . . . . 70 5.7 Correlation And Covariance between variables using aggregated mean andwithoutoutliers............................. 71 6.1 Standard residual error for each model and some vehicles . . . . . . . . . 89 xv LIST OF TABLES xvi Abbreviations FEUP Faculdade de Engenharia da Universidade do Porto IT Instituto de Telecomunicações OBD On-Board Diagnostics UI User Interface EPA Environmental Protection Agency CARB California Air Resources Board ECU Engine Control Unit PID Parameter ID AFR Air-to-Fuel Ratio MDD MyDrivingDroid VSP Vehicle Specific Power VANET Vehicular Ad-hoc NETwork GPS Global Positioning System AP Access Point ANN Artificial Neural Network SVM Support Vector Machine API Application Programming Interface OLS Ordinary Least Squares EM Errors-in-Model SAE Society of Automotive Engineers CAN Controller Area Network NMEA National Marine Electronics Association SDP Service Discovery Protocol CoD Class of Device xvii ABBREVIATIONS xviii Chapter 1 Introduction Mobility has always been an important factor in human history. Considering the rising and continuous change of user requirements, where people increasingly demand everything everywhere [GT08], mobility can be considered one of the foundations of economics and human development [Haa09]. In a world where the evolution of technology has widened the possibilities of innovation, providing new user experiences with mobile platforms like smartphones, tablets, PDAs, and other mobile devices that defined the mobile space, giving the power of mobility at the tip of a finger, this platform has become a hot research topic. Furthermore, mobile devices have become even more powerful with the constant improvement and addition of motion, position and environmental sensors. Harnessing the power of mobile platforms, it is possible to further exploit human mobility by continually improving traffic patterns and driving styles. Parallel to this situation, is the constant improvement of space technology and wireless communications, which aims to provide ubiquitous wireless access [PvHD06]. 1.1 Context In modern cities where infrastructures provide various means of transportation, including private vehicles, buses, taxis, railways and other forms of public transit, a user is sometimes imbued with the task of planing and organizing his travel routes. Because these systems can be large complex networks that usually operate independently instead of collaborating with each other in a holistic fashion [Lit03], the task of coordinating them can prove to be a difficult one as travellers are typically unaware of the current conditions of the various transportation networks, as of the exact time, monetary costs and environmental impact that a certain mobility option incurs. 1 Introduction Since time, money and environmental impact can be directly associated with a route option, this uncoordinated setting limits the minimization of these costs and contributes to an inefficient exploitation of the transportation services as to an unsustainable growth of carbon emissions associated with human mobility [RKBG07]. 1.2 Motivation Considering this scenario, it is difficult to know, for example, exactly how much time and fuel is spent in our travels, what is the most economic route between point A and point B at a certain time, or if there is someone who consumes less fuel or takes less time in the same route and vehicle class and how those differences can be explained. On the other hand, the continuous rise of fuel prices has stimulated the search for alternatives, not only in engine types or driving behaviour [LLL+12], but has also boosted the search for new methods of intelligent transportation [Li12]. There are also on-going researches that aim towards giving vehicle users better fuel efficiency feedback [CQSK11], and lead some to seek driving techniques that maximize fuel economy, also know as hypermiling. The wide availability of smartphones due to its increasingly accessible price has created a potential tool to use in the process of solving these problems. Mobile devices today, besides the obvious portability, possess a high processing power and are equipped with a wide range of sensors like an accelerometer, gyroscope, magnetometer, GPS, Wi-Fi and Bluetooth devices. Using the device’s embedded capabilities, the implementation in these mobile platforms of algorithms that extract information about mobility patterns and ecological efficiency, offer a unique opportunity to leverage information to empower citizens to make informed decisions improving the quality of life of people in a perceivable way [LML10]. 1.3 Goals The general purpose of this thesis is to provide an innovative solution by joining the previous scenarios and present algorithms that enable estimation of fuel consumption from GPS data from a mobile phone that can be placed anywhere. The reasons for choosing only the GPS, besides the ubiquitous availability of the Global Navigation Satellite System (GNSS) in open spaces, is explained in detail on Chapter 5. Among the current proposals, some models provide a solution for generic classes of vehicles with data gathered from an On-Board Diagnostics device (OBD). This will be further addressed in Chapter 2. But since the user must own an OBD device, and these 2 Introduction devices are not widely available due to price and required knowledge of OBD technology, the development of algorithms that can estimate the fuel consumption from mobile sensors enables anyone with only a smartphone to know more about his driving behaviour. A mobile application for Android named MyDrivingDroid (MDD) was developed at Instituto de Telecomunicações that gathers data from a mobile device’s sensors. The first goal of this thesis was to extend that application to be able to gather vehicle data from an On-Board Diagnostics device through a Bluetooth connection. The implementation details of this module are discussed in Chapter 5. A mobile-adapted version of the existing OBD models was implemented, used and improved in the same application to provide automotive data feedback in real time, like fuel consumption. Through the constant use of the extended mobile application in multiple vehicles from various volunteers, a large dataset was collected from the available smartphone’s sensors and external OBD sensor, and fuel efficiency obtained from the OBD dataset was validated. The validation process is explained in Chapter 5. Through constant improvement of the algorithms, this work’s solution can contribute to an enhancement of the existing urban monitor systems. Thus, with the proposed algorithm that estimates fuel consumption, a possibility is presented to provide each user with their own driving statistics, and, in a collaborative fashion, compare it with others and, with this data, recommend alternative mobility solutions, with the help of the mobile platforms’ resources. 1.4 Outline Besides the introduction, this document will detail the related work of the issues described previously regarding engine mechanics and some of the OBD standard’s background, fuel consumption and emission models, among other topics, that were used to provide the proposed solution. Furthermore, chapter 5discusses the implementation details, including the OBD protocols, the chosen Android sensors, the fuel consumption calculation from vehicle data and the OBD Module for MDD, in order to enable the gathering of sensor data. Also in this chapter, the data collection methodology and processing are discussed, and results are presented, along with the methods used to analyse, calibrate and process data. It is also explained how some GPS values were derived and they were used in the obtained fuel consumption algorithms, followed by observations and conclusions. The last chapter discusses regression and some machine learning approaches, and details experimented models and the formula obtained from the data. The introduction has stated the problem which drove this research and provided an insight of what is expected. This, of course, will be further developed in the next chapters. 3 State Of The Art Figure 2.4: Garmin Mechanic Interface error-code scanner for Check Engine Lights issues. The product is closed source and it is not free, although the manuals are and provide some tips for hypermiling. 2.2.3 ScanXL and DashCommand Palmer Performance Engineering is an industry dedicated in manufacturing products for vehicle diagnostic and communications software, notoriously ScanXL and DashCommand. Their products rely on OBDII devices to provide vehicle data and deliver generic diagnostic data in real time 7. It is also possible to view vehicle Diagnostic Trouble Codes. The products are closed source and not free, but the manuals are publicly available at their website. This is a great source of information as the manuals provide instructions on how their technology, DashXL, works in detail. Among the information, it is explained in a model how fuel efficiency is obtained. Since this thesis required a fuel consumption model, these manuals provided very useful information on the OBD model implementation. 2.2.4 Torque Pro Torque 8is an Android application made by Ian Hawkins that connects to an OBDII compliant device via Bluetooth serving as a vehicle performance and diagnostic tool. It allows the user to create vehicle profiles to help improve the accuracy of the outcome. 7http://www.palmerperformance.com 8http://torque-bhp.com/ 10 State Of The Art Figure 2.5: Dash Command Interface It is a commercial product that shows graphs and virtual gauges with readings from the OBD port data, like horsepower, torque, acceleration, etc. Also uses Android sensors like the GPS. Unfortunately, since it is not open source and the documentation in the website is limited, it is not fully known how the data is used. The only open source component is a plugin for developers to improve and upgrade the application, but it does not provide useful information for this project. 2.2.5 alOBD ScanGenPro Alexandre Beloussov has developed some Android Scantool applications 9that provide basic automotive diagnostic by gathering data from an OBDII device via Bluetooth. The products are priced as well as some extensions designed for specific vehicles. It is not as intuitive as Torque as, according to the author, is intended for an audience familiar with the operation of an engine control computer and vehicle sensors. It is also closed source. 2.2.6 OBDroid OBDroid is a very simple diagnostics and vehicle interface tool that uses Bluetooth OBDII devices. It reads data directly from the OBD PIDs and shows it in some form of visual interface. Although is it not free and it does not possess a fuel consumption model, a part of the source code is hosted at github 10 and provides useful information on how to read data from a Bluetooth On-Board Diagnostics with an Android. 9http://www.alobdscanner.com/ 10https://github.com/syntelos/obdroid 11 State Of The Art Figure 2.6: alOBD Interface 2.2.7 OBD II Reader This is a free and simple open source Android application 11 that also uses a Bluetooth OBDII device to read PIDs and present simple data. It is also an interesting reference on how to code OBD data handling. 2.2.8 VoyagerDash VoyagerDash from GTOSoft 12 is an engine trouble code diagnostics and engine monitor. It reads data from an OBD device via Bluetooth and shows it in form of gauges and graphs. It is not free and it is closed source. 2.2.9 DashDyno SPD Auterra DashDyno SPD 13 is a device for instant and average fuel economy measures. It logs data from internal and external sensors. The internal being engine sensors and external are add-ons like GPS or O2 sensors for Air-to-Fuel Ratio measures. From speed and RPM it estimates horsepower and torque. It also serves as a diagnostics tool that warns the user about issues related to the Check Engine Light. Comes with a software for Windows for travel analysis and simulation resorting to Google Earth. It is a commercial product and no source code or model is publicly available. 11http://code.google.com/p/android-obd-reader 12http://www.gtosoft.com/ 13http://www.auterraweb.com/dashdynoseries.html 12 State Of The Art Figure 2.7: DashDyno Interface 2.2.10 Applications Overview Since the source code for most of the applications is not public, it is difficult to comprehend their inner workings in detail. However, it is known that these applications all heavily depend on the available PID’s from an OBD connector to calculate results. Though some applications provide information on how to code OBD handling and how to use the retrieved data to construct a fuel consumption model, it is necessary to further the knowledge about this type of models. The next section will address that issue. 2.3 Fuel Consumption Models Some of the previous mobile applications employ fuel consumption models, but they usually tend to generalize some values like vehicle specifications as not all vehicle related values can be obtained just through an OBD. The vehicle’s structure, engine and parameters can have a great impact on the outcome, and simple models either use constants for generic types of vehicles, or prompt the user about that information. Some applications prompt the user to insert vehicle profiles, which can encompass values like engine displacement, fuel type, tank level and capacity, car size and weight, RPM redline, among others, or even more detailed engine information. Unfortunately, for users not familiar with vehicle specifications, filling that information can be confusing and time-consuming. 13 State Of The Art Furthermore, there are more complex models that not only encompass more detailed vehicle parameters, but are also designed for specific vehicle classes [Fen07] or engines [Ono04], as it is expectable to obtain differences in results between, for example, a gasoline passenger car and a diesel heavy duty truck using just generic models, resulting in low accuracy. There are some models that also use very different approaches, ranging from speed and acceleration [ARTV02] to vehicle specific power [HDY+05] [McC99], and some use external sensors, like a emission sensor on the tailpipe. This section addresses an overview of model types and discusses existing models, pointing some of the advantages and disadvantages between them. Environmental laws and rules from the past decades to the present date have been established due to environmental concerns. In an effort to mitigate this concerns and raise environmental awareness, researches have been and are currently being made related to this topic. So the majority of the existing models are primarily emission related. However, since emissions are directly related to fuel consumption [CCN02], it is plausible to use these models or just specific modules from them. 2.3.1 Emission Model Types Models can be broadly categorized into macroscopic, mesoscopic, and microscopic models [FD12] where, as the name suggests, each is related to a certain scale. Macroscopic are often related to large scale emissions, like in a city. Mesoscopic are medium-scale, such as roads or an individual vehicle’s specific trips. Microscopic are the most detailed and vehicle-oriented and usually take into account more specific vehicle or engine parameters (see Figure 2.8 14). Some of this models use vehicle variables that can be categorized into vehicle parameters and traffic or road parameters. Vehicle parameters are vehicle class, mass and length, fuel type, tank size, engine displacement, accessory use, like air conditioning, among others. Traffic or road parameters are road related characteristics like road grade. Macroscopic models are mainly based on the average speed of the traffic flow. So they do not account for individual operation conditions or speed fluctuations caused by individual vehicles. These models tend to simplify calculations to estimate fuel consumption and pollutant emissions, which finally leads to a reduced accuracy for individual vehicle dynamics. Mesoscopic models use the average or instantaneous speed and acceleration, sometimes to construct virtual driving cycles, since they are more trip oriented. This means that these models usually use traffic or road parameters. 14http://www.cert.ucr.edu/cmem/ 14 State Of The Art Figure 2.8: Transportation Pyramid Microscopic emission models are based on instantaneous individual vehicle variables and can overcome some of the limitations of the macroscopic models since they take into account more detailed individual vehicle parameters and have higher temporal precision. Microscopic models can be classified into three subtypes [CCN02]: •Emission Maps normally based on speed/acceleration lookup tables; •Purely Statistical normally regression-based; •Load-based that use engine load. Emission maps are matrices that contain the average emission rates for every combination of speed and acceleration. Although generally easy to generate and use, they are usually insensitive to traffic or road parameters. Statistical regression-based models predict fuel consumption and emission rates by typically using regression methods on a large number of instantaneous vehicle speed and acceleration values. Load-based models are primarily based on fuel consumption rate which is a surrogate for engine power demand, also referred to as engine load. These models try to simulate the physical phenomena that generate emissions. They are usually more detailed and flexible and can use a lot of vehicle variables. This also means that these models can be more complex and can consume a lot of computational power. Most of the developed models are mesoscopic and microscopic that mostly use speed and acceleration as inputs, as they are closely related to emission outputs [IBL06]. Though 15 State Of The Art there is some discussion on whom is the most accurate and what is the best consumption calculation method [Yue08] [RA03] [DR02] [ZN01]. Huanyu Yue argues that most mesoscopic models use simplified mathematical expressions and ignore transient changes of a vehicle’s speed and acceleration. On the other hand, microscopic models can be very costly and time consuming [Yue08]. While in an article comparing various models, some authors favour mesoscopic model VT-Micro (Section 2.3.4) [RA03] that relies on speed and acceleration, while in another article it is argued that speed values are insufficient [DR02]. The following subsections detail some of the existing models. 2.3.2 EMIT (EMIssions from Traffic) EMIT is statistical model for instantaneous tailpipe emissions and fuel consumption derived from a regression-based and load-based emission modelling approaches [CCN02], composed by two modules as shown in Figure 2.9. Figure 2.9: EMIT Model Structure [CCN02] This model depends on second-by-second speed and acceleration and a vehicle category to predict the corresponding second-by-second fuel consumption. And since fuel consumption is related to emissions as previously stated, with these values the model also predicts second-by-second tailpipe emission rates. This model has been widely referenced in other model proposals as it gives reasonable results with acceptable accuracy for fuel consumption. EMIT presents good results for carbon dioxide emissions, reasonable accuracy for carbon monoxide and nitrogen oxides, but less desirable accuracy for hydrocarbons [CCN02]. The greatest disadvantage of this model is that besides requiring calibration from a large database, it is limited for light-duty vehicles (like a passenger car for common use) and is only suitable for hot-stabilized conditions with zero road grade. However, the fuel consumption module is greatly detailed and provides some useful formulas for fuel rate and allows to comprehend how it is related to emission rate. 16 State Of The Art EO =EI ∗FR (2.1) FR =CO2 MWCO2+CO MWCO∗(MWC +MWH ∗NM)+HC (2.2) EO - engine-out emission rate (g/s) EI - emission index (mass of emission per mass unit of fuel consumed) FR - fuel consumption rate (g/s) CO2- measured Carbon dioxide engine-out emission rate MWCO2- carbon dioxide molecular weight constant (44) CO - measured Carbon monoxide engine-out emission rate MWCO - carbon monoxide molecular weight constant (28) MWC - carbon molecular weight constant (44) MWH - hydrogen molecular weight constant (1) NM - approximate number of moles of hydrogen per mole of carbon in the fuel (1.85) HC - measured hydrocarbon engine-out emission rate Equation 2.1 supports the statement about emissions being closely related to fuel consumption. 2.3.3 aaSIDRA and aaMOTION These instantaneous emission models [AB03] were proposed in 2003 and later in 2007 were integrated in a closed source paid software and subsequently named SIDRA TRIP 15. The model uses microscopic GPS or trip data representing a standard driving cycle and produces vehicle trip assessment data like distance, speed, operating costs, emissions, noise and fuel consumption. If using a GPS, data is easily gathered and it produces reliable results. However, it depends on various vehicle variables like engine, traffic and road parameters. It is also possible to input fuel prices, for cost assessment. SIDRA TRIP provides default vehicle profiles, that can be configured or created, generalizing vehicle classes to facilitate user input, as some parameters can be complex for normal users. 2.3.4 VT-Micro Virginia Tech Micro [AR03] [RA04] is a mesoscopic model for normal and high emitting vehicles that predicts the instantaneous fuel consumption and emission rates of individual vehicles. 15http://www.sidrasolutions.com/ 17 State Of The Art Figure 2.10: Trip Statistics It constructs a synthetic drive cycle based on instantaneous speed and acceleration. For each drive cycle, the model estimates the proportion of time that a vehicle typically spends cruising, decelerating, idling and accelerating. MOEe=       e 3 ∑ i=0 3 ∑ j=0(Le i j∗ui∗aj)for a≥0 e 3 ∑ i=0 3 ∑ j=0(Me i j∗ui∗aj)for a<0 (2.3) MOEeinstantaneous emission rate (mg/s) Lei j - model regression coefficient for MOE ’e’ at speed power ’i’ and acceleration power ’j’ Mei j - model regression coefficient for MOE ’e’ at speed power ’i’ and acceleration power ’j’ uinstantaneous vehicle speed (km/h) ainstantaneous vehicle acceleration (m/s2) The model can provide accurate estimates of distance-based average vehicle fuel consumption and emission rates, but because of the non-linear relationship between vehicle 18 State Of The Art emissions and vehicle speed, the model is also prone to inaccuracies [RA04] [Yue08]. 2.3.5 CMEM (Comprehensive Modal Emission Model) The core of this microscopic emission model is the fuel rate calculation which is a function of engine speed and engine load (power demand) 16. This model consists of several modules as observed in Figure 2.11 Figure 2.11: CMEM Model Architecture Engine speed is determined based on vehicle velocity, gear shift schedule and power demand. The vehicle power demand is determined based on specific vehicle parameters, road grade and second-by-second vehicle speed, from which acceleration is derived. It can also account for accessory use like air conditioning. The model can use a total of 35 parameters to estimate vehicle tailpipe emissions and fuel use. The greatest advantage of CMEM is that it is a public-domain model and it is claimed to be one of the most detailed and best tested models for estimating hot-stabilized vehicle exhaust emissions, and has been used in a variety of projects, like for air quality evaluation in Southern California [Bor07]. For the purpose of this thesis, it not necessary to study all modules of the model. The most appropriate is the Fuel Rate Module on page 48 of the CMEM User’s Guide [BAY+06]. 16http://www.cert.ucr.edu/cmem/ 19 State Of The Art Figure 2.15: MDD Architecture The server is responsible for verifying the integrity of the incoming data. Then it stores that data in a MySQL database and makes it available for analysis and visualization in a website as shown in Figure 2.16. Figure 2.16: MDD Website 26 State Of The Art The goal of this project is to provide trip information, such as travel duration, the overall energy costs and carbon emissions, the areas of excessive traffic, among others. It is also intended that users share this information using this platform in a collaborative environment. As referred in the introduction in Section 1.3, one of the goals of this thesis was to extend MDD to be able to gather vehicle data from a Bluetooth OBDII compliant device. To do so, the previously discussed OBD applications on Section 2.2 and models on Section 2.3 served as important references. This is further addressed in Chapter 5. 2.4.3 Overview The advent of ubiquitous connectivity and powerful, affordable, user-friendly mobile devices offers a unique opportunity to leverage information and communication technologies to improve quality of life. While there are several Android projects that use sensing, none were found that use only the GPS sensor to estimate fuel consumption. Furthermore, MDD offers a great opportunity for expansion since each sensor runs on an individual service. This means that as technology evolves, adding a new sensor is as simply as adding a new service (see Chapter 5for the MDD architecture). Also, the MDD project can be used in a single or in a collaborative way exploiting the advantages that collaborative platforms offer. 2.5 UI Patterns Anyone developing mobile applications intended to function while a user is driving, should be extremely cautious about user interfaces. Not only it is dangerous to operate a phone while driving [Gom12], it is also illegal in some countries, like Portugal 27. That is why some vehicles provide embedded phones or an optional phone dock. Taking those facts in account, ’Car Mode’ mobile applications are available, that use simple interfaces mostly destined to be used while the smartphone is in a dock. They should abide some standard UI patterns 28 for quick interaction. When developing the solution, this will be accounted for. 2.6 Conclusions In this chapter, it is possible to learn the differences between vehicle engines and the reasons behind the appearance of OBD standards, and draw some conclusions about the 27http://www.ansr.pt/ 28http://developer.android.com/design/index.html 27 State Of The Art necessity of emission and fuel models, not only for the environmental impact but also because of the impact they can have on a person’s wallet. There are some On-Board Diagnostics applications that effectively in the long-term can save money, though they possess the handicap that not all users want to or know how to use OBDs. Furthermore, they are focused on individuals and not on a collaborative perspective, so they are mostly for informative than comparative purposes, which is something the proposed solution can eventually provide using the power of mobile platforms. 28 Chapter 3 Problem In Chapter 1it is referred some of the problems caused by mobility that a user faces. One of those is the difficulty of knowing how much fuel a mobility option requires so a person can translate it to costs. So if a person wants to know his fuel consumption, either he owns and understands the OBD technology, and has the availability and knowledge to install an OBD interface in his vehicle, or he uses a software, like the ones presented in Chapter 2, that normally require user input. So, in the end, most of the offers require user intervention. On the other hand, the OBD technology has the limitation that not all vehicles support it. Also it is not as widespread as the mobile technology in the present day. In modern societies it is increasingly more common a person owning a smartphone than owning an OBD device, as this last technology is not usually well known by the average user. So the OBD technology can serve an important, but nevertheless, less popular purpose in our society than the mobile technology. Alo, no fuel consumption parameter is given by the OBD, so most of these softwares use undisclosed fuel consumption formulas, which is something this thesis provides in Chapter 5. It is also important to note that time in a fuel consumption calculation is a very important variable, and it is known that not all vehicles have the same response time. This is also validated in Chapter 5. This means that not all vehicles are capable of providing all the variables necessary to calculate fuel consumption in a one second window. So in order to better control mobility costs, very few free options are available. As referred in Chapter 2, the current solutions are either paid and/or depend on the OBD technology. Also not all vehicles posses an interface with trip assessment information, and even if they do, a user has little chance of using that information in a collaborative fashion. The uploaded aggregated data from different vehicles creates lots of opportunities. Besides contributing to an enhancement of the urban monitor, the usage of this data can eventually help the individual user, as suggested in Chapter 7. 29 Problem 30 Chapter 4 Proposed Solution The proposed solution is an algorithm that estimates fuel consumption from GPS data. The GPS technology, besides its ubiquitous availability, is widely available in many devices, like the modern smartphones. In order to achieve this goal, and know how the GPS data relates to fuel consumption, a dataset of second by second fuel consumption from vehicles is needed along with the corresponding second by second GPS data. Thinking in mathematical terms, what is needed is a statistical technique to estimating the relationships among the dependent variable, fuel consumption, and the independent variables from GPS data. The first step was to understand the GPS technology and what values could be derived from it. Next it was needed to know how fuel consumption could be obtained, and this is where the OBD technology came in. An OBD device connected to a vehicle transmits engine information in real time through a Bluetooth connection. However, as stated in Chapter 2, since the OBD protocol does not offer a specific fuel consumption parameter, it was also necessary to obtain a fuel consumption formula from the available OBD data. So the OBD protocol was also studied in order to know what values a vehicle can provide and how can they be fetched. With those values, emission and fuel consumption models were also studied, so to become possible to know what formula to use. Furthermore, a gathering unit was needed and developed to log OBD and GPS data at the same time. The MDD application for Android was chosen because it is a data gathering application with the opportunity of expansion. To expand the OBD module, research into the Bluetooth communications protocols was also required. By using this data gathering unit, it becomes possible to construct the required dataset. But in order to obtain a dataset large enough to extract conclusions, some volunteers were needed. The more data, and more volunteers, and more vehicles the better. Unfortunately, this was a difficult task because, as stated, not all vehicles support the OBD protocol. As data became available, so did the goal of creating an algorithm that estimates fuel consumption from GPS data through regression analysis, became possible. 31 Proposed Solution So the implemented algorithm consists of two modules that estimate fuel consumption. One that uses gathered OBD data, and another that uses only embedded GPS sensor data. The OBD dataset is used for the calibration of the other module. The approach is presented with the help of the Figures 4.1 and 4.2. Figure 4.1: Model Figure 4.2: Detailed Model Since the data gathering step had to include the broadest public possible, MDD was required to run on a wide range of smartphones. This topic, the process of gathering volunteers, along with the GPS and OBD technology and the module implemented in MyDrivingDroid are detailed in Chapter 5. 32 Chapter 5 Obtaining Data This chapter details the implementation, data collection methodologies and the data processing steps. The beginning of each of these three sections summarizes the respective topic. 5.1 Implementation In this section the choice over the Android’s embedded sensors is discussed and the OBD communication protocols are detailed, as well as the values used from the OBD standards to calculate fuel consumption. The approach used in developing the fuel consumption formulas is also explained in detail, and a background in vehicle mechanics is presented to support the explanation. The integration and implementation details of the module added to MyDrivingDroid to gather OBD data via Bluetooth is also detailed. Finally, an overview of the implementation step is presented. 5.1.1 Android Sensors As mentioned before, the goal of this thesis was to develop an algorithm that estimates fuel consumption from sensor data. Particularly from the GPS sensor present in today’s mobile phones. To do so, it is also important to understand what results are possible to deduct from a sensor or a group of sensors. Most Android-powered devices have built-in sensors capable of providing raw data with high precision and accuracy. While a device can have more than one sensor of a given type, very few have every type of supported sensors embedded. Nevertheless, most of the Android devices have a built-in accelerometer, magnetometer and a GPS 1. 1http://developer.android.com/guide/topics/sensors/sensors_overview.html 33 Obtaining Data The Android platform supports motion, position and environmental sensors. While most of these sensors are hardware-based, some are software-based. The hardware-based sensors are physical components that directly measure specific environmental properties to derive data. Software-based are virtual sensors that mimic hardware-based sensors by deriving data from one or more of the hardware-based sensors. The motion sensors include accelerometers, gyroscopes and virtual gravity and rotational vector sensors to pinpoint and measure acceleration and rotational forces along the three axes. The position sensors include GPS, magnetometers and virtual orientation sensors to measure the physical position of a device. The environmental sensors include barometers, photometers and thermometers to measure illumination, humidity, ambient air temperature and pressure. The following subsections discuss how and what is possible to obtain from some of the Android’s motion and position sensors, and why only the GPS sensor was chosen. Some very useful formulas and documentation on how to handle sensor data on an Android can be consulted at the Android Developers webpage 2. For this project, the most useful ones are referenced below. 5.1.1.1 GPS The Global Positioning System (GPS) 3is a space-based satellite navigation system that provides location, time and satellite information anywhere on Earth. It is freely accessible by anyone with a GPS receiver, which is built-in in any present Android smartphone. With the GPS sensor with is possible to know a user’s location information about latitude and longitude in degrees, altitude in meters and bearing in degrees to the north, also known as azimuth, speed in meters per second and a corresponding GPS timestamp. The GPS system consists of three modules. The satellites that transmit the position information, the receiver that collects data from the satellites and computes its location based on that data, and the ground stations that control and monitor the satellites and update their information guaranteeing their health and atomic clock’s accuracy. To compute a location, the GPS receiver synchronizes with the available satellites and downloads the navigation information, consisting of the satellite’s atomic clock information, ionosphere data, and orbit (ephemeris) data. To maintain a fix, the GPS receiver continuously recalculates the information from the moving satellites. Once it has a fix from sufficient satellites, it is possible to derive much more information than just location or time data, like travel direction (compass heading), distance and speed. The GPS mechanism is very complex. The book Understanding GPS by Kaplan and Hegarty [KH97] is advised to deepen the knowledge of this technological marvel. It 2http://developer.android.com/reference/android/hardware/SensorEvent.html 3http://www.af.mil/information/factsheets/factsheet.asp?id=119 34 Obtaining Data explains that location and velocity are obtained differently, and velocity is actually more precise as it is obtained by the Doppler effect [ZZGD06]. So velocity provides better accuracy for acceleration calculations (see Section 5.3.4 for details) than position. There are various methods and algorithms to compute user location, and they are already implemented in the Android OS. Browsing the Android’s Location Manager source code it is possible to discover an implementation of the Inverse Formula [Vin75] to compute horizontal distance and bearing between latitude and longitude points. However, there are also some other methods implemented by specific manufacturers and for specific ROMs. This lead to some encounters with firmware bugs. Although the position and velocity information on the tested devices presented no problems, on some Android smartphones the GPS timestamp showed a one day difference on leap years, like the year 2012. Since there are no guarantees about the smartphone’s system timestamp reliability (as shown on Section 5.3.3), to correct this bug, a NMEA Listener was also implemented to receive raw NMEA sentences from the GPS to correct time. NMEA is a standard for communicating with marine electronic devices and is a common method for receiving data from a GPS over a serial port [NME02]. So if the GPS timestamp has a bigger difference then a one day offset plus the smartphone’s timestamp and a one day offset plus the NMEA timestamp, it is corrected by subtracting a day’s equivalent time. Satellites have an atomic clock to keep the time very precisely that is used in the location calculations. As previously referred, one of the applications of GPS technology is also to provide the correct time. Still, these calculations are likely to have an error margin. An accuracy error estimative is reported by the Android OS that helps to filter undesired data points. The most common error is receiver clock bias, which affects pseudorange measurements. This can be compensated by additional satellites. Also another common errors are the ionospheric delay due to the sun’s radiation, receiver noise and multipath that occurs when the original GPS signal is reflected before it reaches the receiver [Wan09]. Still, the GPS is the most accurate sensor on position and velocity information, but it only works outdoors and it takes time to obtain a fix. In Android devices it also quickly consumes battery power, so there is a possibility to determine user location using cell tower and Wi-Fi signals, providing location information in a way that works both indoors and outdoors, responds faster, and uses less battery power. However, the Network Location Provider cannot continually pinpoint the exact location and speed that is crucial for this work. So only GPS based locations are used. Although, most of the time spent driving is in an outdoor environment, this means that no data can be provided when the driver is, for example, in a tunnel. In this case, another approach is needed. But because there are always at least four GPS satellites available at any given time and position, since July of 1995, the outdoor data can always be gathered and proved to be sufficient. 35 Obtaining Data Figure 5.6: Average Number Of OBD Responses Per Second In Various Vehicles responses as stated before. The details can be seen on the ’Multiline Responses’ chapter in the ELM327 documentation [ELM11]. This was also resolved with regular expressions. 5.1.3 Fuel Consumption Calculation Before expanding MyDrivingDroid with a service to gather OBD data, it is important to understand how engines and the OBD protocol work, and what is the data needed to calculate fuel consumption, since there are hundreds of codes on the OBD standard. This was done using the information obtained from the state of the art research as a guideline reference. The purpose of an engine is to convert fuel into motion. In this case so that the vehicle moves. There are two kinds of combustion engines, external (like steam engines) and internal like petrol or diesel engines. Currently, most of the vehicles have these two types of internal combustion engines that use a four-stroke combustion cycle to convert fuel into motion. The four-stroke approach is also known as the Otto cycle, and correspond in order of execution to the intake stroke, compression stroke, combustion stroke and exhaust stroke (see Figure 5.7). The algorithm is aimed at this type of engines. 42 Obtaining Data Figure 5.7: Four-Stroke Cycle Fuel consumption, also known as fuel efficiency, is a ratio of fuel consumed per distance travelled. In most European vehicles it is normally represented as litres per 100 kilometres. The fuel consumed, also known as fuel flow or fuel rate, means how many litres are consumed per hour. So, at a given speed in kilometres per hour, to calculate fuel consumption the expression below is used. FuelConsumption(l/100km) = FuelFlow(l/h) Speed(km/h)∗100 (5.3) As stated before, the OBD standard does not explicitly include a fuel economy parameter, but includes PIDs for engine fuel flow (code 015E) and speed (code 010D). However, while speed is mandatorily available, fuel flow is not. Either because the manufacturer chooses not to make it available, or because there is no sensor inserted in the fuel line between the fuel tank and the carburettor of the engine to measure litres per hour. In most of the cases this information is not available, as was the unfortunate case of all the tested cars. So a fuel flow expression is needed. The principle behind an internal combustion engine is that by putting fuel in a small enclosed space and igniting it, an amount of energy is released in the form of expanding gas. In order for combustion to occur, however, a fuel and an oxidant are needed to be burnt for this chemical reactions to take place. The difference between diesel and gasoline engines, besides the type of fuel, is that with gasoline engines, fuel is mixed 43 Obtaining Data with air, compressed by pistons and ignited by a spark plug. In a diesel engine the air is compressed first, which causes it to heat up, and then when the fuel is injected it ignites. It is important to know that there is a ratio for how much a certain type of fuel can be burnt with a certain amount of oxygen. So the air-fuel ratio in the cylinder is a limiting factor. If there is excess oxygen, it is called lean combustion. When there is an excess of fuel, it is a rich combustion. In an incomplete combustion not all the excess fuel is burnt and it will be pushed out through the exhaust valve. The ’perfect’ ratio for both the fuel and the oxygen in the air to be completely consumed is called the stoichiometric air-fuel ratio. In order to be able to judge if an air-fuel mixture has the correct ratio of air and fuel, the composition of a fuel has also to be known. So when there is no information about the amount of fuel entering the combustion chamber (fuel flow in litres per hour), it can be determined by using the instantaneous rate of air entering the combustion chamber, if the air to fuel ratio and fuel density is known. In the OBD standards there is a PID for Mass Air Flow or MAF (code 0110) that gives the mass of air read by the sensor placed just before the intake manifold in grams per second. So it is possible to use the formula below to determine fuel flow. MAF ∗3600 FF =AFRA∗FD <=>FF =MAF ∗3600 AFRA∗FD (5.4) FF - fuel flow (l/h) MAF - mass air flow (g/s) AFRA- actual air to fuel ratio FD - fuel density (g/l) Since diesel engines operate at a higher compression ratio, the mass air flow has to be adjusted. Both spark and compression ignition engines support Calculated Engine Load (code 0104) that is the relative cylinder charge. This value is the airflow divided by the peak air flow at wide open throttle in standard temperature (25 oC) and pressure (101.325 kPa) as a function of RPM. For diesel engines, it is the current output torque divided by peak output torque at current RPM. Engine Load is linearly correlated with engine vacuum. So there is linearity between load, airflow and torque. So for diesel engines MAF becomes: MAFe=MAF ∗LOADCALC (5.5) MAFeadjusted amount of air that is actually used to burn fuel (g/s) MAF - mass air flow (g/s) LOADCALC - calculated load 44 Obtaining Data Fuel density and the stoichiometric air to fuel ratio are constants that depend on fuel type (see Table 5.1 for stoich ratios). Although, as seen before, the actual air to fuel ratio in the engine is not always stoichiometric. Sometimes the are lean or rich mixtures. To correct this ratio, a lambda value is used based on corrections obtained from the Fuel Trim, which is generally calculated by using a wide set of data values. These include oxygen (O2) sensors, intake air temperature and pressure, coolant temperature, anti-knock sensors, engine load, changes in throttle position, and even battery voltage. AFRA=AFRS+ (AFRS∗(Lambda/100)) (5.6) AFRA- actual air to fuel ratio AFRS- stoichiometric air to fuel ratio Lambda - correction value Table 5.1: Stoichiometric Ratios For Common Fuel Types Fuel Type Stoich Ratio Gasoline 14.7 Diesel 14.6 Propane 15.5 Ethanol 9.0 Methanol 6.4 Lambda is a tricky value to obtain, and its calculation is quite different between engine types [AFO]. There are two modes of operation in an engine, that are Open Loop and Closed Loop. When the engine is first started, the system goes into Open Loop operation. It stays in this mode until the engine ’warms up’, which means until the coolant sensor detects a specified temperature has been reached, or a specific amount of time has elapsed after the engine was started. After that, it enters in Closed Loop operation. There are PIDs in the OBD standard to obtain the lambda correction for each specific group of cylinders (also known as banks) in a specific mode of operation. The Long Term Fuel Trim PID (codes 0107 for bank 1 and 0109 for bank 2) indicates the correction being used by the fuel control system in both open and closed loop modes of operation. The Short Term Fuel Trim PID (codes 0106 for bank 1 and 0108 for bank 2) indicates the correction being used by the closed loop fuel algorithm. In open loop it is more difficult to obtain correct readings. The lambda value obtained ranges between -100 and 100. If the value is 0, the actual air to fuel ratio is stoichiometric, less than zero means a rich condition, bigger than zero is a lean condition. 45 Obtaining Data The OBD protocol has also a PID to obtain the fuel system status (code 0103) to know the system’s mode of operation, so it is possible to choose the lambda value from the corresponding Fuel Trim. Unfortunately, as with many PIDs, this one is not always available. If that is the case, the workaround can be a time solution. About five minutes of operation is an acceptable time for engines to enter Closed Loop. When the lambda correction value is also not available, the stoichiometric air to fuel ratio is used for calculation (lambda is zero). So the main PID now becomes the MAF value. But as with the fuel flow measurement, the manufacturer can choose not to make it available, or there is no sensor placed before the intake manifold, and so it is not possible to measure the instantaneous rate of air entering the combustion chamber. So if MAF is not available, another formula to calculate the mass of air is needed. By looking at the Ideal Gas Law, a relation is shown with the amount of air to its pressure, volume and temperature. An ideal gas obeys the relationship below. P∗V=n∗R∗T(5.7) Where Pis pressure in Pascal, Vis volume in m3,nis the number of moles, Ris the ideal gas constant (J/K*mol) and Tis the absolute temperature in Kelvin. The number of moles can also be translated by the following formula. n=m MM (5.8) Where nis the number of moles, mthe mass of the gas in grams and MM the molar mass in grams per mol. To obtain the mass of air, the formula can be rearranged as shown below. mair =P∗V R∗T∗MMair (5.9) Where mair is the mass of the air in grams and MMair is the molar mass of air in grams/mol. As stated before in this section, the algorithm is aimed at four-stroke engine types. The first piston stroke is the intake stroke were a mixture of fuel and air, in the case of petrol engines, and just air in a diesel engine (as fuel is injected afterwards), is forced by pressure into the cylinder through the intake port. The intake valve then closes and the compression stroke takes place. The volume of air and fuel mixture in the cylinder relative to the maximum volume of the cylinder is called the engine’s volumetric efficiency. 46 Obtaining Data In the combustion chamber system, the pressure (P) inside the cylinder is a product of the pressure in the manifold and volumetric efficiency. The intake manifold absolute pressure (code 010B) and the intake air temperature (code 010F) are available in the OBD protocol. P=MAP ∗V E (5.10) Where MAP is the intake manifold absolute pressure and VE the volumetric efficiency. Considering that the volume (V) is known from the displacement of the engine, it is now possible to calculate the mass of air (m) in the cylinder proportional to the number of moles of air. mair =(MAP ∗V E)∗ED R∗(IAT +273.15)∗MMair (5.11) The manifold absolute pressure (MAP) is in kilopascal, the intake air temperature (IAT) is in Celsius, and ED is the cylinder displacement in litres. Knowing the mass of the air, the next step is to determine the rate at which it enters the combustion chamber. The OBD standards states that the air mass per intake stroke is a function of revolutions per minute (RPM), and the maximum air mass is a constant for a given cylinder swept volume. AM =AMT (RPM/60)∗(S∗C)(5.12) AMT=ρ∗ED (5.13) Where AM is the air mass in grams per intake stroke, AMTis the total engine air mass in grams per second, RPM is revolutions per minute, Sis the strokes per revolution and C the number of cylinders. The ρsymbol represents the density of air in grams per intake stroke. The air mass is used to calculate the Absolute Load (code 0143), present in the OBD standard, that is the normalised value of air mass per intake stroke displayed as a percent. This translates to the current air mass and pressure divided by the maximum air mass at standard temperature (25 oC) and atmospheric pressure (101.325 kPa), and at 100% volumetric efficiency at wide open throttle. Absolute Load can then be calculated with the formula below. 47 Obtaining Data LOADABS =AM AMT (5.14) Where LOADABS is the absolute load. Absolute Load is an indicator of the pumping efficiency of the engine, and at the peak value correlates with volumetric efficiency. A four-stroke engine has 2 revolutions per intake stroke. Looking again at the Ideal Gas Law, the rate of air considering the maximum air mass at 100% volumetric efficiency and at standard temperature and pressure can then be calculated. MAF =P∗ED R∗(T+273.15)∗MMair ∗LOADABS ∗(RPM/60)/2 (5.15) Where MAF is mass air flow grams per second, Pis the standard atmospheric pressure in kilopascal and Tis the standard air temperature converted to Celsius. Spark ignition engines are required to support absolute load, while compression ignition engines (like diesel) are not required to do so. So compression engines will not be able to use this formula. If that is the case, it is possible to use the same formula considering the intake manifold absolute pressure (MAP) and the intake air temperature (IAT), but volumetric efficiency (VE) becomes an incognito. MAF =(MAP ∗VE)∗ED R∗(IAT +273.15)∗MMair ∗(RPM/60)/2 (5.16) Volumetric efficiency is the measurement of how close the actual volumetric flow rate is to the theoretical volumetric flow rate. It is very difficult for an engine to use the full volume of a cylinder since there can be friction losses or leaks, or other factors [PDP+12]. The theoretical volumetric flow rate is a function of RPM and engine displacement at full volumetric efficiency (100%). The actual volumetric flow rate is a function of the mass flow rate and the density of the intake air. Even knowing that volumetric efficiency depends on RPM, it is not guaranteed to be the same among different vehicles. So since not enough data is supplied for the calculations, the approach used was to set volumetric efficiency at a value of 75% when unknown, which is a reasonable value for most common engines [Iri10]. There is a PID for fuel type (code 0151) but it is rarely available, so only engine displacement, and fuel type if not available, must be obtained from the user, while all the other values are extracted by the OBD. So to calculate fuel consumption, a series of conditional steps are performed. First the 48 Obtaining Data algorithm tests if speed is available, since with no speed it is not possible to obtain litres per 100 kilometres. Then checks if fuel flow is available. If it is, then the formula is direct (Formula 5.3), if not, it is calculated through Formula 5.4. In this case the algorithm first searches for a Lambda value to correct the stoichiometric air to fuel ratio constant. If no lambda is available then the stoichiometric value is used. Next the algorithm searches for a MAF value. If it is available then the fuel flow can be calculated (Formula 5.4), but first, it checks if the fuel type is diesel to correct the mass air flow with Formula 5.5. If no MAF is provided by the OBD, it has to be calculated. So now there are two ways to calculate MAF: Equation 5.15 and 5.16 and they both need RPM. So in this step, if RPM is not available, then fuel consumption cannot be calculated. If it is available, then the algorithm checks if Absolute Load is available to calculate MAF using Formula 5.15. If Absolute Load is not available, then Equation 5.16 is used with MAP and IAT if both are available. Otherwise, fuel consumption cannot be calculated. If MAF can be calculated with one of these equations, the algorithm checks again if the vehicle is a diesel type to adjust MAF with Formula 5.5. So now it is possible to calculate fuel flow with Formula 5.4 and finally fuel consumption (Formula 5.3). 5.1.4 MyDrivingDroid The MyDrivingDroid project developed at Instituto de Telecomunicações is aimed at data gathering taking advantage of an Android Smartphone’s embedded sensors. In order to gain acceptance by a wide audience in the general public, MDD was required to run on a wide range of smartphones, as mentioned in Chapter 4. Also, the interface was designed to be minimal, intuitive and easy. Enhancing user experience and validating a stable implementation on multiple smartphones constituted a significant time effort. The tested smartphones can be seen on Table 5.2. In respect to the interface (see Figure 5.8), in the beginning of each trip a user presses the play button to start logging and the stop button at the end. While the application is logging data, it is possible to close it and use the smartphone for other daily purposes, as it runs in the background. As seen in Chapter 2, MDD is also constituted of a server that receives and processes data. When the user chooses to synchronize with the server, the data is sent to it and then deleted from the smartphone. This section explains how the OBD data gathering module was implemented by explaining first the project’s architecture and the problems encountered, as well as the improvements made in user experience. Since the latter was very important to convince people to use the application as referred, there was an interest in perfecting the interaction and automating repetitive functions, which also turned out to be a time consuming effort. 49 Obtaining Data Figure 5.8: MDD Interface 5.1.4.1 Architecture The project architecture was designed thinking in expansionism. This is explained with the aid of Figure 5.9. Figure 5.9: MDD Architecture 50 Obtaining Data The application consists of the Main Activity and Main Service for interface and background work respectively, sensor managers for each sensor, a Memory Manager for data handling, an Internet Manager for server communication and a SQLite Manager for SQLite database handling. Each manager has its handler associated with its thread message queue so that all requests are serialized and no data is lost. When the application runs, the first thing it executes is a base class that extends the Android Application super class whose function is to maintain a global application state. Then the Main Activity is invoked showing the interface and, in turn, creates and binds with the Main Service. The reason for this is because the Main Activity is designed to process only the user interface, leaving the background processing to other services, so it does not interfere with the user interaction. The Main Service is always running, so even when the interface is hidden or destroyed, this service can still do its work in the background and hold onto other services. The Main Activity and the Main Service communicate between them so that they remain synchronized and updated. The Main Service is responsible for the Memory Manager. This class is initialized as soon as the service is created. It runs on a separate thread with a handler that receives requests, enqueueing and processing them. When sensor data is available, it is sent to the Memory Manager handler queue, and after a while or when the buffer is full, the Memory Manager flushes this data to the SQLite Manager and Internet Manager. Each sensor has its own service and each runs on a different thread. A sensor service is only responsible for gathering data and sending it to the Memory Manager. So each sensor service can be killed manually at runtime, without harming the application. SQLite Manager is a different thread that receives data, converts it into a single insert statement, and writes it to the SQLite database. Internet Manager is also a different thread that receives data and sends it to the server. The main data objects used are JSONObjects and JSONArrays. When a sensor reads data it creates a new JSONObject and puts the sensor data in it. So in order to expand MDD with the OBD Module, all that is needed is a new service and a new table in the SQLite Manager and the corresponding JSON structure in the Memory Manager. The new table and JSON structure added include the trip id, composed by some digits from the smartphone’s IMEI and the timestamp of the trip’s starting time, imported from the trip table that holds the various trips performed, an OBD command code with the mode of operation and the parameter id, the corresponding response value and a timestamp from when the response was recorded. The timestamps are obtained from the smartphones clock in the UTC standard. MyDrivingDroid was tested in various devices as stated previously. The varying behaviour in different smartphones lead to constant adjustments. One of the most significant 51 Obtaining Data Bluetooth OBD devices, in the rest of the smartphones and vehicles presented in Table 5.2 the application is perfectly functional. 5.2 Data Collection Methodology The data gathering process involved various collaborators with different smartphones and different vehicles. This section will address how the collaborators were gathered and chosen and the reasons behind the vehicle’s selection process. 5.2.1 Managing Collaborators To convince and manage collaborators, a website was firstly created 9. The purpose of this virtual space is to provide instructions, background and detail the advantages of the OBD technology, provide a manual for the MDD application as well as a download link, and discuss the thesis in overall. There is also a section reserved for collaborators. Anyone can register provided that an email and some vehicle specifications are submitted. The minimum vehicle specifications are engine size and fuel. Resorting to the dynamic e-mail service of FEUP, an e-mail was sent to the academic community asking for volunteers to contribute with vehicle and sensor data. From the 23 people interested in collaborating, very few had the requirements for this project. After registering, an interview was conducted to known more about the volunteer’s trip time and to conduct a vehicle inspection in order to find if it was OBD compatible. There was a challenge in finding the OBD socket in each vehicle. As referred before, an OBDII connector as to be located in the vehicle’s cabin within approximately half a metre of the steering wheel. But its exact location varied with each vehicle model. Most of the vehicles of the volunteers were diesel type engines, but unfortunately were pre 2004. In the United States, the OBDII standard exists in all vehicles since 1996, but for the European diesel engines, it was only implemented in 2004. Still, at first, there were not enough OBD devices available for the number of drivers. So each driver gathered data until it reached at least 5 hours of total trip time, then the OBD switched to another driver. As new OBD devices became available, the drivers kept the OBDs. Each driver signed an informed consent to allow the use of their gathered data for investigation purposes. 5.2.2 Data Selection Even after the gathering process, some people had very short trips. In these cases the amount of data gathered resulted in vague scatter plots. Figure 5.11 represents the total number of points contributed by each vehicle where fuel consumption calculation 9http://www.fce.pt.vu/ 58 Obtaining Data was possible, and the vehicle identification is shown on Table 5.3 along with its engine displacement and fuel type. Figure 5.11: Number Of Points For Each Vehicle Table 5.3: Vehicle Identification Vehicle Id Model Displacement Fuel Type V1 Audi A4 1896 cm3Diesel V14 Peugeot 207 1560 cm3Diesel V23 Volkswagen Golf 6 1598 cm3Diesel V42 Renault Clio 1461 cm3Diesel V43 Renault Twingo 1461 cm3Diesel V45 Fiat Punto 1248 cm3Diesel V48 BMW 525d 1995 cm3Diesel V49 BMW 320d 1995 cm3Diesel In data analysis, it is important to recognize that there are a number of inherent limitations on what is possible to learn. One of these limitations is related to the finite amount of data gathered. One of the main practical consequence is the necessity of imposing some working assumptions in the analysis. From the vehicles shown in Table 5.3, only the ones with 30000 or more consumption points proved to have enough data for analysis. So only 59 Obtaining Data the vehicles V1, V14, V23, V42 and V45 were used. Appendix Ashows an evolution of the aggregated mean of speed versus fuel consumption (Figures A.1,A.2,A.3,A.4,A.5) showing a decreasing dispersion of points as expected. 5.3 Processing Data This section discusses the methods used to derive values from the GPS, how those values were used and why an interpolation was applied in the OBD data. Also, it discusses the gathered data, how it was validated and the methods used to analyse and filter it. It also shows that, by means of an exploratory analysis, it was possible to find a structure in the dataset. This chapter also shows a comparative analysis of different results provided by different vehicles, a curious result found between system time and GPS time, and a justification why only system time is used to relate the obtained values. 5.3.1 Deriving Acceleration And Steepness As referred is Chapter 2, with a GPS receiver it is possible to obtain time, latitude, longitude, altitude and azimuth. Through these values and the Doppler effect, using navigation equations, it is possible to know the location, travelled distance and speed. Using speed and time, acceleration can be derived, and using altitude and distance travelled, steepness can be also derived. The information is gathered at a sample second. To obtain both acceleration and steepness a least square error method was used. For a given set of data, the least squares method obtains the best-fit curve that produces the minimum sum of the squared deviations. Since velocity obtained from the GPS has a very low error, as proved in plots 5.14 and 5.15, the acceleration in each second was derived from a least square error curve between just the 3 surrounding speeds. To derive steepness, distance travelled and altitude is needed. Instead of using locations to calculate distance (like the Inverse formula [Vin75] used by Android), speed is used because of its superior precision [KH97] [Cha09]. As the speed is in metres per second and the data is obtained at a sample of one second, the sum of speeds gives distance travelled. While the acceleration curve is time limited, the inclination curve is spatially limited. This means that the number of points used on the inclination least square algorithm varies, contrary to the 3 fixed points in acceleration. From a certain point, to calculate inclination, the number of points used satisfies a minimum radius from the original position and a minimum of 5 points required. This means that the radius gradually increases if the number of obtained points does not meet the required minimum. The Figure 5.12 represents speed versus number of steepness points to help illustrate this point. 60 Obtaining Data Figure 5.12: Speed versus Steepness points So at lower speeds the radius is extended to capture at least 5 points. In lower speeds, the radius is less extended than in highways. At higher speeds, the driver is most likely in a highway and the steepness variations are more easily captured with less points. In the Appendix Ais the source code used to derive acceleration and steepness. 5.3.2 OBD Data Interpolation As referred before, vehicles have different response rates, which means that for a given set of OBD commands it is not guaranteed that all the responses are available in all seconds. Also there is some delay in vehicle response time and OBD response time. For example, a vehicle with an average response rate of 4 codes per second (like vehicle V1), sends a request at millisecond 900. The vehicle takes more than 200 milliseconds to respond which means the smartphone will receive the response on the next second. In order to also prevent seconds in the middle of the data with no fuel consumption values, a simple solution was implemented which consisted in linearly interpolating the OBD response values with the GPS timestamps. Values were only interpolated in empty seconds if they were not more than 2 seconds apart from a valid data point. This means that large gaps 61 Obtaining Data between data, where no information was available, were not used, since interpolating this gaps possibly resulted in bad data accuracy. 5.3.3 System Time And GPS Time In the state of the art (Chapter 2) it is referred that there are no guarantees about the smartphone’s system timestamp reliability as proved in Figure 5.13. If it is either a cheap crystal oscillator or a firmware bug cannot be stated for sure, and it also falls out of the scope of this work. However, some smartphones have the possibility of synchronizing with the operator’s clock. Still, in the interval between synchronizations, the clock either suffers a delay or hastens. Figure 5.13: Delay suffered in System Time minus GPS Time Some Android smartphones also use the GPS atomic clocks, which are incredibly accurate, to adjust the time. Although occasionally a leap second adjustment is applied to the Coordinated Universal Time (UTC) in order to keep its time close to the mean solar time, and although the GPS signal includes the offset difference between GPS time and UTC, some Android smartphones do not compensate this offset, and so this leads to some discrepancy on their GPS time. This lead to the workaround with the raw NMEA messages [NME02]. This bug has been debated for some time now in some question-and- answer computer programming or project hosting websites like the Android project page 10. Although not directly related to this thesis proposition, this was an interesting finding. So in order to relate the data gathered between various sensors, namely the OBD and GPS data, the UTC timestamp was used to relate between the various sources. But if at some point an accurate time for the data is needed, that data timestamp can be related to the correct GPS given time. Namely in the OBD and GPS synchronization for this work, 10htt p ://code.google.com/p/android/issues/detail?id =5485 62 Obtaining Data the OBD data was interpolated onto the GPS timestamp, resulting in a more precisely synchronized dataset. 5.3.4 Calibration It is important to validate the GPS values and the fuel consumption formula. To validate the GPS values, the GPS speed was plotted against the vehicle speed read by the OBD. The Figure 5.14 with all the gathered GPS versus OBD velocity points and plot 5.15 with the aggregated mean, show the relationship between these variables obeys a linear model. It also shows that the error decreases as velocity increases. The disperse points in plot 5.14 are usually caused by abrupt variations in speed, like quickly braking. Figure 5.14: GPS Speed versus OBD Speed in metres per second In order to validate the fuel consumption formula, some vehicles (like V1 and V14) offered a visual interface with fuel consumption estimative values that were used to compare with the calculated values from the formula. In order to do this, an event trigger was also implemented in MyDrivingDroid to record an event description and timestamp. So at the beginning of a trip a user defines an event, like ’Consumption at 6l/100km’. While driving, every time the vehicle’s interface showed 6 litres per 100 kilometres the driver quickly tapped the smartphone’s screen and an event was recorded. So after recording 63 Obtaining Data Figure 5.15: GPS Speed versus OBD Speed (Mean) in metres per second some events, the vehicle’s on-board computer values and calculated values from the formula can be compared using the timestamps from the event and the timestamps from the calculated value. Some of the results of the calibration process can be seen on Table 5.4. Table 5.4: Calculated Versus Shown Values Sample Calculated Shown 8.9 8.9 9.6 8.9 8.7 8.9 8.6 8.9 8.9 8.9 8.5 8.9 Average Std. Dev. 8.9 0.4 Calculated Shown 7.4 7.5 7.4 7.5 7.3 7.5 7.4 7.5 7.4 7.5 7.7 7.5 Average Std. Dev. 7.4 0.0 Calculated Shown 5.3 6.0 6.3 6.0 5.1 6.0 5.4 6.0 5.9 6.0 5.4 6.0 Average Std. Dev. 5.6 0.5 5.3.5 Data Relationships The units used for the variables are litres per 100 kilometres for fuel consumption, metres per second for speed, metres per second squared for acceleration and degrees for inclination. In order to comprehend the relationship of fuel consumption with speed, acceleration and inclination, each was individually plotted. The results, considering all 64 Obtaining Data points, are presented below (see Figures 5.16,5.19,5.22). In each relation, the aggregated mean was also calculated and is shown below (see Figures 5.17,5.20,5.23), along with the corresponding histogram for all points (see Figures 5.18,5.21,5.24). Figure 5.16: Inclination versus Fuel Consumption The plotted means helped in detecting where dispersion occurs. In the inclination plot there are scattered points (see Figures 5.17 and 5.18) below -3 and above 3 degrees, in the acceleration plot (see Figures 5.20 and 5.21) below -1 and above 1, and in the speed plot (see Figures 5.23 and 5.24) a dispersion is detected above 40 metres per second. The reason behind this, is because most of the roads where data was gathered do not have an inclination above 3 degrees, vehicles commonly do not suffer an acceleration bigger than 1 metre per second squared, and there are very few points above 40 metres per second (>140 Km/h). So in overall there are very few points in these areas and they represent less than 10% of the data. This is shown by the histograms and Table 5.5. Table 5.5: Percentile of Data Points Data Points Percent Inclination above 3◦9.61% Inclination below -3◦8.81% Acceleration above 1m/s22.39% Acceleration below -1m/s23.86% These undesired points are considered outliers. An outlier is an anomalous data point inconsistent with most of the rest of the data’s behaviour. Boxplots and histograms are a 65 Obtaining Data Figure 5.17: Inclination versus Fuel Consumption (Mean) Figure 5.18: Inclination Histogram helpful technique to detect outliers, or as Ronald Pearson calls them in his book Exploring Data, ’distributional monsters that lurk in the data’ [Pea11]. He also states that the problem of finding relationships between variables is that, in general, no ’correct answer’ 66 Obtaining Data Figure 5.19: Acceleration versus Fuel Consumption Figure 5.20: Acceleration versus Fuel Consumption (Mean) exists to the questions we pose, as they commonly are approximations instead of a direct and precise, unambiguous mathematical solutions. Looking at these plots, it is shown that the dependence of fuel consumption on the independent variables speed, acceleration 67 Obtaining Data Figure 5.30: Speed versus Fuel Consumption (Mean) for different vehicles Figure 5.31: Inclination versus Acceleration (Mean) is uncoupled to the powertrain (no shift is engaged) or when the throttle pedal is not pressed, the vehicle is at idle speed. In this scenario, the engine is capable of generating 74 Obtaining Data Figure 5.32: Speed versus Fuel Consumption from Vehicle V42 enough power to run, but not enough to increase rotation speed. So when acceleration is below zero in diesel engines, if the vehicle is not braking, it is probably idling and the consumption effectively approximates zero. Figure 5.33: Fuel Consumption Histogram Most of the data from vehicle V42 was gathered in a highway scenario as speed points concentration is at higher velocities as seen in Figure 5.34. It also shows a large con- 75 Obtaining Data centration of fuel consumption points below 1 litre per 100 kilometres supporting the previous statement. The large concentration of red points in highway velocities are also an indicator of the average consumption at those speeds. Figure 5.34: 3D Histogram 5.4 Conclusions In conclusion, while the data gathering stage had some setbacks related to applications bugs referred in Section 5.1.4, and there was a low availability of vehicles that met the requirements, a large amount of data was still gathered from a total of 8 vehicles with a total of 347769 consumption points. With this gathered data available, then the previously discussed formulas were used to derive acceleration and inclination and interpolate the OBD data. Some relevant code snippets about these methods can be found in the Appendix A. The data was then validated as it is explained in Section 5.3.4. The outliers proved to have a high influence on correlation. In the case of the aggregated mean acceleration and inclination versus fuel consumption, by removing the outliers, they proved to have a high correlation. While the acceleration and inclination versus fuel consumption can be linearly fitted, in the case of speed versus fuel consumption it was found that it shares a non-linear relationship and can be modelled as a polynomial function as seen in the next chapter (Chapter 6). It is also hinted in this chapter about the effect the gearbox can have on fuel consumption. 76 Chapter 6 Results This chapter describes the techniques and methods used to fit models to the gathered dataset, and provide a comparative analysis of said models. It also discusses in the first section why regression was chosen instead of other available methods. 6.1 Algorithms for Fuel Consumption Estimation There are various approaches to construct the proposed solution. What was needed was to discover the relationship between OBD and GPS data, so that the OBD device could be discarded. The first approach thought was a Machine Learning Algorithm, so this topic was researched. However, due to its simplicity relatively to Machine Learning algorithms, after analysing the dataset, the conclusion was that methods like regression models would also be worthwhile to explore. Some Machine Learning algorithms use regression techniques, so this part of the research was helpful in taking the first steps into regression analysis. Machine Learning does not necessarily mean a complex algorithm. Where there are large data sets, learning algorithms can be a good approach as they can more easily understand and deal with those sets. The two main types of learning algorithms are supervised learning and unsupervised learning. Since it was intended that the proposed algorithm uses the OBD dataset as a learning example, supervised learning was the first though solution. One of the approaches on supervised learning is regression. With regression it is possible to estimate the relationship among variables. To better understand how regression works, it is shown below an example applied to this work’s problem. As it was detailed previously, a fuel consumption model can have a lot of inputs, like vehicle parameters, velocity, acceleration and inclination. In this 77 Results example, the graph represented in Figure 6.1 plots only the vehicle speed obtained from a GPS against the average fuel-consumption obtained from the OBD, obtained from a diesel vehicle in a trip. A plot showing the same behaviour is presented in the paper A Mobile Sensing Architecture for Massive Urban Scanning [RAV+11]. Figure 6.1: Fuel Efficiency Now lets say through the use of only an Android and its sensors, new velocity values are given. This is where regression appears. If a straight line is put through the data, it is possible to locate where the new values fit in that line as in Figure 6.2. However, this might not be the best option. So instead of sending a straight line to the data, a polynomial function might be a better fit as it is possible to observe in Figure 6.3. The regression method is used to predict a continuous value output. With regression it is possible to use very complex polynomial functions, as a high enough degree polynomial can fit any dataset. However, using high degree polynomial functions does not mean better results. Where simple linear functions may cause ’underfitting’, high degree polynomial functions may cause ’overfitting’. With ’overfitting’, the learned hypothesis may fit the training set very well, but will fail to generalize to new examples. 78 Results Figure 6.2: Fuel Efficiency with Simple Linear Regression Figure 6.3: Fuel Efficiency with Polynomial Regression 6.2 Regression Exploratory data analysis has to be dealt with carefully at the risk of leading to the wrong conclusions, as with the killer potato problem based on historical data [CGH+06] described by Ronald Pearson [Pea11]. For a careful analysis Velleman and Hoaglin offer a guide constituted of four fundamental steps known as ’the four R’s of exploratory analysis’:Revelation, Residuals, Reexpression and Resistance [VHM91]. The first R 79 Results was approached in the previous chapter (Chapter 5) as it consists of seeing what the data reveals and identifying outliers. Residuals are the difference between the observed values and the predictions obtained from a model. They present an useful aid in identifying a model’s fitness. The third Rsuggests another approach in revealing a structure in a dataset by transforming it (with a logarithmic function for example). And the last R,Resistance, refers to robustness and the impact that changes in the dataset can have on a model. Since the goal is to obtain fuel consumption from GPS values, in the dataset obtained, fuel consumption serves as a dependent or response variable, and GPS velocity, acceleration and inclination as independent or stimulus variables. A linear model is the simplest of models and widely used as it yields simple functional forms. Many important mathematical functions are analytic, which favour linear constitutive relations [Pea11]. There are various methods of fitting a line in a dataset. The most common assumption is an errors-in-model formulation (EM) in which the ordinary least squares (OLS) method is based. From the scatter plots shown in Chapter 5, it is hinted that not all relationships are linear. Although due to their inherent simplicity, linear models were fitted, along with their partial residual plots, that provide a good indication of the nature of the relationship between each variable. So instead of fitting the whole model, the partial regression plots are firstly generated. Because the independent variables are statistically independent, in a linear model, the following formulation for this work’s case can be viewed as the following: FuelConsumption ≈a∗Speed +b∗Acceleration +c∗Inclination +d(6.1) Figures 6.4,6.6 and 6.8 represent the component residual plots, which is the partial residual plot plus the fitted line calculated with a linear model. These plots are helpful to evaluate non-linearity. The fitted lines were calculated using the Ordinary Least Squares method. The residuals and the coefficients of each model are also present in these figures, along with the histogram of the studentized residuals to test non-normality. The normality tests are used to determine whether a dataset follows a normal distribution. Considering that the inclination and acceleration versus fuel consumption residual histograms are close to a normal distribution, the predicted to the observed residual error is predominantly low. So linearity is more notorious in the inclination and acceleration versus fuel consumption plots compared with the speed versus fuel consumption. The formulated hypothesis with all the three independent variables also follows a normal distribution of studentized residuals (see Figure 6.10). Still, it is perceptible, with the provided results, that a linear 80 Results Figure 6.4: Component+Residual Plots, Linear Model Summary And Studentized Residuals for Inclination Versus Fuel Consumption 81 Results Figure 6.5: Component+Residual Plots, Linear Model Summary And Studentized Residuals for Inclination Versus Fuel Consumption (Mean) model might not be the best fit, mostly because of the failure of fitting a straight line with speed as an independent variable. 82 Results Figure 6.6: Component+Residual Plots, Linear Model Summary And Studentized Residuals for Acceleration Versus Fuel Consumption 83 Results Figure 6.13: Studentized Residuals for All Values Versus Fuel Consumption 6.3 Conclusions Various methods were tried in order to find the best fit for the presented dataset. The proposed solution shows a standard residual error of about 3 litres and a half for all vehicles but a normal distribution of studentized residuals. So, as referred, it is better fitted with emission maps than with individual vehicles. This solution is therefore a valid contribution to enhance the current urban monitor. Also, as seen in the previous chapter, the gearbox has an impact on consumption. So this formula is only valid when the engine is not in idle speed. The data was filtered by range, but using some techniques to discover the engaged shift, like a function of RPM and speed, could improve the outcome. 90 Chapter 7 Conclusions In a society largely dependent on a finite amount of fossil fuels, which provokes a potential rise in price of this natural resource, an algorithm for fuel consumption estimation from GPS data provides an innovative solution that aims to supply a new way to inform citizens about their driving behaviour, reducing mobility related costs and improving quality of life. This thesis encompassed a broad range of fields, since just the topic of mobility spawns a lot of research. Besides giving a background of vehicle mechanics and OBD history and usefulness, this document discussed fuel consumption models, applications that use OBD, mobile and collaborative sensing platforms, UI patterns, Machine Learning methods, and focused on the GPS and OBD technology and algorithms for fuel consumption estimation using GPS and OBD sensor data. It was possible to conclude that some OBD applications provided useful information for OBD data handling and, alongside the analysed emission models, a valuable guideline for the implementation of the fuel consumption algorithm from OBD data. The ELM327 manual also proved fundamental for the implementation and comprehension of the OBD communication protocols. The models also allowed a better understanding of emissions and vehicle dynamics, led to the learning of products and projects that use them, and allowed to understand their importance and impact on a society that highly depends on mobility. The experience with the Android platform also proved to be challenging but fruitful, as the barriers that emerged resulted in research about the components like Bluetooth, and even different smartphones, and a closer look into the Android source code, which enabled an even bigger comprehension of this platform. Since there is a lot of focus on the GPS, there was also a lot of research into the space technology and its navigation and localization methods. 91 Conclusions The result was a regression model that provides fuel consumption in real time, that implemented in a data gathering mobile application like MyDrivingDroid, and used in a collaborative fashion, can generate even more data. The data mining phase resulted in a lot of stored data. In a global perspective, if users abide the previous proposition, on a wide scale, new researches can also be created because a large scale data mining generates a lot of information. And information is power. Information is a currency of today’s world. And the sharing of information generates new opportunities as it is noted in the mobile and collaborative sensing platforms topic. Not only can users compare their driving behaviour, but this can also be a potential tool in creating fuel friendly driving patterns, possibly recommend alternative mobility solutions, or even generate wide scale emission maps. 92 References [AB03] Rahmi Akçelik and Mark Besley. Operating cost, fuel consumption, and emission models in aaSIDRA and aaMOTION. 25th Conference of Australian Institutes of Transport Research, (December 2003):3–5, 2003. [ABH12] Shahid Ayub, A Bahraminisaab, and B Honary. A Sensor Fusion Method for Smart phone Orientation Estimation. cms.livjm.ac.uk, 2012. [AFO] Adriano Alessandrini, Francesco Filippi, and Fernando Ortenzi. Consumption calculation of vehicles using OBD data. epa.gov. [ALM11] Jong Hoon Ahnn, Uichin Lee, and Hyun Jin Moon. GeoServ: A Distributed Urban Sensing Platform. 2011 11th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, pages 164–173, May 2011. [AR03] Kyoungho Ahn and H Rakha. VT-MICRO FRAMEWORK FOR MODELING OF HIGH EMITTING VEHICLES. E-print Network, (540), 2003. [ARTV02] Kyoungho Ahn, Hesham Rakha, Antonio Trani, and Michel Van Aerde. Estimating Vehicle Fuel Consumption and Emissions based on Instantaneous Speed and Acceleration Levels. Journal of Transportation Engineering, 128(2):182–190, March 2002. [BAY+06] M Barth, F An, T Younglove, G Scora, and C Levine. Comprehensive Modal Emissions Model (CMEM). Center for Environmental Research & Technology, (June), 2006. [Bor07] Kanok Boriboonsomsin. Evaluating Air Quality Benefits of Freeway High- Occupancy Vehicle Lanes in Southern California. Transportation Research Record, (951):137–147, 2007. [CCL10] Francesco Calabrese, Massimo Colonna, and Piero Lovisolo. Real-time urban monitoring using cell phones: A case study in Rome. IEEE Transactions on Intelligent Transportation Systems, 12(1):141–151, March 2010. [CCN02] Alessandra Cappiello, I Chabini, and EK Nam. A statistical model of vehicle emissions and fuel consumption. Intelligent Transportation Systems, 2002. Proceedings. The IEEE 5th International Conference, (617):1–25, 2002. [CGH+06] S. B. Carter, S. S. Gartner, M. R. Haines, A. L. Olmstead, R. Sutch, and G. Wright. Historical Statistics of the United States: Earliest Times to the Present. Historical Statistics of the United States: Earliest Times to the Present: Millennial Editio. Cambridge University Press, 2006. 93 REFERENCES [Cha09] Tom Chalko. Estimating Accuracy of GPS Doppler Speed Measurement using Speed Dilution of Precision (SDOP) parameter. 2009. [CQSK11] T Camacho, F Quintal, Michelle Scott, and V Kostakos. Towards Egocentric Fuel Efficiency Feedback. Proceedings of the workshop on PINC: Persuasion, Influence, Nudge & Coercion through mobile devices at ACM CHI 2011, Vancouver, Canada, pages 6–8, 2011. [DR02] Yonglian Ding and HESHAM RAKHA. Trip-based explanatory variables for estimating vehicle fuel consumption and emission rates. Water, Air, & Soil Pollution: Focus, 2(5-6):61–77, 2002. [ELM11] ELM327 OBD to RS232 Interpreter. Electronics, pages 1–51, 2011. [FD12] Michel Ferreira and Pedro M. D’Orey. On the Impact of Virtual Traffic Lights on Carbon Emissions Mitigation. IEEE Transactions on Intelligent Transportation Systems, 13(1):284–295, March 2012. [Fen07] Chunxia Feng. Transit bus load-based modal emission rate model development. Georgia Institute of Technology, (July), 2007. [Gom12] Maria Gomez. Dangers of driving and using your cell phone. Medical News Today, 2012. [GT08] Judith Gebauer and Y Tang. User requirements of mobile technology: results from a content analysis of user reviews. Information-Knowledge-Systems Management - Enterprise Mobility: Applications, Technologes and Strategies, 7(1,2):101–119, 2008. [Haa09] Hein De Haas. Mobility and Human Development. Oxford: International Migration Institute, University of Oxford., 2009. [HBZC06] Bret Hull, Vladimir Bychkovsky, Yang Zhang, and Kevin Chen. CarTel: a distributed mobile sensor computing system. Proceedings of the 4th international conference on Embedded networked sensor systems, pages 125–138, 2006. [HD09] H.L. Hwang and S.C. Davis. Off-Highway Gasoline Consumption Estimation Models Used in the Federal Highway Administration Attribution and Process. ORNL/TM-2009/222, 2009. [HDY+05] Tao Huai, S Thomas D Durbin, Ted Younglove, George Scora, Matthew Barth, and Joseph M Norbeck. Vehicle specific power approach to estimating on-road NH3 emissions from light-duty vehicles. Environmental science & technology, 39(24):9595–600, December 2005. [IBL06] Luc Int Panis, Steven Broekx, and Ronghui Liu. Modelling instantaneous traffic emission and the influence of traffic speed limits. The Science of the total environment, 371(1-3):270–85, December 2006. [Iri10] Adrian Irimescu. Study of Volumetric Efficiency for Spark Ignition Engines Using Alternative Fuels. (2):149–154, 2010. 94 REFERENCES [ISO] ISO 15031-5. http://www.iso.org/iso/home/store/catalogue_ tc/catalogue_detail.htm?csnumber=50816. [KGG07] Ryan Keefe, James Griffin, and John D. Graham. The Benefits and Costs of New Fuels and Engines for Cars and Light Trucks. Santa Monica, CA: RAND Corporation, 2007. [KH97] ED Kaplan and CJ Hegarty. Understanding GPS. Artech House, 1997. [Li12] Chunxiao Li. An Open Traffic Light Control Model for Reducing Vehicles CO2 Emissions Based on ETC Vehicles. Vehicular Technology, IEEE Transactions, 61(1):97–110, 2012. [Lit03] RG Little. Toward more robust infrastructure: observations on improving the resilience and reliability of critical systems. System Sciences, 2003. Proceedings of the 36th, 2003. [LLL+12] Kun Li, M Lu, Fenglong Lu, Q Lv, and L Shang. Personalized Driving Behavior Monitoring and Analysis for Emerging Hybrid Vehicles. Pervasive 2012: Proceedings of the 10th International Conference on Pervasive Computing, 2012. [LML10] ND Lane, Emiliano Miluzzo, and Hong Lu. A survey of mobile phone sensing. Communications Magazine, IEEE, (September):140–150, 2010. [McC99] PM McClintock. Vehicle specific power: a useful parameter for remote sensing and emission studies. Vehicle Emissions, 1999. [NME02] NMEA 0183 Standard For Interfacing Marine Electronic Devices. 2002. [Ono04] Shigeru Onoda. PSIM-based modeling of automotive power systems: conventional, electric, and hybrid electric vehicles. Vehicular Technology, IEEE Transactions, 53(2):390–400, 2004. [PDP+12] Radivoje B. PEŠI ´ C, Aleksandar Lj. DAVINI ´ C, Snežana D. PETKOVI ´ C, Dragan S. TARANOVI ´ C, and Danijela M. MILORADOVI ´ C. Aspects of volumetric efficiency measurement for reciprocating engines. 4, 2012. [Pea11] Ronald K Pearson. Exploring data in engineering, the sciences and medicine. New York ; Oxford : Oxford University Press, 2011. Formerly CIP. [PvHD06] Adam Pollard, J von Hafen, and M Dottling. Winner-towards ubiquitous wireless access. Vehicular Technology Conference, 2006. VTC 2006-Spring. IEEE 63rd, 00(c):42–46, 2006. [RA03] Hesham Rakha and Kyoungho Ahn. Comparison of MOBILE5a, MOBILE6, VT-MICRO, and CMEM models for estimating hot-stabilized lightduty gasoline vehicle emissions. Canadian Journal of Civil Engineering, 30(6):1010, 2003. [RA04] Hesham Rakha and Kyoungho Ahn. Development of VT-Micro model for estimating hot stabilized light duty vehicle and truck emissions. Transportation Research Part D: Transport and Environment,, 9(0536):49–74(26), 2004. 95 REFERENCES [RAV+11] Joao G. P. Rodrigues, Ana Aguiar, Fausto Vieira, Joao Barros, and Joao P. Silva Cunha. A mobile sensing architecture for massive urban scanning. 2011 14th International IEEE Conference on Intelligent Transportation Systems (ITSC), pages 1132–1137, October 2011. [RKBG07] S Kahn Ribeiro, S Kobayashi, M Beuthe, and J Gasca. Transport and its infrastructure. Climate Change 2007: Mitigation. Contribution of Working Group III to the Fourth Assessment Report of the IPCC, 2007. [SAE] SAE J1979. http://standards.sae.org/j1979_201202/. [Sak92] RM Sakia. The Box-Cox transformation technique: a review. The statistician, 41(2):169, 1992. [SB11] Marcin Seredynski and Pascal Bouvry. A survey of vehicular-based cooperative traffic information systems. 2011 14th International IEEE Conference on Intelligent Transportation Systems (ITSC), pages 163–168, October 2011. [TBG10] Arvind Thiagarajan, James Biagioni, and T Gerlich. Cooperative transit tracking using smart-phones. Proceedings of the 8th ACM Conference on Embedded Networked Sensor Systems, pages 85–98, 2010. [VHM91] P. F. Velleman, D. C. Hoaglin, and D. S. Moore. Perspectives on Contemporary Statistics. Historical Statistics of the United States: Earliest Times to the Present: Millennial Edition. The Mathematical Association of America, 1991. [Vin75] T. Vincenty. Direct and Inverse Solutions of Geodesics on the Ellipsoid with Application of Nested Equations. XXIII(176), 1975. [Wan09] K Wanglund. Evaluation of GPS Velocity and Altitude Data used for Road Grade Estimation. (June), 2009. [Yue08] H Yue. Mesoscopic fuel consumption and emission modeling. PhD thesis, Virginia Tech, 2008. [ZN01] Y Zheng and D Niemeier. A Grid-Based Mobile Sources Emissions Inventory Model. Institute of Transportation Studies, 94274(18), 2001. [ZZGD06] JAson ZHang, KEfei Zhang, RON GRenfell, and ROD DEakin. On the Relativistic Doppler Effect for Precise Velocity Determination using GPS. 80(DoD 1996):104–110, 2006. 96 Appendix A Appendix A.1 Relevant Figures This section is reserved for some relevant figures. A.1.1 Data Gathering Evolution Figure A.1: Speed Vs Fuel Consumption (Mean) on 20 Trips 97 Appendix Figure A.2: Speed Vs Fuel Consumption (Mean) on 40 Trips Figure A.3: Speed Vs Fuel Consumption (Mean) on 60 Trips 98 Appendix Figure A.4: Speed Vs Fuel Consumption (Mean) on 80 Trips Figure A.5: Speed Vs Fuel Consumption (Mean) on 100 Trips 99 Appendix yi[j] = slope[loc] *xi[j] + intercept[loc]; } // System.out.println("Found between: " + loc + " and: " + (loc+1) + ", x[loc]: " + x[loc] + ", x[loc+1]: " + x[loc+1]); } else { yi[j] = y[loc]; } } } return yi; } 106