Depósito de Investigación de la Universidad de Sevilla https://idus.us.es/ This version of the article has been accepted for publication, after peer review and is subject to Springer Nature’s AM terms of use, but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: https://doi.org/10.1007/s11116-0179771-5
1 Estimating traffic volumes on intercity road locations using roadway attributes, socioeconomic features and other work-related activity characteristics Noelia Caceres1 • Luis Romero2 • Francisco Morales2 • Antonio Reyes2 • Francisco G. Benitez2 Abstract Traffic volume data are key inputs to many applications in highway design and planning. But these data are collected in only a limited number of road locations due to the cost involved. This paper presents an approach for estimating daily and hourly traffic volumes on intercity road locations combining clustering and regression modelling techniques. With the aim of applying the procedure to any road location, it proposes the use of roadway attributes and socioeconomic characteristics of nearby cities as explanatory variables, together with a set of previously discovered patterns with the hourly traffic percent distribution. Test results show that the proposed approach significantly produces accurate estimates of daily volumes for most locations. The accuracy at hourly level is a bit more reduced but, for periods when traffic is significant, more than half of the estimates are within 20% of absolute percentage error. Moreover, the main peak period is approximately identified for most cases. These findings together with its great applicability make this approach attractive for planners when no traffic data are available and an estimate is helpful. Keywords Clustering algorithms, traffic volume estimates, socioeconomic characteristics, workrelated activity, roadway attributes, hourly traffic percent distribution, traffic patterns. ________________________________ Noelia Caceres
[email protected] 1 Transportation Engineering, AICIA, Camino de los Descubrimientos s/n. 41092 Seville, Spain. 2 Transportation Engineering, Faculty of Engineering, University of Seville. Camino de los Descubrimientos s/n, 41092 Seville, Spain.
2 Introduction Traffic volume estimation is a crucial topic in traffic planning and operations. Information on daily volume is used by Road Administrators for highway planning or pavement design. Changes over time are also useful during the road design phase or even for the calibration and validation of travel demand models. These data are collected in only a limited number of road locations, usually main roads, due to budgetary restrictions. For roads where data are unavailable, estimates are traditionally made based on statistical methodologies that can lead to large errors (Gastaldi et al. 2014). Nowadays, mobility is becoming more diverse since population activities are more and more numerous and occupy more space and time. In order to predict traffic volumes that can be expected on the road network during specific periods, cognizance should be taken on the fact that traffic volumes changes considerably at each point in time as a function of local conditions. For example, for intercity traffic, larger cities usually attract more people than smaller ones, and cities closer together tend to have a greater attraction. According to these principles, this paper presents an approach for estimating traffic volumes on intercity road locations based on clustering and regression modelling techniques. The study has explored as explanatory variables not only roadway attributes (e.g. number of lanes, speed limit and distance to major cities), but also indicators that reflect the socioeconomic and other work-related activity characteristics surrounding a road location (e.g. population and employment), selecting the most relevant ones. The main advantage is its applicability; it can be applied to any intercity location, which is especially important where there are no count stations. In the remainder of this paper, a literature review on the subject is provided first. After describing the characteristics of the study area, data used, and analyses performed to derive the explanatory variable, the proposed approach is then described. Finally, the results are discussed and conclusions provided. Literature review Due to the level of uncertainty in traffic forecasts (as reviewed in Jong et al. 2007), over many years researchers have tried to find ways to produce better traffic volume estimates. Advanced methodologies (e.g. neural network, fuzzy logic, regression or clustering techniques) have been applied as a means of estimating traffic volumes using traffic information, such as Annual Traffic Census (Lam et.al 2006), peak volumes or times of the peak hours (Weijermars and van Berkum 2005), or even speed and density data (Azimi and Zhang 2010). However, monitoring is necessary to obtain such traffic information. More recent approaches proposed in the Federal Highway Administration Traffic Monitoring Guide (FHWA 2013) combine the use of automatic traffic recorders (e.g. permanent loop detectors) taken on a small number of road locations to estimate average traffic volume on other locations with short period counts (e.g. using microwave radar). This procedure allows reducing the need for extensive monitoring activities and a large amount of papers have been published with interesting findings. Nevertheless, monitoring is still necessary. The big challenge is to estimate traffic volumes at anywhere else in the road network, especially for roads that do not have detector systems. In this regard, new techniques have recently been developed in order to estimate traffic volume using non-traffic information (e.g. roadway characteristics, land-use indicators, or socio-economic and demographic data). The methods developed involve not only clustering but also other statistical methods such as linear regression. Mohamad et al. (1998) used multiple regression analysis to develop annual average daily traffic (AADT) prediction models for county roads from aggregated data at the county level (population, arterial mileage, location and accessibility), and an R-squared of 0.75 was achieved. Similarly, Xia et al. (1999) specified a model for AADT predictions by using explanatory variables such as functional class, number of lanes, area type, auto ownership, accessibility of non-state roads to other county roads, and nearby employment; the resulting R-squared was 0.60. This R-squared value was improved by Zhao and Chung (2001) analyzing land use and accessibility variables more extensively; and by Zhao and Park (2004) through geographically weighted regression techniques to produce more locally specific model parameters. Anderson et al. (2006) also developed a model with non-traffic explanatory variables, including functional classification, number of lanes, population, employment, and whether the road is a through street or destination street, resulting an R-squared equal to 0.82. With the evolution of spatial analysis techniques using geographic information system (GIS) technology and the progress in data mining, researchers have started to explore other methods that exploit the spatial context of traffic and other supplementary data (Gecchele et al. 2011; Song and Miller 2012). The kriging technique, which exploits the spatial aspect of observations, has been explored by Wang and Kockelman
3 (2009) for mining network and count data, over time and space, using highway count data. The method forecasted AADT values at locations where count stations are unavailable. In this regard, Selby and Kockelman (2011) showed that kriging can reduce average-absolute-error by 16%-79%, depending on the data and model specification used. Neural networks have also been applied to estimate AADT, combined with fuzzy set theory (Gastaldi et al. 2014), or using the effect introduced by the principle of demographic gravitation on independent variables based on land-use characteristics for the AADT estimation (Duddu and Pulugurtha 2013). Although methods of estimating volume by not using traffic data may not be necessarily adequate to meet the needs of engineering design and planning (Gastaldi et al. 2014), they are attractive thanks to its cost effectiveness and applicability. They can be applied to any location in the road network, which is especially important for other practical applications that do not require such a high level of accuracy. Inspired by these advantages, this paper exploits clustering analysis and regression modelling for estimating volumes on locations over road network. Of primary interest is not only the total daily volume of vehicle traffic as is often studied, but rather the changes over time or pattern of hourly volume throughout the day. The approach hypothesizes that the traffic volume at a particular road location varies as a function of local conditions regarding not only roadway features, but also regarding characteristics of nearby cities. The idea that socioeconomic characteristics, activity participation and/or other aspects of life can influence travel behavior has been studied previously in the travel modeling field (Kitamura et al. 1997; Allahviranloo and Recker 2015). The approach proposed in this paper is a continued effort following a previous work (Caceres et al. 2012), in which traffic volumes of a given location were estimated using an attractiveness factor based on the characteristics of nearby areas (population and distance), associating the typical pattern (obtained by clustering) with the location. However, the approach did not work well for groups whose members have small similarity in volume profiles, especially because the clustering algorithm did not lead to find homogeneity within the groups in terms of volume profiles but according to the attractiveness factors. To solve this drawback, the proposed approach takes patterns directly derived from the use of hourly traffic distribution as clustering variable. This fact produces more compact groups than in the previous work, and thus prediction models may significantly deliver better results. Input data Explanatory variables In order to predict traffic volumes, of course, physical roadway attributes at a particular point (e.g. number of lanes, capacity, speed limit, and so on) have direct effects on traffic volumes. However, these attributes do not exhibit enough variability to be useful for prediction (e.g. if most of the road locations have two lanes, then “number of lanes” is an inadequate variable). Then, the explanatory variables need to have variability, and such variability needs to be correlated with traffic volumes. Cities where people live or work exert strong influence over intercity mobility. The more people there are, the more mobility there will be; the more shopping centres there are, the more merchandize will arrive and more transportation will happen. Then, socioeconomic and other work-related activity characteristics (population, degree of economic activity, proportion of jobs, and so on) can reflect the spatial variability surrounding a road location, and thus have been included as a part of the possible variables. To conduct the study, the transport network corresponding to the pilot area is abstracted into a graph model consisting of a set of centroids and links, which represent cities and road sections, respectively. Each centroid contains the attributes with the socioeconomic and other work-related activity characteristics of the city; while the links contains the roadway attributes to be analysed. This section describes the set of explored variables. Roadway Characteristics These data are related to the attributes of the roadway at the particular location k, and include the number of lanes (LANk) and the speed limit (SLk). The functional road class (FRCk) has been incorporated into the analysis. The functional classification is the process by which roadways are grouped into classes according to the character of service they are intended to provide. Four classes have been used in this study: main roads (motorway, freeway, or other major road), secondary roads, local roads and others. The study also considers other variables related to the dispersion of population and activity in cities surrounding the location. The more dispersed activities or services are, the more travel is required to reach them. For this purpose, three
4 more variable are defined. The first one is DCPk, used to measure the distance from the location k to the mean center of population. This mean center of population for the location k (called CPk) is determined by finding the spatial mean center of all cities surrounding the location (within a radius of influence of 15 kilometers) weighted by its population. The Xand Ycoordinates for the CPk are calculated as follows: for each city whose network distance to the location , , is less or equal to 15km k k ii i CP i i ik ii i CP i i PX XPi k d PY YP = = where Xi and Yi are, respectively, the xand ycoordinates of the city i; Pi is the population (in inhabitants) of the city i; and dki is the network distance (by road) between the city i and the location k. So that the cities used in determining the center of population for the location k (CPk) are those within a radius of influence of 15 kilometers centred at the road location k. (The use of such a radius id explained in following paragraphs). Then, the variable DCPk is calculated using bird’s eye distances (in a straight line) between the location k and its mean center of population (CPk), measured in meters by means of a GIS tool. This distance is the only one measured this way (in a straight line), since a center of population is usually located in a point not connected to the road network. Hence, it is not possible to obtain the shortest travel distance by road from the GIS tool. Apart from this, two more variables are defined to measure the distance from the location k to the most populated city (DMPCk) and to the nearest city (DNCk). Both of them are based on the network distance, measured (in meters) as the shortest travel distance by road (obtained by means of a GIS tool) between the location k and the centroid representing the nearest/most populated city. A summary of the roadbased variables involved is listed in Table 1. Table 1. Set of variables based on roadway attributes for location k. Roadway feature Variable Number of lanes at the road location k LANk Speed limited at the road location k (kilometres/hour) SLk Functional road class at the road location k FRCk Distance from the location k to the mean center of population (meters) DCPk Distance from the location k to the most populated city (meters) DMPCk Distance from the location k to the nearest city (meters) DNCk Socioeconomic and other work-related activity characteristics In this regard, this study has used data published yearly by La Caixa Banking Foundation (2011), with abundant statistical information and socioeconomic indicators on each of the municipalities in Spain with more than 1,000 inhabitants (representing approximately 96% of the total national population). For this case, the variables explored are linked to the cities located nearby the road location k, but considering that each city affects the traffic load in a different manner. To model such influence on each road location caused by activities in nearby areas, this paper has used the concept of location-based accessibility. This concept has been widely treated in the field of transport and urban planning (Koenig 1980; Geurs and van Wee 2004; Lopez et al. 2009); and it has already been used to estimate AADT at unmeasured locations getting good results on the field (Zhao and Chung 2001; Zhao and Park 2004). According to it, the interaction between locations declines with any increase in disutility (or cost) between them, that usually depends on travel time or travel distance to a given location. This measurement estimates the accessibility of a location k with respect to opportunities in all surrounding zones, in which smaller and/or more distant opportunities provide diminishing influences. Then, the accessibility of road location k is defined as follows: () k i ki i A D F c= (1) where Di is the opportunity (socioeconomic and other work-related activity feature showed in Table 2) at zone or city i and F(cki) is the deterrence function depending on the generalized cost (cki) of reaching city i from location k. The explored variables based on the characteristics are showed in Table 2. In this study the cost (cki) is modelled by the network distance (dki), so that the deterrence function describes the effect of space. Regarding deterrence functional form, several studies have used different
5 functions, such as power-law, exponential, gaussian or logistic functions. The gravity model is a well-known formulation of the spatial interaction, and specially, for the trip distribution (Erlander and Stewart 1990). Gravity models assume the interaction between two locations is proportional to their importance (e.g. population), but it decays with distance. In gravity models, the general deterrence form follows a power-law function for the interaction between two locations; that is, F( ki c )=1/ ki d where α reflects the effect of space. In this study, the distance dki is based on the network distance and measured as the shortest travel distance (in meters) by road from the location k to the city i, which is obtained by means of a GIS tool. The exponent of the deterrence function is a parameter whose value depends on the system (Barthélemy 2011). In order to determine its value, data derived from a survey carried out by the Spanish Ministry of Development in 20062007 (MOVILIA 2006) was used, which contained information to enable understanding of the daily mobility patterns of Spanish residents. In particular, it offered the number of displacements between a pair of zones together with the average distance traveled in their displacements. Using this empirical data and the population information associated to the zones from La Caixa Banking Foundation (2011), the deterrence function was fitted by a power-law with exponent α=1.52. At last but not the least, it is important to highlight that the calculation of these variables only considers the effect generated by cities i that are within a radius of 15 km from a given road location k; that is the network distance from the city i to the location k is less or equal than 15 kilometres (dki≤15km). This radius of influence has been selected taking into account that the average trip distance in a working day is around 15 km (Eurostat 2007). This simplification is needed to reduce the number of cities considered in the calculation of accessibility indicators, and it is totally coherent with the assumption that nearest cities are responsible for most of the traffic supported by a given road. Based on this distance, none road location was isolated; there was always at least one city within the radius of influence of a given road location. Table 2. Set of variables based on socioeconomic and other work-related activity characteristics for location k. Socioeconomic/work-related activity feature Variable Population (inhabitants) of the city i (Pi) () k i ki i AP P F c= Economic activity index* (per 100,000) in city i (Ei) () k i ki i AE E F c= Unemployment rate (per 100) in city i (Ui) () k i ki i AU U F c= Gross floor area (meter2) of city i (GFAi), () k i ki i AGFA GFA F c= Number of industrial establishments in city i (IEi) () k i ki i AIE IE F c= Number of manufacturing establishments in city i (MEi) () k i ki i AME ME F c= Most-populated city (inhabitants) near location k (MPCk) () kk k MPC kMPC AMPC P F c= Nearest city(inhabitants) (NCk) () kk k NC kNC ANC P F c= (*) Rate of economic activity for each city in Spain based on the business and professional economic activities tax collected. This index measures the municipal participation (X of 100,000 parts) regarding the respective total at national level (100,000 parts). Selection of explanatory variables Once the set of variables available are established, the selection of explanatory variables to be used in regression modelling is required. In statistics, selection procedures exist to objectively choose a subset of explanatory variables such as stepwise regression. In this semi-automated process, the choice of explanatory variables is carried out by successively adding (forward) or removing (backward) variables based on userspecified criteria, which usually takes the form of a sequence of F-tests or t-tests, but other techniques are possible (Blanchet et. al 2008). For this study the forward-stepwise regression algorithm is used to analyze the subset of explanatory variables to be used in regression. The proposed approach implements two different regression models (one for modeling group choice and other for estimating daily traffic) using the same input variables; the choice is performed using a combination from both stepwise regression outputs. The modeler’s judgment also plays a key role in selecting variables because of collinearity issues. Collinearity, or excessive correlation among explanatory variables, is a common problem when estimating regression models. High correlation between variables might lead to inflated standard errors of the estimators (collinearity effects), but the opposite is not always true. Various methods for testing collinearity exist in literature; this study has opted by the Belsley collinearity diagnostics (Belsley et al. 1980), examining the
6 variance inflation factor (VIF) and the condition number (CI) of the correlation matrix among the selected variables, together with the coefficient R-squared. A common rule of thumb requires that R-squared < 0.8, VIF < 5, and CI < 10, for considering collinearity as negligible (Friendly and Kwan 2009). According to the outputs of the forward-stepwise regression analysis on the set of variables together with the collinearity diagnostics, the explanatories variables finally selected are: X1: LANk with the number of lanes at the location k, X2: AEk related to economic activity of cities within the influence radius at location k, X3: AGFAk related to the gross floor area of cities within the influence radius at location k, X4: AMPCk related to the most populated city near location k, X5: DCPk with the distance from the location k to the mean center of population. Fig. 1 exhibits the collinearity diagnostics, in which the largest condition index (CI=6) corresponds to a near linear dependency involving X4 and X5, with a smaller contribution of X2. However, the values of VIF, CI and R-squared for these variables are less than the recommended values, and it is concluded that there is no serious problem of multicollinearity. A summary of descriptive statistics of the selected variables is given in Table 3. However, the variables to be used in the models must be standardized by subtracting the mean and dividing by the standard deviation. Standardization of variables is particularly important when variables are measured on different scales/units; otherwise, they do not contribute equally to the analysis. Fig. 1. (a) Correlation matrix with R-squared. (b) Tableplot of CI, VIF and variance proportions (VP) for the selected explanatory variables. In column 1, the square symbols are scaled relative to a maximum CI of 30. In the remaining columns, variance proportions (×100) are shown as circles scaled relative to a maximum of 100. In the last row, VIF values are exposed. Table 3. Mean, standard deviation, maximum and minimum of selected variables. Mean Standard deviation Max Min X1 2.474 0.905 4 1 X2 56481670.225 850892348.833 12848260720.179 0.0000000002 X3 1079531438482.049 16225891817230.148 245008999577218.560 0.0093230397 X4 276.435 3836.297 57942.250 0.00081 X5 5716.187 3159.373 14463.597 732.277 Traffic Data Description User mobility is closely linked to activities which tend to be routinized on working days. For commuting trips the repeatability gives rise to a stable component in hourly traffic volumes. Most of daily activities depend on country-specific habits (e.g. start of working hours, lunch breaks, etc.) which suggest the existence of traffic distribution patterns by hour of the day. These patterns provide valuable information to predict volumes, especially to detect peak periods occurrence.
7 The empirical traffic data used throughout this work comes from permanent automatic counters (these data are collected 24 hours per day, 365 days per year) located in a broad geographic distribution across the Spanish road network, provided by the Spanish Directorate General of Traffic (DGT 2011). The data offer a wide range of traffic statistics including not only measurements for AADT (in veh/day), but also traffic volume measured on hourly intervals (in veh/hour). This information refers on an average day, distinguishing between an average weekday and weekend. This study has taken only weekday data since this kind of day is more appropriate for estimation purposes. Travel behavior is repeated frequently on weekdays, while trips on weekends respond to non-routinized activities (entertainment, shopping or social purposes). To find such patterns, traffic data observed at a large number of sites are required. In particular, the data used in this study come from 455 road locations with different traffic backgrounds (highways, secondary roads and so on) and characteristics (from a single lane to a multilane, single or dual carriageway). All these locations are placed on intercity roads; urban environments are not considered at this stage of the work. Then, these road locations have been randomly split into two subsets, one used for model calibration and the other for model validation. Therefore, 345 road locations (~¾ total sample) have been used as the calibration dataset, and the remaining 110 locations (~¼ total sample) have examined the predictive accuracy of the proposed approach. Fig. 2 shows the road locations coloured by the corresponding set (red and green triangles for calibrating or testing set, respectively), together with the centroids representing cities. Fig. 2. Map of road locations as well as the centroids representing the cities and the road network (only cities with greater than 1000 inhabitants). Clustering Cluster analysis is a statistical procedure to define groups that share similar characteristics. Its goal is to organise objects into different groups or clusters, such that a group is a collection of objects “similar” to each other and are “dissimilar” to the objects belonging to other groups. There are many ways to combine cases into groups, overviews of clustering procedures can be found in the literature. The most commonly used is the hierarchical clustering method, which basically forms groups by clustering cases into larger groups until all the cases are members of a single group. The criteria for deciding groups are based on either a difference or similarity matrix, where the similarity measures the closeness of cases. Among the common methods of doing this (single linkage, average between-groups linkage, Ward's method...), this research has selected the average within-groups linkage. Using this method, the distance is defined as the average of the distances
8 between all pairs of cases in the group that would result if they were combined. This minimises intra-group distances and thus tends to produce tight groups. Therefore it is appropriate when the purpose of the clustering is the homogeneity within the groups. Accurate clustering requires a precise definition of the closeness between a pair of objects in multi-dimensional space, in terms of either the pair-wise similarity or distance. For the sake of simplicity, the proposed approach has taken Euclidean distance to give a numerical value to the amount of similarity between two objects, but several similarity or distance measures has been proposed and widely applied in literature (Rui and Wunsch 2005). The aim is to discover a set of mobility patterns over a region from observed traffic data. Experience shows that although traffic volumes may change over time, the relative variations of traffic at certain hours of the day are often quite consistent. Therefore, this study has employed the hourly distribution in terms of the percentage of the daily traffic within each hour of the day as object to be classified, over the calibrating set of road locations. One of the main difficulties for cluster analysis lies in the determination of the optimal number of groups present in a dataset. A variety of indices have been defined in literature to evaluate the fitness between a dataset and clustering result where the optimal number of clusters produces best fit (Milligan and Cooper 1985). According to their work, Calinski and Harabasz’s (CH) index is the most effective one in identifying the number of clusters. The CH index evaluates the cluster validity based on the average betweenand withincluster sum of squares (Calinski and Harabasz, 1974), so a large CH index indicates homogenous clustering. Fig. 3 exposes that ten clusters is the solution with the highest CH index value. Fig. 3. Calinski and Harabasz index as a function of number of clusters. Hourly patterns Once the clustering stage is performed, the road locations included in the calibrating dataset are categorized into ten clusters or groups. Each group is represented by the corresponding hourly pattern (Fig. 4) with the distribution of traffic by hour-of-day (in percentage) in a working day. The title of each plot indicates the number of locations within each group, Nj, and the cluster compactness (CCj) measure based on variance in order to quantify how closely related the objects in a group are. Lower variance indicates better compactness (or in other words, members have high mutual similarity). Most of groups have high homogeneity and compactness, except G10. A visual analysis of these patterns reveals that, in general, the percentage that takes place between midnight and 4:00 AM, when the majority of people are resting, is reduced for all patterns. Traffic usually starts increasing in early morning, between 5:00 and 6:00 AM, and drops in the evening, between 7:00 and 8:00 PM. Therefore, most of the traffic is usually concentrated between 7:00 AM and 8:00 PM, though the behaviour during such hours varies from group to group. Some of them (G1, G3, G4, G5 and G6) exhibit distinguishable peaks associated with commuter trips. These are the morning-peak period related to home-to-work trips (07:00–09:00 AM) and the evening-peak period for work-to-home return trips (5:00–8:00 PM). It is worth noting that in Spain school time is concentrated in half-day; moreover many people are employed in split shifts or even in part-time jobs. These facts cause a third peak period around 1:00–3:00 PM that can take some load from the evening-peak, and usually makes the morning-peak period the most intense time of the day. The evening-peak period tends to be spread over a longer duration because work-to-home return trips are usually linked with some other trip purposes. There are also groups in which traffic begins to increase at early morning and continues throughout the rest of the morning and into the afternoon, being the evening-peak slightly higher (G2, G7, G8 and G9). The last group (G10) exhibits a pattern totally different, with a remarkable evening-peak in the distribution of daily traffic. This group is the least compact based on its CC measure, being its pattern probably the least representative among its members.
15 Fig. 8. (a) Hourly volumes observed and estimated (using ADT derived by Model 3) for testing locations during the hour periods within 7AM-8PM; and (b) Percentage of hourly estimates with APE≤20% in such periods. Table 8. Error levels obtained for each group between 7:00 AM and 8:00 PM (using ADT estimated by Model 3). Group G1 G2 G3 G4 G5 G6 G7 G8 G9 G10 MAPE for hourly estimates 14% 23% 27% 24% 24% 26% 46% 22% 26% 91% Proportion of hourly estimates with APE≤20% 72% 63% 57% 57% 48% 52% 57% 50% 55% 30% Next, the results after applying the forecast procedure at two locations are presented and compared with observed data. At daily level (Fig. 9a), the estimates reach reduced percentage error levels. Fig. 9b displays the estimated volumes by hour-of-day (broken line) and the volume profile observed at each location (solid line), revealing that the estimates follow the peaks and valleys of the observed curve for most hour periods. Then, the approach can provide an approximation of the hourly evolution of traffic during a day at a particular road location, very useful when no information is available. In this regard, other important characteristic of the proposed approach is the high accuracy forecast of the main peak period for road locations. At the two locations showed in Fig. 9a, the predicted period of the main peak (6h-period for E288-0 ASC, and 17h-period for E-22-0 DESC) nearly matches with the observed one (7-hour period for E288-0 ASC, and 17-hour period for E-22-0 DESC). The analysis of such deviation for all testing locations (Fig. 10) exposes that the differences between predicted period of the main peak and the real one are less or equal than one hour for the 80% of road locations.
16 Fig. 9. Observed and estimated volumes for each hour period (a) and the total day (b) at two locations. Fig. 10. Deviation in hours between the predicted period of main peak and the observed one. Conclusions Traffic volume data are key inputs appreciated by Road Administrators in traffic planning and operations. But these data are available in only a limited number of road locations due to the cost involved for deploying sensors. The approach presented in this paper aims to estimate traffic volumes at anywhere else in the road network. For this purpose, the approach combines clustering techniques to infer country-specific mobility patterns, together with regression modelling using roadway attributes and socioeconomic and other workrelated activity statistics of cities surrounding a given road location. This study has been applied on a set of road locations, distributed across the Spanish road network, where also exist permanent count stations for calibrating and validating purposes. Test results show that the proposed approach significantly produces accurate estimates of daily volumes for most locations. The accuracy at hourly level is a bit more reduced but, for periods when traffic is significant, more than half of the estimates are within 20% of absolute percentage error, which is the allowed limit for fulfilling standards for loop detectors (Lehnhoff 2004). Moreover, for most cases, the main peak period is approximately identified. The differences between the predicted period and the real one are less or equal than one hour for the 80% of road locations. This information is also a significant input for design and planning purposes that take into account present and future uses of the considered road. The proposed approach allows Road Administrators to build hourly traffic volume predictions at a particular road location when no information is available and an estimate is helpful. These findings make this approach attractive for practical applications that do not require a high level of accuracy. The main advantage of the procedure is its applicability since the procedure can be applied to any intercity road location. The design of roadways is also a field for which the knowledge of hourly and daily traffic volumes plays a key role. With respect to highways, design criteria consist of a detailed list of considerations in which traffic requirements in relation to use as well as changes over time should be evaluated. Further research based on this approach can provide a better understanding of traffic loading that will be supported by new intercity roads. As a future study, the complex urban environment should be also investigated by exploring city-based features (e.g. the presence of schools, bus stops, shopping areas, and so on) for the
17 inference of traffic volumes on urban roads. Acknowledgements One of the authors, N. Caceres, thanks the Ministry of Economy and Competitiveness of the Spanish Government for the funds provided through the Torres Quevedo Programme (PTQ-13-06428). References Allahviranloo, M., Recker, W.: Mining activity pattern trajectories and allocating activities in the network. Transportation 42, 561–579 (2015). Anderson, M., Sharfi, K., Gholston, S.: Direct Demand Forecasting Model for Small urban Communities Using Multiple Linear Regression. Transp. Res. Rec. 1981, 114-117 (2006). Azimi, M., Zhang, Y.: Categorizing Freeway Volume Conditions by Using Clustering Methods. Transp. Res. Rec. 2173, 105-114 (2010). Barthélemy, M.: Spatial networks. Phys. Rep. 499(1–3), 1–101 (2011) Belsley, D. A., Kuh, E., Welsch, R. E.: Regression diagnostics: Identifying influential data and sources of collinearity. New York: John Wiley and Sons (1980). Blanchet, F.G., Legendre, P., Borcard, D.: Forward selection of explanatory variables. Ecology 89, 2623– 2632 (2008). Caceres, N., Romero, L., Benitez, F.G.: Estimating traffic flow profiles according to a relative attractiveness factor. Procedia - Social and Behavioral Sciences 54, 1115–1124 (2012). Calinski, R. B., Harabasz, J.: A dendrite method for cluster analysis. Communications in Statistics 3, 1–27 (1974). DGT, Directorate General of Traffic. Traffic Map 2011. Traffic Department of the Spanish Home Office. Ministry of Public Works of Spain (2011). Duddu, V., Pulugurtha, S.: Principle of Demographic Gravitation to Estimate Annual Average Daily Traffic: Comparison of Statistical and Neural Network Models. J. Transp. Eng. 139(6), 585–595 (2013). Erlander, S., Stewart, N.F.: The GravityModel in Transp. Analysis: Theory and Extensions. VSP (1990). Eurostat: Passenger mobility in Europe, Statistics in Focus. Catalogue number: KS-SF-07-087-EN-N (2007). FHWA: Traffic Monitoring Guide, Federal Highway Administration U.S. Department of Transportation (2013). Friendly, M., Kwan, E. Where's Waldo: Visualizing collinearity diagnostics. The American Statistician 63(1), 56-65 (2009). Gastaldi, M., Gecchele, G., Rossi, R.: Estimation of Annual Average Daily Traffic from one-week traffic counts. A combined ANN-Fuzzy approach.” Transp. Res Part C 47(1), 86–99 (2014). Gecchele, G., Caprini, A., Gastaldi, M., Rossi, R.: Data mining methods for traffic monitoring data analysis. A case study. Procedia Social and Behavioral Sciences 20, 455-464 (2011). Geurs, K.T., van Wee, B.: Accessibility evaluation of land-use and transport strategies: review and research directions. J. Transp. Geogr. 12, 127–140 (2004). Kitamura, R., Mokhtarian, P.L., Laidet, L.: A micro-analysis of land use and travel in five neighborhoods in the San Francisco Bay Area., Transportation 24, 125-158. (1997). Koenig, J.G.: Indicators of urban accessibility: theory and applications. Transportation 9, 145–172 (1980). Jong, G. de, Daly, A., Pieters, M., Miller, S., Plasmeijer, R., Hofman, F.: Uncertainty in traffic forecasts: literature review and new results for the Netherlands. Transportation 34, 375–395 (2007). Kohavi, R., Provost, F.: Special Issue on Applications of Machine Learning and the Knowledge Discovery Process. Machine Learning 30(2/3), 271-274 (1998). La Caixa Banking Foundation: Spain Economic Year Book (2011). Lam, W.H.K., Tang, Y.F., Chan, K.S., Tam, M.L.: Short-term Hourly Traffic Forecasts using Hong Kong Annual Traffic Census. Transportation 33, 291–310 (2006). Lehnhoff, N. Quality of automatic data collection with loop detectors. Proceedings 2nd Int. Symp. Networks for Mobility, Stuttgart, Germany (2004). Lopez, E., Monzon, A., Ortega, E., Mancebo, S.: Assessment of Cross-Border Spillover Effects of National Transport Infrastructure Plans: An Accessibility Approach. Transp. Rev 29(4), 515-536 (2009). Milligan G, Cooper M.: An examination of procedures for determining the number of clusters in a data set. Psychometrika 50, 159–179 (1985). Mohamad, D., Sinha, K. C., Kuczek, T.: Annual Average Daily Traffic Prediction Model for County Roads. Transp. Res. Rec. 1617, 69–77 (1998).
18 MOVILIA: Encuesta de Movilidad de las Personas Residentes en España 2006-2007. Subdirección General de Estudios Económicos y Estadísticas del Ministerio de Fomento (2006). Rodriguez, G.: Lecture Notes on Generalized Linear Models. http://data.princeton.edu/wws509/notes/ (2007) Accessed 11 June 2013 Rui, X., Wunsch, II D.. Survey of clustering algorithms. IEEE Transaction on Neural Networks, 16(3), 645677 (2005). Selby, B., Kockelman, K.M.: Spatial Prediction of AADT at Unmeasured Locations by Universal Kriging. Transportation Research Board 90th Annual Meeting. Washington D.C, Paper number: 11-1665 (2011). Song, Y., Miller, H.J.: Exploring traffic flow databases using space-time plots and data cubes. Transportation 39, 215-234 (2012). Wang, X., Kockelman, K.M.: Forecasting Network Data: Spatial Interpolation of Traffic Counts Using Texas Data. Transp. Res. Rec. 2105, 100–108 (2009). Weijermars, W., van Berkum, E.: Analyzing highway volume patterns using cluster analysis. Proceedings of 8th Int. IEEE Conf. on Intelligent Transp. Syst., Vienna, Austria (2005). Xia, Q., Zhao, F., Chen, Z., Shen, L. D., Ospina, D.: Development of a Regression Model for Estimating AADT in a Florida County. Transp. Res. Rec. 1660, 32–40 (1999). Zhao, F., Chung, S.: Contributing Factors of Annual Average Daily Traffic in a Florida County: Exploration with Geographic Information System and Regression Models. Transp. Res. Rec. 1769, 113–122 (2001). Zhao, F., Park, N.: Using Geographically Weighted regression Models to Estimate Annual Average Daily Traffic. Transp. Res. Rec. 1879, 99–107 (2004).