scieee AI-readable full text Open interactive document viewer

Assessment of performance in the presence of undesirable outputs: the promotion of livability and sustainable development of cities and countries using Data Envelopment Analysis

Andreia Zanella

Full text

Assessment of performance in the presence of undesirable outputs: the promotion of livability and sustainable development of cities and countries using Data Envelopment Analysis Andreia Zanella A thesis submitted to Faculdade de Engenharia da Universidade do Porto for the doctoral degree in Industrial Engineering and Management Supervisors Professor Ana Maria Cunha Ribeiro dos Santos Ponces Camanho Professor Maria Teresa Galv˜ao Dias 2014 Abstract This thesis is concerned with the development of innovative models for the assessment and monitoring of performance using Data Envelopment Analysis (DEA) in the presence of undesirable outputs. The models focus on the construction of composite indicators and are applied to the evaluation of cities and countries with the aim to promote livability and sustainable development. The empirical part of the thesis contributes to the definition of better public policies through the identification of best practice examples and areas with more potential for improvements. The thesis includes four main research topics. The first topic discusses two alternative approaches to incorporate undesirable outputs in composite indicators constructed using DEA models. The first is an indirect approach, based on a standard DEA model which includes a transformation in the measurement scale of the undesirable outputs. The second is a direct approach, based on a DEA model specified with a directional distance function. This topic also discusses the incorporation of restrictions to weights in this context, and proposes a novel specification of assurance region type I weight restrictions. The second topic approaches the measurement of productivity change over time in the presence of undesirable outputs using the Malmquist-Luenberger index and the Luenberger index. An enhanced version of the MalmquistLuenberger index is proposed. The results obtained using the different productivity indices are compared and discussed. The third topic assesses the environmental performance of countries worldwide. The assessment is conducted using the composite indicator specified using the indirect approach proposed in chapter 3. It enables benchmarking in such a way that it is possible to identify the strengths and weaknesses of each country, as well as the peers with similar features to the country under assessment. The last topic develops a framework to assess the livability of European cities covering two components: human well-being and environmental impact. It is proposed a conceptual model that extends the concept of urban livability to include a component related to environmental sustainability. Then, the measurement of cities’ livability is conducted using the composite indicator specified with a directional distance function, as developed in chapter 3. Finally, the evolution of cities’ performance over time is assessed using the Luenberger productivity index. Overall, this thesis contributed to the development of robust tools to evaluate and promote livability and sustainable development in cities and countries, with a view to foster better standards of living now and in the future. i ii Resumo Esta tese concentra-se no desenvolvimento de modelos inovadores para a avalia¸c˜ao e monitoriza¸c˜ao de desempenho utilizando a t´ecnica Data Envelopment Analysis (DEA) na presen¸ca de outputs indesej´aveis. Os modelos desenvolvidos visam a constru¸c˜ao de indicadores comp´ositos e s˜ao aplicados na avalia¸c˜ao de cidades e pa´ıses com o objetivo de promover o bem-estar e o desenvolvimento sustent´avel. A parte emp´ırica da tese contribui para a defini¸c˜ao de melhores pol´ıticas p´ublicas atrav´es da identifica¸c˜ao dos exemplos de boas pr´aticas e das ´areas com maior potencial de melhoria. Esta tese divide-se em quatro t´opicos principais. O primeiro t´opico discute duas abordagens alternativas que podem ser usadas para incorporar outputs indesej´aveis em indicadores comp´ositos constru´ıdos com base em modelos de DEA. A primeira ´e uma abordagem indireta, baseada num modelo tradicional de DEA, que inclui uma transforma¸c˜ao na escala de medi¸c˜ao dos indicadores indesej´aveis. A segunda ´e uma abordagem direta, baseada num modelo especificado com uma fun¸c˜ao de distˆancia direcional. Esse t´opico tamb´em discute a incorpora¸c˜ao de restri¸c˜oes de pesos, e prop˜oe uma nova especifica¸c˜ao de restri¸c˜oes de pesos do tipo assurance region type I. O segundo t´opico aborda a an´alise da evolu¸c˜ao da produtividade ao longo do tempo na presen¸ca de outputs indesej´aveis utilizando os´ındices de MalmquistLuenberger eLuenberger. Neste t´opico tamb´em se prop˜oe uma vers˜ao melhorada do ´ındice Malmquist-Luenberger. Os resultados obtidos utilizando os diferentes ´ındices de produtividade s˜ao comparados e discutidos. O terceiro t´opico aborda a avalia¸c˜ao do desempenho ambiental de pa´ıses. A avalia¸c˜ao ´e conduzida por meio do indicador comp´osito proposto no cap´ıtulo 3, baseado na abordagem indireta. Este indicador comp´osito permite identificar as for¸cas e fraquezas de cada pa´ıs, bem como os pa´ıses que possuem caracter´ısticas similares `as do pa´ıs em avalia¸c˜ao e que podem ser considerados exemplos de boas pr´aticas. O ´ultimo t´opico abordado nesta tese desenvolve uma ferramenta para avaliar a habitabilidade das cidades europeias englobando duas componentes: bemestar humano e impacto ambiental. Prop˜oe-se um modelo conceptual que alarga o conceito de habitabilidade urbana de forma a incluir uma componente relacionada com a sustentabilidade ambiental. Tendo por base o modelo conceptual proposto, a habitabilidade das cidades ´e ent˜ao avaliada utilizando o indicador comp´osito desenvolvido no cap´ıtulo 3, especificado com uma fun¸c˜ao de distˆancia direcional. Por fim, a evolu¸c˜ao do desempenho das cidades ´e avaliada utilizando o ´ındice de Luenberger. Em s´ıntese, esta tese contribui para o desenvolvimento de ferramentas robustas para avaliar e promover o bem-estar e o desenvolvimento sustent´avel iii em cidades e pa´ıses, com vista a melhorar os padr˜oes de vida nos dias atuais e no futuro. iv Acknowledgments I would like to thank to my supervisors, Professor Ana Camanho, for her excellent guidance, constant encouragement and for sharing her knowledge, which were crucial to the completion of this thesis, and Professor Maria Teresa Galv˜ao Dias for her comments, encouragement and support during the doctoral research. I am grateful to the Doctoral Program in Industrial Engineering and Management for welcoming me and making this research work possible. I thank the financial support of the Portuguese Foundation for Science and Technology (FCT) through the projects iTeam (Pt/SES-SUES/0041/2008) and Mesur (PTDC/SEN-ENR/111710/2009). Also, I would like to thank my colleagues and friends, especially Isabel and Vera for sharing experiences and knowledge, and Carla and Maria for their friendship and support during these four years. I am grateful to my family for all the kindness and understanding over these years I have been living far away from them. I thank particularly my mother, for the support, encouragement and confidence placed in me. I certainly owe her the important steps I took in my life. Finally, I would like to thank Fabr´ıcio for his patience, help, and love throughout the time of this research and especially for encouraging me to come to Portugal to do this course. v Acronyms ARI Assurance Region type I ARII Assurance Region type II CI Composite Indicator CCPI Climate Change Performance Index CRS Constant Returns to Scale DEA Data Envelopment Analysis DRS Decreasing Returns to Scale DMU Decision Making Unit DDF Directional Distance Function EC Efficiency Change EPI Environmental Performance Index EU European Union GDP Gross Domestic Product IRS Increasing Returns to Scale LLuenberger MI Malmquist Index ML Malmquist-Luenberger NDRS Non-Decreasing Returns to Scale NIRS Non-Increasing Returns to Scale OECD Organisation for Economic Co-operation and Development PPS Production Possibility Set SFA Stochastic Frontier Analysis TC Technical Change VRS Variable Returns to Scale vi Table of Contents Abstract i Resumo iii Acknowledgements v Table of Contents vii List of Figures xii List of Tables xv 1 Introduction 1 1.1 Generalcontext.......................... 1 1.2 Motivation and research objectives . . . . . . . . . . . . . . . 4 1.3 Thesissummary ......................... 7 2 The assessment of performance 11 2.1 Introduction............................ 11 2.2 Production theory . . . . . . . . . . . . . . . . . . . . . . . . 12 2.2.1 Production processes . . . . . . . . . . . . . . . . . . . 12 2.2.2 Production sets and the axioms of production . . . . . 13 2.2.3 Technical efficiency . . . . . . . . . . . . . . . . . . . . 18 2.2.4 Parametric and nonparametric frontiers . . . . . . . . 22 2.3 The Data Envelopment Analysis technique . . . . . . . . . . . 24 vii LIST OF FIGURES 4.2 Production possibility set for the directional distance function 96 4.3 Production possibility set for the illustrative example in time periods0and1..........................103 4.4 Correlation between the ML indices and the Luenberger index 109 5.1 (a) Fixed weights for EPI categories and (b) contributions of the categories to the CI for Ireland . . . . . . . . . . . . . . . 129 6.1 Production possibility set and the different directional vectors usedinthisstudy.........................143 6.2 Virtual weights selected by Berlin in each dimension . . . . . 150 6.3 Weights assigned by Berlin and its peers to livability dimensions151 6.4 Cities’ performance over time . . . . . . . . . . . . . . . . . . 155 xiv List of Tables 3.1 Output indicators for the illustrative example . . . . . . . . . 65 3.2 Composite indicator, peers and targets obtained using model (3.2), with M equal to 30 . . . . . . . . . . . . . . . . . . . . 66 3.3 Composite indicator and ranks for different values of M. . . 67 3.4 Composite indicator, rank, peers and targets obtained from the Directional CI model . . . . . . . . . . . . . . . . . . . . . 70 3.5 Comparison of the composite indicator scores of the Directional CI models, with and without weight restrictions . . . . 77 4.1 Data for the illustrative example . . . . . . . . . . . . . . . . 102 4.2 Value of the Directional Distance Function for the illustrative example..............................104 4.3 Results for the different expressions of the productivity indices105 4.4 Countries’ productivity change for the three versions of the ML index and the Luenberger index . . . . . . . . . . . . . . 107 5.1 EPIindicators ..........................118 5.2 Characteristics of each cluster . . . . . . . . . . . . . . . . . . 124 5.3 Variables used in the construction of CI . . . . . . . . . . . . 125 5.4 Environmental performance scores and ranks for countries of clusterC2.............................127 5.5 Peers for Ireland . . . . . . . . . . . . . . . . . . . . . . . . . 130 6.1 Selected indicators to assess cities’ livability . . . . . . . . . . 141 6.2 Cities that reached the best performance score . . . . . . . . 149 xv LIST OF TABLES 6.3 Results for the assessment of performance change over time for four cities used as examples . . . . . . . . . . . . . . . . . 156 6.4 Average results for the efficiency and technological change and Luenberger index per group of cities . . . . . . . . . . . . 157 6.5 Number of cities that declined, maintained or improved the levels of productivity over time, per group . . . . . . . . . . . 158 6.6 Cities considered innovators . . . . . . . . . . . . . . . . . . . 158 A.1 Data for the 17 countries analyzed in 2005 and 2006 . . . . . 172 C.1 Results obtained in each cluster and comparison between the DEA scores and EPI scores . . . . . . . . . . . . . . . . . . . 175 E.1 Results for all cities assessed . . . . . . . . . . . . . . . . . . . 180 xvi Chapter 1 Introduction 1.1 General context The assessment and promotion of urban livability and sustainable development is an issue with growing importance among scientific and policymaking communities. Efforts from academy and governmental institutions are leading to a better understanding of how local communities, cities, and countries are performing compared to their peers, and encouraging the monitoring of progress over time. Urban livability is nowadays recognized as an important component of competitive advantage. It can be defined as the suitability of a given place for human living (Merriam-Webster, 2013). In the past few years, indicators that can measure urban livability have been developed to guide decisions about where to invest in a new business or where to seek employment (Australian Department of Infrastructure and Transport, 2012). National and local authorities are increasingly supporting efforts to improve the understanding of cities’ progress in terms of productivity, sustainability and livability, such as the yearly report “State of Australian Cities” by the Australian Government (Australian Department of Infrastructure and Transport, 2012). 1 Chapter 1 In addition to promoting competitive advantage by improving urban livability, governments can answer claims of more exigent citizens demanding better quality of life, facilities and infrastructure. Governments also face pressures from international agreements, such as the Kyoto Protocol (United Nations, 1998) or the European Union climate and energy package (European Union, 2008), environmentalists, and society to control and reduce emissions and preserve the natural resources and environment. Therefore, the governments’ challenge lies in providing good standards for human living but with a view to preserving the natural resources and environmental conditions. This requires the definition of economic and social development plans to comply with the sustainable development concept. Since the late 80’s, the definition of sustainable development proposed by the World Commission on Environment and Development is widely accepted. It states that sustainable development implies “meeting the needs of the present without compromising the ability of future generations to meet their own needs” (United Nations, 1987). Sustainable development is often presented as the integration of three interdependent components: economic, social and environmental. In the past few years, efforts to monitor these components have generated a large number of indicators intended to measure the performance in aspects related to emissions, waste production, green space, safety, income, health and education, among others. The majority of countries and cities collects data on these indicators, but the amount of data generated is often too large and it is not sufficiently clear to provide useful information and practical guidance to attend the policymakers needs. So, many of these indicators have been aggregated to build composite indicators (CIs). A composite indicator is given by the aggregation of several individual indicators in a single measure. It has benefits such as the capacity to summarize 2 1.1 General context information, the facility to interpret results compared with a battery of separate indicators, and the capacity to reduce the visible size of a set of indicators without dropping the underlying base information (Nardo et al., 2008). Examples of well-established composite indicators are the Environmental Performance Index (Emerson et al., 2012), Climate Change Performance Index (Burck et al., 2012), and the Human Development Index (United Nations, 2013). Although a considerable amount of research related to the construction of composite indicators has been developed in recent years (as reviewed in Nardo et al. (2008)), they are not effective in providing managerial information to guide improvements. Furthermore, they are prone to criticism regarding the subjectivity inherent in the specification of the relative importance given to the individual indicators in the construction of the CI. The Organisation for Economic Co-operation and Development (OECD) and the European commission provide a handbook for the construction of composite indicators that discusses the range of methodological approaches available to construct CIs (Nardo et al., 2008). The handbook highlights the growing interest in composite indicators by the academic circles, media and policymakers. One point discussed and recognized as a source of contention is the definition of the relative importance of the indicators. The handbook points Data Envelopment Analysis (DEA) as an interesting weighting and aggregation procedure to reduce the inherent subjectivity associated with the specification of weights. As the indicator weights result from an optimizing process based on linear programming, they are less prone to subjectivity and controversy. The broad subject area of this thesis is the use of frontier analysis methods, in particular DEA, for the assessment and promotion of urban livability and sustainable development. The motivation and main objectives of this thesis are discussed in detail in the next section. 3 Chapter 1 1.2 Motivation and research objectives Despite the growing interest and work in the field of performance assessment, robust tools to measure and guide improvements towards more livable and sustainable cities and countries are not available. More sophisticated and robust methods for performance assessment and benchmarking in this context are worth exploring. This research contributes to improve the current methods of livability and sustainability assessment by exploring new methodologies, based on frontier techniques, that can effectively provide enhanced managerial information and guide cities and countries towards sustainable development. Frontier methods have the advantage of allowing comparisons with the best observed performance by constructing a best practice frontier based on empirical data. From the alternative frontier methods available, in this thesis it was chosen to explore in detail the use of Data Envelopment Analysis due to the greater flexibility to incorporate the multidimensional nature of the livability and sustainable development concepts, and the use of minimal assumptions on the shape of the best practice frontier. The DEA technique uses linear programming to evaluate the relative efficiency of an homogeneous set of decision making units (DMUs) in their use of multiple inputs to produce multiple outputs. In addition to providing a single overall measure of performance and being able to fight the subjectivity associated with the specification of the indicators weights, the performance assessment using DEA allows the identification of areas for improvement and best practice examples. These properties make this technique particularly valuable for conducting benchmarking in the context of cities and countries livability and sustainable development. The advantages of the DEA technique that motivate its use in this thesis are described in greater detail in 4 1.2 Motivation and research objectives chapter 2. Standard DEA models assume that the individual output indicators represent good aspects, so they are measured on a scale for which higher values correspond to better performance. However, in real-world applications, both desirable and undesirable outputs indicators may be present. For example, in environmental performance assessment we may have an output indicator related to quality of the water, for which more output corresponds to better performance and another output indicator related to the levels of CO2 emissions, for which less output corresponds to better performance. In this situation, an inefficient DMU should increase the quality of the water and/or decrease the levels of CO2emissions to improve performance. Although the literature addresses the construction of DEA models with undesirable outputs, this issue is not discussed in the context of evaluations using composite indicators. Handling undesirable outputs in performance assessments using DEA requires particular attention as the alternative treatments may have a huge impact on the results. This thesis has two main objectives. The first is to develop innovative models for the assessment of performance in the presence of undesirable outputs using DEA. These models will be focused on applications involving the aggregation of key performance indicators. The fulfilment of this first objective opens the possibility to accomplish the second objective, that consists in undertaking a comprehensive evaluation of cities and countries, aiming to promote urban livability and sustainable development. In addition to assessing performance and monitoring its evolution over time, the models developed in this thesis can be used to identify best practice examples and areas in which cities and countries have more potential for improvements. Although the models are applied to the assessment of urban livability and sustainable development, they are easily transposable to different contexts and replicable over time. 5 Chapter 1 Given these main objectives, the specific objectives of the research described in this thesis are as follows: 1. To review the current approaches available to treat undesirable outputs in DEA models and develop a DEA-based composite indicator for the evaluation of performance in the presence of undesirable outputs; 2. To incorporate in the composite indicator information on the relative importance of individual indicators, in percentual terms, using weight restrictions; 3. To explore the different DEA-based approaches that can be used to accommodate undesirable outputs in the analysis of productivity change over time, namely the ratio-based Malmquist-Luenberger index and the difference-based Luenberger index; 4. To assess countries environmental performance using DEA, based on the aggregation of the indicators that underlie the estimation of the Environmental Performance Index (EPI), identifying the factors corresponding to the best and worst features of each country, as well as the peers with similar features to the countries with worse performance; 5. To define an appropriate set of indicators to assess cities’ livability extending the concept of urban livability to include a component related to environmental sustainability; 6. To assess the livability of European cities and provide managerial information for performance improvement. The managerial information is delivered through the identification of peers whose practices are examples to be followed and the identification of the areas in which each city has the best and worst features; 7. To analyse the cities evolution over time in terms of livability using the Luenberger productivity index, and identify the innovative cities, i.e. those 6 1.3 Thesis summary cities that are responsible for movements of the production frontier towards better productivity levels. 1.3 Thesis summary This thesis is structured in seven chapters, which are briefly described in this section. Chapter 2 presents an introduction to the methods for the evaluation of performance using frontier techniques. Particular emphasis is given to the DEA method, which is the core method used for the achievement of the research objectives stated for this thesis. Chapters 3 and 4 present the theoretical developments of the thesis. Each chapter includes a detailed literature review related to the topic approached. Chapter 3 is organized in three parts. The first reviews and discusses the literature related to the incorporation of undesirable outputs in DEA. The second addresses the incorporation of undesirable outputs in the context of the construction of composite indicators using DEA. Two different approaches are discussed and compared: an indirect approach, based on a standard DEA model that includes a transformation in the measurement scale of the undesirable outputs, and a direct approach, based on a DEA model specified using a directional distance function. Finally, the third part discusses the incorporation of restrictions to weights in the context of assessments involving composite indicators. Restrictions to weights can be included in the model in order to reflect decision-maker preferences about the relative importance of individual indicators, or/and to improve the discrimination among the DMUs’ performance scores. In both cases, the restrictions to weights prevent DMUs from obtaining a high score only owing to a judicious choice of weights. 7 Chapter 2 The postulates for the construction of the production possibility set Tcan be defined as follows (Banker (1984); Banker and Thrall (1992); Fare and Grosskopf (2005)): Postulate 1. Inclusion of observations. All observed DMUs are included in the production possibility set, i.e., (xj, yj)∈ T, for all j= 1, ..., n. Postulate 2. No outputs can be produced without some input. Zero output can be produced by any input vector x∈ <m +, but it is impossible to produce output without any inputs, i.e. if y≥0, y6= 0, x= 0 then (x, y)/∈T. Postulate 3. Inefficiency. (a) If (x, y)∈Tand x0≥x, then (x0, y)∈T. (b) If (x, y)∈Tand y0≤y, then (x, y0)∈T. Postulate 4. Ray unboundedness. If (x, y)∈Tthen (λx, λy)∈T, for all λ > 0. Postulate 5. Closedness: Technology T is a closed set.1 Given a closed correspondence, represented by →, if (xj, yj)→(x0, y0) and (xj, yj)∈Tfor all j= 1, ..., n, then (x0, y0)∈T. Postulate 6. Convexity. If (xj, yj)∈T,j= 1, ..., n, and λjare nonnegative scalars such that Pn j=1 λj= 1, then Pn j=1 λjxj,Pn j=1 λjyj∈T. 1Given the imposition of the ray unboundedness postulate, implying the existence of constant returns to scale, the closedness postulate is not required, as the technology is always closed (Podinovski and Bouzdine-Chameeva, 2011). For other variants of the postulates defining the production possibility set, involving changes to the ray unboundedness postulate, the closedness postulate needs to be explicitly stated. 14 2.2 Production theory Postulate 7. Minimum extrapolation. Tis the intersection of all sets ˆ Tthat satisfies the postulates 1, 2, 3, 4, 5 and 6. The previous postulates allow defining a Constant Returns to Scale (CRS) production possibility set. Dropping the ray unboundedness postulate will lead to the definition of a Variable Returns to Scale (VRS) production possibility set. The returns to scale is a characteristic of the boundary of the production technology. It is used to define how the technology behaves when changes to the scale of operation occur. To explain this concept, consider the case of a DMU that operates on the boundary of the PPS and uses a single input xto produce a single output y. If a change in the scale of operation occurs and the output increases proportionally to the input, it is said that the boundary at that point exhibits Constant Returns to Scale (CRS). If output increases proportionally more than the input, the boundary at that point exhibits Increasing Returns to Scale (IRS), and if the output increases proportionally less than the input, the boundary at that point exhibits Decreasing Returns to Scale (DRS). The term Variable Returns to Scale (VRS) is used to denote a boundary that exhibits any combination of IRS and/or DRS with CRS. The production technology Tcan alternatively be represented by input or output sets. The input requirement set, L(y), describes the set of all input vectors xwhich can be used to produce the output vector y, as shown in (2.2). L(y) = {x:ycan be produced by x}={x: (x, y)∈T(x, y)}(2.2) The input set is bounded from the input isoquant which contains the min15 Chapter 2 imum inputs necessary to secure a certain output. The mathematical formulation of the input isoquant is defined as shown in (2.3). IsoqL(y) = {x:x∈L(y), δx /∈L(y) for δ < 1}(2.3) The isoquant defines a frontier to the input set. Those DMUs that lie on the frontier are efficient in the sense that radial (proportional) reduction of inputs is not possible. For the example shown in Figure 2.2, involving two inputs and a single output, the input set L(y) contains all the input vectors within the isoquant IsoqL(y). This isoquant is formed by the segments linking the points B, Cand D, and the rays parallel to the axes, i.e., the ray starting in Band passing through A, and the ray starting in Dand passing through E. Note that in the rays −−→ BA and −−→ DE, the proportional reduction of both inputs is not possible, although it is possible to reduce only one input by moving along the rays to reach the efficient frontier at Bor D. Figure 2.2: The input set Similarly, the production technology can be represented by an output set 16 2.2 Production theory P(x). It describes the set of all output vectors ywhich can be produced using the input vector x, as shown in (2.4). P(x) = {y:xcan produce y}={y: (x, y)∈T(x, y)}(2.4) The output set is bounded by the output isoquant that contains all the output combinations that cannot be proportionally increased without increasing the input vector x. It is defined as shown in (2.5). IsoqP(x) = {y:y∈P(x), θy /∈P(x) for θ > 1}(2.5) The concepts underlying the representation of the technology for the output set are illustrated in Figure 2.3, using two outputs that are produced with a single input. Figure 2.3: The output set The output set P(x) is formed by all points in the production possibility set bounded by the isoquant, i.e., by the segments linking points B,Cand 17 Chapter 2 Dand the rays parallel to the axes, spanned from points Band D. The points that lie on the rays −−→ BA and −−→ DE cannot be proportionally increased keeping the current input levels, as only one of the outputs can be increased, i.e., output 2 when moving along −−→ BA in direction to B, or output 1 when moving along −−→ DE in direction to D.2 2.2.3 Technical efficiency Efficiency involves a comparison of the actual location of a DMU within the PPS with the optimal input and output levels corresponding to the points located on the production frontier. The first measure of technical efficiency dates back to the work of Debreu (1951) and Farrell (1957). The Debreu’s measure of efficiency was called the “coefficient of resource utilisation”. Farrell extended this previous work by proposing the measurement of efficiency using empirical observations, i.e. by comparing a DMU to the best actually achieved by its peers. The DMUs’ efficiency measure can be obtained from two perspectives, corresponding to an input-reduction or output-expansion orientation. Assuming an input orientation, the Debreu-Farrell technical efficiency measure is defined as the maximum radial (proportional) reduction to all inputs that is feasible to achieve whilst securing a certain output level, within a given technology. On the other hand, assuming an output orientation, this measure is defined as the maximum radial (proportional) expansion to all outputs that is feasible to achieve with a certain input level, within a given technology (Fried et al., 2008). In order to relate the Debreu-Farrell measures to the structure of the pro2The production technologies represented in Figures 2.2 and 2.3 are piecewise linear. While the econometric approach estimates smooth parametric frontiers, the mathematical programming approach estimates piecewise linear frontiers. As this thesis will use a mathematical programming approach, we will present the illustration of piecewise linear frontiers only. 18 2.2 Production theory duction technology, consider the technologies defined in (2.2) and (2.4). The Debreu-Farrel input oriented measure of technical efficiency (TEi) can be formally defined as the value of the function shown in (2.6). TEi(x, y) = min{δ:δx ∈L(y)}(2.6) Note that, for x∈L(y), δ≤1, and for x∈IsoqL(y), δ= 1. Figure 2.4 graphically illustrates the input oriented technical efficiency (TEi) measure using a sample of seven DMUs. DMUs A, B, C, D and E lie on the frontier defined by the isoquant IsoqL(y), and thus are considered efficient in the Debreu-Farrell sense (δ= 1). For DMUs Fand G, the technical efficiency measure is given by the ratios OF∗ OF and OG∗ OG , respectively, and is less than one for both DMUs. Figure 2.4: Technical efficiency measure for the input set Similarly, the Debreu-Farrell output oriented measure of technical efficiency can be defined as the value of the function shown in (2.7). 19 Chapter 2 TEo(x, y) = max{θ:θy ∈P(x)}(2.7) Again, note that for y∈P(x), θ≥1, and for y∈IsoqP(x), θ= 1. In Figure 2.5, the output oriented technical efficiency measure is illustrated for a set of seven DMUs. The DMUs A, B, C, D and E are considered efficient in the Debreu-Farrell sense as they lie on the isoquant IsoqP (x), so θ= 1. The technical efficiency measure for DMUs F and G is given by the ratios OF∗ OF and OG∗ OG , respectively, and its value is greater than one for both DMUs. 3 Figure 2.5: Technical efficiency measure for the output set However, the Debreu-Farrell definition of efficiency, also known as radial efficiency, is not sufficient for defining the “truly” efficient DMUs. A stronger notion of efficiency was provided by Koopmans (1951). The author stated that “a producer is technically efficient if an increase in any output requires a reduction in at least one other output or an increase in at least one input, 3This interpretation follows the convention of defining efficiency as a ratio of optimal to actual. Fried et al. (2008) noted that some authors replace (2.7) with T Eo(x, y) = [max{θ:θy ∈P(x)}]−1. 20 2.2 Production theory and if a reduction in any input requires an increase in at least one other input or a reduction in at least one output”. The Debreu-Farrell’s efficiency is based on the radial contraction (expansion) of inputs (outputs). This implies that after the equiproportional reduction (expansion) of the inputs (outputs) of a DMU leading to a point on the boundary of the PPS, there is no scope for further improvement for at least one input (output). On the other hand, Koopmans’ efficiency investigates further the potential reduction (expansion) of each input (output) beyond the radial movements. Considering the technologies defined in (2.2) and (2.4), the Koopmans’ notion of efficiency can be represented by the input and output efficient subsets, given by the expressions shown in (2.8) and (2.9), respectively. E(y) = {x:x∈L(y), x0≤xand x06=x⇒x0/∈L(y)}(2.8) E(x) = {y:y∈P(x), y0≤yand y06=y⇒y0/∈P(x)}(2.9) Therefore, E(y) and E(x) are subsets of the isoquants IsoqL(y) and IsoqP (x), respectively. The efficient subsets E(y) and E(x) differ from the isoquants as the latter may contain rays parallel to the axes, which are not part of the efficient frontier in Koopmans’ sense. These concepts are illustrated in Figures 2.4 and 2.5. In both figures the points A,Eand G∗satisfy the Debreu-Farrell conditions for technical efficiency but they do not satisfy the Koopmans’ conditions. For point G∗, for example, in the case of the input oriented approach (Figure 2.4), it is possible to reduce the consumption of input 1, without increasing input 2 21 Chapter 2 and keeping the current output level, i.e. to move point G∗horizontally until reaching point D. Similarly, in the case of the output oriented approach (Figure 2.5), it is possible to increase the production of output 1, keeping the current input level and without reducing output 2, i.e. to move point G∗horizontally until reaching point D. Conversely, points B,C,Dand F∗ satisfy both definitions of technical efficiency. In the remainder of this thesis, the concept of technical efficiency adopted will follow the Koopmans’ criteria, as it is consensual in the literature that it corresponds to a stronger notion of efficiency, and thus should underlie the analysis of performance based on quantitative methods. 2.2.4 Parametric and nonparametric frontiers The performance measurement methods that rely on the estimation of a frontier have evolved following two parallel research lines: a parametric and a nonparametric approach. These methods differ in the way the frontier is specified and estimated. The parametric approach specifies the frontier as a function with a precise mathematical form (usually the translog or the Cobb-Douglas functions). It requires an a priori specification of the functional form representing the frontier. On the other hand, the nonparametric approach does not require any assumptions with respect to the functional form of the frontier. The frontier is defined by a set of postulates that the points on the boundary of the production possibility set have to satisfy. Figure 2.6 classifies the various types of parametric and nonparametric frontiers. The most common methods for efficiency evaluation are Data Envelopment Analysis (DEA) in the nonparametric literature, and Stochastic Frontier Analysis (SFA) in the parametric literature. Both the parametric and nonparametric methods can be further divided into stochastic and 22 2.2 Production theory deterministic. Figure 2.6: Classification of production frontiers The stochastic approach allows for random noise and measurement error in the data. Thus, the DMUs’ deviations from the estimated frontier are explained both by the DMUs’ inefficiency and the presence of noise or measurement error in the data. In these cases, the estimation of the production frontier involves the use of statistical techniques. The deterministic approach assumes that there is no random noise affecting the construction of the frontier. Thus, the DMUs’ deviations from the estimated frontier are exclusively explained by inefficiency. In these cases, the estimation of the production frontier involves the use of mathematical programming techniques. As noted by Fried et al. (2008), an advantage of 23 Chapter 2 is the inverse of the maximum factor (1/θ∗) by which its outputs levels can be increased equiproportionally within the PPS, whilst the inputs are held constant. A DMU is considered radially efficient if 1/θ∗is equal to one.5 For a DMUj0to be considered truly efficient, in Koopmans’ sense, it must be radially efficient, and, in addition, the slack variables s∗ iand s∗ rmust be equal to zero. The slack variables indicate the extent to which each input or output can be improved beyond the amount indicated by the factors δor θ. In order to illustrate the concept of slack, useful to understand the efficiency in Koopmans’ sense, Figures 2.7 and 2.8 show the slacks in the technical efficiency measure for an input oriented and an output oriented assessment, respectively. For the seven DMUs illustrated in Figure 2.7, DMU A has a positive slack in the input 2 (S2A). Conversely, DMU E and the projection of DMU G on the frontier (G∗) have positive slacks in input 1 (S1Eand S1G, respectively). To become efficient in Koopmans’ sense, these DMUs should move along the segment parallel to the axes to reach the efficient frontier at points B and D, as explained in section 2.2.3. Figure 2.7: Slacks in the technical efficiency measure for an input oriented assessment 5Note that in the input oriented case, the radial efficiency score, given by δ∗,coincides with the Debreu-Farrell radial efficiency measure, shown in (2.6). However, in output oriented case, the radial efficiency score, given by 1/θ∗, corresponds to the inverse of the Debreu-Farrell radial efficiency measure, shown in (2.7). 30 2.3 The Data Envelopment Analysis technique Figure 2.8: Slacks in the technical efficiency measure for an output oriented assessment In addition to the efficiency measure, the DEA envelopment model is also able to provide other managerial information. As by-products of the efficiency assessment, it is possible to identify the peers for each inefficient DMU and the targets that the inefficient DMUs should aim to achieve in order to become efficient. For each DMUj0, the optimal solution of models (2.13) and (2.14) looks for a comparator, i.e. a composite DMU corresponding to a linear combination of efficient DMUs. This reference DMU (Pn j=1 λ∗ jxij,Pn j=1 λ∗ jyrj) uses the same or lower levels of input and produces equal or higher levels of output than DMUj0. When λ∗ j>0, the corresponding DMUjis a peer to DMUj0. The targets are given by the expressions shown in (2.15) and (2.16), corresponding to the input and output oriented models, respectively.              xIT ij0=δ∗xij0+s∗ i= n X j=1 λ∗ jxij i= 1, ..., m yIT rj0=yrj0+s∗ r= n X j=1 λ∗ jyrj r= 1, ..., s (2.15) 31 Chapter 2              xOT ij0=xij0+s∗ i= n X j=1 λ∗ jxij i= 1, ..., m yOT rj0=θ∗yrj0+s∗ r= n X j=1 λ∗ jyrj r= 1, ..., s (2.16) The models for efficiency measurement described in this section assumed constant returns to scale. The next section will introduce models for efficiency measurement under variable returns to scale. 2.3.1.3 Variable returns to scale models As mentioned in section 2.2.2, the returns to scale is a characteristic of the boundary of the technology of production. It measures the responsiveness of the output to equal proportional changes to the input. The Constant Returns to Scale (CRS) assumption is appropriate when maximal productivity is attainable for all scale size ranges. However, in some cases, maximum productivity is limited to a specific scale size range and, as a consequence, some DMUs may be either too small or too large to achieve the maximum productivity. In theses cases, the appropriate assumption is the existence of variable returns to scale. The Variable Returns to Scale (VRS) specification allows calculating technical efficiency measures accounting for scale effects. It ensures that inefficient DMUs are only compared with DMUs of similar size, such that the productivity levels observed in the peers are achievable by the DMU under assessment. Figure 2.9 illustrates the frontiers assuming constant returns to scale and variable returns to scale for the single-output, single-input case. The CRS frontier is given by the ray starting in the origin and passing through DMU 32 2.3 The Data Envelopment Analysis technique B. Note that DMU B corresponds to the maximum productivity level of the sample. Assuming CRS, a change in the input level results in a equally proportionate change in the output level, and so, in this case, the frontier is spanned by the ray −−→ OB. By relaxing the ray unboundedness assumption used for the construction of the CRS frontier, the frontier of the PPS is defined based on the observed performance of the DMUs given their scale of operation. In the example shown in Figure 2.9, the efficient frontier is redefined as the segments between A, B and C, corresponding to a VRS frontier. The segment linking A to B exhibits increasing returns to scale (IRS), as a change to the input level causes a greater than proportionate change to the output, and the segment linking B to C exhibits decreasing returns to scale (DRS), as a change to the input causes a less than proportionate change to the output. In this case, point B, which is part both of the CRS and VRS frontiers, is the only point of the VRS frontier that is said to exhibit constant returns to scale. Figure 2.9: Returns to scale 33 Chapter 2 Although the returns to scale concept was illustrated for the single-output, single-input case, it can be generalised to the case of multiple inputs and multiple outputs, see Banker et al. (1984). In order to enable the estimation of efficiency assuming VRS, Banker et al. (1984) proposed a modification to the original DEA model of Charnes et al. (1978). The VRS models with input and output orientation are presented in (2.17) and (2.18), respectively. DEA input oriented model under VRS (multiplier formulation): Max ˆej0= s X r=1 uryrj0+ω(2.17) s.t. m X i=1 vixij0= 1 s X r=1 uryrj − m X i=1 vixij +ω≤0j= 1, . . . , n ur≥ r = 1, . . . , s vi≥ i = 1, . . . , m ωis free The variable ωcan be used to impose different types of returns to scale. For non-increasing returns to scale the restriction ω≤0 should be imposed. Conversely, for non-decreasing returns to scale the restriction ω≥0 should be imposed. 34 2.3 The Data Envelopment Analysis technique DEA output oriented model under VRS (multiplier formulation): Min ˆ hj0= m X i=1 vixij0+$(2.18) s.t. s X r=1 uryrj0= 1 − s X r=1 uryrj + m X i=1 vixij +$≥0j= 1, . . . , n ur≥ r = 1, . . . , s vi≥ i = 1, . . . , m $is free In model (2.18), the different types of returns to scale are imposed using the variable $. The restriction $≥0 is used for non-increasing returns to scale, and $≤0 for non-decreasing returns to scale. 6 Under VRS, the efficiency scores may be different depending on the model orientation considered. Thus, although the location of the frontier and the subset of efficient DMUs is the same for both model orientations, for inefficient DMUs, the scores of the input and output oriented models may be different (i.e., e∗ j06= 1/h∗ j0). The dual models corresponding to the DEA weights formulations shown in (2.17) and (2.18) are presented in (2.19) and (2.20), respectively. 6Note that the value of the ωand $in formulations (2.17) and (2.18) represent the intercept corresponding to the facet of the frontier against which DMUj0is evaluated. 35 Chapter 2 DEA input oriented model under VRS (envelopment formulation): Min ˆej0=ˆ δ− m X i=1 si+ s X r=1 sr!(2.19) s.t. ˆ δ xij0− n X j=1 λjxij −si= 0 i= 1, . . . , m n X j=1 λjyrj −sr=yrj0r= 1, . . . , s n X j=1 λj= 1 λj≥0j= 1, . . . , n si≥0i= 1, . . . , m sr≥0r= 1, . . . , s DEA output oriented model under VRS (envelopment formulation): Max ˆ hj0=ˆ θ+ m X i=1 si+ s X r=1 sr!(2.20) s.t. n X j=1 λjxij +si=xij0i= 1, . . . , m ˆ θ yrj0− n X j=1 λjyrj +sr= 0 r= 1, . . . , s n X j=1 λj= 1 λj≥0j= 1, . . . , n si≥0i= 1, . . . , m sr≥0r= 1, . . . , s In both formulations (2.19) and (2.20), the non-increasing returns to scale assumption is obtained by imposing the sum of lambda variables to be smaller 36 2.3 The Data Envelopment Analysis technique than or equal to one (Pn j=1 λj≤1), and the non-decreasing returns to scale is obtained by imposing the sum of lambda variables to be greater than or equal to one (Pn j=1 λj≥1). Regardless of the nature of the returns to scale assumed, the targets for the inefficient DMUs are obtained using the expressions presented in (2.15) and (2.16). The nature of returns to scale of a DMU can be identified by comparing the efficiency measure derived from the different DEA models. If a DMU obtains different efficiency scores in the solution of the CRS and VRS models, then this particular DMU exhibits VRS. However, this does not indicate whether the DMU is operating in an area of increasing or decreasing returns to scale. This issue can be solved by running an additional DEA model imposing non-increasing returns to scale (NIRS). The nature of the scale inefficiencies of the DMU can be identified by comparing the efficiency scores derived from the NIRS technology and the VRS technology. If they are equal then decreasing returns to scale exist. If they are different then increasing returns to scale apply (Fare et al., 1985). In assessments allowing for variable returns to scale, DMUs with extreme scale sizes (very small or very large) may be classified as efficient due to the lack of comparators with a similar scale size. In addition, the VRS frontier will always envelop the data as tightly as possible, regardless of whether it is the correct assumption for a given context of DMU’s activity. This will result in an increase in the value of the efficiency estimate and less discrimination between the DMUs’ efficiency scores whenever the VRS assumption is used. Given these limitations, VRS should be assumed only when there are strong evidences that scale effects actually exist. In cases in which the nature of the production technology (CRS or VRS) is unknown a priori, hypothesis tests can be used to verify the existence of scale effects. The VRS model should only be used when scale effects 37 Chapter 2 are demonstrated. Banker (1996) was the first to use hypothesis tests in a DEA assessment to verify the nature of the returns to scales. Some years later, Simar and Wilson (2002) proposed the use of bootstrapping to test the hypothesis regarding returns to scale. Next, the concept of scale efficiency is introduced. Scale efficiency is a measure of how much the scale of operation of a DMU impacts on its ability to achieve the maximum productivity. Comparing the distance between the CRS and VRS frontiers corresponding to the evaluation of a given DMU, it is possible to define a measure of scale efficiency. The scale efficiency of DMUj0can be calculated as shown (2.21). Scale efficiency of DMUj0=CRS efficiency score of DMUj0 VRS efficiency score of DMUj0 .(2.21) The concept of scale efficiency is illustrated in Figure 2.10 adopting an output orientation for the efficiency assessment. DMU D is not technically efficient as it is operating below the efficient frontier. It could become efficient, in a pure technical sense, by increasing its output level until reaching point D∗. However, at this point, DMU D would be considered scale inefficient because it is not possible to achieve the maximum productivity level (observed at DMU B). The scale efficiency measure, given by the ratio O0D∗ O0D∗∗ , evaluates the distance between the CRS and VRS frontiers, and measures the amount of output loss attributable to having a scale size that prevents attaining maximum productivity. Therefore, the overall efficiency measure for DMU D, which incorporates both pure technical and scale efficiencies, is given by the product of the pure technical and scale efficiency scores, i.e., O0D O0D∗×O0D∗ O0D∗∗ =O0D O0D∗∗ . 38 2.3 The Data Envelopment Analysis technique Figure 2.10: Scale efficiency measurement 2.3.2 Restricting weights in DEA models The original DEA model, developed by Charnes et al. (1978), allows total flexibility in the selection of the weights viand urto be attached to the inputs and outputs, such that each DMU tries to maximise its efficiency score, given the inputs consumed and output levels attained. This total flexibility in the selection of weights is important for the identification of under-performing DMUs. Under these conditions, if a DMU does not achieve the maximum score, even when evaluated with a set of weights that intends to maximise its performance score, it provides irrefutable evidence that other DMUs are performing better. However, this feature can also cause problems. For example, one can argue that the weights assigned to the inputs and outputs are not realistic, and thus the robustness of the efficiency measure and its applicability in real-world context is questionable. The first attempt to restrict the flexibility of the weights in DEA models was made in the mid 80’s. The use of DEA by Thompson et al. (1986) to sup39 Chapter 2 46 Chapter 3 Modeling undesirable outputs in the construction of DEA-based composite indicators 3.1 Introduction In standard Data Envelopment Analysis models, as presented in chapter 2, an inefficient DMU can improve its performance by increasing the levels of outputs (results obtained) or decreasing the levels of inputs (resources used). This point of view makes sense when all outputs are intended to be increased and all inputs are intended to be reduced. However, realworld applications may involve both desirable and undesirable outputs and inputs. The accommodation of undesirable factors in DEA models is not straightforward. Although the issue of dealing with undesirable outputs and inputs in DEA models has been approached in the literature by various authors, they do not address the modeling of undesirable factors in the construction of composite indicators. A composite indicator is given by the aggregation of several 47 Chapter 3 individual indicators, it focuses only in the achievements (results obtained) of a set of DMUs and are intended to reflect multidimensional concepts in a single measure. In composite indicators constructed using DEA models, all individual indicators are specified as outputs and a unitary input underlying the evaluation of every DMU is considered. The use of DEA for performance assessments focusing only on achievements, rather than the conversion of inputs to outputs, was first proposed by Cook and Kress (1990), with the purpose to construct a preference voting model (for aggregating votes in a preferential ballot). Other relevant studies that support the empirical use of DEA models only with outputs (or productivity indicators that aggregate output and input information, such as revenue per employee and GDP per capita) can be found in different fields, such as macroeconomic performance assessment (Lovell et al., 1995), human development (Mahlberg and Obersteiner, 2001; Despotis, 2004, 2005), technology achievement (Cherchye et al., 2008) and evaluation of urban quality of life (Morais and Camanho, 2011). In these studies, all variables were specified as outputs and an identical input level, which for simplicity was assumed to be equal to one, was specified for all DMUs. In this chapter, we approach two main issues: the construction of CIs that include both desirable and undesirable factors with aggregation procedures based on DEA, and the use of weight restrictions in this context. Two aggregation procedures are considered to construct the CI in the presence of undesirable outputs. First, the CI is derived based on a traditional DEA model. In this case the undesirable outputs require a prior transformation in the measurement scale to be accommodated in the CI model. Next, we propose the use of a CI model specified using a directional distance function. This CI model is able to accommodate the undesirable outputs in their original form. The features, weaknesses and advantages of both approaches 48 3.1 Introduction are discussed and illustrated using a small example. Concerning the use of weight restrictions in the context of the estimation of CIs, it is explored the implementation of the two most popular types of weight restrictions: virtual weight restrictions and the assurance regions type I (ARIs). We propose an enhanced formulation of weight restrictions to incorporate the relative importance of individual indicators expressed in percentage, using ARI. We discuss and illustrate the features of each type of weight restriction using a small example. Therefore, from a methodological perspective, the two major contributions of this chapter consist on the construction of DEA-based CIs that can accommodate both desirable and undesirable outputs and provide peers and targets as by-products of the assessment, and the specification of a novel type of weight restriction, using ARIs, to incorporate the relative importance of indicators, expressed as a percentage. The remainder of this chapter proceeds as follows. Section 3.2 provides a literature review of studies that have approached the issue of dealing with undesirable outputs in DEA models. Section 3.3 presents the DEA formulations that can be used for efficiency assessments in the presence of undesirable outputs, and adapts the models to evaluations with composite indicators. Section 3.4 approaches the construction of composite indicators incorporating decision-maker preferences, using weight restrictions of different types. Section 3.5 illustrates the specificities of the models, and discusses their strengths and weaknesses using a small numerical example. Finally, section 3.6 presents the conclusions. 49 Chapter 3 3.2 Review of the literature on undesirable outputs in DEA Although the seminal paper of Koopmans (1951) already mentioned that the production process may also generate undesirable outputs, such as air pollutants or waste, it was only in the 80s that the issue of efficiency measurement in the presence of undesirable outputs was first addressed. One of the earliest studies addressing the incorporation of undesirable outputs in the assessment of production efficiency was developed by Pittman (1983). This study extended the multilateral productivity indicator proposed by Caves et al. (1982) to include measures of both desirable and undesirable outputs. The multilateral productivity indicator developed by Caves et al. (1982) required the specification of the price data, but this information is often unavailable for undesirable factors. Therefore, Pittman (1983) proposed an extension of this indicator, which assigned a value to the undesirable outputs based on estimates of shadow prices instead of market prices. Some years later, Fare et al. (1993) proposed an alternative method to estimate shadow prices based on the distance function defined by Shephard (1970). The specification of the shadow prices of undesirable outputs using a linear programming model allowed enhancing the approach proposed by Pittman (1983). Fare et al. (1989) also proposed a modification of Farrell (1957) approach to efficiency measurement to allow an asymmetric treatment of desirable and undesirable outputs. While the multilateral productivity indicator requires the specification of price information for the undesirable outputs, the nonparametric approach of Fare et al. (1989) only requires data on quantities of the undesirable outputs. The authors proposed a hyperbolic model to efficiency measurement to allow considering different assumptions on the disposability of undesirable outputs. The new constraints state that the 50 3.2 Review of the literature on undesirable outputs in DEA desirable outputs are strongly disposable (i.e. they can be reduced without cost), while the undesirable outputs are weakly disposable (i.e. they can only be reduced in conjunction with a reduction in the other outputs or an increase in the use of inputs). Some years later, Chung et al. (1997) introduced a different approach to deal with undesirable outputs in the efficiency and productivity measurement literature. The authors extended the Chambers et al. (1996) directional distance function to allow expanding the desirable outputs while simultaneously contracting the undesirable ones. The outputs are expanded or contracted along a path that is defined according to a directional vector. The directional distance function has been widely used in the context of environmental performance assessments, in which the production of waste (an undesirable output) is often present. The approaches mentioned above are known as direct approaches to treat undesirable outputs. These approaches allow treating the outputs in their original form, that is, without requiring any modification to the measurement scale. On the other hand, there are indirect approaches that transform the values of the undesirable outputs to allow treating them as normal outputs in traditional DEA models. Scheel (2001), Dyson et al. (2001) and Seiford and Zhu (2002) discussed the different approaches to handle undesirable outputs in DEA models using indirect approaches. One option is to move the variables from the output to the input side. Scheel (2001) pointed out that this approach results in the same technology set as incorporating the undesirable outputs as normal outputs, in the form of their additive inverses (−yund). The incorporation of the undesirable outputs in the form of their additive inverses was first suggested by Koopmans (1951). Regarding this option, Seiford and Zhu (2002) pointed out that to treat undesirable outputs as inputs would not re51 Chapter 3 flect the real production process, as the input-output structure that defines the production process would be lost. Another possibility is to consider the undesirable outputs in the form of their multiplicative inverses (1/yund), as proposed by Golany and Roll (1989). Regarding this option, Dyson et al. (2001) pointed out that this transformation would destroy the ratio or interval scale of the data. The third option is to add to the additive inverses of the undesirable outputs a sufficient large positive number (−yund +c), as first suggested by Ali and Seiford (1990). This transformation is the most frequently used in the literature to deal with undesirable outputs using a traditional DEA formulation (Cook and Green, 2005; Oggioni et al., 2011). It has the advantage of enabling a simple interpretation of results, but it is sensitive to the choice of the constant c, as will be discussed in the next section. In addition to the above mentioned approaches, in Cherchye et al. (2011) the transformation in the measurement scale of the undesirable outputs was performed based on a normalization procedure, which was applied both to desirable and undesirable outputs. This procedure results in indicators varying between 0 and 1. As data normalization leads to a loss of information, this approach is rarely used in DEA studies. It does not take advantage of the ability of DEA to deal with data measured on different scales. 3.3 Undesirable outputs in the construction of DEAbased composite indicators In this section we discuss the main models available to treat undesirable outputs in DEA efficiency assessments. This is followed by the presentation of CI models that can be obtained based on these DEA formulations. As mentioned in the literature review, the models can follow a direct or an indirect approach to handle the undesirable outputs. 52 3.3 Undesirable outputs in the construction of DEA-based composite indicators 3.3.1 Indirect approach The DEA output oriented model by Charnes et al. (1978), presented in section 2.3.1.2, can be used for assessments involving undesirable outputs with a transformation in the measurement scale of the undesirable outputs as proposed by Seiford and Zhu (2002). The resulting model, with a constant returns to scale formulation in shown in (3.1). Min hj0= m X i=1 vixij0(3.1) s.t. s X r=1 uryrj0+ l X k=1 pk(Mk−bkj0)=1 s X r=1 uryrj + l X k=1 pk(Mk−bkj)− m X i=1 vixij ≤0j= 1, . . . , n ur≥0r= 1, . . . , s pk≥0k= 1, . . . , l vi≥0i= 1, . . . , m In formulation (3.1), the efficiency score for DMU j0is given by 1/hj0, and it ranges between zero (worst) and one (best). xij (i= 1, ..., m), yrj (r= 1, ..., s) and bkj (k= 1, ..., l) correspond to the value of the input i, desirable output rand undesirable output k, respectively, for DMU j(j= 1, ..., n). vi,urand pkare the weights attached to the inputs, desirable outputs and undesirable outputs, respectively, in the performance assessment. Mkis a large positive number greater than or equal to the maximum value of the undesirable output kobserved in all DMUs. In the transformation proposed by Seiford and Zhu (2002), all of the translated outputs are positive, so the constant Mkmust be a value larger than the maximum observed for each output indicator. As the model is sensitive to the choice of the Mk 53 Chapter 3 value, a sensitivity analysis of the results for different values of Mkshould be undertaken. In a CI we have only outputs to be aggregated, so we can assume that all DMUs are similar in terms of inputs. Thus, following Koopmans (1951), Lovell (1995) and Lovell et al. (1995), we can have a unitary input underlying the evaluation of every DMU, interpreted as a “helmsman” attempting to steer the DMUs towards the maximization of outputs. By considering a unitary input level for all DMUs in model (3.1) we obtain the CI model presented in (3.2). Min hj0=v(3.2) s.t. s X r=1 uryrj0+ l X k=1 pk(Mk−bkj0)=1 s X r=1 uryrj + l X k=1 pk(Mk−bkj)−v≤0j= 1, . . . , n ur≥0r= 1, . . . , s pk≥0k= 1, . . . , l v≥0 In formulation (3.2), the reciprocal of the value of the objective function represents an efficiency measure that ranges between zero and one, where one correspond to the best level of performance observed in the sample. This is also the value of the composite indicator associated with this model that will be used throughout the thesis. The use of DEA to construct CIs was popularized by Cherchye et al. (2007). This approach is known as “benefit of the doubt” construction of composite indicators. The CI proposed by Cherchye et al. (2007) is equivalent to the input oriented model proposed by Charnes et al. (1978), assuming constant 54 3.3 Undesirable outputs in the construction of DEA-based composite indicators returns to scale and a unitary input level for all DMUs. A CI can also be obtained from the output oriented model, and this was the approach followed in this chapter, leading to the formulation (3.2). As both formulations assume constant return to scale, the performance scores obtained with different orientations are the same. The advantage of using the output oriented formulation, as done in this chapter, is that it leads to a more direct estimation of targets using the dual model, and facilitates the incorporation of weight restrictions. For benchmarking purposes, the identification of the peers and targets for the inefficient DMUs can be done through the envelopment formulation of model (3.2), shown in (3.3). The objective function value at the optimal solution of the model (3.3) corresponds to the factor θby which all outputs of the DMU under assessment can be proportionally improved to reach the target output values. As occurred in the primal formulation shown in (3.2), the performance score, or composite indicator, of DMU j0under assessment is the reciprocal of the objective function value of model (3.3). Therefore, the DMUs with the best performance are those for which there is no evidence that it is possible to expand their outputs, such that the value of θ∗is equal to 1. Max θ(3.3) s.t. θ yrj0− n X j=1 λjyrj ≤0r= 1, . . . , s θ(Mk−bkj0)− n X j=1 λj(Mk−bkj)≤0k= 1, . . . , l n X j=1 λj≤1j= 1, . . . , n λj≥0j= 1, . . . , n 55 Chapter 3    φr≤uryrj0≤ψr φk≤pkbkj0≤ψk (3.14) 3.4.2 ARI restrictions in CIs The most prevalent type of direct weight restrictions used in DEA applications are assurance regions type I (ARI), proposed by Thompson et al. (1990). They usually incorporate information concerning marginal rates of substitution between the outputs (or between the inputs), as explained in section 2.3.2.1. It is worth highlighting that this type of weight restrictions is sensitive to the units of measurement of the inputs and outputs (Allen et al., 1997). As a result, it is often difficult to specify meaningful marginal rates of substitution between the variables in empirical applications. In this chapter we propose an enhanced formulation of ARI weight restrictions that enables expressing the relative importance of the output indicators in percentual terms, instead of specifying marginal rates of substitution. This requires the use of an “artificial” DMU representing the average values of the outputs in the sample analyzed. This type of formulation for the weight restrictions recurring to the use of an “artificial” DMU, equal to the sample average, was originally proposed by Wong and Beasley (1990) as a complement of DMU-specific virtual weight restrictions. If instead of restricting the virtual outputs of a DMUj, as shown in (3.9), the restrictions are imposed to the average DMU (¯yr), as shown in (3.15), all DMUs are assessed with identical restrictions1. Thus, these weight re1Note that using the restrictions shown in (3.11) each DMUj0is assessed with its own weight restrictions, whose bounds depend on the value of the output yrobserved for each DMU. 62 3.4 Incorporating value judgments in composite indicators strictions in fact work as ARIs, as they are no longer DMU-specific. φr≤ur¯yr Ps r=1 ur¯yr ≤ψr(3.15) Besides the advantage of avoiding the problems associated with the DMUspecific weight restrictions previously discussed, the bounds of the restrictions become independent of the units of measurement of the outputs, as the numerator and denominator are the product of the raw weights with the output quantities. Thus, the bounds φror ψrof expression (3.15) may be interpreted as the percentual importance of the output yrin the assessment. Values of φrand ψrequal to one mean that the output yris the only one to be considered in the assessment, whereas values equal to zero mean that the corresponding output should be ignored. The ARI weight restrictions shown in (3.15) can also be added to the CI models (3.2) and (3.8). The specification that accommodates the ARI restrictions to both desirable and undesirable outputs in model (3.2) is shown in (3.16). Note that in this case the denominator no longer coincides with the normalization constraint, so it cannot be omitted.          φr≤ur¯yr Ps r=1 ur¯yr+Pl k=1 pk(Mk−¯ bk)≤ψr φk≤pk(Mk−¯ bk) Ps r=1 ur¯yr+Pl k=1 pk(Mk−¯ bk)≤ψk (3.16) Similarly, the specification of the output restrictions for the model (3.8) is shown in (3.17).          φr≤ur¯yr Ps r=1 ur¯yr+Pl k=1 pk¯ bk ≤ψr φk≤pk¯ bk Ps r=1 ur¯yr+Pl k=1 pk¯ bk ≤ψk (3.17) 63 Chapter 3 Restrictions (3.16) and (3.17) can be also generalised to restrictions to categories of indicators (Cz, z = 1, ..., q), as shown in (3.18) and (3.19), respectively. φz≤Pr∈Czur¯yr+Pk∈Czpk(Mk−¯ bk) Ps r=1 ur¯yr+Pl k=1 pk(Mk−¯ bk)≤ψz(3.18) φz≤Pr∈Czur¯yr+Pk∈Czpk¯ bk Ps r=1 ur¯yr+Pl k=1 pk¯ bk ≤ψz(3.19) In the illustrative application discussed next, we demonstrate the advantages and limitations of using virtual weight restrictions and ARIs in the context of the construction of composite indicators including both desirable and undesirable outputs. 3.5 Illustrative application This section illustrates the application of the two approaches to construct DEA-based CIs including desirable and undesirable outputs, presented in section 3.3. It also discusses the implications of using weight restrictions in this context. Our illustrative example consists of a set of 9 DMUs. To allow a graphical illustration of the models, these DMUs are assessed considering two output indicators: Y, a desirable output, and B, an undesirable output. Table 3.1 shows the data for the 9 DMUs. 64 3.5 Illustrative application Table 3.1: Output indicators for the illustrative example DMU Y(desirable) B(undesirable) A 5 7 B 25 28 C 22 15 D 10 18 E 30 30 F 20 22 G 19.80 17.80 H 20.80 19.50 I 19.08 19.66 3.5.1 CI models with desirable and undesirable outputs 3.5.1.1 Indirect approach illustration Figure 3.1 illustrates the production possibility set for the illustrative example presented in Table 3.1, corresponding to an evaluation using model (3.2). As lower values of the output indicator Bcorrespond to better performance, the technically efficient frontier is given by the segments linking DMUs A, C and E. Figure 3.1: Production possibility set for the indirect approach 65 Chapter 3 Table 3.2 shows the performance scores, peers and targets obtained for the 9 DMUs using model (3.2) and its dual (3.3). The value of Mwas set to be equal to 30 (i.e., the largest value observed for the undesirable output indicator B). The composite indicator (or efficiency score) for a DMU is given by 1/hj0at the optimal solution to model (3.2). For example, for DMU F the CI is 0.809, which corresponds to the ratio O0F O0F∗in Figure 3.1. The target that the DMU F should achieve to improve performance and reach the frontier of the production possibility set is given by the point F*, which corresponds to the value 24.725 for the output indicator Yand 20.110 for the output indicator B. The peers for DMU F are DMUs C and E, with values of λCand λEequal to 0.659 and 0.341, respectively. The values of λjprovide an indication of the degree of similarity between a DMU and its peers. Table 3.2: Composite indicator, peers and targets obtained using model (3.2), with M equal to 30 DMU CI (1/hj0) Peers (λ) Target for YTarget for B A 1 A(1) 5 7 B 0.869 C (0.153); E (0.847) 28.772 27.698 C 1 C(1) 22 15 D 0.659 A (0.401); C (0.599) 15.176 11.789 E 1 E(1) 30 30 F 0.809 C (0.659); E (0.341) 24.725 20.110 G 0.877 C (0.928); E (0.072) 22.580 16.087 H 0.880 C (0.795); E (0.205) 23.636 18.068 I 0.820 C (0.841); E (0.159) 23.273 17.387 With the transformation in the measurement scale of the undesirable output, the computation of the composite indicator uses as basis a point that no longer coincides with the origin. Instead, it corresponds to a new reference point that depends on the value of the constant Mused in model (3.2). In our illustrative example, the composite indicator is estimated considering that the projection starts at the worst value observed in the original mea66 3.5 Illustrative application surement scale of the undesirable output (30, corresponding to DMU E), as shown in Figure 3.1. As mentioned in section 3.3, the transformation consisting on subtracting the values of undesirable outputs from a large positive number (M) has an impact on the results of the DEA model. Table 3.3 presents the results that would be obtained by specifying different values for M. Note that as the value of Mincreases, the discrimination between the DMUs’ scores decreases. In addition, the improvement required for the DMUs to reach the frontier becomes more demanding for the undesirable output and less demanding for the desirable output. This effect can be seen in Figure 3.2, that shows the projections obtained using Mequal to 35 and 60. Table 3.3: Composite indicator and ranks for different values of M DMU M= 30 M= 35 M= 60 CI Rank CI Rank CI Rank A 1 1 1 1 1 1 B 0.869 6 0.880 6 0.914 6 C 1 1 1 1 1 1 D 0.659 9 0.715 9 0.844 9 E 1 1 1 1 1 1 F 0.809 8 0.824 8 0.875 8 G 0.877 5 0.887 5 0.931 4 H 0.880 4 0.890 4 0.928 5 I 0.820 7 0.834 7 0.891 7 By using a value of Mlarger than the maximum observed for the undesirable output, the DMUs’ classification, as efficient or inefficient, remains unchanged, but the DMUs’ efficiency score changes, and the ranking of the DMUs may also be different. For example, Figure 3.2 shows that when the constant Mis specified as equal to 60, DMU G is projected to the segment linking DMUs A and C instead of the segment linking C and E, where it was projected when Mwas specified as equal to 35. This led to a change 67 Chapter 3 in the efficiency ranking of DMUs, as shown in Table 3.3. Thus, using the indirect approach to construct a composite indicator, it is only possible to ensure that the assessment is classification invariant for different values of the constant M, as stated by Seiford and Zhu (2002). Figure 3.2: Sensibility analysis for different values of M 68 3.5 Illustrative application 3.5.1.2 Direct approach illustration This section illustrates the estimation of composite indicators using the Directional CI model. Figure 3.3 shows the production frontier that would be obtained for our illustrative example (Table 3.1) using the Directional CI model. The efficient frontier is defined by the segments linking O, C and E. By setting the directional vector as g= (gy,−gb) = (yrj0,−btj0), i.e. the current value of the outputs for the DMU under assessment, it is possible to simultaneously expand the desirable outputs and contract the undesirable outputs through a path that allows proportional interpretation of improvements. In order to facilitate the interpretation of the DMUs’ projection to the frontier, Figure 3.3 illustrates the directional vectors and the projection of DMUs D and F on the frontier, corresponding to points D* and F*, respectively. Note that, for each DMU, the desirable and undesirable outputs are expanded and contracted, respectively, according to a direction that corresponds to proportional changes to the original levels. Figure 3.3: Production possibility set for the direct approach Table 3.4 shows the composite indicator, peers and targets obtained using 69 Chapter 3 the Directional CI model (3.6). The value of β∗obtained at the optimal solution to the model can be interpreted as the scope for improvement of a given DMU. For example, the value of β∗for DMU F is 0.181, corresponding to FF∗ O0Fin Figure 3.3. The point F* is the target that the DMU F should achieve to become efficient, i.e., to operate at the frontier, which corresponds to the value 23.613 for the output indicator Yand 18.025 for the output indicator B. The peers for DMU F are DMUs C and E, with values of λC and λEequal to 0.798 and 0.202, respectively. Table 3.4: Composite indicator, rank, peers and targets obtained from the Directional CI model DMU βCI Rank Peers (λ) Target for YTarget for B A 0.345 0.744 8 C (0.306) 6.725 4.585 B 0.098 0.910 3 C (0.317); E (0.683) 27.462 25.242 C 0 1 1 C(1) 22 15 D 0.451 0.689 9 C (0.659) 14.505 9.890 E 0 1 1 E(1) 30 30 F 0.181 0.847 6 C (0.798); E (0.202) 23.613 18.025 G 0.126 0.888 5 C (0.963); E (0.037) 22.296 15.556 H 0.115 0.897 4 C (0.850); E (0.150) 23.200 17.250 I0.183 0.845 7 C (0.929); E (0.071) 22.567 16.063 As explained in section 3.3.2, from the factor β∗it is possible to obtain an efficiency measure, given by 1/(1 + β∗), corresponding to the best level of performance observed. Thus, the Directional CI score for the DMU F would be equal to 0.847 (= 1/(1 + 0.181)), which corresponds to the ratio O0F O0F∗in Figure 3.3. If instead of assuming weak disposability of undesirable outputs, it was imposed strong disposability, the frontier would be given by the horizontal extension of E to the Axis Y, as shown in Figure 3.4. Under this condition, the DMU F, for example, would have a score β= 0.5 and it would be projected to the point F* on the frontier with coordinates (30, 11). This 70 3.5 Illustrative application means that, in order to be producing on the frontier, DMU F should reduce by half the production of undesirable outputs whilst keeping the same production level of desirable outputs. Note that depending on the directional vector used, the projection of some DMUs to the frontier could correspond to negative values of undesirable outputs. This would be the case of DMU D, whose score under strong disposability of the undesirable output is equal to β= 2. This requires an unreasonable improvement for the undesirable output, since the DMU would be projected to point D* (30, -18). Therefore, in the context of assessments using a CI, we believe that weak disposability is more appropriate, since it avoids unrealistic projections, i.e., projections involving huge reductions to undesirable outputs and increments to desirable outputs, eventually leading to unrealistic targets corresponding to negative values of the undesirable outputs. Figure 3.4: Production possibility set for the direct approach assuming strong disposability of undesirable outputs 71 Chapter 3 regions. For this reason, we believe they are not the best option to construct composite indicators and ranks. A fair comparison requires the DMUs to be assessed based on similar conditions (Ramon et al., 2011; Hatefi and Torabi, 2010). The fact that the ARI are more conservative than virtual weight restrictions, combined with the fact that the virtual weight restrictions can suggest peer DMUs that are inefficient when evaluated with their own weight restrictions, led us to conclude that the approach based on ARI restrictions is the most appropriate. 3.6 Conclusions The traditional DEA-based composite indicator models cannot be used in the presence of both desirable and undesirable outputs, as they do not seek for reductions to the undesirable indicators. In this chapter we discussed two different approaches that can be used for the construction of CI in this context: an indirect approach, based on a traditional DEA model including a transformation in the measurement scale of undesirable outputs, and a direct approach, based on a DEA model specified with a directional distance function, that allows dealing with the undesirable outputs in their original measurement scale. In order to explain the features of the approaches discussed in this chapter we illustrated their implementation using a small example. It was demonstrated that in the indirect approach, after the transformation in the measurement scale of the undesirable outputs, the reference point used to compute the measure of performance is no longer the origin, as in standard DEA models, but a new reference point whose coordinates are equal to the positive numbers used to transform the measurement scale of the undesirable outputs. As a result, this approach does not allow proportional improvements to both desirable and undesirable outputs. Furthermore, it was shown that 78 3.6 Conclusions the results of the indirect approach are sensitive to the value of the constant used for the transformation of the measurement scale. Different values for the constant have implications both in the performance scores and ranking of the DMUs. Conversely, using the direct approach to deal with the undesirable outputs, it is possible to preserve the proportional interpretability of the improvements by setting the components of the directional vector equal to the values of the desirable and undesirable outputs of the DMU under assessment. This is an important advantage of the direct approach, and thus we argue that it is the most appropriate for constructing composite indicators in the presence of both desirable and undesirable outputs. This chapter also explored different ways to incorporate information on decision-maker preferences about the relative importance of individual indicators aggregated in the CI. The specification of two different types of weight restrictions that can be used in this context (virtual weight restrictions and ARI weight restrictions) was discussed and illustrated using a small example. We suggested an enhanced specification of the ARI weight restrictions that allows incorporating the relative importance of outputs, expressed as a percentage. This formulation also has the advantage of being independent of the units of measurement of the output indicators. The ARI weight restrictions overcome some limitations of virtual weight restrictions that are commonly used for evaluations based on CI. The problem of having peers for inefficient DMUs that are not efficient when assessed with their own set of virtual weights does not occur with the specification of the enhanced ARI restrictions, as they avoid evaluations against different frontiers. Thus, we also conclude that the new restrictions proposed here, in the form of ARI, are the most appropriate approach to reflect the relative importance of outputs in assessments involving the use of CI. 79 Chapter 3 The specification of this novel type of weight restriction and the construction of DEA-based CIs that can accommodate both desirable and undesirable outputs using a directional distance function are the two major methodological contributions of this chapter. 80 Chapter 4 An enhanced Malmquist-Luenberger index to assess productivity change in the presence of undesirable outputs 4.1 Introduction The Malmquist productivity index, introduced by Caves et al. (1982) and then developed by Fare et al. (1994b), is the most frequently used approach to assess productivity change over time. As the Malmquist productivity index is calculated based on ratios of Shephard’s distance functions, it can either have an input or output orientation. As the distance functions cannot consider simultaneous adjustments to inputs and outputs or reductions to a subset of outputs and increment to other outputs, the Malmquist index cannot accommodate assessments with undesirable outputs. In order to overcome these limitations, Chung et al. (1997) proposed an 81 Chapter 4 adaptation to the Malmquist productivity index that can accommodate undesirable outputs. The new index, named Malmquist-Luenberger (ML) index, allows DMUs to pursue simultaneous contractions of inputs and undesirable outputs and expansion of desirable outputs. Instead of estimating Shephard’s distance functions, the ML index is calculated using ratios of directional distance functions. This approach has been extensively applied in the literature to measure changes in productivity over time in the presence of undesirable outputs. A different approach that can also overcome the limitations of the Malmquist index was proposed by Chambers (1996). The authors introduced the Luenberger productivity index to assess productivity change over time using the difference of directional distance functions. While the Malmquist index focuses on either inputs contraction or outputs expansion (cost minimization or revenue maximization), the MalmquistLuenberger index and the Luenberger index can consider simultaneously inputs contraction and outputs expansion (profit maximization). Although these indices were designed to account for simultaneous improvements in inputs and outputs, they can also be specified with an output or input orientation when necessary. In this sense, we can say that the MalmquistLuenberger index and the Luenberger index encompass the Malmquist index. Boussemart et al. (2003) pointed out the three main differences between the Malmquist and Luenberger indices, which can motivate the use of one or another. The first is related to the choice of the distance function (Shephard or the directional distance functions). The second is associated with the economic motivation, which can focus on revenue/cost optimization or on profit maximization. The last difference is related to the nature of the index (multiplicative or additive). The first two issues mentioned by Boussemart et al. (2003) can also differentiate the Malmquist from the Malmquist-Luenberger 82 4.1 Introduction index. Both the Malmquist-Luenberger and Luenberger indices use the directional distance function to estimate the index, and thus, can account for revenue/cost optimization or profit maximization. The main methodological difference between the Malmquist-Luenberger and Luenberger indices is related to the multiplicative or additive nature. Several empirical applications used the ratio-based Malmquist-Luenberger index to assess productivity change (e.g. Zhang et al. (2011); Krautzberger and Wetzel (2012); He et al. (2013)). On the other hand, fewer empirical applications used the Luenberger productivity index, which measures productivity change in terms of differences rather than ratios (e.g. Epure et al. (2011); Briec et al. (2011); Williams et al. (2011)). As noted by Epure et al. (2011), although the ratio-based indices (such as Malmquist and Malmquist-Luenberger indices) are more familiar in the academic community, in the business and accounting communities the difference-based indices may be more obvious, since they measure cost, revenue, or profit differences in monetary terms. This chapter addresses the different approaches that can be used to accommodate undesirable outputs in the analysis of productivity change over time. We start from the approach proposed by Chung et al. (1997), in which the ML index is derived from a standard Malmquist index using the relationship between the directional distance function and the Shephard’s output distance function. Next, we show that an equivalent index can be derived using the relationship between the directional distance function and Shephard’s input distance function. In the context of assessments involving both desirable and undesirable outputs, the two versions of the ML index represent equally good adaptations of the Malmquist index. As the indices provide different results, we propose the use of an enhanced version of the ML index, given by the geometric mean of the two former versions of the ML indices. This approach avoids an arbitrary selection of the input or out83 Chapter 4 put Shephard’s distance function to derive the Malmquist-Luenberger index. In assessments in which improvements in both directions are required (reducing undesirable outputs and increasing desirable outputs), the Average ML proposed in this chapter has the advantage of representing more accurately the changes in DMUs’ productivity over time, as it incorporates both orientations in the computation of the productivity change score. We used a case study to compare the results obtained by the different versions of the ML indices with the Luenberger index, which is an established difference-based measure to assess productivity change over time considering simultaneous adjustments to desirable and undesirable outputs. The different indices are applied to evaluate the performance change over time of the European Commercial Transport Industry. The data concern 17 European countries in years 2005 and 2006. Based on the empirical results, we explore the relationship between the Malmquist-Luenberger index proposed in this chapter and the Luenberger index, both in terms of the values of productivity change estimates and the rankings obtained. It is shown that the Average ML index proposed in this chapter is a robust alternative to the Luenberger index, which can be used when a multiplicative index based on ratios is considered preferable to a difference based index. The limitations of the input and output ML index are highlighted by comparison to the Average ML index. The generalisation of the enhanced ML index proposed in this section to assessments involving composite indicators is straightforward, as it only requires replacing the directional distance function used for the estimation of the ML index by the CI model based on the directional distance function described in the previous section. The remainder of this chapter proceeds as follows. Section 4.2 approaches the measurement of productivity change over time using the Malmquist in84 4.2 Productivity change over time dex. Section 4.3 presents the main nonparametric indices that can be used to measure productivity change over time in the presence of undesirable outputs. These two sections constitute a literature review that provide the foundations for the enhanced Malmquist-Luenberger index proposed in section 4.4. Section 4.5 presents a graphical illustration of the productivity change indices and explores, using a real-world application, the robustness of the different productivity indices approached in this chapter. Finally, section 4.6 concludes the chapter. 4.2 Productivity change over time The Malmquist index is considered in the literature the standard approach to evaluate productivity change over time. Before showing its formulation, we start by presenting the concept of a distance function for a technology involving the use of multiple inputs to produce multiple outputs, which underlies the construction of the index. 4.2.1 Shephard’s distance functions Consider that the production technology Tmodels the transformation of inputs, denoted by x∈ <m +, into outputs, denoted by y∈ <s +, as shown in (4.1). The production technology consists of the set of all feasible input/output vectors for a certain production process. T={(x, y) : xcan produce y}(4.1) Following Shephard (1970) and Fare et al. (1994a), the output distance function for DMU j0in relation to the technology Tis defined as shown in (4.2). 85 Chapter 4 Do(x, y) = min{θ: (x, y θ)∈T}(4.2) This function gives the reciprocal of the maximum factor 1/θ by which the output vector ycan be proportionally expanded, given inputs x. This means that it corresponds to the efficiency score of DMU j0, i.e. Do(x, y)≤1. Note that Do(x, y)≤1 if and only if (x, y)∈T. In particular, Do(x, y) = 1 if and only if (x, y) is on the frontier of the technology, meaning that the production point corresponding to (x, y) is technically efficient. Assuming constant returns to scale, the efficiency of the DMU j0can be determined using the linear programming problem shown in (4.3), proposed by Charnes et al. (1978). Thus, the distance function can be estimated using a Data Envelopment Analysis (DEA) model. Fare et al. (1994b) were the first to note that input and output distance functions could be estimated using DEA models. (Do(x, y))−1= Max θ(4.3) s.t. n X j=1 yrj λj≥θyrj0r= 1, . . . , s n X j=1 xij λj≤xij0i= 1, . . . , m λj≥0j= 1, . . . , n The λjare the intensity variables. The factor 1/θ indicates the DMU’s efficiency. Similarly, the input distance function is defined as shown in (4.4)1. 1A more rigorous definition could be made by replacing min and max (which stands for minimum and maximum) with inf and sup (which stands for infimum and supremum), because the min and max may not be attained. However, in the interests of easy reading, the terms min and max are frequently used (see Coelli et al. (2005, p.49) and Fried et al. (2008, p.22)). 86 4.2 Productivity change over time Di(x, y) = max{δ: (x δ, y)∈T}(4.4) The input distance function gives the reciprocal of the minimum factor 1/δ by which the input vector xcan be proportionally contracted, given outputs y. The input technical efficiency is therefore defined as 1/Di(x, y). For the DMU j0, the input oriented efficiency can be obtained through the linear programming problem shown in (4.5), proposed by Charnes et al. (1978). This DEA model also assumes constant returns to scale. (Di(x, y))−1= Min δ(4.5) s.t. n X j=1 yrj λj≥yrj0r= 1, . . . , s n X j=1 xij λj≤δxij0i= 1, . . . , m λj≥0j= 1, . . . , n Under constant returns to scale, the following relationship holds for the distance functions: Do(x, y) = (Di(x, y))−1. 4.2.2 Directional distance functions Chambers et al. (1996), based on Luenberger (1992a,b) shortage function, proposed a directional distance function that allows a producer to scale input and outputs simultaneously along a path that is defined according to a directional vector g. The general form of the directional distance function is presented in (4.6). ~ D(x, y;gx, gy) = max {β: (x+βgx, y +βgy)∈T}(4.6) 87 Chapter 4 they can be reduced without affecting the production of undesirable outputs. This assumption can be written as shown in (4.23). (y, b)∈P(x) and ˆy≤yimply (ˆy, b)∈P(x)(4.23) Finally, it is modeled the idea that the desirable outputs are jointly produced with the undesirable outputs. This means that the only way to produce zero undesirable outputs is by producing zero desirable outputs. This idea, named null-jointness, was introduced by Shephard and Fare (1974) and it is expressed as shown in (4.24). if (y, b)∈P(x) and b= 0 then y= 0 (4.24) As shown in Chung et al. (1997), following Shephard (1970), for assessments involving both desirable and undesirable outputs, the output distance function in relation to the production technology P(x) can be defined as follows: Do(x, y, b) = min {θ: ((y, b)/θ)∈P(x)}(4.25) However, the Shephard’s output distance function shown in (4.25), seeks to increase both desirable and undesirable outputs simultaneously. One possibility to overcome this problem and credit DMUs for reductions in undesirable outputs is the specification of a directional distance function. Chung et al. (1997) extended the Chambers et al. (1996) approach, presented in section 4.2.2, to allow including undesirable outputs in the evaluation. It can be defined as shown in (4.26). ~ D(x, y, b;g) = max {β: (y, b) + βg ∈P(x)}(4.26) 94 4.3 Productivity change over time in the presence of undesirable outputs Conversely to the distance function (4.25), the directional distance function allows to simultaneously expand the desirable outputs and contract the undesirable ones. The directional distance function (4.26) can be solved by the linear programming problem shown in (4.27). It assumes constant returns to scale and satisfies the conditions (4.22), (4.23) and (4.24). ~ D(x, y, b;g) = max β(4.27) s.t. n X j=1 yrj λj≥yrj0+βgyrj0r= 1, . . . , s n X j=1 bkj λj=bkj0−βgbkj0k= 1, . . . , l n X j=1 xij λj≤xij0i= 1, . . . , m λj≥0j= 1, . . . , n The factor βcorresponds to the maximal feasible expansion of desirable outputs and contraction of undesirable outputs that can be achieved simultaneously. Figure 4.2 shows an illustration of the production frontier that would be obtained using model (4.27) with the components of the directional vector set as g= (yrj0,−bkj0), i.e. the current value of the outputs for the DMU under assessment. As explained in section 3.3.2, by specifying g= (yrj0,−bkj0) the desirable and undesirable outputs are expanded and contracted, respectively, according to a direction that corresponds to proportional changes to the original levels for each DMU. In order to facilitate the interpretation of the DMUs’ projection to the frontier, Figure 4.2 illustrates the directional vectors for DMUs A and B (gA and gB) and the DMUs’ projection on the frontier, corresponding to points A* and B*, respectively. The value of the directional distance function (corresponding to an ineffi95 Chapter 4 Figure 4.2: Production possibility set for the directional distance function ciency measure) is given by βin the optimal solution of model (4.27). For DMU A it corresponds to the ratio AA∗ AA0(taking the axis Xas reference), or equivalently, by the ratio AA∗ AA00 (taking the axis Yas reference). Similarly, for DMU B it is given by the ratio BB∗ BB0, or equivalently, by the ratio BB∗ BB00 . The value of the directional distance function in assessments with undesirable outputs is always greater than or equal to zero, with values greater than zero signalling the existence of inefficiencies, and zero meaning that the production occurs on the frontier of the production possibility set. 4.3.2 The Malmquist-Luenberger index Chung et al. (1997) defined the Malmquist-Luenberger index with the aim to define a productivity index that is able to credit a firm for reductions in undesirable outputs without requiring changes in their original measurement scale, as shown in (4.28). It is derived from the output oriented Malmquist index, shown in (4.14), using the equivalence between the Shephard’s output distance function and the directional distance function, shown in (4.9). 96 4.3 Productivity change over time in the presence of undesirable outputs MLt,t+1 o=h(1+ ~ Dt(xt,yt,bt;yt,−bt)) (1+ ~ Dt(xt+1,yt+1,bt+1;yt+1,−bt+1)) (1+ ~ Dt+1(xt,yt,bt;yt,−bt)) (1+ ~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1)) i1 2 (4.28) The Malmquist-Luenberger index can be decomposed as follows: ECt,t+1 o=1+ ~ Dt(xt,yt,bt;yt,−bt) 1+ ~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1)(4.29) and TCt,t+1 o=h(1+ ~ Dt+1(xt,yt,bt;yt,−bt)) (1+ ~ Dt(xt,yt,bt;yt,−bt)) (1+ ~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1)) (1+ ~ Dt(xt+1,yt+1,bt+1;yt+1,−bt+1)) i1 2 (4.30) The estimation of within-period and mixed-period directional distance functions used to calculate the Malmquist-Luenberger index is obtained using model (4.27), assuming a directional vector equal to g= (yrj0,−bkj0). While the value of the function in the within-period assessment is always greater than or equal to zero, in the mixed-period it can be smaller, equal or greater than zero. Values smaller than zero occur when the output quantities observed in one period are not feasible in the technology corresponding to the other period. Values of the Malmquist-Luenberger index greater than one correspond to productivity improvements, whereas values smaller than one signal productivity decline. Note that the Malmquist-Luenberger index is calculated using directional distance functions specified with a vector that seeks for simultaneous improvements to both desirable and undesirable outputs. The value of the ML index (4.28) is identical to the output oriented Malmquist index when the directional vector is specified to improve desirable outputs only. In case the directional vector is specified to seek for improvements in both types 97 Chapter 4 of outputs, the directional distance function is not equivalent to Shephard’s distance function, as the latter cannot account for simultaneous changes to desirable and undesirable outputs. Therefore, the Malmquist-Luenberger and the Malmquist indices can only be considered comparable, as noted by Chung et al. (1997). 4.3.3 The Luenberger productivity index As an alternative to the Malmquist-Luenberger indices previously discussed, it is possible to use the Luenberger productivity index to assess productivity change over time. Chambers (1996) introduced the Luenberger productivity index as a difference of directional distance functions. The Luenberger productivity index for an assessment involving inputs xij (i= 1, ..., m), desirable outputs yrj (r= 1, ..., s) and undesirable outputs bkj (k= 1, ..., l), is defined as follows: Lt,t+1 =1 2[~ Dt+1(xt,yt,bt;yt,−bt)−~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1) +~ Dt(xt,yt,bt;yt,−bt)−~ Dt(xt+1,yt+1,bt+1;yt+1,−bt+1)] (4.31) Lt,t+1 measures the productivity change between time periods tand t+ 1. Following the idea used to construct the Malmquist index, in order to avoid an arbitrary choice between the base periods, the Luenberger productivity index is given by an arithmetic mean of the indices corresponding to the periods t(the first difference) and t+ 1 (the second difference inside the square brackets). As explained in Fare and Grosskopf (2005), the Luenberger productivity index can be addictively decomposed in two components: efficiency change and technological change, as shown in (4.32) and (4.33), respectively. This decomposition is similar to the one proposed by Fare et al. (1994b) in the context of the Malmquist index. 98 4.4 An enhanced version of the Malmquist-Luenberger index LECt,t+1 =~ Dt(xt,yt,bt;y,−b)−~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1)(4.32) LTCt,t+1 =1 2[~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1)−~ Dt(xt+1,yt+1,bt+1;yt+1,−bt+1) +~ Dt+1(xt,yt,bt;yt,−bt)−~ Dt(xt,yt,bt;yt,−bt)] (4.33) Both the Luenberger productivity index and its components signal improvements with values greater than zero, and declines in productivity with values smaller than zero. Values equal to zero indicate no productivity change. The change in relative efficiency between periods t and t+1 can be interpreted as the change in the distance between the observed production level and the maximum potential production. As in the Malmquist index and the ML indices, improvements in the efficiency change are evidence of catching up to the frontier, and the technological change component captures the shift in technology between the two periods of time considered. 4.4 An enhanced version of the Malmquist-Luenberger index In this chapter, we propose an alternative formulation of the MalmquistLuenberger index, as shown in (4.34), which is derived from the input oriented Malmquist index, shown in (4.18), using the equivalence between the distance functions shown in (4.11). MLt,t+1 i=h(1−~ Dt(xt+1,yt+1,bt+1;yt+1,−bt+1)) (1−~ Dt(xt,yt,bt;yt,−bt)) (1−~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1)) (1−~ Dt+1(xt,yt,bt;yt,−bt)) i1 2 (4.34) The components of the Malmquist-Luenberger index, shown in (4.34), are as follows: 99 Chapter 4 ECt,t+1 i=1−~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1) 1−~ Dt(xt,yt,bt;yt,−bt)(4.35) and TCt,t+1 i=h(1−~ Dt(xt,yt,bt;yt,−bt)) (1−~ Dt+1(xt,yt,bt;yt,−bt)) (1−~ Dt(xt+1,yt+1,bt+1;yt+1,−bt+1)) (1−~ Dt+1(xt+1,yt+1,bt+1;yt+1,−bt+1)) i1 2 (4.36) This alternative version of the Malmquist-Luenberger index (MLt,t+1 i), which is input-oriented, is obtained using a directional distance function specified with a vector equal to g= (yrj0,−bkj0). This index would only be equivalent to the input oriented Malmquist index if the directional vector was specified in order to reduce only the undesirable outputs, keeping the desirable outputs with their current value. Note that to preserve the interpretation of the productivity indices such that values greater than one mean improvements, the input oriented versions of the Malmquist index and Malmquist-Luenberger index use the inverse of the input distance function, corresponding to an efficiency measure. For the output oriented indices, the output distance function is equal to the efficiency measure. The Malmquist-Luenberger indices (MLt,t+1 iand MLt,t+1 o) are derived from the Malmquist indices (MIt,t+1 iand MIt,t+1 o) based on the relations between the directional distance function and the Shephard’s distance functions, presented in (4.9) and (4.11). The estimation of the Shephard’s distance functions can also be visualized in Figure 4.2, which was previously used in section 4.3.1 to illustrate the directional distance function. The value of the Shephard’s output distance function (efficiency measure) used to calculate the MLt,t+1 oindex for the DMU A, is given by the ratio AA0 A∗A0(corresponding to the relation Do(xt, yt)=1/(1 + ~ Dt(xt, yt, bt;yt,−bt))). On the other 100 4.4 An enhanced version of the Malmquist-Luenberger index hand, the inverse of the Shephard’s input distance function (efficiency measure) used to calculate the MLt,t+1 iindex for the DMU A, is given by A∗A00 AA00 (corresponding to the relation (Dt i(xt, yt))−1= 1 −~ Dt(xt, yt, bt;yt,−bt)). Although both ratios are acceptable to estimate an efficiency measure for DMU A (as they only differ in the axis used as reference to estimate efficiency), they do not result in the same values. As a consequence, the results of the input-based and output-based Malmquist-Luenberger indices are different, as they depend on the relation used to convert the directional distance function in a Shephard’s distance function. Recall that, as noted by Chung et al. (1997), the Shephard’s distance functions and the directional distance functions are not equivalent in the presence of both desirable and undesirable outputs, so this implies that the input-based and output-based ML indices are not identical. Following the idea of Fare et al. (1994b), that in order to avoid the use of an arbitrary reference technology proposed the estimation of the Malmquist index as the geometric mean of two Malmquist indices using as reference the technology in different time periods (tand t+ 1), we propose an enhanced version of the Malmquist-Luenberger index that avoids an arbitrary choice between the input or output Shephard’s distance function used to derive the Malmquist-Luenberger index. The new index is specified as the geometric mean of the indices MLt,t+1 oand MLt,t+1 i, as shown in (4.37). MLt,t+1 av =hMLt,t+1 o×MLt,t+1 ii1 2(4.37) The new Malmquist-Luenberger index can be decomposed as follows: ECt,t+1 av =hECt,t+1 o×ECt,t+1 ii1 2(4.38) and 101 Chapter 4 TCt,t+1 av =hTCt,t+1 o×TCt,t+1 ii1 2(4.39) The three versions of the ML indices (4.28), (4.34) and (4.37), corresponding to the input, output or average indices, indicate improvement, stagnation and decline in productivity, by values greater, equal or smaller than one, respectively. Next we illustrative graphically how the productivity indices previously discussed (the three versions of the Malmquist-Luenberger index and the Luenberger index) measure changes in productivity over time in the presence of undesirable outputs. The advantages and limitations of each approached are also discussed. 4.5 Empirical examples 4.5.1 Graphical illustration of the productivity change Our illustrative example is based on a set of 6 DMUs assessed in two time periods: 0 and 1. To allow a graphical illustration of the models, these DMUs are assessed considering two output indicators: Y, a desirable output, and B, an undesirable output, and a unitary input underlying the evaluation of every DMU in both time periods. Table 4.1 shows the data for the 6 DMUs. Table 4.1: Data for the illustrative example DMUs X0Y0(desirable) B0(undesirable) X1Y1(desirable) B1(undesirable) A 1 8 12 1 12 9 B 1 22 30 1 25 34 C 1 18 10 1 30 13 D 1 7 20 1 10 21 E 1 30 34 1 35 29 F 1 20 23 1 24 24 Figure 4.3 illustrates the production possibility sets for the illustrative example, in time periods 0 and 1, corresponding to an evaluation using the 102 4.5 Empirical examples directional distance function model (4.27). The projections to the frontier for DMUs D and F, in time period 0 and 1, are also represented. Note that the DMUs are projected to the frontier according to a direction that corresponds to proportional changes to the original output values of each DMU. As lower values in the output indicator Bcorrespond to better performance, the technically efficient frontier is given by the segments linking the origin and the DMUs C0and E0for the first time period, and the origin and the DMUs C1and E1for the second time period. Figure 4.3: Production possibility set for the illustrative example in time periods 0 and 1 The values of the directional distance functions (that corresponds to the inefficiency measure βin model (4.27)) used to calculate the MalmquistLuenberger indices and the Luenberger index are shown in Table 4.2. For DMU F, in time period 0, the value of the directional distance function ~ D0(x0, y0, b0;y0,−b0) is 0.1429 and ~ D1(x0, y0, b0;y0,−b0) is 0.4526, corresponding to the proportional improvement required to the outputs of this 103 Chapter 4 4.6 Conclusions This study addressed the different approaches that have been used to assess productivity change in the presence of undesirable outputs: the ratio-based Malmquist-Luenberger index and the difference-based Luenberger index. It was demonstrated that the ML index can be derived from a standard Malmquist index using the relationship between the directional distance function and the Shephard’s output distance function, or equivalently, using the relationship between the directional distance function with the Shephard’s input distance function. The two versions of the ML indices represent equally good adaptations of the Malmquist index for assessments involving both desirable and undesirable outputs. In order to avoid the need to arbitrarily choose one of the measures, it was proposed a new version of the ML index, given by the geometric mean of these alternative versions of the ML indices. Both the Malmquist-Luenberger and Luenberger indices are estimated using directional distance functions, and thus they can accommodate undesirable outputs without requiring changes in their original measurement scale. The main methodological difference between the Malmquist-Luenberger and Luenberger indices, which can motivate the use of one or another, is related to their multiplicative or additive nature. In cases in which the researcher has a preference for ratio-based indices, and the assessment involves simultaneous improvements to inputs and desirable and undesirable outputs, we suggest using the new version of the Malmquist-Luenberger index proposed in this study, as it has the advantage of incorporating both orientations in the computation of the productivity change score, and thus represent more accurately the changes in DMUs features. 110 4.6 Conclusions Using an empirical example, we compared the results obtained by the different versions of the ML indices with the Luenberger index, which is an established difference-based measure to assess productivity change over time considering simultaneous adjustments to inputs and desirable and undesirable outputs. The empirical results suggested that the Average ML index is more aligned with the Luenberger index than the other two versions of the ML index, both in terms of the relationship between the values of the productivity estimates as well as in terms of the rankings obtained. The results also showed that the logarithm of the Average Malmquist-Luenberger index is approximately equal to the Luenberger index. Although the ML and the Luenberger indices discussed in this chapter were originally proposed for assessments involving the conversion of inputs to desirable and undesirable outputs, they can be easily adapted for the context of composite indicators. By assuming a unitary level of inputs for all DMUs in the model (4.27), used to calculate both the ML and the Luenberger indices, it becomes the Direction CI model presented in (3.6), as explained in section 3.3.2. 111 Chapter 4 112 Chapter 5 Benchmarking countries’ environmental performance 5.1 Introduction Environmental concerns have increased dramatically in the past few years and are now among the most serious challenges affecting people’s wellbeing. Besides the climate change that has been widely discussed, other environmental problems such as local air and water pollution, soil erosion, water scarcity, deforestation, and loss of biodiversity are also becoming more serious (World Bank, 2008). Countries are facing new challenges to control their waste production and to reduce the consumption of natural resources, in order to achieve the environmental targets imposed by international agreements such as the Kyoto Protocol (United Nations, 1998) or the European Union climate and energy package (European Union, 2008). In this context, it is imperative that countries become able to monitor their environmental performance in order to understand how they are doing compared to others and to identify the potential for improvements. 113 Chapter 5 Environmental performance assessments are often conducted using environmental indicators that are able to measure the pressures on the environment, to appraise the state of the ecosystem and to evaluate the impacts on human activity resulting from changes in environmental quality. These indicators usually measure particular features of the environment and provide a starting point for performance assessments. It is also common to use composite indicators (CI) to aggregate several individual indicators in a summary measure of performance. Although the indicators, individual or aggregated, can be extremely useful to guide discussions on environmental issues and attract public interest, they do not provide guidelines that countries should follow to improve performance. This chapter aims to assess countries environmental performance using an enhanced CI model that, besides assigning a summary measure of performance for each country, can be used for benchmarking purposes. The CI is defined based on the Data Envelopment Analysis (DEA) technique. In the context of an environmental performance assessment, the DEA models can be used either to measure the environmental efficiency of a given set of units, i.e. the ability of convert inputs to outputs, or to provide an environmental effectiveness measure, which aggregates several output indicators in a CI (i.e. looking only at the achievements, rather than the conversion of inputs to outputs). Irrespectively of the approach followed in the environmental performance assessment, both desirable and undesirable outputs may be present. Since the standard DEA models rely on the assumption that outputs are maximized, the undesirable outputs must be dealt with in order to be accommodated in a DEA formulation. As discussed in chapter 3, two different approaches can be followed to treat undesirable outputs, a direct and an indirect approach. This chapter illustrate the application of the indirect approach, which is based on a transformation of the measurement scale of the undesirable outputs. Moreover, 114 5.2 Review of composite indicators for environmental performance assessment it is illustrated the use of the assurance region type I weight restrictions, described in section (3.4.2), to incorporate in the CI model information of the relative importance of indicators, expressed as a percentage. 5.2 Review of composite indicators for environmental performance assessment In the past few years, efforts to assess environmental performance of organizations, cities and countries have generated a large number of indicators related to gas emissions, water quality, green space area, waste production, among others. Due to the large amount of individual indicators available, the process of analyzing and understanding the information available becomes difficult. Therefore, many of these indicators are often aggregated into composite indicators to gain an overall picture of performance that can be used by decision makers for planning and control purposes. According to the Organisation for Economic Co-operation and Development report by Nardo et al. (2008), the construction of composite indicators involves several stages: the selection of sub-indicators, the treatment of missing values, the understanding of the data, and the specification of the weights for the sub-indicators to be used in the aggregation model. Concerning the selection of sub-indicators, they should be selected according to their analytical soundness, measurability, coverage, and relevance to the phenomenon being assessed. Regarding to the treatment of missing values, there are two possibilities: to impute values to replace the missing fields or to remove the whole observation with missing data from the analysis. Concerning the understanding of the data, an exploratory analysis should be performed to study the overall structure of the dataset. It is necessary to understand the relationship among indicators and observations, and to identify outliers and extreme values that can introduce bias in the results. 115 Chapter 5 Concerning the specification of weights for the sub-indicators, they can be specified based on quantitative methods or expert judgment. The weights can reflect policy priorities, or reward the factors deemed to be more important to the performance assessment. The choice of weights inevitably impacts the results of the CI, so the method should be robust to avoid undermining the credibility of the CI results. Since the specification of the weights is often subject of criticism and disagreement, most composite indicators rely on equal weighting to minimize the subjectivity. One such example is the calculation of the Human Development Index (United Nations Development Program, 2000), which is based on the use of equal weights for its component indices, which reflect longevity (measured by life expectancy at birth), education attainment (measured by adult literacy and enrolment rate) and standard of living (measured by GDP per capita). In the context of environmental performance assessments, the Climate Change Performance Index (CCPI), and the Environmental Performance Index (EPI) are examples of well-established composite indicators that aggregate individual indicators in a summary measure. The CCPI is measured via 12 different indicators, which can be classified in the following categories: emissions trend, emissions level and climate policy. This index compares countries that together are responsible for more than 90% of global energy-related CO2emissions. The CCPI needs the predefinition of weights for the indicators. The countries ranking is calculated from the weighted average of the scores achieved by the countries evaluated in the indicators considered (Burck et al., 2009). The EPI provides a global index based in 25 indicators, grouped in 10 categories, covering two core objectives: Environmental Health and Ecosystem Vitality. The first core objective measures the environmental effects on human health, whereas the second measures the state of the ecosystem and the 116 5.2 Review of composite indicators for environmental performance assessment natural resources management. Each of these two core objectives contributes with a weight of 50% to the overall EPI score. The EPI construction requires a predefinition of weights and targets for the indicators. The weights assigned to the indicators are determined through expert judgment. For each country and each indicator, a proximity-to-target value is calculated based on the gap between the current results presented by each country and a target previously identified. The targets are defined based on four sources: treaties or other internationally agreed goals, standards defined by international organizations, national regulatory requirements, or expert judgment. After defining the indicators’ weights and the proximity-to-target of each country, a weighted average is calculated to obtain the EPI score for each country (Emerson et al., 2010). The countries’ performance assessment presented in our study was conducted using the indicators that underlie the estimation of the EPI 2010. The EPI indicators and weights are reported on Table 5.1. For further details on the specification of the EPI indicators see the metadata information in Emerson et al. (2010). Both the EPI and CCPI use a weighted average to provide an overall measure of performance, and rely on expert opinion to specify the weights. An alternative to overcome the difficulties in the selection of weights is to use Data Envelopment Analysis (DEA) to determine the weights. Using DEA, individual indicator weights result from an optimizing process, based on linear programming, so they are less prone to subjectivity and controversy. Although the DEA technique has been used to construct CI in different fields, such as the evaluation of urban quality of life (Morais and Camanho, 2011), human development (Mahlberg and Obersteiner, 2001; Despotis, 2004, 2005), social deprivation (Zaim et al., 2001) , technology achievement (Cherchye et al., 2008), monetary aggregation (Sahoo and Acharya, 2010) or the financial soundness of construction companies (Horta et al., 2010), its application for environmental performance assessment is scarce. This chapter 117 Chapter 5 Table 5.1: EPI indicators Policy Categories Indicators Weight (w) Environmental burden of disease Disability Life Adjusted Years 25% Water (effects on humans) Access to adequate sanitation 6.3% Access to drinking water 6.3% Air pollution (effects on humans) Indoor air pollution 6.3% Outdoor air pollution - Urban Particulates 6.3% Air Pollution (effects on ecosystem) Ozone Exceedance 0.7% Non-methane volatile organic compound emissions 0.7% Sulfur dioxide emissions 2.1% Nitrogen oxides emissions 0.7% Water (effects on ecosystem) Water quality index 2.1% Water stress index 1.0% Water scarcity index 1.0% Biodiversity e Habitat Biome protection 2.1% Critical habitat protection 1.0% Marine protection 1.0% Forestry Growing stock change 2.1% Forest cover change 2.1% Fisheries Marine trophic index 2.1% Trawling intensity 2.1% Agriculture Agricultural water intensity 0.8% Agricultural subsidies 1.3% Pesticide regulation 2.1% Climate Change Greenhouse gas emissions per capita 12.5% Industrial greenhouse gas emissions intensity 6.3% CO2emissions per electricity generation 6.3% contributes to the literature in this field by proposing a methodology to evaluate countries environmental performance based on the construction of a composite indicator whose weights are specified using DEA. 5.3 Methodology Three main features of the DEA technique motivated its use in this chapter to define a CI to assess countries environmental performance. The first is related to the ability to identify best-practice peers corresponding to countries with a similar profile to that of the country under assessment. The second is related to the possibility of assigning weights to individual indicators recurring to optimization. This procedure has the advantage of being less prone to subjectivity, and allowing the identification of the areas in which coun118 5.3 Methodology tries have good or bad performance. Finally, DEA is able to handle data measured in different measurement scales and thus it is possible to use the raw data corresponding to each EPI indicator, without prior normalisations or conversions to similar measurement scales. In the environmental performance assessment presented in this chapter, the EPI indicators shown in Table 5.1 were included as outputs of the CI model (3.2), developed in chapter 3. As all outputs were measured as ratios or indexes, we considered an identical input level for all DMUs. In a first moment, in order to establish a ranking of countries environmental performance, we fixed the weights of all indicators in the CI model to ensure that all countries are evaluated using the same criteria (i.e., using common weights). We adopted weight restrictions that mimic the value judgments implicit in the EPI, which are based on expert opinion (see win Table 5.1). The use of common weights allows obtaining a robust ranking of countries, as it prevents obtaining a high performance score only due to a judicious choice of weights. Furthermore, the use of a unique weighting system for all DMUs improves the discrimination power of the performance assessment. The weight restrictions imposed to the CI model (3.2), leading to the use of a common weighting system, are shown in (5.1).          ur¯yr Ps r0=1 ur0¯yr0+Pl k0=1 pk0(Mk−¯ bk0)=wrr= 1, ..., s pk(Mk−¯ bk) Ps r0=1 ur0¯yr0+Pl k0=1 pk0(Mk−¯ bk0)=wkk= 1, ..., l (5.1) The rational underlying the specification of these restrictions is as follows. We consider an artificial DMU whose outputs are equal to the average value of each output variable (i.e., EPI indicator) in the sample. For the artificial DMU, the virtual weight of a desirable output (r) or undesirable output (k) 119 Chapter 5 the EPI, was observed in a country from Central America (Costa Rica). For the remaining countries, a score lower than 100% indicates that there is potential for improvement, supported by a comparison with Costa Rica, which is the country that presented the maximum score in this cluster. The ranking presents a satisfactory level of discrimination between the countries, which was possible due to the imposition of weight restrictions in the DEA model. As the performance assessment conducted in our study used a system of weights that mimics the one used in the construction of the EPI, it is possible to assess the robustness of the new approach proposed in this chapter by comparing its results with the EPI ranking. Table 5.4 presents the scores and the rank position of countries obtained using DEA and the EPI methodologies. Note that while the EPI evaluated the overall set of countries (163 countries), we evaluated the performance within clusters, so the comparison between the rank positions cannot be direct. For the countries in cluster C2 (43 countries), the Spearman’s rank correlation coefficient between the DEA and EPI rankings is 0.681. In order to test whether the observed correlation is significantly different from zero we performed the Spearman’s rank correlation test. The p-value of the test is close to zero (p-value=0.0000), so the null hypothesis that there is no relation between the results of the approaches was rejected. The correlation analysis for the other three clusters was also tested, and for all clusters were found a significant positive correlation between the approaches. These results are presented in Appendix C. Although the EPI and DEA methodologies use similar indicators and weighting systems, differences in results would be expected as the data used in the two approaches were subject to different normalization procedures and treatment of extreme values and outliers. In the construction of the EPI 126 5.4 Results and discussion Table 5.4: Environmental performance scores and ranks for countries of cluster C2 Country DEA EPI Score Rank Score Rank Costa Rica 100% 1 86.4% 3 Ecuador 98.8% 2 69.3% 30 Iceland 98.7% 3 93.5% 1 France 97.9% 4 78.2% 7 Colombia 97.6% 5 76.8% 10 Germany 97.4% 6 73.2% 17 Portugal 97.3% 7 73.0% 19 Sweden 96.0% 8 86.0% 4 Italy 95.8% 9 73.1% 18 Dominican Republic 95.7% 10 68.4% 36 United Kingdom 94.7% 11 74.2% 14 Spain 94.7% 12 70.6% 25 Cuba 94.5% 13 78.1% 9 New Zealand 93.4% 14 73.4% 15 Panama 93.0% 15 71.4% 24 Norway 92.9% 16 81.1% 5 Japan 92.7% 17 72.5% 20 Switzerland 91.7% 18 89.1% 2 Latvia 91.5% 19 72.5% 21 Finland 91.2% 20 74.7% 12 Lithuania 90.6% 21 68.3% 37 Denmark 90.5% 22 69.2% 32 Romania 90.3% 23 67.0% 45 Canada 89.8% 24 66.4% 46 Croatia 89.7% 25 68.7% 35 Ireland 89.4% 26 67.1% 44 Austria 89.3% 27 78.1% 8 Malaysia 89.0 28 65.0% 54 Russia 88.9% 29 61.2% 69 Netherlands 88.0% 30 66.4% 47 Slovakia 87.2% 31 74.5% 13 Hungary 86.9% 32 69.1% 33 Bulgaria 86.9% 33 62.5% 65 Poland 86.8% 34 63.1% 63 Estonia 86.8% 35 63.8% 57 Slovenia 86.6% 36 65.0% 55 Greece 86.4% 37 60.9% 71 Czech Repuplic 85.9% 38 71.6% 22 Malta 85.5% 39 76.3% 11 South Korea 84.9% 40 57.0% 94 Belgium 83.9% 41 58.1% 88 Luxembourg 81.5% 42 67.8% 41 Cyprus 78.4% 43 56.3% 96 the raw data is transformed in a proximity-to-target value that is calculated based on the gap between the values of the indicators for each country and 127 Chapter 5 a target previously identified. This transformation works as a normalization process that converts all indicators to comparable measurement scales. In addition, for some indicators, a logarithmic transformation is employed to increase discrimination, and a winsorization process is used to trim the tails of distributions that presented extreme values or outliers. In our approach the performance scores were obtained based on the raw data. The DEA approach does note requires any normalization of the indicators values prior to the assessment. In order to analyze the strengths and weaknesses of each country we relaxed the fixed weight restrictions, allowing a flexibility of c= 0.4 around the indicator weight w, as shown in expression (5.2). Different levels of flexibility (c) were tested before choosing c= 0.4. Allowing more flexibility (i.e., for higher values of c), several countries were able to obtain the maximum score, so the level of discrimination for the performance assessment would not be satisfactory. On the other hand, for lower values of c, the weights selected by the countries would become similar, such that it would be difficult to identify the categories of indicators for which the countries are doing better. Through the virtual weights chosen by each country, within the limits allowed, we are able to identify the areas in which countries are specialized and have better environmental performance. Figure 5.1 shows the results obtained for Ireland, the country selected for illustration. The pie chart (a) shows the virtual weight by category of indicators resulting from the use of the fixed weights wspecified by the EPI assessment. The pie chart (b) has the virtual weights selected by Ireland for each category of indicators, given the boundaries of flexibility allowed (c= 0.4). Allowing this weight flexibility Ireland obtained a score of 95.5%. Ireland selected lower virtual weights for the following categories: Environmental burden of disease,Climate Change and Biodiversity & Habitat. These are the categories in which Ireland needs to improve its performance 128 5.4 Results and discussion Figure 5.1: (a) Fixed weights for EPI categories and (b) contributions of the categories to the CI for Ireland the most. On the other hand, Ireland selected higher virtual weights for the categories Water (effects both on the humans and ecosystem),Air pollution (effects both on humans and humans),Agriculture and Forestry. This provides evidence that Ireland is specialized in these categories. In particular for the categories related to Water and Air pollution, Ireland selected the maximum possible weight. Note that most countries from cluster C2 have good performance in indicators related to these categories, as shown by the results of the decision tree analysis reported in Table 5.2. This methodology can also be used to point out, for a country with poor performance, the peers with a similar environmental profile that it should look to improve its performance for each indicator. Table 5.5 presents the values of the indicators of Ireland and its peers. The values of the λjprovide an indication of the degree of similarity between Ireland (country under assessment) and each peer. Based on the values presented by the peer countries it is possible to identify where the best practice examples for Ireland can be found for each of the 129 Chapter 5 Table 5.5: Peers for Ireland Categories Indicators Ireland Peers France Iceland Costa Rica λ=0.689 λ=0.255 λ=0.056 Environmental burden of disease Disability Life Adjusted Years* (O1) 215.34 214 218.34 211 Water (effects on humans) Access to adequate sanitation (O2) 100 100 100 96 Access to drinking water (O3) 100 100 100 98 Air Pollution (effects on humans) Indoor air pollution* (O4) 90 90 90 82.26 Outdoor air pollution - Urban Particulates* (O5) 130.32 132.43 127.66 109.66 Air Pollution (effects on ecosystem) Ozone Exceedance* (O6) 86.13 85.11 86.14 86.14 Non-methane volatile organic compound emissions* (O7) 56.67 52.94 55.05 55.67 Sulfur dioxide emissions* (O8) 93.40 93.31 73.57 93.65 Nitrogen oxides emissions* (O9) 75.07 74.23 67.90 74.59 Water (effects on ecosystem) Water quality index (O10) 91.91 86.51 100 47.72 Water stress index* (O11) 68.55 60.16 67.63 68.55 Water scarcity index* (O12) 4.87 4.87 4.87 4.87 Biodiversity & Habitat Biome protection (O13) 0.93 10 8.73 10 Critical habitat protection (O14) 7.14 50 7.14 75 Marine protection (O15) 0.38 60.53 38.53 56.27 Forestry Growing stock change (O16) 109.36 109.36 111.11 100.4 Forest cover change (O17) 6.12 4.52 8.12 4.32 Fisheries Marine trophic index (O18 ) 4.43 4.75 3.29 6.66 Trawling intensity* (O19) 39.01 75.2 46.51 98.25 Agriculture Agricultural water intensity* (O20) 398.91 396.99 398.91 397.64 Agricultural subsidies* (O21) 35.87 41.94 6.44 54.54 Pesticide regulation (O22) 20 21 20 18 Climate Change Greenhouse gas emissions per capita* (O23) 30.23 36.66 45.89 43.89 Industrial greenhouse gas emissions intensity* (O24) 222.62 202.93 208.96 202.82 CO2emissions per electricity generation* (O25) 845.15 1258.8 1347.55 1277.04 *non-isotonic output indicators 25 indicators. For example, considering specifically the output indicator Disability Life Adjusted Years (O1), which is related with premature death, Ireland can improve its performance applying the policy of Iceland (peer country which presented the best value in this indicator). If we consider the indicator related to Greenhouse gas emissions (O23), Iceland and Costa Rica provide good examples to learn from. A similar analysis can be done for the others indicators. Table 5.5 highlights for each indicator the best value observed in the peers. The identification of best practices for each indicator can help decision makers to define good environmental policies to improve 130 5.5 Conclusions the performance of countries. Allowing for flexibility in the choice of weights, corresponding to a value of c= 0.4, the countries that are examples of best practices in cluster C2 are: Colombia, Costa Rica, Dominican Republic, Ecuador, France, Germany, Iceland, Italy, Portugal and Sweden. 5.5 Conclusions While the climate change and warming issues are usually given more emphasis by the media in environmental reports, the world is, in fact, facing several other environmental problems that threaten the wellbeing of people on a global scale. In order to effectively tackle complex environmental issues, decision makers must first understand exactly where their countries stand in terms of environmental performance, to be able to design the steps that should be taken to improve performance. This chapter conducted a data-driven assessment, that took into account different environmental characteristics, to deliver a robust performance assessment and suggest directions for improvement. Firstly, by employing cluster analysis we were able to group countries into relative homogeneous groups and assess the environmental performance inside these groups. This ensured that each country was compared to other countries with similar features. Moreover, using a decision tree it was possible to identify the features that the countries from the same cluster share. Then, in order to conduct the assessment of countries’ environmental performance, we applied the approach developed in section 3.3.1 for the construction of a composite indicator in the presence of undesirable outputs. The CI model was used with weight restrictions corresponding to assurance regions type I, proposed in section 3.4.2. These restrictions are able to reflect in percentage terms the relative importance of individual indicators. 131 Chapter 5 The environmental performance assessment was conducted in two stages. First, using the ARI weight restrictions, we were able to create a ranking of countries, where all countries were evaluated against a unique frontier based on a common system of weights. The use of common weights enabled a fairer comparison of countries performance, as it prevented the countries to obtain a high score only due to a careful choice of weights. In a second moment, we are concerned with performance management and its improvement. By allowing a range of flexibility for the value of the indicators’ weights, it was possible to identify the environmental strengths and weaknesses of each country. The peers with similar features to the low-performing countries were also identified. These peers provide examples of good environmental practices that the countries with worse performance should follow to improve performance. The information provided in this chapter can support decision makers in understanding where their countries stand in terms of environmental performance, and guide the definition of environmental policies leading to performance improvements. 132 Chapter 6 The assessment of cities’ livability integrating human wellbeing and environmental impact 6.1 Introduction The rapid growth of urban centers and the increasing demand for products and services have led to several social, economic and environmental challenges. Overcrowding, insecurity, unemployment, natural resources depletion and pollution are just a few of the factors that deter human wellbeing and environmental quality. Governments have devoted considerable attention to these issues that directly affect livability and sustainable development of cities. In this study we address the assessment of livability in European cities. The assessment of cities’ livability and the evolution of their performance compared to others play an important role in urban planning and management. Such efforts are expected to lead to better standards of human wellbeing without compromising environmental sustainability in the long-term. 133 Chapter 6 National and local authorities are increasingly supporting efforts to better understand the cities’ progress in terms of economic development, sustainability and livability. Examples of efforts done on this direction are the reports “State of Australian Cities” by the Australian Department of Infrastructure and Transport (2012), and the “State of the English Cities” by the English Department for Communities and Local Government (2006). While it is becoming widely accepted that livability is an increasingly important topic in social sciences, there are components of livability which are not yet sufficiently elaborated. The English Department for Communities and Local Government (2006) report claims that much work remains to develop in order to arrive at an effective method for assessing and monitoring livability. In particular, further research is needed in order to define the appropriate weights to be assigned to the components used in livability assessments. Although these efforts to assess cities’ livability are extremely useful to understand where cities stand and to provide a starting point for discussions on issues affecting urban livability, they are not effective in providing guidelines that cities should follow to improve livability. In this context, the purpose of this chapter is to develop a methodology for conducting a fair and data-driven assessment of cities’ livability taking into account several dimensions, in order to provide viable guidelines for improvement. This was achieved by using a composite indicator that, besides of providing an overall measure of performance for each city, enables benchmarking in such a way that it becomes possible to identify the strengths and weaknesses of each city, as well as the peers with similar features to the cities with worse performance. The CI used in this assessment also has the advantage to address the issue raised by the English Department for Communities and Local Government (2006) concerning the assignment of 134 6.1 Introduction weights to the key performance indicators, as it uses an optimization process to identify the appropriate weights to be assigned to each indicator of livability. This study involves three stages. The first consists in defining the appropriate set of indicators to assess cities’ livability. The indicators are defined based on a literature review of the studies that approached the assessment of cities’ livability. The model proposed extends the concept of urban livability to include a component related to environmental sustainability, which is based on indicators capable of measuring the pressures on the environment. The conceptual model proposed includes twenty four indicators, grouped in eight dimensions of livability: Housing quality, Accessibility and Transportation, Human health, Economic development, Education and Culture and Leisure, representing the human wellbeing component, and Solid waste and Air pollutants representing the environmental impact component. The second stage applies the Directional CI model developed in section 3.3.2 to evaluate cities’ livability. This model is based on a Data Envelopment Analysis model specified with a directional distance function. In addition to providing an overall measure of cities’ performance covering the two components of livability, by using different values for the directional vector it is possible to identify the cities’ potential for improvement considering different perspectives concerning the components of livability. The CI model also included the novel specification of weight restrictions proposed in section 3.4.2, corresponding to assurance regions type I. The weight restrictions imposed ensure that all indicators and dimensions contribute at least with a minimum weight to the cities’ livability assessment. The major advantage of using the Directional CI in the context of the assessment of cities’ livability is that it can guide improvements with different livability objectives, depending on the directional vector specified. This 135