Introduction to Econometric Production Analysis with R
Abstract
This is a collection of my lecture notes for various courses in the field of econometric production analysis. These lecture notes are still incomplete and may contain many typos, errors, and inconsistencies. Please report any problems to [email protected].
Full text
Introduction to Econometric Production Analysis with R (Seventh Draft Version) Arne Henningsen with contributions of Tomasz G. Czekaj Department of Food and Resource Economics University of Copenhagen November 11, 2025 Introduction to Econometric Production Analysis with R ©2025 by Arne Henningsen is licensed under CC BY-NC-SA 4.0 cbna
Foreword This is a collection of my lecture notes for various courses in the field of econometric production analysis. These lecture notes are still incomplete and may contain many typos, errors, and inconsistencies. Please report any problems to [email protected]. I am grateful to Tomasz G. Czekaj who drafted and revised some parts of these lecture notes and to my former students who helped me to improve my teaching and these notes through their questions, suggestions, and comments. Finally, I thank the Rcommunity for providing so many excellent tools for econometric production analysis. November 11, 2025 Arne Henningsen How to cite these lecture notes: Henningsen, Arne (2025): Introduction to Econometric Production Analysis with R. Collection of Lecture Notes. 7th Draft Version. Department of Food and Resource Economics, University of Copenhagen. Available at Zenodo (http://doi.org/10.5281/zenodo.11093657) and Leanpub (http://leanpub.com/ProdEconR/). 2
History First Draft Version (March 9, 2015) •Initial release at Leanpub.com Second Draft Version (February 2, 2018) •added a chapter on distance functions •corrected the equation for calculating direct elasticities of substitution •a large number of further additions, improvements, and corrections Third Draft Version (January 29, 2019) •corrected several typos and minor errors •some minor additions and improvements Fourth Draft Version (May 27, 2019) •added three sections about homotheticity of cost functions (sections 3.1.4,3.2.6, and 3.5.6) •added a section about cost functions with multiple outputs (section 3.1.6) •added a section about cost functions with technical change (section 9.4) •added a proof that cost minimisation and profit maximisation imply that the output elasticities of the inputs divided by the elasticity of scale are equal to their cost shares (section 2.1.11) •added a proof that revenue maximisation and profit maximisation imply that the distance elasticities of the outputs derived from an output distance function are equal to their revenue shares (section 8.1.1.4) •added a proof that cost minimisation and profit maximisation imply that the negative distance elasticities of the inputs divided by the elasticity of scale derived from an output distance function are equal to their cost shares (section 8.1.1.4) •added an approach for obtaining unobserved (shadow) prices of outputs from an output distance function (section 8.1.1.5) •a few minor additions and improvements Fifth Draft Version (October 28, 2020) •added two subsections that discuss the suitability of the production function and the cost function, respectively, in econometric applications (sections 2.1.13 and 3.1.7) •added a comparison of the observed cost shares with the cost shares that would minimize costs according to an estimated Cobb-Douglas production function in section 2.4.14 •added a section about imposing monotonicity on the Translog input distance function (section 8.5.6) •several minor corrections and improvements 3
Sixth Draft Version (April 30, 2024) •several small corrections, improvements, and extensions Seventh Draft Version (November 11, 2025) •several small corrections, improvements, and extensions 4
Contents 1 Introduction 14 1.1 Objectives of the course and the lecture notes .................... 14 1.2 An extremely short introduction to R......................... 14 1.2.1 Some commands for simple calculations ................... 15 1.2.2 Creating objects and assigning values .................... 16 1.2.3 Vectors ..................................... 17 1.2.4 Simple functions ................................ 18 1.2.5 Comparing values and Boolean values .................... 19 1.2.6 Data sets (“data frames”) ........................... 20 1.2.7 Functions .................................... 22 1.2.8 Simple graphics ................................. 22 1.2.9 Other useful commands ............................ 23 1.2.10 Extension packages ............................... 23 1.2.11 Reading data into R.............................. 24 1.2.12 Linear regression ................................ 24 1.3 Rpackages ....................................... 28 1.4 Data sets ........................................ 31 1.4.1 French apple producers ............................ 31 1.4.1.1 Description of the data set ..................... 31 1.4.1.2 Abbreviating name of data set ................... 32 1.4.1.3 Calculation of input quantities ................... 33 1.4.1.4 Calculation of total costs, variable costs, and cost shares . . . . 33 1.4.1.5 Calculation of profit and gross margin ............... 34 1.4.2 Rice producers on the Philippines ...................... 34 1.4.2.1 Description of the data set ..................... 34 1.4.2.2 Mean-scaling quantities ....................... 35 1.4.2.3 Logarithmic mean-scaled quantities ................ 35 1.4.2.4 Mean-adjusting the time trend ................... 36 1.4.2.5 Total costs and cost shares ..................... 36 1.4.2.6 Specifying panel structure ...................... 36 1.5 Mathematical and statistical methods ........................ 37 1.5.1 Exponentiation ................................. 37 1.5.2 Logarithms ................................... 37 5
Contents 1.5.3 Partial derivatives ............................... 38 1.5.4 Aggregating quantities ............................. 38 1.5.5 Concave and convex functions ......................... 40 1.5.6 Quasiconcave and quasiconvex functions ................... 41 1.5.7 Delta method .................................. 42 2 Primal Approach: Production Function 43 2.1 Theory .......................................... 43 2.1.1 Production function .............................. 43 2.1.2 Average products ................................ 43 2.1.3 Total factor productivity ........................... 43 2.1.4 Marginal products ............................... 44 2.1.5 Output elasticities ............................... 44 2.1.6 Elasticity of scale and most productive scale size .............. 44 2.1.7 Marginal rates of technical substitution ................... 45 2.1.8 Relative marginal rates of technical substitution .............. 46 2.1.9 Elasticities of substitution ........................... 46 2.1.9.1 Direct elasticities of substitution .................. 46 2.1.9.2 Allen elasticities of substitution .................. 46 2.1.9.3 Morishima elasticities of substitution ............... 47 2.1.10 Profit maximization .............................. 48 2.1.11 Cost minimization ............................... 49 2.1.12 Derived input demand functions and output supply functions ....... 51 2.1.12.1 Derived from profit maximization ................. 51 2.1.12.2 Derived from cost minimization .................. 52 2.1.13 Suitability of the production function for econometric applications . . . . 52 2.2 Productivity measures ................................. 53 2.2.1 Average products ................................ 53 2.2.2 Total factor productivity ........................... 56 2.3 Linear production function .............................. 58 2.3.1 Specification .................................. 58 2.3.2 Estimation ................................... 59 2.3.3 Properties .................................... 59 2.3.4 Predicted output quantities .......................... 60 2.3.5 Marginal products ............................... 61 2.3.6 Output elasticities ............................... 62 2.3.7 Elasticity of scale ................................ 64 2.3.8 Marginal rates of technical substitution ................... 67 2.3.9 Relative marginal rates of technical substitution .............. 68 2.3.10 First-order conditions for profit maximization ................ 69 6
Contents 2.3.11 First-order conditions for cost minimization ................. 71 2.3.12 Derived input demand functions and output supply functions ....... 73 2.4 Cobb-Douglas production function .......................... 74 2.4.1 Specification .................................. 74 2.4.2 Estimation ................................... 74 2.4.3 Properties .................................... 75 2.4.4 Predicted output quantities .......................... 75 2.4.5 Output elasticities ............................... 76 2.4.6 Marginal products ............................... 77 2.4.7 Elasticity of scale ................................ 79 2.4.8 Marginal rates of technical substitution ................... 81 2.4.9 Relative marginal rates of technical substitution .............. 82 2.4.10 First and second partial derivatives ...................... 83 2.4.11 Elasticities of substitution ........................... 84 2.4.11.1 Direct elasticities of substitution .................. 84 2.4.11.2 Allen elasticities of substitution .................. 85 2.4.11.3 Morishima elasticities of substitution ............... 89 2.4.12 Quasiconcavity ................................. 90 2.4.13 First-order conditions for profit maximization ................ 91 2.4.14 First-order conditions for cost minimization ................. 93 2.4.15 Derived input demand functions and output supply functions ....... 96 2.4.16 Derived input demand elasticities ....................... 99 2.5 Quadratic production function ............................ 101 2.5.1 Specification .................................. 101 2.5.2 Estimation ................................... 101 2.5.3 Properties .................................... 103 2.5.4 Predicted output quantities .......................... 104 2.5.5 Marginal products ............................... 105 2.5.6 Output elasticities ............................... 107 2.5.7 Elasticity of scale ................................ 107 2.5.8 Marginal rates of technical substitution ................... 108 2.5.9 Relative marginal rates of technical substitution .............. 110 2.5.10 Quasiconcavity ................................. 112 2.5.11 Elasticities of substitution ........................... 114 2.5.11.1 Direct elasticities of substitution .................. 114 2.5.11.2 Allen elasticities of substitution .................. 116 2.5.11.3 Comparison of direct and Allen elasticities of substitution . . . . 118 2.5.12 First-order conditions for profit maximization ................ 119 2.5.13 First-order conditions for cost minimization ................. 119 7
Contents 2.6 Translog production function ............................. 123 2.6.1 Specification .................................. 123 2.6.2 Estimation ................................... 123 2.6.3 Statistical significance of individual inputs .................. 125 2.6.4 Properties .................................... 129 2.6.5 Predicted output quantities .......................... 130 2.6.6 Output elasticities ............................... 131 2.6.7 Marginal products ............................... 132 2.6.8 Elasticity of scale ................................ 133 2.6.9 Marginal rates of technical substitution ................... 134 2.6.10 Relative marginal rates of technical substitution .............. 136 2.6.11 Second partial derivatives ........................... 138 2.6.12 Quasiconcavity ................................. 138 2.6.13 Elasticities of substitution ........................... 140 2.6.13.1 Direct elasticities of substitution .................. 140 2.6.13.2 Allen elasticities of substitution .................. 143 2.6.13.3 Comparison of direct and Allen elasticities of substitution . . . . 146 2.6.14 Mean-scaled quantities ............................. 147 2.6.15 First-order conditions for profit maximization ................ 150 2.6.16 First-order conditions for cost minimization ................. 152 2.7 Evaluation of different functional forms ....................... 155 2.7.1 Goodness of fit ................................. 156 2.7.2 Test for functional form misspecification ................... 157 2.7.3 Theoretical consistency ............................ 158 2.7.4 Plausible estimates ............................... 159 2.7.5 Summary .................................... 160 2.8 Non-parametric production function ......................... 161 3 Dual Approach: Cost Functions 167 3.1 Theory .......................................... 167 3.1.1 Cost function .................................. 167 3.1.2 Properties of the cost function ........................ 167 3.1.3 Cost flexibility and elasticity of size ..................... 167 3.1.4 Homotheticity of cost functions ........................ 168 3.1.5 Short-run cost functions ............................ 168 3.1.6 Cost functions with multiple outputs ..................... 169 3.1.7 Suitability of the cost function for econometric applications ........ 170 3.2 Cobb-Douglas cost function .............................. 171 3.2.1 Specification .................................. 171 3.2.2 Estimation ................................... 171 8
Contents 3.2.3 Properties .................................... 172 3.2.4 Estimation with linear homogeneity in input prices imposed ........ 173 3.2.5 Checking concavity in input prices ...................... 177 3.2.6 Homotheticity ................................. 182 3.2.7 Optimal input quantities ........................... 183 3.2.8 Optimal cost shares .............................. 184 3.2.9 Derived input demand functions ....................... 184 3.2.10 Derived input demand elasticities ....................... 186 3.2.11 Cost flexibility and elasticity of size ..................... 188 3.2.12 Marginal costs, average costs, and total costs ................ 188 3.3 Cobb-Douglas short-run cost function ........................ 192 3.3.1 Specification .................................. 192 3.3.2 Estimation ................................... 192 3.3.3 Properties .................................... 193 3.3.4 Estimation with linear homogeneity in input prices imposed ........ 193 3.4 Cobb-Douglas cost function with multiple outputs ................. 195 3.4.1 Specification .................................. 195 3.4.2 Estimation ................................... 195 3.4.3 Properties .................................... 196 3.4.4 Estimation with linear homogeneity in input prices imposed ........ 196 3.4.5 Cost flexibility and elasticity of size ..................... 197 3.5 Translog cost function ................................. 197 3.5.1 Specification .................................. 197 3.5.2 Estimation ................................... 197 3.5.3 Linear homogeneity in input prices ...................... 199 3.5.4 Estimation with linear homogeneity in input prices imposed ........ 201 3.5.5 Cost flexibility and elasticity of size ..................... 206 3.5.6 Homotheticity ................................. 208 3.5.7 Marginal costs and average costs ....................... 210 3.5.8 Derived input demand functions ....................... 213 3.5.9 Derived input demand elasticities ....................... 216 3.5.10 Theoretical consistency ............................ 221 3.6 Translog cost function with multiple outputs .................... 223 3.6.1 Specification .................................. 223 3.6.2 Properties .................................... 223 4 Dual Approach: Profit Function 224 4.1 Theory .......................................... 224 4.1.1 Profit functions ................................. 224 4.1.2 Short-run profit functions ........................... 224 9
1 Introduction [1] 1.414214 > 2^0.5 # also the same [1] 1.414214 > log(3) # natural logarithm [1] 1.098612 > exp(3) # exponential function [1] 20.08554 The commands can span multiple lines. They are executed as soon as the command can be considered as complete. >2+ + 3 [1] 5 >(2 + + + 3 ) [1] 5 1.2.2 Creating objects and assigning values >a<-2 > a [1] 2 >b<-3 > b [1] 3 >a*b [1] 6 Initially, the arrow symbol (<-, consistent of a “smaller than” sign and a dash) was used to assign values to objects. However, in recent versions of R, also the equality sign (=) can be used for this. 16
1 Introduction >a=4 > a [1] 4 >b=5 > b [1] 5 >a*b [1] 20 In these lecture notes, I stick to the traditional assignment operator, i.e. the arrow symbol (<-). Please note that Ris case-sensitive, i.e. Rdistinguishes between upper-case and lower-case letters. Therefore, the following commands return error messages: > A # NOT the same as "a" > B # NOT the same as "b" > Log(3) # NOT the same as "log(3)" > LOG(3) # NOT the same as "log(3)" 1.2.3 Vectors > v <- 1:4 # create a vector with 4 elements: 1, 2, 3, and 4 > v [1]1234 > 2 + v # adding 2 to each element [1]3456 > 2 * v # multiplying each element by 2 [1]2468 > log( v ) # the natural logarithm of each element [1] 0.0000000 0.6931472 1.0986123 1.3862944 > w <- c( 2, 4, 8, 16 ) # concatenate 4 numbers to a vector > w 17
1 Introduction [1] 2 4 8 16 > v + w # element-wise addition [1] 3 6 11 20 > v * w # element-wise multiplication [1] 2 8 24 64 > v %*% w # scalar product (inner product) [,1] [1,] 98 > w[2] # select the second element [1] 4 > w[c(1,3)] # select the first and the third element [1] 2 8 > w[2:4] # select the second, third, and fourth element [1] 4 8 16 > w[-2] # select all but the second element [1] 2 8 16 > length( w ) [1] 4 1.2.4 Simple functions > sum( w ) [1] 30 > mean( w ) [1] 7.5 > median( w ) 18
1 Introduction [1] 6 > min( w ) [1] 2 > max( w ) [1] 16 > which.min( w ) [1] 1 > which.max( w ) [1] 4 1.2.5 Comparing values and Boolean values >a==2 [1] FALSE >a!=2 [1] TRUE >a>4 [1] FALSE >a>=4 [1] TRUE >w>3 [1] FALSE TRUE TRUE TRUE > w == 2^(1:4) [1] TRUE TRUE TRUE TRUE > all.equal( w, 2^(1:4) ) [1] TRUE > w > 3 & w < 6 # ampersand = and [1] FALSE TRUE FALSE FALSE > w < 3 | w > 6 # vertical line = or [1] TRUE FALSE TRUE TRUE 19
1 Introduction 1.2.6 Data sets (“data frames”) The data set “women” is included in R. > data( "women" ) # load the data set into the workspace > women height weight 1 58 115 2 59 117 3 60 120 4 61 123 5 62 126 6 63 129 7 64 132 8 65 135 9 66 139 10 67 142 11 68 146 12 69 150 13 70 154 14 71 159 15 72 164 > names( women ) # display the variable names [1] "height" "weight" > dim( women ) # dimension of the data set (rows and columns) [1] 15 2 > nrow( women ) # number of rows (observations) [1] 15 > ncol( women ) # number of columns (variables) [1] 2 > women[[ "height" ]] # display the values of variable "height" [1] 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 > women$height # short-cut for the previous command 20
1 Introduction [1] 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 > women$height[ 3 ] # height of the third observation [1] 60 > women[ 3, "height" ] # the same [1] 60 > women[ 3, 1 ] # also the same [1] 60 > women[ 1:3, 1 ] # height of the first three observations [1] 58 59 60 > women[ 1:3, ] # all variables of the first three observations height weight 1 58 115 2 59 117 3 60 120 > women$cmHeight <- 2.54 * women$height # new variable: height in cm > women$kgWeight <- women$weight / 2.205 # new variable: weight in kg > women$bmi <- women$kgWeight / ( women$cmHeight / 100 )^2 # new variable: BMI > women height weight cmHeight kgWeight bmi 1 58 115 147.32 52.15420 24.03067 2 59 117 149.86 53.06122 23.62685 3 60 120 152.40 54.42177 23.43164 4 61 123 154.94 55.78231 23.23643 5 62 126 157.48 57.14286 23.04152 6 63 129 160.02 58.50340 22.84718 7 64 132 162.56 59.86395 22.65364 8 65 135 165.10 61.22449 22.46110 9 66 139 167.64 63.03855 22.43112 10 67 142 170.18 64.39909 22.23631 11 68 146 172.72 66.21315 22.19520 12 69 150 175.26 68.02721 22.14711 13 70 154 177.80 69.84127 22.09269 14 71 159 180.34 72.10884 22.17198 15 72 164 182.88 74.37642 22.23836 21
1 Introduction 1.2.7 Functions In order to execute a function in R, the function name has to be followed by a pair of parenthesis (round brackets). The documentation of a function (if available) can be obtained by, e.g., typing at the Rprompt a question mark followed by the name of the function. > ?log One can read in the documentation of the function log, e.g., that this function has a second optional argument base, which can be used to specify the base of the logarithm. By default, the base is equal to the Euler number (e,exp(1)). A different base can be chosen by adding a second argument, either with or without specifying the name of the argument. > log( 100, base = 10 ) [1] 2 > log( 100, 10 ) [1] 2 1.2.8 Simple graphics Histograms can be created with the command hist. The optional argument breaks can be used to specify the approximate number of cells: > hist( women$bmi ) > hist( women$bmi, breaks = 10 ) women$bmi Frequency 22.0 22.5 23.0 23.5 24.0 24.5 02468 women$bmi Frequency 22.0 22.5 23.0 23.5 24.0 01234 Figure 1.1: Histogram of BMIs The resulting histogram is shown in figure 1.1. Scatter plots can be created with the command plot: > plot( women$height, women$weight ) The resulting scatter plot is shown in figure 1.2. 22
1 Introduction 58 60 62 64 66 68 70 72 120 130 140 150 160 women$height women$weight Figure 1.2: Scatter plot of heights and weights 1.2.9 Other useful commands > class( a ) [1] "numeric" > class( women ) [1] "data.frame" > class( women$height ) [1] "numeric" > ls() # list all objects in the workspace [1] "a" "b" "v" "w" "women" > rm(w) # remove an object > ls() [1] "a" "b" "v" "women" 1.2.10 Extension packages Currently (June 12, 2013, 2 pm GMT), 4611 extension packages for Rare available on CRAN (Comprehensive RArchive Network, http://cran.r-project.org). When an extension package is installed, it can be loaded with the command library. The following command loads the R package foreign that includes function for reading data in various formats. > library( "foreign" ) 23
1 Introduction Please note that you should cite scientific software packages in your publications if you used them for obtaining your results (as any other scientific works). You can use the command citation to find out how an Rpackage should be cited, e.g.: > citation( "frontier" ) To cite package 'frontier'in publications use: Tim Coelli and Arne Henningsen (2020). frontier: Stochastic Frontier Analysis. R package version 1.1-8. https://CRAN.R-Project.org/package=frontier. A BibTeX entry for LaTeX users is @Manual{, title = {frontier: Stochastic Frontier Analysis}, author = {Tim Coelli and Arne Henningsen}, year = {2020}, note = {R package version 1.1-8}, url = {https://CRAN.R-Project.org/package=frontier}, } 1.2.11 Reading data into R Rcan read and import data from many different file formats. This is described in the official Rmanual “R Data Import/Export” (http://cran.r-project.org/doc/manuals/r-release/ R-data.pdf). I usually read my data into Rfrom files in CSV (comma separated values) format. This can be done by the function read.csv. The command read.csv2 can read files in the “European CSV format” (values separated by semicolons, comma as decimal separator). The functions read.spss and read.xport (both in package foreign) can read SPSS data files and SAS “XPORT” files, respectively. While the add-on package readstata13 can be used to read binary data files from all STATA versions (including versions 13 and 14), function read.dta (in package foreign) can only read binary data files from STATA versions 5–12. Functions for reading MS-Excel files are available, e.g., in the packages XLConnect and xlsx. 1.2.12 Linear regression The command for estimating linear models in Ris lm. The first argument of the command lm specifies the model that should be estimated. This must be a formula object that consists of the name of the dependent variable, followed by a tilde (˜) and the name of the explanatory variable. Argument data can be used to specify the data set: 24
1 Introduction > olsWeight <- lm( weight ~ height, data = women ) > olsWeight Call: lm(formula = weight ~ height, data = women) Coefficients: (Intercept) height -87.52 3.45 The summary method can be used to display summary statistics of the regression: > summary( olsWeight ) Call: lm(formula = weight ~ height, data = women) Residuals: Min 1Q Median 3Q Max -1.7333 -1.1333 -0.3833 0.7417 3.1167 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) -87.51667 5.93694 -14.74 1.71e-09 *** height 3.45000 0.09114 37.85 1.09e-14 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 1.525 on 13 degrees of freedom Multiple R-squared: 0.991, Adjusted R-squared: 0.9903 F-statistic: 1433 on 1 and 13 DF, p-value: 1.091e-14 The command abline can be used to add a linear (regression) line to a (scatter) plot: > plot( women$height, women$weight ) > abline( olsWeight ) The resulting plot is shown in figure 1.3. This figure indicates that the relationship between the height and the corresponding average weights of the women is slightly nonlinear. Therefore, we add the squared height as additional explanatory regressor. When specifying more than one explanatory variable, the names of the explanatory variables must be separated by plus signs (+): 25
1 Introduction The full panel data set is available in the journal’s data archive: http://www.econ.queensu. ca/jae/1996-v11.6/ivaldi-ladoux-ossard-simioni/.1 The cross-sectional data set that we will predominantly use in the course is available in the R package micEcon. It has the name appleProdFr86 and can be loaded by the command: > data( "appleProdFr86", package = "micEcon" ) The names of the variables in the data set can be obtained by the command names: > names( appleProdFr86 ) [1] "vCap" "vLab" "vMat" "qApples" "qOtherOut" "qOut" [7] "pCap" "pLab" "pMat" "pOut" "adv" The data set includes following variables:2 vCap costs of capital (including land) vLab costs of labor (including remuneration of unpaid family labor) vMat costs of intermediate materials (e.g. seedlings, fertilizer, pesticides, fuel) qOut quantity index of all outputs (apples and other outputs) pCap price index of capital goods pLab price index of labor pMat price index of materials pOut price index of the aggregate output∗ adv use of advisory service∗ Please note that variables indicated by ∗are not in the original data set but are artificially generated in order to be able to conduct some further analyses with this data set. Variable names starting with vindicate volumes (values), variable names starting with qindicate quantities, and variable names starting with pindicate prices. 1.4.1.2 Abbreviating name of data set In order to avoid too much typing, give the data set a much shorter name (dat) by creating a copy of the data set and removing the original data set: > dat <- appleProdFr86 > rm( appleProdFr86 ) 1In order to focus on the microeconomic analysis rather than on econometric issues in panel data analysis, we only use a single year from this panel data set. 2This information is also available in the documentation of this data set, which can be obtained by the command: help( "appleProdFr86", package = "micEcon" ). 32
1 Introduction 1.4.1.3 Calculation of input quantities Our data set does not contain input quantities but prices and costs (volumes) of the inputs. As we will need to know input quantities for many of our analyses, we calculate input quantity indices based on following identity: vi=xi·wi,(1.1) where wiis the price, xiis the quantity and viis the volume of the ith input. In R, we can calculate the input quantities with the following commands: > dat$qCap <- dat$vCap / dat$pCap > dat$qLab <- dat$vLab / dat$pLab > dat$qMat <- dat$vMat / dat$pMat 1.4.1.4 Calculation of total costs, variable costs, and cost shares Total costs are defined as: c= N X i=1 wixi,(1.2) where Ndenotes the number of inputs. We can calculate the apple producers’ total costs by following command: > dat$cost <- with( dat, vCap + vLab + vMat ) Alternatively, we can calculate the costs by summing up the products of the quantities and the corresponding prices over all inputs: > all.equal( dat$cost, with( dat, pCap * qCap + pLab * qLab + pMat * qMat ) ) [1] TRUE Variable costs are defined as: cv=X i∈N1 wixi,(1.3) where N1is a vector of the indices of the variable inputs. If capital is a quasi-fixed input and labor and materials are variable inputs, the apple producers’ variable costs can be calculated by following command: > dat$vCost <- with( dat, vLab + vMat ) The cost shares of the individual inputs in total costs can be calculated as: > dat$sCap <- dat$vCap / dat$cost > dat$sLab <- dat$vLab / dat$cost > dat$sMat <- dat$vMat / dat$cost 33
1 Introduction 1.4.1.5 Calculation of profit and gross margin Profit is defined as: π=p y − N X i=1 wixi=p y −c, (1.4) where all variables are defined as above. We can calculate the apple producers’ profits by: > dat$profit <- with( dat, pOut * qOut - cost ) Alternatively, we can calculate the profit by subtracting the products of the quantities and the corresponding prices of all inputs from the revenues: > all.equal( dat$cost, with( dat, pCap * qCap + pLab * qLab + pMat * qMat ) ) [1] TRUE The gross margin (“variable profit”) is defined as: πv=p y −X i∈N1 wixi=p y −cv,(1.5) where all variables are defined as above. If capital is a quasi-fixed input and labor and materials are variable inputs, the apple producers’ gross margins can be calculated by following command: > dat$vProfit <- with( dat, pOut * qOut - vLab - vMat ) 1.4.2 Rice producers on the Philippines 1.4.2.1 Description of the data set In the last part of this course, we will use a balanced panel data set of annual data collected from 43 smallholder rice producers in the Tarlac region of the Philippines between 1990 and 1997. This data set has the name riceProdPhil and is available in the Rpackage frontier. Detailed information about these data is available in the documentation of this data set. We can load this data set with following command: > data( "riceProdPhil", package = "frontier" ) The names of the variables in the data set can be obtained by the command names: > names( riceProdPhil ) [1] "YEARDUM" "FMERCODE" "PROD" "AREA" "LABOR" "NPK" [7] "OTHER" "PRICE" "AREAP" "LABORP" "NPKP" "OTHERP" [13] "AGE" "EDYRS" "HHSIZE" "NADULT" "BANRAT" The following variables are of particular importance for our analysis: 34
1 Introduction PROD output (tonnes of freshly threshed rice) AREA area planted (hectares). LABOR labor used (man-days of family and hired labor) NPK fertilizer used (kg of active ingredients) YEARDUM time period (1 = 1990, . . . , 8 = 1997) In our analysis of the production technology of the rice producers we will use variable PROD as output quantity and variables AREA,LABOR, and NPK as input quantities. 1.4.2.2 Mean-scaling quantities In some model specifications, it is an advantage to use mean-scaled quantities. Therefore, we create new variables with mean-scaled input and output quantities: > riceProdPhil$area <- riceProdPhil$AREA / mean( riceProdPhil$AREA ) > riceProdPhil$labor <- riceProdPhil$LABOR / mean( riceProdPhil$LABOR ) > riceProdPhil$npk <- riceProdPhil$NPK / mean( riceProdPhil$NPK ) > riceProdPhil$prod <- riceProdPhil$PROD / mean( riceProdPhil$PROD ) As expected, the sample means of the mean-scaled variables are all one so that their logarithms are all zero (except for negligible very small rounding errors): > colMeans( riceProdPhil[ , c( "prod", "area", "labor", "npk" ) ] ) prod area labor npk 1111 > log( colMeans( riceProdPhil[ , c( "prod", "area", "labor", "npk" ) ] ) ) prod area labor npk 0.000000e+00 -1.110223e-16 0.000000e+00 0.000000e+00 1.4.2.3 Logarithmic mean-scaled quantities As we use logarithmic input and output quantities in the Cobb-Douglas and Translog specifications, we can reduce our typing work by creating variables with logarithmic (mean-scaled) input and output quantities: > riceProdPhil$lArea <- log( riceProdPhil$area ) > riceProdPhil$lLabor <- log( riceProdPhil$labor ) > riceProdPhil$lNpk <- log( riceProdPhil$npk ) > riceProdPhil$lProd <- log( riceProdPhil$prod ) Please note that the (arithmetic) mean values of the logarithmic mean-scaled variables are not equal to zero: 35
1 Introduction > colMeans( riceProdPhil[ , c( "lProd", "lArea", "lLabor", "lNpk" ) ] ) lProd lArea lLabor lNpk -0.3263075 -0.2718549 -0.2772354 -0.4078492 1.4.2.4 Mean-adjusting the time trend In some model specifications, it is an advantage to have a time trend variable that is zero at the sample mean. If we subtract the sample mean from our time trend variable, the sample mean of the adjusted time trend is zero: > riceProdPhil$mYear <- riceProdPhil$YEARDUM - mean( riceProdPhil$YEARDUM ) > mean( riceProdPhil$mYear ) [1] 0 1.4.2.5 Total costs and cost shares The following code calculates total costs and the cost shares (ignoring “other inputs”): > riceProdPhil$cost <- with( riceProdPhil, + AREA * AREAP + LABOR * LABORP + NPK * NPKP) > riceProdPhil$sArea <- with( riceProdPhil, AREA * AREAP / cost ) > riceProdPhil$sLabor <- with( riceProdPhil, LABOR * LABORP / cost ) > riceProdPhil$sNpk <- with( riceProdPhil, NPK * NPKP / cost ) Now we check whether these cost shares sum up to one: > all.equal( with( riceProdPhil, sArea + sLabor + sNpk ), + rep( 1, nrow( riceProdPhil ) ) ) [1] TRUE 1.4.2.6 Specifying panel structure This data set does not include any information about its panel structure. Hence, Rwould ignore the panel structure and treat this data set as cross-sectional data collected from 352 different producers. The command pdata.frame of the plm package (Croissant and Millo,2008) can be used to create data sets that include the information on its panel structure. The following commands creates a new data set of the rice producers from the Philippines that includes information on the panel structure, i.e. variable FMERCODE indicates the individual (farmer), and variable YEARDUM indicated the time period (year):3 > library( "plm" ) > pdat <- pdata.frame( riceProdPhil, c( "FMERCODE", "YEARDUM" ) ) 3Please note that the specification of variable YEARDUM as the time dimension in the panel data set pdat converts this variable to a categorical variable. If a numeric time variable is needed, it can be created, e.g., by the command pdat$year <- as.numeric( pdat$YEARDUM ). 36
1 Introduction 1.5 Mathematical and statistical methods 1.5.1 Exponentiation Exponentiation is a calculation with two numbers, e.g., ba, where one number (in this example b) is the base and the other number (in this example a) is the exponent. If the base is equal to the Euler number e≈2.71828, the exponentiation eais called exponential function. In R, exponentiation can be done with the symbol “ˆ”, while values of the exponential function can be obtained by function exp(). The most important rules for calculating with exponentiation (including exponential functions) are: •b0= 1, e.g., e0= 1 •b1=b, e.g., e1=e •b−x= 1/bx, e.g., e−x= 1/ex •bx+y=bx·by, e.g., ex+y=ex·ey •bx−y=bx/by, e.g., ex−y=ex/ey •bx·y= (bx)y, e.g., ex·y= (ex)y •b1/x =x √b •(x·y)a=xa·ya •(x/y)a=xa/ya 1.5.2 Logarithms logb(x)indicates the logarithm of xto the base b. If y= logb(x), then by=x. In economics, we often use the natural logarithm ln(x)≡loge(x), i.e., the logarithm of xto the base of the Euler number e≈2.71828. Some economists write log(x)even if they mean ln(x). In R, function log() calculates logarithms with the base being the Euler number by default, i.e., log(x) returns the natural logarithm of x, whereas, e.g., log(x, base = 10) returns the logarithm of xto the base 10. The most important rules for calculating with logarithms (including natural logarithms) are: •logb(1) = 0, e.g., ln(1) = 0 •logb(b)=1, e.g., ln(e)=1 •logb(bn) = n, e.g., ln(en) = n •limx→0+ logb(x) = −∞, e.g., limx→0+ ln(x) = −∞ •logb(x)and also ln(x)are undefined for x≤0 •logb(1/x) = −logb(x), e.g., ln(1/x) = −ln(x) 37
1 Introduction •logb(x·y) = logb(x) + logb(y), e.g., ln(x·y) = ln(x) + ln(y) •logb(x/y) = logb(x)−logb(y), e.g., ln(x/y) = ln(x)−ln(y) •logb(xy) = ylogb(x), e.g., ln(xy) = yln(x) •logb(x) = ln(x)/ln(b) •logb(x+y)cannot be reasonably transformed 1.5.3 Partial derivatives The most important partial derivatives in microeconomic and econometric production analysis are: •∂a ∂x = 0 •∂xa ∂x =a·xa−1 •∂ex ∂x =ex •∂ln(x) ∂x =1 x •∂(a·f(x)) ∂x =a∂f(x) ∂x •∂(f(x) + g(x)) ∂x =∂f(x) ∂x +∂g(x) ∂x •∂(f(x)−g(x)) ∂x =∂f(x) ∂x −∂g(x) ∂x •∂(f(x)·g(x)) ∂x =∂f(x) ∂x ·g(x)+ ∂g(x) ∂x ·f(x) •∂(f(x)/g(x)) ∂x =∂f(x) ∂x ·g(x)−∂g(x) ∂x ·f(x)(g(x))2 •∂f (g(x)) ∂x =∂f (g(x)) ∂g(x)·∂g(x) ∂x 1.5.4 Aggregating quantities Sometimes, it is desirable to aggregate quantities of different goods to an aggregate quantity. This can be done by a quantity index, e.g. the Laspeyres or Paasche quantity index XL j=Pixij ·pi0 Pixi0·pi0 XP j=Pixij ·pij Pixi0·pij ,(1.6) where subscript iindicates the good, subscript jindicates the observation, xi0is the “base” quantity, and pi0is the “base” price of the ith good, e.g. the sample means. 38
1 Introduction The Paasche and Laspeyres quantity indices of all three inputs in the data set of French apple producers can be calculated by: > dat$XP <- with( dat, + ( vCap + vLab + vMat ) / + ( mean( qCap ) * pCap + mean( qLab ) * pLab + mean( qMat ) * pMat ) ) > dat$XL <- with( dat, + ( qCap * mean( pCap ) + qLab * mean( pLab ) + qMat * mean( pMat ) ) / + ( mean( qCap ) * mean( pCap ) + mean( qLab ) * mean( pLab ) + + mean( qMat ) * mean( pMat ) ) ) In many cases, the choice of the formula for calculating quantity indices does not have a major influence on the result. We demonstrate this with two scatter plots, where we set argument log of the second plot command to the character string "xy" so that both axes are measured in logarithmic terms and the dots (firms) are more equally spread: > plot( dat$XP, dat$XL ) > plot( dat$XP, dat$XL, log = "xy" ) 12345 1 2 3 4 5 XP XL 0.5 1.0 2.0 5.0 0.5 1.0 2.0 5.0 XP XL Figure 1.6: Comparison of Paasche and Laspeyres quantity indices The resulting scatter plots are shown in figure 1.6. As a compromise, one can use the Fisher quantity index, which is the geometric mean of the Paasche quantity index and the Laspeyres quantity index: > dat$X <- sqrt( dat$XP * dat$XL ) We can can also use function quantityIndex from the micEconIndex package to calculate the quantity index: > library( "micEconIndex" ) > dat$XP2 <- quantityIndex( c( "pCap", "pLab", "pMat" ), 39
1 Introduction + c( "qCap", "qLab", "qMat" ), data = dat, method = "Paasche" ) > all.equal( dat$XP, dat$XP2, check.attributes = FALSE ) [1] TRUE > dat$XL2 <- quantityIndex( c( "pCap", "pLab", "pMat" ), + c( "qCap", "qLab", "qMat" ), data = dat, method = "Laspeyres" ) > all.equal( dat$XL, dat$XL2, check.attributes = FALSE ) [1] TRUE > dat$X2 <- quantityIndex( c( "pCap", "pLab", "pMat" ), + c( "qCap", "qLab", "qMat" ), data = dat, method = "Fisher" ) > all.equal( dat$X, dat$X2, check.attributes = FALSE ) [1] TRUE 1.5.5 Concave and convex functions A function f(x) : RN→Ris concave convex over a convex set Dif and only if θf(xu) + (1 −θ)f(xv)≤ ≥f(θxu+ (1 −θ)xv)(1.7) for all combinations of xu, xv∈Dand for all 0≤θ≤1, while it is strictly concave convex over a convex set Dif and only if θf(xu) + (1 −θ)f(xv)< >f(θxu+ (1 −θ)xv)(1.8) for all combinations of xu, xv∈Dand for all 0< θ < 1(Chiang,1984, p. 342). Acontinuous and twice continuously differentiable function f(x) : RN→Ris concave convex over a convex set Dif and only if its Hessian matrix is negative positive semidefinite at all x∈D, while it is strictly concave convex over a convex set Dif (but not only if) its Hessian matrix is negative positive definite at all x∈D(Chiang,1984, p. 347). A symmetric quadratic N×Nmatrix His negative semidefinite if and only if all its ith-order principal minors (not only its leading principal minors) are non-positive for ibeing odd and nonnegative for ibeing even for all i∈ {1, . . . , N}, while it is positive semidefinite if all its principal minors (not only its leading principal minors) are non-negative (Gantmacher,1959, p. 307–308). An ith-order principal minor of a symmetric N×Nmatrix His the determinant of an k×k submatrix of H, whereas N−krows and the corresponding N−kcolumns have been deleted. An N×Nmatrix has N kith-order principal minors and in total PN k=1 N kprincipal minors.4 4The binomial coefficient N k(“Nchoose k”) can be calculated in Rwith the command choose( N, k ). 40
1 Introduction A quadratic N×Nmatrix His negative definite if and only if its first leading principal minor is strictly negative and the following leading principle minors alternate in sign, i.e. |B1|<0, |B2|>0,|B3|<0, . . . , (−1)N|BN|>0, while it is positive definite if and only if all its Nleading principal minors are positive, i.e. |Bi|>0∀i∈ {1,...,}, where |Bi|is the ith leading principal minor of matrix H(Chiang,1984, p. 324). An ith-order leading principal minor of a symmetric N×Nmatrix His the determinant of an k×ksubmatrix of H, whereas last N−krows and the last N−kcolumns have been deleted. An N×Nhas Nleading principal minors. A quadratic N×Nmatrix His negative positive semidefinite if and only if all Neigenvalues of Hare non-positive non-negative and at least one of these eigenvalues is zero (Chiang,1984, p. 330). A quadratic N×Nmatrix His negative positive definite if and only if all Neigenvalues of Hare strictly negative positive (Chiang,1984, p. 330). 1.5.6 Quasiconcave and quasiconvex functions A function f(x) : RN→Ris quasiconcave quasiconvex over a convex set Dif and only if f(θxu+ (1 −θ)xv)≥ ≤min(f(xu), f(xv)) (1.9) for all combinations of xu, xv∈Dand for all 0≤θ≤1, while it is strictly quasiconcave quasiconvex over a convex set Dif and only if f(θxu+ (1 −θ)xv)> <min(f(xu), f(xv)) (1.10) for all combinations of xu, xv∈Dand for all 0< θ < 1(Chiang,1984, p. 389; Chambers,1988, p. 311). An alternative definition of quasiconcavity and quasiconvexity is that a function f(x) : RN→R is quasiconcave quasiconvex over a convex set Dif and only if Su(k) = {x|f(x)≥k} Sl(k) = {x|f(x)≤k}is a convex set (1.11) for any constant k(Chiang,1984, p. 391; Chambers,1988, p. 311). The sets Su(k)and Sl(k)are called upper contour set and lower contour set, respectively. There is no corresponding definition of strict quasiconcavity or strict quasiconvexity (Chiang,1984, p. 391). If f(x) : RN→Ris a continuous and twice continuously differentiable function, a necessary sufficient condition for quasiconcavity is that |B1|≤ <0,|B2|≥ >0,|B3|≤ <0, . . . , (−1)N|BN|≥ >0, while a necessary sufficient condition for quasiconvexity is that |B1|≤ <0,|B2|≤ <0,|B3|≤ <0, . . . , |BN|≤ <0, where |Bi|is the ith 41
2 Primal Approach: Production Function (σM ij =σM ji ). From the above definition of the Morishima elasticities of substitution (2.25), we can derive the relationship between the Morishima elasticities of substitution and the Allen elasticities of substitution: σM ij =fjxj PkfkxkPkfkxk xixj Fij F−fjxj PkfkxkPkfkxk x2 j Fjj F(2.26) =fjxj Pkfkxk σij −fjxj Pkfkxk σjj (2.27) =fjxj Pkfkxk (σij −σjj),(2.28) where σjj can be calculated as the Allen elasticities of substitution with equation (2.21), but does not have an economic meaning. In case of two inputs (x= (x1, x2)), the Morishima elasticities of substitution are symmetric and equal to the direct and Allen elasticities of substitution: σM 12 =f2 x1 F12 F−f2 x2 F22 F(2.29) =F12 Ff2 x1−f2 x2 F22 F12 (2.30) =F12 F f2 x1 +f2 x2 f2 1 f1f2!(2.31) =F12 Ff2 x1 +f1 x2(2.32) =F12 Ff2x2 x1x2 +f1x1 x1x2(2.33) =f1x1+f2x2 x1x2 F12 F(2.34) =σ12 =σD 12.(2.35) 2.1.10 Profit maximization We assume that the firms maximize their profit. The firm’s profit is given by: π=p y −X i wixi,(2.36) where pis the price of the output and wiis the price of the ith input. If the firm faces output price pand input prices wi, we can calculate the maximum profit that can be obtained by the firm by solving following optimization problem: max y,x p y −X i wixi,s.t. y=f(x)(2.37) 48
2 Primal Approach: Production Function This restricted maximization can be transformed into an unrestricted optimization by replacing yby the production function: max xp f(x)−X i wixi(2.38) Hence, the first-order conditions are: ∂π ∂xi =p∂f(x) ∂xi−wi=p MPi−wi= 0 (2.39) so that we get: wi=p MPi=MV Pi(2.40) where MV Pi=p(∂y/∂xi)is the marginal value product of the ith input. 2.1.11 Cost minimization Now, we assume that the firms take total output as given (e.g. because production is restricted by a quota) and try to produce this output quantity with minimal costs. The total cost is given by: c=X i wixi,(2.41) where wiis the price of the ith input. If the firm faces input prices wiand wants to produce yunits of output, the minimum costs can be obtained by: min xX i wixi,s.t. y=f(x)(2.42) This restricted minimization can be solved by using the Lagrangian approach: L=X i wixi+λ(y−f(x)) (2.43) So that the first-order conditions are: ∂L ∂xi =wi−λ∂f(x) ∂xi =wi−λ MPi= 0 (2.44) ∂L ∂λ =y−f(x)=0 (2.45) From the first-order conditions (2.44), we get: wi=λMPi(2.46) and wi wj =λMPi λMPj =MPi MPj =−MRTSji (2.47) 49
2 Primal Approach: Production Function As profit maximization implies producing the optimal output quantity with minimum costs, the first-order conditions for the optimal input combinations (2.47) can be obtained not only from cost minimization but also from the first-order conditions for profit maximization (2.40): wi wj =MV Pi MV Pj =p MPi p MPj =MPi MPj =−MRTSji (2.48) Equations (2.47) and (2.48) can also be expressed in terms of the RMRTS instead of the MRTS: −RMRTSji =−∂xj ∂xi xi xj =−MRTSji xi xj =wi wj xi xj =wixi wjxj = wixi c wjxj c =si sj ,(2.49) where si=wixi/c are the cost shares. Hence, for a profit maximising or cost minimizing producer, the ratio of the cost shares of two inputs must be equal to the negative value of the (reversed) relative marginal rate of technical substitution (RMRTS). As the cost shares sum up to one, i.e., PN i=1 si= 1, we can revise equation (2.49) to: si=−RMRTSji sj=ϵi ϵj sj(2.50) =ϵi ϵj 1−X k∈{1,...,N}\j sk (2.51) =ϵi ϵj 1−X k∈{1,...,N}\j ϵk ϵi si (2.52) =ϵi ϵj 1−si ϵiX k∈{1,...,N}\j ϵk (2.53) =ϵi ϵj−si ϵjX k∈{1,...,N}\j ϵk(2.54) si+si ϵjX k∈{1,...,N}\j ϵk=ϵi ϵj (2.55) si 1 + 1 ϵjX k∈{1,...,N}\j ϵk =ϵi ϵj (2.56) si=ϵi ϵj1 + 1 ϵjPk∈{1,...,N}\jϵk(2.57) =ϵi ϵj+Pk∈{1,...,N}\jϵk (2.58) =ϵi PN k=1 ϵk (2.59) =ϵi ϵ(2.60) Hence, for a profit maximising or cost minimizing producer, the the cost share of each input must 50
2 Primal Approach: Production Function be equal to the output elasticity of this input divided by the elasticity of scale. 2.1.12 Derived input demand functions and output supply functions In this section, we will analyze how profit maximizing or cost minimizing firms react on changing prices and on changing output quantities. 2.1.12.1 Derived from profit maximization If we replace the marginal products in the first-order conditions for profit maximization (2.40) by the equations for calculating these marginal products and then solve this system of equations for the input quantities, we get the input demand functions: xi=xi(p, w),(2.61) where w= [wi]is the vector of all input prices. The input demand functions indicate the optimal input quantities (xi) given the output price (p) and all input prices (w). We can obtain the output supply function from the production function by replacing all input quantities by the corresponding input demand functions: y=f(x(p, w)) = y(p, w),(2.62) where x(p, w)=[xi(p, w)] is the set of all input demand functions. The output supply function indicates the optimal output quantity (y) given the output price (p) and all input prices (w). Hence, the input demand and output supply functions can be used to analyze the effects of prices on the (optimal) input use and output supply. In economics, the effects of price changes are usually measured in terms of price elasticities. These price elasticities can measure the effects of the input prices on the input quantities: ϵij(p, w) = ∂xi(p, w) ∂wj wj xi(p, w),(2.63) the effects of the input prices on the output quantity (expected to be non-positive): ϵyj(p, w) = ∂y(p, w) ∂wj wj y(p, w),(2.64) the effects of the output price on the input quantities (expected to be non-negative): ϵip(p, w) = ∂xi(p, w) ∂p p xi(p, w),(2.65) 51
2 Primal Approach: Production Function and the effect of the output price on the output quantity (expected to be non-negative): ϵyp(p, w) = ∂y(p, w) ∂p p y(p, w).(2.66) The effect of an input price on the optimal quantity of the same input is expected to be nonpositive (ϵii(p, w)≤0). If the cross-price elasticities between two inputs iand jare positive (ϵij(p, w)≥0,ϵji(p, w)≥0), they are considered as gross substitutes. If the cross-price elasticities between two inputs iand jare negative (ϵij(p, w)≤0,ϵji(p, w)≤0), they are considered as gross complements. 2.1.12.2 Derived from cost minimization If we replace the marginal products in the first-order conditions for cost minimization (2.46) by the equations for calculating these marginal products and the solve this system of equations for the input quantities, we get the conditional input demand functions: xi=xi(w, y)(2.67) These input demand functions are called “conditional,” because they indicate the optimal input quantities (xi) given all input prices (w) and conditional on the fixed output quantity (y). The conditional input demand functions can be used to analyze the effects of input prices on the (optimal) input use if the output quantity is given. The effects of price changes on the optimal input quantities can be measured by conditional price elasticities: ϵij(w, y) = ∂xi(w, y) ∂wj wj xi(w, y)(2.68) The effect of the output quantity on the optimal input quantities can also be measured in terms of elasticities (expected to be positive): ϵiy(w, y) = ∂xi(w, y) ∂y y xi(w, y).(2.69) The conditional effect of an input price on the optimal quantity of the same input is expected to be non-positive (ϵii(w, y)≤0). If the conditional cross-price elasticities between two inputs i and jare positive (ϵij(w, y)≥0,ϵji(w, y)≥0), they are considered as net substitutes. If the conditional cross-price elasticities between two inputs iand jare negative (ϵij(w, y)≤0, ϵji(w, y)≤0), they are considered as net complements. 2.1.13 Suitability of the production function for econometric applications Given the microeconomic theory of the production function and the assumptions of standard econometric methods such as ordinary least squares (OLS), it is appropriate to estimate a pro52
2 Primal Approach: Production Function duction function with these econometric methods if several conditions are fulfilled. The most important and relevant conditions are: 1. The inputs and outputs are rather similar across all firms in the data set (within each input or output category). 2. The production conditions are rather similar for all firms in the data set (unless the empirical specification appropriately accounts for these differences). 3. The data set contains variables that indicate—in reasonably precise ways—the quantities of all relevant inputs, the (overall) output quantity, and all relevant ‘environmental’ variables that can be used for controlling for differences in production conditions. 4. Random / representative sample of the firms to be analysed. 5. All firms in the data set produce the maximum obtainable output quantity given the input quantities. 6. No firm in the data set has all input quantities equal to zero but produces a strictly positive output quantity (i.e., weak essentiality must be fulfilled at all observations). 7. The functional form specification of the estimated production function is sufficiently similar to the ‘true’ production function in the ‘real world’. 8. No perfect multicollinearity (i.e., no Leontief production function). 9. Homoscedasticity (often fulfilled to a higher degree when the logarithm of the output quantity instead of the non-transformed output quantity is used as dependent variable). 10. All input quantities (i.e., all explanatory variables) are uncorrelated with the error term, e.g., unobserved heterogeneity between firms (e.g., differences in management quality or production conditions such as soil quality or weather in agricultural production) is unrelated to the input quantities. This assumption is usually not fulfilled if the output quantity is exogenously given so that producers must adjust the input quantities to account for unobserved heterogeneity. 2.2 Productivity measures 2.2.1 Average products We calculate the average products of the three inputs for each firm in the data set by equation 2.2: > dat$apCap <- dat$qOut / dat$qCap > dat$apLab <- dat$qOut / dat$qLab > dat$apMat <- dat$qOut / dat$qMat 53
2 Primal Approach: Production Function We can visualize these average products with histograms that can be created with the command hist. > hist( dat$apCap ) > hist( dat$apLab ) > hist( dat$apMat ) apCap Frequency 0 50 100 150 0 10 20 30 40 50 60 apLab Frequency 0 5 10 15 20 25 0 5 10 15 20 apMat Frequency 0 50 150 250 350 0 10 20 30 40 50 Figure 2.1: Average products The resulting graphs are shown in figure 2.1. These graphs show that average products (partial productivities) vary considerably between firms. Most firms in our data set produce on average between 0 and 40 units of output per unit of capital, between 2 and 16 units of output per unit of labor, and between 0 and 100 units of output per unit of materials. Looking at each average product separately, There are usually many firms with medium to low partial productivity measures and only a few firms with high partial productivity measures. We can find the firms with the highest partial productivities with the commands: > which.max( dat$apCap ) [1] 132 > which.max( dat$apLab ) [1] 7 > which.max( dat$apMat ) [1] 83 Firm number 132 has the highest capital productivity, firm number 7 has the highest labor productivity, and firm number 83 has the highest materials productivity. The relationships between the average products can be visualized by scatter plots: 54
2 Primal Approach: Production Function 0 50 100 150 0 5 10 15 20 25 dat$apCap dat$apLab 0 50 100 150 0 50 100 200 300 dat$apCap dat$apMat 0 5 10 15 20 25 0 50 100 200 300 dat$apLab dat$apMat Figure 2.2: Relationships between average products > plot( dat$apCap, dat$apLab ) > plot( dat$apCap, dat$apMat ) > plot( dat$apLab, dat$apMat ) The resulting graphs are shown in figure 2.2. They show that the average products of the three inputs are positively correlated. As the units of measurements of the input and output quantities in our data set cannot be interpreted in practical terms, the interpretation of the size of the average products is practically not useful. However, they can be used to make comparisons between firms. For instance, the interrelation between average products and firm size can be analyzed. A possible (although not perfect) measure of size of the firms in our data set is the total output. > plot( dat$qOut, dat$apCap, log = "x" ) > plot( dat$qOut, dat$apLab, log = "x" ) > plot( dat$qOut, dat$apMat, log = "x" ) 1e+05 5e+05 5e+06 0 50 100 150 qOut apCap 1e+05 5e+05 5e+06 0 5 10 15 20 25 qOut apLab 1e+05 5e+05 5e+06 0 50 100 200 300 qOut apMat Figure 2.3: Average products for different firm sizes The resulting graphs are shown in figure 2.3. These graphs show that the larger firms (i.e. firms with larger output quantities) produce also a larger output quantity per unit of each input. This 55
2 Primal Approach: Production Function is not really surprising, because the output quantity is in the numerator of equation (2.2) so that the average products are necessarily positively related to the output quantity for a given input quantity. 2.2.2 Total factor productivity After calculating a quantity index of all inputs (see section 1.5.4), we can use equation 2.3 to calculate the total factor productivity, where we arbitrarily choose the Fisher quantity index: > dat$tfp <- dat$qOut / dat$X The variation of the total factor productivities can be visualized as before in a histogram: > hist( dat$tfp ) TFP Frequency 0e+00 2e+06 4e+06 6e+06 0 10 20 30 40 012345 0.0e+00 1.0e+07 2.0e+07 dat$X dat$qOut Figure 2.4: Total factor productivities The resulting histogram is shown in the left panel of figure 2.4. As the total factor productivity is the ratio between the aggregate output quantity and the aggregate input quantity, it can be illustrated by a scatter plot between the aggregate input quantity and the aggregate output quantity, whereas the slope of a line from the origin to a respective point in the scatter plot indicates the total factor productivity. The following commands create a scatter plot between the aggregate input quantity and the aggregate output quantity and add a line through the origin with the slope equal to the maximum total factor productivity as well as a line through the origin with the slope equal to the minimum total factor productivity in the data set: > plot( dat$X, dat$qOut, + xlim = c( 0, max( dat$X ) ), ylim = c( 0, max( dat$qOut ) ) ) > abline( 0, max( dat$tfp ) ) > abline( 0, min( dat$tfp ) ) 56
2 Primal Approach: Production Function The resulting scatter plot is shown in the right panel of figure 2.4. Both parts of figure 2.4 indicate that total factor productivity considerably varies between firms. Where do these large differences in (total factor) productivity come from? We can check the relation between total factor productivity and firm size by a scatter plot. We use two different measures of firm size: total output and aggregate input. The following commands produce scatter plots, where we set argument log of the plot command to the character string "x" so that the horizontal axis is measured in logarithmic terms and the dots (firms) are more equally spread: > plot( dat$qOut, dat$tfp, log = "x" ) > plot( dat$X, dat$tfp, log = "x" ) 1e+05 5e+05 2e+06 1e+07 0e+00 3e+06 6e+06 dat$qOut dat$tfp 0.5 1.0 2.0 5.0 0e+00 3e+06 6e+06 dat$X dat$tfp Figure 2.5: Firm size and total factor productivity The resulting scatter plots are shown in figure 2.5. The scatter plot in the left panel clearly shows that the firms with larger output quantities also have a larger total factor productivity. This is not really surprising, because the output quantity is in the numerator of equation (2.3) so that the total factor productivity is necessarily positively related to the output quantity for given input quantities. The total factor productivity is only slightly positively related to the measure of aggregate input use. We can also analyze whether the firms that use an advisory service have a higher total factor productivity than firms that do not use an advisory service. We can visualize and compare the total factor productivities of the two different groups of firms (with and without advisory service) using boxplot diagrams: > boxplot( tfp ~ adv, data = dat ) > boxplot( log(qOut) ~ adv, data = dat ) > boxplot( log(X) ~ adv, data = dat ) The resulting boxplot graphic is shown on the left panel of figure 2.6. It suggests that the firms that use advisory service are slightly more productive than firms that do not use advisory service (at least when looking at the 25th percentile and the median). 57
2 Primal Approach: Production Function eCapFit eLabFit eMatFit 0.07407719 1.21044421 0.58821500 > hist( dat$eCapFit, 20 ) > hist( dat$eLabFit, 20 ) > hist( dat$eMatFit, 20 ) eCapFit Frequency 0.0 0.5 1.0 1.5 2.0 2.5 3.0 0 20 40 60 80 100 eLabFit Frequency −10 0 10 20 30 40 50 0 20 40 60 80 120 eMatFit Frequency −5 0 5 10 15 20 25 0 20 40 60 80 100 Figure 2.9: Linear production function: output elasticities based on predicted output quantities The resulting graphs are shown in figure 2.9. While the choice of the variable for the output quantity (observed vs. predicted) only has a minor effect on the mean and median values of the output elasticities, the ranges of the output elasticities that are calculated from the predicted output quantities are much larger than the ranges of the output elasticities that are calculated from the observed output quantities. Due to 1 negative predicted output quantity, the output elasticities of this observation are also negative. 2.3.7 Elasticity of scale The elasticity of scale is the sum of all output elasticities ϵ=X i ϵi(2.73) Hence, the elasticities of scale of all firms in the sample can be calculated by: > dat$eScale <- with( dat, eCap + eLab + eMat ) > dat$eScaleFit <- with( dat, eCapFit + eLabFit + eMatFit ) The mean and median values of the elasticities of scale can be calculated by > colMeans( subset( dat, , c( "eScale", "eScaleFit" ) ) ) eScale eScaleFit 3.056945 3.334809 64
2 Primal Approach: Production Function > colMedians( subset( dat, , c( "eScale", "eScaleFit" ) ) ) eScale eScaleFit 1.941536 1.864253 Hence, if a firm increases all input quantities by one percent, the output quantity will usually increase by around 1.9 percent. This means that most firms have increasing returns to scale and hence, the firms could increase productivity by increasing the firm size (i.e. increasing all input quantities). The (variation of the) elasticities of scale can be visualized with histograms: > hist( dat$eScale, 30 ) > hist( dat$eScaleFit, 50 ) > hist( dat$eScaleFit[ dat$eScaleFit > 0 & dat$eScaleFit < 15 ], 30 ) eScale Frequency 0 5 10 15 0 5 10 15 20 25 30 eScaleFit Frequency 0 20 40 60 80 0 20 40 60 0 < eScaleFit < 15 Frequency 2 4 6 8 10 12 14 0 10 20 30 40 Figure 2.10: Linear production function: elasticities of scale The resulting graphs are shown in figure 2.10. As the predicted output quantity of 1 firm is negative, the elasticity of scale of this observation also is negative, if the predicted output quantities are used for the calculation. However, all remaining elasticities of scale that are based on the predicted output quantities are larger than one, which indicates increasing returns to scale. In contrast, 15 (out of 140) elasticities of scale that are calculated with the observed output quantities indicate decreasing returns to scale. However, both approaches indicate that most firms have an elasticity of scale between one and two. Hence, if these firms increase all input quantities by one percent, the output of most firms will increase by between 1 and 2 percent. Some firms even have an elasticity of scale larger than five, which is very implausible and might indicate that the true production technology cannot be reasonably approximated by a linear production function. Information on the optimal firm size can be obtained by analyzing the interrelationship between firm size and the elasticity of scale: > plot( dat$qOut, dat$eScale, log = "x" ) > abline( 1, 0 ) 65
2 Primal Approach: Production Function > plot( dat$X, dat$eScale, log = "x" ) > abline( 1, 0 ) > plot( dat$qOut, dat$eScaleFit, log = "x", ylim = c( 0, 15 ) ) > abline( 1, 0 ) > plot( dat$X, dat$eScaleFit, log = "x", ylim = c( 0, 15 ) ) > abline( 1, 0 ) 0.5 1.0 2.0 5.0 5 10 15 X eScale 1e+05 5e+05 2e+06 5e+06 2e+07 5 10 15 qOut eScale 0.5 1.0 2.0 5.0 0 5 10 15 X eScaleFit 1e+05 5e+05 2e+06 5e+06 2e+07 0 5 10 15 qOut eScaleFit Figure 2.11: Linear production function: elasticities of scale for different firm sizes The resulting graphs are shown in figure 2.11. They indicate that very small firms could enormously gain from increasing their size, while the benefits from increasing firm size decrease with size. Only a few elasticities of scale that are calculated with the observed output quantities indicate decreasing returns to scale so that productivity would decline when these firms increase their size. For all firms that use at least 2.1 times the input quantities of the average firm or produces more than 6,000,000 quantity units (approximately 6,000,000 Euros), the elasticities of scale that are based on the observed input quantities are very close to one. From this observation we could conclude that firms have their optimal size when they use at least 2.1 times the input quantities 66
2 Primal Approach: Production Function of the average firm or produce at least 6,000,000 quantity units (approximately 6,000,000 Euros turn over). In contrast, the elasticities of scale that are based on the predicted output quantities are larger one even for the largest firms in the data set. From this observation, we could conclude that the even the largest firms in the sample would gain from growing in size and thus, the most productive scale size is lager than the size of the largest firms in the sample. The high elasticities of scale explain why we found much higher partial productivities (average products) and total factor productivities for larger firms than for smaller firms. 2.3.8 Marginal rates of technical substitution As the marginal products based on a linear production function are equal to the coefficients, we can calculate the MRTS (2.17) as follows: > mrtsCapLab <- - coef(prodLin)["qLab"] / coef(prodLin)["qCap"] qLab -6.615934 > mrtsLabCap <- - coef(prodLin)["qCap"] / coef(prodLin)["qLab"] qCap -0.1511502 > mrtsCapMat <- - coef(prodLin)["qMat"] / coef(prodLin)["qCap"] qMat -26.09666 > mrtsMatCap <- - coef(prodLin)["qCap"] / coef(prodLin)["qMat"] qCap -0.03831908 > mrtsLabMat <- - coef(prodLin)["qMat"] / coef(prodLin)["qLab"] qMat -3.944516 > mrtsMatLab <- - coef(prodLin)["qLab"] / coef(prodLin)["qMat"] qLab -0.2535165 Hence, if a firm wants to reduce the use of labor by one unit, he/she has to use 6.62 additional units of capital in order to produce the same output as before. Alternatively, the firm can replace the unit of labor by using 0.25 additional units of materials. If the firm increases the use of labor by one unit, he/she can reduce capital by 6.62 units whilst still producing the same output as before. Alternatively, the firm can reduce materials by 0.25 units. 67
2 Primal Approach: Production Function 2.3.9 Relative marginal rates of technical substitution We can calculate the RMRTS (2.18) derived from the linear production function as follows: > dat$rmrtsCapLab <- - dat$eLab / dat$eCap > dat$rmrtsLabCap <- - dat$eCap / dat$eLab > dat$rmrtsCapMat <- - dat$eMat / dat$eCap > dat$rmrtsMatCap <- - dat$eCap / dat$eMat > dat$rmrtsLabMat <- - dat$eMat / dat$eLab > dat$rmrtsMatLab <- - dat$eLab / dat$eMat We can visualize (the variation of) these RMRTSs with histograms: > hist( dat$rmrtsCapLab, 20 ) > hist( dat$rmrtsLabCap ) > hist( dat$rmrtsCapMat ) > hist( dat$rmrtsMatCap ) > hist( dat$rmrtsLabMat ) > hist( dat$rmrtsMatLab ) rmrtsCapLab Frequency −150 −100 −50 0 0 10 20 30 40 50 rmrtsLabCap Frequency −0.20 −0.10 0.00 0 10 20 30 rmrtsCapMat Frequency −60 −40 −20 0 0 10 20 30 40 50 60 rmrtsMatCap Frequency −0.5 −0.4 −0.3 −0.2 −0.1 0.0 0 10 20 30 40 rmrtsLabMat Frequency −1.5 −1.0 −0.5 0.0 0 10 20 30 40 50 rmrtsMatLab Frequency −8 −6 −4 −2 0 0 10 20 30 40 50 Figure 2.12: Linear production function: relative marginal rates of technical substitution (RMRTS) 68
2 Primal Approach: Production Function The resulting graphs are shown in figure 2.12. According to the RMRTS based on the linear production function, most firms need between 20% more capital or around 2% more materials to compensate a 1% reduction of labor. The mean and median values of these RMRTS are: > colMeans( dat[ , c( "rmrtsCapLab", "rmrtsLabCap", "rmrtsCapMat", + "rmrtsMatCap", "rmrtsLabMat", "rmrtsMatLab" ) ] ) rmrtsCapLab rmrtsLabCap rmrtsCapMat rmrtsMatCap rmrtsLabMat rmrtsMatLab -24.24019545 -0.06533753 -10.70239406 -0.14805675 -0.49845972 -2.50782286 > colMedians( dat[ , c( "rmrtsCapLab", "rmrtsLabCap", "rmrtsCapMat", + "rmrtsMatCap", "rmrtsLabMat", "rmrtsMatLab" ) ] ) rmrtsCapLab rmrtsLabCap rmrtsCapMat rmrtsMatCap rmrtsLabMat rmrtsMatLab -17.70447629 -0.05648301 -7.95275036 -0.12574680 -0.43240915 -2.31262549 2.3.10 First-order conditions for profit maximization In this section, we will check to what extent the first-order conditions for profit maximization (2.40) are fulfilled, i.e. to what extent the firms use the optimal input quantities. We do this by comparing the marginal value products of the inputs with the corresponding input prices. We can calculate the marginal value products by multiplying the marginal products by the output price: > dat$mvpCap <- dat$pOut * coef(prodLin)["qCap"] > dat$mvpLab <- dat$pOut * coef(prodLin)["qLab"] > dat$mvpMat <- dat$pOut * coef(prodLin)["qMat"] The command compPlot (package miscTools) can be used to compare the marginal value products with the corresponding input prices: > compPlot( dat$pCap, dat$mvpCap ) > compPlot( dat$pLab, dat$mvpLab ) > compPlot( dat$pMat, dat$mvpMat ) > compPlot( dat$pCap, dat$mvpCap, log = "xy" ) > compPlot( dat$pLab, dat$mvpLab, log = "xy" ) > compPlot( dat$pMat, dat$mvpMat, log = "xy" ) The resulting graphs are shown in figure 2.13. The graphs on the left side indicate that the marginal value products of capital are sometimes lower but more often higher than the capital prices. The four other graphs indicate that the marginal value products of labor and materials are always higher than the labor prices and the materials prices, respectively. This indicates that some firms could increase their profit by using more capital and all firms could increase their 69
2 Primal Approach: Production Function 0 1 2 3 4 5 012345 w Cap MVP Cap 0 5 10 15 20 25 30 35 0 5 10 15 20 25 30 35 w Lab MVP Lab 0 20 40 60 80 120 0 20 40 60 80 100 140 w Mat MVP Mat 0.2 0.5 1.0 2.0 5.0 0.2 0.5 1.0 2.0 5.0 w Cap MVP Cap 0.5 1.0 2.0 5.0 20.0 0.5 1.0 2.0 5.0 10.0 w Lab MVP Lab 5 10 20 50 100 5 10 20 50 100 w Mat MVP Mat Figure 2.13: Marginal value products and corresponding input prices 70
2 Primal Approach: Production Function profit by using more labor and more materials. Given that most firms operate under increasing returns to scale, it is not surprising that most firms would gain from increasing most—or even all—input quantities. Therefore, the question arises why the firms in the sample did not do this. There are many possible reasons for not increasing the input quantities until the predicted optimal input levels, e.g. legal restrictions, environmental regulations, market imperfections, credit (liquidity) constraints, and/or risk aversion. Furthermore, market imperfections might cause that the (observed) average prices are lower than the marginal costs of obtaining these inputs (e.g. Henning and Henningsen,2007), particularly for labor and capital. 2.3.11 First-order conditions for cost minimization As the marginal rates of technical substitution are constant for linear production functions, we compare the input price ratios with the negative inverse marginal rates of technical substitution by creating a histogram for each input price ratio and drawing a vertical line at the corresponding negative marginal rate of technical substitution: > hist( dat$pCap / dat$pLab ) > abline( v = - mrtsLabCap, lwd = 3 ) > hist( dat$pCap / dat$pMat ) > abline( v = - mrtsMatCap, lwd = 3 ) > hist( dat$pLab / dat$pMat ) > abline( v = - mrtsMatLab, lwd = 3 ) > hist( dat$pLab / dat$pCap ) > abline( v = - mrtsCapLab, lwd = 3 ) > hist( dat$pMat / dat$pCap ) > abline( v = - mrtsCapMat, lwd = 3 ) > hist( dat$pMat / dat$pLab ) > abline( v = - mrtsLabMat, lwd = 3 ) The resulting graphs are shown in figure 2.14. The upper left graph shows that the ratio between the capital price and the labor price is larger than the absolute value of the marginal rate of technical substitution between labor and capital (0.151) for the most firms in the sample: wcap wlab >−MRTSlab,cap =MPcap MPlab (2.74) Or taken the other way round, the lower left graph shows that the ratio between the labor price and the capital price is smaller than the absolute value of the marginal rate of technical substitution between capital and labor (6.616) for the most firms in the sample: wlab wcap <−MRTScap,lab =MPlab MPcap (2.75) Hence, the firm can get closer to the minimum of the costs by substituting labor for capital, 71
2 Primal Approach: Production Function w Cap / w Lab Frequency 0 1 2 3 4 5 0 10 20 30 40 w Cap / w Mat Frequency 0.0 0.2 0.4 0.6 0.8 0 10 20 30 40 50 w Lab / w Mat Frequency 0.1 0.2 0.3 0.4 0 10 20 30 40 w Lab / w Cap Frequency 02468 0 20 40 60 80 w Mat / w Cap Frequency 0 10 20 30 40 50 60 0 10 20 30 40 50 60 w Mat / w Lab Frequency 5 10 15 20 0 10 20 30 40 Figure 2.14: First-order conditions for costs minimization 72
2 Primal Approach: Production Function because this will decrease the marginal product of labor and increase the marginal product of capital so that the absolute value of the MRTS between labor and capital increases, the absolute value of the MRTS between capital and labor decreases, and both of the MRTS get closer to the corresponding input price ratios. Similarly, the graphs in the middle column indicate that almost all firms should substitute materials for capital and the graphs on the right indicate that most of the firms should substitute labor for materials. Hence, the firms could reduce production costs particularly by using less capital and more labor. 2.3.12 Derived input demand functions and output supply functions Given a linear production function (2.70), the input quantities chosen by a profit maximizing producer are either zero, indeterminate, or infinity: xi(p, w) = 0if MV Pi< wi indeterminate if MV Pi=wi ∞if MV Pi> wi (2.76) If all input quantities are zero, the output quantity is equal to the intercept, which is zero in case of weak essentiality. Otherwise, the output quantity is indeterminate or infinity: y(p, w) = β0if MV Pi< wi∀i ∞if MV Pi> wi∃i indeterminate otherwise (2.77) A cost minimizing producer will use only a single input, i.e. the input with the lowest cost per unit of produced output (wi/MPi). If the lowest cost per unit of produced output can be obtained by two or more inputs, these input quantities are indeterminate. xi(w, y) = 0if βi wi<βj wj∃j y−β0 βiif βi wi>βj wj∀j=i indeterminate otherwise (2.78) Given that the unconditional and conditional input demand functions and the output supply functions based on the linear production function are non-continuous and often return either zero or infinite values, it does not make much sense to use this functional form to predict the effects of price changes when the true technology implies that firms always use non-zero finite input quantities. 73
2 Primal Approach: Production Function > ESCD <- sum( coef(prodCD)[-1] ) [1] 1.466442 > dESCD <- c( 0, 1, 1, 1 ) [1]0111 > varESCD <- t(dESCD) %*% vcov(prodCD) %*% dESCD [,1] [1,] 0.0118237 > seESCD <- sqrt( varESCD ) [,1] [1,] 0.1087369 Now, we can apply a ttest to test whether the elasticity of scale significantly differs from one. The following commands calculate the tvalue and the critical value for a two-sided ttest based on a 5% significance level: > tESCD <- (ESCD - 1) / seESCD [,1] [1,] 4.289645 > cvESCD <- qt( 0.975, 136 ) [1] 1.977561 Given that the tvalue is larger than the critical value, we can reject the null hypothesis of constant returns to scale and conclude that the technology has significantly increasing returns to scale. The P value for this two-sided ttest is: > pESCD <- 2 * ( 1 - pt( tESCD, 136 ) ) [,1] [1,] 3.372264e-05 Given that the P value is close to zero, we can be very sure that the technology has increasing returns to scale. The 95% confidence interval for the elasticity of scale is: > c( ESCD - cvESCD * seESCD, ESCD + cvESCD * seESCD ) [1] 1.251409 1.681476 80
2 Primal Approach: Production Function 2.4.8 Marginal rates of technical substitution The MRTS based on the Cobb-Douglas production function differ between firms. They can be calculated as follows: > dat$mrtsCapLabCD <- - dat$mpLabCD / dat$mpCapCD > dat$mrtsLabCapCD <- - dat$mpCapCD / dat$mpLabCD > dat$mrtsCapMatCD <- - dat$mpMatCD / dat$mpCapCD > dat$mrtsMatCapCD <- - dat$mpCapCD / dat$mpMatCD > dat$mrtsLabMatCD <- - dat$mpMatCD / dat$mpLabCD > dat$mrtsMatLabCD <- - dat$mpLabCD / dat$mpMatCD Given that the marginal rates of technical substitution are ratios between marginal products, the output quantities in the marginal products cancel out so that it does not matter whether one uses the marginal products that were calculated based on the observed output quantities or the marginal products that were calculated based on the predicted output quantities: > all.equal( dat$mrtsCapLabCD, - dat$mpLabCDFit / dat$mpCapCDFit ) [1] TRUE > all.equal( dat$mrtsLabCapCD, - dat$mpCapCDFit / dat$mpLabCDFit ) [1] TRUE > all.equal( dat$mrtsCapMatCD, - dat$mpMatCDFit / dat$mpCapCDFit ) [1] TRUE > all.equal( dat$mrtsMatCapCD, - dat$mpCapCDFit / dat$mpMatCDFit ) [1] TRUE > all.equal( dat$mrtsLabMatCD, - dat$mpMatCDFit / dat$mpLabCDFit ) [1] TRUE > all.equal( dat$mrtsMatLabCD, - dat$mpLabCDFit / dat$mpMatCDFit ) [1] TRUE We can visualize (the variation of) these MRTSs with histograms: > hist( dat$mrtsCapLabCD ) > hist( dat$mrtsLabCapCD ) > hist( dat$mrtsCapMatCD ) > hist( dat$mrtsMatCapCD ) > hist( dat$mrtsLabMatCD ) > hist( dat$mrtsMatLabCD ) 81
2 Primal Approach: Production Function mrtsCapLabCD Frequency −6 −5 −4 −3 −2 −1 0 0 5 10 15 20 25 30 35 mrtsLabCapCD Frequency −6 −5 −4 −3 −2 −1 0 0 10 20 30 40 50 60 mrtsCapMatCD Frequency −50 −40 −30 −20 −10 0 0 10 20 30 mrtsMatCapCD Frequency −0.6 −0.4 −0.2 0.0 0 10 20 30 40 50 60 mrtsLabMatCD Frequency −35 −25 −15 −5 0 0 10 20 30 40 50 60 mrtsMatLabCD Frequency −0.4 −0.3 −0.2 −0.1 0.0 0 10 20 30 40 50 Figure 2.17: Cobb-Douglas production function: marginal rates of technical substitution (MRTS) The resulting graphs are shown in figure 2.17. According to the MRTS based on the CobbDouglas production function, most firms only need between 0.5 and 2 additional units of capital or between 0.05 and 0.15 additional units of materials to replace one unit of labor. 2.4.9 Relative marginal rates of technical substitution As we do not know the units of measurements of the input quantities, the interpretation of the MRTSs is practically not very useful. To overcome this problem, we calculate the relative marginal rates of technical substitution (RMRTS) by equation (2.18). As the output elasticities based on a Cobb-Douglas production function are equal to the coefficients, we can calculate the RMRTS as follows: > rmrtsCapLabCD <- - coef(prodCD)["log(qLab)"] / coef(prodCD)["log(qCap)"] log(qLab) -4.147897 > rmrtsLabCapCD <- - coef(prodCD)["log(qCap)"] / coef(prodCD)["log(qLab)"] log(qCap) -0.241086 82
2 Primal Approach: Production Function > rmrtsCapMatCD <- - coef(prodCD)["log(qMat)"] / coef(prodCD)["log(qCap)"] log(qMat) -3.847203 > rmrtsMatCapCD <- - coef(prodCD)["log(qCap)"] / coef(prodCD)["log(qMat)"] log(qCap) -0.2599291 > rmrtsLabMatCD <- - coef(prodCD)["log(qMat)"] / coef(prodCD)["log(qLab)"] log(qMat) -0.9275069 > rmrtsMatLabCD <- - coef(prodCD)["log(qLab)"] / coef(prodCD)["log(qMat)"] log(qLab) -1.078159 Hence, if a firm wants to reduce the use of labor by one percent, it has to use 4.15 percent more capital in order to produce the same output as before. Alternatively, the firm can replace one percent of labor by using 1.08 percent more materials. If the firm increases the use of labor by one percent, it can reduce capital by 4.15 percent whilst still producing the same output as before. Alternatively, the firm can reduce materials by 1.08 percent. 2.4.10 First and second partial derivatives For the Cobb-Douglas production function with three inputs (2.79), the first derivatives (marginal products) are: f1=∂y ∂x1 =α1A xα1−1 1xα2 2xα3 3=α1 y x1 (2.84) f2=∂y ∂x2 =α2A xα1 1xα2−1 2xα3 3=α2 y x2 (2.85) f3=∂y ∂x3 =α3A xα1 1xα2 2xα3−1 3=α3 y x3 (2.86) and the second derivatives are: f11 =∂f1 ∂x1 =α1 f1 x1−α1 y x2 1 =f2 1 y−f1 x1 (2.87) f22 =∂f2 ∂x2 =α2 f2 x2−α2 y x2 2 =f2 2 y−f2 x2 (2.88) f33 =∂f3 ∂x3 =α3 f3 x3−α3 y x2 3 =f2 3 y−f3 x3 (2.89) 83
2 Primal Approach: Production Function f12 =∂f1 ∂x2 =α1 f2 x1 =f1f2 y(2.90) f13 =∂f1 ∂x3 =α1 f3 x1 =f1f3 y(2.91) f23 =∂f2 ∂x3 =α2 f3 x2 =f2f3 y.(2.92) Generally, for an N-input Cobb-Douglas function, the first and second derivatives are fi=αi y xi (2.93) fij =fifj y−δij fi xi ,(2.94) where δij denotes Kronecker’s delta with δij = 1if i=j 0if i=j (2.95) In section 2.4.6, we have already computed the first partial derivatives (fi) of the CobbDouglas production function, which are also called marginal products. In section 2.4.6, we also argue that using observed (rather than predicted) output quantities can result in inconsistencies in microeconomic analyses. Therefore, we use marginal products based on predicted output quantities in the following derivations. In order to simplify further computations, we create variables with shorter names for the first derivatives (marginal products): > dat$fCap <- dat$mpCapCDFit > dat$fLab <- dat$mpLabCDFit > dat$fMat <- dat$mpMatCDFit Based on these first derivatives, we can also calculate the second derivatives: > dat$fCapCap <- with( dat, fCap^2 / qOutCD - fCap / qCap ) > dat$fLabLab <- with( dat, fLab^2 / qOutCD - fLab / qLab ) > dat$fMatMat <- with( dat, fMat^2 / qOutCD - fMat / qMat ) > dat$fCapLab <- with( dat, fCap * fLab / qOutCD ) > dat$fCapMat <- with( dat, fCap * fMat / qOutCD ) > dat$fLabMat <- with( dat, fLab * fMat / qOutCD ) 2.4.11 Elasticities of substitution 2.4.11.1 Direct elasticities of substitution We can calculate the direct elasticities of substitution using equation (2.20): 84
2 Primal Approach: Production Function > dat$esdCapLab <- with( dat, + - ( fCap * fLab * ( qCap * fCap + qLab * fLab ) ) / + ( qCap * qLab * + ( fCapCap * fLab^2 - 2 * fCapLab * fCap * fLab + fLabLab * fCap^2 ) ) ) > dat$esdCapMat <- with( dat, + - ( fCap * fMat * ( qCap * fCap + qMat * fMat ) ) / + ( qCap * qMat * + ( fCapCap * fMat^2 - 2 * fCapMat * fCap * fMat + fMatMat * fCap^2 ) ) ) > dat$esdLabMat <- with( dat, + - ( fLab * fMat * ( qLab * fLab + qMat * fMat ) ) / + ( qLab * qMat * + ( fLabLab * fMat^2 - 2 * fLabMat * fLab * fMat + fMatMat * fLab^2 ) ) ) > range( dat$esdCapLab ) [1] 1 1 > range( dat$esdCapMat ) [1] 1 1 > range( dat$esdLabMat ) [1] 1 1 All direct elasticities of substitution are exactly one for all observations and for all combinations of the inputs. This is no surprise and confirms that our calculations have been done correctly, because the Cobb-Douglas functional form implies that the direct elasticities of substitution are always equal to one, irrespective of the input and output quantities and the estimated parameters. 2.4.11.2 Allen elasticities of substitution In order to calculate the Allen elasticities of substitution, we need to construct the bordered Hessian matrix. As the first and second derivatives of the Cobb-Douglas function differ between observations, also the bordered Hessian matrix differs between observations. As a starting point, we construct the bordered Hessian Matrix just for the first observation: > bhm <- matrix( 0, nrow = 4, ncol = 4 ) > bhm[ 1, 2 ] <- bhm[ 2, 1 ] <- dat$fCap[ 1 ] > bhm[ 1, 3 ] <- bhm[ 3, 1 ] <- dat$fLab[ 1 ] > bhm[ 1, 4 ] <- bhm[ 4, 1 ] <- dat$fMat[ 1 ] > bhm[ 2, 2 ] <- dat$fCapCap[ 1 ] > bhm[ 3, 3 ] <- dat$fLabLab[ 1 ] > bhm[ 4, 4 ] <- dat$fMatMat[ 1 ] 85
2 Primal Approach: Production Function > bhm[ 2, 3 ] <- bhm[ 3, 2 ] <- dat$fCapLab[ 1 ] > bhm[ 2, 4 ] <- bhm[ 4, 2 ] <- dat$fCapMat[ 1 ] > bhm[ 3, 4 ] <- bhm[ 4, 3 ] <- dat$fLabMat[ 1 ] > print(bhm) [,1] [,2] [,3] [,4] [1,] 0.000000 6.229014e+00 6.031225e+00 59.0909133861 [2,] 6.229014 -6.202845e-05 1.169835e-05 0.0001146146 [3,] 6.031225 1.169835e-05 -5.423455e-06 0.0001109752 [4,] 59.090913 1.146146e-04 1.109752e-04 -0.0006462733 Based on this bordered Hessian matrix, we can calculate the co-factors Fij: > FCapLab <- - det( bhm[ -2, -3 ] ) [1] -0.06512713 > FCapMat <- det( bhm[ -2, -4 ] ) [1] -0.006165438 > FLabMat <- - det( bhm[ -3, -4 ] ) [1] -0.02641227 Now, we can use equation (2.21) to calculate the Allen elasticities of substitution: > esaCapLab <- with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) / ( qCap * qLab ) ) * + FCapLab / det( bhm ) [1] 1 > esaCapMat <- with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) / ( qCap * qMat ) ) * + FCapMat / det( bhm ) [1] 1 > esaLabMat <- with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) / ( qLab * qMat ) ) * + FLabMat / det( bhm ) [1] 1 86
2 Primal Approach: Production Function All elasticities of substitution are exactly one at the first observation. We can check whether condition (2.24) is fulfilled: > FCapCap <- det( bhm[ -2, -2 ] ) > FLabLab <- det( bhm[ -3, -3 ] ) > FMatMat <- det( bhm[ -4, -4 ] ) > esaCapCap <- with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) / ( qCap * qCap ) ) * + FCapCap / det( bhm ) > esaLabLab <- with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) / ( qLab * qLab ) ) * + FLabLab / det( bhm ) > esaMatMat <- with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) / ( qMat * qMat ) ) * + FMatMat / det( bhm ) > k1 <- dat$qCap[ 1 ] * bhm[ 1, 2 ] / with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) ) > k2 <- dat$qLab[ 1 ] * bhm[ 1, 3 ] / with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) ) > k3 <- dat$qMat[ 1 ] * bhm[ 1, 4 ] / with( dat[1,], + ( qCap * fCap + qLab * fLab + qMat * fMat ) ) > k1 * esaCapCap + k2 * esaCapLab + k3 * esaCapMat [1] -8.326673e-16 > k1 * esaCapLab + k2 * esaLabLab + k3 * esaLabMat [1] 5.551115e-17 > k1 * esaCapMat + k2 * esaLabMat + k3 * esaMatMat [1] 3.330669e-16 As all three values are virtually zero (the small deviations from zero are caused by rounding errors), we can conclude that condition (2.24) is fulfilled for all three inputs. In order to calculate the Allen elasticities of substitution for all firms, we need to construct the bordered Hessian matrices for all observations. We can stack the bordered Hessian matrices on top of each other, which results in a three-dimensional array: > bhmArray <- array( 0, c( 4, 4, nrow( dat ) ) ) > bhmArray[ 1, 2, ] <- bhmArray[ 2, 1, ] <- dat$fCap > bhmArray[ 1, 3, ] <- bhmArray[ 3, 1, ] <- dat$fLab 87
2 Primal Approach: Production Function > bhmArray[ 1, 4, ] <- bhmArray[ 4, 1, ] <- dat$fMat > bhmArray[ 2, 2, ] <- dat$fCapCap > bhmArray[ 3, 3, ] <- dat$fLabLab > bhmArray[ 4, 4, ] <- dat$fMatMat > bhmArray[ 2, 3, ] <- bhmArray[ 3, 2, ] <- dat$fCapLab > bhmArray[ 2, 4, ] <- bhmArray[ 4, 2, ] <- dat$fCapMat > bhmArray[ 3, 4, ] <- bhmArray[ 4, 3, ] <- dat$fLabMat Based on these bordered Hessian matrices, we can calculate the co-factors and the determinants of the bordered Hessian matrices for each observation: > dat$FCapLab <- - apply( bhmArray[ -2, -3, ], 3, det ) > dat$FCapMat <- apply( bhmArray[ -2, -4, ], 3, det ) > dat$FLabMat <- - apply( bhmArray[ -3, -4, ], 3, det ) > dat$bhmDet <- apply( bhmArray, 3, det ) Finally, we can calculate the Allen elasticities of substitution: > dat$esaCapLab <- with( dat, + ( qCap * fCap + qLab * fLab + qMat * fMat ) / + ( qCap * qLab ) * FCapLab / bhmDet ) > dat$esaCapMat <- with( dat, + ( qCap * fCap + qLab * fLab + qMat * fMat ) / + ( qCap * qMat ) * FCapMat / bhmDet ) > dat$esaLabMat <- with( dat, + ( qCap * fCap + qLab * fLab + qMat * fMat ) / + ( qLab * qMat ) * FLabMat / bhmDet ) > range( dat$esaCapLab ) [1] 1 1 > range( dat$esaCapMat ) [1] 1 1 > range( dat$esaLabMat ) [1] 1 1 Also the Allen elasticities of substitution are equal to one for all firms and for all input combinations. This is again no surprise and confirms that our calculations have been done correctly, because the Cobb-Douglas functional form implies that the Allen elasticities of substitution are always equal to one, irrespective of the input and output quantities and the estimated parameters. 88
2 Primal Approach: Production Function 2.4.11.3 Morishima elasticities of substitution In order to calculate the Morishima elasticities of substitution, we need to calculate the co-factors of the diagonal elements of the bordered Hessian matrix for each observation: > dat$FCapCap <- apply( bhmArray[ -2, -2, ], 3, det ) > dat$FLabLab <- apply( bhmArray[ -3, -3, ], 3, det ) > dat$FMatMat <- apply( bhmArray[ -4, -4, ], 3, det ) We use equation (2.25) to calculate the Morishima elasticities of substitution for all observations in our data set: > dat$esmCapLab <- with( dat, ( fLab / qCap ) * FCapLab / bhmDet - + ( fLab / qLab ) * FLabLab / bhmDet ) > range( dat$esmCapLab ) [1] 1 1 > dat$esmLabCap <- with( dat, ( fCap / qLab ) * FCapLab / bhmDet - + ( fCap / qCap ) * FCapCap / bhmDet ) > range( dat$esmLabCap ) [1] 1 1 > dat$esmCapMat <- with( dat, ( fMat / qCap ) * FCapMat / bhmDet - + ( fMat / qMat ) * FMatMat / bhmDet ) > range( dat$esmCapMat ) [1] 1 1 > dat$esmMatCap <- with( dat, ( fCap / qMat ) * FCapMat / bhmDet - + ( fCap / qCap ) * FCapCap / bhmDet ) > range( dat$esmMatCap ) [1] 1 1 > dat$esmLabMat <- with( dat, ( fMat / qLab ) * FLabMat / bhmDet - + ( fMat / qMat ) * FMatMat / bhmDet ) > range( dat$esmLabMat ) [1] 1 1 > dat$esmMatLab <- with( dat, ( fLab / qMat ) * FLabMat / bhmDet - + ( fLab / qLab ) * FLabLab / bhmDet ) > range( dat$esmMatLab ) 89
2 Primal Approach: Production Function 2.4.15 Derived input demand functions and output supply functions Given a Cobb-Douglas production function (2.79), the input quantities chosen by a profit maximizing producer are xi(p, w) = αi wi P A Y j αj wj!αj 1 1−α if α < 1 0∨∞ if α= 1 ∞if α > 1 (2.98) and the output quantity is y(p, w) = A PαY j αj wj!αj 1 1−α if α < 1 0∨∞ if α= 1 ∞if α > 1 (2.99) with α=Pjαj. Hence, if the Cobb-Douglas production function exhibits increasing returns to scale (ϵ=α > 1), the optimal input and output quantities are infinity. As our estimated Cobb-Douglas production function has increasing returns to scale, the optimal input quantities are infinity. Therefore, we cannot evaluate the effect of prices on the optimal input quantities. A cost minimizing producer would choose the following input quantities: xi(w, y) = y AY j=i αiwj αjwi!αj 1 α (2.100) For our three-input Cobb-Douglas production function, we get following conditional input demand functions xcap(w, y) = y A αcap wcap !αlab+αmat wlab αlab αlab wmat αmat αmat 1 αcap+αlab+αmat (2.101) xlab(w, y) = y A wcap αcap !αcap αlab wlab αcap+αmat wmat αmat αmat !1 αcap+αlab+αmat (2.102) xmat(w, y) = y A wcap αcap !αcap wlab αlab αlab αmat wmat αcap+αlab !1 αcap+αlab+αmat (2.103) We can use these formulas to calculate the cost-minimizing input quantities based on the observed input prices (w) and the predicted output quantities (f(x)). Alternatively, we could calculate the cost-minimizing input quantities based on the observed input prices (w) and the observed 96
2 Primal Approach: Production Function output quantities (y). However, in the latter case, the predicted output quantities based on the cost-minimizing input quantities would differ from the predicted output quantities based on the observed input quantities (i.e. y=f(x(w, y)) =f(x)) so that a comparison of the cost-minimizing input quantities (x(w, y)) with the observed input quantities (x) would be less useful. As the coefficients of the Cobb-Douglas function repeatedly occur in the formulas for calculating the cost-minimizing input quantities, it is convenient to define short-cuts for them: > A <- exp( coef( prodCD )[ "(Intercept)" ] ) > aCap <- coef( prodCD )[ "log(qCap)" ] > aLab <- coef( prodCD )[ "log(qLab)" ] > aMat <- coef( prodCD )[ "log(qMat)" ] Now, we can calculate the cost-minimizing input quantities: > dat$qCapCD <- with( dat, + ( ( qOutCD / A ) * ( aCap / pCap )^( aLab + aMat ) + * ( pLab / aLab )^aLab * ( pMat / aMat )^aMat + )^(1/( aCap + aLab + aMat ) ) ) > dat$qLabCD <- with( dat, + ( ( qOutCD / A ) * ( pCap / aCap )^aCap + * ( aLab / pLab )^( aCap + aMat ) * ( pMat / aMat )^aMat + )^(1/( aCap + aLab + aMat ) ) ) > dat$qMatCD <- with( dat, + ( ( qOutCD / A ) * ( pCap / aCap )^aCap + * ( pLab / aLab )^aLab * ( aMat / pMat )^( aCap + aLab ) + )^(1/( aCap + aLab + aMat ) ) ) Before we continue, we will check whether it is indeed possible to produce the predicted output with the calculated cost-minimizing input quantities: > dat$qOutTest <- with( dat, + A * qCapCD^aCap * qLabCD^aLab * qMatCD^aMat ) > all.equal( dat$qOutCD, dat$qOutTest ) [1] TRUE Given that the output quantities predicted from the cost-minimizing input quantities are all equal to the output quantities predicted from the observed input quantities, we can be pretty sure that our calculations are correct. Now, we can use scatter plots to compare the cost-minimizing input quantities with the observed input quantities: > compPlot( dat$qCapCD, dat$qCap ) > compPlot( dat$qLabCD, dat$qLab ) 97
2 Primal Approach: Production Function > compPlot( dat$qMatCD, dat$qMat ) > compPlot( dat$qCapCD, dat$qCap, log = "xy" ) > compPlot( dat$qLabCD, dat$qLab, log = "xy" ) > compPlot( dat$qMatCD, dat$qMat, log = "xy" ) 0e+00 2e+05 4e+05 6e+05 0e+00 2e+05 4e+05 6e+05 qCapCD qCap 200000 600000 1000000 200000 600000 1000000 qLabCD qLab 20000 60000 100000 20000 60000 100000 qMatCD qMat 5e+03 2e+04 1e+05 5e+05 5e+03 2e+04 1e+05 5e+05 qCapCD qCap 5e+04 2e+05 5e+05 5e+04 2e+05 5e+05 qLabCD qLab 5e+03 2e+04 5e+04 5e+03 2e+04 5e+04 qMatCD qMat Figure 2.22: Optimal and observed input quantities The resulting graphs are shown in figure 2.22. As we already found out in section 2.4.14, many firms could reduce their costs by substituting materials for capital. We can also evaluate the potential for cost reductions by comparing the observed costs with the costs when using the cost-minimizing input quantities: > dat$costProdCD <- with( dat, + pCap * qCapCD + pLab * qLabCD + pMat * qMatCD ) > mean( dat$costProdCD / dat$cost ) [1] 0.9308039 Our model predicts that the firms could reduce their costs on average by 7% by using costminimizing input quantities. The variation of the firms’ cost reduction potentials are shown by a histogram: 98
2 Primal Approach: Production Function > hist( dat$costProdCD / dat$cost ) costProdCD / cost Frequency 0.75 0.80 0.85 0.90 0.95 1.00 0 5 15 25 Figure 2.23: Minimum total costs as share of actual total costs The resulting graph is shown in figure 2.23. While many firms have a rather small potential for reducing costs by reallocating input quantities, there are some firms that could save up to 25% of their total costs by using the optimal combination of input quantities. We can also compare the observed input quantities with the cost-minimizing input quantities and the observed costs with the minimum costs for each single observation (e.g. when consulting individual firms in the sample): > round( subset( dat, , c("qCap", "qCapCD", "qLab", "qLabCD", "qMat", "qMatCD", + "cost", "costProdCD") ) )[1:5,] qCap qCapCD qLab qLabCD qMat qMatCD cost costProdCD 1 84050 33720 360066 405349 34087 38038 846329 790968 2 39663 18431 249769 334442 40819 36365 580545 545777 3 37051 14257 140286 135701 24219 32176 306040 281401 4 21222 13300 83427 69713 18893 25890 199634 191709 5 44675 28400 89223 108761 14424 13107 226578 221302 2.4.16 Derived input demand elasticities We can measure the effect of the input prices and the output quantity on the cost-minimizing input quantities by calculating the conditional price elasticities based on the partial derivatives of the conditional input demand functions (2.100) with respect to the input prices and the output quantity. In case of two inputs, we can calculate the demand elasticities of the first input by: x1(w, y) = y Aα1 α2 w2 w1α21 α(2.104) ϵ11(w, y) =∂x1(w, y) ∂w1 w1 x1(w, y)(2.105) 99
2 Primal Approach: Production Function =1 αy Aα1 α2 w2 w1α21 α−1y Aα2α1 α2 w2 w1α2−1−α1 α2 w2 w2 1w1 x1 (2.106) =−1 αy Aα1 α2 w2 w1α21 α−1y Aα1 α2 w2 w1α2−1α1 α2 w2 w1 α2 x1 (2.107) =−1 αy Aα1 α2 w2 w1α21 α−1y Aα1 α2 w2 w1α2α2 x1 (2.108) =−1 αy Aα1 α2 w2 w1α21 αα2 x1 (2.109) =−1 αx1 α2 x1 =−α2 α=α1−α α=α1 α−1(2.110) ϵ12(w, y) =∂x1(w, y) ∂w2 w2 x1(w, y)(2.111) =1 αy Aα1 α2 w2 w1α21 α−1y Aα2α1 α2 w2 w1α2−1α1 α2 1 w1 w2 x1 (2.112) =1 αy Aα1 α2 w2 w1α21 α−1y Aα1 α2 w2 w1α2−1α1 α2 w2 w1 α2 x1 (2.113) =1 αy Aα1 α2 w2 w1α21 α−1y Aα1 α2 w2 w1α2α2 x1 (2.114) =1 αy Aα1 α2 w2 w1α21 αα2 x1 (2.115) =1 αx1 α2 x1 =α2 α(2.116) ϵ1y(w, y) =∂x1(w, y) ∂y y x1(w, y)(2.117) =1 αy Aα1 α2 w2 w1α21 α−11 Aα1 α2 w2 w1α2y x1 (2.118) =1 αy Aα1 α2 w2 w1α21 α−1y Aα1 α2 w2 w1α21 x1 (2.119) =1 αy Aα1 α2 w2 w1α21 α1 x1 (2.120) =1 αx1 1 x1 =1 α(2.121) and analogously the demand elasticities of the second input: x2(w, y) = y Aα2 α1 w1 w2α11 α(2.122) ϵ22(w, y) =∂x2(w, y) ∂w2 w2 x2(w, y)=−α1 α=α2−α α=α2 α−1(2.123) ϵ21(w, y) =∂x2(w, y) ∂w1 w1 x2(w, y)=α1 α(2.124) ϵ2y(w, y) =∂x2(w, y) ∂y y x2(w, y)=1 α.(2.125) 100
2 Primal Approach: Production Function One can similarly derive the input demand elasticities for the general case of Ninputs: ϵij(w, y) = ∂xi(w, y) ∂wj wj xi(w, y)=αj α−δij (2.126) ϵiy(w, y) = ∂xi(w, y) ∂y y xi(w, y)=1 α,(2.127) where δij is (again) Kronecker’s delta (2.95). We have calculated all these elasticities based on the estimated coefficients of the Cobb-Douglas production function; these elasticities are presented in table 2.1. If the price of capital increases by one percent, the cost-minimizing firm will decrease the use of capital by 0.89% and increase the use of labor and materials by 0.11% each. If the price of labor increases by one percent, the cost-minimizing firm will decrease the use of labor by 0.54% and increase the use of capital and materials by 0.46% each. If the price of materials increases by one percent, the cost-minimizing firm will decrease the use of materials by 0.57% and increase the use of capital and labor by 0.43% each. If the cost-minimizing firm increases the output quantity by one percent, (s)he will increase all input quantities by 0.68%. Table 2.1: Conditional demand elasticities derived from Cobb-Douglas production function wcap wlab wmat y xcap -0.89 0.46 0.43 0.68 xlab 0.11 -0.54 0.43 0.68 xmat 0.11 0.46 -0.57 0.68 2.5 Quadratic production function 2.5.1 Specification A quadratic production function is defined as y=β0+X i βixi+1 2X iX j βijxixj,(2.128) where the restriction βij =βji is required to identify all coefficients, because xixjand xjxiare the same regressors. Based on this general form, we can derive the specification of a quadratic production function with three inputs: y=β0+β1x1+β2x2+β3x3+1 2β11x2 1+1 2β22x2 2+1 2β33x2 3+β12x1x2+β13x1x3+β23x2x3(2.129) 2.5.2 Estimation We can estimate this quadratic production function with the command > prodQuad <- lm( qOut ~ qCap + qLab + qMat + + I( 0.5 * qCap^2 ) + I( 0.5 * qLab^2 ) + I( 0.5 * qMat^2 ) 101
2 Primal Approach: Production Function + + I( qCap * qLab ) + I( qCap * qMat ) + I( qLab * qMat ), + data = dat ) > summary( prodQuad ) Call: lm(formula = qOut ~ qCap + qLab + qMat + I(0.5 * qCap^2) + I(0.5 * qLab^2) + I(0.5 * qMat^2) + I(qCap * qLab) + I(qCap * qMat) + I(qLab * qMat), data = dat) Residuals: Min 1Q Median 3Q Max -3928802 -695518 -186123 545509 4474143 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) -2.911e+05 3.615e+05 -0.805 0.422072 qCap 5.270e+00 4.403e+00 1.197 0.233532 qLab 6.077e+00 3.185e+00 1.908 0.058581 . qMat 1.430e+01 2.406e+01 0.595 0.553168 I(0.5 * qCap^2) 5.032e-05 3.699e-05 1.360 0.176039 I(0.5 * qLab^2) -3.084e-05 2.081e-05 -1.482 0.140671 I(0.5 * qMat^2) -1.896e-03 8.951e-04 -2.118 0.036106 * I(qCap * qLab) -3.097e-05 1.498e-05 -2.067 0.040763 * I(qCap * qMat) -4.160e-05 1.474e-04 -0.282 0.778206 I(qLab * qMat) 4.011e-04 1.112e-04 3.608 0.000439 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 1344000 on 130 degrees of freedom Multiple R-squared: 0.8449, Adjusted R-squared: 0.8342 F-statistic: 78.68 on 9 and 130 DF, p-value: < 2.2e-16 Although many of the estimated coefficients are statistically not significantly different from zero, the statistical significance of some quadratic and interaction terms indicates that the linear production function, which neither has quadratic terms not interaction terms, is not suitable to model the true production technology. As the linear production function is “nested” in the quadratic production function, we can apply a “Wald test” or a likelihood ratio test to check whether the linear production function is rejected in favor of the quadratic production function. These tests can be done by the functions waldtest and lrtest (package lmtest): > library( "lmtest" ) 102
2 Primal Approach: Production Function > waldtest( prodLin, prodQuad ) Wald test Model 1: qOut ~ qCap + qLab + qMat Model 2: qOut ~ qCap + qLab + qMat + I(0.5 * qCap^2) + I(0.5 * qLab^2) + I(0.5 * qMat^2) + I(qCap * qLab) + I(qCap * qMat) + I(qLab * qMat) Res.Df Df F Pr(>F) 1 136 2 130 6 8.1133 1.869e-07 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > lrtest( prodLin, prodQuad ) Likelihood ratio test Model 1: qOut ~ qCap + qLab + qMat Model 2: qOut ~ qCap + qLab + qMat + I(0.5 * qCap^2) + I(0.5 * qLab^2) + I(0.5 * qMat^2) + I(qCap * qLab) + I(qCap * qMat) + I(qLab * qMat) #Df LogLik Df Chisq Pr(>Chisq) 1 5 -2191.3 2 11 -2169.1 6 44.529 5.806e-08 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 These tests show that the linear production function is clearly inferior to the quadratic production function and hence, should not be used for analyzing the production technology of the firms in this data set. 2.5.3 Properties We cannot see from the estimated coefficients whether the monotonicity condition is fulfilled. Unless all coefficients are non-negative (but not necessarily the intercept), quadratic production functions cannot be globally monotone, because there will always be a set of input quantities that result in negative marginal products. We will check the monotonicity condition at each observation in section 2.5.5. Our estimated quadratic production function does not fulfill the weak essentiality assumption, because the intercept is different from zero (but its deviation from zero is not statistically signif103
2 Primal Approach: Production Function icant). The production technology described by a quadratic production function with more than one (relevant) input never shows strict essentiality. The input requirement sets derived from quadratic production functions are always closed and non-empty. The quadratic production function always returns finite, real, and single values but the nonnegativity assumption is only fulfilled, if all coefficients (including the intercept), are non-negative. All quadratic production functions are continuous and twice-continuously differentiable. 2.5.4 Predicted output quantities We can obtain the predicted output quantities with the fitted method: > dat$qOutQuad <- fitted( prodQuad ) We can evaluate the “fit” of the model by comparing the observed with the fitted output quantities: > compPlot( dat$qOut, dat$qOutQuad ) > compPlot( dat$qOut, dat$qOutQuad, log = "xy" ) 0.0e+00 1.0e+07 2.0e+07 0.0e+00 1.0e+07 2.0e+07 observed fitted 1e+05 5e+05 5e+06 1e+05 5e+05 2e+06 1e+07 observed fitted Figure 2.24: Quadratic production function: fit of the model The resulting graphs are shown in figure 2.24. While the graph in the left panel uses a linear scale for the axes, the graph in the right panel uses a logarithmic scale for both axes. Hence, the deviations from the 45°-line illustrate the absolute deviations in the left panel and the relative deviations in the right panel. The fit of the model looks okay in the scatter plot on the left-hand side, but if we use a logarithmic scale on both axes (as in the graph on the right-hand side), we can see that the output quantity is over-estimated if the the observed output quantity is small. As negative output quantities would render the corresponding output elasticities useless, we have carefully check the sign of the predicted output quantities: 104
2 Primal Approach: Production Function > table( dat$qOutQuad >= 0 ) TRUE 140 Fortunately, not a single predicted output quantity is negative. 2.5.5 Marginal products In case of a quadratic production function, the marginal products are MPi=βi+X j βijxj(2.130) We can simplify the code for computing the marginal products and some other figures by using short names for the coefficients: > b1 <- coef( prodQuad )[ "qCap" ] > b2 <- coef( prodQuad )[ "qLab" ] > b3 <- coef( prodQuad )[ "qMat" ] > b11 <- coef( prodQuad )[ "I(0.5 * qCap^2)" ] > b22 <- coef( prodQuad )[ "I(0.5 * qLab^2)" ] > b33 <- coef( prodQuad )[ "I(0.5 * qMat^2)" ] > b12 <- b21 <- coef( prodQuad )[ "I(qCap * qLab)" ] > b13 <- b31 <- coef( prodQuad )[ "I(qCap * qMat)" ] > b23 <- b32 <- coef( prodQuad )[ "I(qLab * qMat)" ] Now, we can use the following commands to calculate the marginal products in R”: > dat$mpCapQuad <- with( dat, + b1 + b11 * qCap + b12 * qLab + b13 * qMat ) > dat$mpLabQuad <- with( dat, + b2 + b21 * qCap + b22 * qLab + b23 * qMat ) > dat$mpMatQuad <- with( dat, + b3 + b31 * qCap + b32 * qLab + b33 * qMat ) We can visualize (the variation of) these marginal products with histograms: > hist( dat$mpCapQuad, 15 ) > hist( dat$mpLabQuad, 15 ) > hist( dat$mpMatQuad, 15 ) The resulting graphs are shown in figure 2.25. If the firms increase capital input by one unit, the output of most firms will increase by around 2 units. If the firms increase labor input by one unit, the output of most firms will increase by around 5 units. If the firms increase material 105
2 Primal Approach: Production Function the RMRTS: > colMedians( subset( dat, monoQuad, + c( "rmrtsCapLabQuad", "rmrtsLabCapQuad", "rmrtsCapMatQuad", + "rmrtsMatCapQuad", "rmrtsLabMatQuad", "rmrtsMatLabQuad" ) ) ) rmrtsCapLabQuad rmrtsLabCapQuad rmrtsCapMatQuad rmrtsMatCapQuad rmrtsLabMatQuad -5.5741780 -0.1793986 -4.2567577 -0.2349206 -0.7745132 rmrtsMatLabQuad -1.2911336 Given that the median relative marginal rate of technical substitution between capital and labor is -5.57, a typical firm that reduces the use of labor by one percent, has to use around 5.57 percent more capital in order to produce the same amount of output as before. Alternatively, the typical firm can replace one percent of labor by using 1.29 percent more materials. 2.5.10 Quasiconcavity In order to check whether the estimated quadratic production function is quasiconcave, we need to construct the bordered Hessian matrix. As the first derivatives of the quadratic production function differ between observations, also the bordered Hessian matrix differs between observations. As a starting point, we construct the bordered Hessian matrix just for the first observation: > bhm <- matrix( 0, nrow = 4, ncol = 4 ) > bhm[ 1, 2 ] <- bhm[ 2, 1 ] <- dat$mpCapQuad[ 1 ] > bhm[ 1, 3 ] <- bhm[ 3, 1 ] <- dat$mpLabQuad[ 1 ] > bhm[ 1, 4 ] <- bhm[ 4, 1 ] <- dat$mpMatQuad[ 1 ] > bhm[ 2, 2 ] <- b11 > bhm[ 3, 3 ] <- b22 > bhm[ 4, 4 ] <- b33 > bhm[ 2, 3 ] <- bhm[ 3, 2 ] <- b12 > bhm[ 2, 4 ] <- bhm[ 4, 2 ] <- b13 > bhm[ 3, 4 ] <- bhm[ 4, 3 ] <- b23 > print(bhm) [,1] [,2] [,3] [,4] [1,] 0.000000 -3.068305e+00 6.042883e+00 9.063140e+01 [2,] -3.068305 5.032000e-05 -3.096510e-05 -4.160252e-05 [3,] 6.042883 -3.096510e-05 -3.084032e-05 4.011378e-04 [4,] 90.631400 -4.160252e-05 4.011378e-04 -1.895501e-03 Based on this bordered Hessian matrix, we can check whether the estimated quadratic production function is quasiconcave at the first observation: 112
2 Primal Approach: Production Function > det( bhm[ 1:2, 1:2 ] ) [1] -9.414493 > det( bhm[ 1:3, 1:3 ] ) [1] -0.0003988879 > det( bhm ) [1] 3.541548e-05 Quasiconcavity requires that the first determinant calculated above (i.e. the determinant of the upper-left 2×2submatrix) is negative and the following determinants alternate in sign. As this condition is not fulfilled for the second and third determinant, we can conclude that the estimated quadratic production function is not quasiconcave at the first observation. Now, we check whether our estimated quadratic production function is quasiconcave at each observation by stacking the bordered Hessian matrices of all observation on top of each other, which results in a three-dimensional array: > bhmQuad <- array( 0, c( 4, 4, nrow( dat ) ) ) > bhmQuad[ 1, 2, ] <- bhmQuad[ 2, 1, ] <- dat$mpCapQuad > bhmQuad[ 1, 3, ] <- bhmQuad[ 3, 1, ] <- dat$mpLabQuad > bhmQuad[ 1, 4, ] <- bhmQuad[ 4, 1, ] <- dat$mpMatQuad > bhmQuad[ 2, 2, ] <- b11 > bhmQuad[ 3, 3, ] <- b22 > bhmQuad[ 4, 4, ] <- b33 > bhmQuad[ 2, 3, ] <- bhmQuad[ 3, 2, ] <- b12 > bhmQuad[ 2, 4, ] <- bhmQuad[ 4, 2, ] <- b13 > bhmQuad[ 3, 4, ] <- bhmQuad[ 4, 3, ] <- b23 > dat$quasiConcQuad <- apply( bhmQuad[ 1:2, 1:2, ], 3, det ) <= 0 & + apply( bhmQuad[ 1:3, 1:3, ], 3, det ) >= 0 & + apply( bhmQuad, 3, det ) <= 0 > table( dat$quasiConcQuad ) FALSE 140 Our estimated quadratic production function is quasiconcave at none of the 140 observations, which implies that the isoquants are convex not at a single observation. We can check whether the (partial) isoquants are convex if we hold one input fixed and only vary the other two inputs: 113
2 Primal Approach: Production Function > bhmQuadCapLab <- bhmQuad[ -4, -4, ] > dat$quasiConcCapLabQuad <- + apply( bhmQuadCapLab[ 1:2, 1:2, ], 3, det ) <= 0 & + apply( bhmQuadCapLab, 3, det ) >= 0 > table( dat$quasiConcCapLabQuad ) FALSE TRUE 121 19 > bhmQuadCapMat <- bhmQuad[ -3, -3, ] > dat$quasiConcCapMatQuad <- + apply( bhmQuadCapMat[ 1:2, 1:2, ], 3, det ) <= 0 & + apply( bhmQuadCapMat, 3, det ) >= 0 > table( dat$quasiConcCapMatQuad ) FALSE TRUE 114 26 > bhmQuadLabMat <- bhmQuad[ -2, -2, ] > dat$quasiConcLabMatQuad <- + apply( bhmQuadLabMat[ 1:2, 1:2, ], 3, det ) <= 0 & + apply( bhmQuadLabMat, 3, det ) >= 0 > table( dat$quasiConcLabMatQuad ) TRUE 140 While the isoquants between labor and materials holding capital constant are convex at all 140 observations, the isoquants between capital and labor holding materials constant and between capital and materials holding labor constant are non-convex at most of the observations. The non-convex isoquants between capital and the other two inputs at many observations are partly caused by a positive second derivative of the estimated quadratic production function with respect to capital (coefficient βcap,cap >0), which indicates increasing returns to capital. 2.5.11 Elasticities of substitution 2.5.11.1 Direct elasticities of substitution We can calculate the direct elasticities of substitution using equation (2.20): > dat$esdCapLabQuad <- with( dat, - ( mpCapQuad * mpLabQuad * + ( qCap * mpCapQuad + qLab * mpLabQuad ) ) / + ( qCap * qLab * ( b11 * mpLabQuad^2 - 114
2 Primal Approach: Production Function + 2 * b12 * mpCapQuad * mpLabQuad + b22 * mpCapQuad^2 ) ) ) > dat$esdCapMatQuad <- with( dat, - ( mpCapQuad * mpMatQuad * + ( qCap * mpCapQuad + qMat * mpMatQuad ) ) / + ( qCap * qMat * ( b11 * mpMatQuad^2 - + 2 * b13 * mpCapQuad * mpMatQuad + b33* mpCapQuad^2 ) ) ) > dat$esdLabMatQuad <- with( dat, - ( mpLabQuad * mpMatQuad * + ( qLab * mpLabQuad + qMat * mpMatQuad ) ) / + ( qLab * qMat * ( b22 * mpMatQuad^2 + - 2 * b23 * mpLabQuad * mpMatQuad + b33 * mpLabQuad^2 ) ) ) As the elasticities of substitution measure changes in the marginal rates of technical substitution (MRTS) and the MRTS are meaningless if the monotonicity conditions are not fulfilled, also the elasticities of substitution are meaningless if the monotonicity conditions are not fulfilled. Hence, we visualize (the variation of) the direct elasticities of substitution only for the observations, where the monotonicity condition is fulfilled: > hist( dat$esdCapLabQuad[ dat$monoQuad ], 30 ) > hist( dat$esdCapMatQuad[ dat$monoQuad ], 30 ) > hist( dat$esdLabMatQuad[ dat$monoQuad ], 30 ) esdCapLabQuad Frequency −8 −6 −4 −2 0 2 4 0 10 20 30 40 esdCapMatQuad Frequency 0 50 150 250 0 20 40 60 80 esdLabMatQuad Frequency 0.0 0.2 0.4 0.6 0.8 0 2 4 6 8 10 12 14 Figure 2.31: Quadratic production function: direct elasticities of substitution The resulting graphs are shown in figure 2.31. The estimated direct elasticities of substitution between capital and labor are almost always negative, the estimated direct elasticities of substitution between capital and materials are partly negative and partly positive, and the estimated direct elasticities of substitution between labor and materials are always positive. As the direct elasticities of substitution assume that only the quantities of the two considered inputs are varied, it is expected that all direct elasticities of substitution are non-negative. However, the estimated quadratic production function is not quasiconcave, which implies that the isoquants are not convex. In case of non-convex isoquants, the elasticities of substitution do not indicate substitutability between inputs. 115
2 Primal Approach: Production Function As only the (partial) isoquants between labor and materials holding capital constant are convex at all observations (see section 2.5.10) and the direct elasticities of substitution between labor and materials also assume that capital is held constant, we focus on these elasticities of substitution. For all firms in our sample, the estimated values of the direct elasticities of substitution between labor and materials lie between the value implied by the Leontief production function (σ= 0) and the value implied by the Cobb-Douglas production function (σ= 1). Hence, the substitutability between labor and materials seems to be between very low and moderate. In fact, the direct elasticity of substitution between labor and materials is for most of the firms around 0.4. Hence, if the firms keep the capital quantity and the output quantity unchanged and substitute labor for materials (or vice versa) so that the ratio between the labor quantity and the quantity of materials increases (decreases) by around 0.4 percent, the MRTS between labor and materials increases (decreases) by one percent. If the ratio between the materials price and the labor price increases by one percent, firms that keep the capital input and output quantity unchanged and minimize their costs will substitute labor for materials so that ratio between the labor quantity and the quantity of materials increases by around 0.4 percent. Hence, the relative change of the ratio of the input quantities is considerably smaller than the relative change of the ratio of the input prices, which indicates a low substitutability between labor and materials. 2.5.11.2 Allen elasticities of substitution In the following, we calculate the Allen elasticities of substitution.5In order to check whether our calculations are correct, we can use equation (2.24) to derive the following conditions: X i xiMPiσij = 0 ∀j(2.132) In order to check this condition, we need to calculate not only (normal) elasticities of substitution (σij;i=j) but also economically not meaningful “elasticities of self-substitution” (σii): > dat$FCapLabQuad <- - apply( bhmQuad[ -2, -3, ], 3, det ) > dat$FCapMatQuad <- apply( bhmQuad[ -2, -4, ], 3, det ) > dat$FLabMatQuad <- - apply( bhmQuad[ -3, -4, ], 3, det ) > dat$FCapCapQuad <- apply( bhmQuad[ -2, -2, ], 3, det ) > dat$FLabLabQuad <- apply( bhmQuad[ -3, -3, ], 3, det ) > dat$FMatMatQuad <- apply( bhmQuad[ -4, -4, ], 3, det ) > dat$bhmDetQuad <- apply( bhmQuad, 3, det ) > dat$numeratorQuad <- with( dat, + qCap * mpCapQuad + qLab * mpLabQuad + qMat * mpMatQuad ) > dat$esaCapLabQuad <- with( dat, 5We do not calculate the Morishima elasticities of substitution for the quadratic production function. The calculation of the Morishima elasticities of substitution requires only minimal changes of the code for calculating the Allen elasticities of substitution (compare sections 2.4.11.2 and 2.4.11.3). 116
2 Primal Approach: Production Function + numeratorQuad / ( qCap * qLab ) * FCapLabQuad / bhmDetQuad ) > dat$esaCapMatQuad <- with( dat, + numeratorQuad / ( qCap * qMat ) * FCapMatQuad / bhmDetQuad ) > dat$esaLabMatQuad <- with( dat, + numeratorQuad / ( qLab * qMat ) * FLabMatQuad / bhmDetQuad ) > dat$esaCapCapQuad <- with( dat, + numeratorQuad / ( qCap * qCap ) * FCapCapQuad / bhmDetQuad ) > dat$esaLabLabQuad <- with( dat, + numeratorQuad / ( qLab * qLab ) * FLabLabQuad / bhmDetQuad ) > dat$esaMatMatQuad <- with( dat, + numeratorQuad / ( qMat * qMat ) * FMatMatQuad / bhmDetQuad ) Before we take a look at and interpret the elasticities of substitution, we check whether the conditions (2.132) are fulfilled: > range( with( dat, qCap * mpCapQuad * esaCapCapQuad + + qLab * mpLabQuad * esaCapLabQuad + qMat * mpMatQuad * esaCapMatQuad ) ) [1] -1.788139e-07 2.235174e-08 > range( with( dat, qCap * mpCapQuad * esaCapLabQuad + + qLab * mpLabQuad * esaLabLabQuad + qMat * mpMatQuad * esaLabMatQuad ) ) [1] -1.629815e-09 2.561137e-09 > range( with( dat, qCap * mpCapQuad * esaCapMatQuad + + qLab * mpLabQuad * esaLabMatQuad + qMat * mpMatQuad * esaMatMatQuad ) ) [1] -1.862645e-09 1.303852e-08 The extremely small deviations from zero are most likely caused by rounding errors that are unavoidable on digital computers. This test does not prove that all our calculations are done correctly but if we had made a mistake, we would have discovered it with a very high probability. Hence, we can be rather sure that our calculations are correct. As for the direct elasticities of substitution, we visualize (the variation of) the Allen elasticities of substitution only for the observations, where the monotonicity condition is fulfilled: > hist( dat$esaCapLabQuad[ dat$monoQuad ], 30 ) > hist( dat$esaCapMatQuad[ dat$monoQuad ], 30 ) > hist( dat$esaLabMatQuad[ dat$monoQuad ], 30 ) The resulting graphs are shown in figure 2.32. As the Allen elasticities of substitution allow the other inputs to adjust, a meaningful interpretation requires that not only the partial isoquants of the two considered inputs are convex but that the production is quasiconcave. As the estimated quadratic production is not quasiconcave at any observation, we do not interpret the obtained Allen elasticities of substitution. 117
2 Primal Approach: Production Function esaCapLabQuad Frequency −8 −6 −4 −2 0 0 10 20 30 40 50 esaCapMatQuad Frequency −2 −1 0 1 2 3 4 0 5 10 15 20 esaLabMatQuad Frequency 0.0 0.5 1.0 1.5 2.0 2.5 0 5 10 15 20 25 Figure 2.32: Quadratic production function: Allen elasticities of substitution 2.5.11.3 Comparison of direct and Allen elasticities of substitution In the following, we use scatter plots to compare the estimated direct elasticities of substitution with the estimated Allen elasticities of substitution: > compPlot( dat$esdCapLabQuad[ dat$monoQuad ], + dat$esaCapLabQuad[ dat$monoQuad ], lim = c( -4, 1 ) ) > compPlot( dat$esdCapMatQuad[ dat$monoQuad ], + dat$esaCapMatQuad[ dat$monoQuad ], lim = c( -10, 5) ) > compPlot( dat$esdLabMatQuad[ dat$monoQuad ], + dat$esaLabMatQuad[ dat$monoQuad ] ) −4 −3 −2 −1 0 1 −4 −3 −2 −1 0 1 esdCapLabQuad esaCapLabQuad −10 −5 0 5 −10 −5 0 5 esdCapMatQuad esaCapMatQuad 0.0 0.5 1.0 1.5 2.0 2.5 0.0 0.5 1.0 1.5 2.0 2.5 esdLabMatQuad esaLabMatQuad Figure 2.33: Quadratic production function: Comparison of direct and Allen elasticities of substitution The resulting graphs are shown in figure 2.33. We focus on the elasticities of substitution between labor and materials, because the elasticities of substitution between capital and labor and the elasticities of substitution between capital and material cannot be meaningful interpreted (see section 2.5.11.1). In contrast to the direct elasticities of substitution, the Allen elasticities of substitution allow the other inputs to adjust. Therefore, it is not surprising that the Allen 118
2 Primal Approach: Production Function elasticities of substitution between labor and materials are about the same or a little larger than the direct elasticities of substitution between labor and materials. However, even the Allen elasticities of substitution between labor and materials indicate a low substitutability between labor and materials (see also section 2.5.11.1). 2.5.12 First-order conditions for profit maximization In this section, we will check to what extent the first-order conditions for profit maximization (2.40) are fulfilled, i.e. to what extent the firms use the optimal input quantities. We do this by comparing the marginal value products of the inputs with the corresponding input prices. We can calculate the marginal value products by multiplying the marginal products by the output price: > dat$mvpCapQuad <- dat$pOut * dat$mpCapQuad > dat$mvpLabQuad <- dat$pOut * dat$mpLabQuad > dat$mvpMatQuad <- dat$pOut * dat$mpMatQuad The command compPlot (package miscTools) can be used to compare the marginal value products with the corresponding input prices. As the logarithm of a non-positive number is not defined, we have to limit the comparisons on the logarithmic scale to observations with positive marginal products: > compPlot( dat$pCap, dat$mvpCapQuad ) > compPlot( dat$pLab, dat$mvpLabQuad ) > compPlot( dat$pMat, dat$mvpMatQuad ) > compPlot( dat$pCap[ dat$monoQuad ], dat$mvpCapQuad[ dat$monoQuad ], log = "xy" ) > compPlot( dat$pLab[ dat$monoQuad ], dat$mvpLabQuad[ dat$monoQuad ], log = "xy" ) > compPlot( dat$pMat[ dat$monoQuad ], dat$mvpMatQuad[ dat$monoQuad ], log = "xy" ) The resulting graphs are shown in figure 2.34. They indicate that the marginal value products of most firms are higher than the corresponding input prices. This indicates that most firms could increase their profit by using more of all inputs. Given that the estimated quadratic function shows that (almost) all firms operate under increasing returns to scale, it is not surprising that most firms would gain from increasing all input quantities. Therefore, the question arises why the firms in the sample did not do this. This questions has already been addressed in section 2.3.10. 2.5.13 First-order conditions for cost minimization As the marginal rates of technical substitution differ between observations for the three other functional forms, we use scatter plots for visualizing the comparison of the input price ratios with the negative inverse marginal rates of technical substitution. As the marginal rates of technical substitution are meaningless if the monotonicity condition is not fulfilled, we limit the comparisons to the observations, where all monotonicity conditions are fulfilled: 119
2 Primal Approach: Production Function −60 −40 −20 0 −60 −40 −20 0 w Cap MVP Cap 0 10 20 30 40 0 10 20 30 40 w Lab MVP Lab 0 100 300 500 0 100 200 300 400 500 w Mat MVP Mat 0.1 0.5 2.0 5.0 0.1 0.5 2.0 5.0 w Cap MVP Cap 0.5 1.0 2.0 5.0 20.0 0.5 1.0 2.0 5.0 10.0 w Lab MVP Lab 5 10 20 50 100 5 10 20 50 100 w Mat MVP Mat Figure 2.34: Marginal value products and corresponding input prices 120
2 Primal Approach: Production Function > compPlot( ( dat$pCap / dat$pLab )[ dat$monoQuad ], + - dat$mrtsLabCapQuad[ dat$monoQuad ] ) > compPlot( ( dat$pCap / dat$pMat )[ dat$monoQuad ], + - dat$mrtsMatCapQuad[ dat$monoQuad ] ) > compPlot( ( dat$pLab / dat$pMat )[ dat$monoQuad ], + - dat$mrtsMatLabQuad[ dat$monoQuad ] ) > compPlot( ( dat$pCap / dat$pLab )[ dat$monoQuad ], + - dat$mrtsLabCapQuad[ dat$monoQuad ], log = "xy" ) > compPlot( ( dat$pCap / dat$pMat )[ dat$monoQuad ], + - dat$mrtsMatCapQuad[ dat$monoQuad ], log = "xy" ) > compPlot( ( dat$pLab / dat$pMat )[ dat$monoQuad ], + - dat$mrtsMatLabQuad[ dat$monoQuad ], log = "xy" ) 0 5 10 15 0 5 10 15 w Cap / w Lab − MRTS Lab Cap 0 1 2 3 4 5 6 0123456 w Cap / w Mat − MRTS Mat Cap 0123 0 1 2 3 w Lab / w Mat − MRTS Mat Lab 0.02 0.10 0.50 2.00 10.00 0.02 0.10 0.50 2.00 10.00 w Cap / w Lab − MRTS Lab Cap 0.001 0.010 0.100 1.000 0.001 0.010 0.100 1.000 w Cap / w Mat − MRTS Mat Cap 0.02 0.10 0.50 2.00 0.02 0.05 0.20 0.50 2.00 w Lab / w Mat − MRTS Mat Lab Figure 2.35: First-order conditions for costs minimization The resulting graphs are shown in figure 2.35. Furthermore, we use histograms to visualize the (absolute and relative) differences between the input price ratios and the corresponding negative inverse marginal rates of technical substitution: > hist( ( - dat$mrtsLabCapQuad - dat$pCap / dat$pLab )[ dat$monoQuad ] ) 121
2 Primal Approach: Production Function I(0.5 * log(qLab)^2) + I(0.5 * log(qMat)^2) + I(log(qCap) * log(qLab)) + I(log(qCap) * log(qMat)) + I(log(qLab) * log(qMat)) Model 2: log(qOut) ~ log(qCap) + log(qMat) + I(0.5 * log(qCap)^2) + I(0.5 * log(qMat)^2) + I(log(qCap) * log(qMat)) #Df LogLik Df Chisq Pr(>Chisq) 1 11 -131.25 2 7 -146.70 -4 30.905 3.201e-06 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > # materials > linearHypothesis( prodTL, c( "log(qMat) = 0", "I(0.5 * log(qMat)^2) = 0", + "I(log(qCap) * log(qMat)) = 0", "I(log(qLab) * log(qMat)) = 0" ) ) Linear hypothesis test: log(qMat) = 0 I(0.5 * log(qMat)^2) = 0 I(log(qCap) * log(qMat)) = 0 I(log(qLab) * log(qMat)) = 0 Model 1: restricted model Model 2: log(qOut) ~ log(qCap) + log(qLab) + log(qMat) + I(0.5 * log(qCap)^2) + I(0.5 * log(qLab)^2) + I(0.5 * log(qMat)^2) + I(log(qCap) * log(qLab)) + I(log(qCap) * log(qMat)) + I(log(qLab) * log(qMat)) Res.Df RSS Df Sum of Sq F Pr(>F) 1 134 68.178 2 130 53.447 4 14.731 8.9578 2.02e-06 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > waldtest( prodTL, log( qOut ) ~ log( qCap ) + log( qLab ) + + I( 0.5 * log( qCap )^2 ) + I( 0.5 * log( qLab )^2 ) + + I( log( qCap ) * log( qLab ) ) ) Wald test Model 1: log(qOut) ~ log(qCap) + log(qLab) + log(qMat) + I(0.5 * log(qCap)^2) + I(0.5 * log(qLab)^2) + I(0.5 * log(qMat)^2) + I(log(qCap) * log(qLab)) + I(log(qCap) * log(qMat)) + I(log(qLab) * log(qMat)) Model 2: log(qOut) ~ log(qCap) + log(qLab) + I(0.5 * log(qCap)^2) + I(0.5 * 128
2 Primal Approach: Production Function log(qLab)^2) + I(log(qCap) * log(qLab)) Res.Df Df F Pr(>F) 1 130 2 134 -4 8.9578 2.02e-06 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > lrtest( prodTL, log( qOut ) ~ log( qCap ) + log( qLab ) + + I( 0.5 * log( qCap )^2 ) + I( 0.5 * log( qLab )^2 ) + + I( log( qCap ) * log( qLab ) ) ) Likelihood ratio test Model 1: log(qOut) ~ log(qCap) + log(qLab) + log(qMat) + I(0.5 * log(qCap)^2) + I(0.5 * log(qLab)^2) + I(0.5 * log(qMat)^2) + I(log(qCap) * log(qLab)) + I(log(qCap) * log(qMat)) + I(log(qLab) * log(qMat)) Model 2: log(qOut) ~ log(qCap) + log(qLab) + I(0.5 * log(qCap)^2) + I(0.5 * log(qLab)^2) + I(log(qCap) * log(qLab)) #Df LogLik Df Chisq Pr(>Chisq) 1 11 -131.25 2 7 -148.28 -4 34.081 7.173e-07 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 All inputs have a statistically significant effect at the 10% significance level, while labour and materials also have statistically significant effects at the 5% and 1% (and at even higher) significance levels. 2.6.4 Properties We cannot see from the estimated coefficients whether the monotonicity condition is fulfilled. The Translog production function cannot be globally monotone, because there will be always a set of input quantities that result in negative marginal products.7The Translog function would only be globally monotone, if all first-order coefficients are positive and all second-order coefficients are zero, which is equivalent to a Cobb-Douglas function. We will check the monotonicity condition at each observation in section 2.6.6. All Translog production functions fulfill the weak and the strong essentiality assumption, because as soon as a single input quantity approaches zero, the right-hand side of equation (2.134) approaches minus infinity (if monotonicity is fulfilled), and thus, the output quantity y= exp(ln y) approaches zero. Hence, if a data set includes observations with a positive output quantity but 7Please note that ln xjis a large negative number if xjis a very small positive number. 129
2 Primal Approach: Production Function at least one input quantity that is zero, strict essentiality cannot be fulfilled in the underlying true production technology so that the Translog production function is not a suitable functional form for analyzing this data set. The input requirement sets derived from Translog production functions are always closed and non-empty. The Translog production function always returns finite, real, non-negative, and single values as long as all input quantities are strictly positive. All Translog production functions are continuous and twice-continuously differentiable. 2.6.5 Predicted output quantities As before, we can easily obtain the predicted output quantities with the fitted method. As we used the logarithmic output quantity as dependent variable in our estimated model, we must use the exponential function to obtain the output quantities measured in levels: > dat$qOutTL <- exp( fitted( prodTL ) ) Now, we can evaluate the “fit” of the model by comparing the observed with the fitted output quantities: > compPlot( dat$qOut, dat$qOutTL ) > compPlot( dat$qOut, dat$qOutTL, log = "xy" ) 0.0e+00 1.0e+07 2.0e+07 0.0e+00 1.0e+07 2.0e+07 observed fitted 1e+05 5e+05 5e+06 1e+05 5e+05 2e+06 1e+07 observed fitted Figure 2.37: Translog production function: fit of the model The resulting graphs are shown in figure 2.37. While the graph in the left panel uses a linear scale for the axes, the graph in the right panel uses a logarithmic scale for both axes. Hence, the deviations from the 45°-line illustrate the absolute deviations in the left panel and the relative deviations in the right panel. The fit of the model looks rather okay, but there are some observations, at which the predicted output quantity is not very close to the observed output quantity. 130
2 Primal Approach: Production Function 2.6.6 Output elasticities The output elasticities calculated from a Translog production function are: ϵi=∂ln y ∂ln xi =αi+X j αij ln xj(2.135) We can simplify the code for computing these output elasticities by using short names for the coefficients: > a1 <- coef( prodTL )[ "log(qCap)" ] > a2 <- coef( prodTL )[ "log(qLab)" ] > a3 <- coef( prodTL )[ "log(qMat)" ] > a11 <- coef( prodTL )[ "I(0.5 * log(qCap)^2)" ] > a22 <- coef( prodTL )[ "I(0.5 * log(qLab)^2)" ] > a33 <- coef( prodTL )[ "I(0.5 * log(qMat)^2)" ] > a12 <- a21 <- coef( prodTL )[ "I(log(qCap) * log(qLab))" ] > a13 <- a31 <- coef( prodTL )[ "I(log(qCap) * log(qMat))" ] > a23 <- a32 <- coef( prodTL )[ "I(log(qLab) * log(qMat))" ] Now, we can use the following commands to calculate the output elasticities in R: > dat$eCapTL <- with( dat, + a1 + a11 * log(qCap) + a12 * log(qLab) + a13 * log(qMat) ) > dat$eLabTL <- with( dat, + a2 + a21 * log(qCap) + a22 * log(qLab) + a23 * log(qMat) ) > dat$eMatTL <- with( dat, + a3 + a31 * log(qCap) + a32 * log(qLab) + a33 * log(qMat) ) We can visualize (the variation of) these output elasticities with histograms: > hist( dat$eCapTL, 15 ) > hist( dat$eLabTL, 15 ) > hist( dat$eMatTL, 15 ) The resulting graphs are shown in figure 2.38. If the firms increase capital input by one percent, the output of most firms will increase by around 0.2 percent. If the firms increase labor input by one percent, the output of most firms will increase by around 0.5 percent. If the firms increase material input by one percent, the output of most firms will increase by around 0.7 percent. These graphs also show that the monotonicity condition is not fulfilled for all observations: > table( dat$eCapTL >= 0 ) FALSE TRUE 32 108 131
2 Primal Approach: Production Function eCap Frequency −0.4 0.0 0.4 0.8 0 5 10 15 20 25 eLab Frequency −1.0 0.0 0.5 1.0 1.5 2.0 0 5 10 15 20 25 eMat Frequency 0.0 0.5 1.0 1.5 2.0 0 10 20 30 Figure 2.38: Translog production function: output elasticities > table( dat$eLabTL >= 0 ) FALSE TRUE 14 126 > table( dat$eMatTL >= 0 ) FALSE TRUE 8 132 > dat$monoTL <- with( dat, eCapTL >= 0 & eLabTL >= 0 & eMatTL >= 0 ) > table( dat$monoTL ) FALSE TRUE 48 92 32 firms have a negative output elasticity of capital, 14 firms have a negative output elasticity of labor, and 8 firms have a negative output elasticity of materials. In total the monotonicity condition is not fulfilled at 48 out of 140 observations. Although the monotonicity conditions are fulfilled for a large part of firms in our data set, these frequent violations indicate a possible model misspecification. 2.6.7 Marginal products The first derivatives (marginal products) of the Translog production function with respect to the input quantities are: MPi=∂y ∂xi =y xi ∂ln y ∂ln xi =y xi αi+X j αij ln xj (2.136) We can calculate the marginal products based on the output elasticities that we have calculated above. As argued in section 2.4.11.1, we use the predicted output quantities in this calculation: 132
2 Primal Approach: Production Function > dat$mpCapTL <- with( dat, eCapTL * qOutTL / qCap ) > dat$mpLabTL <- with( dat, eLabTL * qOutTL / qLab ) > dat$mpMatTL <- with( dat, eMatTL * qOutTL / qMat ) We can visualize (the variation of) these marginal products with histograms: > hist( dat$mpCapTL, 15 ) > hist( dat$mpLabTL, 15 ) > hist( dat$mpMatTL, 15 ) mpCapTL Frequency −10 0 10 20 0 5 10 15 20 mpLabTL Frequency −5 0 5 10 15 20 25 0 5 10 15 20 mpMatTL Frequency 0 50 100 0 5 10 15 20 25 Figure 2.39: Translog production function: marginal products The resulting graphs are shown in figure 2.39. If the firms increase capital input by one unit, the output of most firms will increase by around 4 units. If the firms increase labor input by one unit, the output of most firms will increase by around 4 units. If the firms increase material input by one unit, the output of most firms will increase by around 70 units. 2.6.8 Elasticity of scale The elasticity of scale can—as always—be calculated as the sum of all output elasticities. > dat$eScaleTL <- dat$eCapTL + dat$eLabTL + + dat$eMatTL The (variation of the) elasticities of scale can be visualized with a histogram. > hist( dat$eScaleTL, 30 ) > hist( dat$eScaleTL[ dat$monoTL ], 30 ) The resulting graphs are shown in figure 2.40. All firms experience increasing returns to scale and most of them have an elasticity of scale around 1.45. Hence, if these firms increase all input quantities by one percent, the output of most firms will increase by around 1.45 percent. These elasticities of scale are realistic and on average close to the elasticity of scale obtained from the Cobb-Douglas production function (1.47). 133
2 Primal Approach: Production Function eScaleTL Frequency 1.2 1.3 1.4 1.5 1.6 1.7 0 5 10 15 eScaleTL[ monoTL ] Frequency 1.2 1.3 1.4 1.5 1.6 1.7 02468 Figure 2.40: Translog production function: elasticities of scale Information on the optimal firm size can be obtained by analyzing the relationship between firm size and the elasticity of scale. We can either use the observed or the predicted output: > plot( dat$qOut, dat$eScaleTL, log = "x" ) > plot( dat$X, dat$eScaleTL, log = "x" ) > plot( dat$qOut[ dat$monoTL ], dat$eScaleTL[ dat$monoTL ], log = "x" ) > plot( dat$X[ dat$monoTL ], dat$eScaleTL[ dat$monoTL ], log = "x" ) The resulting graphs are shown in figure 2.41. Both of them indicate that the elasticity of scale slightly decreases with firm size but there are considerable increasing returns to scale even for the largest firms in the sample. Hence, all firms in the sample would gain from increasing their size and the optimal firm size seems to be larger than the largest firm in the sample. 2.6.9 Marginal rates of technical substitution We can calculate the marginal rates of technical substitution (MRTS) based on our estimated Translog production function by following commands: > dat$mrtsCapLabTL <- with( dat, - mpLabTL / mpCapTL ) > dat$mrtsLabCapTL <- with( dat, - mpCapTL / mpLabTL ) > dat$mrtsCapMatTL <- with( dat, - mpMatTL / mpCapTL ) > dat$mrtsMatCapTL <- with( dat, - mpCapTL / mpMatTL ) > dat$mrtsLabMatTL <- with( dat, - mpMatTL / mpLabTL ) > dat$mrtsMatLabTL <- with( dat, - mpLabTL / mpMatTL ) As the marginal rates of technical substitution are meaningless if the monotonicity condition is not fulfilled, we visualize (the variation of) these MRTS only for the observations, where the monotonicity condition is fulfilled: 134
2 Primal Approach: Production Function 1e+05 5e+05 2e+06 1e+07 1.2 1.4 1.6 observed output eScaleTL 0.5 1.0 2.0 5.0 1.2 1.4 1.6 quantity index of inputs eScaleTL 1e+05 5e+05 2e+06 1e+07 1.2 1.3 1.4 1.5 1.6 1.7 observed output eScaleTL[ monoTL ] 0.5 1.0 2.0 5.0 1.2 1.3 1.4 1.5 1.6 1.7 quantity index of inputs eScaleTL[ monoTL ] Figure 2.41: Translog production function: elasticities of scale at different firm sizes > hist( dat$mrtsCapLabTL[ dat$monoTL ], 30 ) > hist( dat$mrtsLabCapTL[ dat$monoTL ], 30 ) > hist( dat$mrtsCapMatTL[ dat$monoTL ], 30 ) > hist( dat$mrtsMatCapTL[ dat$monoTL ], 30 ) > hist( dat$mrtsLabMatTL[ dat$monoTL ], 30 ) > hist( dat$mrtsMatLabTL[ dat$monoTL ], 30 ) The resulting graphs are shown in figure 2.43. As some outliers hide the variation of the majority of the MRTS, we use function colMedians (package miscTools) to show the median values of the MRTS: > colMedians( subset( dat, monoTL, + c( "mrtsCapLabTL", "mrtsLabCapTL", "mrtsCapMatTL", + "mrtsMatCapTL", "mrtsLabMatTL", "mrtsMatLabTL" ) ) ) mrtsCapLabTL mrtsLabCapTL mrtsCapMatTL mrtsMatCapTL mrtsLabMatTL mrtsMatLabTL -0.83929283 -1.19196521 -12.72554396 -0.07858435 -12.79850828 -0.07813810 Given that the median marginal rate of technical substitution between capital and labor is -0.84, a typical firm that reduces the use of labor by one unit, has to use around 0.84 additional units of capital in order to produce the same amount of output as before. Alternatively, the typical firm can replace one unit of labor by using 0.08 additional units of materials. 135
2 Primal Approach: Production Function mrtsCapLabTL Frequency −150 −100 −50 0 0 20 40 60 mrtsLabCapTL Frequency −60 −40 −20 0 0 10 20 30 40 50 60 mrtsCapMatTL Frequency −800 −600 −400 −200 0 0 10 20 30 40 50 60 70 mrtsMatCapTL Frequency −0.6 −0.4 −0.2 0.0 0 5 10 15 mrtsLabMatTL Frequency −1200 −800 −400 0 0 20 40 60 80 mrtsMatLabTL Frequency −4 −3 −2 −1 0 0 10 20 30 40 50 Figure 2.42: Translog production function: marginal rates of technical substitution (MRTS) 2.6.10 Relative marginal rates of technical substitution As we do not have a practical interpretation of the units of measurement of the input quantities, the relative marginal rates of technical substitution (RMRTS) are practically more meaningful than the MRTS. The following commands calculate the RMRTS: > dat$rmrtsCapLabTL <- with( dat, - eLabTL / eCapTL ) > dat$rmrtsLabCapTL <- with( dat, - eCapTL / eLabTL ) > dat$rmrtsCapMatTL <- with( dat, - eMatTL / eCapTL ) > dat$rmrtsMatCapTL <- with( dat, - eCapTL / eMatTL ) > dat$rmrtsLabMatTL <- with( dat, - eMatTL / eLabTL ) > dat$rmrtsMatLabTL <- with( dat, - eLabTL / eMatTL ) As the (relative) marginal rates of technical substitution are meaningless if the monotonicity condition is not fulfilled, we visualize (the variation of) these RMRTS only for the observations, where the monotonicity condition is fulfilled: > hist( dat$rmrtsCapLabTL[ dat$monoTL ], 30 ) > hist( dat$rmrtsLabCapTL[ dat$monoTL ], 30 ) > hist( dat$rmrtsCapMatTL[ dat$monoTL ], 30 ) > hist( dat$rmrtsMatCapTL[ dat$monoTL ], 30 ) 136
2 Primal Approach: Production Function > hist( dat$rmrtsLabMatTL[ dat$monoTL ], 30 ) > hist( dat$rmrtsMatLabTL[ dat$monoTL ], 30 ) rmrtsCapLabTL Frequency −500 −300 −100 0 0 20 40 60 80 rmrtsLabCapTL Frequency −25 −20 −15 −10 −5 0 0 20 40 60 rmrtsCapMatTL Frequency −400 −300 −200 −100 0 0 20 40 60 80 rmrtsMatCapTL Frequency −3.0 −2.0 −1.0 0.0 0 5 10 15 20 rmrtsLabMatTL Frequency −50 −40 −30 −20 −10 0 0 10 20 30 40 50 60 rmrtsMatLabTL Frequency −35 −25 −15 −5 0 0 10 20 30 40 50 Figure 2.43: Translog production function: relative marginal rates of technical substitution (RMRTS) The resulting graphs are shown in figure 2.43. As some outliers hide the variation of the majority of the RMRTS, we use function colMedians (package miscTools) to show the median values of the RMRTS: > colMedians( subset( dat, monoTL, + c( "rmrtsCapLabTL", "rmrtsLabCapTL", "rmrtsCapMatTL", + "rmrtsMatCapTL", "rmrtsLabMatTL", "rmrtsMatLabTL" ) ) ) rmrtsCapLabTL rmrtsLabCapTL rmrtsCapMatTL rmrtsMatCapTL rmrtsLabMatTL -2.8357239 -0.3539150 -3.0064237 -0.3331325 -1.3444115 rmrtsMatLabTL -0.7439008 Given that the median relative marginal rate of technical substitution between capital and labor is -2.84, a typical firm that reduces the use of labor by one percent, has to use around 2.84 percent more capital in order to produce the same amount of output as before. Alternatively, the typical firm can replace one percent of labor by using 0.74 percent more materials. 137
2 Primal Approach: Production Function > dat$FCapLabTL <- - apply( bhmTL[ -2, -3, ], 3, det ) > dat$FCapMatTL <- apply( bhmTL[ -2, -4, ], 3, det ) > dat$FLabMatTL <- - apply( bhmTL[ -3, -4, ], 3, det ) > dat$FCapCapTL <- apply( bhmTL[ -2, -2, ], 3, det ) > dat$FLabLabTL <- apply( bhmTL[ -3, -3, ], 3, det ) > dat$FMatMatTL <- apply( bhmTL[ -4, -4, ], 3, det ) > dat$bhmDetTL <- apply( bhmTL, 3, det ) > dat$numeratorTL <- with( dat, + qCap * mpCapTL + qLab * mpLabTL + qMat * mpMatTL ) > dat$esaCapLabTL <- with( dat, + numeratorTL / ( qCap * qLab ) * FCapLabTL / bhmDetTL ) > dat$esaCapMatTL <- with( dat, + numeratorTL / ( qCap * qMat ) * FCapMatTL / bhmDetTL ) > dat$esaLabMatTL <- with( dat, + numeratorTL / ( qLab * qMat ) * FLabMatTL / bhmDetTL ) > dat$esaCapCapTL <- with( dat, + numeratorTL / ( qCap * qCap ) * FCapCapTL / bhmDetTL ) > dat$esaLabLabTL <- with( dat, + numeratorTL / ( qLab * qLab ) * FLabLabTL / bhmDetTL ) > dat$esaMatMatTL <- with( dat, + numeratorTL / ( qMat * qMat ) * FMatMatTL / bhmDetTL ) Before we take a look at and interpret the elasticities of substitution, we check whether the conditions (2.132) are fulfilled: > range( with( dat, qCap * mpCapTL * esaCapCapTL + + qLab * mpLabTL * esaCapLabTL + qMat * mpMatTL * esaCapMatTL ) ) [1] -5.215406e-08 9.536743e-07 > range( with( dat, qCap * mpCapTL * esaCapLabTL + + qLab * mpLabTL * esaLabLabTL + qMat * mpMatTL * esaLabMatTL ) ) [1] -2.980232e-08 5.960464e-08 > range( with( dat, qCap * mpCapTL * esaCapMatTL + + qLab * mpLabTL * esaLabMatTL + qMat * mpMatTL * esaMatMatTL ) ) [1] -2.235174e-08 5.960464e-07 The extremely small deviations from zero are most likely caused by rounding errors that are unavoidable on digital computers. This test does not prove that all of our calculations are done 144
2 Primal Approach: Production Function correctly but if we had made a mistake, we probably would have discovered it. Hence, we can be rather sure that our calculations are correct. As explained above in section 2.6.13.1, the elasticities of substitution are meaningless if the monotonicity condition is violated and they are difficult to interpret if the quasiconcavity condition is violated. Hence, we visualize (the variation of) the Allen elasticities of substitution only for the observations, where both the monotonicity condition and the quasiconcavity condition are fulfilled: > hist( dat$esaCapLabTL[ dat$monoTL & dat$quasiConcTL ], 30 ) > hist( dat$esaCapMatTL[ dat$monoTL & dat$quasiConcTL ], 30 ) > hist( dat$esaLabMatTL[ dat$monoTL & dat$quasiConcTL ], 30 ) > hist( dat$esaCapLabTL[ dat$monoTL & dat$quasiConcTL & + abs( dat$esaCapLabTL ) < 10 ], 30 ) > hist( dat$esaCapMatTL[ dat$monoTL & dat$quasiConcTL & + abs( dat$esaCapMatTL ) < 10 ], 30 ) > hist( dat$esaLabMatTL[ dat$monoTL & dat$quasiConcTL & + abs( dat$esaLabMatTL ) < 10 ], 30 ) esaCapLabTL Frequency −400 −300 −200 −100 0 0 10 20 30 40 esaCapMatTL Frequency 0 500 1000 1500 0 10 20 30 40 50 esaLabMatTL Frequency 0 50 100 150 0 10 20 30 40 abs( esaCapLabTL ) < 10 Frequency −8 −6 −4 −2 0 0 2 4 6 8 10 12 14 abs( esaCapMatTL ) < 10 Frequency 2 3 4 5 6 7 8 01234567 abs( esaLabMatTL ) < 10 Frequency 0246 0 5 10 15 Figure 2.46: Translog production function: Allen elasticities of substitution The resulting graphs are shown in figure 2.46. The estimated elasticities of substitution between capital and labor suggest that capital and labor are complements for the majority of firms. In 145
2 Primal Approach: Production Function contrast, capital and materials are always substitutes and labor and materials are substitutes for the majority of firms. In order to avoid the effects of outliers, we use function colMedians (package miscTools) to obtain the median values of the Allen elasticities of substitution: > colMedians( subset( dat, monoTL & quasiConcTL, + c( "esaCapLabTL", "esaCapMatTL", "esaLabMatTL" ) ) ) esaCapLabTL esaCapMatTL esaLabMatTL -0.8018194 3.6322399 0.5848447 The median elasticity of substitution between labor and materials (0.58) lies between the elasticity of substitution implied by the Leontief production function (σ= 0) and the elasticity of substitution implied by the Cobb-Douglas production function (σ= 1). Hence, the substitutability between labor and materials seems to be rather low. If a typical firm keeps the output quantity constant, substitutes materials for labor (or vice versa) so that the ratio between the quantity of materials and the labor quantity increases (decreases) by 0.58 percent, and adjusts the capital quantity accordingly, the MRTS between materials and labor will increase (decrease) by one percent. If the ratio between the labor price and the materials price increases by one percent, a firm that keeps the output quantity constant, adjusts all input quantities and minimizes costs, will substitute materials for labor so that the ratio between the quantity of materials and the labor quantity increases by 0.58 percent. Hence, the relative change of the ratio of the input quantities is considerably smaller than the relative change of the ratio of the input prices, which indicates a low substitutability between labor and materials. In contrast, the median elasticity of substitution between capital and materials is larger than one (3.63), which indicates that it is much easier to substitute between capital and materials. 2.6.13.3 Comparison of direct and Allen elasticities of substitution In the following, we use scatter plots to compare the estimated direct elasticities of substitution with the estimated Allen elasticities of substitution: > compPlot( dat$esdCapLabTL[ dat$monoTL & dat$quasiConcTL ], + dat$esaCapLabTL[ dat$monoTL & dat$quasiConcTL ], + lim = c( -2, 2 ) ) > compPlot( dat$esdCapMatTL[ dat$monoTL & dat$quasiConcTL ], + dat$esaCapMatTL[ dat$monoTL & dat$quasiConcTL ], + lim = c( -10, 10 ) ) > compPlot( dat$esdLabMatTL[ dat$monoTL & dat$quasiConcTL ], + dat$esaLabMatTL[ dat$monoTL & dat$quasiConcTL ], + lim = c( -2, 2 ) ) 146
2 Primal Approach: Production Function −2 −1 0 1 2 −2 −1 0 1 2 esdCapLabTL esaCapLabTL −10 −5 0 5 10 −10 −5 0 5 10 esdCapMatTL esaCapMatTL −2 −1 0 1 2 −2 −1 0 1 2 esdLabMatTL esaLabMatTL Figure 2.47: Translog production function: Comparison of direct and Allen elasticities of substitution The resulting graphs are shown in figure 2.47. As can also be seen from comparing the histograms in Figures 2.45 and 2.46, the variation of the Allen elasticities of substitution is much larger than the variation of the corresponding direct elasticities of substitution. 2.6.14 Mean-scaled quantities The Translog functional form is often estimated with mean-scaled variables. In the case of a Translog production function, it is estimated with mean-scaled input quantities and sometimes also with mean-scaled output quantities: y∗=y/¯y(2.143) x∗ i=xi/¯xi∀i(2.144) The following commands create variables with mean-scaled output and input quantities: > dat$qmOut <- with( dat, qOut / mean( qOut ) ) > dat$qmCap <- with( dat, qCap / mean( qCap ) ) > dat$qmLab <- with( dat, qLab / mean( qLab ) ) > dat$qmMat <- with( dat, qMat / mean( qMat ) ) This implies that the logarithms of the mean values of these variables are zero (except for negligible very small rounding errors): > log( colMeans( dat[ , c( "qmOut", "qmCap", "qmLab", "qmMat" ) ] ) ) qmOut qmCap qmLab qmMat -1.110223e-16 -1.110223e-16 0.000000e+00 0.000000e+00 Please note that mean-scaling does not imply that the mean values of the logarithmic variables are zero: 147
2 Primal Approach: Production Function > colMeans( log( dat[ , c( "qmOut", "qmCap", "qmLab", "qmMat" ) ] ) ) qmOut qmCap qmLab qmMat -0.4860021 -0.3212057 -0.1565112 -0.2128551 The following derivation explores the relationship between a Translog production function with mean-scaled quantities and a Translog production function with the original quantities: ln y∗=α∗ 0+X i α∗ iln x∗ i+1 2X iX j α∗ ij ln x∗ iln x∗ j(2.145) ln (y/¯y) = α∗ 0+X i α∗ iln (xi/¯xi) + 1 2X iX j α∗ ij ln (xi/¯xi) ln (xj/¯xj)(2.146) ln y−ln ¯y=α∗ 0+X i α∗ i(ln xi−ln ¯xi) + 1 2X iX j α∗ ij (ln xi−ln ¯xi) (ln xj−ln ¯xj)(2.147) ln y=α∗ 0+ ln ¯y+X i α∗ iln xi−X i α∗ iln ¯xi(2.148) +1 2X iX j α∗ ij ln xiln xj−X iX j α∗ ij ln xiln ¯xj+1 2X iX j α∗ ij ln ¯xiln ¯xj ln y= α∗ 0+ ln ¯y−X i α∗ iln ¯xi+1 2X iX j α∗ ij ln ¯xiln ¯xj (2.149) +X i α∗ i−X j α∗ ij ln ¯xj ln xi+1 2X iX j α∗ ij ln xiln xj Thus, the relationship between the coefficients of the Translog function with the original quantities and the coefficients of the Translog function with mean-scaled quantities is: α0=α∗ 0+ ln ¯y−X i α∗ iln ¯xi+1 2X iX j α∗ ij ln ¯xiln ¯xj(2.150) αi=α∗ i−X j α∗ ij ln ¯xj∀i(2.151) αij =α∗ ij ∀i, j. (2.152) Accordingly, the reciprocal relationship is: α∗ 0=α0−ln ¯y+X i αiln ¯xi+1 2X iX j αij ln ¯xiln ¯xj(2.153) α∗ i=αi+X j αij ln ¯xj∀i(2.154) α∗ ij =αij ∀i, j. (2.155) The following command estimates the Translog production function with mean-scaled quantities: 148
2 Primal Approach: Production Function > prodTLm <- lm( log( qmOut ) ~ log( qmCap ) + log( qmLab ) + log( qmMat ) + + I( 0.5 * log( qmCap )^2 ) + I( 0.5 * log( qmLab )^2 ) + + I( 0.5 * log( qmMat )^2 ) + I( log( qmCap ) * log( qmLab ) ) + + I( log( qmCap ) * log( qmMat ) ) + I( log( qmLab ) * log( qmMat ) ), + data = dat ) > summary( prodTLm ) Call: lm(formula = log(qmOut) ~ log(qmCap) + log(qmLab) + log(qmMat) + I(0.5 * log(qmCap)^2) + I(0.5 * log(qmLab)^2) + I(0.5 * log(qmMat)^2) + I(log(qmCap) * log(qmLab)) + I(log(qmCap) * log(qmMat)) + I(log(qmLab) * log(qmMat)), data = dat) Residuals: Min 1Q Median 3Q Max -1.68015 -0.36688 0.05389 0.44125 1.26560 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) -0.09392 0.08815 -1.065 0.28864 log(qmCap) 0.15004 0.11134 1.348 0.18013 log(qmLab) 0.79339 0.17477 4.540 1.27e-05 *** log(qmMat) 0.50201 0.16608 3.023 0.00302 ** I(0.5 * log(qmCap)^2) -0.02573 0.20834 -0.124 0.90189 I(0.5 * log(qmLab)^2) -1.16364 0.67943 -1.713 0.08916 . I(0.5 * log(qmMat)^2) -0.50368 0.43498 -1.158 0.24902 I(log(qmCap) * log(qmLab)) 0.56194 0.29120 1.930 0.05582 . I(log(qmCap) * log(qmMat)) -0.40996 0.23534 -1.742 0.08387 . I(log(qmLab) * log(qmMat)) 0.65793 0.42750 1.539 0.12623 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 0.6412 on 130 degrees of freedom Multiple R-squared: 0.6296, Adjusted R-squared: 0.6039 F-statistic: 24.55 on 9 and 130 DF, p-value: < 2.2e-16 As expected, the intercept and the first-order coefficients have adjusted to the new units of measurement, while the second-order coefficients of the Translog function remain unchanged (compare with estimates in section 2.6.2): > all.equal( coef(prodTL)[-c(1:4)], coef(prodTLm)[-c(1:4)], 149
2 Primal Approach: Production Function + check.attributes = FALSE ) [1] TRUE In case of functional forms that are invariant to the units of measurement (e.g. linear, CobbDouglas, quadratic, Translog), mean-scaling does not change the relative indicators of the technology (e.g. output elasticities, elasticities of scale, relative marginal rates of technical substitution, elasticities of substitution). As the logarithms of the mean values of the mean-scaled input quantities are zero, the first-order coefficients are equal to the output elasticities at the sample mean (see equation 2.135), i.e. the output elasticity of capital is 0.15, the output elasticity of labor is 0.793, the output elasticity of materials is 0.502, and the elasticity of scale is 1.445 at the sample mean. 2.6.15 First-order conditions for profit maximization In this section, we will check to what extent the first-order conditions for profit maximization (2.40) are fulfilled, i.e. to what extent the firms use the optimal input quantities. We do this by comparing the marginal value products of the inputs with the corresponding input prices. We can calculate the marginal value products by multiplying the marginal products by the output price: > dat$mvpCapTL <- dat$pOut * dat$mpCapTL > dat$mvpLabTL <- dat$pOut * dat$mpLabTL > dat$mvpMatTL <- dat$pOut * dat$mpMatTL The command compPlot (package miscTools) can be used to compare the marginal value products with the corresponding input prices. As the logarithm of a non-positive number is not defined, we have to limit the comparisons on the logarithmic scale to observations with positive marginal products: > compPlot( dat$pCap, dat$mvpCapTL ) > compPlot( dat$pLab, dat$mvpLabTL ) > compPlot( dat$pMat, dat$mvpMatTL ) > compPlot( dat$pCap[ dat$monoTL ], dat$mvpCapTL[ dat$monoTL ], log = "xy" ) > compPlot( dat$pLab[ dat$monoTL ], dat$mvpLabTL[ dat$monoTL ], log = "xy" ) > compPlot( dat$pMat[ dat$monoTL ], dat$mvpMatTL[ dat$monoTL ], log = "xy" ) The resulting graphs are shown in figure 2.48. They indicate that the marginal value products of most firms are higher than the corresponding input prices. This indicates that most firms could increase their profit by using more of all inputs. Given that the estimated Translog function shows that all firms operate under increasing returns to scale, it is not surprising that most firms would gain from increasing all input quantities. Therefore, the question arises why the firms in the sample did not do this. This questions has already been addressed in section 2.3.10. 150
2 Primal Approach: Production Function −40 0 20 40 60 −40 −20 0 20 40 60 w Cap MVP Cap −5 0 5 10 15 20 25 30 −5 0 5 10 15 20 25 30 w Lab MVP Lab 0 50 100 150 0 50 100 150 w Mat MVP Mat 5e−02 5e−01 5e+00 5e+01 5e−02 5e−01 5e+00 5e+01 w Cap MVP Cap 0.05 0.20 1.00 5.00 20.00 0.05 0.20 1.00 5.00 20.00 w Lab MVP Lab 5 10 20 50 100 5 10 20 50 100 200 w Mat MVP Mat Figure 2.48: Marginal value products and corresponding input prices 151
2 Primal Approach: Production Function 2.6.16 First-order conditions for cost minimization As the marginal rates of technical substitution differ between observations for the three other functional forms, we use scatter plots for visualizing the comparison of the input price ratios with the negative inverse marginal rates of technical substitution: As the marginal rates of technical substitution are meaningless if the monotonicity condition is not fulfilled, we limit the comparisons to the observations, where all monotonicity conditions are fulfilled: > compPlot( ( dat$pCap / dat$pLab )[ dat$monoTL ], + - dat$mrtsLabCapTL[ dat$monoTL ] ) > compPlot( ( dat$pCap / dat$pMat )[ dat$monoTL ], + - dat$mrtsMatCapTL[ dat$monoTL ] ) > compPlot( ( dat$pLab / dat$pMat )[ dat$monoTL ], + - dat$mrtsMatLabTL[ dat$monoTL ] ) > compPlot( ( dat$pCap / dat$pLab )[ dat$monoTL ], + - dat$mrtsLabCapTL[ dat$monoTL ], log = "xy" ) > compPlot( ( dat$pCap / dat$pMat )[ dat$monoTL ], + - dat$mrtsMatCapTL[ dat$monoTL ], log = "xy" ) > compPlot( ( dat$pLab / dat$pMat )[ dat$monoTL ], + - dat$mrtsMatLabTL[ dat$monoTL ], log = "xy" ) The resulting graphs are shown in figure 2.49. Furthermore, we use histograms to visualize the (absolute and relative) differences between the input price ratios and the corresponding negative inverse marginal rates of technical substitution: > hist( ( - dat$mrtsLabCapTL - dat$pCap / dat$pLab )[ dat$monoTL ] ) > hist( ( - dat$mrtsMatCapTL - dat$pCap / dat$pMat )[ dat$monoTL ] ) > hist( ( - dat$mrtsMatLabTL - dat$pLab / dat$pMat )[ dat$monoTL ] ) > hist( log( - dat$mrtsLabCapTL / ( dat$pCap / dat$pLab ) )[ dat$monoTL ] ) > hist( log( - dat$mrtsMatCapTL / ( dat$pCap / dat$pMat ) )[ dat$monoTL ] ) > hist( log( - dat$mrtsMatLabTL / ( dat$pLab / dat$pMat ) )[ dat$monoTL ] ) The resulting graphs are shown in figure 2.50. The graphs in the middle column of figures 2.49 and 2.50 show that the ratio between the capital price and the materials price is larger than the absolute value of the marginal rate of technical substitution between materials and capital for a majority of the firms in the sample: wcap wmat >−MRTSmat,cap =MPcap MPmat (2.156) Hence, these firms can get closer to the minimum of their production costs by substituting materials for capital, because this will decrease the marginal product of materials and increase the marginal product of capital so that the absolute value of the MRTS between materials and 152
2 Primal Approach: Production Function 0 20 40 60 0 20 40 60 w Cap / w Lab − MRTS Lab Cap 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.0 0.1 0.2 0.3 0.4 0.5 0.6 w Cap / w Mat − MRTS Mat Cap 01234 0 1 2 3 4 w Lab / w Mat − MRTS Mat Lab 1e−02 1e+00 1e+02 1e−02 1e−01 1e+00 1e+01 1e+02 w Cap / w Lab − MRTS Lab Cap 0.001 0.005 0.050 0.500 0.001 0.005 0.020 0.100 0.500 w Cap / w Mat − MRTS Mat Cap 0.001 0.010 0.100 1.000 0.001 0.010 0.100 1.000 w Lab / w Mat − MRTS Mat Lab Figure 2.49: First-order conditions for costs minimization 153
6 Stochastic Frontier Analysis sigmaSq 1.000040 0.202456 4.9396 7.830e-07 *** gamma 0.896664 0.070952 12.6375 < 2.2e-16 *** sigmaSqU 0.896700 0.241715 3.7097 0.0002075 *** sigmaSqV 0.103340 0.055831 1.8509 0.0641777 . sigma 1.000020 0.101226 9.8791 < 2.2e-16 *** sigmaU 0.946942 0.127629 7.4195 1.176e-13 *** sigmaV 0.321465 0.086838 3.7019 0.0002140 *** lambdaSq 8.677179 6.644542 1.3059 0.1915829 lambda 2.945705 1.127835 2.6118 0.0090061 ** varU 0.325843 NA NA NA sdU 0.570827 NA NA NA gammaVar 0.759217 NA NA NA --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -133.8893 cross-sectional data total number of observations = 140 mean efficiency: 0.5379937 The additionally returned parameter are defined as follows: sigmaSqU =σ2 u=σ2·γ,sigmaSqV =σ2 v=σ2·(1 −γ) = V ar (v),sigma =σ=√σ2,sigmaU =σu=pσ2 u,sigmaV =σv=pσ2 v, lambdaSq =λ2=σ2 u/σ2 v,lambda =λ=σu/σv,varU =V ar (u),sdU =pV ar (u), and gammaVar =V ar (u)/(V ar (u) + V ar (v)). 6.1.3.3 Statistical tests for inefficiencies If there would be no inefficiencies, i.e. u= 0 for all observations, coefficient γwould be equal to zero. Hence, one can test the null hypothesis of no inefficiencies by simply testing whether γis equal to (does not significantly deviate from) zero. However, a t-test of the null hypothesis γ= 0 (e.g. reported in the output of the summary method) is not valid, because γis bound to the interval [0,1] and hence, cannot follow a t-distribution. Instead, we can use a likelihood ratio test to check whether adding the inefficiency term usignificantly improves the fit of the model. If the lrtest method is called just with a single stochastic frontier model, it compares the stochastic frontier model with the corresponding OLS model (i.e. a model with γequal to zero): > lrtest( prodCDSfa ) Likelihood ratio test 256
6 Stochastic Frontier Analysis Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 5 -137.61 2 6 -133.89 1 7.4387 0.003192 ** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Under the null hypothesis (no inefficiency, only noise), the test statistic asymptotically follows a mixed χ2-distribution (Coelli,1995).3The rather small P-value indicates that the data clearly reject the OLS model in favor of the stochastic frontier model, i.e. there is significant technical inefficiency. 6.1.3.4 Obtaining technical efficiency estimates As neither the noise term vnor the inefficiency term ubut only the total error term ε=−u+v is known, the technical efficiencies TE =e−uare generally unknown. However, given that the parameter estimates (including the parameters σ2and γor σ2 vand σ2 u) and the total error term ε are known, it is possible to determine the expected value of the technical efficiency (see, e.g. Coelli et al.,2005, p. 255): d TE =Ee−u(6.12) These efficiency estimates can be obtained by the efficiencies method: > dat$effCD <- efficiencies( prodCDSfa ) The following two commands create a scatter plot that visualizes the relationship between the residuals (y−f(x) = ε=−u+v)and the efficiency estimates (d TE =E[e−u]) and a histogram that visualizes the variation of the efficiency estimates, respectively: > plot( residuals( prodCDSfa ), dat$effCD ) > hist( dat$effCD, 15 ) The resulting graphs are shown in figure 6.3. The efficiency estimates are rather low: the firms only produce between 10% and 90% of the maximum possible output quantities. In order to illustrate how efficiency estimates are obtained, e.g., with the efficiencies() method, we manually calculate the approximate efficiency estimate of the first observation: > # choose the observation number > obsNo <- 1 > # print the automatically obtained efficiency estimate for this observation > dat$effCD[ obsNo ] 3As a standard likelihood ratio test assumes that the test statistic follows a (standard) χ2-distribution under the null hypothesis, a test that is conducted by the command lrtest( prodCD, prodCDSfa ) returns an incorrect P-value. 257
6 Stochastic Frontier Analysis −2.5 −1.5 −0.5 0.5 0.2 0.4 0.6 0.8 residCD effCD effCD Frequency 0.2 0.4 0.6 0.8 0 5 10 15 Figure 6.3: Efficiency estimates of Cobb-Douglas production frontier [1] 0.262434 > # obtain the residual of this observation > r <- residuals( prodCDSfa )[ obsNo ] > # create a vector with a range of technical efficiencies > TE <- c( 1:100 ) / 100 > # calculate values of the inefficiency term u that correspond to the values of TE > u <- - log( TE ) > # calculate values of the noise term v that correspond to the values of TE >v<-r+u > # obtain sigma_U > su <- coef( summary( prodCDSfa, extraPar = TRUE ) )[ "sigmaU", 1 ] > # sigma_V > sv <- coef( summary( prodCDSfa, extraPar = TRUE ) )[ "sigmaV", 1 ] > # relative probabilities of the values of TE > prTE <- dnorm( v, sd = sv ) * ( 2 * dnorm( u, sd = su ) ) > # probabilities of the values of TE > pTE <- prTE / sum( prTE ) > # approximate expected TE > sum( TE * pTE ) [1] 0.287908 We can also illustrate the probabilities of the technical efficiencies of the first observation in a density plot: 258
6 Stochastic Frontier Analysis > plot( TE, pTE, type = "l" ) 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.01 0.02 0.03 0.04 0.05 TE pTE Figure 6.4: Probabilities of the technical efficiencies of the first observation The resulting graph is shown in figure 6.4. We explore the correlation between firm size (measured as output quantity as well as aggregate input quantity indicated by a Fisher quantity index of all inputs) and the efficiency estimates: > plot( dat$qOut, dat$effCD, log = "x" ) > plot( dat$X, dat$effCD, log = "x" ) The resulting graphs are shown in figure 6.5. As the efficiency directly influences the output quantity, it is not surprising that the efficiency estimates are highly correlated with the output quantity. On the other hand, the efficiency estimates are only slightly correlated with firm size measured as aggregate input quantity. However, the largest firms all have an above-average efficiency estimate. 6.1.3.5 Truncated normal distribution of the inefficiency term Instead of assuming that the inefficiency term ufollows a half-normal distribution, we can assume that it follows a truncated norm distribution: u∼N+(µ, σ2 u)(6.13) 259
6 Stochastic Frontier Analysis 1e+05 5e+05 5e+06 0.2 0.4 0.6 0.8 qOut effCD 0.5 1.0 2.0 5.0 0.2 0.4 0.6 0.8 X effCD Figure 6.5: Firm size and efficiency estimates of Cobb-Douglas production frontier If the location parameter µis equal to zero, the truncated normal distribution is identical to the half-normal distribution. If argument truncNorm of function sfa() of the frontier package is set to TRUE, this function estimates a stochastic frontier model assuming a truncated normal distribution of the inefficiency term: > prodCDSfaTn <- sfa( log( qOut ) ~ log( qCap ) + log( qLab ) + log( qMat ), + data = dat, truncNorm = TRUE ) > summary( prodCDSfaTn, extraPar = TRUE ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged Iterative ML estimation terminated after 13 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 0.227323 1.250360 0.1818 0.8557349 log(qCap) 0.160632 0.079397 2.0232 0.0430568 * log(qLab) 0.688275 0.148699 4.6286 3.681e-06 *** log(qMat) 0.464421 0.132878 3.4951 0.0004739 *** sigmaSq 0.934405 0.599621 1.5583 0.1191562 gamma 0.895931 0.069525 12.8865 < 2.2e-16 *** 260
6 Stochastic Frontier Analysis mu 0.130129 1.145962 0.1136 0.9095912 sigmaSqU 0.837163 0.557136 1.5026 0.1329373 sigmaSqV 0.097242 0.077929 1.2478 0.2120920 sigma 0.966646 0.310155 3.1167 0.0018292 ** sigmaU 0.914966 0.304457 3.0052 0.0026537 ** sigmaV 0.311837 0.124951 2.4957 0.0125720 * lambdaSq 8.609044 6.419490 1.3411 0.1798948 lambda 2.934117 1.093939 2.6822 0.0073149 ** varU 0.331133 NA NA NA sdU 0.575442 NA NA NA gammaVar 0.772998 NA NA NA --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -133.8836 cross-sectional data total number of observations = 140 mean efficiency: 0.5272536 The results still indicate statistical significant technical inefficiency: > lrtest( prodCDSfaTn ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 5 -137.61 2 7 -133.88 2 7.45 0.0092 ** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 However, neither the t-test presented in the summary output above nor a likelihood ratio test can reject the model with the half-normal distribution of the inefficiency term (i.e., with µ= 0) in favour of the model with truncated normal distribution of the inefficiency term: > lrtest( prodCDSfa, prodCDSfaTn ) Likelihood ratio test 261
6 Stochastic Frontier Analysis Model 1: prodCDSfa Model 2: prodCDSfaTn #Df LogLik Df Chisq Pr(>Chisq) 1 6 -133.89 2 7 -133.88 1 0.0113 0.9153 Given the high P-values of these two tests, it seems to be reasonable to use the model with the half-normal distribution of the error term. 6.1.4 Translog production frontier 6.1.4.1 Estimation As the Cobb-Douglas functional form is very restrictive, we additionally estimate a Translog stochastic production frontier: > prodTLSfa <- sfa( log( qOut ) ~ log( qCap ) + log( qLab ) + log( qMat ) + + I( 0.5 * log( qCap )^2 ) + I( 0.5 * log( qLab )^2 ) + + I( 0.5 * log( qMat )^2 ) + I( log( qCap ) * log( qLab ) ) + + I( log( qCap ) * log( qMat ) ) + I( log( qLab ) * log( qMat ) ), + data = dat ) > summary( prodTLSfa, extraPar = TRUE ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged Iterative ML estimation terminated after 23 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) -8.8110424 19.9181626 -0.4424 0.6582271 log(qCap) -0.6332521 2.0855273 -0.3036 0.7614012 log(qLab) 4.4511064 4.4552358 0.9991 0.3177593 log(qMat) -1.3976309 3.8097808 -0.3669 0.7137284 I(0.5 * log(qCap)^2) 0.0053258 0.1866174 0.0285 0.9772324 I(0.5 * log(qLab)^2) -1.5030433 0.6812813 -2.2062 0.0273700 * I(0.5 * log(qMat)^2) -0.5113559 0.3733348 -1.3697 0.1707812 I(log(qCap) * log(qLab)) 0.4187529 0.2747251 1.5243 0.1274434 I(log(qCap) * log(qMat)) -0.4371561 0.1902856 -2.2974 0.0215978 * I(log(qLab) * log(qMat)) 0.9800294 0.4216637 2.3242 0.0201150 * 262
6 Stochastic Frontier Analysis sigmaSq 0.9587307 0.1968009 4.8716 1.107e-06 *** gamma 0.9153387 0.0647478 14.1370 < 2.2e-16 *** sigmaSqU 0.8775633 0.2328364 3.7690 0.0001639 *** sigmaSqV 0.0811674 0.0497448 1.6317 0.1027476 sigma 0.9791480 0.1004960 9.7432 < 2.2e-16 *** sigmaU 0.9367835 0.1242744 7.5380 4.771e-14 *** sigmaV 0.2848989 0.0873025 3.2634 0.0011010 ** lambdaSq 10.8117751 9.0334818 1.1969 0.2313628 lambda 3.2881264 1.3736518 2.3937 0.0166789 * varU 0.3188892 NA NA NA sdU 0.5647027 NA NA NA gammaVar 0.7971103 NA NA NA --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -128.0684 cross-sectional data total number of observations = 140 mean efficiency: 0.5379939 6.1.4.2 Statistical test for inefficiencies A likelihood ratio test confirms that the stochastic frontier model fits the data much better than an average production function estimated by OLS: > lrtest( prodTLSfa ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 11 -131.25 2 12 -128.07 1 6.353 0.005859 ** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 6.1.4.3 Testing against the Cobb-Douglas functional form A further likelihood ratio test indicates that it is not really clear whether the Translog stochastic frontier model fits the data significantly better than the Cobb-Douglas stochastic frontier model: 263
6 Stochastic Frontier Analysis > lrtest( prodCDSfa, prodTLSfa ) Likelihood ratio test Model 1: prodCDSfa Model 2: prodTLSfa #Df LogLik Df Chisq Pr(>Chisq) 1 6 -133.89 2 12 -128.07 6 11.642 0.07045 . --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 While the Cobb-Douglas functional form cannot be rejected at the 5% significance level, it is rejected in favor of the Translog functional form at the 10% significance level. 6.1.4.4 Obtaining technical efficiency estimates The efficiency estimates based on the Translog stochastic production frontier can be obtained (again) by the efficiencies method: > dat$effTL <- efficiencies( prodTLSfa ) The following commands illustrate their variation, their correlation with the output level, and their correlation with the firm size (measured as input use): > hist( dat$effTL, 15 ) > plot( dat$qOut, dat$effTL, log = "x" ) > plot( dat$X, dat$effTL, log = "x" ) effTL Frequency 0.2 0.4 0.6 0.8 0 2 4 6 8 10 12 1e+05 5e+05 5e+06 0.2 0.4 0.6 0.8 qOut effTL 0.5 1.0 2.0 5.0 0.2 0.4 0.6 0.8 X effTL Figure 6.6: Efficiency estimates of Translog production frontier The resulting graphs are shown in figure 6.6. These efficiency estimates are rather similar to the efficiency estimates based on the Cobb-Douglas stochastic production frontier. This is confirmed by a direct comparison of these efficiency estimates: 264
6 Stochastic Frontier Analysis > compPlot( dat$effCD, dat$effTL ) 0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 effCD effTL Figure 6.7: Efficiency estimates of Cobb-Douglas and Translog production frontier The resulting graph is shown in figure 6.7. Most efficiency estimates only slightly differ between the two functional forms but a few efficiency estimates are considerably higher for the Translog functional form. The inflexibility of the Cobb-Douglas functional form probably resulted in an insufficient adaptation of the frontier to some observations, which lead to larger negative residuals and hence, lower efficiency estimates in the Cobb-Douglas model. 6.1.4.5 Truncated normal distribution of the inefficiency term As explained in section 6.1.3.5, we can estimate a stochastic frontier model based on the assumption that the inefficiency term ufollows a truncated normal distribution instead of a half-normal distribution: > prodTLSfaTn <- sfa( log( qOut ) ~ log( qCap ) + log( qLab ) + log( qMat ) + + I( 0.5 * log( qCap )^2 ) + I( 0.5 * log( qLab )^2 ) + + I( 0.5 * log( qMat )^2 ) + I( log( qCap ) * log( qLab ) ) + + I( log( qCap ) * log( qMat ) ) + I( log( qLab ) * log( qMat ) ), + data = dat, truncNorm = TRUE ) > summary( prodTLSfaTn, extraPar = TRUE ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged 265
6 Stochastic Frontier Analysis (Intercept) 6.74924420 0.74118625 9.1060 < 2.2e-16 *** log(pCap/pMat) 0.07241390 0.04565845 1.5860 0.1127 log(pLab/pMat) 0.44642025 0.07955197 5.6117 2.004e-08 *** log(qOut) 0.37415329 0.03009634 12.4319 < 2.2e-16 *** sigmaSq 0.11117595 0.01442194 7.7088 1.270e-14 *** gamma 0.00018685 0.06501769 0.0029 0.9977 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -44.87812 cross-sectional data total number of observations = 140 mean efficiency: 0.9963738 The parameter γ, which indicates the proportion of the total residual variance that is caused by inefficiency is close to zero and a t-test suggests that it is statistically not significantly different from zero. As the t-test for the parameter γis not always reliable, we use a likelihood ratio test to verify this result: > lrtest( costCDHomSfa ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 5 -44.878 2 6 -44.878 1 0 0.4991 This test confirms that the fit of the OLS model (which assumes that γis zero and hence, that there is no inefficiency) is not significantly worse than the fit of the stochastic frontier model. In fact, the cost efficiency estimates are all very close to one. By default, the efficiencies() method calculates the efficiency estimates as E[e−u], which means that we obtain estimates of Farrell-type cost efficiencies (6.17). Given that E[eu]is not equal to 1/E [e−u](as the expectation operator is an additive operator), we cannot obtain estimates of Shepard-type cost efficiencies (6.16) by taking the inverse of the estimates of the Farrell-type cost efficiencies (6.17). However, we can obtain estimates of Shepard-type cost efficiencies (6.16) by setting argument minusU of the efficiencies() method equal to FALSE, which tells the efficiencies() method to calculate the efficiency estimates as E[eu]. 272
6 Stochastic Frontier Analysis > dat$costEffCDHomFarrell <- efficiencies( costCDHomSfa ) > dat$costEffCDHomShepard <- efficiencies( costCDHomSfa, minusU = FALSE ) > hist( dat$costEffCDHomFarrell, 15 ) > hist( dat$costEffCDHomShepard, 15 ) costEffCDHomFarrell Frequency 0.99632 0.99636 0.99640 0 4 8 12 costEffCDHomShepard Frequency 1.00360 1.00364 1.00368 0 4 8 12 Figure 6.9: Efficiency estimates of Cobb-Douglas cost frontier The resulting graphs are shown in figure 6.9. While the Farrell-type cost efficiencies are all slightly below one, the Shepard-type cost efficiencies are all slightly above one. Both graphs show that we do not find any relevant cost inefficiencies, although we have found considerable technical inefficiencies. 6.2.4 Short-run cost frontiers > hist( residuals( costCDSRHom ) ) residuals costCDSRHom Frequency −0.5 0.0 0.5 0 10 20 30 Figure 6.10: Residuals of Cobb-Douglas short-run cost function The resulting graph is shown in figure 6.10. > costCDSRHomSfa <- sfa( log( vCost / pMat ) ~ log( pLab / pMat ) + + log( qCap ) + log( qOut ), data = dat, ineffDecrease = FALSE ) > summary( costCDSRHomSfa ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency increases the endogenous variable (as in a cost function) The dependent variable is logged 273
6 Stochastic Frontier Analysis iteration failed final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 5.62206 1.00000 5.6221 1.887e-08 *** log(pLab/pMat) 0.53487 1.00000 0.5349 0.5927 log(qCap) 0.18774 1.00000 0.1877 0.8511 log(qOut) 0.29010 1.00000 0.2901 0.7717 sigmaSq 0.10122 1.00000 0.1012 0.9194 gamma 0.05000 1.00000 0.0500 0.9601 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -36.05615 cross-sectional data total number of observations = 140 mean efficiency: 0.9456755 > costCDSRHomSfa <- sfa( log( vCost / pMat ) ~ log( pLab / pMat ) + + log( qCap ) + log( qOut ), data = dat, ineffDecrease = FALSE, + startVal = c( coef( costCDSRHom ), 0.1, 0.01 ) ) > summary( costCDSRHomSfa ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency increases the endogenous variable (as in a cost function) The dependent variable is logged Iterative ML estimation terminated after 45 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 5.6711615 0.3617424 15.6773 < 2.2e-16 *** log(pLab/pMat) 0.5348738 0.0637620 8.3886 < 2.2e-16 *** log(qCap) 0.1877450 0.0367325 5.1111 3.202e-07 *** log(qOut) 0.2901018 0.0308632 9.3996 < 2.2e-16 *** sigmaSq 0.0980563 0.0150503 6.5152 7.258e-11 *** gamma 0.0009243 0.1326111 0.0070 0.9944 --- 274
6 Stochastic Frontier Analysis Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -36.05537 cross-sectional data total number of observations = 140 mean efficiency: 0.9924491 > dat$costEffCDSRHom <- efficiencies( costCDSRHomSfa ) > hist( dat$costEffCDSRHom, 15 ) costEffCDSRHom Frequency 0.9922 0.9923 0.9924 0.9925 0.9926 0.9927 0 5 15 25 Figure 6.11: Efficiency estimates of Cobb-Douglas short-run cost frontier The resulting graphs are shown in figure 6.11. 6.2.5 Profit frontiers ln π= ln π(p, w)−u+vwith u≥0,(6.18) where −u≤0accounts for profit inefficiency and vaccounts for statistical noise. This model can be re-written as: π=π(p, w)e−uev(6.19) Profit efficiency according to Farrell: PE =π π(w, y)ev=π(w, y)e−uev π(w, y)ev=e−u(6.20) 6.3 Analyzing the effects of zvariables In many empirical cases, the output quantity does not only depend on the input quantities but also on some other variables, e.g. the manager’s education and experience and in agricultural production also the soil quality and rainfall. If these factors influence the production process, they must be included in applied production analyses in order to avoid an omitted-variables bias. Our data set on French apple producers includes the variable adv, which is a dummy variable 275
6 Stochastic Frontier Analysis and indicates whether the apple producer uses an advisory service. In the following, we will apply different methods to figure out whether the production process differs between users and non-users of an advisory service. 6.3.1 Production functions with zvariables Additional factors that influence the production process (z) can be included as additional explanatory variables in the production function: y=f(x, z).(6.21) This function can be used to analyze how the additional explanatory variables (z) affect the output quantity for given input quantities, i.e. how they affect the productivity. In case of a Cobb-Douglas functional form, we get following extended production function: ln y=α0+X i αiln xi+αzz(6.22) Based on this Cobb-Douglas production function and our data set on French apple producers, we can check whether the apple producers who use an advisory service produce a different output quantity than non-users with the same input quantities, i.e. whether the productivity differs between users and non-users. This extended production function can be estimated by following command: > prodCDAdv <- lm( log( qOut ) ~ log( qCap ) + log( qLab ) + log( qMat ) + adv, + data = dat ) > summary( prodCDAdv ) Call: lm(formula = log(qOut) ~ log(qCap) + log(qLab) + log(qMat) + adv, data = dat) Residuals: Min 1Q Median 3Q Max -1.7807 -0.3821 0.0022 0.4709 1.3323 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) -2.33371 1.29590 -1.801 0.0740 . log(qCap) 0.15673 0.08581 1.826 0.0700 . log(qLab) 0.69225 0.15190 4.557 1.15e-05 *** log(qMat) 0.62814 0.12379 5.074 1.26e-06 *** 276
6 Stochastic Frontier Analysis adv 0.25896 0.10932 2.369 0.0193 * --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 0.6452 on 135 degrees of freedom Multiple R-squared: 0.6105, Adjusted R-squared: 0.599 F-statistic: 52.9 on 4 and 135 DF, p-value: < 2.2e-16 The estimation result shows that users of an advisory service produce significantly more than non-users with the same input quantities. Given the Cobb-Douglas production function (6.22), the coefficient of an additional explanatory variable can be interpreted as the marginal effect on the relative change of the output quantity: αz=∂ln y ∂z =∂ln y ∂y ∂y ∂z =∂y ∂z 1 y(6.23) Hence, our estimation result indicates that users of an advisory service produce approximately 25.9% more output than non-users with the same input quantity but the large standard error of this coefficient indicates that this estimate is rather imprecise. Given that the change of a dummy variable from zero to one is not marginal and that the coefficient of the variable adv is not close to zero, the above interpretation of this coefficient is a rather poor approximation. In fact, our estimation results suggest that the output quantity of apple producers with advisory service is on average exp(αz)= 1.296 times as large as (29.6% larger than) the output quantity of apple producers without advisory service given the same input quantities. As users and non-users of an advisory service probably differ in some unobserved variables that affect the productivity (e.g. motivation and effort to increase productivity), the coefficient azis not necessarily the causal effect of the advisory service but describes the difference in productivity between users and non-users of the advisory service. 6.3.2 Production frontiers with zvariables A production function that includes additional factors that influence the production process (6.21) can also be estimated as a stochastic production frontier. In this specification, it is assumed that the additional explanatory variables influence the production frontier. The following command estimates the extended Cobb-Douglas production function (6.22) using the stochastic frontier method: > prodCDAdvSfa <- sfa( log( qOut ) ~ log( qCap ) + log( qLab ) + log( qMat ) + adv, + data = dat ) > summary( prodCDAdvSfa ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) 277
6 Stochastic Frontier Analysis The dependent variable is logged Iterative ML estimation terminated after 14 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) -0.247751 1.357917 -0.1824 0.8552301 log(qCap) 0.156906 0.081337 1.9291 0.0537222 . log(qLab) 0.695977 0.148793 4.6775 2.904e-06 *** log(qMat) 0.491840 0.139348 3.5296 0.0004162 *** adv 0.150742 0.111233 1.3552 0.1753583 sigmaSq 0.916031 0.231604 3.9552 7.648e-05 *** gamma 0.861029 0.114087 7.5471 4.450e-14 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -132.8679 cross-sectional data total number of observations = 140 mean efficiency: 0.5545099 The estimation result still indicates that users of an advisory service have a higher productivity than non users, but the coefficient is smaller and no longer statistically significant. The result of the t-test is confirmed by a likelihood-ratio test: > lrtest( prodCDSfa, prodCDAdvSfa ) Likelihood ratio test Model 1: prodCDSfa Model 2: prodCDAdvSfa #Df LogLik Df Chisq Pr(>Chisq) 1 6 -133.89 2 7 -132.87 1 2.0428 0.1529 The model with advisory service as additional explanatory variable indicates that there are significant inefficiencies (at 5% significance level): > lrtest( prodCDAdvSfa ) 278
6 Stochastic Frontier Analysis Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 6 -134.76 2 7 -132.87 1 3.78 0.02593 * --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 The following commands compute the technical efficiency estimates and compare them to the efficiency estimates obtained from the Cobb-Douglas production frontier without advisory service as an explanatory variable: > dat$effCDAdv <- efficiencies( prodCDAdvSfa ) > compPlot( dat$effCD[ dat$adv == 0 ], + dat$effCDAdv[ dat$adv == 0 ] ) > points( dat$effCD[ dat$adv == 1 ], + dat$effCDAdv[ dat$adv == 1 ], pch = 20 ) 0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 Production frontier without advisory service Production frontier with advisory service Figure 6.12: Technical efficiency estimates of Cobb-Douglas production frontier with and without advisory service as additional explanatory variable (circles = producers who do not use an advisory service, solid dots = producers who use an advisory service The resulting graph is shown in figure 6.12. It appears as if the non-users of an advisory service became somewhat more efficient. This is because the stochastic frontier model that includes 279
6 Stochastic Frontier Analysis the advisory service as an explanatory variable has in fact two production frontiers: a lower frontier for the non-users of an advisory service and a higher frontier for the users of an advisory service. The coefficient of the dummy variable adv, i.e. αadv, can be interpreted as a quick estimate of the difference between the two frontier functions. In our empirical case, the difference is approximately 15.1%. However, a precise calculation indicates that the frontier of the users of the advisory service is exp (αadv)=1.163 times (16.3% higher than) the frontier of the non-users of advisory service. And the frontier of the non-users of the advisory service is exp (−αadv) = 0.86 times (14% lower than) the frontier of the users of advisory service. As the non-users of an advisory service are compared to a lower frontier now, they appear to be more efficient now. While it is reasonable to have different frontier functions for different soil types, it does not seem to be too reasonable to have different frontier functions for users and non-users of an advisory service, because there is no physical reasons, why users of an advisory service should have a maximum output quantity that is different from the maximum output quantity of non-users. 6.3.3 Efficiency effects production frontiers As explained above, it does not seem to be too reasonable to have different frontier functions for users and non-users of an advisory service. However, it seems to be reasonable to assume that users of an advisory service have on average different efficiencies than non-users. A model that can account for this has been proposed by Battese and Coelli (1995). In this stochastic frontier model, the efficiency level might be affected by additional explanatory variables: The inefficiency term ufollows a positive truncated normal distribution with constant scale parameter σ2 uand a location parameter µthat depends on additional explanatory variables: u∼N+(µ, σ2 u)with µ=δ z, (6.24) where δis an additional parameter (vector) to be estimated. Function sfa can also estimate these “efficiency effects frontiers”. The additional variables that should explain the efficiency level must be specified at the end of the model formula, where a vertical bar separates them from the (regular) input variables: > prodCDSfaAdvInt <- sfa( log( qOut ) ~ log( qCap ) + log( qLab ) + log( qMat ) | + adv, data = dat ) > summary( prodCDSfaAdvInt ) Efficiency Effects Frontier (see Battese & Coelli 1995) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged Iterative ML estimation terminated after 19 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit 280
6 Stochastic Frontier Analysis final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) -0.090700 1.235454 -0.0734 0.941476 log(qCap) 0.168623 0.081284 2.0745 0.038034 * log(qLab) 0.653860 0.146054 4.4768 7.576e-06 *** log(qMat) 0.513533 0.132236 3.8835 0.000103 *** Z_(Intercept) -0.016812 1.255298 -0.0134 0.989314 Z_adv -1.077590 1.053765 -1.0226 0.306492 sigmaSq 1.096521 0.789600 1.3887 0.164922 gamma 0.863095 0.099424 8.6809 < 2.2e-16 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -130.516 cross-sectional data total number of observations = 140 mean efficiency: 0.6004358 One can use the lrtest() method to test the statistical significance of the entire inefficiency model, i.e. the null hypothesis is H0:γ= 0 and δj= 0 ∀j: > lrtest( prodCDSfaAdvInt ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Efficiency Effects Frontier (EEF) #Df LogLik Df Chisq Pr(>Chisq) 1 5 -137.61 2 8 -130.52 3 14.185 0.001123 ** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 The test indicates that the fit of this model is significantly better than the fit of the OLS model (without advisory service as explanatory variable). The coefficient of the advisory service in the inefficiency model is negative but statistically insignificant. By default, an intercept is added to the inefficiency model but it is completely statistically insignificant. In many econometric estimations of the efficiency effects frontier model, the intercept of the inefficiency model (δ0) is only weakly identified, because the values of δ0can 281
7 Data Envelopment Analysis (DEA) 7.1 Preparations We load the R package “Benchmarking” in order to use it for Data Envelopment Analysis: > library( "Benchmarking" ) We create a matrix of input quantities and a vector of output quantities: > xMat <- cbind( dat$qCap, dat$qLab, dat$qMat ) > yVec <- dat$qOut 7.2 DEA with input-oriented efficiencies The following command conducts an input-oriented DEA with VRS: > deaVrsIn <- dea( xMat, yVec ) > hist( eff( deaVrsIn ) ) Display the “peers” of the first 14 observations: > peers( deaVrsIn )[ 1:14, ] peer1 peer2 peer3 peer4 [1,] 44 73 80 135 [2,] 80 100 126 NA [3,] 44 54 73 100 [4,] 4 NA NA NA [5,] 17 54 81 NA [6,] 41 73 126 132 [7,] 7 NA NA NA [8,] 44 54 80 83 [9,] 100 126 132 NA [10,] 38 73 80 135 [11,] 54 81 100 NA [12,] 44 54 81 100 [13,] 38 73 80 135 [14,] 44 54 81 100 288
7 Data Envelopment Analysis (DEA) Display the λs of the first 14 observations: > lambda( deaVrsIn )[ 1:14, ] L4 L7 L17 L19 L38 L41 L44 L54 L61 L64 [1,] 0 0 0.00000000 0 0.00000000 0.0000000 0.08707089 0.00000000 0 0 [2,] 0 0 0.00000000 0 0.00000000 0.0000000 0.00000000 0.00000000 0 0 [3,] 0 0 0.00000000 0 0.00000000 0.0000000 0.05466873 0.34157362 0 0 [4,] 1 0 0.00000000 0 0.00000000 0.0000000 0.00000000 0.00000000 0 0 [5,] 0 0 0.07874218 0 0.00000000 0.0000000 0.00000000 0.62716635 0 0 [6,] 0 0 0.00000000 0 0.00000000 0.9520817 0.00000000 0.00000000 0 0 [7,] 0 1 0.00000000 0 0.00000000 0.0000000 0.00000000 0.00000000 0 0 [8,] 0 0 0.00000000 0 0.00000000 0.0000000 0.39228600 0.34818591 0 0 [9,] 0 0 0.00000000 0 0.00000000 0.0000000 0.00000000 0.00000000 0 0 [10,] 0 0 0.00000000 0 0.06541405 0.0000000 0.00000000 0.00000000 0 0 [11,] 0 0 0.00000000 0 0.00000000 0.0000000 0.00000000 0.52820862 0 0 [12,] 0 0 0.00000000 0 0.00000000 0.0000000 0.44076458 0.09749327 0 0 [13,] 0 0 0.00000000 0 0.01725343 0.0000000 0.00000000 0.00000000 0 0 [14,] 0 0 0.00000000 0 0.00000000 0.0000000 0.35937585 0.44329381 0 0 L73 L74 L78 L80 L81 L83 L100 L103 [1,] 0.243735897 0 0 0.6423537 0.00000000 0.00000000 0.0000000 0 [2,] 0.000000000 0 0 0.5147430 0.00000000 0.00000000 0.3620871 0 [3,] 0.153372277 0 0 0.0000000 0.00000000 0.00000000 0.4503854 0 [4,] 0.000000000 0 0 0.0000000 0.00000000 0.00000000 0.0000000 0 [5,] 0.000000000 0 0 0.0000000 0.29409147 0.00000000 0.0000000 0 [6,] 0.002769034 0 0 0.0000000 0.00000000 0.00000000 0.0000000 0 [7,] 0.000000000 0 0 0.0000000 0.00000000 0.00000000 0.0000000 0 [8,] 0.000000000 0 0 0.2101886 0.00000000 0.04933947 0.0000000 0 [9,] 0.000000000 0 0 0.0000000 0.00000000 0.00000000 0.6917918 0 [10,] 0.068686498 0 0 0.2825911 0.00000000 0.00000000 0.0000000 0 [11,] 0.000000000 0 0 0.0000000 0.25455055 0.00000000 0.2172408 0 [12,] 0.000000000 0 0 0.0000000 0.29388540 0.00000000 0.1678567 0 [13,] 0.383969646 0 0 0.5669254 0.00000000 0.00000000 0.0000000 0 [14,] 0.000000000 0 0 0.0000000 0.04033289 0.00000000 0.1569974 0 L126 L129 L132 L135 L137 [1,] 0.000000000 0 0.00000000 0.02683954 0 [2,] 0.123169836 0 0.00000000 0.00000000 0 [3,] 0.000000000 0 0.00000000 0.00000000 0 [4,] 0.000000000 0 0.00000000 0.00000000 0 [5,] 0.000000000 0 0.00000000 0.00000000 0 [6,] 0.008468157 0 0.03668108 0.00000000 0 289
7 Data Envelopment Analysis (DEA) [7,] 0.000000000 0 0.00000000 0.00000000 0 [8,] 0.000000000 0 0.00000000 0.00000000 0 [9,] 0.249102366 0 0.05910586 0.00000000 0 [10,] 0.000000000 0 0.00000000 0.58330837 0 [11,] 0.000000000 0 0.00000000 0.00000000 0 [12,] 0.000000000 0 0.00000000 0.00000000 0 [13,] 0.000000000 0 0.00000000 0.03185153 0 [14,] 0.000000000 0 0.00000000 0.00000000 0 The following commands display the “slack” of the first 14 observations in an input-oriented DEA with VRS: > deaVrsIn <- dea( xMat, yVec, SLACK = TRUE ) > table( deaVrsIn$slack ) FALSE TRUE 78 62 > deaVrsIn$sx[ 1:14, ] sx1 sx2 sx3 [1,] 0 0 0.00000 [2,] 0 0 345.70719 [3,] 0 0 0.00000 [4,] 0 0 0.00000 [5,] 0 0 38.54949 [6,] 0 0 0.00000 [7,] 0 0 0.00000 [8,] 0 0 0.00000 [9,] 0 0 1624.33417 [10,] 0 0 0.00000 [11,] 0 0 12993.07250 [12,] 0 0 0.00000 [13,] 0 0 0.00000 [14,] 0 0 0.00000 > deaVrsIn$sy[ 1:14, ] [1]00000000000000 The following command conducts an input-oriented DEA with CRS: > deaCrsIn <- dea( xMat, yVec, RTS = "crs" ) > hist( eff( deaCrsIn ) ) 290
7 Data Envelopment Analysis (DEA) We can calculate the scale efficiencies by: > se <- eff( deaCrsIn ) / eff( deaVrsIn ) > hist( se ) The following command conducts an input-oriented DEA with DRS > deaDrsIn <- dea( xMat, yVec, RTS = "drs" ) > hist( eff( deaDrsIn ) ) And we check if firms are too small or too large. This is the number of observations that produce at the scale below the optimal scale size: > sum( eff( deaVrsIn ) - eff( deaDrsIn ) > 1e-4 ) [1] 117 7.3 DEA with output-oriented efficiencies The following command conducts an output-oriented DEA with VRS: > deaVrsOut <- dea( xMat, yVec, ORIENTATION = "out" ) > hist( efficiencies( deaVrsOut ) ) The following command conducts an output-oriented DEA with CRS: > deaCrsOut <- dea( xMat, yVec, RTS = "crs", ORIENTATION = "out" ) > hist( eff( deaCrsOut ) ) In case of CRS, input-oriented efficiencies are equivalent to output-oriented efficiencies: > all.equal( eff( deaCrsIn ), 1 / eff( deaCrsOut ) ) [1] TRUE 7.4 DEA with “super efficiencies” The following command obtains “super efficiencies” for an input-oriented DEA with CRS: > sdeaVrsIn <- sdea( xMat, yVec ) > hist( eff( sdeaVrsIn ) ) 7.5 DEA with graph hyperbolic efficiencies The following command conducts a DEA with graph hyperbolic efficiencies and VRS: > deaVrsGraph <- dea( xMat, yVec, ORIENTATION = "graph" ) > hist( eff( deaVrsGraph ) ) > plot( eff( deaVrsIn ), eff( deaVrsGraph ) ) > abline(0,1) 291
8 Distance Functions 8.1 Theory 8.1.1 Output distance functions The Shepard output distance function is defined as: Do(x, y) = min{λ > 0|y/λ ∈P(x)}(8.1) or (equivalently) as: Do(x, y) = min{λ > 0|(x, y/λ)∈T},(8.2) where xis a vector of input quantities, yis a vector of output quantities, P(x)is the production possibility set, and Tis the technology set. The Shepard output distance function, Do(x, y), defined in (8.1) and (8.2) returns the Shepard output-oriented technical efficiencies defined in (5.1) and (5.5). Thus, it returns a value of one for fully efficient sets of inputs and outputs (x, y), whereas it returns a non-negative value smaller than one for inefficient sets of inputs and outputs (x, y).1 8.1.1.1 Properties It is usually assumed that the Shepard output distance function Do(x, y)given in (8.1) fulfills the following properties (see, e.g., Färe and Primont,1995;Coelli et al.,2005): 1. Do(x, 0) = 0 for all non-negative x 2. Do(x, y)is non-increasing in x, i.e. ∂Do(x, y)/∂xi≤0∀i= 1, . . . , N 3. Do(x, y)is non-decreasing in y, i.e. ∂Do(x, y)/∂yi≥0∀i= 1, . . . , M 4. Do(x, y)is linearly homogeneous in y, i.e. Do(x, k y) = k Do(x, y)∀k > 0 5. Do(x, y)is quasiconvex in x⇒convex isoquants2 1The Farrell output distance function is defined as the inverse of the Shepard output distance function: DoF (x, y) = max{λ > 0|λ y ∈P(x)}= 1/Do(x, y). It returns the Farrell output-oriented technical efficiencies defined in (5.2) and (5.6). 2O’Donnell and Coelli (2005) clarified a typographical error in Färe and Primont (1995, p. 152) who state by mistake that the output distance function is quasiconcave in input quantities x. The quasiconcavity of the production function is equivalent to the quasiconvexity of the output distance function, whereas the ‘conversion’ between quasiconcavity and quasiconvexity is caused by the opposite ‘direction’ regarding the input quantities 292
8 Distance Functions 6. Do(x, y)is convex in y⇒concave transformation curves (i.e., joint production is advantageous compared to specialized production for a given set of input quantities) 7. Do(x, y)≤1indicates that ybelongs to the production possibility set of x, i.e. y∈P(x), while Do(x, y)>1indicates that ydoes not belong to the production possibility set of x. 8. Do(x, y)=1indicates that yis at the boundary of the production possibility set P(x), i.e. ylies on the transformation curve (given x) and xlies on the isoquant (given y). We will illustrate these properties using a very simple output distance function with two inputs x= (x1, x2)′and two outputs y= (y1, y2)′:3 Do(x, y) = qy2 1+y2 2x−0.4 1x−0.4 2(8.3) In the following, we illustrate that this simple output distance function fulfills the abovementioned properties 1–4: •no output Do(x, 0) = p02+ 02x−0.4 1x−0.4 2= 0 (8.4) •monotonically non-increasing in x ∂Do(x, y) ∂xi =−0.4qy2 1+y2 2x−0.4 1x−0.4 2 xi≤0; i∈ {1,2}(8.5) •monotonically non-decreasing in y ∂Do(x, y) ∂yi =x−0.4 1x−0.4 2yi qy2 1+y2 2≥0; i∈ {1,2}(8.6) •linear homogeneous in y Do(x, ky) = q(ky1)2+ (ky2)2x−0.4 1x−0.4 2(8.7) =qk2y2 1+k2y2 2x−0.4 1x−0.4 2(8.8) =qk2y2 1+y2 2x−0.4 1x−0.4 2(8.9) =√k2qy2 1+y2 2x−0.4 1x−0.4 2(8.10) =kqy2 1+y2 2x−0.4 1x−0.4 2(8.11) =kDo(x, y)(8.12) (i.e., the production function is increasing in input quantities, while the output distance function is decreasing in input quantities). 3This output distance function is a very much simplified version of the multiple-output “ray frontier production function” suggested by Löthgren (1997,2000). 293
8 Distance Functions •quasiconvexity in x A sufficient condition for quasiconvexity is that all leading principal minors of the bordered Hessian matrix are strictly negative (see section 1.5.6). The first derivatives of the output distance function (8.3) with respect to the input quantities are given in (8.5). The remaining elements of the bordered Hessian matrix, i.e. the second derivatives with respect to the input quantities, are: ∂2Do(x, y) ∂xi∂xj =0.4(∆ij + 0.4)qy2 1+y2 2x−0.4 1x−0.4 2 xixj ;i∈ {1,2}(8.13) ∆ij = 1if i=j 0if i=j (8.14) The first leading principal minor of the bordered Hessian Bis: |B1|=−∂Do(x, y) ∂x12 (8.15) =−−0.4qy2 1+y2 2x−1.4 1x−0.4 22 (8.16) =−0.42y2 1+y2 2x−2.8 1x−0.8 2(8.17) <0∀x1, x2>0; y1, y2≥0; y1+y2>0(8.18) The determinant of the (3 ×3) bordered Hessian is: |B|=2 ∂Do(x, y) ∂x1 ∂Do(x, y) ∂x2 ∂2Do(x, y) ∂x1∂x2 (8.19) −∂Do(x, y) ∂x12∂2Do(x, y) ∂x2 2−∂Do(x, y) ∂x22∂2Do(x, y) ∂x2 1 = 2 −0.4qy2 1+y2 2x−1.4 1x−0.4 2−0.4qy2 1+y2 2x−0.4 1x−1.4 2(8.20) 0.42qy2 1+y2 2x−1.4 1x−1.4 2 −−0.4qy2 1+y2 2x−1.4 1x−0.4 220.4·1.4qy2 1+y2 2x−0.4 1x−2.4 2 −−0.4qy2 1+y2 2x−0.4 1x−1.4 220.4·1.4qy2 1+y2 2x−2.4 1x−0.4 2 =0.4qy2 1+y2 2x−0.4 1x−0.4 23 (8.21) 2−x−1 1−x−1 20.4x−1 1x−1 2−−x−1 121.4x−2 2−−x−1 221.4x−2 1 =0.4qy2 1+y2 2x−0.4 1x−0.4 23h0.8x−2 1x−2 2−1.4x−2 1x−2 2−1.4x−2 1x−2 2i(8.22) 294
8 Distance Functions =−20.4qy2 1+y2 2x−0.4 1x−0.4 23 x−2 1x−2 2(8.23) =−2·0.43qy2 1+y2 23 x−3.2 1x−3.2 2(8.24) <0∀x1, x2>0; y1, y2≥0; y1+y2>0(8.25) If both of the input quantities are strictly positive and at least one of the output quantities is strictly positive, the first leading principal minor and the determinant of the bordered Hessian matrix are both strictly negative, which means that both the necessary condition and the sufficient condition for quasiconvexity are fulfilled. Figure 8.1 illustrates the relationship between the two input quantities x1and x2and the output distance measure Do(x, y)holding the output quantities y1and y2constant so that qy2 1+y2 2= 1, e.g. y1=y2=√0.5(please note that the origin, i.e. x1=x2= 0, is in the back). The grey part of the surface indicates distance measures that are larger than one (Do(x, y)>1), i.e. combinations of the input quantities x1and x2that are too small to produce the output quantities y1and y2. The dark green line is an isoquant that indicates all combinations of the input quantities x1and x2that a fully efficient firm (i.e. Do(x, y)=1) needs to produce the output quantities y1and y2. Given that the output distance function is quasiconvex in x, the lower contour sets, i.e. the sets of input combinations that give an output distance measure (Do(x, y)) smaller than a certain value (e.g. 1), are convex sets (see section 1.5.6). As the origin (x1=x2= 0) is not inside the lower contour sets, the quasiconvexity of the output distance function in xresults in isoquants that are convex to the origin. The light green part of the surface indicates distance measures (Do(x, y)) that are slightly smaller than one, indicating small technical inefficiency. The red part of the surface indicates distance measures (Do(x, y)) that are considerably smaller than one, indicating considerable technical inefficiency. •convexity in y Convexity requires that the Hessian matrix is positive semidefinite (see section 1.5.5). The elements of the Hessian matrix, i.e. the second derivatives with respect to the output quantities, are: ∂2Do(x, y) ∂yi∂yj =x−0.4 1x−0.4 2 qy2 1+y2 2∆ij −yiyj y2 1+y2 2;i, j ∈ {1,2}(8.26) ∆ij = 1if i=j 0if i=j (8.27) Equation (8.26) shows that all diagonal elements of the Hessian matrix (i.e. i=j) are non-negative, because 1−y2 i/(y2 1+y2 2)is non-negative for i∈ {1,2}. 295
8 Distance Functions x1 0 1 2 x2 0 1 2 Do ( x, y ) 0.5 1.0 Figure 8.1: Output distance function for different input quantities The determinant of the (2 ×2) Hessian matrix is: |H|=∂2Di(x, y) ∂x2 1·∂2Di(x, y) ∂x2 2− ∂2Di(x, y) ∂x1∂x2!2 (8.28) x−0.4 1x−0.4 2 qy2 1+y2 2 1−y2 1 y2 1+y2 2!x−0.4 1x−0.4 2 qy2 1+y2 2 1−y2 2 y2 1+y2 2!− x−0.4 1x−0.4 2 qy2 1+y2 2 y1y2 y2 1+y2 2 2 (8.29) = x−0.4 1x−0.4 2 qy2 1+y2 2 2 1−y2 1 y2 1+y2 2! 1−y2 2 y2 1+y2 2!− x−0.4 1x−0.4 2 qy2 1+y2 2 2y1y2 y2 1+y2 22 (8.30) = x−0.4 1x−0.4 2 qy2 1+y2 2 2" 1−y2 1 y2 1+y2 2! 1−y2 2 y2 1+y2 2!−y1y2 y2 1+y2 22#(8.31) = x−0.4 1x−0.4 2 qy2 1+y2 2 2"y2 2 y2 1+y2 2 y2 1 y2 1+y2 2−y2 1y2 2 y2 1+y2 22#(8.32) = x−0.4 1x−0.4 2 qy2 1+y2 2 2"y2 2y2 1 y2 1+y2 22−y2 1y2 2 y2 1+y2 22#= 0 ≥0(8.33) As all three principal minors (the two diagonal elements and the determinant) of the Hessian matrix are non-negative, we can conclude that the Hessian matrix is positive semidefinite and, thus, the output distance function 8.3 is convex in y. 296
8 Distance Functions It is indeed not surprising that the determinant of the Hessian matrix is exactly zero, because the output distance function is linearly homogeneous in yand the determinants of linearly homogeneous functions are always zero. Figure 8.2 illustrates the relationship between the two output quantities y1and y2and the output distance measure Do(x, y)holding the input quantities x1and x2constant so that x0.4 1x0.4 2= 1, e.g. x1=x2= 1 (please note that the origin, i.e. y1=y2= 0, is in the front). The grey part of the surface indicates distance measures that are larger than one (Do(x, y)>1), i.e. combinations of the output quantities y1and y2that cannot be produced by the input quantities x1and x2. The dark green line is a transformation curve that indicates all combinations of the output quantities y1and y2that a fully efficient firm (i.e. Do(x, y)=1) can produce from the input quantities x1and x2. Given that the output distance function is convex in y, it is also quasiconvex in yso that the lower contour sets, i.e. the sets of output combinations that give an output distance measure (Do(x, y)) smaller than a certain value (e.g. 1), are convex sets (see section 1.5.6). As the origin (y1=y2= 0) is inside the lower contour sets, the convexity of the output distance function in yresults in transformation curves that are concave to the origin. The light green part of the surface indicates distance measures (Do(x, y)) that are slightly smaller than one, indicating small technical inefficiency. The red part of the surface indicates distance measures (Do(x, y)) that are considerably smaller than one, indicating considerable technical inefficiency. y1 0.0 0.5 1.0 1.5 y2 0.0 0.5 1.0 1.5 Do ( x, y ) 0.0 0.5 1.0 1.5 Figure 8.2: Output distance function for different output quantities The following code generates figures 8.1 and 8.2: > nVal <- 100 > colRedGreen <- colorRampPalette( c( "red", "green" ) )( 100 ) 297
8 Distance Functions Equation (8.67) shows that all diagonal elements of the Hessian matrix (i.e. i=j) are non-positive, because (0.5−∆ii)is negative for i∈ {1,2}, while all other terms on the right-hand side of (8.67) are always non-negative. The determinant of the (2 ×2) Hessian matrix is: |H|=∂2Di(x, y) ∂x2 1·∂2Di(x, y) ∂x2 2− ∂2Di(x, y) ∂x1∂x2!2 (8.69) = 0.5 (−0.5) x0.5 1x0.5 2 x2 1qy2 1+y2 2·0.5 (−0.5) x0.5 1x0.5 2 x2 2qy2 1+y2 2− 0.5·0.5x0.5 1x0.5 2 x1x2qy2 1+y2 2 2 (8.70) =0.54 x1x2y2 1+y2 2−0.54 x1x2y2 1+y2 2= 0 ≥0(8.71) As all first-order principal minors (i.e. the two diagonal elements) are non-positive and the second-order principal (i.e. the determinant) is non-negative, we can conclude that the Hessian matrix is negative semidefinite and, thus, the input distance function (8.60) is concave in x. It is indeed not surprising that the determinant of the Hessian matrix is exactly zero, because the input distance function is linearly homogeneous in xand the determinants of linearly homogeneous functions are always zero. Figure 8.3 illustrates the relationship between the two input quantities x1and x2and the input distance measure Di(x, y)holding the output quantities y1and y2constant so that qy2 1+y2 2= 1, e.g. y1=y2=√0.5(please note that the origin, i.e. x1=x2= 0, is in the front). The grey part of the surface indicates distance measures that are smaller than one (Di(x, y)<1), i.e. combinations of the input quantities x1and x2that are too small to produce the output quantities y1and y2. The dark green line is an isoquant that indicates all combinations of the input quantities x1and x2that a fully efficient firm (i.e. Di(x, y)=1) needs to produce the output quantities y1and y2. Given that the input distance function is concave in x, it is also quasiconcave in xso that the upper contour sets, i.e. the sets of input combinations that give an input distance measure (Di(x, y)) larger than a certain value (e.g. 1), are convex sets (see section 1.5.6). As the origin (x1=x2= 0) is not inside the upper contour sets (only on its border if Di(x, y)=0), the concavity of the input distance function in xresults in isoquants that are convex to the origin. The light green part of the surface indicates distance measures (Di(x, y)) that are slightly larger than one, indicating small technical inefficiency. The red part of the surface indicates distance measures (Di(x, y)) that are considerably larger than one, indicating considerable technical inefficiency. •quasiconcavity in y A sufficient condition for quasiconcavity of a function with two arguments is that the first leading principal minor of the bordered Hessian matrix is strictly negative, while its second 304
8 Distance Functions x1 0 2 4 x2 0 2 4 Di ( x, y ) 0 1 2 Figure 8.3: Input distance function for different input quantities leading principle minor is strictly positive (see section 1.5.6). The first derivatives of the input distance function (8.60) with respect to the output quantities are given in (8.62). The remaining elements of the bordered Hessian matrix, i.e. the second derivatives with respect to the output quantities, are: ∂2Di(x, y) ∂yi∂yj =x0.5 1x0.5 2 y2 1+y2 21.53yiyj y2 1+y2 2−∆ij;i∈ {1,2}(8.72) ∆ij = 1if i=j 0if i=j (8.73) The first leading principal minor of the bordered Hessian Bis: |B1|=−∂Do(x, y) ∂x12 (8.74) =− −x0.5 1x0.5 2y1 y2 1+y2 21.5!2 (8.75) =−x1x2y2 1 y2 1+y2 23(8.76) <0∀x1, x2>0; y1>0; y2≥0(8.77) The determinant of the (3 ×3) bordered Hessian is: |B|=2 ∂Di(x, y) ∂y1 ∂Di(x, y) ∂y2 ∂2Di(x, y) ∂y1∂y2 (8.78) 305
8 Distance Functions − ∂Di(x, y) ∂y1!2∂2Di(x, y) ∂y2 2− ∂Di(x, y) ∂y2!2∂2Di(x, y) ∂y2 1 = 2 −x0.5 1x0.5 2y1 y2 1+y2 21.5! −x0.5 1x0.5 2y2 y2 1+y2 21.5! 3x0.5 1x0.5 2y1y2 y2 1+y2 22.5!(8.79) − −x0.5 1x0.5 2y1 y2 1+y2 21.5!2 x0.5 1x0.5 2 y2 1+y2 21.5 3y2 2 y2 1+y2 2−1!! − −x0.5 1x0.5 2y2 y2 1+y2 21.5!2 x0.5 1x0.5 2 y2 1+y2 21.5 3y2 1 y2 1+y2 2−1!! = x0.5 1x0.5 2 y2 1+y2 21.5!3"6y2 1y2 2 y2 1+y2 2−y2 1 3y2 2 y2 1+y2 2−1!−y2 2 3y2 1 y2 1+y2 2−1!# (8.80) = x1.5 1x1.5 2 y2 1+y2 24.5!hy2 1+y2 2i(8.81) =x1.5 1x1.5 2 y2 1+y2 23.5(8.82) >0∀x1, x2>0; y1, y2≥0; y1+y2>0(8.83) If both of the input quantities are strictly positive and at least one of the output quantities is strictly positive. the first leading principal minor of the bordered Hessian matrix is nonpositive and its determinant is non-negative, which means that the necessary conditions for quasiconvexity are fulfilled. If the quantities of both of the inputs and of the first output are strictly positive and the quantity of the second output is non-negative, the first leading principal minor of the bordered Hessian matrix is strictly negative and its determinant is strictly positive, which means that also the sufficient conditions for quasiconvexity are fulfilled. As the order of the outputs is arbitrary, we can rearrange the order of the outputs so that the quantity of the first input is strictly positive as long as at least one output quantity is strictly positive. Hence, if at least one of the output quantities is strictly positive, not only the necessary conditions but also the sufficient conditions for quasiconvexity are fulfilled. Figure 8.4 illustrates the relationship between the two output quantities y1and y2and the input distance measure Di(x, y)holding the input quantities x1and x2constant so that x0.5 1x0.5 2= 1, e.g. x1=x2= 1 (please note that the origin, i.e. y1=y2= 0, is in the back). The grey part of the surface indicates distance measures that are smaller than one (Di(x, y)<1), i.e. combinations of the output quantities y1and y2that cannot be produced by the input quantities x1and x2. The dark green line is a transformation curve that indicates all combinations of the output quantities y1and y2that a fully efficient firm (i.e. Di(x, y) = 1) can produce from the input quantities x1and x2. Given that the input distance function is quasiconcave in y, the upper contour sets, i.e. the sets of output combinations that give an input distance measure (Di(x, y)) larger than a certain value 306
8 Distance Functions (e.g. 1), are convex sets (see section 1.5.6). As the origin (y1=y2= 0) is inside the upper contour sets, the quasiconcavity of the input distance function in yresults in transformation curves that are concave to the origin. The light green part of the surface indicates distance measures (Di(x, y)) that are slightly larger than one, indicating small technical inefficiency. The red part of the surface indicates distance measures (Di(x, y)) that are considerably larger than one, indicating considerable technical inefficiency. y1 0.0 0.5 1.0 1.5 y2 0.0 0.5 1.0 1.5 Do ( x, y ) 0.5 1.0 1.5 2.0 Figure 8.4: Input distance function for different output quantities The following code generates figures 8.3 and 8.4: > nVal <- 100 > colRedGreen <- colorRampPalette( c( "green", "red" ) )( 100 ) > x1 <- seq( 0, 4, length.out = nVal ) > x2 <- seq( 0, 4, length.out = nVal ) > dxMat <- outer( x1, x2, function( a, b ) a^( 0.5 ) * b^( 0.5 ) ) > dxFacet <- ( dxMat[-1, -1] + dxMat[-1, -nVal] + dxMat[-nVal, -1] + + dxMat[-nVal, -nVal] ) / 4 > dxInt <- cut( dxFacet, + seq( 1, 3, length.out = length( colRedGreen ) ) ) > dxCol <- colRedGreen[ dxInt ] > dxCol[ dxFacet < 1 ] <- "grey" > dxMat[ dxMat > 3 ] <- NA > ppx <- persp( x1, x2, dxMat, xlim = c( 0, max( x1 ) ), ylim = c( 0, max( x2 ) ), + zlim = c( min( dxMat, na.rm = TRUE ), max( dxMat, na.rm = TRUE ) ), 307
8 Distance Functions + zlab = "Di ( x, y )", ticktype = "detailed", nticks = 3, + theta = -25, phi = 20, col = dxCol, border = NA ) > xa <- seq( 1/max(x2), max(x1), length.out = nVal ) > lines( trans3d( xa, 1/xa, 1, ppx ), col = "darkgreen", lwd = 5 ) > y1 <- seq( 0, 1.5, length.out = nVal ) > y2 <- seq( 0, 1.5, length.out = nVal ) > dyMat <- outer( y1, y2, function( a, b ) 1 / sqrt( a^2 + b^2 ) ) > dyFacet <- ( dyMat[-1, -1] + dyMat[-1, -nVal] + dyMat[-nVal, -1] + + dyMat[-nVal, -nVal] ) / 4 > dyInt <- cut( dyFacet, + seq( 1, 2, length.out = length( colRedGreen ) ) ) > dyCol <- colRedGreen[ dyInt ] > dyCol[ dyFacet < 1 ] <- "grey" > # dyMat[ dyMat > dyMat[ 1, nVal ] ] <- NA > dyMat[ dyMat > 2 ] <- NA > ppy <- persp( y1, y2, dyMat, zlab = "Do ( x, y )", + ticktype = "detailed", nticks = 3, + theta = 155, phi = 20, col = dyCol, border = NA ) > ya <- seq( 0, 1, length.out = nVal ) > lines( trans3d( ya, sqrt(1-ya^2), 1, ppy ), col = "darkgreen", lwd = 5 ) 8.1.2.2 Distance elasticities A distance elasticity of the input distance function Di(x, y)with respect to an input quantitiy xi indicates the percentage change in the input distance measure Di(x, y)(and, hence, in the firm’s inefficiency) given a one percent increase in this input quantity. The distance elasticity of the ith input quantity is defined as: ϵI i=∂Di(x, y) ∂xi xi Di=∂ln Di(x, y) ∂ln xi≥0(8.84) Linear homogeneity of the input distance function Di(x, y)in input quantities implies that its distance elasticities with respect to the input quantities sum to one: N X i=1 ϵI i= 1 (8.85) A distance elasticity of the input distance function Di(x, y)with respect to an output quantity yiindicates the percentage change in the input distance measure Di(x, y)(and, hence, in the firm’s inefficiency) given a one percent increase in this output quantity. The distance elasticity 308
8 Distance Functions of the ith output quantity is defined as: ϵO i=∂Di(x, y) ∂yi yi Di=∂ln Di(x, y) ∂ln yi≤0(8.86) The linear homogeneity of the input distance function Di(x, y)in input quantities xallows an alternative interpretation of the distance elasticities of the outputs. The interpretation above equation (8.86) assumes that the remaining output quantities and all input quantities remain unchanged, while the input distance measure Di(x, y)and, thus, the technical inefficiency adjusts to the change in one output quantity. Alternatively, we can assume the remaining output quantities and the input distance measure Di(x, y)and, thus, the technical inefficiency remain unchanged, while the input quantities adjust to the change in one output quantity. Under this assumption, if the quantity of the ith output increases by one percent, all input quantities must simultaneously increase by −ϵO i% if the remaining output quantities and the technical efficiency should remain unchanged. This interpretation includes the case, where the technical (in)efficiency remains constant at one so that the change in the input quantities refers to the ‘frontier’, i.e., the minimum required input quantities. 8.1.2.3 Elasticity of scale The elasticity of scale of an input distance function Di(x, y)equals the inverse of the negative sum of distance elasticities with respect to the output quantities: ϵ= − M X i=1 ϵO i!−1 (8.87) 8.2 Cobb-Douglas output distance function 8.2.1 Specification The general form of the Cobb-Douglas output distance function is Do(x, y) = A N Y i=1 xαi i! M Y i=1 yβi i!,(8.88) which can be linearized to: ln Do(x, y) = α0+ N X i=1 αiln xi+ M X i=1 βiln yi(8.89) with α0= ln A. 309
8 Distance Functions 8.2.2 Estimation The Cobb-Douglas output distance function can be expressed (and then estimated) as traditional stochastic frontier model. In the following, we will derive, under which conditions the Cobb-Douglas output distance function defined in (8.88) and (8.89) is linear homogeneous in output quantities (y): kDo(x, y) = Do(x, ky)(8.90) ln(kDo(x, y)) = ln Do(x, ky)(8.91) ln(kDo(x, y)) = α0+ N X i=1 αiln xi+ M X i=1 βiln(k yi)(8.92) ln k+ ln Do(x, y) = α0+ N X i=1 αiln xi+ M X i=1 βiln k+ M X i=1 βiln yi(8.93) ln k+ ln Do(x, y) = ln Do(x, y) + ln k M X i=1 βi(8.94) ln k= ln k M X i=1 βi(8.95) 1 = M X i=1 βi(8.96) Hence, the Cobb-Douglas output distance function is linear homogeneous in output quantities if the coefficients of the output quantities (βi) sum up to one. We can impose linear homogeneity in output quantities by re-arranging equation (8.96) to: βM= 1 − M−1 X i=1 βi(8.97) and by substituting the right-hand side of (8.97) for βMin (8.89): ln Do(x, y) = α0+ N X i=1 αiln xi+ M−1 X i=1 βiln(yi) + 1− M−1 X i=1 βi!ln(yM)(8.98) ln Do(x, y) = α0+ N X i=1 αiln xi+ M−1 X i=1 βi(ln yi−ln yM) + ln yM(8.99) −ln yM=α0+ N X i=1 αiln xi+ M−1 X i=1 βiln(yi/yM)−ln Do(x, y).(8.100) We can assume that u≡ −ln(Do(x, y)) ≥0follows a half normal or truncated normal distribution (i.e. u∼N+(µ, σ2 u)) and add a disturbance term v, that accounts for statistical noise and 310
8 Distance Functions follows a normal distribution (i.e. v∼N(0, σ2 v)) so that we get: −ln(yM) = α0+ N X i=1 αiln xi+ M−1 X i=1 βiln(yi/yM) + u+v. (8.101) This specification is equivalent to the specification of stochastic frontier models so that we can use the stochastic frontier methods introduced in Chapter 6to estimate this output distance function. In the following, we will estimate a Cobb-Douglas output distance function with our data set of French apple producers. This data set distinguishes between two outputs: the quantity of apples produced (qApples) and the quantity of the other outputs (qOtherOut). The following command estimates the Cobb-Douglas output distance function using the command sfa() for stochastic frontier analysis: > odfCD <- sfa( -log(qApples) ~ log( qCap ) + log( qLab ) + log( qMat ) + + log( qOtherOut/qApples), data = dat, ineffDecrease = FALSE ) > summary( odfCD ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency increases the endogenous variable (as in a cost function) The dependent variable is logged Iterative ML estimation terminated after 11 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 10.776633 1.393814 7.7318 1.061e-14 *** log(qCap) -0.087286 0.086890 -1.0046 0.315106 log(qLab) -0.501954 0.157395 -3.1891 0.001427 ** log(qMat) -0.451892 0.116505 -3.8787 0.000105 *** log(qOtherOut/qApples) 0.621158 0.044476 13.9663 < 2.2e-16 *** sigmaSq 1.306200 0.219877 5.9406 2.840e-09 *** gamma 0.923605 0.039930 23.1307 < 2.2e-16 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -147.8225 cross-sectional data total number of observations = 140 311
8 Distance Functions mean efficiency: 0.5069105 8.2.3 Distance elasticities The distance elasticities of a Cobb-Douglas output distance function with respect to input quantities (x), ϵI, can be calculated as: ϵI i=∂Do(x, y) ∂xi xi Do=∂ln Do(x, y) ∂ln xi =αi(8.102) The distance elasticities of a Cobb-Douglas output distance function with respect to output quantities (y), ϵO, can be calculated as: ϵO i=∂Do(x, y) ∂yi yi Do=∂ln Do(x, y) ∂ln yi =βi(8.103) Since the distance elasticities of Cobb-Douglas output distance functions are equal to some of the estimated parameters, they can be directly obtained from the estimation output. The distance elasticity with respect to the capital input (ϵI qCap =−0.087) indicates that a 1% increase in the capital input results (ceteris paribus) in a 0.09% decrease of the distance measure (i.e. efficiency decreases and inefficiency increases). The distance elasticity with respect to the labor input (ϵI qLab =−0.502) indicates that a 1% increase in the labor input results (ceteris paribus) in a 0.5% decrease of the distance measure (i.e. efficiency decreases and inefficiency increases). The distance elasticity with respect to the material input (ϵI qMat =−0.452) indicates that a 1% increase in the materials input results (ceteris paribus) in a 0.45% decrease of the distance measure (i.e. efficiency decreases and inefficiency increases). According to the alternative interpretation of the distance elasticities of the inputs (see section 8.1.1.2), the distance elasticity with respect to the capital input (ϵI qCap =−0.087) indicates that a 1% increase in the capital input increases the maximum feasible quantities of both outputs (simultaneously) by 0.09%. Similarly, the distance elasticity with respect to the labor input (ϵI qLab =−0.502) indicates that a 1% increase in the labor input increases the maximum feasible quantities of both outputs (simultaneously) by 0.5%, while the distance elasticity with respect to the material input (ϵI qMat =−0.452) indicates that a 1% increase in the materials input increases the maximum feasible quantities of both outputs (simultaneously) by 0.45%. The distance elasticity with respect to the other output (ϵO qOtherOut = 0.621) indicates that a 1% increase in other outputs results (ceteris paribus) in a 0.62% increase of the distance measure (i.e. efficiency increases and inefficiency decreases). Based on the homogeneity property (8.97), we can calculate the distance elasticity of the apple quantity ϵO qApples = 1 −ϵO qOtherOut = 0.379, which indicates that a 1% increase in the output of apples results (ceteris paribus) in a 0.38% increase of the distance measure (i.e. efficiency increases and inefficiency decreases). 312
8 Distance Functions 8.2.4 Elasticity of scale The elasticity of scale for an output distance function Do(x, y)equals the negative sum of distance elasticities with respect to the input quantities: ϵ=−X i ϵI i(8.104) For our Cobb-Douglas output distance function, the elasticity of scale can be obtained by using following command: > -sum(coef(odfCD)[c("log(qCap)", "log(qLab)", "log(qMat)")]) [1] 1.041133 The estimated elasticity of scale ϵ= 1.04 indicates slightly increasing returns to scale. 8.2.5 Properties 8.2.5.1 Non-increasing in input quantities Do(x, y)is non-increasing in xif its partial derivatives with respect to the input quantities are non-positive: ∂Do(x, y) ∂xi =∂ln Do(x, y) ∂ln xi Do(x, y) xi =αi Do(x, y) xi≤0(8.105) Given that Do(x, y)and xiare always non-negative, a Cobb-Douglas output distance function Do(x, y)is non-increasing in xif all coefficients αiare non-positive. We can see from the summary() output of our estimated Cobb-Douglas output distance function that all coefficients related to input quantities are negative. This means that monotonicity in input quantities is fulfilled. 8.2.5.2 Non-decreasing in output quantities Do(x, y)is non-decreasing in yif its partial derivatives with respect to the output quantities are non-negative: ∂Do(x, y) ∂yi =∂ln Do(x, y) ∂ln yi Do(x, y) yi =βi Do(x, y) yi≥0(8.106) Given that Do(x, y)and yiare always non-negative, a Cobb-Douglas output distance function Do(x, y)is non-decreasing in yif all coefficients βiare non-negative. For our Cobb-Douglas output distance function, the coefficient of the (normalized) quantity of “other outputs” is positive (0.621). The coefficient of the apple quantity can be recovered from the homogeneity condition (1-0.621 = 0.379). Since the coefficients of both output quantities are positive, we can conclude that monotonicity in output quantities is fulfilled. 313
8 Distance Functions + N X i=1 M X j=1 ζij ln xiln(k yj) ln k+ ln Do(x, y) =α0+ N X i=1 αiln xi+1 2 N X i=1 N X j=1 αij ln xiln xj(8.121) + M X i=1 βiln(k) + M X i=1 βiln yi +1 2 M X i=1 M X j=1 βij ln(k) ln(k) + 1 2 M X i=1 M X j=1 βij ln(k) ln(yj) +1 2 M X i=1 M X j=1 βij ln(yi) ln(k) + 1 2 M X i=1 M X j=1 βij ln(yi) ln(yj) + N X i=1 M X j=1 ζij ln xiln k+ N X i=1 M X j=1 ζij ln xiln yj ln k+ ln Do(x, y) = ln Do(x, y) + ln(k) M X i=1 βi+ ln(k) ln(k)1 2 M X i=1 M X j=1 βij (8.122) + ln(k)1 2 M X j=1 ln(yj) M X i=1 βij + ln(k)1 2 M X i=1 ln(yi) M X j=1 βij + ln k N X i=1 ln xi M X j=1 ζij ln k= ln(k) M X i=1 βi+ ln(k) ln(k)1 2 M X i=1 M X j=1 βij (8.123) + ln(k)1 2 M X j=1 ln(yj) M X i=1 βij + ln(k)1 2 M X i=1 ln(yi) M X j=1 βij + ln k N X i=1 ln xi M X j=1 ζij 1 = M X i=1 βi+ ln(k)1 2 M X i=1 M X j=1 βij (8.124) +1 2 M X j=1 ln(yj) M X i=1 βij +1 2 M X i=1 ln(yi) M X j=1 βij + N X i=1 ln xi M X j=1 ζij This condition is only fulfilled for all values of k,yi, and xi, if the parameters fulfill the following conditions: M X i=1 βi= 1 (8.125) 320
8 Distance Functions M X i=1 βij = 0 ∀jβij =βji ←−−−−→ M X j=1 βij = 0 ∀i(8.126) M X j=1 ζij = 0 ∀i(8.127) In order to impose linear homogeneity in output quantities, we can rearrange these restrictions to get: βM= 1 − M−1 X i=1 βi(8.128) βMj =− M−1 X i=1 βij ∀j= 1, . . . , M (8.129) βiM =− M−1 X j=1 βij ∀i= 1, . . . , M (8.130) ζiM =− M−1 X j=1 ζij ∀i= 1, . . . , N (8.131) By substituting the right-hand sides of equations (8.128) to (8.131) for βM,βMj,βiM , and ζiM in equation (8.117) and re-arranging this equation, we get: −ln yM=α0+ N X i=1 αiln xi+1 2 N X i=1 N X j=1 αij ln xiln xj(8.132) + M−1 X i=1 βiln yi yM +1 2 M−1 X i=1 M−1 X j=1 βij ln yi yM ln yj yM + N X i=1 M−1 X j=1 ζij ln xiln yj yM−ln Do(x, y). We can assume that u≡ −ln(Do(x, y)) ≥0follows a half-normal or truncated normal distribution (i.e. u∼N+(µ, σ2 u)) and add a disturbance term vthat accounts for statistical noise and follows a normal distribution (i.e. v∼N(0, σ2 v)) so that we get: −ln yM=α0+ N X i=1 αiln xi+1 2 N X i=1 N X j=1 αij ln xiln xj(8.133) + M−1 X i=1 βiln yi yM +1 2 M−1 X i=1 M−1 X j=1 βij ln yi yM ln yj yM + N X i=1 M−1 X j=1 ζij ln xiln yj yM +u+v. (8.134) This specification is equivalent to the specification of stochastic frontier models so that we 321
8 Distance Functions can use the stochastic frontier methods introduced in Chapter 6to estimate this output distance function. In the following, we use command sfa() to estimate a Translog output distance function for our data set of French apple producers: > odfTL <- sfa( -log(qApples) ~ log( qCap ) + log( qLab ) + log( qMat ) + + I( 0.5 * log( qCap )^2 ) + I( 0.5 * log( qLab )^2 ) + + I( 0.5 * log( qMat )^2 ) + I( log( qCap ) * log( qLab ) ) + + I( log( qCap ) * log( qMat ) ) + I( log( qLab ) * log( qMat ) ) + + log( qOtherOut/qApples) + I( 0.5 * log( qOtherOut/qApples )^2 ) + + I( log( qCap ) * log( qOtherOut/qApples ) ) + + I( log( qLab ) * log( qOtherOut/qApples ) ) + + I( log( qMat ) * log( qOtherOut/qApples ) ), + data = dat, ineffDecrease = FALSE) > summary( odfTL ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency increases the endogenous variable (as in a cost function) The dependent variable is logged Iterative ML estimation terminated after 28 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 38.882431 21.847091 1.7798 0.0751163 log(qCap) 0.901531 2.138098 0.4217 0.6732797 log(qLab) -5.046845 4.402014 -1.1465 0.2515943 log(qMat) -1.509307 3.738472 -0.4037 0.6864165 I(0.5 * log(qCap)^2) 0.068372 0.186448 0.3667 0.7138388 I(0.5 * log(qLab)^2) 1.185172 0.654824 1.8099 0.0703096 I(0.5 * log(qMat)^2) 0.204052 0.368668 0.5535 0.5799313 I(log(qCap) * log(qLab)) -0.475477 0.275217 -1.7276 0.0840518 I(log(qCap) * log(qMat)) 0.401302 0.194155 2.0669 0.0387422 I(log(qLab) * log(qMat)) -0.458295 0.409950 -1.1179 0.2635971 log(qOtherOut/qApples) -0.076972 0.784813 -0.0981 0.9218712 I(0.5 * log(qOtherOut/qApples)^2) 0.130865 0.018159 7.2065 5.740e-13 I(log(qCap) * log(qOtherOut/qApples)) -0.022233 0.052472 -0.4237 0.6717711 I(log(qLab) * log(qOtherOut/qApples)) 0.028802 0.081368 0.3540 0.7233620 I(log(qMat) * log(qOtherOut/qApples)) 0.057849 0.071400 0.8102 0.4178180 sigmaSq 0.677662 0.175928 3.8519 0.0001172 gamma 0.820876 0.135616 6.0530 1.422e-09 322
8 Distance Functions (Intercept) . log(qCap) log(qLab) log(qMat) I(0.5 * log(qCap)^2) I(0.5 * log(qLab)^2) . I(0.5 * log(qMat)^2) I(log(qCap) * log(qLab)) . I(log(qCap) * log(qMat)) * I(log(qLab) * log(qMat)) log(qOtherOut/qApples) I(0.5 * log(qOtherOut/qApples)^2) *** I(log(qCap) * log(qOtherOut/qApples)) I(log(qLab) * log(qOtherOut/qApples)) I(log(qMat) * log(qOtherOut/qApples)) sigmaSq *** gamma *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -116.8208 cross-sectional data total number of observations = 140 mean efficiency: 0.6032496 We can use a likelihood ratio test to compare the Translog output distance function with the corresponding Cobb-Douglas output distance function: > lrtest( odfCD, odfTL ) Likelihood ratio test Model 1: odfCD Model 2: odfTL #Df LogLik Df Chisq Pr(>Chisq) 1 7 -147.82 2 17 -116.82 10 62.003 1.511e-09 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 323
8 Distance Functions The likelihood ratio test clearly rejects the Cobb-Douglas output distance function in favor of the Translog output distance function. 8.3.3 Distance elasticities The distance elasticities of a Translog output distance function with respect to input quantities (x), ϵI, can be obtained by: ϵI i=∂Do(x, y) ∂xi xi Do=∂ln Do(x, y) ∂ln xi =αi+ N X j=1 αij ln xj+ M X j=1 ζij ln yj(8.135) The distance elasticities of a Translog output distance function with respect to output quantities (y), ϵO, can be obtained by: ϵO i=∂Do(x, y) ∂yi yi Do=∂ln Do(x, y) ∂ln yi =βi+ M X j=1 βij ln yj+ N X j=1 ζji ln xj(8.136) In order to facilitate the calculation of the distance elasticities, we create short-cuts for the estimated coefficients: > aCap <- coef( odfTL )["log(qCap)"] > aLab <- coef( odfTL )["log(qLab)"] > aMat <- coef( odfTL )["log(qMat)"] > aCapCap <- coef( odfTL )["I(0.5 * log(qCap)^2)"] > aCapLab <- aLabCap <- coef( odfTL )["I(log(qCap) * log(qLab))"] > aCapMat <- aMatCap <- coef( odfTL )["I(log(qCap) * log(qMat))"] > aLabLab <- coef( odfTL )["I(0.5 * log(qLab)^2)"] > aLabMat <- aMatLab <- coef( odfTL )["I(log(qLab) * log(qMat))"] > aMatMat <- coef( odfTL )["I(0.5 * log(qMat)^2)"] > bOther <- coef( odfTL )["log(qOtherOut/qApples)"] > bApples <- 1 - bOther > bOtherOther <- coef( odfTL )["I(0.5 * log(qOtherOut/qApples)^2)"] > bOtherApples <- bApplesOther <- - bOtherOther > bApplesApples <- - bApplesOther > zCapOther <- coef( odfTL )["I(log(qCap) * log(qOtherOut/qApples))"] > zCapApples <- - zCapOther > zLabOther <- coef( odfTL )["I(log(qLab) * log(qOtherOut/qApples))"] > zLabApples <- - zLabOther > zMatOther <- coef( odfTL )["I(log(qMat) * log(qOtherOut/qApples))"] > zMatApples <- - zMatOther The following Rcode computes the distance elasticities of the inputs and of the outputs: 324
8 Distance Functions > dat$eCapOdfTL <- with( dat, aCap + + aCapCap * log( qCap ) + aCapLab * log( qLab ) + aCapMat * log( qMat ) + + zCapApples * log( qApples ) + zCapOther * log( qOtherOut ) ) > dat$eLabOdfTL <- with( dat, aLab + + aLabCap * log( qCap ) + aLabLab * log( qLab ) + aLabMat * log( qMat ) + + zLabApples * log( qApples ) + zLabOther * log( qOtherOut ) ) > dat$eMatOdfTL <- with( dat, aMat + + aMatCap * log( qCap ) + aMatLab * log( qLab ) + aMatMat * log( qMat ) + + zMatApples * log( qApples ) + zMatOther * log( qOtherOut ) ) > dat$eApplesOdfTL <- with( dat, bApples + + bApplesApples * log( qApples ) + bApplesOther * log( qOtherOut ) + + zCapApples * log( qCap ) + zLabApples * log( qLab ) + + zMatApples * log( qMat ) ) > dat$eOtherOdfTL <- with( dat, bOther + + bOtherApples * log( qApples ) + bOtherOther * log( qOtherOut ) + + zCapOther * log( qCap ) + zLabOther * log( qLab ) + + zMatOther * log( qMat ) ) > summary( dat[ , + c( "eCapOdfTL", "eLabOdfTL", "eMatOdfTL", "eApplesOdfTL", "eOtherOdfTL" ) ] ) eCapOdfTL eLabOdfTL eMatOdfTL eApplesOdfTL Min. :-0.75715 Min. :-1.5312 Min. :-1.6399 Min. :-0.5368 1st Qu.:-0.25962 1st Qu.:-0.8143 1st Qu.:-0.7657 1st Qu.: 0.2643 Median :-0.13894 Median :-0.4487 Median :-0.5881 Median : 0.3948 Mean :-0.11768 Mean :-0.4782 Mean :-0.5788 Mean : 0.3935 3rd Qu.: 0.01858 3rd Qu.:-0.1502 3rd Qu.:-0.3657 3rd Qu.: 0.5289 Max. : 0.38125 Max. : 0.8224 Max. : 0.3553 Max. : 1.0616 eOtherOdfTL Min. :-0.06165 1st Qu.: 0.47106 Median : 0.60518 Mean : 0.60655 3rd Qu.: 0.73566 Max. : 1.53677 We can check whether the distance elasticities with respect to the output quantities always sum up to one: > range( dat$eApplesOdfTL + dat$eOtherOdfTL ) [1] 1 1 325
8 Distance Functions 8.3.4 Elasticity of scale We can use equation (8.37) in section 8.1.1.3 to calculate the elasticity of scale for our estimated Translog output distance function: > dat$eScaleOdfTL <- -rowSums( dat[ , c( "eCapOdfTL", "eLabOdfTL", "eMatOdfTL" ) ] ) > summary( dat$eScaleOdfTL ) Min. 1st Qu. Median Mean 3rd Qu. Max. 0.6621 1.0412 1.1951 1.1747 1.3053 1.7523 On average, the estimated elasticity of scale is ϵ= 1.175, which indicates increasing returns to scale. The following code plots the elasticities of scale against the two measures of firm size: > plot( dat$qOut, dat$eScaleOdfTL, log = "x" ) > abline( 1, 0 ) > plot( dat$X, dat$eScaleOdfTL, log = "x" ) > abline( 1, 0 ) 1e+05 5e+05 2e+06 1e+07 0.8 1.0 1.2 1.4 1.6 dat$qOut dat$eScaleOdfTL 0.5 1.0 2.0 5.0 0.8 1.0 1.2 1.4 1.6 dat$X dat$eScaleOdfTL Figure 8.5: Estimated elasticities of scale for different farm sizes The resulting graphs are shown in figure 8.5. According to these estimates of the elasticity of scale, farms with an aggregate output quantity of less than 1,000,000 units and with an aggregate input quantity of less than the sample mean generally have increasing returns to scale so that increasing their size would increase their (total factor) productivity. In contrast, farms with an aggregate output quantity of more than 5,000,000 units and with an aggregate input quantity of 326
8 Distance Functions more than 1.5 times the sample mean generally have decreasing returns to scale so that increasing their size would decrease their (total factor) productivity. 8.3.5 Properties 8.3.5.1 Non-increasing in input quantities In order to investigate the monotonicity of Translog output distance function in input quantities x, we calculate the partial derivatives (fi;i= 1, . . . , N) of the Translog output distance function with respect to the input quantities (xi): fi=∂Do(x, y) ∂xi =∂ln Do(x, y) ∂ln xi Do(x, y) xi =ϵI i Do(x, y) xi∀i. (8.137) As the output distance measure Do(x, y)and the input quantities xiare always non-negative, we can check the monotonicity in input quantities by checking the sign of the distance elasticities with respect to the input quantities: > table( dat$eCapOdfTL <= 0 ) FALSE TRUE 36 104 > table( dat$eLabOdfTL <= 0 ) FALSE TRUE 20 120 > table( dat$eMatOdfTL <= 0 ) FALSE TRUE 5 135 > dat$monoInOdfTL <- dat$eCapOdfTL <= 0 & dat$eLabOdfTL <= 0 & dat$eMatOdfTL <= 0 > table( dat$monoInOdfTL ) FALSE TRUE 57 83 8.3.5.2 Non-decreasing in output quantities In a similar way, we check the monotonicity of the Translog output distance function in output quantities yby calculating the partial derivatives (hi;i= 1, . . . , M) of the Translog output distance function with respect to the output quantities (yi): hi=∂Do(x, y) ∂yi =∂ln Do(x, y) ∂ln yi Do(x, y) yi =ϵO i Do(x, y) yi∀i. (8.138) 327
8 Distance Functions As the output distance measure Do(x, y)and the output quantities yiare always non-negative, we can check the monotonicity in output quantities by checking the sign of the distance elasticities with respect to the output quantities: > table( dat$eApplesOdfTL >= 0 ) FALSE TRUE 7 133 > table( dat$eOtherOdfTL >= 0 ) FALSE TRUE 2 138 > dat$monoOutOdfTL <- dat$eApplesOdfTL >= 0 & dat$eOtherOdfTL >= 0 > table( dat$monoOutOdfTL ) FALSE TRUE 9 131 Combining the monotonicity conditions for inputs and outputs, we can investigate how many observations fulfill all monotonicity conditions: > dat$monoOdfTL <- dat$monoInOdfTL & dat$monoOutOdfTL > table( dat$monoOdfTL ) FALSE TRUE 63 77 8.3.5.3 Quasiconvex in input quantities In the following, we will check if our estimated Translog output distance function is quasiconvex in input quantities. A sufficient condition for quasiconvexity of Do(x, y)in xis that |Bi|<0; i= 1, . . . , N, where |Bi|is the ith leading principal minor of the bordered Hessian matrix Bwith B= 0f1f2. . . fN f1f11 f12 . . . f1N f2f12 f22 . . . f2N . . .. . .. . ..... . . fNf1Nf2N. . . fNN ,(8.139) where fi=∂Do(x, y)/∂xiare the first-order partial derivatives and fij are the second-order partial derivatives of the Translog output distance function Do(x, y)with respect to the input quantities x. 328
8 Distance Functions We use equation (8.137) and the following Rcode to calculate the (first-order) partial derivatives of the Translog output distance function with respect to the input quantities xat the “frontier” (i.e. Do(x, y)=1) for each observation: > dat$fCapOdfTL <- with( dat, eCapOdfTL / qCap ) > dat$fLabOdfTL <- with( dat, eLabOdfTL / qLab ) > dat$fMatOdfTL <- with( dat, eMatOdfTL / qMat ) > summary( dat[ , c( "fCapOdfTL", "fLabOdfTL", "fMatOdfTL" ) ] ) fCapOdfTL fLabOdfTL fMatOdfTL Min. :-8.168e-05 Min. :-1.941e-05 Min. :-2.370e-04 1st Qu.:-4.015e-06 1st Qu.:-4.461e-06 1st Qu.:-4.396e-05 Median :-1.759e-06 Median :-2.165e-06 Median :-2.649e-05 Mean :-4.040e-06 Mean :-3.099e-06 Mean :-3.557e-05 3rd Qu.: 2.070e-07 3rd Qu.:-5.804e-07 3rd Qu.:-1.203e-05 Max. : 8.224e-06 Max. : 4.679e-06 Max. : 7.724e-06 We can calculate the second-order derivatives with respect to input quantities with following formula: fij =∂2Do(x, y) ∂xi∂xj =∂fi ∂xj =∂(∂ln Do(x,y) ∂ln xi Do(x,y) xi) ∂xj (8.140) =∂αi+PN k=1 αik ln xk+PM k=1 ζik ln ykDo(x,y) xi ∂xj (8.141) =αij xi Do(x, y) xi +∂ln Do(x, y) ∂ln xi ∂ln Do(x, y) ∂ln xj Do(x, y) xixj (8.142) −δij ∂ln Do(x, y) ∂ln xj Do(x, y) xixj (8.143) =αij Do(x, y) xixj +ϵI iϵI j Do(x, y) xixj (8.144) −δijϵI j Do(x, y) xixj (8.145) =αij +ϵI iϵI j−δijϵI jDo(x, y) xixj (8.146) We use the following Rcode to compute the second-order derivatives of the Translog output distance function with respect to input quantities xfor each observation: > dat$fCapCapOdfTL <- with( dat, + ( aCapCap + eCapOdfTL^2 - eCapOdfTL ) / qCap^2 ) > dat$fLabLabOdfTL <- with( dat, + ( aLabLab + eLabOdfTL^2 - eLabOdfTL ) / qLab^2 ) > dat$fMatMatOdfTL <- with( dat, 329
8 Distance Functions which can be linearized to: ln Di(x, y) = α0+ N X i=1 αiln xi+ M X i=1 βiln yi(8.156) with α0= ln A. 8.4.2 Estimation The Cobb-Douglas input distance function can be expressed (and then estimated) as traditional stochastic frontier model. In the following, we will derive, under which conditions the Cobb-Douglas input distance function defined in (8.155) and (8.156) is linear homogeneous in input quantities (x): kDi(x, y) = Di(kx, y)(8.157) ln(kDi(x, y)) = ln Di(kx, y)(8.158) ln(kDi(x, y)) = α0+ N X i=1 αiln kln xi+ M X i=1 βiln yi(8.159) ln k+ ln Di(x, y) = α0+ N X i=1 αiln k+ N X i=1 αiln xi+ M X i=1 βiln yi(8.160) ln k+ ln Di(x, y) = ln Di(x, y) + ln k N X i=1 αi(8.161) ln k= ln k N X i=1 αi(8.162) 1 = N X i=1 αi(8.163) Hence, the Cobb-Douglas input distance function is linear homogeneous in input quantities if the coefficients of the input quantities (αi) sum up to one. We can impose linear homogeneity in input quantities by re-arranging equation (8.163) to: αN= 1 − N−1 X i=1 αi(8.164) and by substituting the right-hand side of (8.164) for αNin (8.156): ln Di(x, y) = α0+ N−1 X i=1 αiln xi+ (1 − N−1 X i=1 αi) ln xN+ M X i=1 βiln yi(8.165) ln Di(x, y) = α0+ N−1 X i=1 αi(ln xi−ln xN) + ln xi+ M X i=1 βiln yi(8.166) 336
8 Distance Functions ln Di(x, y)−ln xN=α0+ N−1 X i=1 αi(ln xi/xN) + M X i=1 βiln yi(8.167) −ln xN=α0+ N−1 X i=1 αi(ln xi/xN) + M X i=1 βiln yi−ln Di(x, y)(8.168) We can assume that u≡ln(Di(x, y)) ≥0follows a half normal or truncated normal distribution (i.e. u∼N+(µ, σ2 u)) and add a disturbance term vthat accounts for statistical noise and follows a normal distribution (i.e. v∼N(0, σ2 v)) so that we get: −ln(xN) = α0+ N−1 X i=1 αiln(xi/xN) + M X i=1 βiln yi−u+v(8.169) This specification is equivalent to the specification of stochastic frontier models so that we can use the stochastic frontier methods introduced in Chapter 6to estimate this input distance function. In the following, we will estimate a Cobb-Douglas input distance function with our data set of French apple producers. This data set distinguishes between two outputs: the quantity of apples produced (qApples) and the quantity of the other outputs (qOtherOut). The following command estimates the Cobb-Douglas input distance function using the command sfa() for stochastic frontier analysis: > idfCD <- sfa( -log(qMat) ~ log( qCap/qMat ) + log( qLab/qMat ) + + log( qApples ) + log ( qOtherOut ), + ineffDecrease = TRUE, data = dat ) > summary( idfCD ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged Iterative ML estimation terminated after 10 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) -11.084030 0.135311 -81.9155 < 2.2e-16 *** log(qCap/qMat) 0.029982 0.040619 0.7381 0.4604 log(qLab/qMat) 0.684137 0.065195 10.4938 < 2.2e-16 *** log(qApples) -0.158068 0.022508 -7.0229 2.173e-12 *** log(qOtherOut) -0.129456 0.028820 -4.4918 7.061e-06 *** sigmaSq 0.375209 0.065546 5.7243 1.038e-08 *** gamma 0.925826 0.044667 20.7271 < 2.2e-16 *** 337
8 Distance Functions --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -60.46615 cross-sectional data total number of observations = 140 mean efficiency: 0.6594516 8.4.3 Distance elasticities The distance elasticities of a Cobb-Douglas input distance function Di(x, y)with respect to input quantities (x), ϵI, can be calculated as: ϵI i=∂Di(x, y) ∂xi xi Di=∂ln Di(x, y) ∂ln xi =αi(8.170) The distance elasticities of a Cobb-Douglas input distance function Di(x, y)with respect to output quantities (y), ϵO, can be calculated as: ϵO i=∂Di(x, y) ∂yi yi Di=∂ln Di(x, y) ∂ln yi =βi(8.171) Since the distance elasticities of Cobb-Douglas input distance functions are equal to some of the estimated parameters, they can be directly obtained from the estimation output. The distance elasticity with respect to the capital input (ϵI qCap = 0.03) indicates that a 1% increases in the capital input results (ceteris paribus) in a 0.03% increase of the distance measure (i.e., inefficiency increases and efficiency decreases). The distance elasticity with respect to the labor input (ϵI qLab = 0.684) indicates that a 1% increases in the labor input results (ceteris paribus) in a 0.68% increase of the distance measure (i.e., inefficiency increases and efficiency decreases). The distance elasticity with respect to the material input (ϵI qMat = 0.286) indicates that a 1% increase in the material input results (ceteris paribus) in a 0.29% increase of the distance measure (i.e., inefficiency increases and efficiency decreases). The distance elasticity with respect to the apples output (ϵO qApples =−0.158) indicates that a 1% increase in the apple output results (ceteris paribus) in a 0.16% decrease of the distance measure (i.e., inefficiency decreases and efficiency increases). The distance elasticity with respect to the other outputs (ϵO qOther =−0.129) indicates that a 1% increase in the other outputs results (ceteris paribus) in a 0.13% decrease of the distance measure (i.e., inefficiency decreases and efficiency increases). According to the alternative interpretation of the distance elasticities of the outputs (see section 8.1.2.2), the distance elasticity with respect to the apples output (ϵO qApples =−0.158) indicates that a 1% increase in the apple output increases the minimum required quantities of all inputs 338
8 Distance Functions (simultaneously) by 0.16%. Similarly, the distance elasticity with respect to the other outputs (ϵO qOther =−0.129) indicates that a 1% increase in the other outputs increases the minimum required quantities of all inputs (simultaneously) by 0.13%. 8.4.4 Elasticity of scale The elasticity of scale for an input distance function Di(x, y)equals the inverse of the negative sum of the distance elasticities with respect to the output quantities: ϵ= − M X i=1 ϵO i!−1 (8.172) For our Cobb-Douglas input distance function, the elasticity of scale can be obtained by using following command: > (-sum(coef(idfCD)[c("log(qApples)", "log(qOtherOut)")]))^(-1) [1] 3.477971 The estimated elasticity of scale ϵ= 3.478 indicates highly increasing returns to scale. 8.4.5 Properties 8.4.5.1 Non-decreasing in input quantities Di(x, y)is non-decreasing in xif its partial derivatives with respect to the input quantities are non-negative: ∂Di(x, y) ∂xi =∂ln Di(x, y) ∂ln xi Di(x, y) xi =αi Di(x, y) xi≥0.(8.173) Given that Di(x, y)and xiare always non-negative, a Cobb-Douglas input distance function fulfills the monotonicity condition, if all coefficients of the (normalized) input quantities are positive. This is the case for our estimated Cobb-Douglas input distance function, where the coefficient of the quantity of the material input (qMat) that was not directly estimated can be obtained from the homogeneity condition, i.e., 1−0.03 −0.684 = 0.286. 8.4.5.2 Non-increasing in output quantities Di(x, y)is non-increasing in yif its partial derivatives with respect to the output quantities are non-positive: ∂Di(x, y) ∂yi =∂ln Di(x, y) ∂ln yi Di(x, y) yi =βi Di(x, y) yi≤0.(8.174) 339
8 Distance Functions Given that Di(x, y)and yiare always non-negative, a Cobb-Douglas input distance function is non-increasing in output quantities if the coefficients of all output quantities are non-positive. This is the case for our estimated Cobb-Douglas input distance function so that we can conclude that it is non-increasing in output quantities. 8.4.5.3 Concave in input quantities In the following, we check if the estimated Cobb-Douglas input distance function is concave in input quantities (x). Di(x, y)is concave in xif and only if its Hessian matrix Hwith respect to the input quantities is negative semidefinite. A sufficient condition for negative semidefiniteness of a symmetric matrix is that all its ith-order principal minors (not only its leading principal minors) are non-positive for ibeing odd and non-negative for ibeing even for all i∈ {1, . . . , N} (see section 1.5.5). The Hessian matrix with respect to the input quantities is H= f11 f12 . . . f1N f12 f22 . . . f2N . . .. . ..... . . f1Nf2N. . . fNN ,(8.175) where fij are the second derivatives of Di(x, y)with respect to input quantities x: fij =∂2Di(x, y) ∂xi∂xj =∂fi ∂xj =∂(∂ln Di(x,y) ∂ln xi Di(x,y) xi) ∂xj (8.176) =αi xi ∂Di(x, y) ∂xi−δijαi Di(x, y) x2 i (8.177) =αi xi αj Do(x, y) xj−δijαi Di(x, y) x2 i (8.178) =αi(αj−δij)Di(x, y) xixj (8.179) where δij is (again) Kronecker’s delta (2.95). As all elements of the Hessian matrix (all second derivatives) contain the input distance measure Di(x, y)as a multiplicative element, we can factor it out of the Hessian matrix. As the input distance measure Di(x, y)is non-negative, it does not affect the signs of the principal minors so that we can ignore it in our further calculations. Again, for the sake of simplicity of the code, we create short-cuts for the coefficients: > iCap <- coef(idfCD)["log(qCap/qMat)"] > iLab <- coef(idfCD)["log(qCap/qMat)"] > iMat <- 1-(iCap+iLab) The following code computes the second derivatives of our estimated Cobb-Douglas input 340
8 Distance Functions distance function with respect to input quantities (x) (with the input distance measures being factored out of the Hessian matrix or assuming that the firms are fully efficient, i.e. Di(x, y)=1): > dat$fCapCap <- iCap * ( iCap - 1 ) * 1 / dat$qCap^2 > dat$fLabLab <- iLab * ( iLab - 1 ) * 1 / dat$qLab^2 > dat$fMatMat <- iMat * ( iMat - 1 ) * 1 / dat$qMat^2 > dat$fCapLab <- iCap * iLab * 1 / ( dat$qCap * dat$qLab ) > dat$fCapMat <- iCap * iMat * 1 / ( dat$qCap * dat$qMat ) > dat$fLabMat <- iLab * iMat * 1 / ( dat$qLab * dat$qMat ) The following code generates a three-dimensional array of the Hessian matrices stacked for all observations: > hessianArray <- array( NA, c( 3, 3, nrow( dat ) ) ) > hessianArray[ 1, 1, ] <- dat$fCapCap > hessianArray[ 1, 2, ] <- hessianArray[ 2, 1, ] <- dat$fCapLab > hessianArray[ 1, 3, ] <- hessianArray[ 3, 1, ] <- dat$fCapMat > hessianArray[ 2, 2, ] <- dat$fLabLab > hessianArray[ 2, 3, ] <- hessianArray[ 3, 2, ] <- dat$fLabMat > hessianArray[ 3, 3, ] <- dat$fMatMat > print( hessianArray[ , , 1 ] ) [,1] [,2] [,3] [1,] -4.116775e-12 2.970247e-14 9.837352e-12 [2,] 2.970247e-14 -2.243230e-13 2.296342e-12 [3,] 9.837352e-12 2.296342e-12 -4.851371e-11 In the following, we compute all principal minors at the first observation: > diag( hessianArray[ , , 1 ] ) [1] -4.116775e-12 -2.243230e-13 -4.851371e-11 > det( hessianArray[ -3, -3, 1 ] ) [1] 9.226048e-25 > det( hessianArray[ -2, -2, 1 ] ) [1] 1.029465e-22 > det( hessianArray[ -1, -1, 1 ] ) [1] 5.609554e-24 341
8 Distance Functions > det( hessianArray[ , , 1 ] ) [1] -3.936263e-51 As all first-order principal minors (diagonal elements) are negative, all second-order principal minors are positive, and the determinant (third-order principal minor) is virtually zero, the conditions for concavity are fulfilled at the first observation. The following code checks the concavity at all observations: > dat$CDidfConcaveX <- apply( hessianArray, 3, semidefiniteness, + positive = FALSE ) > table( dat$CDidfConcaveX ) TRUE 140 The estimated Cobb-Douglas input distance function is concave in input quantities at all observations. 8.4.5.4 Quasiconcave in output quantities In the following, we check if the estimated Cobb-Douglas input distance function Di(x, y)is quasiconcave in output quantities (y). A sufficient condition for quasiconcavity in yis that |B1|<0,|B2|>0,|B3|<0, . . . , (−1)N|BN|>0, where |Bi|is the ith principal minor of the bordered Hessian matrix Bwith B= 0f1f2. . . fN f1f11 f12 . . . f1N f2f12 f22 . . . f2N . . .. . .. . ..... . . fNf1Nf2N. . . fNN ,(8.180) where fi=∂Di(x, y)/∂yiare the first partial derivatives and fij are the second partial derivatives of the input distance function with respect to the output quantities y: fij =∂2Di(x, y) ∂yi∂yj =∂∂Di(x,y) ∂yi ∂yj =∂βiDi(x,y) yi ∂yj (8.181) =βi yi ∂Di(x, y) ∂yj−δijβi Di(x, y) y2 i (8.182) =βi yi βj Di(x, y) yj−δijβi Di(x, y) y2 i (8.183) =βi(βj−δij)Di(x, y) yiyj ,(8.184) 342
8 Distance Functions where δij is (again) Kronecker’s delta (2.95). As all elements of the bordered Hessian matrix (both the first derivatives and the second derivatives) contain the input distance measure Di(x, y)as a multiplicative element, we can factor it out of the bordered Hessian matrix. As the output distance measure Di(x, y)is nonnegative, it does not affect the signs of the principal minors so that we can ignore it in our further calculations. In order to facilitate further computations, we create short-cuts for the coefficients regarding the output quantities y: > iApples <- coef(idfCD)["log(qApples)"] > iOther <- coef(idfCD)["log(qOtherOut)"] The following code computes the first-order and second-order partial derivatives of the CobbDouglas input distance function with respect to the output quantities (with the input distance measures being factored out of the bordered Hessian matrix or assuming that the firms are fully efficient, i.e. Di(x, y)=1): > dat$fApples <- iApples * 1 / dat$qApples > dat$fOther <- iOther * 1 / dat$qOtherOut > dat$fApplesApples <- iApples * ( iApples - 1 ) * 1 / dat$qApples^2 > dat$fApplesOther <- iApples * iOther * 1 / (dat$qApples * dat$qOtherOut) > dat$fOtherOther <- iOther * ( iOther - 1 ) * 1 / dat$qOtherOut^2 The following code creates a three-dimensional array with the bordered Hessian matrices at all observations stacked upon each other: > bhmArray <- array( 0, c( 3, 3, nrow( dat ) ) ) > bhmArray[ 1, 2, ] <- bhmArray[ 2, 1, ] <- dat$fApples > bhmArray[ 1, 3, ] <- bhmArray[ 3, 1, ] <- dat$fOther > bhmArray[ 2, 2, ] <- dat$fApplesApples > bhmArray[ 2, 3, ] <- bhmArray[ 3, 2, ] <- dat$fApplesOther > bhmArray[ 3, 3, ] <- dat$fOtherOther > print( bhmArray[ , , 1 ] ) [,1] [,2] [,3] [1,] 0.0000000 -0.11357085 -0.13246637 [2,] -0.1135708 0.09449814 0.01504432 [3,] -0.1324664 0.01504432 0.15309443 We check quasiconcavity in output quantities at the first observation by calculating the first and second leading principal minor: > det( bhmArray[ 1:2, 1:2, 1 ] ) 343
8 Distance Functions [1] -0.01289834 > det( bhmArray[ , , 1 ] ) [1] -0.003180192 As the determinant (second leading principal minor) of the bordered Hessian matrix is negative, we can conclude that the estimated Cobb-Douglas input distance function is not quasiconcave in output quantities at the first observation. The following code checks this at all observations: > dat$CDidfQuasiConcaveY <- apply( bhmArray[1:2,1:2, ], 3, det ) < 0 & + apply( bhmArray, 3, det ) > 0 > table( dat$CDidfQuasiConcaveY ) FALSE 140 This estimated Cobb-Douglas input distance function is quasiconcave in output quantities (y) not at a single observation in our data set. Indeed, Cobb-Douglas input distance functions cannot be quasiconcave in output quantities, because this functional form implies transformation curves that are convex to the origin. 8.4.6 Efficiency estimates If there were no inefficiencies the input distance Diwould be equal to 1 (what would imply that ln Di(1) = 0, i.e., u= 0) and all observations would be on the frontier. We can use likelihood ratio test to investigate if there is inefficiency > lrtest( idfCD ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 6 -66.739 2 7 -60.466 1 12.546 0.0001986 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 The small p-value indicates that there is statistically significant inefficiency. 344
8 Distance Functions The efficiency estimates can be obtained by the efficiencies method. By default, the efficiencies method calculates the efficiency estimates as d TE =E[e−u]. If we set argument minusU of the efficiencies method to FALSE, it calculates the efficiency estimates as d TE =E[eu].7Hence, in order to obtain Shepard input-oriented efficiency estimates from an input distance function, we need to set argument minusU of the efficiencies method to FALSE: > dat$effIdfCD <- efficiencies( idfCD, minusU = FALSE ) > summary( dat$effIdfCD ) efficiency Min. :1.069 1st Qu.:1.244 Median :1.493 Mean :1.718 3rd Qu.:1.887 Max. :5.207 A median Shepard input-oriented efficiency estimate of 1.493 indicates that this apple producer uses 49.3% more of each input quantity than would be necessary to produce its output quantities. 8.5 Translog input distance function 8.5.1 Specification The general form of the Translog input distance function is: ln Di(x, y) =α0+ N X i=1 αiln xi+1 2 N X i=1 N X j=1 αij ln xiln xj(8.185) + M X i=1 βiln yi+1 2 M X i=1 M X j=1 βij ln yiln yj + N X i=1 M X j=1 ζij ln xiln yj, with αij =αji ∀i, j = 1, . . . , N and βij =βji ∀i, j = 1, . . . , M. 8.5.2 Estimation The Translog input distance function can be expressed (and then estimated) as traditional stochastic frontier model. This is facilitated by imposing linear homogeneity in the input quantities (x). 7Please note that E[eu]is not equal to E[e−u]−1. 345
8 Distance Functions Median : 0.06433 Median :0.6240 Median : 0.3072 Median :-0.2571 Mean : 0.04577 Mean :0.6124 Mean : 0.3418 Mean :-0.2309 3rd Qu.: 0.16061 3rd Qu.:0.6818 3rd Qu.: 0.5105 3rd Qu.:-0.1431 Max. : 0.35954 Max. :0.8099 Max. : 1.0109 Max. : 0.3607 eOtherTLidf Min. :-0.29949 1st Qu.:-0.18255 Median :-0.15620 Mean :-0.15859 3rd Qu.:-0.13553 Max. :-0.07973 We can check whether the distance elasticities with respect to the input quantities always sum up to one: > range( dat$eCapTLidf + dat$eLabTLidf + dat$eMatTLidf ) [1] 1 1 8.5.4 Elasticity of scale We can use equation (8.87) in section 8.1.2.3 to calculate the elasticity of scale for our estimated Translog input distance function: > dat$eScaleTLidf <- ( - ( dat$eApplesTLidf + dat$eOtherTLidf ) )^(-1) > summary( dat$eScaleTLidf ) Min. 1st Qu. Median Mean 3rd Qu. Max. -53.173 2.028 2.368 3.242 3.224 69.218 8.5.5 Efficiency estimates If there were no inefficiencies, the input distance measure Diof all observations would be equal to 1 (which implies u= ln Di= 0) and all observations would be on the frontier. We can use a likelihood ratio test to investigate whether there is inefficiency: > lrtest( idfTL ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 352
8 Distance Functions 1 16 -31.980 2 17 -29.399 1 5.1635 0.01153 * --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 The test rejects the hypothesis that all firms are fully technically efficient at 5% significance level. The Shepard input-oriented efficiency estimates can be obtained by the efficiencies method: > dat$effIdfTL <- efficiencies( idfTL, minusU = FALSE ) > summary( dat$effIdfTL ) efficiency Min. :1.059 1st Qu.:1.190 Median :1.348 Mean :1.498 3rd Qu.:1.719 Max. :2.937 8.5.6 Imposing monotonicity The monotonicity conditions can be written as: ∂ln Di(x, y) ∂ln xi =αi+ N X j=1 αij ln xj+ M X j=1 ζij ln yj≥0∀i= 1, . . . , N (8.205) ∂ln Di(x, y) ∂ln yi =βi+ M X j=1 βij ln yj+ N X j=1 ζji ln xj≤0∀i= 1, . . . , M (8.206) Replacing the non-estimated coefficients by the right-hand sides of equations (8.197) to (8.200) and multiplying the monotonicity restrictions regarding the outputs by −1, we get: ∂ln Di(x, y) ∂ln xi =αi+ N−1 X j=1 αij ln xj+αiN ln xN+ M X j=1 ζij ln yj(8.207) =αi+ N−1 X j=1 αij ln xj+ − N−1 X j=1 αij ln xN+ M X j=1 ζij ln yj(8.208) =αi+ N−1 X j=1 αij ln xj xN + M X j=1 ζij ln yj≥0∀i= 1, . . . , N (8.209) ∂ln Di(x, y) ∂ln xN =αN+ N−1 X j=1 αNj ln xj xN + M X j=1 ζNj ln yj(8.210) = 1 − N−1 X i=1 αi+ N−1 X j=1 − N−1 X i=1 αij!ln xj xN + M X j=1 − N−1 X i=1 ζij!ln yj(8.211) 353
8 Distance Functions = 1 − N−1 X i=1 αi− N−1 X i=1 N−1 X j=1 αij ln xj xN− N−1 X i=1 M X j=1 ζij ln yj≥0(8.212) ⇔ − N−1 X i=1 αi− N−1 X i=1 N−1 X j=1 αij ln xj xN− N−1 X i=1 M X j=1 ζij ln yj≥ −1(8.213) −∂ln Di(x, y) ∂ln yi =−βi− M X j=1 βij ln yj− N−1 X j=1 ζji ln xj−ζNi ln xN(8.214) =−βi− M X j=1 βij ln yj− N−1 X j=1 ζji ln xj− − N−1 X j=1 ζji ln xN(8.215) =−βi− M X j=1 βij ln yj− N−1 X j=1 ζji ln xj xN≥0(8.216) We can write these monotonicity restrictions using matrix notation as: R θ ≥r, (8.217) where R= [Rα, Rβ, Rζ]with (8.218) Rα= 0 1 0 . . . 0 ln x1 xNln x2 xN. . . ln xN−1 xN0. . . 0. . . 0 0 0 1 . . . 0 0 ln x1 xN. . . 0 ln x2 xN. . . ln xN−1 xN. . . 0 . . .. . .. . ..... . .. . .. . ..... . .. . ..... . ..... . . 0 0 0 . . . 1 0 0 . . . ln x1 xN0. . . x2 xN. . . xN−1 xN 0−1−1. . . −1−ln x1 xN−ln x1 xN+ ln x2 xN. . . −ln x1 xN+ ln xN−1 xN−ln x2 xN. . . −x2 xN+xN−1 xN. . . −xN−1 xN 0 0 0 . . . 0 0 0 . . . 0 0 . . . 0. . . 0 0 0 0 . . . 0 0 0 . . . 0 0 . . . 0. . . 0 . . .. . .. . ..... . .. . .. . ..... . .. . ..... . ..... . . 0 0 0 . . . 0 0 0 . . . 0 0 . . . 0. . . 0 (8.219) Rβ= 0 0 . . . 0 0 0 . . . 0 0 . . . 0. . . 0 0 0 . . . 0 0 0 . . . 0 0 . . . 0. . . 0 . . .. . ..... . .. . .. . ..... . .. . ..... . ..... . . 0 0 . . . 0 0 0 . . . 0 0 . . . 0. . . 0 0 0 . . . 0 0 0 . . . 0 0 . . . 0. . . 0 −1 0 . . . 0−ln y1−ln y2. . . −ln yM0. . . 0. . . 0 0−1. . . 0 0 −ln y1. . . 0−ln y2. . . −lnyM. . . 0 . . .. . ..... . .. . .. . ..... . .. . ..... . ..... . . 0 0 . . . −1 0 0 . . . −ln y10. . . −lny2. . . −ln yM (8.220) 354
8 Distance Functions Rζ= ln y1ln y2. . . ln yM0 0 . . . 0. . . 0 0 . . . 0 0 0 . . . 0 ln y1ln y2. . . ln yM. . . 0 0 . . . 0 . . .. . ..... . .. . .. . ..... . .. . ..... . ..... . . 0 0 . . . 0 0 0 . . . 0. . . ln y1ln y2. . . ln yM −ln y1−ln y2. . . −ln yM−ln y1−ln y2. . . −ln yM. . . −ln y1−ln y2. . . −ln yM −ln x1 xN0. . . 0−ln x2 xN0. . . 0. . . −ln xN−1 xN0. . . 0 0−ln x1 xN. . . 0 0 −ln x2 xN. . . 0. . . 0−ln xN−1 xN. . . 0 . . .. . ..... . .. . .. . ..... . .. . ..... . ..... . . 0 0 . . . −ln x1 xN0 0 . . . −ln x2 xN. . . 0 0 . . . −ln xN−1 xN (8.221) is the restriction matrix for the monotonicity restrictions (∂ln Di(x, y)/∂ ln x1,∂ln Di(x, y)/∂ ln x2, . . . , ∂ln Di(x, y)/∂ ln xN−1,∂ln Di(x, y)/∂ ln xN,∂ln Di(x, y)/∂ ln y1,∂ln Di(x, y)/∂ ln y2, . . . , ∂ln Di(x, y)/∂ ln yM), θ=α′, β′, ζ′′with (8.222) α= (α0, α1, α2, . . . , αN−1, α11, α12, . . . , α1,N−1, α22, . . . , α2,N−1, . . . , αN−1,N−1)′(8.223) β= (β1, β2, . . . , βM, β11, β12, . . . , β1M, β22, . . . , β2M, . . . , βMM )′(8.224) ζ= (ζ11, ζ12, . . . , ζ1M, ζ21, ζ22, . . . , ζ2M, . . . , ζN−1,1, ζN−1,2, . . . , ζN−1,M ,)′(8.225) is the vector of estimated coefficients, and r= (0,0,...,0,−1,0,0,...,0) is the vector of righthand-side values. If one has Ninputs and Moutputs and one wants to impose monotonicity at nobservations, the restriction matrix Rmust have n·(N+M)rows and 1+(N−1)(1+N/2)+M(3+M)/2+(N−1)M columns, where the number of columns is equal to the number of estimated coefficients. As our data set has 140 observations and we have 3 inputs and 2 outputs, our restriction matrix must have 140 ·(3 + 2) = 700 rows and 1+2·2.5+2·5/2+2·2 = 15 columns, The following code generates the restriction matrix and the restriction vector: > RMat <- matrix( 0, nrow = 5 * nrow( dat ), ncol = 15 ) > colnames( RMat ) <- names( coef( idfTL )[1:15] ) > rVec <- rep( 0, 5 * nrow( dat ) ) > # regarding capital input > rowsCap <- 1:nrow(dat) > RMat[ rowsCap, "log(qCap/qMat)" ] <- 1 > RMat[ rowsCap, "I(0.5 * log(qCap/qMat)^2)" ] <- log( dat$qCap / dat$qMat ) > RMat[ rowsCap, "I(log(qCap/qMat) * log(qLab/qMat))" ] <- log( dat$qLab / dat$qMat ) > RMat[ rowsCap, "I(log(qCap/qMat) * log(qApples))" ] <- log( dat$qApples ) > RMat[ rowsCap, "I(log(qCap/qMat) * log(qOtherOut))" ] <- log( dat$qOtherOut ) > # regarding labor input > rowsLab <- ( nrow(dat) + 1 ):( 2 * nrow(dat) ) > RMat[ rowsLab, "log(qLab/qMat)" ] <- 1 > RMat[ rowsLab, "I(0.5 * log(qLab/qMat)^2)" ] <- log( dat$qLab / dat$qMat ) 355
8 Distance Functions > RMat[ rowsLab, "I(log(qCap/qMat) * log(qLab/qMat))" ] <- log( dat$qCap / dat$qMat ) > RMat[ rowsLab, "I(log(qLab/qMat) * log(qApples))" ] <- log( dat$qApples ) > RMat[ rowsLab, "I(log(qLab/qMat) * log(qOtherOut))" ] <- log( dat$qOtherOut ) > # regarding materials input > rowsMat <- ( 2 * nrow(dat) + 1 ):( 3 * nrow(dat) ) > RMat[ rowsMat, "log(qCap/qMat)" ] <- -1 > RMat[ rowsMat, "log(qLab/qMat)" ] <- -1 > RMat[ rowsMat, "I(0.5 * log(qCap/qMat)^2)" ] <- -log( dat$qCap / dat$qMat ) > RMat[ rowsMat, "I(0.5 * log(qLab/qMat)^2)" ] <- -log( dat$qLab / dat$qMat ) > RMat[ rowsMat, "I(log(qCap/qMat) * log(qLab/qMat))" ] <- + - ( log( dat$qCap / dat$qMat ) + log( dat$qLab / dat$qMat ) ) > RMat[ rowsMat, "I(log(qCap/qMat) * log(qApples))" ] <- -log( dat$qApples ) > RMat[ rowsMat, "I(log(qCap/qMat) * log(qOtherOut))" ] <- -log( dat$qOtherOut ) > RMat[ rowsMat, "I(log(qLab/qMat) * log(qApples))" ] <- -log( dat$qApples ) > RMat[ rowsMat, "I(log(qLab/qMat) * log(qOtherOut))" ] <- -log( dat$qOtherOut ) > rVec[ rowsMat ] <- -1 > # regarding apple output > rowsApp <- ( 3 * nrow(dat) + 1 ):( 4 * nrow(dat) ) > RMat[ rowsApp, "log(qApples)" ] <- -1 > RMat[ rowsApp, "I(0.5 * log(qApples)^2)" ] <- -log( dat$qApples ) > RMat[ rowsApp, "I(log(qApples) * log(qOtherOut))" ] <- -log( dat$qOtherOut ) > RMat[ rowsApp, "I(log(qCap/qMat) * log(qApples))" ] <- -log( dat$qCap / dat$qMat ) > RMat[ rowsApp, "I(log(qLab/qMat) * log(qApples))" ] <- -log( dat$qLab / dat$qMat ) > # regarding other outputs > rowsOth <- ( 4 * nrow(dat) + 1 ):( 5 * nrow(dat) ) > RMat[ rowsOth, "log(qOtherOut)" ] <- -1 > RMat[ rowsOth, "I(log(qApples) * log(qOtherOut))" ] <- -log( dat$qApples ) > RMat[ rowsOth, "I(0.5 * log(qOtherOut)^2)" ] <- -log( dat$qOtherOut ) > RMat[ rowsOth, "I(log(qCap/qMat) * log(qOtherOut))" ] <- -log( dat$qCap / dat$qMat ) > RMat[ rowsOth, "I(log(qLab/qMat) * log(qOtherOut))" ] <- -log( dat$qLab / dat$qMat ) As R θ −rshould be equal to the distance elasticities of the inputs and the negative distance elasticities of the outputs, we can use the following code to check whether we have specified the restriction matrix Rand the restriction vector rcorrectly: > all.equal( c( RMat %*% coef( idfTL )[1:15] - rVec ), + c( dat$eCapTLidf, dat$eLabTLidf, dat$eMatTLidf, + -dat$eApplesTLidf, -dat$eOtherTLidf ) ) [1] TRUE We can obtain restricted coefficients as suggested by Henningsen and Henning (2009): 356
8 Distance Functions > uCoef <- coef( idfTL )[1:15] > uCovInv <- solve( vcov( idfTL )[ 1:15,1:15] ) > library( "quadprog" ) > minDistResult <- solve.QP( Dmat = uCovInv, dvec = rep( 0, length( uCoef ) ), + Amat = t( RMat ), bvec = - RMat %*% uCoef + rVec ) > rCoef <- minDistResult$solution + uCoef We can calculate the distance elasticities of the inputs and the (negative) distance elasticities of the outputs based on the restricted coefficients by: > elaIdfTl <- RMat %*% rCoef - rVec > dat$eCapTLRidf <- elaIdfTl[ 1:nrow( dat ) ] > dat$eLabTLRidf <- elaIdfTl[ ( nrow( dat ) + 1 ):( 2 * nrow( dat ) ) ] > dat$eMatTLRidf <- elaIdfTl[ ( 2 * nrow( dat ) + 1 ):( 3 * nrow( dat ) ) ] > dat$eApplesTLRidf <- -elaIdfTl[ ( 3 * nrow( dat ) + 1 ):( 4 * nrow( dat ) ) ] > dat$eOtherTLRidf <- -elaIdfTl[ ( 4 * nrow( dat ) + 1 ):( 5 * nrow( dat ) ) ] > summary( dat[ , c( "eCapTLRidf", "eLabTLRidf", "eMatTLRidf", + "eApplesTLRidf", "eOtherTLRidf" ) ] ) eCapTLRidf eLabTLRidf eMatTLRidf eApplesTLRidf Min. :0.00000 Min. :0.5105 Min. :0.1150 Min. :-0.2958 1st Qu.:0.06868 1st Qu.:0.6228 1st Qu.:0.1987 1st Qu.:-0.2229 Median :0.10036 Median :0.6596 Median :0.2434 Median :-0.1943 Mean :0.09914 Mean :0.6571 Mean :0.2437 Mean :-0.1864 3rd Qu.:0.13391 3rd Qu.:0.6950 3rd Qu.:0.2737 3rd Qu.:-0.1593 Max. :0.20573 Max. :0.7971 Max. :0.4360 Max. : 0.0000 eOtherTLRidf Min. :-0.19191 1st Qu.:-0.15044 Median :-0.13416 Mean :-0.13396 3rd Qu.:-0.11993 Max. :-0.06537 We can see that all distance elassticities of the three are non-negative and all distance elasticities of the two outputs are non-positive. The homogeneity condition is—of course—still fulfilled: > range( dat$eCapTLRidf + dat$eLabTLRidf + dat$eMatTLRidf ) [1] 1 1 And the elasticities of scale can be calculated by: 357
8 Distance Functions > dat$eScaleTLRidf <- ( - ( dat$eApplesTLRidf + dat$eOtherTLRidf ) )^(-1) > summary( dat$eScaleTLRidf ) Min. 1st Qu. Median Mean 3rd Qu. Max. 2.050 2.741 2.992 3.418 3.515 15.298 358
9 Panel Data and Technological Change Until now, we have only analyzed cross-sectional data, i.e. all observations refer to the same period of time. Hence, it was reasonable to assume that the same technology is available to all firms (observations). However, when analyzing time series data or panel data, i.e. when observations can originate from different time periods, different technologies might be available in the different time periods due to technological change. Hence, the state of the available technologies must be included as an explanatory variable in order to conduct a reasonable production analysis. Often, a time trend is used as a proxy for a gradually changing state of the available technologies. We will demonstrate how to analyze production technologies with data from different time periods by using a balanced panel data set of annual data collected from 43 smallholder rice producers in the Tarlac region of the Philippines between 1990 and 1997. We loaded this data set (riceProdPhil) in section 1.4.2. As it does not contain information about the panel structure, we created a copy of the data set (pdat) that includes information on the panel structure. 9.1 Average production functions with technological change In case of an applied production analysis with time-series data or panel data, usually the time (t) is included as additional explanatory variable in the production function: y=f(x, t).(9.1) This function can be used to analyze how the time (t) affects the (available) production technology. The average production technology (potentially depending on the time period) can be estimated from panel data sets by the OLS method (i.e. “pooled”) or by any of the usual panel data methods (e.g. fixed effects, random effects). 9.1.1 Cobb-Douglas production function with technological change In case of a Cobb-Douglas production function, usually a linear time trend is added to account for technological change: ln y=α0+X i αiln xi+αtt(9.2) 359
9 Panel Data and Technological Change Given this specification, the coefficient of the (linear) time trend can be interpreted as the rate of technological change per unit of the time variable t: αt=∂ln y ∂t =∂ln y ∂y ∂y ∂t ≈ ∆y y ∆t(9.3) 9.1.1.1 Pooled estimation of the Cobb-Douglas production function with technological change The pooled estimation can be done by: > riceCdTime <- lm( log( PROD ) ~ log( AREA ) + log( LABOR ) + log( NPK ) + + mYear, data = riceProdPhil ) > summary( riceCdTime ) Call: lm(formula = log(PROD) ~ log(AREA) + log(LABOR) + log(NPK) + mYear, data = riceProdPhil) Residuals: Min 1Q Median 3Q Max -1.83351 -0.16006 0.05329 0.22110 0.86745 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) -1.665096 0.248509 -6.700 8.68e-11 *** log(AREA) 0.333214 0.062403 5.340 1.71e-07 *** log(LABOR) 0.395573 0.066421 5.956 6.48e-09 *** log(NPK) 0.270847 0.041027 6.602 1.57e-10 *** mYear 0.010090 0.008007 1.260 0.208 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 0.3299 on 339 degrees of freedom Multiple R-squared: 0.86, Adjusted R-squared: 0.8583 F-statistic: 520.6 on 4 and 339 DF, p-value: < 2.2e-16 The estimation result indicates an annual rate of technical change of approximately 1%, but this is not statistically different from 0%, i.e., we cannot reject that there has been no technological change. The output elasticities are equal to the coefficients of the (logarithmic) input quantities and the elasticity of scale is almost exactly one (0.99963), indicating (approximately) constant returns to scale. 360
9 Panel Data and Technological Change The command above can be simplified by using the pre-calculated logarithmic (and meanscaled) quantities: > riceCdTimeS <- lm( lProd ~ lArea + lLabor + lNpk + mYear, data = riceProdPhil ) > summary( riceCdTimeS ) Call: lm(formula = lProd ~ lArea + lLabor + lNpk + mYear, data = riceProdPhil) Residuals: Min 1Q Median 3Q Max -1.83351 -0.16006 0.05329 0.22110 0.86745 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) -0.015590 0.019325 -0.807 0.420 lArea 0.333214 0.062403 5.340 1.71e-07 *** lLabor 0.395573 0.066421 5.956 6.48e-09 *** lNpk 0.270847 0.041027 6.602 1.57e-10 *** mYear 0.010090 0.008007 1.260 0.208 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 0.3299 on 339 degrees of freedom Multiple R-squared: 0.86, Adjusted R-squared: 0.8583 F-statistic: 520.6 on 4 and 339 DF, p-value: < 2.2e-16 The intercept has changed because of the mean-scaling of the input and output quantities but all slope parameters are unaffected by using the pre-calculated logarithmic (and mean-scaled) quantities: > all.equal( coef( riceCdTime )[-1], coef( riceCdTimeS )[-1], + check.attributes = FALSE ) [1] TRUE 9.1.1.2 Panel data estimations of the Cobb-Douglas production function with technological change The panel data estimation with fixed individual effects can be done by: > riceCdTimeFe <- plm( lProd ~ lArea + lLabor + lNpk + mYear, data = pdat ) > summary( riceCdTimeFe ) 361
9 Panel Data and Technological Change Res.Df Df F Pr(>F) 1 339 2 333 6 5.1483 4.451e-05 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 The Cobb-Douglas specification is clearly rejected in favor of the Translog specification for the pooled estimation. 9.1.2.2 Panel-data estimations of the Translog production function with constant and neutral technological change The following command estimates a Translog production function that can account for constant and neutral technical change with fixed individual effects: > riceTlTimeFe <- plm( lProd ~ lArea + lLabor + lNpk + + I( 0.5 * lArea^2 ) + I( 0.5 * lLabor^2 ) + I( 0.5 * lNpk^2 ) + + I( lArea * lLabor ) + I( lArea * lNpk ) + I( lLabor * lNpk ) + mYear, + data = pdat, model = "within" ) > summary( riceTlTimeFe ) Oneway (individual) effect Within Model Call: plm(formula = lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear, data = pdat, model = "within") Balanced Panel: n = 43, T = 8, N = 344 Residuals: Min. 1st Qu. Median 3rd Qu. Max. -1.012473 -0.144573 0.019129 0.167687 0.745525 Coefficients: Estimate Std. Error t-value Pr(>|t|) lArea 0.5828102 0.1173298 4.9673 1.16e-06 *** lLabor 0.0473355 0.0848594 0.5578 0.577402 lNpk 0.1211928 0.0610114 1.9864 0.047927 * I(0.5 * lArea^2) -0.8543901 0.2861292 -2.9860 0.003067 ** 368
9 Panel Data and Technological Change I(0.5 * lLabor^2) -0.6217163 0.2935429 -2.1180 0.035025 * I(0.5 * lNpk^2) 0.0429446 0.0987119 0.4350 0.663849 I(lArea * lLabor) 0.5867063 0.2125686 2.7601 0.006145 ** I(lArea * lNpk) 0.1167509 0.1461380 0.7989 0.424995 I(lLabor * lNpk) -0.2371219 0.1268671 -1.8691 0.062619 . mYear 0.0165309 0.0069206 2.3887 0.017547 * --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Total Sum of Squares: 43.632 Residual Sum of Squares: 21.912 R-Squared: 0.49781 Adj. R-Squared: 0.40807 F-statistic: 28.8456 on 10 and 291 DF, p-value: < 2.22e-16 And the panel data estimation with random individual effects can be done by: > riceTlTimeRan <- plm( lProd ~ lArea + lLabor + lNpk + + I( 0.5 * lArea^2 ) + I( 0.5 * lLabor^2 ) + I( 0.5 * lNpk^2 ) + + I( lArea * lLabor ) + I( lArea * lNpk ) + I( lLabor * lNpk ) + mYear, + data = pdat, model = "random" ) > summary( riceTlTimeRan ) Oneway (individual) effect Random Effect Model (Swamy-Arora's transformation) Call: plm(formula = lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear, data = pdat, model = "random") Balanced Panel: n = 43, T = 8, N = 344 Effects: var std.dev share idiosyncratic 0.07530 0.27440 0.79 individual 0.01997 0.14130 0.21 theta: 0.434 Residuals: 369
9 Panel Data and Technological Change Min. 1st Qu. Median 3rd Qu. Max. -1.393176 -0.162097 0.045567 0.184209 0.798214 Coefficients: Estimate Std. Error z-value Pr(>|z|) (Intercept) 0.0213211 0.0347371 0.6138 0.539357 lArea 0.6831045 0.0922069 7.4084 1.278e-13 *** lLabor 0.0974523 0.0804060 1.2120 0.225511 lNpk 0.1708366 0.0546853 3.1240 0.001784 ** I(0.5 * lArea^2) -0.4275328 0.2468086 -1.7322 0.083230 . I(0.5 * lLabor^2) -0.6367899 0.2872825 -2.2166 0.026651 * I(0.5 * lNpk^2) 0.0307547 0.0957745 0.3211 0.748122 I(lArea * lLabor) 0.5666863 0.2059076 2.7521 0.005921 ** I(lArea * lNpk) 0.1037657 0.1421739 0.7299 0.465481 I(lLabor * lNpk) -0.2055786 0.1277476 -1.6093 0.107560 mYear 0.0142202 0.0070184 2.0261 0.042752 * --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Total Sum of Squares: 114.08 Residual Sum of Squares: 26.624 R-Squared: 0.76662 Adj. R-Squared: 0.75961 Chisq: 1093.86 on 10 DF, p-value: < 2.22e-16 The Translog production function cannot be estimated by a variable-coefficient model for panel model with our data set, because the number of time periods in the data set is smaller than the number of the coefficients. A pooled estimation can be done by > riceTlTimePool <- plm( lProd ~ lArea + lLabor + lNpk + + I( 0.5 * lArea^2 ) + I( 0.5 * lLabor^2 ) + I( 0.5 * lNpk^2 ) + + I( lArea * lLabor ) + I( lArea * lNpk ) + I( lLabor * lNpk ) + mYear, + data = pdat, model = "pooling" ) > summary(riceTlTimePool) Pooling Model Call: plm(formula = lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + 370
9 Panel Data and Technological Change I(lArea * lNpk) + I(lLabor * lNpk) + mYear, data = pdat, model = "pooling") Balanced Panel: n = 43, T = 8, N = 344 Residuals: Min. 1st Qu. Median 3rd Qu. Max. -1.521838 -0.181205 0.043555 0.222979 0.870190 Coefficients: Estimate Std. Error t-value Pr(>|t|) (Intercept) 0.0137557 0.0246454 0.5581 0.5771201 lArea 0.5880972 0.0851622 6.9056 2.542e-11 *** lLabor 0.1917638 0.0808764 2.3711 0.0183052 * lNpk 0.1978747 0.0516045 3.8344 0.0001505 *** I(0.5 * lArea^2) -0.4355466 0.2474913 -1.7598 0.0793520 . I(0.5 * lLabor^2) -0.7422415 0.3032362 -2.4477 0.0148916 * I(0.5 * lNpk^2) 0.0203673 0.0979072 0.2080 0.8353358 I(lArea * lLabor) 0.6786472 0.2165937 3.1333 0.0018822 ** I(lArea * lNpk) 0.0639200 0.1456135 0.4390 0.6609677 I(lLabor * lNpk) -0.1782859 0.1386111 -1.2862 0.1992559 mYear 0.0126820 0.0077947 1.6270 0.1046801 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Total Sum of Squares: 263.52 Residual Sum of Squares: 33.761 R-Squared: 0.87189 Adj. R-Squared: 0.86804 F-statistic: 226.623 on 10 and 333 DF, p-value: < 2.22e-16 This gives the same estimated coefficients as the model estimated by lm: > all.equal( coef( riceTlTime ), coef( riceTlTimePool ) ) [1] TRUE A Hausman test can be used to check the consistency of the random-effects estimator: > phtest( riceTlTimeRan, riceTlTimeFe ) Hausman Test 371
9 Panel Data and Technological Change data: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + ... chisq = 66.071, df = 10, p-value = 2.528e-10 alternative hypothesis: one model is inconsistent The Hausman test clearly rejects the consistency of the random-effects estimator. The following command tests the poolability of the model: > pooltest( riceTlTimePool, riceTlTimeFe ) F statistic data: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + ... F = 3.7469, df1 = 42, df2 = 291, p-value = 1.525e-11 alternative hypothesis: unstability The pooled model (riceCdTimePool) is clearly rejected in favor of the model with fixed individual effects (riceCdTimeFe), i.e. the individual effects are statistically significant. The following commands test if the fit of Translog specification is significantly better than the fit of the Cobb-Douglas specification: > waldtest( riceCdTimeFe, riceTlTimeFe ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + mYear Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear Res.Df Df Chisq Pr(>Chisq) 1 297 2 291 6 39.321 6.191e-07 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > waldtest( riceCdTimeRan, riceTlTimeRan ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + mYear Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear 372
9 Panel Data and Technological Change Res.Df Df Chisq Pr(>Chisq) 1 339 2 333 6 30.077 3.8e-05 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > waldtest( riceCdTimePool, riceTlTimePool ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + mYear Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear Res.Df Df Chisq Pr(>Chisq) 1 339 2 333 6 30.89 2.66e-05 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 The Cobb-Douglas functional form is rejected in favor of the Translog functional for for all three panel-specifications that we estimated above. The Wald test for the pooled model differs from the Wald test that we did in section 9.1.2.1, because waldtest by default uses a finite sample F statistic for models estimated by lm but uses a large sample Chi-squared statistic for models estimated by plm. The test statistic used by waldtest can be specified by argument test. 9.1.3 Translog production function with non-constant and non-neutral technological change Technological change is not always constant and is not always neutral (unbiased). Therefore, it might be more suitable to estimate a production function that can account for increasing or decreasing rates of technological change as well as biased (e.g. labor saving) technological change. This can be done by including a quadratic time trend and interaction terms between time and input quantities: ln y=α0+X i αiln xi+1 2X iX j αij ln xiln xj+αtt+X i αti tln xi+1 2αtt t2(9.7) In this specification, the rate of technological change depends on the input quantities and the time period: ∂ln y ∂t =αt+X i αti ln xi+αtt t(9.8) 373
9 Panel Data and Technological Change and the output elasticities might change over time: ϵi=∂ln y ∂ln xi =αi+X j αij ln xj+αti t. (9.9) 9.1.3.1 Pooled estimation of a translog production function with non-constant and non-neutral technological change The following command estimates a Translog production function that can account for nonconstant rates of technological change as well as biased technological change: > riceTlTimeNn <- lm( lProd ~ lArea + lLabor + lNpk + + I( 0.5 * lArea^2 ) + I( 0.5 * lLabor^2 ) + I( 0.5 * lNpk^2 ) + + I( lArea * lLabor ) + I( lArea * lNpk ) + I( lLabor * lNpk ) + + mYear + I( mYear * lArea ) + I( mYear * lLabor ) + I( mYear * lNpk ) + + I( 0.5 * mYear^2 ), data = riceProdPhil ) > summary( riceTlTimeNn ) Call: lm(formula = lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear + I(mYear * lArea) + I(mYear * lLabor) + I(mYear * lNpk) + I(0.5 * mYear^2), data = riceProdPhil) Residuals: Min 1Q Median 3Q Max -1.54976 -0.17245 0.04623 0.21624 0.87075 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) 0.001255 0.031934 0.039 0.96867 lArea 0.579682 0.085892 6.749 6.73e-11 *** lLabor 0.187505 0.081359 2.305 0.02181 * lNpk 0.207193 0.052130 3.975 8.67e-05 *** I(0.5 * lArea^2) -0.468372 0.265363 -1.765 0.07849 . I(0.5 * lLabor^2) -0.688940 0.308046 -2.236 0.02599 * I(0.5 * lNpk^2) 0.055993 0.099848 0.561 0.57533 I(lArea * lLabor) 0.676833 0.223271 3.031 0.00263 ** I(lArea * lNpk) 0.082374 0.151312 0.544 0.58654 I(lLabor * lNpk) -0.226885 0.145568 -1.559 0.12005 mYear 0.008746 0.008513 1.027 0.30497 I(mYear * lArea) 0.003482 0.028075 0.124 0.90136 374
9 Panel Data and Technological Change I(mYear * lLabor) 0.034661 0.029480 1.176 0.24054 I(mYear * lNpk) -0.037964 0.020355 -1.865 0.06305 . I(0.5 * mYear^2) 0.007611 0.007954 0.957 0.33933 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 0.3184 on 329 degrees of freedom Multiple R-squared: 0.8734, Adjusted R-squared: 0.868 F-statistic: 162.2 on 14 and 329 DF, p-value: < 2.2e-16 We conduct a Wald test to test whether the Translog production function with non-constant and non-neutral technological change outperforms the Cobb-Douglas production function and the Translog production function with constant and neutral technological change: > waldtest( riceCdTimeS, riceTlTimeNn ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + mYear Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear + I(mYear * lArea) + I(mYear * lLabor) + I(mYear * lNpk) + I(0.5 * mYear^2) Res.Df Df F Pr(>F) 1 339 2 329 10 3.488 0.00022 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > waldtest( riceTlTime, riceTlTimeNn ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear + I(mYear * lArea) + I(mYear * lLabor) + I(mYear * lNpk) + I(0.5 * mYear^2) Res.Df Df F Pr(>F) 375
9 Panel Data and Technological Change 1 333 2 329 4 0.9976 0.4089 The fit of the Translog specification with non-constant and non-neutral technological change is significantly better than the fit of the Cobb-Douglas specification but it is not significantly better than the fit of the Translog specification with constant and neutral technological change. In order to simplify the calculation of the output elasticities (with equation 9.9) and the annual rates of technological change (with equation 9.8), we create shortcuts for the estimated coefficients: > a1 <- coef( riceTlTimeNn )[ "lArea" ] > a2 <- coef( riceTlTimeNn )[ "lLabor" ] > a3 <- coef( riceTlTimeNn )[ "lNpk" ] > at <- coef( riceTlTimeNn )[ "mYear" ] > a11 <- coef( riceTlTimeNn )[ "I(0.5 * lArea^2)" ] > a22 <- coef( riceTlTimeNn )[ "I(0.5 * lLabor^2)" ] > a33 <- coef( riceTlTimeNn )[ "I(0.5 * lNpk^2)" ] > att <- coef( riceTlTimeNn )[ "I(0.5 * mYear^2)" ] > a12 <- a21 <- coef( riceTlTimeNn )[ "I(lArea * lLabor)" ] > a13 <- a31 <- coef( riceTlTimeNn )[ "I(lArea * lNpk)" ] > a23 <- a32 <- coef( riceTlTimeNn )[ "I(lLabor * lNpk)" ] > a1t <- at1 <- coef( riceTlTimeNn )[ "I(mYear * lArea)" ] > a2t <- at2 <- coef( riceTlTimeNn )[ "I(mYear * lLabor)" ] > a3t <- at3 <- coef( riceTlTimeNn )[ "I(mYear * lNpk)" ] Now, we can use the following commands to calculate the partial output elasticities: > riceProdPhil$eArea <- with( riceProdPhil, + a1 + a11 * lArea + a12 * lLabor + a13 * lNpk + a1t * mYear ) > riceProdPhil$eLabor <- with( riceProdPhil, + a2 + a21 * lArea + a22 * lLabor + a23 * lNpk + a2t * mYear ) > riceProdPhil$eNpk <- with( riceProdPhil, + a3 + a31 * lArea + a32 * lLabor + a33 * lNpk + a3t * mYear ) We can calculate the elasticity of scale by taken the sum over all partial output elasticities: > riceProdPhil$eScale <- with( riceProdPhil, eArea + eLabor + eNpk ) We can visualize (the variation of) the output elasticities and the elasticity of scale with histograms: > hist( riceProdPhil$eArea, 15 ) > hist( riceProdPhil$eLabor, 15 ) > hist( riceProdPhil$eNpk, 15 ) > hist( riceProdPhil$eScale, 15 ) 376
9 Panel Data and Technological Change eArea Frequency −0.5 0.0 0.5 1.0 0 20 40 60 eLabor Frequency −0.5 0.0 0.5 1.0 1.5 0 20 40 eNpk Frequency −0.1 0.1 0.3 0.5 0 20 40 eScale Frequency 0.8 1.0 1.2 1.4 0 20 60 Figure 9.1: Output elasticities and elasticities of scale 377
9 Panel Data and Technological Change Finally, we test whether the fit of Translog specification with non-constant and non-neutral technological change is significantly better than the fit of Translog specification with constant and neutral technological change: > waldtest( riceTlTimeNnFe, riceTlTimeFe ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear + I(mYear * lArea) + I(mYear * lLabor) + I(mYear * lNpk) + I(0.5 * mYear^2) Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear Res.Df Df Chisq Pr(>Chisq) 1 287 2 291 -4 2.3512 0.6715 > waldtest( riceTlTimeNnRan, riceTlTimeRan ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear + I(mYear * lArea) + I(mYear * lLabor) + I(mYear * lNpk) + I(0.5 * mYear^2) Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear Res.Df Df Chisq Pr(>Chisq) 1 329 2 333 -4 3.6633 0.4535 > waldtest( riceTlTimeNnPool, riceTlTimePool ) Wald test Model 1: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear + I(mYear * lArea) + I(mYear * lLabor) + I(mYear * lNpk) + I(0.5 * mYear^2) 384
9 Panel Data and Technological Change Model 2: lProd ~ lArea + lLabor + lNpk + I(0.5 * lArea^2) + I(0.5 * lLabor^2) + I(0.5 * lNpk^2) + I(lArea * lLabor) + I(lArea * lNpk) + I(lLabor * lNpk) + mYear Res.Df Df Chisq Pr(>Chisq) 1 329 2 333 -4 3.9905 0.4073 The tests indicate that the fit of Translog specification with constant and neutral technological change is not significantly worse than the fit of Translog specification with non-constant and non-neutral technological change. The difference between the Wald tests for the pooled model and the Wald test that we did in section 9.1.3.1 is explained at the end of section 9.1.2.2. 9.2 Frontier production functions with technological change The frontier production technology can be estimated by many different specifications of the stochastic frontier model. We will focus on four specifications that are all nested in the general specification: ln ykt = ln f(xkt, t, k)−ukt +vkt,(9.10) where the subscript k= 1, . . . , K indicates the firm, t= 1, . . . , T indicates the time period, and all other variables are defined as before. We will apply the following four model specifications: 1. the same frontier for all firms, i.e., f(xkt, t, k) = f(xkt, t)∀k, t, and time-invariant individual efficiencies, i.e., ukt =uk∀k, t, which means that each firm has an individual fixed efficiency that remains constant over time; 2. the same frontier for all firms, i.e., f(xkt, t, k) = f(xkt, t)∀k, t, and time-variant individual efficiencies, with ukt =ukexp(−η(t−T)) ∀k, t, which means that each firm has an individual efficiency and the inefficiency terms ukt of all firms can change over time with the same rate (and in the same direction) as indicated by the additional coefficient η; 3. the same frontier for all firms, i.e., f(xkt, t, k) = f(xkt, t)∀k, t, and observation-specific efficiencies, i.e., no restrictions on ukt, which means that the efficiency term of each observation is estimated independently from the other efficiencies of the firm so that basically the panel structure of the data is ignored; and 4. a different frontier for each firm, i.e., f(xkt, t, k) = f(xkt, t)eδk∀k, t (i.e., individual fixed effects in the frontier), and observation-specific efficiencies, i.e., no restrictions on ukt. 9.2.1 Cobb-Douglas production frontier with technological change We will use the specification in equation (9.2). 385
9 Panel Data and Technological Change 9.2.1.1 Time-invariant individual efficiencies We start with estimating a Cobb-Douglas production frontier with time-invariant individual efficiencies. The following commands estimate two Cobb-Douglas production frontiers with timeinvariant individual efficiencies, the first does not account for technological change, while the second does: > riceCdSfaInv <- sfa( lProd ~ lArea + lLabor + lNpk, data = pdat ) > summary( riceCdSfaInv ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged Iterative ML estimation terminated after 9 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 0.182636 0.035340 5.1680 2.366e-07 *** lArea 0.453900 0.064382 7.0501 1.788e-12 *** lLabor 0.288922 0.063853 4.5248 6.045e-06 *** lNpk 0.227542 0.040644 5.5983 2.164e-08 *** sigmaSq 0.155377 0.024144 6.4354 1.232e-10 *** gamma 0.464317 0.088270 5.2602 1.439e-07 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -86.43042 panel data number of cross-sections = 43 number of time periods = 8 total number of observations = 344 thus there are 0 observations not in the panel mean efficiency: 0.8187935 > riceCdTimeSfaInv <- sfa( lProd ~ lArea + lLabor + lNpk + mYear, data = pdat ) > summary( riceCdTimeSfaInv ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) 386
9 Panel Data and Technological Change The dependent variable is logged Iterative ML estimation terminated after 11 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 0.1832751 0.0345895 5.2986 1.167e-07 *** lArea 0.4625174 0.0644245 7.1792 7.011e-13 *** lLabor 0.3029415 0.0641323 4.7237 2.316e-06 *** lNpk 0.2098907 0.0418709 5.0128 5.364e-07 *** mYear 0.0116003 0.0071758 1.6166 0.106 sigmaSq 0.1556806 0.0242951 6.4079 1.475e-10 *** gamma 0.4706143 0.0869549 5.4122 6.227e-08 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -85.0743 panel data number of cross-sections = 43 number of time periods = 8 total number of observations = 344 thus there are 0 observations not in the panel mean efficiency: 0.8176333 In the Cobb-Douglas production frontier that accounts for technological change, the monotonicity conditions are globally fulfilled and the (constant) output elasticities of land, labor and fertilizer are 0.463, 0.303, and 0.21, respectively. The estimated (constant) annual rate of technological progress is around 1.2%. However, both the t-test for the coefficient of the time trend and a likelihood ratio test give rise to doubts whether the production technology indeed changes over time (P-values around 10%): > lrtest( riceCdTimeSfaInv, riceCdSfaInv ) Likelihood ratio test Model 1: riceCdTimeSfaInv Model 2: riceCdSfaInv #Df LogLik Df Chisq Pr(>Chisq) 1 7 -85.074 387
9 Panel Data and Technological Change 2 6 -86.430 -1 2.7122 0.09958 . --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Further likelihood ratio tests show that OLS models are clearly rejected in favor of the corresponding stochastic frontier models (no matter whether the production frontier accounts for technological change or not): > lrtest( riceCdSfaInv ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 5 -104.91 2 6 -86.43 1 36.953 6.051e-10 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > lrtest( riceCdTimeSfaInv ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 6 -104.103 2 7 -85.074 1 38.057 3.434e-10 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 This model estimates only a single efficiency estimate for each of the 43 firms. Hence, the vector returned by the efficiencies method only has 43 elements by default: > length( efficiencies( riceCdSfaInv ) ) [1] 43 One can obtain the efficiency estimates for each observation by setting argument asInData equal to TRUE: > pdat$effCdInv <- efficiencies( riceCdSfaInv, asInData = TRUE ) Please note that the efficiency estimates for each firm still do not vary between time periods. 388
9 Panel Data and Technological Change 9.2.1.2 Time-variant individual efficiencies Now we estimate a Cobb-Douglas production frontier with time-variant individual efficiencies. Again, we estimate two Cobb-Douglas production frontiers, the first does not account for technological change, while the second does: > riceCdSfaVar <- sfa( lProd ~ lArea + lLabor + lNpk, + timeEffect = TRUE, data = pdat ) > summary( riceCdSfaVar ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged Iterative ML estimation terminated after 11 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 0.182016 0.035251 5.1635 2.424e-07 *** lArea 0.474919 0.066213 7.1726 7.360e-13 *** lLabor 0.300094 0.063872 4.6983 2.623e-06 *** lNpk 0.199461 0.042740 4.6669 3.058e-06 *** sigmaSq 0.129957 0.021098 6.1598 7.285e-10 *** gamma 0.369639 0.104045 3.5527 0.0003813 *** time 0.058909 0.030863 1.9087 0.0563017 . --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -84.55036 panel data number of cross-sections = 43 number of time periods = 8 total number of observations = 344 thus there are 0 observations not in the panel mean efficiency of each year 12345678 0.7848433 0.7950303 0.8048362 0.8142652 0.8233226 0.8320146 0.8403483 0.8483313 mean efficiency: 0.817874 389
9 Panel Data and Technological Change > riceCdTimeSfaVar <- sfa( lProd ~ lArea + lLabor + lNpk + mYear, + timeEffect = TRUE, data = pdat ) > summary( riceCdTimeSfaVar ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged Iterative ML estimation terminated after 13 iterations: log likelihood values and parameters of two successive iterations are within the tolerance limit final maximum likelihood estimates Estimate Std. Error z value Pr(>|z|) (Intercept) 0.1817471 0.0360859 5.0365 4.741e-07 *** lArea 0.4761177 0.0657003 7.2468 4.267e-13 *** lLabor 0.2987917 0.0647805 4.6124 3.981e-06 *** lNpk 0.1991399 0.0428877 4.6433 3.429e-06 *** mYear -0.0031907 0.0155009 -0.2058 0.83692 sigmaSq 0.1255592 0.0295753 4.2454 2.182e-05 *** gamma 0.3478660 0.1507342 2.3078 0.02101 * time 0.0711165 0.0674356 1.0546 0.29162 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -84.52871 panel data number of cross-sections = 43 number of time periods = 8 total number of observations = 344 thus there are 0 observations not in the panel mean efficiency of each year 12345678 0.7780285 0.7905809 0.8025753 0.8140187 0.8249202 0.8352907 0.8451431 0.8544916 mean efficiency: 0.8181311 In the Cobb-Douglas production frontier that accounts for technological change, the monotonicity conditions are globally fulfilled and the (constant) output elasticities of land, labor and fertilizer are 0.476, 0.299, and 0.199, respectively. The estimated (constant) annual rate of technological change is around -0.3%, which indicates technological regress. However, the t-test for the 390
9 Panel Data and Technological Change coefficient of the time trend and a likelihood ratio test indicate that the production technology (frontier) does not change over time, i.e. there is neither technological regress nor technological progress: > lrtest( riceCdTimeSfaVar, riceCdSfaVar ) Likelihood ratio test Model 1: riceCdTimeSfaVar Model 2: riceCdSfaVar #Df LogLik Df Chisq Pr(>Chisq) 1 8 -84.529 2 7 -84.550 -1 0.0433 0.8352 A positive sign of the coefficient η(named time) indicates that efficiency is increasing over time. However, in the model without technological change, the t-test for the coefficient ηand the corresponding likelihood ratio test indicate that the effect of time on the efficiencies only is significant at the 10% level: > lrtest( riceCdSfaInv, riceCdSfaVar ) Likelihood ratio test Model 1: riceCdSfaInv Model 2: riceCdSfaVar #Df LogLik Df Chisq Pr(>Chisq) 1 6 -86.43 2 7 -84.55 1 3.7601 0.05249 . --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 In the model that accounts for technological change, the t-test for the coefficient ηand the corresponding likelihood ratio test indicate that the efficiencies do not change over time: > lrtest( riceCdTimeSfaInv, riceCdTimeSfaVar ) Likelihood ratio test Model 1: riceCdTimeSfaInv Model 2: riceCdTimeSfaVar #Df LogLik Df Chisq Pr(>Chisq) 1 7 -85.074 2 8 -84.529 1 1.0912 0.2962 391
9 Panel Data and Technological Change Finally, we can use a likelihood ratio test to simultaneously test whether the technology and the technical efficiencies change over time: > lrtest( riceCdSfaInv, riceCdTimeSfaVar ) Likelihood ratio test Model 1: riceCdSfaInv Model 2: riceCdTimeSfaVar #Df LogLik Df Chisq Pr(>Chisq) 1 6 -86.430 2 8 -84.529 2 3.8034 0.1493 All together, these tests indicate that there is no significant technological change, while it remains unclear whether the technical efficiencies significantly change over time. In econometric estimations of frontier models, where one variable (e.g. time) can affect both the frontier and the efficiency, the two effects of this variable can often be hardly separated, because the corresponding parameters can be simultaneous adjusted with only marginally reducing the log-likelihood value. This can be checked by taking a look at the correlation matrix of the estimated parameters: > round( cov2cor( vcov( riceCdTimeSfaVar ) ), 2 ) (Intercept) lArea lLabor lNpk mYear sigmaSq gamma time (Intercept) 1.00 0.18 -0.12 0.01 0.06 0.44 0.47 -0.19 lArea 0.18 1.00 -0.68 -0.39 -0.06 0.04 0.06 0.07 lLabor -0.12 -0.68 1.00 -0.27 0.08 -0.07 -0.09 0.00 lNpk 0.01 -0.39 -0.27 1.00 0.01 0.02 0.00 -0.11 mYear 0.06 -0.06 0.08 0.01 1.00 0.71 0.70 -0.88 sigmaSq 0.44 0.04 -0.07 0.02 0.71 1.00 0.94 -0.85 gamma 0.47 0.06 -0.09 0.00 0.70 0.94 1.00 -0.85 time -0.19 0.07 0.00 -0.11 -0.88 -0.85 -0.85 1.00 The estimate of the parameter for technological change (mYear) is highly correlated with the estimate of the parameter that indicates the change of the efficiencies (time). Again, further likelihood ratio tests show that OLS models are clearly rejected in favor of the corresponding stochastic frontier models: > lrtest( riceCdSfaVar ) Likelihood ratio test Model 1: OLS (no inefficiency) 392
9 Panel Data and Technological Change Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 5 -104.91 2 7 -84.55 2 40.713 4.489e-10 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > lrtest( riceCdTimeSfaVar ) Likelihood ratio test Model 1: OLS (no inefficiency) Model 2: Error Components Frontier (ECF) #Df LogLik Df Chisq Pr(>Chisq) 1 6 -104.103 2 8 -84.529 2 39.149 9.85e-10 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 In case of time-variant efficiencies, the efficiencies method returns a matrix, where each row corresponds to one of the 43 firms and each column corresponds to one of the 0 time periods: > dim( efficiencies( riceCdSfaVar ) ) [1] 43 8 One can obtain a vector of efficiency estimates for each observation by setting argument asInData equal to TRUE: > pdat$effCdVar <- efficiencies( riceCdSfaVar, asInData = TRUE ) 9.2.1.3 Observation-specific efficiencies In this section, we estimate a Cobb-Douglas production frontier with observation-specific efficiencies. The following commands estimate two Cobb-Douglas production frontiers, the first does not account for technological change, while the second does: > riceCdSfa <- sfa( lProd ~ lArea + lLabor + lNpk, data = riceProdPhil ) > summary( riceCdSfa ) Error Components Frontier (see Battese & Coelli 1992) Inefficiency decreases the endogenous variable (as in a production function) The dependent variable is logged 393
9 Panel Data and Technological Change I(log(labor) * log(npk)) -0.1370538 0.1407360 -0.9738 0.3301377 mYear 0.0151111 0.0069164 2.1848 0.0289024 * sigmaSq 0.2217092 0.0251305 8.8223 < 2.2e-16 *** gamma 0.8835549 0.0367095 24.0688 < 2.2e-16 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 log likelihood value: -74.40992 cross-sectional data total number of observations = 344 mean efficiency: 0.7294192 In the Translog production frontier that accounts for constant and neutral technological change, the monotonicity conditions are fulfilled at the sample mean and the estimated output elasticities of land, labor and fertilizer are 0.531, 0.231, and 0.203, respectively, at the sample mean. The estimated (constant) annual rate of technological progress is around 1.5%. A likelihood ratio test confirms the t-test for the coefficient of the time trend, i.e. the production technology (frontier) significantly changes over time: > lrtest( riceTlTimeSfa, riceTlSfa ) Likelihood ratio test Model 1: riceTlTimeSfa Model 2: riceTlSfa #Df LogLik Df Chisq Pr(>Chisq) 1 13 -74.410 2 12 -76.954 -1 5.0884 0.02409 * --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Two further likelihood ratio tests indicate that the Translog specification is superior to the CobbDouglas specification, no matter whether the two models allow for technological change or not. > lrtest( riceTlSfa, riceCdSfa ) Likelihood ratio test Model 1: riceTlSfa Model 2: riceCdSfa #Df LogLik Df Chisq Pr(>Chisq) 400
[Document text truncated for crawler view.]