Endogeneity in stochastic frontier models with 'wrong' skewness: copula approach without external instruments
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Haschka, Rouven E. Article — Published Version Endogeneity in stochastic frontier models with 'wrong' skewness: copula approach without external instruments Statistical Methods & Applications Provided in Cooperation with: Springer Nature Suggested Citation: Haschka, Rouven E. (2024) : Endogeneity in stochastic frontier models with 'wrong' skewness: copula approach without external instruments, Statistical Methods & Applications, ISSN 1613-981X, Springer, Berlin, Heidelberg, Vol. 33, Iss. 3, pp. 807-826, https://doi.org/10.1007/s10260-024-00750-4 This Version is available at: https://hdl.handle.net/10419/315072 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) Statistical Methods & Applications (2024) 33:807–826 https://doi.org/10.1007/s10260-024-00750-4 1 3 ORIGINAL PAPER Endogeneity instochastic frontier models with’wrong’ skewness: copula approach withoutexternal instruments RouvenE.Haschka1,2 Accepted: 25 February 2024 / Published online: 3 April 2024 © The Author(s) 2024 Abstract Stochastic frontier models commonly assume positive skewness for the inefficiency term. However, when this assumption is violated, efficiency scores converge to unity. The potential endogeneity of model regressors introduces another empirical challenge, impeding the identification of causal relationships. This paper tackles these issues by employing an instrument-free estimation method that extends joint estimation through copulas to handle endogenous regressors and skewness issues. The method relies on the Gaussian copula function to capture dependence between endogenous regressors and composite errors with a simultaneous consideration of positively or negatively skewed inefficiency. Model parameters are estimated through maximum likelihood, and Monte Carlo simulations are employed to evaluate the performance of the proposed estimation procedures in finite samples. This research contributes to the stochastic frontier models and production economics literature by presenting a flexible and parsimonious method capable of addressing wrong skewness of inefficiency and endogenous regressors simultaneously. The applicability of the method is demonstrated through an empirical example. Keywords Stochastic frontier analysis· Skewness· Endogenous regressors· Copula function· Maximum likelihood JEL Classification C13· C14· C21· C51 * Rouven E. Haschka [email protected]; [email protected] 1 Chair ofBusiness Analytics & Data Science, Zeppelin University, Am Seemoser Horn 20, 88045Friedrichshafen, Germany 2 Institute ofStrategy andManagement, Corvinus University, Fővám tér 8, 1093Budapest, Hungary
808 R.E.Haschka 1 3 1 Introduction The classical assumption in stochastic frontier (SF) production models is that inefficiency exhibits positive skewness, resulting in the composite error, i.e., the regression residuals, having negative skewness. However, in empirical applications, the residuals may exhibit positive skewness, which contradicts the assumption of positively skewed inefficiency. Waldman (1982) first demonstrated that if the residuals from the SF model have ‘wrong’ skewness, i.e., positive, inefficiency variance is biased towards zero. Consequently, efficiency scores tend to be one, leading to false conclusions of high efficiency (Hafner etal. 2018). Green and Mayes (1991) argue that this either indicates ‘super efficiency’ (all firms in the industry operate close to the frontier) or the inappropriateness of the SF analysis technique to measure inefficiencies. Thus, implausibly high efficiency scores obtained under the classical SF specification can indicate misspecification of inefficiency skewness (Haschka and Wied 2022). Another significant empirical challenge arises in the case of regressor endogeneity. Endogeneity can be loosely defined as dependence between regressors and errors, which is particularly important for SF models, as this dependence may stem from inefficiency, idiosyncratic noise, or both (Tran and Tsionas 2015). Linking endogenous covariate information with composite errors can lead to biased estimates for causal effects if the methods used assume regressor exogeneity (Haschka and Herwartz 2022, 2020). The standard approach to handling the endogeneity problem in SF models is to use likelihood-based instrumental variable (IV) estimation methods (Amsler etal. 2016; Kutlu 2010; Prokhorov etal. 2020; Tran and Tsionas 2013). However, a general drawback of such methods is their reliance on the availability of external information to construct instruments. Instruments, if available at all, are often subject to potential pitfalls if they fail to adequately meet the two required conditions: they must be sufficiently correlated with the endogenous regressors and uncorrelated with the composite errors term. Thus, a potential difficulty in implementing IV-based estimators arises when there is no external information available to construct appropriate instruments (Haschka etal. 2020). Given that suitable instrumental information in SF models is often scant, unavailable, or weak, this study proposes an SF model with a data-driven choice of correct or wrong skewness, which we conceptualize as an extension of the IVfree joint estimation model using copulas introduced by Park and Gupta (2012). Copula techniques have been successfully adopted for classical SF models (Tran and Tsionas 2015), and we extend them to address cases of wrong skewness. The method relies on a copula function to directly model dependencies between endogenous regressors and composite errors, thus eliminating the need for external information. Specifically, copulas allow for the separate modeling of marginal distributions of endogenous regressors and composite errors, while capturing their dependency. We construct the joint distribution of the endogenous regressor and composite error, which accommodates mutual dependency between them. Subsequently, we use this joint distribution to derive the likelihood function,
809 1 3 Endogeneity instochastic frontier models with’wrong’… which distinguishes between correct and wrong skewness through an indicator function, and maximize it to obtain consistent estimates of model parameters. The empirical analysis aims to impartially explore the determinants of firm performance based on data from 16,641 Vietnamese firms in 2015. A key observation is that the presence of wrong skewness is compounded by endogeneity, while the suggested estimator provides an unbiased perspective. The empirical findings indicating wrong skewness imply a growing number of inefficient firms persisting in the market, challenging the assumption of a competitive landscape. This persistence in inefficiency is likely influenced by factors such as corruption and the constraints imposed by the communist regime, hindering the establishment of liberal and competitive market conditions. Consequently, policy interventions are deemed necessary to incentivize firms to optimize their processes and improve efficiency. The paper is organized as follows. Section2 presents the model and discusses the copula approach to deal with regressor endogeneity in SF models with wrong skewness. In Sect.3, we examine the finite sample performance of the proposed approach through Monte Carlo simulations. An empirical application is provided in Sect.4. We revisit the methodology in Sects.5, and 6 concludes with a summary of our contribution. 2 Methodology Consider the standard SF model given by: Here, yi represents the output of producer i, xi is an L×1 vector of exogenous inputs, zi is a K×1 vector of endogenous inputs, 𝛽 and 𝛿 are L×1 and K×1 vectors of unknown parameters to be estimated, respectively. Additionally, vi is a symmetric random error, ui is a one-sided random disturbance representing technical efficiency, and the composite error is denoted as ei=vi−ui . It is assumed that xi is uncorrelated with ei , while zi is allowed to be correlated with ei , giving rise to endogeneity, but we make no assumption regarding the source of endogeneity, being it correlation with vi or with ui . Furthermore, it is assumed that ui and vi are independent, and the skewness of ui is left unrestricted. This discussion can be readily extended to cases where the (exogenous) environmental variables are included in the distribution of ui (Battese and Coelli 1995; Haschka and Herwartz 2022). 2.1 ‘Wrong’ skewness ofinefficiency distribution Adhering to standard SF practices, we presume that vi ∼N(0, 𝜎 2 v) represents symmetric, two-sided idiosyncratic noise. In this context, our model operates under the assumption that skewness issues arise from the inefficiency term, while vi itself is (1) y i=x � i𝛽+z � i𝛿+vi−ui ⏟⏟⏟ e i ,i=1, …,n ,
810 R.E.Haschka 1 3 symmetric.1 Following the approach of Hafner etal. (2018), we delineate two cases for ui that characterise the distributional shape: The assumption in (2) is widely acknowledged, characterising the correct skewness of ui (and consequently ei ) due to the strictly decreasing density of ui in the interval [0, ∞) (Kumbhakar and Lovell 2003), and constitutes the classical SF model. On the contrary, wrong skewness is induced by (3), where a0≈1.389 represents the nontrivial solution of 𝜙 (0) Φ( 0)=a0+𝜙 ( a0 ) −𝜙(0) Φ ( a 0) −Φ(0 ) , and the density of ui is strictly increasing and bounded in [0, a0|𝛾|] . It is noteworthy that the expectations of both ui and ei remain unaffected by the skewness’ sign (Hafner etal. 2018). Thus, the inefficiency variance and the sign of skewness are directly linked because 𝛾>0 ( 𝛾<0 ) induces correct (wrong) skewness, but 𝔼[ui] and 𝔼[ei] are not influenced by the sign of 𝛾 . By convolution of v with u and integration, the density of ei=vi−ui obtains as: with 𝜎2 =𝛾 2 +𝜎 2 v and ∫ eg + e (e)de = ∫ eg − e (e) de . It is important to note that under correct skewness, e follows a (1x1)-dimensional closed skew normal (CSN) distribution, denoted as e+ ∼CSN(0, 𝜎 2 ,− 𝛾 𝜎v𝜎 ,0,1 ) , while, as demonstrated by Haschka and Wied (2022), under wrong skewness, e follows a (1x2)-dimensional CSN distribution, denoted as e −∼CSN1×2 ( a0𝛾,𝜎2, ( 𝛾∕𝜎 −𝛾∕𝜎 ) , ( −a0𝜎 0 ) , ( 𝜎2 v0 0𝜎2 v)) . Con- (2) ‘Correct’ skewness: u i ∼N [0,∞) (0, 𝛾2),𝛾> 0 (3) ‘Wrong’ skewness: u i ∼N [0,a0|𝛾|) (a 0| 𝛾 | ,𝛾2),𝛾< 0 (4) ‘Correct’ skewness: g+ e(e)= 2 𝜎 𝜙 ( e 𝜎) Φ ( −e𝛾 𝜎𝜎v), (5) ‘Wrong’ skewness: g− e(e)= 1 𝜎(Φ(a0)− Φ(0))𝜙(e−a0𝛾 𝜎)[Φ(Aw+a0𝜎 𝜎v)−Φ (Aw)] , Aw=e−a0𝛾 𝜎 𝛾 𝜎v , 1 Within the SF literature, there are discussions on various sources of wrong skewness. While it is commonly attributed to the inefficiency term (as in our case), asymmetry of idiosyncratic noise (Badunenko and Henderson 2023; Bonanno etal. 2017; Horrace etal. 2023; Son etal. 1993), or dependence between noise inefficiency and idiosyncratic noise (Smith 2008; Bonanno et al. 2017; Bonanno and Domma 2022) is also considered.
811 1 3 Endogeneity instochastic frontier models with’wrong’… sequently, it becomes evident that another advantage of the distribution in (3) is its association with the well-known CSN distribution. As a result, properties of the CSN distribution, such as moments, can be derived (Flecher etal. 2009). Although Li (1996) argues that one-sided distributions with an unbounded range always exhibit positive skewness, Johnson etal. (1995) demonstrates that the twoparameter Weibull distribution can have both positive and (small) negative skewness for specific parameter combinations. Accordingly, this distribution could be an alternative without needing to impose an upper boundary. However, since the upper bound of the distribution in (3) depends on 𝛾 , the advantage of using the truncated normal distribution in (3) is that it entails negative skewness without necessitating the identification of additional parameters. Moreover, other parsimonious distributions for u that can produce wrong skewness, such as the negative skew-exponential distribution discussed in Hafner et al. (2018), may also be considered. However, in such cases, it would not be clear if the distribution of e will be known or has a closed-form expression. The configuration of the inefficiency distribution under both correct and wrong skewness is illustrated in Panel (a) of Fig.1, with the corresponding distributions of composite errors presented in Panel (b). When 𝛾>0 , the distribution of u exhibits positive skewness, while the distribution of e has negative skewness. Conversely, for 𝛾<0 , the skewness of u becomes negative, and that of e is positive. It is important to emphasize that the adopted one-sided distribution for inefficiency, which can result in both correct and wrong skewness, is parsimonious. Alternative approaches allowing for a data-driven selection of either correct or wrong skewness often involve identifying multiple parameters determining the inefficiency distribution (see, e.g., Tsionas 2017, for Weibull inefficiency), may not result in a well-known distribution for composite errors, or require a priori determination of the skewness sign (corrected OLS or modified OLS). However, such determinations are challenging in the presence of endogeneity. 2.2 Joint estimation using copulas Consider F(z1,…,zK,e) and f(z1,…,zK,e) as the joint distribution and joint density of endogenous regressors (z1,…,zK) and composite errors e, respectively. In Fig. 1 Densities of u and e for 𝛾 = 2.5 , i.e., correct skewness (dotted lines) and 𝛾 =− 2.5 , i.e., wrong skewness (solid lines); with 𝜎v = .5
812 R.E.Haschka 1 3 practical applications, F( ⋅ ) and f( ⋅ ) are unknown due to the unobservability of e and thus require estimation. In line with the methodology proposed by Park and Gupta (2012), we employ a copula approach to approximate this joint density. The copula serves as a tool to capture dependence within the joint distribution of endogenous regressors and composite errors. Let 𝝎z,i =(F z1 (z 1i ),…,F zK (z Ki )) � and 𝜔e,i=G(ei;𝜎v,𝛾) represent the margins ( 𝝎 z,i , 𝜔e,i ) � ∈[0, 1] K+1 based on a probability integral transform. Here, the F’s signify the respective marginal cumulative distribution functions of the observed endogenous regressors, and G(ei;𝜎v,𝛾) is the cumulative distribution function of the CSN distribution for errors, subject to the sign of 𝛾 . Drawing on the approach by Haschka (2021) and Tran and Tsionas (2015), we substitute F1( z 1i) ,…,F p( z pi) with their respective empirical counterparts in a first stage. Given observed samples of zji , j=1, …,p ; i=1, …,n , we utilise the empirical cumulative distribution function (ecdf) of zj , denoted as F j= 1 n+1∑n i=1 1 � zji ≤ z0j � . Utilising a Gaussian copula, 𝝃z,i = (Φ−1( F z1 (z 1i )),…,Φ−1( F zK (zKi))) � , and 𝜉e,i =Φ −1 (G(e i ;𝜎 v ,𝛾 )) , both follow a standard multivariate normal distribution with a dimension of (K+1) and a correlation matrix Ξ . Subsequently, the joint density can be expressed as where 𝜉e,i and g(ei;𝜎v,𝛾) are again subject to either correct or wrong skewness. The copula density in the first row establishes a connection between the error and all endogenous variables. Meanwhile, the densities in the second row describe the marginal behaviour. The marginal densities f zk( z ki) in (6) do not involve any parameters of interest and can be omitted when deriving the likelihood since they function as normalising constants. To simultaneously determine the choice of correct or wrong skewness based on the sign of 𝛾 , we adopt the approach of Hafner etal. (2018) and incorporate an indicator function into the likelihood. While Hafner etal. (2018) propose choosing correct or wrong skewness a priori by examining the skewness of the OLS residual, our method estimates the sign of 𝛾 simultaneously with all other parameters. This is preferred because any pre-determination of residual skewness could be influenced by (potential) endogeneity. Consequently, the likelihood function is given by: (6) f(zi,ei)= 1 √det(Ξ) exp � −1 2 � 𝝃z,i 𝜉e,i � � �Ξ−1−I� � 𝝃z,i 𝜉e,i �� ×g(ei;𝜎v,𝛾)× K � k=1 fzk � zki � ,
813 1 3 Endogeneity instochastic frontier models with’wrong’… To explicitly account for the scenario of only fully efficient firms, the likelihood also accommodates 𝛾=0 .2 In this case, the marginal distribution of the errors follows a normal distribution with mean zero and variance 𝜎2 v , represented as 𝜉0 e,i =ei∕𝜎 v . Notably, our approach encompasses those of Hafner etal. (2018), Tran and Tsionas (2015), and Park and Gupta (2012). Specifically, under the assumption of exogeneity of all regressors ( Ξ=I ), the likelihood in (7) reduces to that in Hafner etal. (2018). In the case of correct skewness ( 𝛾>0 ), corresponding to the traditional SF model, the likelihood collapses to that in Tran and Tsionas (2015). Finally, for the scenario of only fully efficient firms ( 𝛾=0 ), it reduces to that in Park and Gupta (2012). Subsequently, the likelihood is logarithms and maximised with respect to the vector of unknown parameters 𝜽=(𝛽,𝛿,𝜎2v,𝛾, vechl[Ξ]) , where vechl[Ξ] = (𝜌1,…,𝜌K)� stacks the lower diagonal elements of the correlation matrix Ξ into a column vector. With parameter estimates, technical inefficiency ui can be predicted using Jondrow etal. (1982) as follows: (7) L (𝜽∣y,z,x)∝1(𝛾>0) n � i=1 1 √det(Ξ) exp �−1 2� 𝝃z,i 𝜉+ e,i�� �Ξ−1−I�� 𝝃z,i 𝜉+ e,i��g+(ei;𝜎v,𝛾) +1(𝛾<0) n � i=1 1 √det(Ξ) exp �−1 2� 𝝃z,i 𝜉− e,i�� �Ξ−1−I�� 𝝃z,i 𝜉− e,i��g−(ei;𝜎v,𝛾 ) +1(𝛾=0) n � i =1 1 √ det(Ξ) exp � −1 2 � 𝝃z,i 𝜉0 e,i � � � Ξ−1−I �� 𝝃z,i 𝜉0 e,i �� 𝜙(ei;𝜎v). (8) ‘Correct’ skewness: ui= E�ui∣ei�=�𝛾 2+𝜎 2 v(𝛾 ∕𝜎 v) 1+(𝛾 ∕𝜎 v)2⎡ ⎢ ⎢ ⎢ ⎣ 𝜙�(𝛾 ∕𝜎 v)ei∕√𝛾 +𝜎 v)� 1−Φ �(𝛾 ∕𝜎 v)ei∕√𝛾 +𝜎 v) � −(𝛾 ∕𝜎 v)ei � 𝛾 2+𝜎 2 v) ⎤ ⎥ ⎥ ⎥ ⎦ (9) ‘Wrong’ skewness: ui= E ( ui∣ei ) = ∫ a0 | 𝛾 | 0 exp{−ui}f−(ui∣ei)du i 2 In case of exogeneity, the likelihood function is continuous in 𝛾 , as shown in Appendix A.1 of Hafner etal. (2018). Intuitively, this should also hold for nonzero diagonal elements in Ξ .
814 R.E.Haschka 1 3 Note that in the case of correct skewness, the predicted technical inefficiency has a closed-form expression, while under wrong skewness, the integral needs to be solved numerically (Hafner etal. 2018). 2.2.1 Model identifiability In our framework, model identification hinges on two key conditions: (i) the distribution of endogenous regressors must differ from that of the composite error (for an in-depth discussion on identification in copula-based endogeneity-correction models, see Haschka 2022b; Papadopoulos 2022; Park and Gupta 2012). Consequently, the model maintains identification as long as 𝛾 is not zero (or very close to zero), and endogenous regressors deviate from a normal distribution. However, identification collapses if 𝛾=0 (resulting in a normal composite error, in which case our model aligns with that of Park and Gupta 2012), and endogenous regressors follow a normal distribution. In such a scenario, the joint distribution of endogenous regressors and the composite error becomes multivariate normal, implying a linear relationship between the endogenous regressors and the composite error. Consequently, we would be unable to distinguish the linear effect of the endogenous regressor on the outcome. To address this, a formal separation of 𝔼[y|z] from 𝔼[e|z] without IV information is impossible if both are linear functions, as indicated by joint normality (Haschka 2022b). In contrast, when the joint distribution is not multivariate normal (with 𝛾=0 , indicating the non-normality of endogenous regressors), the relationship between regressors and errors becomes non-linear. Identification in this context does not require IVs and can be accomplished through the joint distribution of e and z. Although 𝔼[y|z] remains a linear function due to the specification of a linear regression model, non-normality implies that 𝔼[e|z] becomes a non-linear function. This non-linearity facilitates the separation of variation attributed to endogenous regressors from the variation due to the composite error, as discussed in Tran and Tsionas (2015). Consequently, in empirical applications, it is essential to assess the marginal distribution of endogenous regressors before estimation-an approach commonly adopted in empirical literature utilizing copula-based identification (e.g., Haschka 2022b; Papadopoulos 2022; Park and Gupta 2012). Through the utilization of a Gaussian copula, model identification also requires (ii) a linear dependence between 𝝃z and 𝜉e , enabling the correlation to be expressed through pairwise Pearson coefficients. Consequently, the model exclusively accommodates (iii) continuous endogenous regressors (or discrete with numerous distinct outcomes). In instances where zk is continuous, 𝜉z,k = (Φ −1 (F zk (z k))) follows a standard normal distribution. Conversely, if zk is binary, 𝜉z,k would also be binary, and multivariate normality with 𝜉e would not be guaranteed. While the continuity of zk is crucial and its verification is straightforward, confirming the assumption of Gaussian-type dependence is more challenging empirically due to the unobservability of errors. However, since any copula capable of modeling multivariate dependency structures can be employed, this identification assumption can be readily substituted when opting for a different copula. Despite this, in cases where the true dependence deviates from the Gaussian copula assumption, existing literature has demonstrated
821 1 3 Endogeneity instochastic frontier models with’wrong’… employ one-year lagged assets and one-year lagged wages as instruments, following a similar approach used in previous studies (Haschka and Herwartz 2022). However, it is important to note that such internal instrumentation may suffer from weak instruments and might not be entirely suitable for handling endogeneity. Table3 presents the estimation results, which consistently indicate human capital as the primary driver of firm performance in Vietnam across all employed estimators. This is evident from the significantly higher coefficient associated with log wages compared to log assets. The consideration of endogeneity through GMM, Copula, and the proposed estimator reduces the observed difference in coefficients. GMM and Copula estimators may still face remaining endogeneity issues when wrong skewness is present, with GMM encountering additional challenges due to potential weak instrumentation. MLE indicates increasing returns to scale, as reflected in the sum of the elasticities associated with wages and assets being above one. However, for Copula and GMM, this sum significantly falls below one, suggesting decreasing returns to scale. This implies that scaling up is challenging for firms. By contrast, the proposed estimator yields constant returns to scale, which aligns with the dataset dominated by small firms (O’Toole and Newman 2017). The discrepancy raises concerns about the MLE results being flawed due to endogeneity. Further evidence supporting endogeneity is found in the significant estimates of correlations between production inputs and errors when using Copula and the proposed estimators. In terms of firm efficiency, MLE yields surprisingly high mean efficiency, with an average score of .7280. However, accounting for endogeneity while assuming correct skewness through GMM and Copula estimators decreases mean efficiency scores to .4169 (GMM) and .3903 (Copula). The proposed approach reveals a mean efficiency of .6126. While GMM and Copula estimators exhibit minimal differences among all coefficients, the results undergo significant changes when the proposed estimator is applied. Notably, the proposed Table 3 Estimation outcomes are presented employing MLE (Hafner etal. 2018), GMM (Tran and Tsionas 2013), copula-based estimation (Tran and Tsionas 2015), and our proposed estimator. Standard errors for copula-based estimators (copula and proposed) are determined through bootstrap procedures with 1,999 replications. Efficiency scores are computed utilising the estimator by Jondrow etal. (1982) MLE GMM Copula Proposed Est SE Est SE Est SE Est SE const 1.497 .0625 1.241 .0798 1.147 .0802 .9822 .0803 log wages .9229 .0120 .6440 .0204 .6024 .0209 .6833 .0214 log assets .2163 .0099 .2515 .0193 .2904 .0205 .3102 .0211 𝜎v 1.082 .0142 .9914 .0295 1.194 .0308 .8993 .0301 𝛾 .4151 .0312 1.241 .0396 1.290 .0411 -.5809 .0411 𝜌e,log wages .3025 .0403 .2627 .0401 𝜌e,log assets .1790 .0371 .2298 .0366 Mean 𝔼 [ u|e] .3174 .8747 .9407 .4900 Mean Efficiency .7280 .4169 .3903 .6126
822 R.E.Haschka 1 3 estimator indicates the presence of wrong skewness after accounting for endogeneity, which is challenging for MLE to detect under such conditions. The insights derived from the proposed estimator contribute evidence supporting the existence of endogenous regressors and wrong skewness. While the former has been previously emphasized in empirical development literature, the latter has not received recognition. To comprehend the economic reasons and market mechanisms behind the occurrence of wrong skewness, it is crucial to delve into the contributing factors. As illustrated in Panel (a) of Fig.1, the presence of correct skewness (dotted line) indicates that most firms should operate near the efficiency frontier, aligning with competitive market conditions where inefficient firms are likely to exit due to a lack of competitiveness. Conversely, empirical evidence supporting wrong skewness (straight line) implies a growing number of inefficient firms persisting in the market, contradicting the assumption of a competitive market situation. This suggests a lack of incentives for firms to optimize efficiency, attributed to factors such as widespread corruption and the constraints imposed by the communist regime in Vietnam. Despite ongoing reforms, these challenges persist, necessitating policy interventions to incentivize firms to optimize processes and improve efficiency, as market forces alone seem insufficient to generate such incentives. 5 Discussion ofthemethodology The proposed approach offers several advantages that make it useful for handling both endogeneity and skewness issues. Firstly, it overcomes the need for instrumental variables, eliminating the challenge of obtaining and validating such instruments. By employing a copula function to directly connect endogenous regressors and errors, the model exhibits parsimony and requires the identification of only one additional parameter for each endogenous regressor - the correlation coefficient depicting regressor-error dependence. This parsimony is further bolstered by the one-parameter inefficiency distribution, which can accommodate both correct and wrong skewness with only one parameter to estimate. The sign of this parameter determines the skewness. Additionally, the likelihood function is readily obtained even under wrong skewness, as the errors follow a CSN distribution for which certain properties can be derived (Flecher etal. 2009). This simplifies the estimation process and enhances the model’s computational efficiency. The proposed approach does have certain limitations, which stem from either the SF specifications or the copula function employed for endogeneity correction. On the one hand, while we attribute wrong skewness to the inefficiency component, other potential sources, such as asymmetry in idiosyncratic noise or dependence between idiosyncratic noise and inefficiency, could also be responsible. Additionally, we restrict our attention to the (truncated) half-normal distribution for inefficiency. While this choice simplifies the analysis by ensuring a CSN distribution after convolution with the normal distribution assumed for idiosyncratic noise, alternative distributions for inefficiency could be considered, such as the negative skew exponential distribution (Hafner etal. 2018). While the choice of appropriate distribution assumptions is a key consideration in any SF model, in the context of wrong
823 1 3 Endogeneity instochastic frontier models with’wrong’… skewness, we would also need to establish whether the density of e=v−u , resulting from the convolution of v with u, follows a known parametric distribution. On the other hand, while the Gaussian copula exhibits considerable flexibility, particularly when dealing with multiple endogenous regressors as in the empirical application, it is rooted in the assumption of linear dependency between its margins. However, prior studies have demonstrated the robustness of the Gaussian copula in capturing various non-Gaussian dependencies (Haschka 2022b; Haschka and Herwartz 2022; Park and Gupta 2012). By virtue, any copula-based endogeneity correction generally disregards the potential sources of endogeneity, be it omitted variables, reverse causality, or simultaneity, as they are employed to address the symptoms of endogeneity. Regarding the SF specification, it is further complicated by the fact of discerning whether endogeneity arises from correlation with the inefficiency term or idiosyncratic noise because we model the joint distribution of endogenous regressors and composite errors. Moreover, the proposed approach only allows for continuous endogenous variables. Since model identification necessitates marginal distributions of endogenous regressors to be different from that of the composite errors (which follow a CSN distribution), empirical applications demand an a priori assessment of the marginal distribution of explanatory variables. For instance, in cases where the model incorporates only fully efficient firms, the endogenous regressors must exhibit a non-normal distribution. While instrumental variable estimation remains the preferred method for addressing endogeneity when strong and valid instruments are available, weak instrumentation or skewness misspecifications pose empirical challenges to such fully parametric approaches. Nevertheless, we expect that the proposed method offers valuable alternatives for numerous empirical SF models that encounter regressor endogeneity and/or skewness issues. 6 Conclusion Traditional stochastic frontier (SF) models typically assume that inefficiency follows a half-normal distribution with positive skewness. However, when true inefficiency exhibits negative skewness, efficiency scores are biased toward one, leading to misleading conclusions of high efficiency (Waldman 1982). While recent literature has highlighted the importance of addressing the ’wrong’ skewness problem in SF analysis (Curtiss etal. 2021; Choi etal. 2021; Daniel etal. 2019), existing studies have not considered the potential endogeneity of regressors. This paper fills this gap by proposing an instrument-free approach for estimating SF models with endogenous regressors and allowing for a simultaneous choice of inefficiency skewness. Building upon the work of Park and Gupta (2012), we employ a copula function to directly construct the joint density of endogenous regressors and composite errors, enabling us to capture mutual dependency without the need for instrumental variables. Our model distinguishes between correct and wrong skewness without imposing a priori restrictions on the sign of inefficiency skewness or requiring the identification of additional parameters governing its direction. We evaluate the finite sample performance of the approach through Monte Carlo simulations. The
824 R.E.Haschka 1 3 simulation results demonstrate that the estimator performs well in finite samples, exhibiting desirable properties in terms of bias and mean squared error. These findings are further validated in an empirical application. The following contributions of this article are also worth mentioning. On the methodological front, we advance the understanding of copula-based endogeneity corrections in scenarios with non-Gaussian outcomes. Joint estimation using copulas still occupies a niche in endogeneity-robust modeling (Papies etal. 2023; Papadopoulos 2021). Our work contributes to bridging this gap by understanding copulabased endogeneity corrections in SF settings. The existing methodological literature about dealing with skewness issues in SF models predominantly addresses the symptoms but fails to delve into the underlying causes, often attributing skewness issues to weak samples (Almanidis and Sickles 2011; Hafner etal. 2018; Simar and Wilson 2009). Empirically, Papadopoulos and Parmeter (2023) note that only two studies in the past 25 years have considered skewness issues in SF analyses. However, these studies remain silent on the potential economic explanations for skewness. Our study breaks new ground by attributing skewness to the inefficiency component of the SF model, providing an economic explanation for this phenomenon. Regarding the empirical application, this study is the first to explicitly address and account for skewness issues when applying SF models in a development economics context. Previous studies on firm growth in Vietnam using SF analysis, such as Haschka etal. (2023), have overlooked potential skewness issues in their findings. By incorporating an economic explanation for wrong skewness and proposing an endogeneity-robust methodology, we offer a novel perspective on growth dynamics in Vietnam and shed light on underlying competition levels. Funding Open access funding provided by Corvinus University of Budapest. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/ licenses/by/4.0/. References Almanidis P, Sickles RC (2011) The skewness issue in stochastic frontiers models: fact or fiction? In: van Keilegom I, Wilson PW (eds) Exploring research frontiers in contemporary statistics and econometrics. Springer, Berlin, pp 201–227 Amsler C, Prokhorov A, Schmidt P (2016) Endogeneity in stochastic frontier models. J Econom 190(2):280–288 Badunenko O, Henderson DJ (2023) Production analysis with asymmetric noise. J Product Anal 61:1–18 Battese GE, Coelli TJ (1995) A model for technical inefficiency effects in a stochastic frontier production function for panel data. Empir Econ 20(2):325–332
825 1 3 Endogeneity instochastic frontier models with’wrong’… Becker J-M, Proksch D, Ringle CM (2021) Revisiting gaussian copulas to handle endogenous regressors. J Acad Market Sci (forthcoming) Bonanno G, Domma F (2022) Analytical derivations of new specifications for stochastic frontiers with applications. Mathematics 10(20):3876 Bonanno G, De Giovanni D, Domma F (2017) The ‘wrong skewness’ problem: a re-specification of stochastic frontiers. J Product Anal 47(1):49–64 Breitung J, Mayer A, Wied D (2023) Asymptotic properties of endogeneity corrections using nonlinear transformations. Econom J, utae002 Chen Y-Y, Schmidt P, Wang H-J (2014) Consistent estimation of the fixed effects stochastic frontier model. J Econom 181(2):65–76 Choi K, Kang HJ, Kim C (2021) Evaluating the efficiency of Korean festival tourism and its determinants on efficiency change: parametric and non-parametric approaches. Tourism Manag 86:104348 Curtiss J, Jelínek L, Medonos T, Hruška M, Hüttel S (2021) Investors’ impact on Czech farmland prices: a microstructural analysis. Eur Rev Agric Econ 48(1):97–157 Daniel BC, Hafner CM, Simar L, Manner H (2019) Asymmetries in business cycles and the role of oil prices. Macroecon Dyn 23(4):1622–1648 Flecher C, Naveau P, Allard D (2009) Estimating the closed skew-normal distribution parameters using weighted moments. Stat Probab Lett 79(19):1977–1984 Genest C, Ghoudi K, Rivest L-P (1995) A semiparametric estimation procedure of dependence parameters in multivariate families of distributions. Biometrika 82(3):543–552 Green A, Mayes D (1991) Technical inefficiency in manufacturing industries. Econ J 101(406):523–538 Hafner CM, Manner H, Simar L (2018) The wrong skewness problem in stochastic frontier models: a new approach. Econom Rev 37(4):380–400 Haschka RE (2021) Exploiting between-regressor correlation to robustify copula correction models for handling endogeneity. SSRN Working Paper. https:// ssrn. com/ abstr act= 42228 08 Haschka RE (2022a) Bayesian inference for joint estimation models using copulas to handle endogenous regressors. SSRN Working Paper. https:// ssrn. com/ abstr act= 42351 94 Haschka RE (2023) Endogeneity-robust estimation of nonlinear regression models using copulas: a Bayesian approach with an application to demand modelling. SSRN Working Paper. https:// papers. ssrn. com/ sol3/ papers. cfm? abstr act_ id= 44515 91 Haschka RE, Wied D (2022) Estimating fixed effects stochastic frontier panel models under ‘wrong’ skewness with an application to health care efficiency in Germany. SSRN Working Paper. https:// papers. ssrn. com/ sol3/ papers. cfm? abstr act_ id= 40796 60 Haschka RE (2022) Handling endogenous regressors using copulas: A generalisation to linear panel models with fixed effects and correlated regressors. J Market Res 59(4):860–881 Haschka RE, Herwartz H (2020) Innovation efficiency in European high-tech industries: evidence from a Bayesian stochastic frontier approach. Res Policy 49:104054 Haschka RE, Herwartz H (2022) Endogeneity in pharmaceutical knowledge generation: an instrumentfree copula approach for Poisson frontier models. J Econ Manag Strategy 31(4):942–960 Haschka RE, Schley K, Herwartz H (2020) Provision of health care services and regional diversity in Germany: insights from a Bayesian health frontier analysis with spatial dependencies. Eur J Health Econ 21:55–71 Haschka RE, Herwartz H, Struthmann P, Tran VT, Walle YM (2021) The joint effects of financial development and the business environment on firm growth: evidence from Vietnam. J Comp Econ 50(2):486–506 Haschka RE, Herwartz H, Silva Coelho C, Walle YM (2023) The impact of local financial development and corruption control on firm efficiency in Vietnam: evidence from a geoadditive stochastic frontier analysis. J Product Anal 60(2):203–226 Horrace WC, Parmeter CF, Wright IA (2023) On asymmetry and quantile estimation of the stochastic frontier model. J Product Anal 61(1):19–36 Joe H, Xu JJ (1996) The estimation method of inference functions for margins for multivariate models. Technical Report: The University of British Columbia, Canada Johnson NL, Kotz S, Balakrishnan N (1995) Continuous univariate distributions, vol 2. Wiley, New York Jondrow J, Lovell CK, Materov IS, Schmidt P (1982) On the estimation of technical inefficiency in the stochastic frontier production function model. J Econom 19(2–3):233–238 Kumbhakar SC, Lovell CK (2003) Stochastic frontier analysis. Cambridge University Press, Cambridge Kutlu L (2010) Battese–Coelli estimator with endogenous regressors. Econ Lett 109(2):79–81
826 R.E.Haschka 1 3 Li Q (1996) Estimating a stochastic production frontier when the adjusted error is symmetric. Econ Lett 52(3):221–228 O’Toole C, Newman C (2017) Investment financing and financial development: Vidence from Vietnam. Rev Finance 21(4):1639–1674 Papadopoulos A (2021) Measuring the effect of management on production: a two-tier stochastic frontier approach. Empir Econ 60(6):3011–3041 Papadopoulos A (2022) Accounting for endogeneity in regression models using Copulas: a step-by-step guide for empirical studies. J Econom Methods 11(1):127–154 Papadopoulos A, Parmeter CF (2023) The wrong skewness problem in stochastic frontier analysis: a review. J Product Anal. https:// doi. org/ 10. 1007/ s1112302300708-w Papies D, Ebbes P, Feit EM (2023) Endogeneity and causal inference in marketing. In: Winder RS, Neslin SA (eds) The history of marketing science. World Scientific Publishing Co. Pte. Ltd., Singapore, pp 253–300 Park S, Gupta S (2012) Handling endogenous regressors by joint estimation using Copulas. Market Sci 31(4):567–586 Prokhorov A, Schmidt P (2009) Likelihood-based estimation in a panel setting: robustness, redundancy and validity of copulas. J Econom 153(1):93–104 Prokhorov A, Tran KC, Tsionas MG (2020) Estimation of semi-and nonparametric stochastic frontier models with endogenous regressors. Empir Econ 60:3043–3068 Simar L, Wilson PW (2009) Estimation and inference in cross-sectional, stochastic frontier models. Econom Rev 29(1):62–98 Smith MD (2008) Stochastic frontier models with dependent error components. Econom J 11(1):172–192 Son TVH, Coelli T, Fleming E (1993) Analysis of the technical efficiency of state rubber farms in Vietnam. Agric Econ 9(3):183–201 Tran KC, Tsionas MG (2021) Efficient semiparametric copula estimation of regression models with endogeneity. Econom Rev (forthcoming) Tran KC, Tsionas EG (2013) GMM estimation of stochastic frontier models with endogenous regressors. Econ Lett 118:233–236 Tran KC, Tsionas EG (2015) Endogeneity in stochastic frontier models: copula approach without external instruments. Econ Lett 133:85–88 Tsionas MG (2017) When, where, and how of efficiency estimation: improved procedures for stochastic frontier modeling. J Am Stat Assoc 112(519):948–965 Waldman DM (1982) A stationary point for the stochastic frontier likelihood. J Econom 18(2):275–279 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.