An application of mixture distributions in modelization of length of hospital stay
Abstract
Length of hospital stay (LOS) is an important indicator of the hospital activity and management of health care. The skewness exhibited by this variable poses problems in statistical modeling. The aim of this work is to model the variable LOS within diagnosis-related groups (DRG) through finite mixtures of distributions. A mixture of the union of Gamma, Weibull and Lognormal families is used in the model, instead of a mixture of a unique distribution family. Some theoretical questions regarding the model, such as the identifiability and study of asymptotic properties of ML estimators, are analyzed. The EM algorithm is proposed for performing these estimators. Finally, this new proposed model is illustrated by using data from different DRGs.
Full text
An application of mixture distributions in modelization of length of hospital stay N. Atienza1, ∗, †,J. Garc´ıa-Heras2,J. M.Mu˜noz-Pichardo2 and R. Villa3 1Departamento de Matem´atica Aplicada I, Universidad de Sevilla, Spain 2Departamento de Estad´ıstica e I.O, Universidad de Sevilla, Spain 3Departamento de An´alisis Matem´atico, Universidad de Sevilla, Spain SUMMARY Length of hospital stay (LOS) is an important indicator of the hospital activity and management of health care. The skewness exhibited by this variable poses problems in statistical modeling. The aim of this work is to model the variable LOS within diagnosis-related groups (DRG) through finite mixtures of distributions. A mixture of the union of Gamma, Weibull and Lognormal families is used in the model, instead of a mixture of a unique distribution family. Some theoretical questions regarding the model, such as the identifiability and study of asymptotic properties of ML estimators, are analyzed. The EM algorithm is proposed for performing these estimators. Finally, this new proposed model is illustrated by using data from different DRGs. KEY WORDS: length of stay; diagnosis-related groups; finite mixture model; Gamma distribution; Weibull distribution; Lognormal distribution 1. INTRODUCTION Length of stay (LOS) is an easily available indicator of hospital activity. It is used for various purposes, such as management of hospital care, quality control, appropriateness of hospital use and hospital planning [1, 2]. LOS is an indirect estimator of resources consumption and of the efficiency of one of the aspects of hospital patient care: bed management. Lee et al. [3] state that ‘comprehensive and accurate information about inpatient LOS should be a high priority for ∗Correspondence to: N. Atienza, Departamento de Matem´ atica Aplicada I, Universidad de Sevilla, E.T.S. Ingenier´ ıa Inform´ atica. Avd. Reina Mercedes s/n. 41012-Sevilla, Spain. †E-mail: [email protected] Contract/grant sponsor: Ministerio de Educaci´ on y Ciencia; contract/grant number: MTM2004-01433 Contract/grant sponsor: DGES; contract/grant number: BFM2006-1409
health planners and administrators in the strategic planning and deployment of financial, human and physical resources’. Since the early 1980s, health-care systems in industrial countries have exhibited deep changes in order to reduce overspending on health care, particularly hospital expenditures. Among all the proposed modifications, the hospital stay classification given by Fetter et al. [4]received the widest approval for assessing hospital output. In this classification, the main diagnosis established at the end of the stay and the procedures performed during the hospitalization are used to retrospectively classify a patient into a diagnosis-related group (DRG) that determines the amount of payment allocated to the hospital. The U.S.A. Congress decided in October 1983 to implement a Medicare prospective system based on the classification of Fetter et al. [4]. Later on, many countries such as Australia or France introduced DRGs to reduce the health-care budget. The DRGs provide a classification system of episodes of hospitalization with clinically recognized definitions, where it is expected that patients in the same class consume similar quantities of resources, as a result of a process of similar hospital care. The mean of the LOS is used as an indicator of the consumption of resources because of its availability and good relation with the raised costs. Hence, we may say that DRGs have been partially created in order to get homogeneous groups with respect to the consumption of services and costs, closely related to the LOS. Moreover, the comparison of the expected LOS by DRG raises conclusions of a given aspect of the management for each specific kind of patients. The expectation is only a localization measure of the variable LOS. Therefore, a more complete study of the distribution of this variable is convenient. To this aim, in the past years, many authors proposed different technics to analyze the variable LOS. However, the empirical distribution of LOS is well established to be positively skewed, plurimodal, to contain outliers and to significantly vary between DRGs [5, 6], etc. This heterogeneity of LOS poses a problem for statistical analysis, limiting the use of inference techniques based on normality assumptions. Since a large number of DRGs must be analyzed routinely, automatic procedures are needed for conveniently treating skewness. Different transformations (e.g. the logarithmic one) of LOS have been attempted to attain normality, and subsequently to apply the corresponding tests [5, 7, 8], etc.). However, as Xiao et al.[9]show, these approaches rely on the unrealistic homogeneity assumption on the entire sample. Marazzi et al. [6]assessed the adequacy of three conventional parametric models, Lognormal (long-tailed), Weibull and Gamma (short-tailed), for describing the LOS distribution. But, as Lee et al. [10]point out, none of them seemed to fit satisfactorily in a wide variety of samples. The main issue is that the assumption of heterogeneous sub-populations would be more appropriate than single DRG populations. Mixture distribution analysis can clarify whether or not a skewed distribution is composed of heterogeneous components [9, 11]. With this new focus, several works model the LOS variable through finite mixture distributions. For example, Quantin et al. [12] analyzed the cost and the hospital stay of DRG 589 and 590 (lymphoma and leukemia with or without complications) using mixtures of Weibull distributions. The same model was also applied by Quantin et al. [11]to analyze DRG 316 (renal failure). A Gamma mixture was proposed to analyze heterogeneity of maternity LOS in Lee et al. [10, 13]. The model was applied to different DRGs related to cesarean delivery (DRG 670,671,672 and 687) and to DRGs related to vaginal delivery (DRG 674,675,676 and 688). It is then clear that the proposed mixture model depends on which DRG or group of DRGs is considered. The objective of this work is to analyze the variable LOS unifying the approach given by Marazzi and this new focus, tackling the problem of the hospital stay through a finite mixture of
distributions in the union of Gamma, Weibull and Lognormal families. This new posing can be used to model the variable LOS for most of the DRGs. The results obtained in this work prove the validity of this new approach. The paper is organized as follows. Section 2 contains a theoretical exposition of the mixture distributions: the identifiability problem, the asymptotic properties of the maximum likelihood (ML) estimators and the adaptation of the EM algorithm, needed for the parametric estimation. In Section 3, a simulation study is presented to analyze the performance of our method. We apply this method in Section 4 to several DRGs with the aim of illustrating the goodness of fit of the proposed model. Finally, we have included a short discussion on the proposed methodology. 2. METHODS A finite mixture distribution can be considered as a convex combination of distribution functions, which can be applied to model two different practice situations, according to Titterington et al. [14]. One in the case of a variable with several underlying categories or with different sources of origin (direct application), and the other, when these categories do not exist or are not physically interpretable. In this case, the finite mixture model is used as a statistical tool to provides more flexibility in the fittings and get better results (indirect application). This work deals with the second case. We have obtain a sample of the LOS variables obtained from patients in a given DRG. Moreover, our distributions have specific parametric forms as follows. Let FL,FGand FWbe, respectively, the Lognormal, Gamma and Weibull distributions families: FL=F:F(x;,)=x 0 1 √2uexp −1 2log u− 2du;∈R,>0,x>0 FG=F:F(x;a,b)=x 0 b−a (a)ua−1exp −u bdu;a,b>0,x>0 FW=F:F(x;c,d)=x 0 c dcuc−1exp −uc dcdu;c,d>0,x>0 We propose to model the LOS variable with a threefold mixture distribution: Q(x;)=1F1(x;,)+2F2(x;a,b)+3F3(x;c,d) where 1,2,3>0, 1+2+3=1, F1∈FL,F2∈FG,F3∈FWand =(1,2,,,a,b,c,d) is the parametric vector determinate by the given distributions. Let us denote this class of mixture distributions by C. Obviously, the main problem is the estimation of the parameters. However, before solving this problem, it is necessary to tackle the problem of identifiability, the question on the unique representation of the class of models being considered. Generally, the problem of estimation of parameters is undertaken through the ML method, because of its desirable statistical properties, such as efficiency, consistency and asymptotic normality under some uniform integrability assumptions on the mixture and its derivatives.
However, despite the good properties of the ML estimators, in our case it is not possible to obtain explicit solutions. Hence, numerical methods must be applied. The EM algorithm, proposed by Dempster et al. [15], is the more used method. Next, we study these three aspects in the finite mixtures considered: identifiability, asymptotic properties and the EM algorithm. 2.1. Identifiability Identifiability problems concerning finite and countable mixtures have been studied widely. Teicher [16]gave a sufficient condition for a finite mixture to be identifiable. This condition is based on the existence of a linear and one-to-one mapping defined on the distribution family and a total ordering on this family. As an application, he proved the identifiability of finite mixtures of normal distributions and of Gamma distributions by using the bilateral Laplace transform and the lexicographic order. This sufficient condition was subsequently modified by Chandra [17] using the moment-generating function of log X. This result was used by Khalaf [18]to prove the identifiability of finite mixtures of Weibull, Lognormal, Chi-squared and Pareto distributions. Different modifications with application to specific families of distributions have been proposed by several authors (see Brandorff-Nielsen [19]or Henna [20]). Atienza et al. [21]provide a sufficient condition for the identifiability of finite mixtures, which is applied to the class of all finite mixtures generated by the union of Lognormal, Gamma and Weibull distributions, where Teicher’s and Henna’s conditions are not applicable. Thus, if we denote by U=FL∪FG∪FWthe family obtained by the union of these three families, then the class HUof all finite mixtures of distributions from Uis identifiable. Obviously, since C⊂HU, the class Cis identifiable. 2.2. Asymptotic properties of ML estimators In finite mixture models, ML estimators have good properties (efficiency, consistency and asymptotic normality) under some uniform integrability assumptions on the mixture and its derivatives up to the third order. The problem of consistency has been the center of interest for many authors in the study of solutions for both general and specific families. The works of Chanda [22]and Redner and Walker [23]should be cited because of their importance and influence in subsequent works. In particular, these works establish, for a sample of size n, under two conditions, that there exists in any sufficiently small neighborhood of the value of the parameter ∗a unique strongly consistent solution nof the likelihood equations and that this solution at least locally maximizes the loglikelihood function and is asymptotically normally distributed. The first condition establishes the uniform integrability for the partial derivatives up to the third order of the mixture. The second condition deals with the positive-definite character of the Fisher information matrix. More precisely, the Fisher information matrix I()given by I()=[∇log q(x;)][∇log q(x;)]tq(x;)dx has to be well defined and positive definite at ∗, where ∇denotes the gradient of the first partial derivatives with respect to the components of . Atienza et al. [24]prove that under another two conditions related to the integrability of the components of the mixtures, for any compact subset of the parametric space which contains
∗, with probability 1, nis a ML estimator in for sufficiently large n. Furthermore, this ML estimator is unique in , that is, there is no other ML estimator besides nwhich leads to a limiting density different from q(x,∗). The integrability conditions are proved in Atienza et al. [25]are to be verified for the union of the so-called W-type families. Since the distribution families FL,FGand FWare W-type families, the authors establish that the class HUof all finite mixtures of distributions from U verifies the integrability conditions. The inclusion C⊂HUclearly implies that the class Calso verifies these conditions. It remains to prove that the second condition is also verified, that is, I()is well defined and positive definite at ∗. Applying Proposition 1 (see Appendix), it is easily proved that the information matrix of any element from the class Cis well defined. The definite positive character is obtained by application of Proposition 2. The density function of the mixture distribution can be expressed in the form q(x;)=1f1(x;)+2f2(x;)+(1−1−2)f3(x;) where f1(x;)=1 √2xexp −1 2log x− 2 f2(x;)=b−a (a)xa−1exp −x b and f3(x;)=c dcxc−1exp −xc dc with x>0, ,a,b,c,d>0and∈R. The set of partial derivatives with respect to the parameter is provided in Table I. In order to prove its linear independence character, the following partial order is defined on that set: g1≺g2⇔lim x→∞ g2(x) g1(x)=0andg1∼g2⇔lim x→∞ g2(x) g1(x)∈R\{0} The arrangement of the set of partial derivatives {wi(x;), i=1,...,8}is provided in Table II. Let us consider any linear combination 8 i=1iwi(x;)of this set. The linear independence of the set is proved by showing that the equality 8 i=1iwi(x;)=0 holds provided that all coefficients i’s are zero. Dividing by w4(x;)and taking limit x→∞ give 4=0. We may continue in this fashion for cases (1), (2) and (4) in Table II to obtain that the rest of the coefficients are null. In case (3), an analogous procedure yields 4=3=1=7=0. Dividing by w2(x;)and taking limit now give 2+8lim x→∞ w8(x;) w2(x;)=0
Table I. Partial derivatives of q(x;). Parameter Partial derivatives of q(x;) 1w1(x;)=f1(x;)−f3(x;) 2w2(x;)=f2(x;)−f3(x;) w3(x;)=1log x− 2f1(x;) w4(x;)=1(log x−)2 3−1 f1(x;) aw5(x;)=2[log x−log b−(a)/(a)]f2(x;) bw6(x;)=2x b2f2(x;) cw7(x;)=(1−1−2)1 c−log x d−x dclog x df3(x;) dw8(x;)=(1−1−2)xc−1 d−c xd f3(x;) Table II. Partial ordering depending on the values of the parameters. Case Arrangement (1) c<1w4≺w3≺w1≺w7≺w2≺w8≺w6≺w5 (2) (c>1)or (c=1,b>d)or (c=1,b=d,a>2)w 4≺w3≺w1≺w6≺w5≺w2≺w7≺w8 (3) (c=1,b<d)w 4≺w3≺w1≺w7≺w2≺w8≺w6≺w5 (4) (c=1,b=d,1<a<2)w 4≺w3≺w1≺w6≺w7≺w5≺w2≺w8 (5) (c=1,b=d,a<1)w 4≺w3≺w1≺w7≺w6≺w2≺w8≺w5 (6) (c=1,b=d,a=2)w 4≺w3≺w1≺w6≺w5≺w7≺w2≺w8 or, equivalently, 8=(b/(1−1−2))2. Consequently, 2w2(x;)+b 1−1−2 w8(x;)+6w6(x;)+5w5(x;)=0 According to the defined order, w2+b 1−1−2 w8≺w6≺w5 Dividing by w2+(b/(1−1−2))w8and proceeding analogously, we obtain the nullity of the rest of the coefficients. A similar reasoning can be applied to cases (5) and (6). Consequently, the information matrix is positive definite and, as a conclusion, we can assure that ML estimators of finite mixture models of Chave good asymptotic properties. 2.3. Estimation through EM algorithm In the previous subsections, we have analyzed the problem of identifiability and the properties of ML estimators. The aim of this section is to study their calculation. Mixtures of distributions under our study constitute a parametric family depending on eight parameters, and the calculation
of ML estimators needs to deal with a quite complex system of nonlinear equations. According to the recommendations of Mclachlan and Krishnan [26]we have tackled this problem by using the EM algorithm. We will just sketch the adaptation of this algorithm to our case. The first problem is the election of the initial values for ((0)), in order to assure the convergence with an acceptable speed to the biggest local maximum. It consists in the election of the initial weights of the component functions ((0) 1and (0) 2) and their parameters ((0),(0),a(0), b(0),c(0)and d(0)). For (0) 1and (0) 2, we propose the strategy suggested by Karlis and Xekalaki [27]: we take several initial values running over the range of possible values, 1,2∈(0,1),and choose those that reach the biggest likelihood. This strategy prevents falling in regions where the likelihood is nearly constant, and detect, if they exist, several local maxima and choose, consequently, the global maximum. For the initial values of the parameters of the component functions, we consider independently the ML estimators of each family. •Lognormal estimators: (0)=1 n n i=1 log xi,(0)=1 n n i=1 (log xi−(0))21/2 •Gamma estimators (following Greenwood and Durand [28]): a(0)=slog x xand b(0)=1 na(0) n i=1 xi where xand xare, respectively, the arithmetic and geometric means, and s(t)= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ 0.5000876 +0.1648852t−0.0544276t2 t,0⩽t<0.5772 8.898919 +9.05995t+0.9775373t2 t(17.79728 +11.968477t+t2),0.5772⩽t<17 1 t,17⩽t •Weibull estimators (applying Bain and Antle’s method [29]): c(0)= n i=1 x1/n (i)n i=1 1/nd(0) iand d(0)=n i=1log ilog x(i)−1 nn i=1log in i=1log x(i) n i=1log2x(i)−1 n(n i=1log x(i))2 where i=i j=11/(n−j+1)and (x(1),x(2),...,x(n))is the ordered sample. The second problem is to improve the low speed of convergency of the algorithm. We propose to use the conjugate gradient acceleration method [30].
Finally, the stop criterion is determined by the parameter variation between two consecutive steps, the variation measured by the sum of the absolute values of each variation parameter. Following Karlis and Xekalaki [27],thevalue10 −5gives an appropriate stop criterion. 3. SIMULATION This section presents a simulation study to analyze the performance of the proposed method. Two different strategies have been used: •Simulation A: For a set of 128 combinations of values of the parameters, 128 samples have been generated (with size n=100) and used to apply the proposed methodology. •Simulation B: Fifteen combinations have been selected from among the 128 previous ones, and 100 samples of size n=500 have been generated and used to apply the proposed methodology to each individual combination. The combinations of values of the parameters have been selected by fitting several samples of the length of hospital stay variable to each component function separately. The goal of Simulation A is to get conclusions about the adequacy of the proposed methodology in a wide range of possible values of the eight parameters of the model, so that the possible values related to the variable LOS are covered. Simulation B is carried out to analyze the performance of the proposed EM algorithm. 3.1. Simulation A The considered mixture model has eight parameters, =(,,a,b,c,d,1,2). We have considered two different values for each of them, obtaining 64 possible combinations: =2,3;=0.5,0.7;a=7,10;b=2,3;c=1,2;d=15,25 For each of these 64 combinations, we have randomly generated two pairs of values for the weights 1and 2. For a given parameter g=(g 1,g 2,g,g,ag,bg,cg,dg), the associated sample has been generated using the following procedure. Consider the intervals I1=[0,g 1],I2=(g 1,g 1+g 2]and I3=(g 1+g 2,1]; generate nrandom values {ui:i=1,...,n}from a uniform distribution on the interval [0,1], and for each value ui: •if ui∈I1then generate a random value xifrom a Lognormal distribution with parameters g and g; •if ui∈I2then generate a random value xifrom a Gamma distribution with parameters ag and bg; •if ui∈I3then generate a random value xifrom a Weibull distribution with parameters ag and bg. As a discrepancy measure between the original F(x,g)and the estimated distribution function F(x, ), the uniform measure [31], also called Kolmogorov measure [32], is proposed: d(g, )=sup x∈R|F(x,g)−F(x, )|
Figure 1. Discrepancy measure: histogram (Simulation A). Figure 1 represents the obtained results for the 128 generated samples, which should be considered satisfactory, since for 93.8 per cent of the samples the discrepancy measure is less than 0.03, and for none of them is this value bigger than 0.05. 3.2. Simulation B Fifteen among the 128 combinations generated in Simulation A are selected. For each of them: •the expected value () and the variance (2) of the associated distribution mixture are calculated. Both values are functions of the parameters =g1()and 2=g2(). •B=100 samples of size n=500 are generated. The following statistics are calculated for each sample: ◦The ML estimator of using the method proposed in this paper, { k:k=1,...,B}. ◦The expected value and the variance of the associated mixture distribution, {(k,2 k): k=1,...,B}, where k=g1( k)and 2 k=g2( k). ◦The absolute deviation b(k)=|k−|and the variance ratio r(2 k)=2 k/2. ◦The discrepancy measure d( k)=d( k,). •The following values have been calculated: estimated absolute bias (EAB), maximum deviation (MD), mean variance ratio (MVR), minimum variance ratio (MINVR), maximum variance
Using the bound q(x;)⩾hfh(x,hh), h=1,...,K(A1) yields fh(x,hh)fl(x,hl) q(x;)⩽1 h fl(x,hl) fh(x,hh)l q(x;) *fk(x,hl) *lr ⩽l h *fl(x,hl) *lr and both of the right-hand functions are integrable. Finally, using the inequality ab⩽(a2+b2)/2, a,b∈R, we obtain hl*fh(x,hh) *hs *fl(x,hl) *lr q(x;)⩽hl 2q(x;)*fh(x,hh) *hs 2 +*fl(x,hl) *lr 2 applying inequality (A1) again, we can bound it by 1 2⎡ ⎢ ⎢ ⎢ ⎣*fh(x,hh) *hs 2 fh(x,hh)+*fl(x,hl) *lr 2 fl(x,hl)⎤ ⎥ ⎥ ⎥ ⎦ which is integrable by hypothesis. Proposition 2 Let f(x;h)be a density function. Its Fisher information matrix in h0∈Rd,I(h0), is positive definite if and only if the set of functions {w1(x,h0),...,w d(x,h0)} is linear independent a.e. in the set Sf={x:f(x;h)=0}, where wi(x,h0)=(*/*i)f(x;h)|h=h0. Proof Let M(x;h)denote the matrix [∇hlog f(x;h)][∇hlog f(x;h)]t, with components mij(x;h)=*log f(x,h) *i*log f(x,h) *j Then, ftI(h)f=0 if and only if i,jijmij(x;h)f(x;h)=0 for almost all x∈Rp. It means that there exists a null set Z⊂Rpsuch that for x/∈Z,(ftM(x;h)f)f(x;h)=0. The latter is equivalent to ftM(x;h)f=0 for x∈Sf\Z. Now, this equality is equivalent to ((1/f(x;h))ft[∇hf(x;h)])2=0, which finally is iiwi(x;h)=0. Taking into account that ftI(h)f⩾0, the conclusion of the proposition follows.
ACKNOWLEDGEMENTS The first three authors have been supported by Ministerio de Educaci´ on y Ciencia of Spain, Plan Nacional I+D, Ref.MTM2004-01433. The last author thanks the DGES Spanish grant BFM2006-1409 for financial support. REFERENCES 1. Fernonw LC, Mccoll I, Mackie C. Firm, patient, and process variables associated with length of stay in four diseases. British Medical Journal 1978; 1:556–559. 2. Lagoe RJ. A community-based analysis of regional differences in hospital stay by DRGs. Inquiry 1986; 23(2): 183–190. 3. Lee AH, Gracey M, Wang K, Kelvin KW. A robustified modelling approach to analyze pediatric length of stay. Annals of Epidemiology 2005; 15(9):637–677. 4. Fetter RB, Youngsoo S, Freeman JL, Averill RF, Thomson JD. Case-mix; definition by diagnosis relates groups. Medical Care 1980; 18(Suppl. 2):1–53. 5. Shachtman RH, Snapinn SM, Quade D, Freund DA, Kronhaus AK. A method for constructing case-mix indexes, with application to hospital length of stay. Health Services Research 1986; 20:737–762. 6. Marazzi A, Paccaud F, Ruffieux C, Beguin C. Fitting the distributions of length of stay by parametric models. Medical Care 1998; 36(6):915–927. 7. Silberbach M, Shumaker D, Menashe V, Cobanoglu A, Morris C. Predicting hospital charge and length of stay for congenital heart disease surgery. American Journal of Cardiology 1993; 72:958–963. 8. Wolfe MW, Roubin GS, Schweiger M. Length of hospital stay and complications after percutaneous transluminal coronary angioplasty. Circulation 1995; 92:311–319. 9. Xiao J, Lee AH, Vemuri SR. Mixture distribution analysis of length of hospital stay for efficient funding. Socio-economic Planning Sciences 1999; 33(1):39–59. 10. Lee AH, Ng ASK, Yau KKW. Determinants of maternity length of stay: a gamma mixture risk-adjusted model. Health Care Management Science 2001; 4:249–255. 11. Quantin C, Sauleau E, Bolard P, Mousson C, Kerkri M, Brunet Lecomte P, Moreau T, Duserre L. Modeling of high-cost patient distribution within renal failure diagnosis related group. Journal of Clinical Epidemiology 1999; 52(3):251–258. 12. Quantin C, Entezam F, Brunet-Lecomte P, Lepage E, Guy H, Duserre L. High cost factors for leukaemia and lymphoma patients: a new analysis of costs within these diagnosis related groups. Journal of Epidemiology and Community Health 1999; 53:24–31. 13. Lee AH, Xiao J, Codde JP, Ng ASK. Public versus private hospital maternity length of stay: a gamma mixture modelling approach. Health Services Management Research 2002; 15(1):46–54. 14. Titterington DM, Smith AFM, Makov UE. Statistical Analysis of Finite Mixture Distributions. Wiley: New York, 1985. 15. Dempster AP, Laird NM, Rubin DB. Maximum likelihood from incomplete data via the EM algorithm (with discussion). Journal of the Royal Statistical Society Series B (Statistical Methodology)1977; 39:1–38. 16. Teicher H. Identifiability of finite mixtures. The Annals of Mathematical Statistics 1963; 38(4):1300–1302. 17. Chandra S. On the mixtures of probability distributions. Scandinavian Journal of Statistics 1977; 4:105–112. 18. Khalaf EA. Identifiability of finite mixtures using a new transform. Annals of the Institute of Statistical Mathematics 1988; 40:261–265. 19. Brandorff-Nielsen O. Identifiability of mixtures of exponential families. Journal of Mathematical Analysis and Applications 1965; 12(1):115–121. 20. Henna J. Examples of identifiable mixture. Journal of the Japan Statistical Society 1994; 24:193–200. 21. Atienza N, Garc´ ıa-Heras J, Mu˜ noz-Pichardo JM. A new condition for identifiability of finite mixture distributions. Metrika 2005; 63(2):215–221. 22. Chanda KC. A Note on the consistency and maxima of the roots of likelihood equations. Biometrika 1954; 41:56–61. 23. Redner RA, Walker HF. Mixture densities, maximum likelihood and the EM algorithm. SIAM Review 1984; 26:195–239. 24. Atienza N, Garc´ ıa-Heras J, Mu˜ noz-Pichardo JM, Villa R. On the consistency of MLE in finite mixture models of exponential families. Journal of Inference Statistics Planning 2006; 137(2):496–505. DOI: 10.1016/j.jspi.2005.12.014.
25. Atienza N, Garc´ ıa-Heras J, Mu˜ noz-Pichardo JM. Consistency of maximum likelihood estimators in finite mixture models of the union of W-type families. Communications in Statistics,Theory and Methods 2005; 34:1471–1485. 26. Mclachlan GJ, Krishnan T. The EM Algorithm and Extensions. Wiley: New York, 1997. 27. Karlis D, Xekalaki E. Choosing initial values for the EM algorithm for finite mixtures. Computational Statistics and Data Analysis 2002; 41:577–590. 28. Greenwood JA, Durand D. Aids for fitting the Gamma distribution by maximum likelihood. Technometrics 1960; 2:55–56. 29. Bain LJ, Antle CE. Estimation of parameters in the Weibull distribution. Technometrics 1967; 9:621–627. 30. Jamshidian M, Jenrich RI. Conjugate gradient acceleration of the EM algorithm. Journal of the American Statistical Association 1993; 88:221–228. 31. Zolotarev VM. Probability metrics. Theory of Probability and its Applications 1983; 28:278–302. 32. Kolmogorov AN. Sulla determinazione empririca di una legge di distribuzione. Giornalle dell’Instituto Italiano degli Attuari 1933; 4:83–91.