SURVIVAL REGRESSION MODELS FOR SINGLE EVENTS AND COMPETING RISKS BASED ON PSEUDOOBSERVATIONS
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Wycinka, Ewa; Jurkiewicz, Tomasz Article SURVIVAL REGRESSION MODELS FOR SINGLE EVENTS AND COMPETING RISKS BASED ON PSEUDOOBSERVATIONS Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Wycinka, Ewa; Jurkiewicz, Tomasz (2019) : SURVIVAL REGRESSION MODELS FOR SINGLE EVENTS AND COMPETING RISKS BASED ON PSEUDOOBSERVATIONS, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, NY, Vol. 20, Iss. 1, pp. 171-188, https://doi.org/10.21307/stattrans-2019-010 This Version is available at: https://hdl.handle.net/10419/207930 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/
STATISTICS IN TRANSITION new series, March 2019 171 STATISTICS IN TRANSITION new series, March 2019 Vol. 20, No. 1, pp. 171–188, DOI 10.21307/stattrans-2019-010 SURVIVAL REGRESSION MODELS FOR SINGLE EVENTS AND COMPETING RISKS BASED ON PSEUDOOBSERVATIONS Ewa Wycinka 1 , Tomasz Jurkiewicz 2 ABSTRACT Survival data is a special type of data that measures the time to an event of interest. The most important feature of survival data is the presence of censored observations. An observation is said to be right-censored if the time of the observation is, for some reason, shorter than the time to the event. If no censoring occurs in the data, standard statistical models can be used to analyse the data. Pseudo-observations can replace censored observations and thereby allow standard statistical models to be used. In this paper, a pseudo-observation approach was applied to single-event and competing-risks analysis, with special attention paid to the properties of the pseudo-observations. In the empirical part of the study, the use of regression models based on pseudo-observations in credit-risk assessment was investigated. Default, defined as a delay in payment, was considered to be the event of interest, while prepayment of credit was treated as a possible competing risk. Credits that neither default nor are prepaid during the follow-up were censored observations. Typical application characteristics of the credit and creditor were the covariates in the regression model. In a sample of retail credits provided by a Polish financial institution, regression models based on pseudo-observations were built for the single-event and competing-risks approaches. Estimates and discriminatory power of these models were compared to the Cox PH and Fine-Gray models. Key words: generalised estimating equations, cumulative incidence function, probability of default, credit risk, survival analysis. 1. Introduction In the past few decades, survival analysis methods have become more widely used, not only in biostatistics, where their roots are, but also in many other branches of science, including economics, and the social sciences. Survival analysis is a term that covers a vast collection of different methods that focus on timing and duration prior to an event’s occurrence (Mills, 2011). Among these methods are parametric and non-parametric estimation of survival time 1 University of Gdańsk, Faculty of Management. E-mail: ewa.w[email protected]. ORCID ID: https://orcid.org/0000-0002-5237-3488. 2 University of Gdańsk, Faculty of Management. E-mail: [email protected]. ORCID ID: https://orcid.org/0000-0001-7066-5196.
172 E. Wycinka, T. Jurkiewicz: Survival regression models… distributions, and parametric and semiparametric regression models. The common goal of these methods is to handle censored observations that are inevitable in time-to-event analysis. A quite new and innovative approach to the problem of censoring is the idea of pseudo-observations that can replace both complete and censored actual observations. Pseudo-observations can be applied to many different objectives; this paper focuses on the usefulness of pseudoobservations in the development of regression models for survival functions in the case that a single event is analysed and in the competing risk analysis. The first objective was to review the properties of pseudo-observations in these two situations. The second goal was to compare the results and performance of the regression models for pseudo-observations with some of the more classical survival models that are currently most popular – in this case, the Cox Proportional Hazards model for single events (Cox, 1972) and the Fine-Gray model for competing risks (Fine and Gray, 1999). 2. Pseudo-observations for single events and competing risks The methodology of pseudo-observations was first proposed by Andersen et al. (2003). The main idea of this approach is to replace censored observations by the function of event times 𝑓(𝑇), for which an expected value is 𝐸(𝑓(𝑇)). The condition is that an unbiased estimator 𝜃 of 𝜃 = 𝐸(𝑓(𝑇)) exists. Let 𝑛 be the sample size (𝑖 = 1,…,𝑛). A pseudo-observation for 𝑓(𝑇) for individual 𝑖 at a predefined series of time points 𝑡 = 1,…,𝐻 is defined as 𝜃 𝑖(𝑡) = 𝑛𝜃 (𝑡)−(𝑛−1)𝜃 (−𝑖)(𝑡) (1) and is evaluated by the leave-one-out method. 𝜃 (𝑡) is the estimator in the sample of size 𝑛 at time 𝑡, and 𝜃 (−𝑖)(𝑡) is the estimator at time 𝑡 in the sample of size 𝑛−1, consisting of all units except the 𝑖-th individual. The pseudo-observation is then a contribution of the 𝑖-th unit to the 𝐸(𝑓(𝑇)) estimate in the sample of size 𝑛. Although the aim of using pseudo-observations is to replace the censored observations, pseudo-observations are calculated for all units in the sample (both completed and censored observations). Therefore, an 𝑛 × 𝐻 matrix of pseudoobservations is obtained. Subsequently, pseudo-observations are used as dependent variables in a generalised regression model with some link function 𝑔: 𝑔(𝐸(𝑓(𝑡)|𝑋)) = 𝛽0+∑𝛽𝑗𝑋𝑗= 𝛽𝑇𝑋. (2) For each unit 𝐻 pseudo-observations are calculated. Multiple measurement is a source of correlation in the data set; a possible solution to this deficiency would be to use generalised estimating equations (GEE), which are the generalisation of regression models for the case of correlated data (Andersen et al., 2003). 2.1. Single event Assume that there is only one type of event and 𝑇 is the time to that event, while 𝑇𝐶 is the time to censoring. Due to the right censoring, we can observe min(𝑇,𝑇𝐶). The survival function is the probability that the unit does not experience the event until time 𝑡 𝑆(𝑡) = 𝑃(𝑇 > 𝑡). (3)
STATISTICS IN TRANSITION new series, March 2019 173 In the survival analysis to the assumed sole type of event (single event), the survival function 𝑆(𝑡) can be estimated with the use of the Kaplan-Meier (KM) estimator 𝑆 (𝑡)=∏(1−𝐷𝑗 𝑁𝑗) 𝑡𝑗≤𝑡 , (4) where 𝐷𝑗 is the number of events at time 𝑡𝑗, 𝑁𝑗 is the number at risk just prior to time 𝑡𝑗, and 𝑡𝑗 for 𝑗 = 1,…,𝑟 (𝑟 ≤ 𝑛) are distinct event times. The KM estimator is a maximum likelihood estimator (Klein and Moeschberger, 2003). The 𝑖-th pseudo-observation based on the survival function is 𝜃 𝑖(𝑡) = 𝑛𝑆 (𝑡)−(𝑛 −1)𝑆 (−𝑖)(𝑡), (5) where 𝑆 (𝑡) is the estimated survival function at time 𝑡 in a sample of size 𝑛 and 𝑆 −𝑖(𝑡) is the estimated survival function derived from the 𝑛 −1 sample (without the 𝑖-th observation) (Andersen and Perme 2010). At 𝑡 = 0, the pseudoobservations for survival functions for all units are equal to one. As 𝑡 increases, the values of pseudo-observations for units in the cohort increase at each event time observed in the cohort (see Figure 1). Between any two successive event times, the values of pseudo-observations do not change. As a result, the curve of pseudo-observations over time for a particular unit is a step function with a varying length of steps depending on the successive event times. If the event for a unit is observed, the pseudo-observation drops below zero at the event time. At the subsequent time points, the unit that has just been excluded from the cohort has negative and increasing pseudo-values. If the unit is censored, then, beginning at the next event time after censoring, the values of pseudo-observations for that unit start decreasing. They remain, however, positive until the end of the follow-up (see Figure 1). Figure 1. The pseudo-observations for the survival function over time in a censored data set for the individual with event time 𝑡=7 (risk 1) and the individual with censored time 𝑡𝑐=7 (censored) As long as the units are in the cohort, they have similar pseudo-values. The values of pseudo-observations increase at each event time observed in the cohort. Therefore, the value of the pseudo-observation for the unit at its event
174 E. Wycinka, T. Jurkiewicz: Survival regression models… time is greater if the event occurred later in time. The later the event occurs, the greater the drop is (see Figure 2). The same pattern is observed if the observation is censored (see Figure 3). Figure 2. Pseudo-observations for the survival function over time for the units with event times 𝑡=4,..9 Figure 3. Pseudo-observations for the survival function over time for the units with censoring times 𝐶=4,..9 In the absence of censoring, the pseudo-value at time 𝑡 reduces to the indicator that 𝑇 > 𝑡. Therefore, the pseudo-observations are equal as long as the unit is observed in the cohort; after the event, the value of the pseudo-observation falls to zero and is constant until the end of the follow-up (see Figure 4). In this case, pseudo-observations are also independent.
STATISTICS IN TRANSITION new series, March 2019 175 Figure 4. Pseudo-observations for the survival function over time for an individual with a survival time 𝑡=7 in a data set with no censoring 2.2. Competing risks Let (𝑇,𝐶) be a bivariate random variable, such that 𝑇 is a continuous variable representing the time of the first event, and 𝐶 = 𝑘 (𝑘 = 1,…,𝑝) is a discrete variable denoting the type of event. If the time of the observation for some units is shorter than the time of the first event, we encounter right censoring. In such a situation, 𝐶 = 0 and 𝑇𝑐 is the time at which the observation was censored; what we only know is that 𝑇 > 𝑇𝑐. Due to the right censoring, the variable (𝑇,𝐶) is only partially observable, and we observe a pair (min{𝑇,𝑇𝑐},𝐶). As a result, the joint distribution of (𝑇,𝐶) is difficult to identify and can be estimated only by making some unverifiable assumptions (Pintilie, 2006, p. 41). The subdistribution of event 𝑘 (cumulative incidence function, CIF) is the probability, until time 𝑡, that event 𝑘 will occur 𝐹𝑘(𝑡)= 𝑃(𝑇 ≤ 𝑡,𝐶 = 𝑘). (6) The subdistribution is not a proper distribution because 𝑙𝑖𝑚 𝑡→∞ 𝐹𝑘(𝑡)= 𝑃(𝐶 = 𝑘)≤ 1. (7) The equality 𝑃(𝐶 = 𝑘)= 1 holds if there is only one type of event (no competing risks). The sum of the subdistributions for all types of events is a marginal distribution of the variable 𝑇 𝐹(𝑡)= 𝑃(𝑇 ≤ 𝑡)=∑𝐹𝑘(𝑡) 𝑝 𝑘=1 . (8) The maximum likelihood estimator of the subdistribution is 𝐹 𝑘(𝑡)=∑ℎ 𝑘𝑗𝑆 (𝑡𝑗−1) 𝑡𝑗≤𝑡 , (9) where ℎ 𝑘𝑗 is the cause-specific hazard at time 𝑡𝑗 for event 𝑘. This can be defined as ℎ 𝑘𝑗 =𝐷𝑘𝑗 𝑁𝑗, where 𝐷𝑘𝑗 is the number of events of type 𝑘 at time 𝑡𝑗, 𝑁𝑗 is the number at risk just prior to time 𝑡𝑗, and 𝑡𝑗 , for 𝑗 = 1,…,𝑟 (𝑟 ≤ 𝑛) are distinct event times. 𝑆 (𝑡𝑗−1) is the survival function for all types of events just before time 𝑡𝑗.
176 E. Wycinka, T. Jurkiewicz: Survival regression models… It is worth noting that the estimate depends not only on the number of individuals who have experienced the 𝑘-th type of event, but also on the number of individuals who have not experienced any type of event (Binder et al., 2014). Usually, only one type of event is of interest and other types of event are treated as competing risks to it. In such a situation, it is reasonable to consider only two types of event: the event of interest (risk 1) and every other event combined (risk 2). This approach will be considered later in this paper. The pseudo-observation for the unit 𝑖, at time 𝑡, for the event type 𝑘, based on the CIF, has the form 𝜃 𝑖𝑘(𝑡) = 𝑛𝐹 𝑘(𝑡)−(𝑛 −1)𝐹 𝑘(−𝑖)(𝑡). (10) Here, 𝐹 𝑘(𝑡) is the estimated CIF for the 𝑘-th event at time 𝑡 using all observations, and 𝐹 𝑘(−𝑖)(𝑡) is the estimated CIF derived from all but the 𝑖-th observation. When units are in a cohort, have the same pseudo-observation values for the CIF at subsequent times. At 𝑡 = 0, a pseudo-observation for the CIF equals zero. Then, as time increases, pseudo-observations decrease, taking negative values. Figure 4 shows pseudo-observations over time for a unit that leaves the cohort at time 7. If the unit leaves the cohort due to an event of type 1, the pseudo-observation jumps above one at the time of the event, and then, at subsequent times, gradually decreases towards one. When the unit leaves the cohort due to an event of type 2, the pseudo-values remain negative and decreasing at all subsequent times (Andersen and Perme, 2010). If an individual is censored at time 𝑡, the pseudo-observations start increasing as of the next event time recorded in the data set (see Figure 5). Figure 5. A comparison of the development of pseudo-values over time for three units that leave a cohort at time 7 due to either risk 1, risk 2, or censoring As we compare the pseudo-observations for units with the same cause of leaving a cohort but at different times, we can see greater changes in the pseudovalues for later departures (see Figure 6). Jumps in the values of pseudoobservations are higher if the event of type 1 happens later, due to the reduction in time of the risk set.
STATISTICS IN TRANSITION new series, March 2019 177 Figure 6. Pseudo-observations for units that experienced (a) risk 1 or (b) risk 2 or were (c) censored at different times (t=4,...9). Different scales are used in each figure In a special case with no competing risks, the estimated CIF for the type-1 event reduces to the estimation of the distribution function (𝐹 𝐶𝐼𝐹), and the survival function can be estimated as 𝑆 𝐶𝐼𝐹(𝑡)= 1−𝐹 𝐶𝐼𝐹(𝑡). If survival functions are estimated directly with a Kaplan-Meier estimator or as 𝑆 𝐶𝐼𝐹(𝑡), the estimations are equal. However, as we show in the empirical part of the study, in the case of pseudo-observations based on these estimators, this equality no longer holds.
178 E. Wycinka, T. Jurkiewicz: Survival regression models… If no censoring occurs in the data set, the pseudo-observations for risk 1 reduce to the indicator 𝐹1(𝑡)= 1[𝑇1≤ 𝑡]. They equal zero as long as the unit is in the cohort and rise towards one as the event of type 1 (risk 1) happens. The pseudo-observations for risk 2 equal zero at all time points, even after the occurrence of the event of type 2 (risk 2) (see Figure 7). Figure 7. Pseudo-observations for the CIF for risk 1 and risk 2 over time in a data set with no censoring 3. Regression models based on pseudo-observations For each unit there are 𝐻 pseudo-observations – one for each predefined point in time. As a result, the data transformed into pseudo-observations is no longer independent, and generalised linear models (GLM) cannot be applied. Generalised estimating equations (GEEs) are the generalisation of GLM models for correlated data, as introduced by Liang and Zeger (1986). This is a method for analysing data collected in clusters where observations within a cluster may be correlated, but observations from different clusters are independent. The variance is a function of the expectation, and a monotone transformation of the expectation is linearly related to the explanatory variable (Højsgaard et al., 2005). The pseudo-observations are dependent variables in GLMs for a given link function 𝑔(.). The regression model is 𝑔(𝜃 𝑘(𝑡)|𝑋) = 𝛽0+∑𝛽𝑗𝑋𝑗∗𝑚+𝐻 𝑗=1 (11) Here, the vector 𝑋∗ includes indicators of time points 𝑋 = (𝑋𝑚+1,…,𝑋𝑚+𝐻) for 𝑡 = 1,…,𝐻 (as dummy variables), as well as the covariates 𝑋 = (𝑋1,…,𝑋𝑚). When a complementary log-log link function is used, such as 𝑔(𝑥)= log (−log(𝑥)) for a single event, then the regression model has the form log (−log(𝑆(𝑡|𝑋))) = 𝛽0+∑𝛽𝑗𝑋𝑗∗𝑚+𝐻 𝑗=1 (12) and can be depicted as 𝑆(𝑡|𝑋)= exp (−exp (𝛽0+∑𝛽𝑗𝑋𝑗∗)) 𝑚+𝐻 𝑗=1 . (13) Estimated coefficients for time points can be put into the model as timedependent coefficients 𝛽0(𝑡): 𝑆(𝑡|𝑋)= exp (−exp (𝛽0+𝛽0(𝑡)+∑𝛽𝑗𝑋𝑗)) 𝑚 𝑗=1 . (14)
STATISTICS IN TRANSITION new series, March 2019 185 Table 4. Measures of the performance of models for the survival function. Month H 95% CI K-S 95% CI AUC 95% CI GEE F-G GEE F-G GEE F-G p-value 3 0.199 0.130-0.302 0.203 0.148-0.317 0.433 0.314-0.526 0.418 0.329-0.528 0.751 0.680-0.800 0.749 0.694-0.805 0.832 4 0.229 0.163-0.296 0.226 0.175-0.309 0.454 0.365-0.525 0.449 0.384-0.534 0.773 0.720-0.803 0.773 0.734-0.811 0.923 5 0.207 0.157-0.267 0.207 0.171-0.277 0.42 0.346-0.484 0.414 0.359-0.492 0.762 0.718-0.788 0.761 0.728-0.797 0.892 6 0.199 0.148-0.252 0.195 0.163-0.264 0.406 0.332-0.459 0.401 0.351-0.468 0.751 0.708-0.777 0.752 0.722-0.787 0.778 7 0.195 0.152-0.244 0.197 0.166-0.259 0.403 0.339-0.456 0.403 0.360-0.469 0.753 0.714-0.777 0.755 0.729-0.785 0.531 8 0.193 0.152-0.241 0.193 0.166-0.251 0.391 0.333-0.445 0.390 0.352-0.458 0.751 0.714-0.773 0.753 0.729-0.781 0.489 9 0.184 0.147-0.23 0.184 0.161-0.24 0.380 0.327-0.433 0.379 0.347-0.447 0.746 0.711-0.769 0.748 0.725-0.775 0.452 10 0.179 0.142-0.222 0.180 0.159-0.235 0.370 0.318-0.421 0.369 0.337-0.434 0.74 0.706-0.763 0.744 0.722-0.771 0.226 11 0.173 0.138-0.218 0.175 0.155-0.225 0.356 0.306-0.405 0.357 0.325-0.420 0.733 0.698-0.755 0.738 0.717-0.764 0.101 12 0.169 0.133-0.211 0.171 0.151-0.223 0.35 0.301-0.400 0.353 0.320-0.412 0.729 0.695-0.751 0.734 0.714-0.761 0.056 13 0.172 0.139-0.215 0.174 0.154-0.225 0.354 0.305-0.403 0.358 0.325-0.415 0.731 0.700-0.753 0.736 0.715-0.762 0.070 14 0.17 0.140-0.214 0.172 0.152-0.221 0.35 0.303-0.400 0.354 0.323-0.411 0.727 0.697-0.751 0.732 0.714-0.759 0.064 15 0.171 0.141-0.215 0.169 0.149-0.218 0.353 0.307-0.401 0.354 0.321-0.408 0.728 0.700-0.751 0.731 0.713-0.757 0.210 GEEgeneralized estimating equations, F-G – Fine-Gray model, 95% CI – 95% confidence intervals as percentiles form 1000 bootstrapped samples. For the single-event approach, we also applied the method based on a reduction of the CIF to the case of one type of event. This led to the use of the 𝑆 𝐶𝐼𝐹(𝑡)= 1−𝐹 𝐶𝐼𝐹(𝑡)) estimator of the survival function, instead of the 𝑆 𝐾𝑀(𝑡) estimator, in the calculation of the pseudo-observations. We observed that, for the pseudo-observations, the relation 𝑆 𝐾𝑀(𝑡)= 1−𝐹 𝐶𝐼𝐹(𝑡) does not hold. The differences between estimates were very low, but irregular. Figure 8 shows boxplots for the differences 𝑆 𝐾𝑀(𝑡)−(1−𝐹 𝐶𝐼𝐹(𝑡)) for all the units at all time points. However, one should note that all the differences are very close to zero.
186 E. Wycinka, T. Jurkiewicz: Survival regression models… Figure 9. Distribution of differences between the KM estimator and the complement to CIF estimator for the survival function, both estimated by pseudo-observations The GEE models for survival functions built for pseudo-observations based on both estimators gave exactly the same results. Thus, in spite of the above differences, the methods are fully interchangeable. 5. Conclusions Pseudo-observations are a method that can be considered competitive with other survival analysis techniques. As shown in section 2, the values of pseudoobservations depend both on the type and time of event. Regression models for pseudo-observations correctly evaluate the whole survival curve, and the use of the log(-log) link function causes the GEE models for both single and competing approaches to simply mimic the results of the Cox PH and Fine-Gray models, respectively. This observation is consistent with the results of earlier studies by other authors and argues against the use of a more cumbersome pseudo-values approach instead of more classic methods. However, because the independence matrix happened to be the best choice for the GEE model in all of the studies, it is suggested that pseudo-observations could be used as dependent variables in other methods for complete, independent data, such as classification trees. In application to credit-risk assessment, competing-risks models had more discriminatory power than single-event models, which supports the use of competing-risks models in preference to models for single events. Further studies should focus on the variable-selection method that could be applied to the GEE models. Acknowledgments The authors gratefully acknowledge helpful feedback from an anonymous reviewer.
STATISTICS IN TRANSITION new series, March 2019 187 REFERENCES AGRESTI, A., (2007). Logistic Regression, in An Introduction to Categorical Data Analysis, Second Edition, John Wiley & Sons, Inc., Hoboken, NJ, USA. AKAIKE, H., (1974). A new look at the statistical model identification, IEEE Transactions on Automatic Control, 19 (6), pp. 716–723. ANDERSEN, P. K., KLEIN, J. P., ROSTHØJ, S., (2003). Generalised linear models for correlated pseudo-observations, with applications to multi-state models, Biometrika, 90 (1), pp. 15–27. ANDERSEN, P. K., PERME, M., (2010). Pseudo-observations in survival analysis, Statistical Methods in Medical Research 19 (1), pp. 71–99. BINDER, N., GERDS, T. A., ANDERSEN, P. K., (2014). Pseudo-observations for competing risks with covariate dependent censoring, Lifetime data analysis, 20(2), pp. 303–315. COX, D., (1972). Regression Models and Life-Tables, Journal of the Royal Statistical Society, Series B (Methodological), 34 (2), pp. 187–220 DELONG, E., DELONG, D., CLARKE-PEARSON, D., (1988). Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach, Biometrics 44, pp. 837–845. DIRICK, L., CLAESKENS, G., BAESENS, B., (2017). Time to default in credit scoring using survival analysis: a benchmark study, J Oper Res Soc 6, pp. 652–655. FINE, J., GRAY, R., (1999). A Proportional Hazards Model for the Subdistribution of a Competing Risk, Journal of the American Statistical Association, 94 (446), pp. 496–509. HALLER, B., SCHMIDT, G., ULM, K., (2013). Applying competing risks regression models: an overview, Lifetime Data Anal 19, pp. 33–58. HAND, D. J., (2009). Measuring classifier performance: a coherent alternative to the area under the ROC curve, Mach Learn 77, pp. 103–123. HØJSGAARD, S., HALEKOH, U., YAN, J., (2005). The R Package geepack for Generalized Estimating Equations. Journal of Statistical Software, 15:2, pp. 1–11. KLEIN, J., MOESCHBERGER, M., (2003). Survival Analysis: Techniques for Censored and Truncated Data, Statistics for Biology and Health, 2nd ed., Springer, New York. KLEIN, J. P., ANDERSEN, P. K., (2005). Regression Modelling of Competing Risks Data Based on Pseudovalues of the Cumulative Incidence Function, Biometrics, 61 (1), pp. 223–229. KUK, D., VARADHAN R., (2013). Model selection in competing risks regression, Statistics in Medicine 32, pp. 3077–3088.
188 E. Wycinka, T. Jurkiewicz: Survival regression models… LIANG, K., ZEGER, S., (1986). Longitudinal Data Analysis Using Generalized Linear Models. Biometrika, 73 (1), pp. 13–22. MILLS, M., (2011). Introducing Survival and Event History Analysis, Sage, Los Angeles. PINTILIE, M., (2006). Competing Risks: A Practical Perspective, Wiley. VENABLES, W. N., RIPLEY, B. D., (2002). Modern Applied Statistics with S. Fourth edition, Springer. WATKINS, J. G. T., VASNEV, A. L., GERLACH, R., (2014). Multiple Event Incidence and Duration Analysis for Credit Data Incorporating Non-Stochastic Loan Maturity, J. Appl. Econ., 29, pp. 627–648.