scieee AI-readable full text Open interactive document viewer

Statistical models to study subtoxic concentrations for some standard mutagens in three colon cancer cell lines

Bardina, Xavier; Fernandez Gomez-Recuero, Laura; Piñeiro, Elisabet; Surralles, Jordi; Velázquez Henar, Antonia

Abstract

The aim of this work is to propose models to study the toxic effect of different concentrations of some standard mutagens in different colon cancer cell lines. We find estimates and, by means of an inverse regression problem, confidence intervals for the subtoxic concentration, that is the concentration that reduces by thirty percent the number of colonies obtained in the absence of mutagen.

Full text

Statistics & Operations Research Transactions SORT 30 (2) July-December 2006, 193-204 Statistics & Operations Research Transactions Statistical models to study subtoxic concentrations for some standard mutagens in three colon cancer cell lines c Institut d’Estad´ ıstica de Catalunya [email protected] ISSN: 1696-2281 www.idescat.net/sort Xavier Bardina1, Laura Fern´ andez2, Elisabet Pi˜ neiro2, Jordi Surrall´ es2& Antonia Vel´ azquez2 Universitat Aut`onoma de Barcelona Abstract The aim of this work is to propose models to study the toxic effect of different concentrations of some standard mutagens in different colon cancer cell lines. We find estimates and, by means of an inverse regression problem, confidence intervals for the subtoxic concentration, that is the concentration that reduces by thirty percent the number of colonies obtained in the absence of mutagen. MSC: 62J05 Keywords: Inverse regression problem, subtoxic concentration, confidence interval. 1 Introduction Human populations are exposed to a variety of environmental agents, including biological, chemical and physical entities, that can injure the DNA and cause adverse health consequences, such as cancer. It is therefore extremely important to detect these mutagenic agents, unravel their mechanisms of action, and define what type of injury they produce. All this information is crucial to determine the genetic risk of exposed population. 1Departament de Matem` atiques, Universitat Aut` onoma de Barcelona, 08193 Bellaterra, Barcelona, Spain. [email protected].cat 2Departament de Gen` etica, Universitat Aut` onoma de Barcelona, 08193 Bellaterra, Barcelona, Spain. [email protected].es, [email protected].es, [email protected].es, a[email protected].es Received: September 2006 Accepted: November 2006 194 Statistical models to study subtoxic concentrations for some standard... Mutagenicity assays are specifically designed to detect DNA damaging agents and to analyze their biological effects. Most of these assays are performed in vitro with a variety of cell types and under controlled conditions of cell growth and viability. The range of the treatment concentrations is usually determined in a previous toxicity study. The chosen concentrations for the mutagenicity study must be subtoxic in order to ensure biological effects without extreme cell death. This is why a well performed toxicity assay is absolutely required as a previous routine in all mutagenicity assays. Tandem repeated DNA sequences of few nucleotides, the so-called microsatellites, are known to be highly unstable in colon cancer cells defective in DNA mismatch repair (MMR). In addition, expansion of tandem repeated sequences have been causally related to a number of degenerative diseases including miotonic dystrophy, fragile X syndrome and Huntington’s disease. The final aim of the present investigation was to determine whether microsatellite instability is inducible in vitro by a set of mutagens of different mode of action. To do so, subtoxic concentrations, i.e. those inducing a reduction of thirty percent in cell viability, had to be previously determined for the following standard mutagens: bleomycin (BLEO), N-methyl-N-nitrosourea (MNU), ethoposide (ETO), mitomycin C (MMC) and ethidium bromide (EtBr). The toxicity assay was carry out in three human fibroblast cell lines derived from colon tumours: the wild-type cell line SW480, and lines HCT116 and LoVo which are both defective in MMR. The toxicity data were obtained by the colony forming efficiency method. About 200 cells from exponentially growing cell cultures were platted in triplicate on 25cm2 plates (falcons). After allowing for attachment to the plate for 24 hours, the medium was replaced with fresh medium containing the test chemical at different concentration for each replica. Cell lines were maintained in these conditions for 10 days, replacing the medium every 3 days. The plates were washed with phosphate buffer saline, fixed with methanol, and stained with Giemsa. Colonies with more than 50 growing cells were counted. A reduction in the number of colonies after treatment is interpreted to be a consequence of the chemical toxicity. In this article we will propose different types of statistical models to study the effect of the concentration of the mutagens in three colon cancer cell lines. We have considered linear and exponential models and in each case we propose the model that approximates better the data. When we consider a regression linear model, we need to assume that the errors are additive and normally distributed. When the model considered is an exponential one, we assume that the errors are multiplicative and their logarithms are normally distributed. In all the models proposed, the analysis of the residuals does not give evidence that controvert this hypothesis. For each mutagen, cell line and concentration we have three values. Then, we will consider models with weights, where the weights are calculated by the inverse of the estimated variances. For each mutagen and each cell line we will obtain an estimation and a confidence interval for the subtoxic concentration, that is the concentration that reduces a thirty percent the initial number of colonies. Xavier Bardina, Laura Fern´ andez, Elisabet Pi˜ neiro, Jordi Surrall´ es & Antonia Vel´ azquez 195 This article is organized as follows. In Section 2 we provide the mathematical justification for the calculation of the confidence intervals used in this study. In Section 3 we present some models for the standard mutagens: bleomycin, N-methyl-N-nitosourea, ethoposide, mitomycin C and ethidium bromide respectively. We also provide for each mutagen and each cell line the estimate of the subtoxic concentration and a confidence interval for this concentration. In Section 4 the use of weighted models is justified. Finally, in the appendix the data used in the study are shown. 2 Mathematical justification of the confidence intervals used in this study In this section we will explain the method used in order to obtain a confidence interval of the subtoxic concentration, that is the concentration that reduces by thirty percent the initial number of colonies. Consider the linear regression model y=β0+β1x, where yis the number of colonies and xthe concentration of the mutagen. We will estimate the concentration that reduces to 70 percent the number of colonies for x=0 (that is, the concentration xsuch that y=0.7β0)by ˆx=−0.3·ˆ β0 ˆ β1 , where ˆ β=ˆ β0,ˆ β1is the usual estimate of (β0,β 1)obtained by the least-squares method. That is, ˆ β=XX−1Xy, where X=⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ 1x1 1x2 . . .. . . 1xn ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ,y=⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ y1 y2 . . . yn ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ and (y1,x1),···,(yn,xn) are the data of the number of colonies and concentrations, respectively. If we consider the vector λ=(λ1,λ 2), it is well known that we can obtain a 100(1 −α)%confidence region for λβby using the following expression: 196 Statistical models to study subtoxic concentrations for some standard... λβ=λˆ β±tα/2,n−2S√λΛλ, where Λ=(XX)−1,Sis the estimate of the standard deviation of the errors and tα/2,n−2 is the critical point such that P|T|>tα/2,n−2=αwhere Tis a Student’s tdistribution with n−2 degrees of freedom. On the other hand, from the linear regression model, the concentration xthat reduces to 70 percent the inicial number of colonies satisfies that 0.7β0=β0+β1x, that is 0.3β0+xβ1=0 and we can write this expression as λβ=0 with λ=(0.3,x). So, in order to obtain a sort of 100(1−α)%confidence interval for the concentration xthat reduces a 30 % the inicial number of colonies, we can solve the following system: ⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ λβ=λˆ β±tα/2,n−2S√λΛλ λβ=0 (1) whith λ=(0.3,x). That is, we have to find the intersections (if there exists) between the line λβ=0 and the curves given by λβ=λˆ β±tα/2,n−2S√λΛλ.Thiskindof problems are called inverse regression problems (see for example Draper and Smith, 1981). When we use a linear model with weights, the matrix Λnow is given by Λ=XV−1X−1 and the estimate of βis ˆ β=XV−1X−1XV−1y, where V−1=⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ w1... 0 . . ..... . . 0... wn ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ and w1,w2,...,wnare the weights of the data y1,y2,...,yn, respectively (see for example Xavier Bardina, Laura Fern´ andez, Elisabet Pi˜ neiro, Jordi Surrall´ es & Antonia Vel´ azquez 197 Montgomery, 1992, or Draper and Smith, 1981). Observe that in our examples the confidence interval are approximated because the weights are estimated from the data. When we consider an exponential model, that is, y=eβ0+β1x, we assume that the residuals are multiplicative and that their logarithms are normally distributed. That is, we suppose that if we consider the transformation ln y=β0+β1x we can proceed like in a linear model. But in this case, we have to estimate the concentration xsuch that y=0.7eβ0. Thus, we will estimate xby ˆx=ln 0.7 ˆ β1 . Then, we obtain the confidence interval solving the system: ⎧ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎩ λβ−ln 0.7=λˆ β−ln 0.7±tα/2,n−2S√λΛλ λβ−ln 0.7=0, where λ=(0,x). We remark that this confidence interval has an exact confidence level of 100(1−α) % only when the weights are perfectly known. If the weights are estimated from the sample then the confidence levels are approximated. 3Examples Using the method described in Section 2 we will obtain now a confidence interval for the subtoxic concentration of the Bleomycin in the cell line LoVo. From the data given in the Appendix we have that X= ⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ 10.0000 10.0000 10.0000 10.0001 . . .. . . 10.0500 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ,y= ⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ 26.5 29.0 31.0 32.5 . . . 6.5 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ,V−1= ⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ 0.1967 0 ... 0 00.1967 ... 0 . . .. . ..... . . 00... 0.2449 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ . 198 Statistical models to study subtoxic concentrations for some standard... Thus, with an easy computation we obtain that Λ=XV−1X−1=1.039 −22.082 −22.082 1008.297 , and the estimate of βis given by ˆ β=XV−1X−1XV−1y=27.205 −440.677 . In this example the estimate of the standard deviation of the errors is S=1.083 and the critical point of the Student’s tis t0.05 2,11 =2.201. Then in order to obtain the confidence interval explained in this section we have to solve the system (1). In this example we get 0.3β0+β1x=8.161 −440.677x±2.383 √0.093 −13.249x+1008.297x2 0.3β0+β1x=0. So,wehavetofind the roots of the following equation: 440.677x−8.161 =±2.383 0.093 −13.249x+1008.297x2, that is, the roots of the quadratic equation 33182.804x2−1253.193x+11.634 =0, that are given by the values x1=0.0164 x2=0.0213. Then, (0.0164,0.0213) is a confidence interval of approximately 95 % for the subtoxic concentration of the Bleomycin in the cell line Lovo. 3.1 Bleomycin (BLEO) We have considered a regression line with weights for each cell line. Recall that yis the number of colonies and xthe concentration of bleomycin. The models obtained using the SAS system are: cell line regression line HCT116 y=67.597 - 1168.925 x+ε LoVo y=27.205 - 440.677 x+ε SW480 y=53.488 - 566.464 x+ε Xavier Bardina, Laura Fern´ andez, Elisabet Pi˜ neiro, Jordi Surrall´ es & Antonia Vel´ azquez 199 The regression coefficients R-square were 0.9285, 0.9372 and 0.9452, respectively. Finally, with these models we obtain the following estimates, and the following confidence intervals, for the subtoxic concentration in each cell line. cell line concentration estimate 95% confidence interval HCT116 0.0173 (0.0152, 0.0203) LoVo 0.0185 (0.0164, 0.0213) SW480 0.0283 (0.0259, 0.0315) 3.2 N-methyl-N-nitrosourea (MNU) We have considered a regression line with weights for the cell lines LoVo and SW480. In this case, the models obtained using the SAS system are the following: cell line regression line LoVo y=22.117 - 0.218 x+ε SW480 y=67.826 - 0.663 x+ε The regression coefficients R-square were 0.9330 and 0.9512, respectively. With these models we obtain the following estimates, and the following confidence intervals, for the subtoxic concentration in each cell line. cell line concentration estimate 95% confidence interval LoVo 30.477 (29.299, 31.777) SW480 30.708 (29.249, 32.344) 3.3 Ethoposide (ETO) We have considered a regression line with weights for each cell line. The models obtained using the SAS system are: cell line regression line HCT116 y=130.075 - 860.173 x+ε LoVo y=43.509 - 233.079 x+ε SW480 y=104.453 - 423.569 x+ε The regression coefficients R-square were 0.9032, 0.7189 and 0.8282, respectively. With these models we obtain the following estimates, and the following confidence intervals, for the subtoxic concentration in each cell line. 200 Statistical models to study subtoxic concentrations for some standard... cell line concentration estimate 95% confidence interval HCT116 0.0454 (0.0380, 0.0563) LoVo 0.0560 (0.0432, 0.0838) SW480 0.0740 (0.0593, 0.0996) 3.4 Mitomycin C (MMC) For this data sets we have fitted a regression line with weights for each cell line. The models obtained using the SAS system are: cell line regression line HCT116 y=85.321 - 6334.859 x+ε LoVo y=13.989 - 714.984 x+ε SW480 y=48.554 - 4174.451 x+ε The regression coefficients R-square were 0.9253, 0.5364 and 0.9314, respectively. With these models we obtain the following estimates, and the following confidence intervals, for the subtoxic concentration in each cell line. cell line concentration estimate 95% confidence interval HCT116 0.0040 (0.0036, 0.0046) LoVo 0.0059 (0.0040, 0.0115) SW480 0.0035 (0.0032, 0.0039) 3.5 Ethidium bromide (EtBr) In this situation we have considered a exponential model with weights for the cell lines HCT116 and SW480, i.e., a regression linear model with weights for the logarithms of the data. The models obtained using the SAS system are the following: cell line model HCT116 y=exp(4.477 - 19.474 x)×ε SW480 y=exp(3.722 - 24.330 x)×ε The regression coefficients R-square were 0.8584 and 0.8601, respectively. With these models we obtain the following estimates, and the following confidence intervals, for the subtoxic concentration in each cell line. cell line concentration estimate 95% confidence interval HCT116 0.0183 (0.0147 , 0.0242) SW480 0.0147 (0.0118, 0.0193) Xavier Bardina, Laura Fern´ andez, Elisabet Pi˜ neiro, Jordi Surrall´ es & Antonia Vel´ azquez 201 4 Conclusions Our exploration of the data sets (see the Appendix) demonstrates the non-homogeneity of the variances of the errors of the corresponding models, and consequently the non-adequateness of the classical least squares method to estimate the parameters. For instance, for the Ethoposide line HCT116, the variances corresponding to each concentration are 2.583, 288.250, 64.333, 376.333 and 82.583. The Levene test, that can be performed with the SPSS statistical package, rejects the equaliy of variances with p=.029. Then a possible option would be to transform the dependent variable by using a suitable stabilizing variance transformation. However the usual transformations (powers, logarithms, etc.) are employed when certain specific patterns are observed between the sampling variances and their respective means. This is not the situation as can be shown in the example above by plotting the variances against their corresponding means. Another option (those considered in this report) is to use the weighted linear regression models. These models can be implemented by using any of the more usual statistical packages such as SAS or SPSS. The ideal setting is when the weight (the inverse of the variance) for each observation is perfectly known. There are real examples where it happens (see Draper and Smith, 1981) but this is not our case here – for the data sets studied in this report the weights are estimated. For the linear regression model we have expressed the value of the variable y(number of colonies) corresponding to the subtoxic concentration xas a linear combination of the coefficients of the regression line. The same has happened with the exponential model and the variable ln y. This is the key point that has allowed us to find the confidence intervals for the subtoxic concentrations. That is, in a linear regression model (respectively, in an exponential model), this method can be applied to find confidence intervals for the value of the variable xfor which the variable y(respectively, the variable ln y) can be expressed as a linear combination of the coefficients of the regression line. Appendix In this Appendix the data used in this work are provided. The methodology used in order to obtain these results has been explained in Section 1. We express also, between parenthesis, the weight of each datum. Recall from the theory of linear models that the weights are given by the inverse of the estimated variances of the three data obtained for each concentration. Observe that some data corresponding to the number of colonies are not integer numbers. The reason is that the number of colonies was counted twice, and we have used their average.