scieee AI-readable full text Open interactive document viewer

Decision making association rules for recognition of differential gene expression profiles

Rubio Escudero, Cristina; Val, Coral del; Cordón, Óscar; Zwir, Igor

Abstract

The rapid development of methods that select over/under expressed genes from RNA microarray experiments have not yet satisfied the need for tools that identify differential profiles that distinguish between experimental conditions such as time, treatment and phenotype. We evaluate several microarray analysis methods and study their performance, finding that none of the methods alone identifies all observable differential profiles, nor subsumes the results obtained by the other methods. Therefore, we propose a machine learning based methodology that identifies and combines the abilities of microarray analysis methods to recognize differential profiles. We encode the results of this methodology in decision making association rules able to decide which method or method-aggregation is optimal to retrieve a set of genes exhibiting a common profile. These solutions are optimal in the sense that they constitute partial ordered subsets of all method-aggregations bounded by the most specific and the most sensitive available solution. This methodology was successfully applied to a study of inflammation and host response to injury data set derived from the analysis of longitudinal blood microarray profiles of human volunteers treated with intravenous endotoxin compared to placebo. Our approach was able to uncover a cohesive set of differentially expressed genes and novel members exhibiting previously studied differential profiles. This guideline serves as a means to support decisions on new microarray problems.

Full text

See discussions, stats, and author profiles for this publication at: https://www.researchgate.net/publication/221253303 Decision Making Association Rules for Recognition of Differential Gene Expression Profiles Conference PaperinLecture Notes in Computer Science · September 2006 DOI: 10.1007/11875581_135·Source: DBLP CITATION 1 READS 42 4 authors: Some of the authors of this publication are also working on these related projects: Software for Soft Computing View project Improving Breath Test sensitivity through Neural Networks and other Machine Learning systems. View project Cristina Rubio-Escudero Universidad de Sevilla 70 PUBLICATIONS564 CITATIONS SEE PROFILE Coral del Val University of Granada 89 PUBLICATIONS2,077 CITATIONS SEE PROFILE Oscar Cordon University of Granada 359 PUBLICATIONS11,626 CITATIONS SEE PROFILE Igor Zwir University of Granada 135 PUBLICATIONS1,844 CITATIONS SEE PROFILE All content following this page was uploaded by Coral del Val on 21 May 2014. The user has requested enhancement of the downloaded file. Decision Making Association Rules for Recognition of Differential Gene Expression Profiles C. Rubio-Escudero1, Coral del Val1, O. Cordón1,2, and I. Zwir1,3 1 Department of Computer Science and Artificial Intelligence, University of Granada, Spain 2 European Center for Soft Computing, Mieres, Spain 3 Howard Hughes Medical Institute, Washington University School of Medicine, St. Louis, MO {crubio, delval, ocordon, zwir}@decsai.ugr.es Abstract. The rapid development of methods that select over/under expressed genes from RNA microarray experiments have not yet satisfied the need for tools that identify differential profiles that distinguish between experimental conditions such as time, treatment and phenotype. We evaluate several microarray analysis methods and study their performance, finding that none of the methods alone identifies all observable differential profiles, nor subsumes the results obtained by the other methods. Therefore, we propose a machine learning based methodology that identifies and combines the abilities of microarray analysis methods to recognize differential profiles. We encode the results of this methodology in decision making association rules able to decide which method or method-aggregation is optimal to retrieve a set of genes exhibiting a common profile. These solutions are optimal in the sense that they constitute partial ordered subsets of all method-aggregations bounded by the most specific and the most sensitive available solution. This methodology was successfully applied to a study of inflammation and host response to injury data set derived from the analysis of longitudinal blood microarray profiles of human volunteers treated with intravenous endotoxin compared to placebo. Our approach was able to uncover a cohesive set of differentially expressed genes and novel members exhibiting previously studied differential profiles. This guideline serves as a means to support decisions on new microarray problems. 1 Background Advances in molecular biology and computational techniques permit the systematical study of molecular processes that underlie biological systems [1]. Particularly, microarray technology has revolutionized modern biomedical research by its capacity to monitor changes in RNA abundance for thousands of genes simultaneously [2]. To address the statistical challenge of analyzing these large data sets, new methods have emerged ([3], [4], [5], [6], [7]). However, there is a dearth of computational methods to facilitate understanding of differential gene expression profiles (e.g., profiles that change over time and/or over treatments and/or over patients) and to decide which is the most reliable method to identify differences across profiles. We develop a detailed evaluation of the performance of several commonly used statistical methods to identify differential expression profiles. We found that the application of these methods return different results applied over the same set of data: the methods do not identify all observable differential profiles (genes exhibiting a common behavior throughout experimental conditions). Moreover, none of the methods subsume the results obtained by the other methods. Our study reveals how some methods are able to recognize some differential profiles and not others and that some of the not retrieved profiles might contain significant genes for the experiment under study. Therefore, we propose a methodology that combines the properties of each method into a set of decision making association rules ([8], [9], [10]) devoted to discover optimal aggregations of microarray analysis methods in an effort to identify differential gene expression profiles. The association rules allow users to query for the most appropriate method or aggregation of them to retrieve significant genes based on the differential profiles they exhibit. To create such set of decision association rules we perform the following steps over a set of microarray gene expression data (Fig. 1). First, we extract from the data set all genes which behave in a different way from one experimental condition to the others (i.e., genes that change over time, treatments and phenotype). We apply several classical microarray analysis methods (T-Tests [11], Permutation Tests [6], Analysis of Variance [5] and Repeated Measures ANOVA [12]). Second, we create a database containing distinct types of differential profiles over time, experiment and subjects from previously recovered genes. DIFFERENTIALLY EXPRESSED GENES MICROARRAY RAW DATA MICROARRAY PREPROCESSED DATA PREPROCESSING (1) (4) IDENTIFICATION OF DIFFERENTIAL PROFILES DIFFERENTIAL EXPRESSION PROFILES METHOD EVALUATION SCALING (2) IDENTIFICATION OF DIFF. EXPRESSED GENES (6) PREDICTION QUERY PROFILES + MICROARRAY DATA (OPTIONAL) METHOD SELECTION (5) CREATION OF METHOD ASS. RULES METHOD ASSOCIATION RULES (3) ASSOCIATION OF STATISTICAL METHODS LATTICE OF METHOD ASSOCIATIONS STATISTICAL METHODS CONCEPTUAL CLUSTERING PROFILE IDENTIF. PROFILE PRUNING NORMALIZING Fig. 1. Graphical representation of the methodology. The squared boxes represent the phases of the methodology, the round cornered boxes correspond to the input/output data at each step, and the ellipses the operations performed at each phase. Third, we create decision making association rules, where the antecedents are differential profiles and the consequents are methods or aggregations of them capable to identify the profiles. Fourth, we arrange the association rules into a lattice, where the rules are ordered from the most general (top) to the most specific solution (bottom). We use this structure to evaluate the performance of the rules by analyzing their specificity, sensitivity and cost, applying multiobjective optimization techniques. Fifth, we use a selected set of optimal rules as a framework to support new decisions about the applicability of microarray analysis methods to retrieve differential gene expression profiles. 2 Results The results are obtained from the application of our procedure to a data set derived from longitudinal blood expression profiles of human volunteers treated with intravenous endotoxin compared to placebo. The motivation of these experiments is to provide insight to the host response to injury as part of a Large-scale Collaborative Research Project sponsored by the National Institute of General Medical Sciences (www.gluegrant.org) [13]. Analysis of the set of gene expression profiles obtained from this experiment is complex, given the number of samples taken and variance due to treatment, time, and subject phenotype. Therefore, we believe this problem is typical and informative as a microarray case study. The data were acquired from blood samples collected from eight normal human volunteers, four treated with intravenous endotoxin (i.e., patients 1 to 4) and four with placebo (i.e., patients 5 to 8). Complementary RNA was generated from circulating leukocytes at 0, 2, 4, 6, 9 and 24 hours after the i.v. infusion and hybridized with GeneChips® HG-U133A v2.0 from Affymetrix Inc., which contains 22216 probe sets, analyzing the expression level of 18400 transcripts and variants, including 14500 well-characterized human genes. 2.1 Accuracy of the Statistical Methods We investigate the performance of several commonly used statistical methods in identifying differential expression profiles that change over time, treatments and phenotype. We name T-Test as 1 M, T-Test considering time as 2 M, Permutation Test as 3 M, Permutation Test considering time as 4 M, ANOVA over treatment as 5 M, ANOVA over time as 6 M, ANOVA over treatment and time as 7 M, RMANOVA over treatment as 8 M, RMANOVA over time as 9 Mand RMANOVA over treatment and time as 10 M, where considering time refers to the fact that the tests have been specifically applied to find differences between time points. For our set of data, we found that these methods do not identify all observable distinct profiles. Moreover, none of them subsumes the results obtained by other methods (Table 1). Different methods retrieve different amounts of probe sets (e.g., the application of 1 Mover the microarray dataset retrieves 962 genes as differentially expressed, whereas 5 Mretrieves 1734 genes, and 3 M retrieves 612 genes). The concordance rates between the sets of genes retrieved also varies widely, indicating that none of the methods subsumes the others (Table 1)(e.g., from the genes retrieved by 3 M, only 31.11% are also retrieved by 5 M, and 52.29% by 1 M. 2.2 Statistical Methods and Differential Profiles We found that there is a relationship between the statistical methods and the differential profiles they are able to identify, having differential profiles identified by some methods and not by others. This type of relation is what we encode in the set of decision making association rules that we obtain from the application of our methodology. In our particular problem, there are genes highly related with the inflammation problem which exhibit profiles that would not be retrieved applying some of the classic microarray analysis methods individually. That is the case of probe set 206011_at, which is related in behavior and in function (apoptosis-related cysteine peptidase) to probe sets 211367_s_at and 211368_s_at (Fig. 2(a)), stated as relevant for the inflammation problem in [13]. For these particular probe set, the isolated application of classical methods such as 1 Mor 3 M with either default p-value or false discovery, rate, depending on what each method uses, would not retrieve such probe set as differentially expressed. The same situation applies to probe sets 202076_at and 210538_s_at, related both in behavior and in function (inhibitor of apoptosis protein 2 and 1 respectively) (Fig. 2(b)). Table 1. Intersection of the results between methods recognizing differentially expressed genes. The number in each cell represents the ratio of coincidence between genes retrieved by the statistical method in the column and in the row relative to the total number of genes recovered by the method in the row RowColumnRow /)(( ∩). % 1 M2 M3 M 4 M 5 M6 M7 M 8 M9 M10 M 1 M -- 92.20 52.29 75.05 96.48 69.23 85.55 70.06 61.33 50.52 2 M56.06 -- 34.07 57.84 85.27 59.54 71.11 62.64 50.57 42.98 3 M 82.19 88.07 -- 96.24 94.77 57.35 78.75 72.87 56.86 46.73 4 M67.22 85.19 54.84 -- 95.16 55.49 73.65 70.20 51.49 42.83 5 M55.20 77.80 33.45 58.94 -- 50.28 66.72 66.38 46.42 38.93 6 M59.04 83.51 31.11 52.84 77.30 -- 89.63 56.56 60.64 49.38 7 M58.36 79.79 34.18 56.10 82.05 71.70 -- 62.34 57.23 49.07 8 M 57.36 84.34 37.96 64.17 95.96 54.30 74.80 -- 49.62 40.51 9 M62.10 84.21 36.63 58.21 84.74 72.00 84.95 61.36 -- 72.31 10 M 59.56 83.34 35.05 56.37 82.72 68.26 84.80 58.34 84.19 -- In contrast, some other available methods retrieve profiles that do not differ between the considered experimental conditions. For example, ANOVA, perhaps based on the violation of statistical constraints ([14]), retrieves a 43% of genes lacking an observable change with the default parameter values. The increase of the specificity of these parameters generates severe effects on the sensitivity of other true changes. These findings reveal that there are desired and undesired differential profiles termed positive and negatively, respectively. For example, some profiles exhibiting similarly arranged patterns but shifted over time may be relevant for a specific experiment but not for other. In addition, we also found that methods applied to microarray profiles are focused on identifying differences among expression patterns over treatment and/or time since biological replicates are averaged in the same experimental group. However, we might also need to detect differences among subjects. We create a database with all possible differential profiles derived from genes retrieved in our inflammation problem by all available methods (i.e., 28 differential profiles). This database contains differential profiles that can be labeled as positive or negative according to their interest to be retrieved for a particular study. To validate biologically these profiles we calculate the coincidence the coincidence between our retrieved differential profiles and external information provided by the Gene Ontology database ([15]) showing that genes sharing behaviour are related in function ([16]). 1000 2000 3000 4000 5000 PAT1 PAT2 PAT3 PAT4 (a) 1000 2000 3000 4000 5000 PAT1 PAT2 PAT3 PAT4 (b) Fig. 2. Probe sets in blue are stated as relevant for the inflammation problem in ([13]). Probe sets in red are detected by application of our methodology but not by applying some classical microarray analysis methods individually. In (a) the probe set in red, 206011_at, is related to probe sets 211367_s_at and 211368_s_at (blue) both in expression throughout time and in function (apoptosis-related cysteine peptidase). In (b) we see the same situation between 202076_at in red and 210538_s_at in blue, which have correlated level of expression throughout time and share their function (inhibitor of apoptosis protein 2 and 1 respectively). The temporal expression data in our database can be averaged or sequentially represented for each biological replicate. Originally, the database was built based on the inflammatory response patterns, which is based on a very robust microarray experiment ([13]). Now, it is being updated with experiments provided from different sources such as the Ventilator Associated Pneumonia (unpublished results). The application of our methodology to the database of differential profiles allowed the optimal retrieval of the desired differential profiles. For example, if we were interested in probe exhibiting any of 27 of the profiles in the database and not exhibiting one of the profiles, we are able to do applying 65 MM ∪with specificity and sensitivity levels of 94% and 92% respectively. In Fig. 3 we show the method-aggregations from the optimal rules to retrieve individually each of the 28 profiles in our database. Fig. 3. Microarray analysis methods are able to retrieve some differential profiles and not others. Rows correspond to method-aggregations using the union operator and columns to each of the 28 individual differential profiles from our example. The coloring scheme corresponds to the sensitivity in retrieving the differential profile: from green, the lowest, to black, the highest. M6.M9.M10 M6.M10 M5.M6.M7.M8.M10 M5.M6 M4.M6.M8.M10 M4.M6.M10 M3.M8 M3.M4 M2.M6.M9 M2.M5.M6.M7.M8.M10 M2.M5 M2.M4.M7.M8.M10 M2.M4.M5.M6.M7.M10 M2.M3.M10 M10 M1.M6.M9 M1.M4.M6.M7.M8 M1.M4 M1.M2 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 Our approach also recovers probe sets with related behavior to other probe sets with already known profiles which might have related functionalities ([16]). For instance, probe set 206011_at is related in behavior and in function (apoptosis-related cysteine peptidase) to probe sets 211367_s_at and 211368_s_at (Fig. 2(a)), stated as relevant for the inflammation problem in [13]. For these particular probe set, the isolated application of classical methods such as 1 Mor 3 M with the default p-value and false discovery rate would not retrieve such probe set as differentially expressed. We retrieve such probe set applying the rule that implies the method aggregation 107MM ∪ with values (1, 0.25, 0.8) for sensitivity, specificity and cost respectively. The same situation applies to probe sets 202076_at and 210538_s_at, related both in behavior and in function (inhibitor of apoptosis protein 2 and 1 respectively) (Fig. 2(b)). It is retrieved applying the rule of methods 63 MM ∪with values (0.93, 0.35, 0.8). In addition, the representation used in the inflammation problem (Fig. 4) allows us to independently examine the gene behavior in each subject, helping to uncover individual tendencies among biological replicates that could represent conditions not previously considered such as gender or age (e.g., differential profile #15 (Fig. 5), where some of the probe sets from patient 1 exhibit a very different behavior than the rest of the patients). We illustrate the obtained association rules for Profile #19 from our database (Table 2) and the Pareto-optimal front for the three objectives corresponding to the selected rules (Fig. 6). Fig. 4. Profile #19: the expression profiles have been represented separately for each subject the experimental group and patients are arranged individually 3 Methods Most machine learning techniques are applied to mine into datasets to discover concepts involving objects which share a common methodological framework, even though they employ distinct metrics, heuristics or probability interpretations ([17], [18]): (1) identification of a database, different data types can be efficiently organized by taking advantage of a naturally occurring structure over feature space. (2) learning rules from the database, searching through the feature space for potential relationships 5000 10000 HOUR 02469240246924 0246924 0246924 PAT1 PAT2 PAT3 PAT4 5000 10000 HOUR 0246924 02924 0246924 0246924 PAT5 PAT6 PAT7 PAT8 0 2 4 6 9 24 5000 10000 0 2 4 6 9 24 5000 10000 among data, and either returning the best one found or an optimal sample of them. This learning process would result in the generation of many rules with small extent, as it is easier to explain or match small data subsets than those that constitute a significant portion of the dataset. For this reason, any successful methodology should also consider additional criteria ([19]) to extract broader or more comprehensive rules as a multiobjective optimization problem, based on their specificity, sensitivity and cost as a measures of the rule quality. (3) Inference, where new observations can be predicted from previously learned rules by using classifiers that optimize their matching to the rules based on distance ([18]) or probabilistic metrics ([20], [21]). We propose a method, inspired on conceptual clustering and optimization techniques ([9], [10], [17]), that identifies a database of gene profiles that change their expression over time and/or over treatments and/or over subjects, learns associations rules and make decisions about the microarray analysis method or the best aggregation of methods capable of detecting a desired set of differential profiles, and finally uses these rules to make decisions based on new situations (Fig.1). 0.5 1 1.5 2 2.5 x 10 4 PAT1 PAT2 PAT3 PAT4 0.5 1 1.5 2 2.5 x 10 4 PAT5 PAT6 PAT7 PAT8 Fig. 5. Profile #15: patient 1 behaves different than patients 2, 3 and 4 for the treatment group Table 2. Set of decision making association rules generated to retrieve Profile #19. The axes (X,Y,Z) represent the number of methods, specificity and sensitivity for each of the 35 solutions generated. RULES Sensitivity Specificity Cost R1:IF x1 IS ( PTPC )19 THEN Z1 IS M1 0.878378 0.0675676 0.9 R2:IF x1 IS ( PTPC )19 THEN Z2 IS M2 0.972973 0.045512 0.9 R3:IF x1 IS ( PTPC )19 THEN Z3 IS M7 0.905405 0.0475177 0.9 R4:IF x1 IS ( PTPC )19 THEN Z4 IS M1 ∩ M2 0.864865 0.0721533 0.8 R5:IF x1 IS ( PTPC )19 THEN Z5 IS M2 ∪ M10 1 0.0430733 0.8 R7:IF x1 IS ( PTPC )19 THEN Z7 IS M1 ∩ M10 0.594595 0.090535 0.8 R8:IF x1 IS ( PTPC )19 THEN Z8 IS M2 ∩ M7 0.891892 0.0538776 0.8 R9:IF x1 IS ( PTPC )19 THEN Z9 IS M3 ∩ M9 0.472973 0.100575 0.8 R10:IF x1 IS ( PTPC )19 THEN Z10 IS M3 ∩ M10 0.405405 0.104895 0.8 R11:IF x1 IS ( PTPC )19 THEN Z11 IS M1 ∩ M2 ∩ M7 0.797297 0.0732919 0.7 R12:IF x1 IS ( PTPC )19 THEN Z12 IS M1 ∩ M2 ∩ M9 0.662162 0.0853659 0.7 R13:IF x1 IS ( PTPC )19 THEN Z13 IS M1 ∩ M3 ∩ M7 ∩ M10 0.459459 0.112211 0.6 R14:IF x1 IS ( PTPC )19 THEN Z14 IS M3 ∩ M6 ∩ M9 ∩ M10 0.364865 0.135 0.6 R15:IF x1 IS ( PTPC )19 THEN Z15 IS M1 ∩ M3 ∩ M6 ∩ M9 ∩ M10 0.364865 0.140625 0.6 R16:IF x1 IS ( PTPC )19 THEN Z16 IS M1 ∩ M3 ∩ M4 ∩ M6 ∩ M9 ∩ M10 0.364865 0.141361 0.4 3.1 Identification of the Database Our database is composed of differential profiles obtained from the probe sets differentially expressed retrieved from the expression datasets. The probe sets are obtained using several classical microarray analysis methods. These methods include Student’s T-Test proposed in [11], with variants that distinguish changes in the abundance of RNA occurring not only over treatment but also over time; Permutation Test described in ([6]), also including a time approach; Analysis of Variance described in ([5]); and Longitudinal Data approach using Repeated Measures Analysis of Variance described in ([12]). Fig. 6. Pareto-front representation for the set of rules generated to retrieve Profile #19. The axes (X,Y,Z) represent the number of methods, specificity and sensitivity for each of the 35 solutions generated. The probe sets identified by the statistical methods serve as a means to create differential expression profiles (i.e., sets of genes with coordinate changes in RNA abundance) expressed from one experimental condition to the others (i.e., probe sets that change over time, treatments and phenotype). We group separately probe sets for different experimental conditions, treatment and PT control PC by applying the Kmeans clustering algorithm ([22]), which takes three input parameters: first, number of resulting clusters K, which is estimated by application of the Davies-Bouldin validity index ([23]); second, the similarity measure applied, Euclidean distance, which yields the best results in the clustering of this problem and third, the initialization strategy, random generation of the cluster centroids. Particularly, in our inflammation problem, we consider the temporal expression data sequentially represented for each biological replicate (i.e., patients in the same experimental group) instead of averaging them to uncover phenotype differences. We identify differential profiles by applying a coincidence index (CI) based on the hypergeometric distribution (p-value <0.05), which determines the statistical significance of overlap between pairwise profile association in treatment and control conditions ([16]): )/ 0 (1),( ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎝ ⎛ ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎝ ⎛ − − ∑=⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎝ ⎛ −= h g qn hq p qq h C P T PCI (1) that gives the chance probability of observing at least p candidates from a set T Pof size h within another set C Pof size n, in a universe of g candidates. Therefore, probe sets belonging to a cluster in treatment, T P, can fit in more than one cluster in control, Z X Y