Full text
ESCUELA T´ ECNICA SUPERIOR DE INGENIER´ IA INFORM´ ATICA GRADO EN INGENIER´ IA DE LA SALUD MENCI´ ON BIOINFORM´ ATICA Estudio y b´usqueda de marcadores gen´eticos mediante el uso de Deep Neural Networks Deep Neural Networks to find genetics signatures Realizado por Fernando Moreno Jabato Tutorizado por Jos´e Manuel Jerez Aragon´es Departamento Dpto. Lenguajes y Ciencias de la Comunicaci´on UNIVERSIDAD DE M´ ALAGA M´ ALAGA, Septiembre 2016 Fecha defensa: El Secretario del Tribunal
Keywords: omics, Machine Learning, Deep Learning, DNN, ANN, cancer, microarray, data mining, Big Data. Abstract This document contais the final dissertation ot the degree student Fernando Moreno Jabato for the studies Grade in Health Engeneering, speciality on Bioinformatics, of University of M´alaga. This dissertation have been performed with the supervision of Dr. Jos´e Manuel Jerez Aragon´es from the Departament of Lenguajes y Ciencias de la Comunicaci´on. The project title is Deep Neural Networks to find genetics signatures and is focused on the development of a bioinformatic tool oriented to identification of relationships between an attribute set and concret factor of interest on medicine. To do this, a tool was designed with the capacity of handle data sets from microarrays of different types. Microarrays was selected as preferent technology because it’s the most extended and accessible techonologie on health and biology fields nowadays. Once implemented the tool, an experiment was performed to evaluate the efficiency of this tool. The experiment uses prostate cancer related datasets from trascriptomics microarrays containing patients of prostate cancer and some normal individues. The results obtained in the experiment shows an improvement offered by the new Deep Learning algoritms (specifically, Deep Neural Networks) to analyze and obtain knowledgement from microarrays data. Besides, has been observed an improvement of efficiency and the beat of computational barriers that traditional Artifical Neural Networks suffered allowing apply this bioinformatics tools of new generation to masiva data sets. Palabras clave: ciencias ´omicas, Aprendizaje Computacional, Deep Learning, DNN, ANN, c´ancer, microarray, miner´ıa de datos, Big Data. Resumen Este documento contiene el Trabajo de Fin de Grado del alumno Fernando Moreno Jabato, estudiante del Grade in Health Engeneering, speciality on Bioinformatics, en la University of M´alaga. Dicho proyecto se ha realizado con la tutorizaci´on de Dr. Jos´e Manuel Jerez Aragon´es, profesor perteneciente al Departament of Lenguajes y Ciencias de la Comunicaci´on. El proyecto recive el t´ıtulo de Estudio y b´usqueda de marcadores gen´eticos mediante el uso de Deep Neural Networks y se centra en el desarrollo de una herramienta bioinform´atica orientada a la identificaci´on de relaciones entre un set de atributos y un factor concreto de inter´es en la medicina. Para ello se ha disa˜no una herramienta capaz de manejar datos procedentes de microarrays de diferentes tipos ya que es la tecnolog´ıa m´as expandida y accesible en la actualidad para este campo del conocimiento. Una vez implementada la herramienta se ha realizado un experimento para probar la eficacia de la misma. El experimento ha utlizado los resultados obtenidos de un microarray del ´ambito de la transcript´omica y el set de datos en cuesti´on correspond´ıa a un grupo de estudio con individuos normales y otros individuos que padecen de tumores de cancer de prostata.
Los resultados obtenidos en el experimento muestran una clara mejor´ıa de los nuevos algoritmos de Deep Learning, en concreto, de las Deep Neural Networks, para analizar y obtener conocimiento de datos obtenidos de microarrays. Adem´as se ha observado una mejora de la eficaciencia y la rotura de las barreras computacionales que los algoritmos tradicionales (Artificial Neural Networks, ANNs) padec´ıan, permitiendo poder aplicar estas herramientas bioinform´aticas de nueva generaci´on a conjuntos de datos masivos.
Contents 1 Introduction 1 1.1 Motivation.................................... 1 1.2 Stateoftheart ................................. 2 1.3 Objectives.................................... 3 1.4 Metodology ................................... 3 1.5 Licence...................................... 4 2 Problem study 5 2.1 Functional requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.2 Non-functional requirements . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.3 Software..................................... 6 3 Design 7 3.1 Datamanagement................................ 7 3.1.I Inputs .................................. 7 Dataloading .............................. 8 Attributesloading ........................... 8 Datasetsloading............................ 9 3.1.II Datadivision .............................. 10 KFolddivision.............................. 10 3.2 Variablehandling................................ 11 3.2.I Filtering................................. 11 Variable filtering and selection . . . . . . . . . . . . . . . . . . . . . 11 3.2.II Formula generation . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 Simpleformula ............................. 13 Complexformula ............................ 14 3.3 DNNhandling.................................. 15 Fit Deep Neural Network . . . . . . . . . . . . . . . . . . . . . . . . 15 4 Implementation 17 4.1 Datamanagement................................ 17 4.1.I Inputs .................................. 17 Dataloading .............................. 17 Attributesloading ........................... 18
Datasetloading ............................ 19 4.1.II Datadivision .............................. 20 KFolddivision.............................. 20 4.2 Variablehandling................................ 21 4.2.I Formula generation . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 Simpleformula ............................. 21 Complexformula ............................ 21 4.3 DNNhandling.................................. 22 Fit Deep Neural Network . . . . . . . . . . . . . . . . . . . . . . . . 22 4.4 Variablehandling(2).............................. 24 Variablefiltering ............................ 24 5 Guided experiment 29 5.1 Thedataset................................... 29 5.2 Motivationyobjectives............................. 30 5.3 Experimentexecution.............................. 30 5.4 Resultados.................................... 35 6 Project analysis and conclusions (EN) 41 6.1 Review of requirements compliance . . . . . . . . . . . . . . . . . . . . . . 41 6.2 Review of objectives compliance . . . . . . . . . . . . . . . . . . . . . . . . 42 6.3 Enhancement opportunities . . . . . . . . . . . . . . . . . . . . . . . . . . 43 6.4 Utilities and applicability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 6.5 Bioethical implications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 6.6 Futurelinesofwork............................... 45 7 An´alisis y conclusiones (ES) 46 7.1 Revision del cumplimiento de los requisitos . . . . . . . . . . . . . . . . . . 46 7.2 Revisi´on del cumplimiento de los objetivos . . . . . . . . . . . . . . . . . . 48 7.3 Oportunidades de mejora . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 7.4 Utilidades y aplicabilidad . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 7.5 Implicaciones bio´eticas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 7.6 L´ıneas futuras de trabajo . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 8 Concepts 52 8.1 Omicssciences ................................. 52 8.2 Deep Learning and Deep Neural Networks . . . . . . . . . . . . . . . . . . 52 8.3 Cross-validation method and KFold strategy . . . . . . . . . . . . . . . . . 53 8.4 Sensibility and specificity . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 8.5 Variablescorrelation .............................. 54 8.6 ROC curve and area under ROC curve (AUC) . . . . . . . . . . . . . . . . 54 8.7 Microarrays ................................... 55 9 Bibliography 56
1 - Introduction This document is a dissertation of the degree student Fernando Moreno Jabato for the studies in Grade in Health Engeneering, speciality on Bioinformatics, of University of M´alaga. The project has the title “Deep Neural Networks to find genetics signatures” wich wants prove the viability of use Deep Neural Networks (DNN) to identify relationships between genes and clinical symptoms and create new reliable diagnosis and prediction tools. To do this, in this project a specifical software has been implemented that can handle all needed variables for identifying genes relationship with a spcific issue and generate DNN models to make predictions using those genes. The entire project has been implement on R language. 1.1 Motivation Nowadays masive amounts of data are generated and stored constantly in all society fields. It’s known that this data can be stored on different variables and be studied to improve and obtain knowledgment in several fields like economy, sociology or medicine. Currently, this way of masive minning of information is known as Big Data and it’s being implemented on several business and social purposes companies. In Healt sector, data is being generated massively each day. On this information we can find diseases presences by location, blood type, metabolic activity registers of a specific patient and, increasingly, genomes, genotypes and other genomics data from specific patients. This new use of Omics on medicine are part of the new trend to personalized medicine that are increasing on the recent years. This personalized studies also can offer information useful for global health creating realtionships between data with symptoms, treatments or any other useful information to clinical medicine and society. Is on this point where Bioinformatics can offer tools for both medicine ways (personal and global). In first place, bioinformatics generate and contribute to diagnosis and decision making with biological data analysis tools. Nowadays the use of tools as electronic phonendoscopes or electrocardiographs with analysis software integrated, on hospitals is habitual. These two examples are a sample of the current tools that automate common process on clinical medicine that give the traditional information and several analysis of this information which generate more knowledgement that the sanitary can can use to making decisions. In other way, link this generated knowledgement is a newflanged concept that is being 1
implemented on other fields using ontologies which generate more knowledgement using reasoners already implemented. My project proposal is about the first way. I’ve observed a huge potential on genetic data to relate symptoms, inmunities or whatever other classifiable factor. That are the reasons because I propose implement a bioinformatical tool which can handle genetical data and relate it with our interest factors. This tool target will be the identification of high related variables on these genetic data sets to generate predictors and diagnostic tools which could be applied on medicine. To implement this tool I’ll use already known statistical methods to select and filter variables. To improve the current process I also will use some new machine learning algorythms known as Deep Learning algorythms. Particulary I will use Deep Neural Networks (DNN) because their excellent results in other fields like Automatic Speech Recognition (ASR)1or image identification2. Traditional Artificial Neural Networks (ANN) are already used to generate cancer predictors3or other medicine interest factors, that’s the reason because I look for improve the current tools using DNNs to generate new more precise predictors. If reults are good enough, this application would be a new tool to integrate bioinformatic on medicine and, also, to increase the knowledgement that Big Data and personal medicine can contribute to this one and to global health. 1.2 State of the art Currently, Deep Neural Networks have been used in several areas as speed automatic recognition(ASR)4or in image recognition5,6 as more remarkable examples. It’s difficult find publications of real applications of DNN in omics field. More easy is find publications about real applications of traditional artificial neural net1Geoffrey Hinton, et.al. “ Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups “. I EEE Signal Processing Magazine (Volume:29 , Issue: 6 ),Pages 8297,doi: 10.1109/MSP.2012.2205597 2Alex Krizhevsky, et.al. “ImageNet Classification with Deep Convolutional Neural Networks” 3Jos´e M. Jerez Aragon´es, et.al. “A combined neural network and decision trees model for prognosis of breast cancer relapse”. Artificial Intelligence in Medicine 27 (2003) 45–63. 4Geoffrey Hinton, et.al. “ Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups “. I EEE Signal Processing Magazine (Volume:29 , Issue: 6 ),Pages 8297,doi: 10.1109/MSP.2012.2205597 5Alex Krizhevsky, et.al. “ImageNet Classification with Deep Convolutional Neural Networks” 6Dan Ciregan, et.al. Multi-column deep neural networks for image classiffcation. Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on 16-21 June. doi 10.1109/CVPR.2012.6248110. 2
works (ANN) in omic fields7,8 or medicine fields9,10,11. Thats the reason because we look for improve the current results obtained in these studies using the new DNN technology as has been used in other fields. 1.3 Objectives This project objectives are: •Implement a support software to identify genetics signatures using DNNs. –Filtering and identification of genetics variables. –Generate DNN models. –GeneraTe preodictions using DNN model. •Research the reliability and effectiveness of DNN models generated to be used with genetic data. •Research effectiveness differences between Artificial Neural Networks (ANNs) and DNNs to find genetic signatures. Besides these particular objectives, there’re generals objectives for all dissertations. These are: •Interrelate different concepts learned during grade studies. •Perform a project related with one of the specialties of the grade studied. •Perform a project autonomously. 1.4 Metodology THe metodology that will be used on this project to implement the software and study the results will be the waterfall model. This model is a sequential (non-iterative) model where each project fase is not performed until the precedent is finished. Our project phases will be: 1.- Identify the project needs. 2.- Research about the State of the art. 3.- Design functional and non-function requirement. 4.- Logical design of the software. 5.- Software implementation. 6.- Software evaluation. 7.- Generating results. 7Jos´e M. Jerez Aragon´es, et.al. “A combined neural network and decision trees model for prognosis of breast cancer relapse”. Artificial Intelligence in Medicine 27 (2003) 45–63. 8Mateo Seregni, et.al. Real-time tumor tracking with an artffcial neural networksbased method: A feasibility study. Acta Medica Iranica, 2013 9Farin Soleimani, et.al. Predicting Developmental Disorder in Infants Using an Artiffcial Neural Network. Acta Medica Iranica51.6 (2013): 347-52. 10Hon-Yi Shi, et.al. Comparison of Artffcial Neural Network and Logistic Regression Models for Predicting In-Hospital Mortality after Primary Liver Cancer Surgery. 11Filippo Amato, et.al. Artiffcial neural networks in medical diagnosis. Journal of Applied Biomedicine, vol.11, issue 2, pages 47-58, 2013 3
8.- Analysis of results. These are the project phases to whitch is added in parallel the documentation development. 1.5 Licence This work is licensed under a Creative Commons “AttributionNonCommercial-NoDerivs 3.0 Unported” license. 4
3.2 Variable handling This section includes the variable handling functionalities. The functional requirements related to these funcionalities are: 1.- Filter genetic variables. (Filtering) 2.- Generation of formulas using a set of variables. (Formula generation) Each requisites will be studied separately. 3.2.I Filtering The software must can filter and select variables from a massive variable set. To do this a statistical filtering and selection function must be implemented. This function must return a subset with a configurable minimum and maximum size that satisfies some requirements. Variable filtering and selection This function must can filter variables using different statistical methods to calculate correlation between variables and the main factor. Necessary parameters to do this activity are: •Data: data set to be filtered and used to train DNN models. •Attribute set: attribute names set used to select and filter variables from data set. Default: data set given names. (OPTIONAL) •Filtering method: statistical filtering method to be used. (OtCIONAL) •Maximum size: maximum size of the returned attributes subset. •Minimum size: minimum size of the returned attributes subset. Default: zero. (OPTIONAL) •Group: indicates if the best minimum or best maximum subset is wanted. (OPTIONAL) •Store: variable where generated vector will be stored. •DNN store: variable where generated best DNN model will be stored. (OPTIONAL) •AUC: variable where AUC obtained on predictions over the main data set. (OtCIONAL) •Testing AUC: variable where AUC obtained on predictions over the testing data set. (OPTIONAL) •Collapse ratio: value used to decide if the subset model accuracy have collapsed. Default: 0.03. (OPTIONAL) •Correlation minimum: minimum correlation value with the main factor to accept an attribute. Default: 0.5. (OPTIONAL) •Testing data: testing data set. (OPTIONAL) •Activation: activation function that will be used in DNNs training. Default: Tanh. Possible values: “Tanh”, “TanhWithDropout”, “Rectifier”, “RectifierWithDropout”, “Maxout” or “MaxoutWithDropout”. (OPTIONAL) •Neurons: hidden neural layers that will be used on DNNs. 11
•Iterations: iterations that will be done over training data set on DNN train process. Default: 100. (OPTIONAL) •Seed: seed used to start the DNN training process. (OPTIONAL) •Rho: factor of learning ratio adaptation on time. Used on DNN training. (OPTIONAL) •Epsilon: adaptative learning ratio used on DNNs training process. (OPTIONAL) •Threads: number of CPU threads that will be used. Default: 2. (OPTIONAL) Functios must return the following values: •Subset generated size. •ERROR IN DATA PARAMETER. •ERROR IN TRAINING PARAMETER. •ERROR IN FILTERING PARAMETER. •ERROR IN GROUP PARAMETER. •ERROR IN MAXIMUM PARAMETER. •ERROR IN MINIMUM PARAMETER. •ERROR IN ATTRIBUTE PARAMETER. •ERROR MAXIMUM LESS THAN MINIMUM. •IMPOSIBLE TO OBTAIN SATISFACTORY SUBSET. •Signals thrown by DNNs function.. Function use case will be: 1.- If data parameter is null: abort process and throw ERROR IN DATA PARAMETER signal. 2.- If filtering parameter is not allowed (unless it’s null): abort process and throw ERROR IN FILTERING PARAMETER. 3.- If group parameter is not allowed (unless it’s null): abort process and throw ERROR IN GROUP PARAMETER signal. 4.- If maximum attribute is less than one: abort process and throw ERROR IN MAXIMUM PARAMETER signal. 5.- If minimum parameter is negative: abort process and throw ERROR IN MINIMUM PARAMETER signal. 6.- If minimum parameter is greater than maximum parameter: abort process and throw ERROR MAXIMUM LESS THAN MINIMUM signal. 7.- If attributes parameter is null: abort process and throw ERROR ATTRIBUTE PARAMETER signal. 8.- Filter variable set using selected statistical filtering method. 9.- Sort obtained results by correlation value. 10.- Select attributes with a correlation value greater than minimum given. 11.- If subset size is greater than maximum given: obtein maximum subset possible. 12.- Check group type selected: •Case Best Minimum: 12.1.- Obtain minimum subset. 12.2.- Generate model with minimum subset. 12.2.1.- If any exception is thrown: throw exception to upper level. 12.3.- Evaluate generated model. 12
12.4.- Store evaluation value on current best model value variable. 12.5.- Add next best attribute to formula. 12.6.- Generate model with current subset. 12.6.1.- If any exception is thrown: throw exception to upper level. 12.7.- Evaluate generated model. 12.8.- If: –Upgrading greater than collapse ratio: go to 12.5. –Upgradin decrease or less than collapse ratio: continue. 12.9.- Delete last attribute added. 12.10.- Store results on best result variables. •Case Best Maximum: 12.1.- Obtain maximum subset. 12.2.- Generate model with maximum subset. 12.2.1.- If any exception is thrown: throw exception to upper level. 12.3.- Evaluate generated model. 12.4.- Store evaluation value on current best model value variable. 12.5.- Delete worst attribute. 12.6.- Generate model with current subset. 12.6.1.- If any exception is thrown: throw exception to upper level. 12.7.- Evaluate generated model. 12.8.- If: –Upgrading greater than collapse ratio: go to 12.5. –Upgradin decrease or less than collapse ratio: continue. 12.9.- Add last deleted attribute. 12.10.- Store results on best result variables. 13.- Store subset selected on store variable given. 14.- Store best DNN model generated on given store variable. 15.- Store AUC values obtained on store variables given. 16.- Return generated subset size. 3.2.II Formula generation The software mus can generate formula instances using an attribute set given. An R formula can have a high complexity, in the current version only addtion an deletion formula types will be implemented. Simple formula This function must generate instances of formula using a given attribute set. This formulas only will have addition factors. The necessary parameters to do this activity are: •Main factor: main attribute of the formula. 13
•Attributes: attribute set. Functios must return the following values: •Generated formula. •ERROR IN MAIN FACTOR PARAMETER. •ERROR IN ATTRIBUTES PARAMETER. Function use case will be: 1.- If main factor paremeter is null: abort process and throw ERROR IN MAIN FACTOR PARAMETER signal. 2.- If main factor isn’t text: abort process and throw ERROR IN MAIN FACTOR PARAMETER signal. 3.- If attributes parameter is null: abort process and throw ERROR IN ATTRIBUTES PARAMETER. 4.- If attributes aprameter isn’t a string vector: abort process and throw ERROR IN ATTRIBUTES PARAMETER. 5.- Generate string with main factor and attributes using the correct mathematical sign. 6.- Parse string to formula. 7.- Return generated formula. Complex formula This function must generate instances of formula using a given attribute set. This formulas only will have addition and sustraction factors. The necessary parameters to do this activity are: •Main factor: main attribute of the formula. •Attributes: attribute set. •Signs: signs set used to do a addition or sustraction factor. Functios must return the following values: •Generated formula. •ERROR IN MAIN FACTOR PARAMETER. •ERROR IN ATTRIBUTES PARAMETER. •ERROR IN SIGNS PARAMETER Function use case will be: 1.- If main factor paremeter is null: abort process and throw ERROR IN MAIN FACTOR PARAMETER signal. 2.- If main factor isn’t text: abort process and throw ERROR IN MAIN FACTOR PARAMETER signal. 3.- If attributes parameter is null: abort process and throw ERROR IN ATTRIBUTES PARAMETER. 4.- If attributes aprameter isn’t a string vector: abort process and throw ERROR IN ATTRIBUTES PARAMETER. 5.- If signs parameter is null: abort process and throw ERROR IN SIGNS PARAMETER signal. 6.- If signs parameter isn’t an integer array: abort process and throw ERROR IN SIGNS PARAMETER signal. 14
7.- Generate string with main factor and attributes using the correct mathematical sign. 8.- Parse string to formula. 9.- Return generated formula. 3.3 DNN handling This section includes the DNN handling functionalities. The functional requirements related to these funcionalities are: 1.- Generate a Deep Neural Network using a formula. To do this, the function must train a DNN model using a given formula and training data set using the bigger relevant parameters set for DNN learning process configuration. Also must evalueate predicitons done over the training data set and a testing data set if it’s possible returning these evaluations. The function must follow this model: Fit Deep Neural Network The necessary parameters to do this activity are: •Formula: formula used to generate the DNN model. •Data: training data set. •Testing data: testing data set. (OPTIONAL) •Activation: activation function that will be used in DNNs training. Default: Tanh. Possible values: “Tanh”, “TanhWithDropout”, “Rectifier”, “RectifierWithDropout”, “Maxout” or “MaxoutWithDropout”. (OPTIONAL) •Neurons: hidden neural layers that will be used on DNNs. •Iterations: iterations that will be done over training data set on DNN train process. Default: 100. (OPTIONAL) •Seed: seed used to start the DNN training process. (OPTIONAL) •Rho: factor of learning ratio adaptation on time. Used on DNN training. (OPTIONAL) •Epsilon: adaptative learning ratio used on DNNs training process. (OPTIONAL) •Threads: number of CPU threads that will be used. Default: 2. (OPTIONAL) •AUC: variable where AUC obtained on predictions over training data set. (OtCIONAL) •Testing AUC: variable where AUC obtained on predictions over the testing data set. (OPTIONAL) Functios must return the following values: •DNN generated model. •ERROR IN FORMULA PARAMETER. •ERROR IN DATA PARAMETER. •ERROR IN TEST PARAMETER. •ERROR IN ACTIVATION PARAMETER. •ERROR IN NEURONS PARAMETER. 15
•ERROR IN ITERATIONS PARAMETER. •ERROR IN SEED PARAMETER. •ERROR IN RHO PARAMETER. •ERROR IN EPSILON PARAMETER. •ERROR IN THREADS PARAMETER. •ERROR GENERATING MODEL. Function use case will be: 1.- If formula parameter isn’t a formula instance: abort process and throw ERROR IN FORMULA PARAMETER signal. 2.- If formula parameter doesn’t contains the main factor: abort process and throw ERROR IN FORMULA PARAMETER signal. 3.- If data parameter is null: abort process and throe ERRO IN DATA PARAMETER signal. 4.- If main factor isn’t on training data set: abort process and throe ERRO IN DATA PARAMETER signal. 5.- If formula attributes aren’t on training data set: abort process and throe ERRO IN DATA PARAMETER signal. 6.- If testing parameter is null: abort process and throw ERROR IN TEST PARAMETER signal. 7.- If main factor isn’t on testing data set: abort process and throw ERROR IN TEST PARAMETER signal. 8.- If formula attributes aren’t on testing data set: abort process and throw ERROR IN TEST PARAMETER signal. 9.- If neurons parameter is less than one or null: abort process and throw ERROR IN NEURONS PARAMETER signal. 10.- If iterations parameter is less than one: abort process and throw ERROR IN ITERATIONS PARAMETER signal. 11.- If seed parameter is negative: abort process and throw ERROR IN SEED PARAMETER signal. 12.- If rho parameter is negative: abort process and throw ERROR IN RHO PARAMETER signal. 13.- If epsilon parameter is negative: abort process and throw ERROR IN EPSILON PARAMETER signal. 14.- If threads parameter is lees than 1: abort process and throw ERROR IN THREADS PARAMETER signal. 15.- Generate model using the training data set. 15.1.- If any exception is thrown: abort process and throw ERROR GENERATING MODEL signal. 16.- Evaluate model making predictions over training set. 17.- Store AUC obtained on store variable given. 18.- If testing and testing AUC parameters are not null: 18.1.- Evaluate model making predictions over testing set. 18.2.- Store testing AUC obtained on testing AUC store variable given. 19.- Return generated model. 16
4 - Implementation This chapter includes the software implementation based on the logical design described in the previous chapter. The full implementatios have been done using R language and grouping functions in files by functionality. The implementation order is the same used on design chapter unless for functions with dependencies of other functions. The order will be: •Data management: –Inputs. –Data division. •Variable handling: –Formula generation. •DNN handling: –Fit Deep Neural Network. •Variable handling (2): –Variable filtering (Dependencies: formula generation and DNN handling). Remark: it’s important to know that all functions have the package dependecy of R.oo. 4.1 Data management The name of the file used to store the data mangement implementation is funcs data management.R and will includes the following functions: 4.1.I Inputs Data loading To implement data loading function, the designed interface will be implemented. The error signals will be implemented as exceptions that will be thrown and the FINALIZATION WITHOUT ERRORS signal will be just implemented as a normal function without return anything. The function name will be read.dataset.table and the implementation is: 1read.dataset.table <- function (dir ="", 2file , 3store = NULL , 17
4dir . sep = "/", 5file. sep = " ,"){ 6# CHECK ATTRIBUTES 7if(is .null(dir )){ 8throw ( ’DIR is not specified ( NULL ) ’) 9} 10 11 if(is .null(file )){ 12 throw ( ’FILE is not specified ( NULL ) ’) 13 } 14 15 # VARIABLES 16 data.table <- "" 17 data <- "" 18 19 # READ AND PARSE 20 data.table <- read .table (file=paste (dir ,file ,sep=dir.sep ) ,sep = file. sep ) 21 data <- data.frame (data.table )# Store data 22 23 # FREE UNNECESSARY SPACE 24 rm(data.table ) 25 26 # STORE 27 eval.parent(substitute ( store <- data)) 28 } Attributes loading To implement attributes loading function, the designed interface will be implemented. The error signals will be implemented as exceptions that will be thrown and the FINALIZATION WITHOUT ERRORS signal will be just implemented as a normal function without return anything. The function name will be read.dataset.attr and the implementation is: 1read.dataset.attr <- function (dir ="" , 2file , 3store = NULL , 4column = TRUE , 5dir . sep = "/", 6file. sep = " ,"){ 7# CHECK ATTRIBUTES 8if(is .null(dir )){ # Check directory 9throw ( ’DIR is not specified ( NULL ) ’) 10 } 11 12 if(is .null(file )){ # Check file 13 throw ( ’FILE is not specified ( NULL ) ’) 14 } 15 16 if(!is.logical( column )){ # Check column 17 throw ( ’COLUMN isn \’ta boolean value ’) 18 } 19 20 # VARIABLES 21 attr . table <- "" 22 attr <- "" 23 24 # READ AND PARSE 25 attr . table <- read . table ( file = paste (dir , file , sep = dir . sep ), sep = file . sep ) 26 if ( column ) 27 attr <- rownames ( attr . table ) # Store attribute names 28 else 18
29 attr <- colnames ( attr . table ) # Store attribute names 30 31 # FREE UNNECESSARY SPACE 32 rm ( attr . table ) 33 34 # STORE 35 eval . parent ( substitute ( store <- attr )) 36 } Data set loading To implement data sets loading function, the designed interface will be implemented. The error signals will be implemented as exceptions that will be thrown and the FINALIZATION WITHOUT ERRORS signal will be just implemented as a normal function without return anything. The function name will be read.dataset and the implementation is: 1read.dataset <- function (dir.data , 2file.test , 3file.train , 4file.attr = NULL , 5store . test = NULL , 6store . train = NULL , 7store . attr = NULL , 8dir . sep = "/", 9file. test . sep = ",", 10 file. train . sep = ",", 11 file.attr. sep = ":", 12 file.attr. column = TRUE ){ 13 14 # CHECK ATTRIBUTES 15 if(is .null(dir )){ # Check directory 16 throw ( ’DIR is not specified ( NULL ) ’) 17 } 18 19 if(is .null(file . test ) | is.null (file . train ) ){ 20 throw ( ’TEST or TRAINING files are not correct files path ’) 21 } 22 23 if(!is.logical(file.attr. column )){ # Check column 24 throw ( ’COLUMN isn \’ta boolean value ’) 25 } 26 27 # VARIABLES 28 data.test <- "" 29 data . train <- "" 30 data.attr <- "" 31 32 33 # READ FILES 34 # Training file 35 read . dataset . table ( dir =dir .data , 36 file =file .train , 37 store = data . train , 38 dir . sep = dir .sep , 39 file . sep = file . train . sep ) 40 41 42 # Test file 43 read . dataset . table ( dir =dir .data , 44 file =file .test , 45 store = data . test , 19
46 dir . sep = dir .sep , 47 file .sep = file . test . sep ) 48 49 if(!is.null (file .attr )){ 50 read . dataset . attr ( dir = dir .data , 51 file =file .attr , 52 store = data . attr , 53 dir . sep = dir .sep , 54 file .sep = file . attr .sep , 55 column = file . attr . column ) 56 57 names ( data . train ) <- data . attr # Give attribute names 58 names ( data . test ) <- data . attr # Give attribute names 59 } 60 61 62 # Set read values 63 eval . parent ( substitute ( store . test <- data . test ) ) 64 65 eval . parent ( substitute ( store . train <- data . train )) 66 67 if(!is.null (data .attr )) 68 eval . parent ( substitute ( store . attr <- data . attr ) ) 69 } 4.1.II Data division KFold division To implement kfold division function, the designed interface will be implemented. The error signals will be implemented as exceptions that will be thrown and returning the subsets vector generated in case the program end without errors. The function name will be kfold and the implementation is: 1kfold <- function (data , 2folds , 3k = 10) { 4# CHECK ATTRIBUTES 5if(is .null(data )){ 6throw ( ’Data is null ’) 7} 8 9if(k <=0) { 10 throw ( ’K is negative or zero ’) 11 } 12 13 if(is .null( folds )){ 14 throw ( ’Store variable given is null ’) 15 } 16 17 # GENERATE PARTITIONS 18 generatedFolds <- createFolds(data , k=k) 19 eval.parent(substitute ( folds <- generatedFolds)) 20 } This functios has a package dependecy: caret. 20
155 store . auc . train = new.AUC . train , 156 connect = FALSE ) 157 if(new . AUC . train - current . AUC . train < collapse ){ # Collapsed 158 bestmin <- current . best [c(1:(length( current . best ) -1)) ] 159 break ; 160 }else{# Update values 161 current . AUC . train <- new . AUC . train 162 current . AUC . test <- new . AUC . test 163 current.dnn <- new.dnn 164 } 165 } 166 167 BEST <- bestmin # STORE 168 BEST .dnn <- current.dnn 169 }else{# BestMax 170 # Select initial minimum set 171 if(max. size <= length( correlation . selection )){ 172 bestmax = correlation . selection [c(1:max . size )] 173 }else{ 174 bestmax = correlation . selection 175 } 176 177 # Generate DNN and calculate AUC ( firsttime , initialize ) 178 current.formula <- variable .formula. generator ( main . class,names (bestmax)) 179 current.dnn <- dnn ( formula = current.formula , 180 data =data , 181 data. test = testingset , 182 hidden = hidden , 183 store . auc . test = current . AUC . test , 184 store . auc . train = current . AUC . train , 185 connect = FALSE ) 186 187 current . best <- bestmax 188 new . AUC . train <- "" 189 new . AUC . test <- "" 190 new . dnn <- "" 191 # Start to search 192 while (length( current . best ) > min .size & length( current . best ) > 1){ 193 current . best <- correlation . selection [c(1:(length( current . best ) -1) )] # Delete a variable variable 194 current.formula <- variable .formula. generator ( main . class,names ( current . best )) # Generate formula 195 new . dnn <- dnn ( formula = current.formula ,# Generate DNN 196 data =data , 197 data. test = testingset , 198 hidden = hidden , 199 store . auc . test = new. AUC .test , 200 store . auc . train = new.AUC . train , 201 connect = FALSE ) 202 if(new . AUC . train - current . AUC . train < collapse ){ # Collapsed 203 bestmin <- correlation . selection [c(1:(length( current . best ) +1) )] 204 break ; 205 }else{# Update values 206 current . AUC . train <- new . AUC . train 207 current . AUC . test <- new . AUC . test 208 current.dnn <- new.dnn 209 } 210 } 211 212 BEST <- bestmax # STORE 213 BEST .dnn <- current.dnn 214 } 215 216 # STORE 217 eval.parent(substitute ( store <- BEST )) 218 if(!is.null( dnn . store )) { 27
219 eval.parent(substitute ( dnn . store <- BEST . dnn )) 220 } 221 if(!is.null( AUC . store )) { 222 eval.parent(substitute ( AUC . store <- current . AUC . train ) ) 223 } 224 225 if(!is.null( testingset ) & !is.null( AUC . test . store )){ 226 eval.parent(substitute ( AUC . test . store <- current . AUC . test )) 227 } 228 229 # Close H2O connection 230 h2o . shutdown ( prompt = FALSE ) # Close H2O environment 231 232 # RETURN FINAL SELECTION VARIABLE 233 return(length( BEST )) 234 } 28
5 - Guided experiment This chapter includes a guiged experiment to test the implemented tool. Also, at the end, eficiency of DDNs and ANNs will be compared. 5.1 The data set In this experiment several individues have been studied using microarrays technology. These individues have been catalogated as normal or tumor classes. This patients (tumor class) have developed a tumor related to prostate cancer. We will understand tumor instances as patients that have developed a prostate cancer tumor and we will understand normal instances as people that haven’t developed a prostate cancer tumor. Not other information about normal individues is known. The microarray has been designed to study a group of candidates gen to be related with prostate cancer tumoring. The total amount of genes studied are 12600. The transcript gene levels read from microarray have been ponderated and stored on data sets with several files. The data set was generated to be used in machine learning experiments, for that reason the files included are: •prostate TumorVSNormal test.data: includes a data set that must be used to test models (gold pattern). Contains 34 patients instances: – Normal: 34. – Tumor: 9. •prostate TumorVSNormal train.data: includes a data set that must be used to train models. Contains 102 patients instances: – Normal: 50. – Tumor: 52. •prostate TumorVSNormal.names: contains metadata about data set atributes. In this case, 12600 are transcripts levels and are of type continuous, there are one more that is the main class (tumor or normal) attribute, this is of type factor. The first file will be used to evaluate the accuracy of generated predictors and models. The second one will be the main data set and will be used to filter, select, and train DNN models. The last one will be used to obtain the gene names and link it to the other two data sets. 29
5.2 Motivation y objectives Prostate cancer is, nowadays, the most frequent cancer on males, except of skin cancer. Most of prostate cancer are glandular cancers (starts at mocusoa and other secreters cells) and commonly affect males with ages over 65 years. The major prostate cancer diagnosis problem is that, usually, doen’t show early symptoms and the evolution is slow. making harder the diagnosis and detection. For that reason, in medicine, some known studies are used to classify the risk leves of patients. One of the main factors used to classify the prostate cancer risk (or other prostate alterations) is the Prostatic Specific Antigen (PSA) levels in blood. PSA is a protein substance specific of prostate. The presence of PSA in blod is low for healthy patients but it encreases for prostate cancer patients or for other prostate alterations. Besides, gene alteration have been observed for prostate cancer patients. Some examples are the lost of p53, amplification of MYC or lost of PTEN. Our objectives ar focused on predict prostate cancer risk level using gene expression levels. Genes selected as candidates are those have shown alterations on prostate cancer patients. Our target is identify and calculate relationships between these expression levels and tumor development risks. To do this, the tool implemented on this project will be used. Besides, we will use the results obtained to generate DNN and ANN models and to compare their reliability on this kind of experiments. 5.3 Experiment execution Our experiment workflow will be the following: 1.- Data sets loading. 2.- Filter&Select variables. 3.- Find best training parameters for train a DNN with our data set. 4.- Efficiency study of DNN in our experiment. 5.- Efficiency study of ANN in our experiment. Before start, there are several variables that have been instantiated to avoid write big lines of code or to not repeat the same information several times. These variables are: 1project . root . dir <- "~PERSONALPATH/tfg" 2project . fun . dir <- "~PERSONALPATH/tfg /src/funcs /" 3project.data.dir <- "~PERSONALPATH/tfg/data/ProstateCancer" 4test.file <- " prostate _TumorVSNormal_test . data " 5train . file <- " prostate _TumorVSNormal_train . data " 6attr.file <- " prostate _TumorVSNormal . names " To begin, our project path will be stablished as working directory on our R engine. Also, necessary packages will be loaded and a seed will be selected to permit experiment replications. To do this we will use these commands: 1# Set directory 2setwd ( project . root . dir) 3## Set a seed for experiment replications 30
4set . seed (1200) 5 6## IMPORT necessary libraries 7library( caret ) ## Friendly way to create boxes 8library( FSelector ) ## Var selection functions 9library( h2o) ## Library for DNNs 10 library(R.oo) ## Oriented Objects Now, implemented functions will be loaded. All functions files have .R extension, for that reason, load all of them is as easy as use this command: 1## Load necessary external functions 2sapply(list. files ( pattern ="[.]R$",path= project . fun .dir, full . names = TRUE ), source) With this, all functios are loaded and ready to be used. The next step is instantiate and initialize necessary variables to load data sets. We have to do it because our functions overwrite this variables functions but, if they’re not initilized, a null pointer exception is thrown. 1# Declare data variables 2test.data <- "" 3train . data <- "" 4attr.data <- "" Now we can load our data set: 1# Load data 2read.dataset(dir.data = project.data.dir, 3file.test = test.file, 4file. train = train . file , 5file.attr =attr.file , 6store . test = test . data , 7store . train = train . data, 8store . attr =attr.data) For our dataset there are a problem with gene names. Several gene names contains reserved characters or starts with a digit, for that reason a specific function has been implemented to swap this character for other characters allowed. This function is called variable.name.correcto and have been implemented on file funcs variable handling.R to be used on other experiments. The implementatios is: 1variable .names . corrector <- function (names ){ 2# VARIABLES 3names .aux <- "" 4# CORRECT TEXT 5names .aux <- gsub("-","__",names ) 6names .aux <- gsub("/",".",names . aux ) 7names .aux <- gsub(" ","",names . aux ) 8# STORE 9eval.parent(substitute (names <- names . aux )) 10 } Now we can change special characters and add the prefix GG to the gene names (not to Class name). To do this we will launch the following set of commads: 1# Correct posible errors 2variable .names . corrector ( attr.data) 3for (i in 1: length(attr.data)) # Add GG_to all genes 31
4if(attr.data[i] != " Class " ) 5attr.data[i] <- paste ("GG_",attr .data[i],sep = "") 6 7# Restore names 8names (test. data)<- attr.data 9names ( train . data)<- attr .data 10 colnames ( test . data )<- attr.data 11 colnames ( train . data)<- attr.data Until now, we have loaded the data sets and have overwritting special characters. If we want know more information about our dataset we can use the R function summary() but it’s not recommendable because it will shown 12601 information column entries. Now we can filter our attribute data set. To do this we ned instantiate and initialize some variables for the same reson than before. 1# Result variables 2var . selected . max <- "" 3var . selected . min <- "" 4AUC . train . max <- "" 5AUC . train . min <- "" 6AUC . test . max <- "" 7AUC . test . min <- "" The next step is call the selector function. To this experiment we will call to the selecto using the minimum set of parameters to observe the efficiency without adjust all parameters of DNN training. Also we will call the tool in his two modes (BestMin and BestMax). The commands to call selector in both modes are: 1# FILTER AND SELECT 2# Best max 3variable . selector ( data = train . data , 4variables = attr .data , 5main.class =" Class " , 6testingset = test .data , 7extra . eval = NULL , 8filter.method = " ChiSquared ", 9max . size = 200 , 10 store = var. selected . max, 11 group = "BestMax", 12 AUC . store = AUC . train . max, 13 AUC . test . store = AUC . test . max) 14 15 # Best min 16 variable . selector ( data = train . data , 17 variables = attr .data , 18 main.class =" Class " , 19 testingset = test .data , 20 extra . eval = NULL , 21 filter.method = " ChiSquared ", 22 max . size = 200 , 23 store = var. selected . min, 24 group = "BestMin", 25 AUC . store = AUC . train . min, 26 AUC . test . store = AUC . test . min) Results can be found on Results section, following this one. As you can see, default values has been used in most of the configurable attributes letting H2O environment select hidden neurons layers size to our data set. Finally, the minimum best gene list obtained is: 32
1[1] "GG_37639 _at" "GG_32598 _at" "GG_38406 _f_at" "GG _41288 _at" "GG _37720 _at" "GG _ 38634 _at" "GG _37366 _at" 2[8] "GG_40282 _s_at" "GG_40856 _at" "GG_41468 _at" "GG _39031 _at" And the maximum best is: 1[1] "GG_37639 _at" "GG_32598 _at" "GG_38406 _f_at" "GG _41288 _at" "GG _37720 _at" "GG _ 38634 _at" "GG _37366 _at" 2[8] "GG_40282 _s_at" "GG_40856 _at" "GG_41468 _at" "GG _39031 _at" "GG _556_s_at" "GG_ 32243 _g_at" "GG_31538 _at" 3[15] "GG_1767_s_at " "GG_39315 _at " "GG_575_s_at " "GG_36601 _at" "GG _36491 _at" "GG _ 38028 _at" "GG _39545 _at" 4[22] "GG_35702 _at" "GG _37068 _at" "GG _40436 _g_at" "GG_40435 _at" "GG_36533 _at" "GG _ 34678 _at" "GG _33121 _g_at" 5[29] "GG_38044 _at" "GG _39939 _at" "GG _37251 _s_at" "GG_31527 _at" "GG_36666 _at" "GG _ 31444 _s_at" "GG_914 _g_at" 6[36] "GG_36589 _at" "GG _39756 _g_at" "GG _37754 _at" "GG_41385 _at" "GG_39755 _at" "GG _ 35742 _at" "GG _34950 _at" 7[43] "GG_32206 _at" "GG _38042 _at" "GG _33198 _at" "GG_581_at" "GG _34840 _at" "GG _ 36569 _at" "GG _216_at " 8[50] "GG_33614 _at" "GG _33674 _at" "GG _41104 _at" "GG_39243 _s_at" "GG_38291 _at" "GG _ 34820 _at" "GG _36495 _at" 9[57] "GG_34775 _at" "GG _38338 _at" "GG _33819 _at" "GG_769_s_at" "GG _31568 _at" "GG _ 41755 _at" "GG _33668 _at" 10 [64] "GG_2046_at " "GG_32718 _at " "GG_39123 _s_at" "GG _33396 _at" "GG _1664_at " "GG_ 37005 _at" "GG _32076 _at" 11 [71] "GG_31545 _at" "GG _35644 _at" "GG _34592 _at" "GG_38814 _at" "GG_36149 _at" "GG _ 38950 _r_at" "GG_291 _s_at" 12 [78] "GG_38827 _at" "GG _33412 _at" "GG _256_s_at" "GG _37203 _at" "GG _36780 _at" "GG _ 35277 _at" "GG _40607 _at" 13 [85] "GG_1740_g_at " "GG_33904 _at " "GG_38642 _at" "GG _1676_s_at " "GG_33820 _g_at" "GG _ 40071 _at" "GG _38669 _at" 14 [92] "GG_37716 _at" "GG _37000 _at" "GG _34407 _at" "GG_39054 _at" "GG_32695 _at" "GG _ 1736_at" "GG_39341 _at" 15 [99] "GG_36555 _at" "GG _38087 _s_at" "GG _34369 _at" "GG_40567 _at" "GG_37406 _at" "GG _ 38057 _at" "GG _33362 _at" 16 [106] "GG_40024 _at" "GG _37741 _at" "GG _36030 _at" "GG_34608 _at" "GG_41530 _at" "GG _37630 _at" "GG _38408 _at" 17 [113] "GG_1513_at " "GG_41504 _s_at " "GG _38322 _at" "GG _32242 _at" "GG _40125 _at" "GG _33137 _at" "GG _41485 _at" 18 [120] "GG_1276_g_at " "GG_41178 _at " "GG_1521_at " "GG_1897_at" "GG _39830 _at" "GG _33408 _at" "GG _39366 _at" 19 [127] "GG_37347 _at" "GG _34784 _at" "GG _41768 _at" "GG_37743 _at" "GG_39634 _at" "GG _40301 _at" "GG _41242 _at" 20 [134] "GG_34791 _at" "GG _36624 _at" "GG _38051 _at" "GG_36587 _at" "GG_36786 _at" "GG _31907 _at" "GG _2041_i_at " 21 [141] "GG_38279 _at" "GG _36943 _r_at" "GG _38429_at" "GG _36095 _at" "GG _31385 _at" "GG _40060 _r_at" "GG _1831_at " 22 [148] "GG_942_at" "GG_39750 _at" "GG_37573 _at" "GG_38435 _at " "GG_35146 _at " "GG _38740 _at" "GG _32780 _at" 23 [155] "GG_41013 _at" "GG _1980_s_at " "GG _39551 _at" "GG _37929 _at" "GG _33355 _at" "GG _36918 _at" "GG _38780 _at" 24 [162] "GG_38098 _at" "GG _496_s_at" "GG_32435 _at" "GG _829 _s_at" "GG_34646 _at" "GG _36668 _at" "GG _254_at" 25 [169] "GG_41732 _at" "GG _37035 _at" "GG _33716 _at" "GG_32109 _at" "GG_37958 _at" "GG _37582 _at" "GG _36864 _at" 26 [176] "GG_33405 _at" "GG _38033 _at" "GG _1257 _s_at" "GG_32315 _at" "GG _31583 _at" "GG _32123 _at" "GG _33328 _at" 27 [183] "GG_35354 _at" "GG _39798 _at" "GG _38410 _at" "GG_31791 _at" "GG_1708_at" "GG _38385 _at" "GG _33134 _at" 28 [190] "GG_32667 _at" "GG _34304 _s_at" "GG _41531_at" "GG _34376 _at" "GG _38986 _at" "GG _32800 _at" "GG _41214 _at" 29 [197] "GG_39070 _at" "GG _40063 _at" "GG _37708 _r_at" "GG_32412 _at" Now, to compare ANN efficiency against DNNs, we will train a traditional artificial neural network. 33
To do this, we will use nnet package to generate ANN models and AUC package to calculate AUC of generated predictions. For that reason we have to load these packages: 1# Import necessary packages 2library( nnet ) ## ANNs tools 3library( AUC) ## AUC and ROC functions Current step consists on generate a formula with the filtered attribute set and train the ANN model. To do this we use nnet function from nnet loaded package. Remark: several experiments have been performed to find the best configuration that is shown below. 1# Generate formula 2best.formula.max <- variable .formula. generator (" Class " ,var . selected . max ) 3best.formula.min <- variable .formula. generator (" Class " ,var . selected . min ) 4 5# Fit ANN 6nn.fit . min <- nnet(formula= best . formula.min ,# our model 7data= train . data,# our training set 8size =50 , # number of hidden neurons 9maxit =1000 , # max number of iterances to try converge 10 decay =5e -8, # avoid overfitting ( value : a little more than 0) 11 trace = FALSE ) # don ’t print the process 12 13 nn.fit . max <- nnet(formula= best . formula.max ,# our model 14 data= train . data,# our training set 15 size =4 , # number of hidden neurons 16 maxit =1000 , # max number of iterances to try converge 17 decay =5e -4, # avoid overfitting ( value : a little more than 0) 18 trace = FALSE ) # don ’t print the process In maximum set case we can boserve that neuron number is 4. That is because for higger values a too high amount of weights exception is thrown. That show the inability of the tool to handle big data sets. Now we evaluate the AUC for predictions over training and testing data sets: 1# Training set 2nn.pred . min . train <- predict( nn.fit .min, train . data , type = ’raw ’) 3nn.roc . min. train <- roc (as.numeric( nn. pred . min. train ) ,train . data [," Class " ]) 4nn.auc . min. train <- auc (nn.roc . min. train ) 5nn.pred . max . train <- predict( nn.fit .max, train . data , type = ’raw ’) 6nn.roc . max. train <- roc (as.numeric( nn. pred . max. train ) ,train . data [," Class " ]) 7nn.auc . max. train <- auc (nn.roc . max. train ) 8# Testing set 9nn.pred . min .test <- predict(nn.fit .min , test .data , type =’raw ’) 10 nn.roc . min.test <- roc (as .numeric( nn. pred . min. test ) ,train . data[," Class " ]) 11 nn.auc . min.test <- auc (nn . roc . min.test ) 12 nn.pred . max .test <- predict(nn.fit .max , test .data , type =’raw ’) 13 nn.roc . max.test <- roc (as .numeric( nn. pred . max. test ) ,train . data[," Class " ]) 14 nn.auc . max.test <- auc (nn . roc . max.test ) Results obtained are: •Grupo m´ınimo: –AUC training: 1 –AUC testing: 0 •Grupo m´aximo: –AUC training: 1 –AUC testing: 0 34
As you can see, in both cases we obtain an overfitting of the model and the predictors cannot classify correctly instances that not fit on training set. This expose the low capacity of nnet ANNs model to adjust models on this kind of experiments (transcriptomic field) without perform an expensive study to adjust properly the ANN model. Unlike this, DNN models adjust offer excellent results without a exhaustive configuration and handle easily the attribute sets selected. Conclusions and analysis are of these results are in the section Results. 5.4 Resultados The results obtained table for filter and selection process with the implemented tool have been: Tipo Min Max Devuelto AUC.test AUC.train Tiempo BestMax 0 200 200 0,9689 0,7948 120,86 BestMax 0 200 200 1 0,9517 114,72 BestMax 0 200 200 0,9889 0,9092 117,72 BestMax 0 200 200 0,9689 0,9383 119,32 BestMax 0 200 200 0.5556 0.8742 120.11 BestMax 0 200 200 0.9978 0.945 113.5 BestMax 0 200 200 1 0.9571 112.31 BestMax 0 200 200 0.9956 0.9254 115.7 BestMax 0 200 200 0.5556 0.7975 114.78 BestMax 0 200 200 0.5 0.8156 113.52 BestMax 0 200 200 1 0.9402 115.27 BestMax 0 200 200 1 0.9413 112.73 BestMax 0 200 200 0.9956 0.9273 114.27 BestMax 0 200 200 1 0.9706 117.11 BestMax 0 200 200 0.9867 0.8917 115.11 BestMax 0 200 200 0.9822 0.8846 118.47 BestMax 0 200 200 0.9956 0.9438 123.8 BestMax 0 200 200 1 0.9288 119.87 BestMax 0 200 200 0.8333 0.9053 130.35 BestMax 0 200 200 0.9289 0.9386 115.79 BestMax 0 200 200 1 0.9242 116.22 BestMax 0 200 200 0.56 0.9467 116.82 BestMax 0 200 200 0.5 0.8246 116.87 BestMax 0 200 200 1 0.9231 121.25 BestMax 0 200 200 1 0.9096 119.69 BestMax 0 200 200 1 0.8901 112.67 BestMin 0 200 10 0.9444 0.9681 133.42 BestMin 0 200 10 0.8844 0.9804 113.87 35
BestMin 0 200 10 0.9911 0.9638 118.19 BestMin 0 200 10 0.8711 0.9685 126.25 BestMin 0 200 10 0.9911 0.9508 109.11 BestMin 0 200 11 0.96 0.9427 165.3 BestMin 0 200 10 0.9733 0.9563 115.3 BestMin 0 200 10 1 0.971 114.42 BestMin 0 200 10 0.9378 0.9479 127.49 BestMin 0 200 10 0.9956 0.9585 116.78 BestMin 0 200 10 0.9822 0.9685 122.98 BestMin 0 200 10 0.9911 0.9565 116.82 BestMin 0 200 10 0.9422 0.9662 118.17 BestMin 0 200 10 0.9956 0.9681 122.23 BestMin 0 200 10 0.9244 0.9746 114.36 BestMin 0 200 10 0.9867 0.9402 113.1 BestMin 0 200 10 1 0.9765 114.42 BestMin 0 200 10 0.9911 0.9577 123.22 BestMin 0 200 10 0.9867 0.9785 114.28 BestMin 0 200 10 0.9956 0.9585 126.38 BestMin 0 200 10 0.9956 0.9758 128.74 BestMin 0 200 10 0.8333 0.9831 124.58 BestMin 0 200 10 0.6111 0.9467 110.31 BestMin 0 200 10 1 0.9481 110.54 BestMin 0 200 10 0.9911 0.9731 109.81 BestMin 0 200 10 0.9822 0.9629 111.19 BestMin 50 200 50 0.9644 0.9769 111.5 BestMin 50 200 50 1 0.9852 121.75 BestMin 50 200 50 0.9667 0.9652 112.38 BestMin 50 200 50 1 0.9246 115.61 BestMin 50 200 50 1 0.9654 130.42 BestMin 50 200 50 0.9956 0.9759 113 BestMin 50 200 50 1 0.9802 118.1 BestMin 50 200 50 0.96 0.9738 112.08 BestMin 50 200 50 1 0.9654 116.08 BestMin 50 200 50 1 0.9823 124.58 BestMin 50 200 50 1 0.9869 137.26 BestMin 50 200 50 0.9778 0.9727 119.42 BestMin 50 200 50 1 0.9754 118.97 BestMin 50 200 50 0.9333 0.9181 118.08 BestMin 50 200 50 1 0.9669 110.72 BestMin 50 200 50 0.9956 0.9785 110.7 BestMin 50 200 50 0.8689 0.9594 120.99 BestMin 50 200 50 0.8689 0.94 117.12 36
•Implement a support software to identify genetics signatures using DNNs. –Filtering and identification of genetics variables. –Generate DNN models. –GeneraTe preodictions using DNN model. •Research the reliability and effectiveness of DNN models generated to be used with genetic data. •Research effectiveness differences between Artificial Neural Networks (ANNs) and DNNs to find genetic signatures. •Interrelate different concepts learned during grade studies. •Perform a project related with one of the specialties of the grade studied. •Perform a project autonomously. The first one was the main guideline of the functiona requirements. Has been acomplished and implemented in several functions that shape the tool. The second objective has been acomplished with the guided experiment execution. The use of DNNs to analyze transcriptomic data have returned excellent results in acceptable times and with suitable resources consumptions. Their efficiency have been demonstrated, for the experiment performed, with the results obtained from predictions over our testing data set (gold pattern) identifying an attribute set higly related with the interest factor, in this case, the develop of prostate cancer tumors. The third objective don’t have been studiend in depth, understanding study in depth as the study of resources and time consumptions. This study have been evoided consciusly because the difficulties and computationals barriers observed in already implemented ANN tools that have been superseded by DNNs. This affirmation has been generated after see that these named ANN tools couldn’t handled correctly a data set that, in omics sciences field, is small. DNNs didn’t show this difficulties and computational barriers, for thath reasons DNNs are, definitely, a new potetntial tool in medicine and omics studies with huge data sets. The last three objectives are the objectives associated to all final dissertations. The project has been selected an designed taking into account the studies cursed (Grade in Health Engeneering, speciality on Bioinformatics) and the knowledgement obtained during this, related to software, design, data mining, statistics, biology, algorithms and biotechnology. Furthemore, the project have been developed with the higgest autonomy by the author. 6.3 Enhancement opportunities During the project development some concepts and functionalities non-implemented have been realized. These can offer a enhacement opportunities and must be taken in account. These are: •Specificity and sensibility evaluation implementation: specificity and sensibility are concepts basics on medicine and must be taken in account if we want design a tool with applications in prognoses and diagnosis. •More statistical filtering methods: currently there are several efficient statistical filtering and selection methods, but some less commons could be better in 43
specific situations, for that reason the amount of implemented filtering methods must be increased. •H20 exceptions handling: it’s unsual but, sometimes H2O environment can throw connection exceptions that kill the DNN training process. This rarely exceptions must be handled by the tool to avoid killing the process when connection fails. •H20 best configuration searcher: H2O offers a random and grid search methods to find the best configuration for our DNN models. Coordinate this functionality with out tool could help to find better results or the best possible. •More complex formulas generator: currently only addition and deletion formulas are generated by the tool. The improvement of the formula generator to allow more complex formulas could offer better results, but harder to understand. 6.4 Utilities and applicability The main objective of this tool isn’t generate research content, the main objective was and is apply this tool in a real medical envrionment. To favor this apllicability real economical conditions and technologies have been taken in account. Based on this, the following utilites can be performed with this tool: •Genetic signatures identification: it’s the main utility and the reason that impulses this project. This tool can be used to find relationships between genes and relevant medical factors increasing the current knowledgement about these factors. •Developing prognoses and diagnosis tools: the knowledgement generated by this tool can promote the development of new diagnoses and prognosis tools that could accelerate the diseases identification and decrease the costs associated to prevention medicine. •Personalized medicine: genetic signatures identification can improve the personalized medicine making easily the individual studies and helping to make strategic decission using population information. •Pharmacogenomic: pharmacogenomic search accuracy of their drugs to alter specific genes or biological process. Identification of genetic signatures can offer information useful to design new drugs. •Autorregulatory gene networks enhancement: autorregulatory gene networks is a fact and this tool, with this networks algorithms, can offer beter adjustments of these networks and could be applied on medicine, pharmacogenomic and in biological process monitoring. For all these application we must take in account the bioethical implications and I suggest to study derived tools in a Bioethical Council. 6.5 Bioethical implications The knowledgement in depth, althought it’s an statistical value, about ourself, about our future or our offspring future, can generate personal conflicts of religious, moral, 44
existential, social or economical type. For this reason, it’s important to highlight that conclusions derived by this project and their applications can provoke bioethical problems that must be resolved on specifics councils or advice groups once they arise. 6.6 Future lines of work After projec analysis about requirements and objectives compliment and the study of the applicability of the impelemented tool, the project is considered ended having satisfied the bojectives imposed at the begining. Now, as future objectives, we can think in several develoment lines, one of them have been already named on enhacemenet opprotunities section. Other development lines could be: •Use of more and more diverse microarray data sets. •Specificity and sensibilit evaluators implementation. •Export functionalities module implementation. •Include this tool in an R package. •Implements a pipleine that search information in biological data bases using the calculated relationships. •Study, on laboratory, of the results obtained with this tool. 45
7 - An´alisis y conclusiones (ES) Para este proyecto se han realizado todas las fases de proyecto normal de desarrollo de software, entre las que se encuentra el estudio del problema, dise˜no de la soluci´on, dise˜no del software, implementaci´on y testeo del mismo. Adem´as se han realizado tareas referentes a la investigaci´on con la comparaci´on de los resultados con herramientas de la tecnolog´ıa tradicional del ´area (ANN). En estas etapas nombradas tambi´en se han integrado aspectos y etapas referentes a un proyecto de bioinform´atica, vi´endose afectada la parte de elecci´on de requisitos y dise˜no l´ogico por factores de decisi´on biol´ogicos o orientando los resultados de la herramienta a datos biol´ogicos. Todo esto lo vamos a analizar apartado por apartado. 7.1 Revision del cumplimiento de los requisitos Comenzaremos revisando el cumplimiento de los requisitos funcionales y no funcionales impuestos al comienzo del proyecto. Los requisitos funcionales eran: 1.- Importaci´on de datos en diferentes formatos. 2.- Filtrado de variables. 3.- Generaci´on de f´ormulas a partir de set de variables. 4.- Divisi´on del set de datos en subconjuntos aleatorios (estrategia KFold). 5.- Generaci´on de un modelo de Deep Neural Networks a partir de una f´ormula. Para el caso de los requisitos funcionales todos han sido cumplidos. Cada uno ha sido implementado en una o varias funciones. Las funciones que implementan el primer y cuarto requisito han sido codificadas en funciones y almacenadas en el fichero funcs data management.R. El segundo y tercer requisito han sido implementados y almacenados en el fichero funcs variable handling.R. Por ´ultimo, el quinto requisito ha sido implementado en una s´ola funci´on que se encuentra en el fichero funcs dnn.R. Adem´as, durante la ejecuci´on de la herramienta ya implementada se identific´o un requisito propio del lenguaje utilizado, dicho requisito era derivado del impedimento de utilizar caracteres reservados en los nombres de las variables. Debido a que el set de datos utilizado para genes presentaba varios nombres de atributos que inclu´ıan dichos caracteres reservados, se implement´o una funci´on que sustituye dichos caracteres por otros no reservados. ´ Esta fue implementada y almacenada en el fichero funcs variable handling.R que agrupaba las funcionalidades m´as semejantes a la que esta funci´on satisfac´ıa. 46
Vistos los requisitos funcionales pasamos a revisar los requisitos no funcionales. Estos eran: 1.- El problema debe poder cargar sets de datos procedentes de microarrays (chips ADN) para favorecer su implantaci´on (tecnolog´ıa con mayor implantaci´on en la actualidad). 2.- El programa debe implementarse en el lenguaje R para aprovechar su capacidad en ´areas como Big Data, Estad´ıstica y Machine Learning. 3.- El programa debe gestionar las excepciones que se generen para evitar el aborto no controlado del workflow. 4.- El programa debe funcionar en cualquier versi´on de R igual o superior a la 3.3.1. 5.- El programa debe atomizar las funcionalidades e implementarlas como funciones para favorecer el paralelismo en futuras versiones. 6.- El programa debe atomizar las funcionalidades e implementarlas como funciones para favorecer la mantenibilidad del c´odigo sin afectar al usuario. 7.- El programa debe ofrecer funciones con la mayor cantidad de atributos posible para favorecer el absoluto control sobre la configuraci´on del programa. El primero de estos requisitos no funcionales es uno de los m´as importantes. Si se observa el estado actual de las tecnolog´ıas de secuenciaci´on y estudios gen´eticos se puede concluir que los microarrays son una tecnolog´ıa estudiada y expandida. Cierto es que est´a siendo eclipsada por la tecnolog´ıa de RNAseq pero no hay que olvidar que la mayor parte de las instalaciones operativas en el medio en el que queremos incidir (Medicina) trabajan con la tecnolog´ıa de microarrays. Es por esto que si quer´ıamos que esta herramienta tuviera una aplicabilidad real en el entorno m´edico actual deb´ıa poder trabajar con los datos obtenidos de microarrays. Siguiendo este requisito la herramienta carga y modela utilizando datos de matrices num´ericas (datos de microarrays) con variables de caracter cont´ınuo (o acotado) que se relaciona a posteriori con variables de caracter finito (enfermedades, malformaciones u otros factores de inter´es m´edico). Los requitos segundo y cuarto hacen referencia al lenguaje de implementaci´on y versi´on m´ınima del mismo. Ambos se han cumplido al implementar la herramienta usando la versi´on de R 3.3.1. Esta decisi´on se tom´o por el gran potencial que ofrece el motor de R para el manejo de grandes cantidades de informaci´on y por la cantidad de herramientas ya implementadas que ofrece la comunidad de R. El tercero de los requisitos se ha implementado utilizando las funcionalidades del paquete para incluir caracter´ısticas de programaci´on Orientada a Objetos en R que es R.oo. En el caso de nuestra herramienta, se manejan excepciones y se lanzan excepciones propias para los casos en los que la llamada a la herramienta contiene errores o falta de argumentos que impiden el normal funcionamiento de la misma. El quinto y sexto requisito hacen referencia a la estrategia de dise˜no del c´odigo a seguir, esta nos lleva a atomizar las funcionalidades implement´andolas en funciones con el fin de favorecer su mantenibilidad y futura paralelizaci´on si conviene. Esta es la estrategia que se ha seguido y se puede ver reflejada en el dise˜no l´ogico del c´odigo. Por ´ultimo encontramos el s´eptimo requisito que busca favorecer el mayor control del usuario sobre la herramienta. Este requisito se ha cumplido a excepci´on de las funciones que manejan DNNs. Esto ha sido as´ı por la decisi´on consciente de limitar el n´umero de 47
par´ametros de dichas funciones a un set de variables relevantes para el entrenamiento y configuraci´on de una DNN ya que este algoritmo, como la gran mayor´ıa de algoritmos pertenecientes a los conocidos como Deep Learning, tienen una cantidad muy grande de par´ametros configurables, de los cuales no todos provocan cambios significativos en los resultados. 7.2 Revisi´on del cumplimiento de los objetivos Para este proyecto se hab´ıan marcado unos objetivos en base a la motivaci´on del proyecto. Dichos objetivos han guiado el estudio del problema, la selecci´on de requisitos funcionales y no funcionales y, por ende, el dise˜no y la implementaci´on de la herramienta. Estos objetivos eran: •Implementar un software de apoyo a la identificaci´on de marcadores gen´eticos haciendo uso de DNNs. –Identificaci´on y filtrado de variables gen´eticas. –Generaci´on de modelos de DNNs. –Generaci´on de predictores mediante el uso de modelos de DNNs. •Estudiar la eficacia y fiabilidad de los modelos generados con DNNs para su uso con datos gen´eticos. •Estudiar la diferencia de eficacia entre el uso de las Artificial Neural Networks (ANNs) y las Deep Neural Networks para la identificaci´on de marcadores gen´eticos. •Interelacionar los conocimientos adquiridos durante los estudios de grado. •Realizar un proyecto relacionado con una de las menciones de los estudios del grado realizado. •Realizar un proyecto de forma aut´onoma. El primero de los objetivos es el que ha guiado los requisitos funcionales. Ha sido cumplido e implementado en diferentes funciones que conforman la herramienta en s´ı. El segundo de los objetivos se ha cumplido con el experimento guiado realizado. El uso de DNNs para an´alisis de datos de transcript´omica ha dado resultados excelentes en tiempos razonables y con consumos de recursos adecuados. La eficacia queda demostrada, para el caso estudiado, con los resultados de predicci´on obtenidos para el set de testeo (patr´on oro) consiguiendo identificar un set de variables fuertemente relacionadas con el factor de estudio, en este caso, la generaci´on de tumores de cancer de pr´ostata. El tercero de los objetivos no ha sido estudiado en profundidad, entendiendo como estudio en profundidad un estudio del consumo de recursos y de tiempo para obtener un mismo resultado. Este estudio se ha evitado de forma consciente al observar las dificultades y barreras computacionales que padec´ıan las ANNs ya implementadas en R que han sido desbancadas por las DNNs. Esta afirmaci´on se genera al tener en cuenta que dichas herramientas no fueron capaces de manejar con soltura un set de datos que, dentro de las ciencias ´omicas, es peque˜no. Al no presentar estas dificultades y barreras el uso de DNNs las convierten, sin duda, en una nueva herramienta con gran potencial dentro del ´area de la medicina y los estudios de datos ´omicos incluso a gran escala. Los ´ultimos tres objetivos son los referentes al propio objetivo de realizar un Trabajo Final de Grado. El trabajo se ha seleccionado y dise˜nado teniendo en cuenta los estudios 48
cursados (Grade in Health Engeneering, speciality on Bioinformatics) y los conocimientos adquiridos durante ´estos interrelacionando conceptos de software, dise˜no, miner´ıa de datos, estad´ıstica, biolog´ıa, algoritmia y biotecnolog´ıa. Adem´as el proyecto se ha realizado con la mayor autonom´ıa posible por parte del alumno. 7.3 Oportunidades de mejora Durante el desarrollo del proyecto se han observado algunos conceptos no implementados y algunas funcionalidades que ampliar´ıan la calidad y utilidad de la herramientas. Estas cosas representan oportunidades de mejora de la herramieta y son: •Implementar evaluaciones de especificidad y sensibilidad: los conceptos prueba sensible y espec´ıfica son b´asicos en la medicina y deben ser tenido en cuenta si se desea dise˜nar una herramienta aplicable para la diagnosis o prognosis de pacientes. •Implementaci´on de m´as m´etodos de filtraci´on estad´ıstica: en la actualidad existen varios m´etodos de filtrado y estudio de correlaci´on entre variables. Aunque se implementen los m´as comunes y eficaces, es recomendable implementar otros que sean mejores para casos espec´ıficos que se puedan dar seg´un el tipo de set de datos. •Gesti´on de excepciones de H2O: aunque es inusual, el entorno de H2O puede lanzar excepciones de conexi´on que acaban con el proceso de entrenamiento de DNNs. Ser´ıa interesante implementar un proceso de gesti´on de excepciones que detecte cu´al ha sido el error y lo subsane para no perder la iteraci´on. •Implementaci´on de la b´usqueda del mejor de H2O: el paquete H2O ofrece un sistema de b´usqueda aleatoria y por grid del mejor conjutno de par´ametros de configuraci´on para un set de datos y su herramienta de entrenamiento de DNNs. Coordinar esta funcionalidad con nuestra herramienta permitir´ıa reducir el conjunto de atributos obligatorios de las funciones de la herramienta y permitir´ıa obtener el mejor resultado posible. •Generador de f´ormulas m´as complejas: actualmente solo formulas con adici´on y delecci´on son generadas por la herramienta. La ampliaci´on de las posibilidades en la generaci´on de f´ormulas podr´ıa ofrecer mejores resultados aunque m´as dificiles de entender. 7.4 Utilidades y aplicabilidad Lo importante de esta herramienta no es s´olo crear nuevo contenido de investigaci´on sino tener aplicabilidad real en el entorno m´edico. Para favorecer esta aplicabilidad se han tenido en cuenta las condiciones reales y actuales de las tecnolog´ıas que se encuentras implementadas en los centros sanitarios. En base a esto se considera que las utilidades pr´acticas de esta herramienta ser´ıan: •Identificaci´on de marcadores gen´eticos de factores de inter´es: este es la utilidad principal para la que se ha dise˜nado la herramienta. ´ Esta puede servir para identificar la relaci´on entre genes y factores de inter´es m´edico ampliando el conocimiento actual de dichos factores. 49
•Creaci´on de pruebas de diagn´ostico y prognosis: la generaci´on de herramientas de diagn´ostico y progn´osis con el conocimiento derivado de la identificaci´on de marcadores gen´eticos podr´ıa suponer una aceleraci´on en la identificaci´on de enfermedades adem´as de un abaratamiento de los costes asociados a la medicina de prevenci´on. •Medicina personalizada: la identificaci´on de marcadores gen´eticos servir´ıa para mejorar la orientaci´on de la medicina personalizada sirviendo tanto para mejorar el estudio de individuos concretos como para tomar decisiones estrat´egicas seg´un la gen´etica de la poblaci´on seg´un zonas geogr´aficas. •Farmacogen´omica: la farmacogen´omica busca la precisi´on en sus f´armacos para afectar a genes concretos o para alterar los procesos que estos controlan. La identificaci´on de marcadores gen´eticos ofrecer´ıa informaci´on de gran utilidad para el dise˜no de estos f´armacos. •Ampliaci´on del conocimiento de redes de autoregulaci´on g´enica: la autoregulaci´on g´enica es un hecho. Esta herramienta, en conjunto con algoritmos de c´alculo de redes de autoregulaci´on g´enica, podr´ıan ofrecer mejores ajustes de las mismas, las cuales son aplicadas en la medicina, la farmacogen´omica y el control de procesos biol´ogicos en general. Para todas estas aplicaciones habr´ıa que tener en cuenta sus implicaciones bio´eticas y se recomiendan que todas las herramientas derivadas de ´esta sean estudiadas por un Consejo de Bio´etica. 7.5 Implicaciones bio´eticas El conocimiento en profundidad, aunque sea de forma probabil´ıstica, de nuestro propio ser, de nuestro futuro, del de nuestra descendencia, puede producir un conflicto personal de tipo religioso, moral, existencial, o a nivel econ´omico o social. Por lo tanto, es de destacar que las conclusiones que se deriven de este trabajo y sus aplicaciones pr´acticas pueden conllevar problemas bio´eticos que deber´ıan resolverse en los respectivos comit´es dedicados al tema, una vez que fueran surgiendo. 7.6 L´ıneas futuras de trabajo Despu´es de la realizaci´on del proyecto y del an´alisis del cumplimiento de los requisitos y la aplicabilidad de la herramienta se considera por satisfechos los objetivos marcados para este proyecto. Como objetivos futuros se pueden marcar varios, entre los que se incluyen las nombradas en el apartado de oportunidades de mejora, la siguientes se consideran como interesantes y ´utiles para la mejora y desarrollo de esta herramienta: •Estudio de m´as y m´as variados sets de datos. •Implementaci´on de evaluadores de sensibilidad y especificidad. •Implementar m´odulo de exportaci´on de datos en diferentes formatos. •Crear paquete en R con la herramienta. 50
•Implementaci´on de pipeline que busque informaci´on en bases de datos utilizando las relaciones calculadas por la herramienta. •Estudio, en laboratorio h´umedo, de los resultados obtenidos por esta herramienta. 51
8 - Concepts 8.1 Omics sciences In biomedicine. omics are all disciplines, technologies and research areas that study a set or the whole biological system. Term omics is a new-flanged sufix that is added to another concept to define a biological system, understandin as biological system a whole or a functional part of an organism. The most popular example and used is genomics that includes all knowledge areas and technologies that research about an individue or specie genome. Based on this, genomics includes all researches related with sequentation and genomes annotation, related with functional genome regulation, with the mutatuions and genetic modifications studies, etc. As you can appreciate, omics are important disciplines by their own, but also as source of new information tools that permit increas the knowledgement about more specific areas. Other omics are transcriptomic (genes transcripts study in an organism), proteomic (proteins presence study in a biological system) and metabolomic (molecular study of metabolics systems). To have a more complete set we must include conectomic, epigenomic, filogenomic and metagenomic. To know more you can see this web1. 8.2 Deep Learning and Deep Neural Networks For a long time, several algorithms have been developed in the machine learning field. Wit this algorithms we want teach (adjust parameters and models of algorithm) to a computer to take advantage of his computational capacity to detect patterns and make predictions using the identified patterns. In recent years a new family of machine learning algortihms have been developed. These are characterized by their capacity to learn about the data representation. This family is known as Deep Learning (DL) algorithms. These algorithms are a result of increasing the complexity of themselves using non-linear multiple composed transformations architectures. Furthemore, the current boom of these technologies it’s given by the relevant enhacement of results regarding its homologues of non-deep (traditional) machine learning algorithms. An specific case of these DL algorithms is Deep Neural Networks (DNN) based on tra1https://en.wikipedia.org/wiki/Omics 52