Full text
Boosting ICD multi-label classification of health records with contextual embeddings and label-granularity Alberto Blancoa,∗, Olatz P´erez de Vi˜nasprea, Alicia P´ereza, Arantza Casillasa aIXA Taldea. UPV-EHU. Manuel Lardizabal Ibilbidea, 1, Donostia 20018 Spain Abstract Background and Objective This work deals with clinical text mining, a field of Natural Language Processing applied to biomedical informatics. The aim is to classify Electronic Health Records with respect to the International Classification of Diseases, which is the foundation for the identification of international health statistics, and the standard for reporting diseases and health conditions. Within the framework of data mining, the goal is the multi-label classification, as each health record has assigned multiple International Classification of Diseases codes. We investigate five Deep Learning architectures with a dataset obtained from the Basque Country Health System, and six different perspectives derived from shifts in the input and the output. Methods We evaluate a Feed Forward Neural Network as the baseline and several Recurrent models based on the Bidirectional GRU architecture, putting our research focus on the text representation layer and testing three variants, from standard word embeddings to meta word embeddings techniques and contextual embeddings. Results The results showed that the recurrent models overcome the non-recurrent model. The meta word embeddings techniques are capable of beating the ∗Corresponding author: Alberto Blanco. IXA Taldea. UPV-EHU. Manuel Lardizabal Ibilbidea, 1, Donostia 20018 Spain Email address: [email protected] (Alberto Blanco) This is the accepted manuscript of the article that appeared in final form in Computer Methods and Programs in Biomedicine 188 : (2020) // Article ID 105264, which has been published in final form at https://doi.org/10.1016/ j.cmpb.2019.105264 © 2019 Elsevier under CC BY-NC-ND license (http://creativecommons.org/licenses/by-ncnd/4.0/)
standard word embeddings, but the contextual embeddings exhibit as the most robust for the downstream task overall. Additionally, the label-granularity alone has an impact on the classification performance. Conclusions The contributions of this work are a) a comparison among five classification approaches based on Deep Learning on a Spanish dataset to cope with the multi-label health text classification problem; b) the study of the impact of document length and label-set size and granularity in the multi-label context; and c) the study of measures to mitigate multi-label text classification problems related to label-set size and sparseness. Keywords: Electronic Health Record, International Classification of Diseases, Multi-label classification, Recurrent Neural Networks, Contextual embeddings, Label-granularity 1. Introduction1 Methodical documentation of healthcare data is fundamental for pub-2 lic health. The International Classification of Diseases (ICD) is the3 standard diagnoses coding system for Electronic Health Records (EHR)4 classification. ICD serves, worldwide, for epidemiology, health management5 and documentation purposes. Over time, several versions have been devel-6 oped, being the ICD-10th the current version. Regarding the hospital net-7 work associated with the Spanish “Ministerio de Sanidad, Servicios Sociales8 e Igualdad”, from January the 1st 2016, the clinical modification of the ICD-9 10th is the reference version, adopting the Spanish translated CIE-10-ES10 variant as the coding standard. The ICD-10 is designed as an alphanumeric11 code and it is arranged hierarchically [1]. Each code is built by a set from 312 to 7 alphanumeric characters as shown in figure 1.13 Figure 1: ICD-10 code structure. 2
In this paper we tackle the task of automatically coding the diagnostic14 terms present in a free-text medical record according to the ICD coding15 system. The task is framed within the Natural Language Processing (NLP)16 field. The purpose is to determine which classes are present in the input17 text. Our approach rests on machine learning, specifically on supervised18 multi-label classification.19 Classification based solely on text is an open challenge in artificial in-20 telligence [2,3,4]. We aim to solve a text classification problem on medical21 free-text, EHRs that present medical jargon and clinical-specific language.22 Furthermore, EHRs often contain abbreviations (frequently non-standard),23 and misspellings are also common. The length of the texts plays an impor-24 tant role, here we face a broad spectrum, ranging from a few words to several25 tens of lines. EHRs seldom express clinical diagnoses as in the standard ICD.26 The length of the texts plays an important role, here we face a broad27 spectrum, ranging from a few words to several tens of lines. EHRs seldom28 express clinical diagnoses as in the standard ICD.29 An EHR could entail many diagnostics henceforth, multiple ICD labels30 should be assigned. This task, Multi-label Classification, can be seen as a31 multi-class classification (not binary) task in which the classes are not mutu-32 ally exclusive. Multi-label classification tends to be, by far, more challenging33 than mere multi-class classification. Its complexity lies in the exponential34 growth of label combinations. Note, as well, that the number of labels asso-35 ciated to each EHR is variable.36 Multi-label classification can be tackled with the so-called binary rele-37 vance approach. This simplistic approach consists of using as many binary38 classifiers as ICD codes to determine if each ICD code is present or absent39 from the EHR. The drawback of this approach rests on the fact that the40 model is not able to capture label-dependencies. While some diagnostics41 are prone to co-appear others are incompatible. Learning label-dependencies42 is crucial to this task. To this end, we explore approaches based on Deep43 Learning [5,6,7].44 The contribution of this work is to explore the impact of dataset char-45 acteristics, such as the characterization of the input text focused on (either46 full document or a part of it), on the predictive ability of the multi-label47 neural models and also to assess the performance with respect to label-set48 cardinality and granularity. We deal with real EHRs from Osakidetza (the49 3
Basque Public Health System) written in Spanish1.50 2. Related work51 Text classification of EHRs is a demanding task, hence most works have52 focused on short English texts, though on this work we deal with novel chal-53 lenges including long EHRs written in Spanish with thousands of words.54 Multi-label classification is a challenging task, especially when the num-55 ber of labels is high [8,9,10]. The binary relevance approach transforms56 the multi-label problem in multiple binary classification problems [11], but57 disregard the dependencies among labels. Several works have addressed the58 EHRs classification according to the ICD [12,13,14,15]. Yet, little at-59 tention was paid to dense features and to the approaches that could take60 advantage of them. Furthermore, much uncertainty still exists about the61 inter-dependency of labels, that could enhance the prediction performance62 avoiding incongruities such as, for example, assigning an adult-specific dis-63 ease simultaneously with a childhood condition. On this work, we tackle64 the model and capture of label dependencies through Deep Learning models,65 leveraging the dense output layer with Sigmoid activation function.66 The text classification field has leapt forward, from linear and proba-67 bilistic models over hand-crafted engineered features [16,17] to non-linear68 Neural Network models and end-to-end learnt inherent high-level text rep-69 resentations. It is shown good performance with NN [18], as Convolutional70 Neural Networks [19], Recurrent Neural Networks [20] and Bidirectional Long71 Short-Term Memory [21].72 Methods of meta-embeddings aim to conduct a complementary combi-73 nation of information from an ensemble of distinct word embeddings to yield74 an embedding set with enhanced quality and characteristics of the semantics75 captured. Yin and Sch¨utze [22] presented, among others, the “concatenation”76 method, wherein the meta-embedding is the concatenation of several embed-77 dings. Coates and Bollegala [23] assured that direct averaging of embedding78 can provide an approximation of the efficiency of concatenation without in-79 creasing the dimension of the embeddings.80 Context representations are vital to NLP tasks such as text classifica-81 tion. To alleviate this weakness present in generic word embeddings the con-82 textual embeddings emerged. Melamud et al. presented an unsupervised83 1The dataset contains sensitive, confidential data, and therefore can not be released. 4
model for learning context embedding of wide contexts of sentences using84 bidirectional LSTMs. These embeddings are dependent on the entire corpus85 from which they were inferred and carry reinforced contextual meaning. The86 ELMo [25] and BERT [26] have become state-of-the-art in contextual word87 representations. Much uncertainty still exists about the advantages of apply-88 ing meta and contextual embeddings over the standard options for clinical89 text classification tasks, and we have found that the contextual embeddings90 may give an extra edge on the ICD classification.91 In the automatic ICD coding, there are also works that point towards92 the Neural Network trend but seems to fall short on the field. These models93 manage to handle large amounts of text through a dense representation of94 words. Nigam [27] took advantage of Recurrent Neural Networks to perform95 multi-label classification. Both works were carried out with discharge sum-96 maries from the MIMIC-III [28] corpus. Recently, this task has gained more97 attention through the CLEF eHealth evaluation labs. Suominen et al. [29]98 presented an overview of the sixth annual edition. The goal of one of the99 tasks is to automatically assign ICD-10 codes to few words length texts from100 free-text descriptions of causes of death as reported by physicians [30,31].101 The task is similar to what we have presented on this work with the Di-102 agnostic input perspective, and the finding is that the performance of the103 classifiers could be improved employing the full documents.104 Spanish NLP is under strong growth, among others, driven by the Plan105 de Tecnolog´ıas del Lenguage2. EHRs in Spanish are currently being collected106 [32,33], as well as complimentary corpora including abstracts [34]. These107 data sets enable to develop several tasks e.g., Negation Extraction [35], Ex-108 traction of Adverse Drug Reactions [36], Text Classification [30,37,31], and109 Negation Cue Detection [38].110 3. Methods111 We explored four unique RNN model instances plus the baseline model, a112 Feed Forward Neural Network with Neural-Net Language Model (NNLM) as113 the text representation layer. The core architecture is a Bidirectional Recur-114 rent Neural Network with GRU units and pooling techniques [39] (explained115 in section 3.1). The cornerstone of the model is the word embedding layer,116 2https://www.plantl.gob.es/tecnologias-lenguaje/actividades/ infraestructuras/Paginas/infraestructuras-linguisticas.aspx 5
as it is responsible for the expressiveness of the input. Thus, we explored117 three variants: standard embeddings, meta-embeddings and contextual em-118 beddings (explained in depth in section 3.2). Together with this work, in an119 attempt to promote reproducibility, we released the software package that120 we implemented3.121 3.1. Bidirectional Recurrent Neural Network with GRU units and pooling122 We applied a Bidirectional layer with GRU units, which leverages se-123 quences of text in forward and reverse order with separate hidden states,124 and whose mathematical formulation for the forward and backward hidden125 state and its combination is shown in (1).126 −→ h(t)=σ(−→ W x(t)+−→ V−→ h(t−1) +−→ b) ←− h(t)=σ(←− W x(t)+←− V←− h(t−1) +←− b) h(t)= [−→ h(t),←− h(t)] (1) The parameters are the weight matrices [−→ W , ←− W] and [−→ V , ←− V], and the127 bias terms [−→ b , ←− b]. The hidden-states are computed through the non-linear128 activation (σ) applied to the weighted sum between previous hidden-states129 [−→ h(t−1),←− h(t−1)] and current input (x(t)) with their corresponding matrices.130 Then, both hidden states are combined with concatenation to provide the131 resulting hidden state (ht).132 The output of the Bidirectional RNN layer could be fed to the dense133 layer. However, this can be computationally challenging, due to the high134 number of parameters. Learning a classifier with too many parameters can135 be unwieldy, and can also be prone to over-fitting. A popular technique to136 deal with the high dimensionality of the Bidirectional RNN layer output is137 Pooling [40]. We applied average and max-pooling, known as 1-dimensional138 global pooling. The pooled features are concatenated and fed into a final139 fully-connected layer. This layer is responsible for computing the probability140 estimation of the labels i.e. ICD codes. Figure 2shows the full architecture141 3The software is available at http://ixa2.si.ehu.es/prosamed/cmpICD_soft and can be downloaded with user CMPB and password IXAcmpb. Provided that the software is used anyhow, this article should be cited. 6
of the Bidirectional Recurrent Neural Network with GRU units and pooling142 techniques, i.e. BiGru. The figure shows a forward pass for an example text.143 The output of the Sigmoid function is the probability estimation of each144 label. The depth of every layer indicates the batch size. The Recurrent layer145 is unrolled, so si∀i∈sbrings the embedded representation of the input146 token {s1=emb(“patient”),s2=emb(“had”),s3=emb(“achalasia”)}.147 GRUGRU GRUGRU GRUGRU GRUGRU ["patient", "had", "achalasia"]["patient", "had", "achalasia"] the suers gallstone gout achalasia from hospital the suers gallstone gout achalasia from hospital [-0.18, 1.86, ..., -0.22] [-0.02, 1.88, ..., 0.63] [ 1.34, 1.14, ..., 0.05] [-0.39, 2.09, ..., -0.25] [ 0.49, 1.40, ..., -0.07] [-1.35, 1.56, ..., 2.82] [-0.87, -0.45, ..., -1.67] [-0.18, 1.86, ..., -0.22] [-0.02, 1.88, ..., 0.63] [ 1.34, 1.14, ..., 0.05] [-0.39, 2.09, ..., -0.25] [ 0.49, 1.40, ..., -0.07] [-1.35, 1.56, ..., 2.82] [-0.87, -0.45, ..., -1.67] [-0.18, ..., -0.22] [ 1.34, ..., 0.05] [-1.35, ..., 2.82] [-0.18, ..., -0.22] [ 1.34, ..., 0.05] [-1.35, ..., 2.82] [ 0.75, ..., 1.25] [ 0.75, ..., 1.25] ← steps → ← steps → [-1.63, ..., 2.11] [-1.63, ..., 2.11] [ 0.75, ..., 1.25, ..., -1.63, ..., 2.11] [ 0.75, ..., 1.25, ..., -1.63, ..., 2.11] [-0.36, ..., -0.23] [ 1.64, ..., 0.32] [-1.63, ..., 2.11] [-0.36, ..., -0.23] [ 1.64, ..., 0.32] [-1.63, ..., 2.11] [0.12, 0.54, 0.85] [0.12, 0.54, 0.85] GRUGRU GRUGRU INPUT INPUT INPUT EMBEDDING LAYEREMBEDDING LAYEREMBEDDING LAYER ← embed_size → ← embed_size → ↑ v o c a b ↓ ↑ v o c a b ↓ ← embed_size → ← embed_size → BIDIRECTIONAL GRU-RNNBIDIRECTIONAL GRU-RNNBIDIRECTIONAL GRU-RNN ← hidden_size →← hidden_size →← hidden_size → ← hidden_size → CONCAT CONCAT CONCAT ← 2 * hidden_size → ← 2 * hidden_size → FULLYFULLY CONNECTEDCONNECTED FULLY CONNECTED SIGMOIDSIGMOIDSIGMOID ← num_classes → ← num_classes → ↑ s ↓ ↑ s ↓ ← hidden_size → ← hidden_size → s₁s₁s₁s₂s₂s₂s₃s₃s₃ MAX POOLINGMAX POOLINGMAX POOLING AVERAGE AVERAGE POOLINGPOOLING AVERAGE POOLING ↑ s ↓ ↑ s ↓ EMBEDDED INPUTEMBEDDED INPUTEMBEDDED INPUT Figure 2: Architecture: Bidirectional RNN with GRU units and pooling model. The BiGru model can handle all the labels at once, instead of following148 a binary relevance approach, training independent classifiers for each label.149 The final dense layer is able to capture and model the label dependencies,150 producing a non-mutually exclusive probability estimation for each label with151 the Sigmoid activation function [41].152 7
3.2. Comprehensive input characterization: embedding layer variations153 A comprehensive input characterization is crucial for attaining competi-154 tive performance. In the training stage, the embedding layer holds more than155 90% of the model’s complexity in terms of parameter count. What is more,156 the predictive capacity rests on the ability of the model to extract knowledge157 from the source provided in the input stage. Thus, we paid special attention158 to this layer. The embedding layer from the figure 2shows just a vanilla159 embedding layer that we enhanced later. Indeed, in this work we explored160 three variations of the embedding layer: i) Standard embeddings. ii) Meta161 embeddings (sections 3.2.1-3.2.2). iii) Contextual embeddings (section 3.2.3)162 Moreover, according to Yin et al. [42] and Coates and Bollegala [23],163 different pre-trained word embeddings have substantial differences in quality164 and characteristics of the word representations. The consequence is some165 word embeddings performing better on some tasks than in others. Bearing166 all this in mind, in addition to a standard pre-trained embedding, we tried167 meta-embeddings, which are ensemble approaches (embedding concatenation168 and blending) with the hope to get an embedding set with the improved169 overall quality.170 We turned to embeddings derived from fastText [43] as the standard171 embeddings setup. As for meta-embeddings setup, we employed fastText,172 Word2Vec [44] and GloVe [45]. Every embedding set is trained on the same173 corpus, the Spanish Billion Word Corpus [46].174 3.2.1. Embedding Concatenation175 The meta-embedding is computed as the concatenation of word embed-176 dings, based on the work by [22]. Before the concatenation, each embedding177 set must be L2-normalized [6], so that all the values are in the range [−1,1]178 and, therefore, every set contributes equally.179 The dimensionality of the resulting meta-embeddings is ˆ dsk=ds1+· · · +180 dsi+dsnwith dsibeing the dimension of the i-th set concatenated. It is181 important to note that the model’s complexity increases with each added182 embedding set, as it increases the dimension of the features of the embedding183 layer.184 3.2.2. Embeddings blending185 The meta-embedding variant is computed as the average of the embed-186 dings involved, based on the work by Coates and Bollegala [23]. Note that187 even having embedding sets with matching number of dimensions (dsi=188 8
dsj∀i, j), each dimension among embeddings is not related. In any case, av-189 eraging can provide an approximation of the performance of concatenation190 without the expense of increasing the dimension [23].191 3.2.3. Contextual embeddings192 Recently, approaches that improve the semantic word representation by193 leveraging the context to encode syntactical meaning and handle polysemy194 are pushing the state-of-the-art. Regular word embedding techniques use all195 the occurrences of a word to extract a joint representation. However, de-196 pending on the context, words could have different meanings. Recent models197 exploit this reasoning and propose contextual word embeddings. There is no198 longer a lookup table between words and dense representations. Instead, the199 word embedding is computed on the fly, taking advantage of the context.200 Embeddings from Language Models (ELMo) [25] representations are ob-201 tained from a bidirectional Language Model (biLM) that has recently pro-202 duced state-of-the-art results in several NLP tasks like Coreference Resolu-203 tion [47] or Natural Language Inference and Sentiment Analysis [25]. The204 embedding for a given word varies from one sentence or document to another205 with its context. As it cannot be pre-computed, the embedding computation206 is done computing a forward propagation of the model for each token of each207 input sequence [48].208 4. Experimental framework209 4.1. Data210 The datasets used in our experiments consist of EHRs written in Spanish211 from the Basque public health system (Osakidetza). Specifically, emergency212 services discharge summaries from hospitals. The EHRs are not structured213 and were not written using templates with sections. Table 1introduces the214 details of the dataset used. There are 10,707 EHRs. As revealed by the table,215 we considered several perspectives of the dataset by varying two factors, the216 input and the output explained in what follows.217 9
Focusing on the document input, we can observe that the behaviour for297 every model is also similar, improving results as the granularity decreases.298 One key finding is that the granularity has an impact alone. With less299 granularity, the performance increases, even with more number of labels.300 This finding is depicted by the situation between the full labels (n= 16)301 and the block labels (n= 19), where with the block labels the performance302 improves despite having 3 more labels. This suggests that is possible to get303 models performing better with the same number of labels by just decreasing304 the label granularity.305 4.2.3. Discussion306 With this work we gained the following insights: Despite is a difficult307 task, Deep Learning recurrent models exhibit strong predictive capabilities308 and can be enhanced by more robust text representation techniques such as309 the meta or contextual embeddings.310 We argue that our experimental results throw one key finding: the gran-311 ularity of the labels alone has an impact on performance. The significance lies312 in the possibility of performance improvement by reducing the granularity313 without reducing the label-set size.314 BiGru powered by ELMo is the dominant model in practically every situ-315 ation from both the input and output perspectives (shown in table 2and fig-316 ure 5). Accordingly, a per-class evaluation on the best-performing dataset317 perspective is shown in figure 6.318 16
0 2000 4000 6000 numberofsamples Classfrequency I E Z J N F K R C G D M T B 0 20 40 60 80 100 f1score 88.3% 73.5% 65.1% 85.5% 80.5% 77.0% 75.7% 51.4% 91.0% 60.5% 67.2% 54.5% 41.7% 54.3% Figure 6: Per-class evaluation of BiGru ELMo based on F-score and class frequency for the {Inp: Document, Out: Chapter}subtask. Draw attention to the fact that the worse performing label T(“Injury,319 poisoning and certain other consequences of external causes”) gets 41.7%320 F-score, while the best-performing label C(“Neoplasms”) reaches an out-321 standing 91%. Half of the labels are above 70% and the ≈30% of labels are322 above 80%.323 To assess the stability of the models and the statistical significance of the324 results, we performed five runs repeating the experimental set with random325 seeds and found that Stdev. among runs remained under 0.5 for precision326 and recall and 0.25 for F-score for every model and setup, which means that327 the given experimental results are both reproducible and representative.328 5. Conclusions329 We presented a set of Deep Learning methods to tackle the NLP challenge330 of multi-label text classification with medical free-text: EHRs written in331 Spanish with datasets from the Basque Country Health System and classified332 according to the ICD. Each EHR is assigned multiple ICD codes, leading to333 multi-label classification of text.334 In this work we turned to deep neural models and we found that contex-335 tual information conveyed by the BiGru ELMo achieved competitive results.336 17
BiGru, by contrast to main approaches seen in the literature, has a mecha-337 nism to cope with label co-apparitions and regard diseases as related.338 We wondered if the neural models were able to extract the information339 from entire EHRs of nearly a thousand words or could be boosted by select-340 ing a small though representative section (diagnoses). Experimental results341 showed that it is worthy providing the model with the full document as342 it might convey meaningful information. Particularly, BiGru powered with343 contextual embeddings form Elmo (BiGru+Elmo) outperformed the rest of344 the models explored. In fact, BiGru+Elmo outperformed every model in345 all the setups. The difficulty of correctly predicting a label is not the same346 across labels. A per-class evaluation revealed the competitive performance347 of this approach on minority classes. That is, BiGru+Elmo resulted robust348 regarding the class imbalance and, obviously, leveraged frequent ICDs.349 Finally, we explored the performance attained varying the output la-350 bel granularity (fully-specified code, block, chapter) and label-set cardinality351 (from 14 to 19). This is of interest to decide whether to create a fully auto-352 matic ICD classification engine or, depending on the performance required,353 make the decision to let the model just predict a higher order in the hierarchy.354 There are several open directions for future work. First, our models355 leverage ELMo based contextual embeddings, but there are other novel ap-356 proaches to contextual embeddings based on Language Models, like BERT357 [26]. Second, the core architecture of this work is the Recurrent Neural Net-358 work, but there are other intriguing architectures like Convolutional Neural359 Networks, especially Capsule Network [52] or the architecture behind BERT,360 the new RNN alternative promising approach called Transformer [53]. Third,361 the methods to address the relation among labels, such as statistical driven362 approaches (e.g., correlation analysis [54]) and strategies leveraging the hi-363 erarchically structured ICD and related ontologies (e.g., Hierarchical Multi-364 label Classification [55] and SNOMED-CT [56]).365 References366 [1] W. H. Organization, International statistical classification of diseases367 and related health problems, volume 1, World Health Organization,368 2004.369 [2] H. T. Madabushi, M. Lee, High accuracy rule-based question classifica-370 tion using question syntax and semantics, in: Proceedings of COLING371 18
2016, the 26th International Conference on Computational Linguistics:372 Technical Papers, pp. 1220–1230.373 [3] D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. St. John, N. Con-374 stant, M. Guajardo-Cespedes, S. Yuan, C. Tar, B. Strope, R. Kurzweil,375 Universal sentence encoder for English, in: Proceedings of the 2018 Con-376 ference on Empirical Methods in Natural Language Processing: System377 Demonstrations, Association for Computational Linguistics, Brussels,378 Belgium, 2018, pp. 169–174.379 [4] J. Howard, S. Ruder, Universal language model fine-tuning for text clas-380 sification, in: Proceedings of the 56th Annual Meeting of the Association381 for Computational Linguistics (Volume 1: Long Papers), volume 1, pp.382 328–339.383 [5] I. Goodfellow, Y. Bengio, A. Courville, Deep Learning, MIT Press, 2016.384 http://www.deeplearningbook.org.385 [6] Y. Goldberg, A primer on neural network models for natural language386 processing, Journal of Artificial Intelligence Research 57 (2016) 345–420.387 [7] Y. LeCun, Y. Bengio, G. Hinton, Deep learning, Nature 521 (2015)388 436–444.389 [8] K. Bhatia, H. Jain, P. Kar, M. Varma, P. Jain, Sparse local embeddings390 for extreme multi-label classification, in: Advances in neural information391 processing systems, pp. 730–738.392 [9] H. Jain, Y. Prabhu, M. Varma, Extreme multi-label loss functions for393 recommendation, tagging, ranking & other missing label applications,394 in: Proceedings of the 22nd ACM SIGKDD International Conference on395 Knowledge Discovery and Data Mining, ACM, pp. 935–944.396 [10] K. Jasinska, K. Dembczynski, R. Busa-Fekete, K. Pfannschmidt,397 T. Klerx, E. Hullermeier, Extreme F-measure maximization using sparse398 probability estimates, in: International Conference on Machine Learn-399 ing, pp. 1435–1444.400 [11] M.-L. Zhang, Y.-K. Li, X.-Y. Liu, X. Geng, Binary relevance for multi-401 label learning: an overview, Frontiers of Computer Science 12 (2018)402 191–202.403 19
[12] P. Franz, A. Zaiss, S. Schulz, U. Hahn, R. Klar, Automated coding404 of diagnoses–three methods compared., in: Proceedings of the AMIA405 Symposium, American Medical Informatics Association, p. 250.406 [13] A. Perotte, R. Pivovarov, K. Natarajan, N. Weiskopf, F. Wood, N. El-407 hadad, Diagnosis code assignment: models and evaluation metrics, Jour-408 nal of the American Medical Informatics Association 21 (2013) 231–237.409 [14] M. Saeed, M. Villarroel, A. T. Reisner, G. Clifford, L.-W. Lehman,410 G. Moody, T. Heldt, T. H. Kyaw, B. Moody, R. G. Mark, Multiparam-411 eter Intelligent Monitoring in Intensive Care II (MIMIC-II): a public-412 access intensive care unit database, Critical care medicine 39 (2011)413 952.414 [15] J. P´erez, A. P´erez, A. Casillas, K. Gojenola, Cardiology record multi-415 label classification using Latent Dirichlet Allocation, Computer methods416 and programs in biomedicine 164 (2018) 111–119.417 [16] T. Joachims, Text categorization with support vector machines: Learn-418 ing with many relevant features, in: European conference on machine419 learning, Springer, pp. 137–142.420 [17] A. McCallum, K. Nigam, et al., A comparison of event models for naive421 bayes text classification, in: AAAI-98 workshop on learning for text422 categorization, volume 752, Citeseer, pp. 41–48.423 [18] J. Nam, J. Kim, E. L. Menc´ıa, I. Gurevych, J. F¨urnkranz, Large-424 scale multi-label text classification revisiting neural networks, in: Joint425 european conference on machine learning and knowledge discovery in426 databases, Springer, pp. 437–452.427 [19] Y. Kim, Convolutional neural networks for sentence classification, in:428 Proceedings of the 2014 Conference on Empirical Methods in Natural429 Language Processing (EMNLP), Association for Computational Linguis-430 tics, Doha, Qatar, 2014, pp. 1746–1751.431 [20] D. Tang, B. Qin, X. Feng, T. Liu, Target-dependent sentiment classifi-432 cation with long short term memory, CoRR, abs/1512.01100 (2015).433 [21] P. Zhou, Z. Qi, S. Zheng, J. Xu, H. Bao, B. Xu, Text classification434 improved by integrating bidirectional LSTM with two-dimensional max435 20
pooling, in: Proceedings of COLING 2016, the 26th International Con-436 ference on Computational Linguistics: Technical Papers, The COLING437 2016 Organizing Committee, Osaka, Japan, 2016, pp. 3485–3495.438 [22] W. Yin, H. Sch¨utze, Learning word meta-embeddings, in: Proceed-439 ings of the 54th Annual Meeting of the Association for Computational440 Linguistics (Volume 1: Long Papers), Association for Computational441 Linguistics, Berlin, Germany, 2016, pp. 1351–1360.442 [23] J. Coates, D. Bollegala, Frustratingly easy meta-embedding – computing443 meta-embeddings by averaging source word embeddings, in: Proceed-444 ings of the 2018 Conference of the North American Chapter of the Asso-445 ciation for Computational Linguistics: Human Language Technologies,446 Volume 2 (Short Papers), Association for Computational Linguistics,447 New Orleans, Louisiana, 2018, pp. 194–198.448 [24] O. Melamud, J. Goldberger, I. Dagan, Context2Vec: Learning generic449 context embedding with bidirectional LSTM, in: Proceedings of The450 20th SIGNLL Conference on Computational Natural Language Learn-451 ing, pp. 51–61.452 [25] M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee,453 L. Zettlemoyer, Deep contextualized word representations, in: Pro-454 ceedings of the 2018 Conference of the North American Chapter of the455 Association for Computational Linguistics: Human Language Technolo-456 gies, Volume 1 (Long Papers), Association for Computational Linguis-457 tics, New Orleans, Louisiana, 2018, pp. 2227–2237.458 [26] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training459 of deep bidirectional transformers for language understanding, CoRR460 abs/1810.04805 (2018).461 [27] P. Nigam, Applying deep learning to ICD-9 multi-label classification462 from medical records, 2016.463 [28] A. E. Johnson, T. J. Pollard, L. Shen, H. L. Li-wei, M. Feng, M. Ghas-464 semi, B. Moody, P. Szolovits, L. A. Celi, R. G. Mark, MIMIC-III, a465 freely accessible critical care database, Scientific data 3 (2016) 160035.466 [29] H. Suominen, L. Kelly, L. Goeuriot, A. N´ev´eol, L. Ramadier, A. Robert,467 E. Kanoulas, R. Spijker, L. Azzopardi, D. Li, et al., Overview of the468 21
CLEF eHealth Evaluation Lab 2018, in: International Conference of the469 Cross-Language Evaluation Forum for European Languages, Springer,470 pp. 286–301.471 [30] A. Atutxa, A. Casillas, N. Ezeiza, V. Fresno, I. Goenaga, K. Gojenola,472 R. Mart´ınez, M. O. Anchordoqui, O. Perez-de Vi˜naspre, IxaMed at473 CLEF eHealth 2018 Task 1: ICD10 coding with a Sequence-to-Sequence474 Approach., in: CLEF (Working Notes), p. 1.475 [31] M. Almagro, S. Montalvo, A. D. de Ilarraza, A. P´erez, MAMTRA-MED476 at CLEF eHealth 2018: A Combination of Information Retrieval Tech-477 niques and Neural Networks for ICD-10 Coding of Death Certificates.,478 in: CLEF (Working Notes), p. 1.479 [32] M. Oronoz, K. Gojenola, A. P´erez, A. D. de Ilarraza, A. Casillas, On the480 creation of a clinical gold standard corpus in spanish: Mining adverse481 drug reactions, Journal of biomedical informatics 56 (2015) 318–332.482 [33] M. Marimon, B. Fisas, N. Bel, J. Vivaldi, S. Torner, M. Lorente,483 S. V´azquez, M. Villegas, The IULA Treebank., in: Lrec, pp. 1920–484 1926.485 [34] A. Duque, M. Stevenson, J. Martinez-Romo, L. Araujo, Co-occurrence486 graphs for word sense disambiguation in the biomedical domain, Artifi-487 cial intelligence in medicine 87 (2018) 9–19.488 [35] S. M. Jim´enez-Zafra, M. Taul´e, M. T. Mart´ın-Valdivia, L. A. Ure˜na-489 L´opez, M. A. Mart´ı, SFU Review SP-NEG: a Spanish corpus annotated490 with negation for sentiment analysis. a typology of negation patterns,491 Language Resources and Evaluation 52 (2018) 533–569.492 [36] S. Santiso, A. P´erez, A. Casillas, Exploring Joint AB-LSTM with em-493 bedded lemmas for Adverse Drug Reaction discovery, IEEE journal of494 biomedical and health informatics (2018).495 [37] M. Almagro, R. Mart´ınez Unanue, V. Fresno Fern´andez, S. Mon-496 talvo Herranz, Estudio preliminar de la anotaci´on autom´atica de c´odigos497 CIE-10 en informes de alta hospitalarios, SEPLN (2018).498 22
[38] H. Fabregat, A. Duque, J. Martinez-Romo, L. Araujo, Extending a Deep499 Learning Approach for Negation Cues Detection in Spanish, in: Pro-500 ceedings of the Iberian Languages Evaluation Forum (IberLEF 2019).501 CEUR Workshop Proceedings, CEUR-WS, Bilbao, Spain, p. 1.502 [39] H. Sak, A. Senior, F. Beaufays, Long short-term memory recurrent neu-503 ral network architectures for large scale acoustic modeling, in: Fifteenth504 annual conference of the international speech communication associa-505 tion, p. 1.506 [40] Y.-T. Zhou, R. Chellappa, Computation of optical flow using a neu-507 ral network, in: IEEE International Conference on Neural Networks,508 volume 1998, pp. 71–78.509 [41] J. Liu, W.-C. Chang, Y. Wu, Y. Yang, Deep learning for extreme multi-510 label text classification, in: Proceedings of the 40th International ACM511 SIGIR Conference on Research and Development in Information Re-512 trieval, ACM, pp. 115–124.513 [42] Y. Yin, Y. Song, M. Zhang, Nnembs at semeval-2017 task 4: Neural514 twitter sentiment classification: a simple ensemble method with different515 embeddings, in: Proceedings of the 11th International Workshop on516 Semantic Evaluation (SemEval-2017), pp. 621–625.517 [43] P. Bojanowski, E. Grave, A. Joulin, T. Mikolov, Enriching word vectors518 with subword information, Transactions of the Association for Compu-519 tational Linguistics 5 (2017) 135–146.520 [44] T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, J. Dean, Distributed521 representations of words and phrases and their compositionality, in:522 Advances in neural information processing systems, pp. 3111–3119.523 [45] J. Pennington, R. Socher, C. Manning, Glove: Global vectors for word524 representation, in: Proceedings of the 2014 conference on empirical525 methods in natural language processing (EMNLP), pp. 1532–1543.526 [46] C. Cardellino, Spanish Billion Words Corpus and Embeddings, 2016.527 [47] K. Lee, L. He, L. Zettlemoyer, Higher-order coreference resolution with528 coarse-to-fine inference, in: Proceedings of the 2018 Conference of the529 23
North American Chapter of the Association for Computational Linguis-530 tics: Human Language Technologies, Volume 2 (Short Papers), Associ-531 ation for Computational Linguistics, New Orleans, Louisiana, 2018, pp.532 687–692.533 [48] M. Fares, A. Kutuzov, S. Oepen, E. Velldal, Word vectors, reuse, and534 replicability: Towards a community repository of large-text resources,535 in: Proceedings of the 21st Nordic Conference on Computational Lin-536 guistics, Association for Computational Linguistics, Gothenburg, Swe-537 den, 2017, pp. 271–276.538 [49] M. Dermouche, J. Velcin, R. Flicoteaux, S. Chevret, N. Taright, Su-539 pervised topic models for diagnosis code assignment to discharge sum-540 maries, in: International Conference on Intelligent Text Processing and541 Computational Linguistics, Springer, pp. 485–497.542 [50] C. Manning, P. Raghavan, H. Sch¨utze, Introduction to information543 retrieval, Natural Language Engineering 16 (2010) 100–103.544 [51] R. Pascanu, T. Mikolov, Y. Bengio, On the difficulty of training recur-545 rent neural networks, in: International conference on machine learning,546 pp. 1310–1318.547 [52] S. Sabour, N. Frosst, G. E. Hinton, Dynamic routing between capsules,548 in: Advances in neural information processing systems, pp. 3856–3866.549 [53] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez,550 L. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in551 Neural Information Processing Systems, pp. 5998–6008.552 [54] Y. Zhang, J. Schneider, Multi-label output codes using canonical corre-553 lation analysis, in: Proceedings of the fourteenth international confer-554 ence on artificial intelligence and statistics, pp. 873–882.555 [55] J. Wehrmann, R. Cerri, R. Barros, Hierarchical multi-label classification556 networks, in: International Conference on Machine Learning, pp. 5225–557 5234.558 [56] K. Donnelly, Snomed-ct: The advanced terminology and coding system559 for ehealth, Studies in health technology and informatics 121 (2006)560 279.561 24