Global reconstruction of language models with linguistic rules – Explainable AI for online consumer reviews
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Binder, Markus; Heinrich, Bernd; Hopf, Marcus; Schiller, Alexander Article — Published Version Global reconstruction of language models with linguistic rules – Explainable AI for online consumer reviews Electronic Markets Provided in Cooperation with: Springer Nature Suggested Citation: Binder, Markus; Heinrich, Bernd; Hopf, Marcus; Schiller, Alexander (2022) : Global reconstruction of language models with linguistic rules – Explainable AI for online consumer reviews, Electronic Markets, ISSN 1422-8890, Springer, Berlin, Heidelberg, Vol. 32, Iss. 4, pp. 2123-2138, https://doi.org/10.1007/s12525-022-00612-5 This Version is available at: https://hdl.handle.net/10419/307273 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) 1 3 Electronic Markets (2022) 32:2123–2138 https://doi.org/10.1007/s12525-022-00612-5 RESEARCH PAPER Global reconstruction oflanguage models withlinguistic rules – Explainable AI foronline consumer reviews MarkusBinder1· BerndHeinrich1· MarcusHopf1· AlexanderSchiller1 Received: 30 May 2022 / Accepted: 27 October 2022 / Published online: 13 December 2022 © The Author(s) 2022 Abstract Analyzing textual data by means of AI models has been recognized as highly relevant in information systems research and practice, since a vast amount of data on eCommerce platforms, review portals or social media is given in textual form. Here, language models such as BERT, which are deep learning AI models, constitute a breakthrough and achieve leading-edge results in many applications of text analytics such as sentiment analysis in online consumer reviews. However, these language models are “black boxes”: It is unclear how they arrive at their predictions. Yet, applications of language models, for instance, in eCommerce require checks and justifications by means of global reconstruction of their predictions, since the decisions based thereon can have large impacts or are even mandatory due to regulations such as the GDPR. To this end, we propose a novel XAI approach for global reconstructions of language model predictions for token-level classifications (e.g., aspect term detection) by means of linguistic rules based on NLP building blocks (e.g., part-of-speech). The approach is analyzed on different datasets of online consumer reviews and NLP tasks. Since our approach allows for different setups, we further are the first to analyze the trade-off between comprehensibility and fidelity of global reconstructions of language model predictions. With respect to this trade-off, we find that our approach indeed allows for balanced setups for global reconstructions of BERT’s predictions. Thus, our approach paves the way for a thorough understanding of language model predictions in text analytics. In practice, our approach can assist businesses in their decision-making and supports compliance with regulatory requirements. Keywords Explainable AI· Text analytics· Language models· BERT· Linguistic rules· Online consumer reviews JEL Classification C80 Introduction Huge amounts of unstructured textual data are generated across various channels of information systems (IS) such as eCommerce platforms, review portals or social media every second (Potnis, 2018). Consequently, the need for techniques that automatically analyze textual data is increasing: Until 2028, the revenues from the natural language processing (NLP) market worldwide are expected to increase at a compound annual growth rate of almost 30% to over 100 billion USD, with text analytics expected to have the highest growth (Fortune Business Insights, 2021). As text analytics facilitate diverse applications such as sentiment analysis or text summarization (Young etal., 2018), various organizations in different business areas benefit from techniques of text analytics (Coheur, 2020; Zhang etal., 2020). For instance, product or service providers can use such techniques to analyze consumer sentiments in large amounts of online consumer reviews. Using this consumer feedback enables organizations Responsible Editor: Fethi Abderrahmane Rabhi. * Bernd Heinrich [email protected] Markus Binder [email protected] Marcus Hopf [email protected] Alexander Schiller [email protected] 1 University ofRegensburg, Germany, attheFaculty ofInformatics andData Science, Regensburg, Germany
2124 M.Binder et al. 1 3 to effectively improve their products and services (Chatterjee, 2019; Heinrich etal., 2022; Heinrich etal., 2020). The state-of-the-art techniques of text analytics are language models, such as the popular deep learning AI model ‘Bidirectional Encoder Representations from Transformers’ (BERT) (Devlin etal., 2019) or its descendants (e.g., ALBERT; Lan etal., 2020), as they have achieved leadingedge results in many tasks such as aspect-based sentiment analysis (Wang etal., 2018). Language models enable a contextualized representation of textual data by assessing the conditional probability of each token (e.g., a word) given the contextual tokens surrounding it (Peters etal., 2018a). Besides coarser classification tasks for sentences, for example, these language model representations can then be used, in particular, as basis for central token-level classifications such as aspect term and sentiment term detection. Since the language model BERT is already incorporated in a plethora of business IS applications, we demonstrate our approach by means of BERT as leading exponent of language models in this paper. Amongst others, popular application scenarios of BERT in electronic markets are eCommerce, chatbots, finance or online recruiting (Coheur, 2020; Dastin, 2018; Luo etal., 2022; Repke & Krestel, 2021; Shrestha etal., 2021; S. Xu etal., 2020; Yang etal., 2020; Zhang etal., 2020). However, similar to most other state-of-the-art deep learning models, BERT is a “black box”. That is, over 100 million learned parameters (Devlin etal., 2019) and various hidden layers contribute to BERT’s immense complexity, making it hardly (if at all) possible to comprehend why and how BERT arrives at its predictions (Kovaleva etal., 2019). To address this black box nature of AI models, a vastly increasing focus on explainable AI (XAI) in IS research and practice has emerged (Adadi & Berrada, 2018; Förster etal., 2021; Förster etal., 2020b). Literature agrees that the need for reconstructions and justifications is urgent and a ‘huge open scientific challenge’ (Guidotti etal., 2018). It is even expected that “algorithmic auditing and ‘data protection by design’ practices will likely become the new gold standard for enterprises deploying machine learning systems” (Casey etal., 2019). Thereby, regulations such as the General Data Protection Regulation (GDPR) in the European Union impose an extensive ‘right to explanation’ for automated data processing systems in general and thereby lay the foundation to enforce algorithmic auditing in companies. In particular, algorithmic auditing is highly relevant for domain experts, managers and data scientists that utilize the language models’ predictions for business-critical decisions or implementations and need to justify their actions. This is especially the case for application scenarios (AS) in electronic markets, as exemplarily outlined in the following and captured later on: • eCommerce (AS1): In eCommerce, BERT is used to conduct token-level classification in the course of sentiment analyses of online consumer reviews on online platforms such as Airbnb, Yelp or TripAdvisor for product development, services offerings and forecasting future demand (Heidari & Rafatirad, 2020; Shrestha etal., 2021; S. Xu etal., 2020). Since these analyses and decisions have large impacts, they require additional validation checks and justifications, far beyond measuring only the prediction accuracy of BERT. For instance, it needs to be ensured that specific groups of consumers are not discriminated against by assigning a negative sentiment to certain countries, ethnicities or genders. • Chatbots (AS2): In applications in consumer services (Luo etal., 2022), BERT-based chatbots conduct direct consumer interaction and embody the company’s voice. Thereby, reconstructions and justifications regarding the underlying BERT model are mandatory to prevent unhelpful, rude or misleading dialogues and thus, to support consumer satisfaction. • Financial applications (AS3): BERT descendants such as FinBERT (Yang etal., 2020) enable token-level classifications of financial entities, sentiments and their relations from texts such as social media posts (e.g., tweets from CEOs or other experts) or contract documents. The extracted information is used for key tasks in finance such as accounting, auditing, compliance and risk assessment. Furthermore, language models enable to automatically process millions of documents as contained in data leaks such as the Panama Papers (O’Donovan etal., 2019) for tax fraud detection. In particular, if legal actions are initiated based on predictions from language models (e.g., tax prosecution based on data leaks), validation checks are mandatory. • Online recruiting (AS4): Supporting text analytics of application documents (Schiller, 2019), language models such as BERT enable pre-processing und pre-filtering of applications and candidates on online job platforms. Here, auditing and validation are required as such automated recruitment may lead to discrimination (e.g., by gender or origin; Dastin, 2018). Reconstructions of models help to avoid such discriminations. These application scenarios show that it is crucial to reconstruct BERT’s predictions to be able to justify the decisions based thereon. Here, the reconstructions and explanations in these scenarios are required on a global level as in all those application scenarios the predictions of language models are used in ongoing operations on a daily basis. This means that a vast number of decisions are made based on these predictions day-by-day for newly generated and hitherto unknown textual data (e.g., chatbots or review summarizations are applied in real-time on consumer texts). Therefore, it is not feasible to use local approaches for reconstruction, as this would require huge efforts for manual checks of each local reconstruction and could practically only be
2125 Global Reconstruction of Language Models with Linguistic Rules 1 3 done a-posteriori if at all. Therefore, global approaches are essential for reconstructions of language model predictions in many applications. Here, we focus on global reconstructions of BERT’s predictions for token-level classifications in this work, since this constitutes popular application scenarios of BERT (e.g., AS1, AS3) and since BERT also establishes text representations based on tokens. Moreover, as Zafar etal. (2021) and Yan etal. (2022) indicate, a reconstruction approach for token-level classifications can also serve as a basis for reconstructions of coarser classification tasks, for instance, for sentence-level classifications (e.g., AS2, AS4). A promising way to obtain such a reconstruction and thus justify BERT’s predictions is to conduct a rule-based XAI approach. On the one hand, rules are highly concrete, which also has been emphasized by Förster etal. (2020a) as decisive XAI characteristic. Indeed, studies have shown that users “prefer, trust and understand rules better than alternatives” (Ribeiro etal., 2018; cf. also Arrieta etal., 2020). On the other hand, rule-based approaches preserve the AI model itself and thus, its high performance, while offering post-hoc reconstructions for explanations (Adadi and Berrada, 2018). Here, local rulebased approaches focus on explaining each prediction for a specific input separately, for instance, by using specific words to predict the sentiment term in a single sentence of an online consumer review. In contrast, global approaches aim at reconstructing the model’s predictions as a whole (Danilevsky etal., 2020). A global approach ideally requires a smaller rule set for reconstructing multiple predictions of a language model compared to local approaches that establish a separate and highly specific rule for each individual prediction and therefore are not really generalizable (Danilevsky etal., 2020). To enable such a global approach, our idea is to build rules based on linguistic information (so-called linguistic rules) which generalize specific words and sentences and can be modeled by NLP building blocks such as part-of- speech tags or dependency relations (Qi etal., 2020). Using NLP building blocks instead of single words as rule arguments is promising for global reconstruction, as they allow for rule arguments and rules analyzing (much) more than, for instance, one single sentence in an online consumer review. Moreover, NLP relation building blocks allow to account for the contextual information in a sentence (i.e., relations between words), which is crucial for the reconstruction of language model predictions for token-level classifications, since language models also use contextual information. Thus, we focus on the following main research question: RQ1: How can language model predictions for tokenlevel classifications be globally reconstructed by means of an XAI approach based on linguistic rules? Analogous to local reconstructions, a global reconstruction has to be analyzed regarding its fidelity (Danilevsky etal., 2020; Gilpin etal., 2018) and comprehensibility (Guidotti etal., 2018). In case of rule-based approaches, the comprehensibility of the rule set depends on the complexity (with respect to the length of the rules; cf. Guidotti etal., 2018) and the generalizability (words vs. NLP building blocks as discussed above) of the rules. Thereby, our approach allows for different setups regarding the comprehensibility of the rule set (e.g., by varying rule length), which is in general outlined as an important requirement of an XAI approach (Gilpin etal., 2018). This enables to analyze the trade-off between these two objectives in a reconstruction, which further supports adoption in IS. Thus, the second research question is as follows: RQ2: How can the trade-off between fidelity and comprehensibility of global reconstructions of language model predictions by linguistic rules be analyzed? Hence, our contribution is twofold: (1) We are the first to propose a global XAI approach for reconstructing predictions of language models by linguistic rules. In particular, (2) this paper is thus the first to analyze the trade-off between fidelity and comprehensibility (i.e., complexity and generalizability) in this setting. For our analysis, we focus on the highly relevant tasks of aspect term detection and sentiment term detection in online consumer reviews. To that end, we use two recognized online consumer review datasets from the domains of laptops and restaurants to account for different types of goods (i.e., laptops as search goods and restaurants as experience goods). We find that our linguistic rules are indeed suited for a global reconstruction of BERT’s predictions in online consumer reviews and in particular allow for balanced setups with respect to the trade-off between comprehensibility and fidelity of the reconstruction. The remainder of this paper is structured as follows. The next section presents the background of our research. Subsequently, we discuss how to globally reconstruct language models such as BERT with linguistic rules. Thereafter, we analyze different global reconstructions of BERT, discuss their results and outline implications for research and practice. Finally, we summarize the paper and provide an outlook on future research directions. Background In this section, we first outline which different types of XAI approaches exist in the context of language models. Second, several NLP building blocks recognized by literature are introduced forming the basis for our approach. The section concludes with a discussion of related work yielding the addressed research gap.
2126 M.Binder et al. 1 3 Types ofXAI approaches inthecontext oflanguage models To clarify the notion of XAI (i.e., what explainable AI really means), a characterization in opaque systems, interpretable systems and comprehensible systems has been proposed (Doran etal., 2017). Here, opaque systems offer no insights into the system’s reasoning on how inputs are mapped to the corresponding outputs. In that line, modern language models such as BERT are opaque systems, as it is not possible to comprehend its mappings, for instance, comprising over 100 million learned parameter values in the case of BERT. Based on that, there are two separate notions of addressing this problem. First, interpretable systems allow to understand how inputs are mapped to outputs by subdividing the mapping. This is not feasible for language models such as BERT due to its large amount of parameters and layers, which results in highly complex concatenated functions (Devlin etal., 2019). Second, comprehensible systems allow to relate properties of the inputs, for instance, single terms of an input sentence, to their output such as a classification of sentiment terms (Doran etal., 2017). While research in both areas is important, it has to be pointed out that the resulting XAI approaches are not “actually” explanation systems (Doran etal., 2017). For instance, rule-based approaches mostly give insights on how, but not why specific predictions are made (Doran etal., 2017). That is, causality cannot be directly established. To account for these different notions, we deliberately refer to “reconstructing” BERT rather than “explaining” in this paper. Related to the two notions of interpretable and comprehensible systems, there are, in general, two main approaches in XAI (Adadi and Berrada 2018): On the one hand, intrinsic XAI approaches ‘force’ the AI model (during training) to produce interpretable mappings from input to output (Adadi and Berrada 2018). The drawback of these intrinsic approaches is that they are limited in the type of interpretations they can provide, as they need to restrict the model to obtain interpretable mappings, thus usually worsening the model’s performance (Adadi and Berrada 2018). Due to its complexity, BERT would have to be extremely simplified to enable interpretable mappings. On the other hand, post-hoc XAI approaches aim to comprehensibly reconstruct the mappings from input to output of an AI model. These approaches do not require to restrict the model during training (Adadi and Berrada 2018). Here, a popular method is rule extraction, since rules can potentially exhibit a high degree of comprehensibility (Ribeiro etal., 2018). In general, there are two categories of rule extraction techniques (Adadi and Berrada 2018): 1)Decompositional rule extraction aims at extracting rules at selected, often single nodes within a neural network. To comprehend the predictions of a language model, it is then necessary to concatenate multiple extracted rules for various hidden layers. Thus, the drawback of this technique is that concatenations of rules are highly complex for deep neural networks such as BERT (Augasta & Kathirvalavakumar, 2012). Since the resulting rules would again be difficult to comprehend, decompositional rule extraction is not feasible for comprehensibly reconstructing language models. 2) In contrast,pedagogical rule extraction aims at extracting rules considering only the inputs and outputs. In particular, rules are extracted based on properties of the inputs and the corresponding outputs to reconstruct the mappings of the AI model. Thus, this approach can contribute to a comprehensible reconstruction even for language models such as BERT, since the extracted rules do not have to be concatenated through the various hidden layers. Additionally, a further important differentiation within post-hoc XAI research is between global and local approaches (Danilevsky etal., 2020). Here, global approaches aim at reconstructing the predictions of an AI model by means of one single global model (Danilevsky etal., 2020). In contrast, local approaches create separate, highly specific reconstruction models for each prediction (e.g., in a single sentence of an online consumer review). To enable local reconstructions for IS text analytics applications, rules solely based on specific words are used by extant literature (e.g., Ribeiro etal., 2018). However, such rules lack the ability to generalize. In contrast, linguistic rules based on NLP building blocks are more promising for the global reconstruction of language models. Indeed, rule arguments with NLP building blocks generalize much better than rule arguments with specific words, and NLP relation building blocks enable to incorporate contextual information, which is a main component of language models. Both objectives fidelity and comprehensibility are crucial for global post-hoc XAI approaches (Arrieta etal., 2020; Guidotti etal., 2018; Szczepański etal., 2021). Indeed, on the one hand, a global reconstruction needs to match the predictions of an AI model to avoid false conclusions, which is measured by fidelity (Gilpin etal., 2018). On the other hand, comprehensibility (i.e., complexity and generalizability; commonly measured in terms of model size) enables the use of the reconstruction (Guidotti etal., 2018). Thus, we analyze the reconstruction of BERT regarding its fidelity and its comprehensibility and strive to enable different setups between the two objectives. NLP building blocks To enable a reconstruction using linguistic rules, our idea is to use different semantical and syntactical NLP building blocks (cf. Introduction). Thus, we briefly outline NLP building blocks that are widely recognized in the literature (Fellbaum, 2013; Kamps etal., 2004; Tenney etal., 2019b) and that constitute a basis for our reconstruction. Table1
2127 Global Reconstruction of Language Models with Linguistic Rules 1 3 summarizes these different building blocks. Thereby, the column ‘type’ characterizes a building block as tag or relation (as described in the following). In addition, the column ‘linguistic information’ shows whether a building block provides semantic or syntactic information. For each building block, an example is given in the last column. A tag building block provides tag labels for selected tokens (e.g., words or punctuation marks) of a sentence. Tag labels describe a certain syntactic or semantic information of tokens in consideration of the whole sentence. Part-of-speech (POS) tags provide information on the syntactic structure of a sentence. Thereby, the POS tag, such as noun (NN), adjective (JJ) or verb (VB), is assigned to a single token. The building block synsets (SYN) considers the semantic information of tokens. In particular, SYN labels (e.g., derived from the lexical database WordNet) indicate words which share the same or a similar meaning (Fellbaum, 2013) taking into account its word context in a sentence. A relation building block provides a label for a pair of tokens in a sentence describing a certain syntactic or semantic relation between these tokens. These relation building blocks enable to account for the contextual information in a sentence (i.e., the relation between tokens in a sentence), which is crucial for a reconstruction of BERT as BERT also considers contextual information. A basic syntactic information is the distance between two tokens, which is covered by the proximity (PROX) building block. For instance, if two tokens are next to each other in a sentence, their distance is 1. Dependencies (DEP) also link two tokens based on their syntactical relationship, such as the adjectival modifier (amod) or nominal subject (nsubj) dependencies (Manning etal., 2014). Semantic information is provided by the building blocks semantic role labeling (SRL) and coreference (COREF). SRL relations identify combinations of predicates and semantic arguments in a sentence (Tenney etal. 2019b). COREF links two tokens referring to the same entity (Tenney etal. 2019a; b). Consequently, information referring to one part of the relation can be traced back to the other part. Related work Our goal is to reconstruct the language model BERT by means of linguistic (pedagogical) rules composed of NLP building blocks. Hence, XAI approaches analyzing language models regarding NLP building blocks (category A), XAI approaches analyzing pedagogical rules for reconstructing language models (category B) and XAI approaches for language models based on other techniques (category C) constitute the related work. In contrast, general rule-based XAI approaches (cf. Adadi and Berrada 2018) and XAI approaches (Ramon etal., 2020; Sushil etal., 2018) relying on a simple ‘bag-of-words’ analysis – both without any focus on language models – are not in the scope for our research. Ad category A): Several existing works analyze language models by using their (contextualized) word embeddings or internal states as input to predict NLP building blocks (Coenen etal., 2019; Hewitt & Manning, 2019; Jumelet & Hupkes, 2018; Kim etal., 2019; Peters etal., 2018b; Tenney etal., 2019a; Tenney etal. 2019b; Van Aken etal., 2019). Then, the quality of these predictions is used as an indication whether a certain NLP building block is encoded in particular word embeddings (i.e., vector representations) or specific layers of the language models. That is, instead of reconstructing predictions of language models for NLP tasks in IS (e.g., sentiment term detection), an analysis of the general word embeddings themselves is aimed for in these works. For instance, different NLP building blocks have been predicted by word embeddings of the language models ELMo (Peters etal. 2018b) and BERT (Tenney etal. 2019a; b). However, the aim of our research is a different one. As discussed in the Introduction, our focus is to better comprehend BERT’s predictions on NLP tasks in IS, for instance, to be able to justify decisions made based on its results. To enable that, it is necessary to reconstruct the predictions of BERT for relevant NLP tasks (such as the extracted sentiment terms in online consumer reviews), since these predictions and not particular word embeddings in form of vector representations are the foundation for further decisions. In Table 1 Overview of NLP building blocks Building block Type Linguistic information Example labels for the sentence “The waiter of The Burger House was nice, he smiled at us.” Part-of-speech tags (POS) Tags Syntactic POS-label (“waiter”) = NN (Noun) Synsets (SYN) Tags Semantic SYN-label (“nice”) = nice.a.01 (Synset description: “pleasant or pleasing or agreeable in nature or appearance”) Dependencies (DEP) Relations Syntactic DEP-label (“waiter”, “nice”) = amod (adjectival modifier) Semantic role labeling (SRL) Relations Semantic SRL-label (“he”, “smiled”) = agent-predicate-relation Coreferences (COREF) Relations Semantic COREF-label (“waiter”, “he”) = True (referring to the same entity) Proximity (PROX) Relations Syntactic PROX-label (“waiter”, “nice”) = 6
2128 M.Binder et al. 1 3 that line, none of the approaches in this category considers pedagogical rules to enable a reconstruction of predictions of a language model for NLP tasks in IS. Ad category B): There also exist recent, interesting works that analyze language models by means of pedagogical rules in a local manner (i.e., for single predictions). In Ribeiro etal. (2018), individual predictions of simple recurrent neural network-based language models are reconstructed by separate if–then rules. Building on this work, BERT’s predictions in an application of fake news detection on social media are analyzed in Szczepański etal. (2021). Both works hardly incorporate contextual information for reconstructions. That is, only information of the previous token is considered to obtain local reconstruction rules. Thus, both works consider only short rules of low complexity. In addition, rules based on individual tokens (e.g., specific words) are used. Hence, both works do not discuss the composition of tag and relation building blocks when extracting rules for reconstruction and as a result, the proposed rules exhibit only low generalizability. In particular, relation building blocks such as DEP or COREF, which enable rules to comprise vital contextual information, are not considered. Ad category C): Moreover, local non-rule-based XAI approaches have been proposed to reason language model predictions. In Malkiel etal. (2022), saliency maps are used to reason similarity predictions of online consumer reviews by a BERT-based model, aiming to highlight important word-pairs for specific similarity predictions. Moreover, different visualizations with respect to neuron activations in the hidden layers have been applied to reason specific language model predictions (Brasoveanu & Andonie, 2022). In Kokalj etal. (2021), the known feature importance XAI approach ‘shapley additive explanations’ (Lundberg & Lee, 2017) has been adapted to account for the contextualized (token-based) text representation in language models. Further, approaches based on the attention weights in language models have been recently proposed (Ali etal., 2022; S. Liu etal., 2021), similarly establishing feature importance scores for language model predictions. As an application case, these approaches aim to determine important words for sentence sentiment classifications of a BERT-based model. However, all of these works focus on local reconstructions, for instance, for individual sentences, for which they do not consider NLP building blocks. That is, global (token-level) reconstructions by linguistic rules are out of their scope. Overall, while the approaches in category A) give interesting indications on how NLP building blocks may be encoded in contextualized word embeddings, they do not enable to reconstruct the predictions of language models in NLP tasks in IS. In contrast, the approaches in category B) indeed analyze rules for reconstructing specific predictions, but only enable local reconstructions and do not incorporate different NLP building blocks comprising contextual linguistic information. Thus, they exhibit only low generalizability. Similarly, the approaches in category C) focus on reasoning specific language model predictions locally by non-rule-based approaches and do not incorporate different NLP building blocks either. Summing up, there are very interesting contributions in the field of XAI regarding language models. However, literature lacks an approach for global reconstructions of language model predictions for NLP tasks in IS (e.g., sentiment term detection in online consumer reviews) based on pedagogical rules. To address this research gap, this paper proposes, to the best of our knowledge, the first global XAI approach for reconstructing token-level language model predictions by linguistic (pedagogical) rules. In particular, this paper is thus the first to enable an analysis of the trade-off between fidelity and comprehensibility (i.e., complexity and generalizability) in this setting. Global reconstruction ofBERT withlinguistic rules In this section, we introduce our approach by postulating the formal structure of linguistic rules for the global reconstruction of BERT’s predictions and then outline appropriate measures to analyze this reconstruction. Formal structure oflinguistic rules forreconstructing BERT’s predictions We begin by deriving the formal structure of linguistic rules. Thereby, for illustration purpose, the language model BERT is applied for the token classification tasks aspect term detection and sentiment term detection that are frequently used in online consumer reviews (Dai & Song, 2019; Sun etal., 2019; H. Xu etal. 2019). More precisely, each sentence in a document comprises a string value and can be split up by tokenization into disjunct substrings (so-called tokens), which have a linguistic meaning, such as (sub)words or punctuation marks. The precise tokenization of sentences depends on specific tokenization policies. For this work, we used w.l.o.g. the widely applied tokenization of the python package NLTK (cf. https:// www. nltk. org). The goal of the token classification tasks performed by BERT is to assign class labels to such tokens. For example, the second token ‘fish’ in the tokenized sentence (‘The’, ‘fish’, ‘was’, ‘good’, ‘!’) is assigned with the class label ASP indicating an aspect term. The following postulates P1)-P3) provide the foundation for linguistic rules based on NLP building blocks, which enable a global reconstruction of BERT's predictions (i.e., the predicted class labels for the tokens of a sentence). P1) “Label assignments”: In our approach, we assign labels only to single tokens or token pairs. Hence, we do not
2129 Global Reconstruction of Language Models with Linguistic Rules 1 3 consider label assignments for whole sentences, documents nor for single character values. This focus is promising for reconstructing BERT, as BERT internally also establishes text representations on a token level. P1.1) “tag label assignments”: A tag building block tbb ∈TBB (where TBB is the set of tag building blocks) assigns at most one tag label tbb(ti)∈Ltbb to a token ti ( Ltbb is the set of all labels from tbb ). For instance, the tag building block POS with LPOS ={NN,VB,JJ,…} assigns the label POS( t 2) == NN (= ‘noun’) to the token t2= ‘fish’ in the exemplary sentence above. P1.2) “Relation label assignments”: A relation building block rbb ∈RBB (where RBB is the set of relation building blocks) assigns at most one relation label rbb(ti,tj)∈Lrbb to a token pair ( t i ,t j) ( Lrbb is the set of all labels from rbb ). For example, the relation building block DEP with LDEP ={amod,nsubj,…} assigns the label DEP( t 2 ,t 4) == nsubj (= ‘nominal subject’) to the token pair ( t2,t4 ) = (‘fish’, ‘good’). In particular, relation building blocks enable to capture contextual information in a sentence, which is a main component of language models such as BERT. P1.3) “Class label assignments”: BERT assigns a class label l 𝜏 ( t i) ∈L 𝜏 to each token ti ( L𝜏 is the set of all class labels in a token classification task 𝜏 ). For instance, in the aspect term detection task with class labels L ASP = { ASP,ASP } , the token t2= ‘fish’ is assigned with the class label ASP by BERT indicating that ‘fish’ is an aspect term. P2) “Feasible aRguments FoR Rules”: In our approach, feasible arguments in the antecedent and consequents of a rule only reference to labels for tokens or token pairs as postulated in P1). P2.1) “Feasible aRguments in Rule antecedents”: A feasible argument in the rule antecedent only contains conditions regarding tag labels of tokens (cf. P1.1)) and relation labels of token pairs (cf. P1.2)). P2.2) “Feasible aRguments in Rule consequents”: A feasible argument in the rule consequent only contains class label assignments of tokens (cf. P1.3). Considering the classification task of sentiment term detection, the argument lSENT ( t4 ) → SENT assigns the class label SENT to the token t4= ‘good’, indicating that ‘good’ is labelled as a sentiment term by BERT in the sentence ‘The fish was good!’. P3) “ConFlicting classiFication Results oF multiple Rules”: Multiple rules R 1 ,…,RnR ( nR∈ℕ ) may result in conflicting classification results l1 𝜏( t i) ,…,l n R 𝜏 ( t i) ∈L𝜏 for the same token ti . To resolve such conflicting classification results for a token ti , it is sensible to assign the class of the rule with the highest precision (cf. next section). Given the postulates P1)-P3), the structure of linguistic rules can be defined. A linguistic rule R is an “if–then-else” rule in the form of IF antecedent THEN “then”-consequent (ELSE “else”-consequent). Here, the antecedent is an arbitrary combination of feasible arguments as postulated in P2.1) by means of logical operators such as AND (i.e., “ ∧ ”), OR (i.e., “ ∨ ”) and NOT (i.e., “ ¬ ”). Further, each “then”-con- sequent and each “else”-consequent consists of one feasible argument as postulated in P2.2). Thus, a rule R outputs the class assignments of the “then”-consequent in case that the antecedent is tRue (otherwise and if an “else”-consequent is contained in the rule, it outputs the class assignments of the “else”-consequent). Moreover, rules can be characterized by their length, which is given by the number of tokens that are connected by a relation building block in the antecedent of a rule. A brief example of a rule of length two is given by: IF ([ POS ( t i) == NN ] ∨¬ [ POS ( t j) == VB ]) ∧ [ DEP ( t i ,t j) == nsubj ] THEN lASP( t i) → ASP This rule can be applied to the tokenized sentence (‘The’, ‘fish’, ‘was’, ‘good’, ‘!’) from above. For this sentence, the antecedent of the rule is only tRue if ti=t2= ‘fish’ and tj=t4= ‘good’. For any other selection of ti and tj , the antecedent is False since only the token pair (‘fish’, ‘good’) has the relation “nsubj” in this sentence. Hence, this linguistic rule correctly detects the aspect term ‘fish’. Rules of the outlined formal structure based on the postulates P1)-P3) constitute the foundation for our approach for reconstructing BERT. Assessing fidelity andcomprehensibility ofglobal reconstructions To globally reconstruct BERT, all predictions of BERT for a token classification task have to be considered. Here, fidelity and comprehensibility are the most relevant measures (cf. Section “Types of XAI approaches in the context of language models”) and assessing both measures is required to analyze the trade-off between fidelity and comprehensibility. Since we focus on global reconstructions of language models, we outline in detail how both measures can be assessed for global reconstructions in the following. To measure fidelity, we consider the predictions of BERT for each class label. More precisely, the set of token ids (i.e., the positions of tokens in the text corpus) predicted by BERT as class C∈L𝜏 is given by IC,BERT = { i∈I | l BERT ( t i) =C } , where I is the set of all token ids. These token ids IC , BERT are used as the basis for extracting the linguistic rules on training data Itrain , C , BERT and validation data Ivalidation , C , BERT as well as for assessing their fidelity of globally reconstructing BERT on test data Itest , C , BERT . Once a set Σ of linguistic rules is extracted, the F1 score is appropriate to assess the fidelity of the rule set (Sushil etal., 2018) as - in contrast to the accuracy measure - it accounts for imbalanced class distributions. The F1 score (i.e., based on precision and recall) of
2130 M.Binder et al. 1 3 the rule set Σ for reconstructing BERT’s predictions IC , BERT is given by: Here, Itest,C,Σ = { i∈I test | l Σ( t i) == C } is the set of token ids from the test data that are assigned with class C by the rule set Σ . In case of multiclass classification the fidelity is then assessed by the average F1 score per class label C , denoted as F1(Σ) (i.e., by the macro-averaged F1 score (Sushil etal., 2018)). In contrast to the regular formulas for classifier evaluation, which aim to evaluate the predictions of a classifier regarding the true class labels, the formulas (1)-(3) enable to evaluate the linguistic rules regarding the predicted class labels by BERT and hence, to assess the fidelity of reconstructing BERT by certain sets of linguistic rules Σ . In contrast to the comprehensibility of local reconstructions (e.g., complexity of single rules), literature suggests to assess the comprehensibility of a global reconstruction by its model size (Guidotti etal., 2018). Since our model is a set of rules Σ , both the number of rules NR(Σ) in the rule set and the number of unique argument values NUAV (Σ) in the antecedents in the rule set (Vilone & Longo, 2021) determine its comprehensibility. These measures are given by: Here, AAV = LPOS ∪ LSYN ∪ LDEP ∪ LSRL ∪ LCOREF ∪ LPROX is the set of all argument values of all NLP building blocks. For both measures, a lower value indicates higher comprehensibility. That is, we leverage two different measures which capture two important perspectives on comprehensibility. Overall, based on the measures (1) – (5) the fidelity and comprehensibility of global reconstructions can be assessed. Analysis In this section we analyze the reconstruction of BERT’s predictions by our approach. First, we outline the selected tasks, datasets and the conducted automated extraction of linguistic rules for global reconstruction. Then, we demonstrate how our approach based on linguistic rules can reconstruct (1) Pr C(Σ)= | | Itest,C,BERT ∩Itest,C,Σ | | | | I test,C,Σ| | (2) Rec C(Σ)= | | Itest,C,BERT ∩Itest,C,Σ | | | | I test,C,BERT | | (3) F 1C(Σ)= 2∗Pr C (Σ)∗Rec C (Σ) Pr C (Σ)+Rec C (Σ) (4) NR(Σ)=|Σ| (5) NUAV (Σ)=|{v∈AAV |∃R∈Σ∶v∈R}| predictions of BERT. After that, we present and discuss the results as well as implications for research and practice. Task selection, data preparation andrule extraction For a meaningful analysis of the reconstruction of BERT’s predictions, we selected the NLP tasks aspect term detection and sentiment term detection as these tasks are frequently analyzed in the IS field and constitute common applications for BERT and text analytics (Dai and Song, 2019; Sun etal., 2019; H. Xu etal., 2019), in particular in electronic markets (Chatterjee etal., 2021; Steur etal., 2022). Also, we chose two publicly available datasets that exhibit different characteristics – with restaurants reviews from the platform Yelp (Yelp Dataset Challenge; cf. https:// www. yelp. com/ datas et) as experience goods vs. laptop reviews from the platform Amazon (Ni etal., 2019) as search goods – to enable broader insights independent of specific item domains. To extract linguistic rules based on the formal structure postulated in the previous section, we used state-of-the-art toolkits for annotating both datasets with the NLP building blocks discussed in Section “NLP building blocks” and leveraged and extended rule generation and rule selection techniques from the literature. The following paragraphs provide more details. The goal of aspect term detection and sentiment term detection is to classify tokens in online consumer reviews that express aspects or sentiments. An aspect term (e.g., ‘laptop screen’) represents an item aspect for which an opinion polarity is expressed by a sentiment term (e.g., ‘very good’) (Sun etal., 2019). The task of token classification is to assign a class label C∈L𝜏 (i.e., L ASP = { ASP,ASP } and L SENT = { SENT,SENT } ) to tokens of a sentence. To conduct aspect term detection and sentiment term detection, we used the publicly available state-of-the-art language model BERT. In particular, we used pre-trained BERT models, which were specifically adapted to the domains of restaurant reviews and laptop reviews, respectively (H. Xu etal., 2019). We fine-tuned these BERT models for the tasks aspect term and sentiment term detection on both domains using the publicly available, labeled dataset SemEval2014 provided by Fan etal. (2019). After that, the fine-tuned BERT models were used in this work to predict aspect terms and sentiment terms in the two review datasets. That is, the tokens of both review datasets were assigned with the class labels of BERT’s predictions. An overview of the (randomly sampled) dataset excerpts used for analysis, including the predictions of BERT regarding both tasks, is given in Table2. For annotation of NLP building blocks on these datasets, we used the state-of-the-art toolkits Stanza (Qi etal., 2020) and AllenNLP (Gardner etal., 2018) as well as the lexical database WordNet (Fellbaum, 2013). More precisely, POS tags and DEP relations were annotated based
2137 Global Reconstruction of Language Models with Linguistic Rules 1 3 explanations. Proceedings of the 54th Hawaii International Conference on System Sciences (p. 1274). Förster, M., Klier, M., Kluge, K., & Sigler, I. (2020a). Evaluating explainable artifical intelligence‐What users really appreciate. Proceedings of the 28th European Conference on Information Systems (ECIS). Förster, M., Klier, M., Kluge, K., & Sigler, I. (2020b). Fostering human agency: A process for the design of user-centric XAI systems. ICIS 2020 Proceedings. Fortune Business Insights (2021). Natural Language Processing (NLP) Market size, share and Covid-19 impact analysis. Retrieved from https:// www. fortu nebus iness insig hts. com/ indus tryrepor ts/ natur allangu ageproce ssingnlp- market- 101933. Accessed 30 Aug2022. Gardner, M., Grus, J., Neumann, M., Tafjord, O., Dasigi, P., Liu, N., Peters, M., Schmitz, M., & Zettlemoyer, L. (2018). AllenNLP: A deep semantic natural language processing platform. ArXiv Preprint. https:// doi. org/ 10. 48550/ arXiv. 1803. 07640 Geng,Z., Zhang,Y. [Yanhui], & Han,Y. (2021). Joint entity and relation extraction model based on rich semantics. Neurocomputing, 429, 132–140. https:// doi. org/ 10. 1016/j. neucom. 2020. 12. 037 Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., & Kagal, L. (2018). Explaining explanations: An overview of interpretability of machine learning. 2018 IEEE 5th International Conference on data science and advanced analytics (DSAA) (pp. 80–89). IEEE. Goeken, T., Tsekouras, D., Heimbach, I., & Gutt, D. (2020). The rise of robo-reviews-The effects of chatbot-mediated review elicitation on review valence. ECIS 2020 Proceedings. Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., & Pedreschi, D. (2018). A survey of methods for explaining black box models. ACM Computing Surveys (CSUR), 51(5), 1–42.https:// doi. org/ 10. 1145/ 32360 09 Heidari, M., & Rafatirad, S. (2020). Semantic convolutional neural network model for safe business investment by using BERT. 2020 Seventh International Conference on Social Networks Analysis, Management and Security (SNAMS) (pp. 1–6). IEEE. https:// doi. org/ 10. 1109/ SNAMS 52053. 2020. 93365 75 Heinrich, B., Hollnberger, T., Hopf, M., & Schiller, A. (2022). Longterm sequential and temporal dynamics in online consumer ratings. ECIS 2022 Proceedings. Heinrich, B., Hopf, M., Lohninger, D., Schiller, A., & Szubartowicz, M. (2020). Something’s missing? A procedure for extending item content data sets in the context of recommender systems. Information Systems Frontiers, 24, 267–286. https:// doi. org/ 10. 1007/ s10796- 020- 10071-y Heinrich, B., Hopf, M., Lohninger, D., Schiller, A., & Szubartowicz, M. (2021). Data quality in recommender systems: the impact of completeness of item content data on prediction accuracy of recommender systems. Electronic Markets, 31(2), 389–409. https:// doi. org/ 10. 1007/ s12525- 019- 00366-7 Hewitt, J., & Manning, C. D. (2019). A structural probe for finding syntax in Word representations. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4129–4138). Jumelet, J., & Hupkes, D. (2018). Do language models understand anything? On the ability of LSTMs to understand negative polarity items. Proceedings of the Workshop: Analyzing and Interpreting Neural Networks for NLP (BlackboxNLP@EMNLP 2018) (pp. 222–231). ACL. Kamps, J., Marx, M., Mokken, R. J., & de Rijke, M. (2004). Using WordNet to measure semantic orientations of adjectives. In LREC (Vol. 4, pp. 1115–1118). ACL. Kim, N., Patel, R., Poliak, A., Wang, A., Xia, P., McCoy, R. T., Tenney, I., Ross, A., Linzen, T., Van Durme, B., Bowman, S. R., & Pavlick, E. (2019). Probing what different NLP tasks teach machines about function word comprehension. Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019). ACL. Kokalj, E., Škrlj, B., Lavrač, N., Pollak, S., & Robnik-Šikonja, M. (2021). BERT meets shapley: Extending SHAP explanations to transformer-based classifiers. Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation (pp. 16–21). Kovaleva, O., Romanov, A., Rogers, A., & Rumshisky, A. (2019). Revealing the dark secrets of BERT. In EMNLP-IJCNLP (pp. 4365–4374). ACL. https:// doi. org/ 10. 18653/ v1/ D19- 1445 Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2020). ALBERT: A Lite BERT for self-supervised learning of language representations. Proceedings of the International Conference on Learning Representations 2020 (ICLR). Liu, Q., Gao, Z., Liu, B., & Zhang, Y. [Yuanlin] (2015). Automated rule selection for aspect extraction in opinion mining. Twenty- Fourth international joint conference on artificial intelligence. AAAI. Liu, S., Le, F., Chakraborty, S., & Abdelzaher, T. (2021). On exploring attention-based explanation for transformer models in text classification. 2021 IEEE International Conference on Big Data (Big Data) (pp. 1193–1203). IEEE. Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in neural information processing systems, 30. Luo,B., Lau,R.Y.K., Li,C., & Si,Y.-W. (2022). A critical review of state‐of‐the‐art chatbot designs and applications. WIREs Data Mining and Knowledge Discovery, 12(1). https:// doi. org/ 10. 1002/ widm. 1434 Malkiel, I., Ginzburg, D., Barkan, O., Caciularu, A., Weill, J., & Koenigstein, N. (2022). Interpreting BERT-based text similarity via activation and saliency maps. Proceedings of the ACM Web Conference 2022 (pp. 3259–3268). Manning, C. D., Surdeanu, M., Bauer, J., Finkel, J., Bethard, S. J., & McClosky, D. (2014). The Stanford CoreNLP natural language processing toolkit. In ACL System Demonstrations (pp. 55–60). ACL. Retrieved from http:// www. aclweb. org/ antho logy/P/ P14/ P14- 5010. Accessed 30 Aug2022. Ni, J., Li, J., & McAuley, J. (2019). Justifying recommendations using distantly-labeled reviews and fine-grained aspects. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 188–197). O’Donovan, J., Wagner, H. F., & Zeume, S. (2019). The value of offshore secrets: Evidence from the Panama Papers. The Review of Financial Studies, 32(11), 4117–4155. https:// doi. org/ 10. 1093/ rfs/ hhz017 Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018a). Deep contextualized word representations. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (pp. 2227–2237). Peters, M. E., Neumann, M., Zettlemoyer, L., & Yih, W. (2018b). Dissecting contextual word embeddings: Architecture and representation. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. ACL. Potnis, A. (2018). Illuminating insight for unstructured data at scale. Retrieved from https:// www. ibm. com/ downl oads/ cas/ Z2ZBA Y6R. Accessed 30 Aug2022. Qi, P., Zhang, Y. [Yuhao], Zhang, Y. [Yuhui], Bolton, J., & Manning, C. D. (2020). Stanza: A Python natural language processing toolkit for many human languages. In ACL System Demonstrations
2138 M.Binder et al. 1 3 (pp. 101–108). ACL. Retrieved from https:// ar xiv. org/ pdf/ 2003. 07082. Accessed 30 Aug2022. Ramon, Y., Martens, D., Evgeniou, T., & Praet, S. (2020). Metafeatures-based rule-extraction for classifiers on behavioral and textual data. ArXiv Preprint. Accessed 30 Aug2022.https:// doi. org/ 10. 48550/ arXiv. 2003. 04792 Repke, T., & Krestel, R. (2021). Extraction and representation of financial entities from text. In S. Consoli, D. Reforgiato Recupero, & M. Saisana (Eds.), Springer eBook Collection. Data science for economics and finance: Methodologies and applications (pp. 241–263). Cham, Switzerland: Springer k. https:// doi. org/ 10. 1007/ 978-3- 030- 66891-4_ 11 Ribeiro, M. T., Singh, S., & Guestrin, C. (2018). Anchors: Highprecision model-agnostic explanations. Proceedings of the AAAI conference on artificial intelligence (Vol. 32, No. 1). Schiller, A. (2019). Knowledge discovery from CVs: A topic modeling procedure. Proceedings of the 14th International Conference on business informatics (Wirtschaftsinformatik). Shrestha, Y. R., Krishna, V., & von Krogh, G. (2021). Augmenting organizational decision-making with deep learning algorithms: Principles, promises, and challenges. Journal of Business Research, 123, 588–603. https:// doi. org/ 10. 1016/j. jbusr es. 2020. 09. 068 Steur, A. J., Fritzsche, F., & Seiter, M. (2022). It’s all about the text: An experimental investigation of inconsistent reviews on restaurant booking platforms. Electronic Markets, 32(3), 1187–1220. https:// doi. org/ 10. 1007/ s12525- 022- 00525-3 Sun, C., Huang, L., & Qiu, X. (2019). Utilizing BERT for aspect-based sentiment analysis via constructing auxiliary sentence. Conference of the North American Chapter of the ACL (pp. 380–385). ACL. https:// doi. org/ 10. 18653/ v1/ N19- 1035 Sushil, M., Šuster, S., & Daelemans, W. (2018). Rule induction for global explanation of trained models. Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP (pp. 82–97). ACL. Szczepański, M., Pawlicki, M., Kozik, R., & Choraś, M. (2021). New explainability method for BERT-based model in fake news detection. Nature Scientific Reports, 11(1), 1–13.https:// doi. org/ 10. 1038/ s41598- 021- 03100-6 Tenney, I., Das, D., & Pavlick, E. (2019a). Bert rediscovers the classical nlp pipeline. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. ACL. Tenney, I., Xia, P., Chen, B., Wang, A., Poliak, A., McCoy, R. T., Kim, N., Van Durme, B., Bowman, S. R., Das, D., & Pavlick, E. (2019b). What do you learn from context? Probing for sentence structure in contextualized word representations. International Conference on Learning Representations 2019 (ICLR). Van Aken, B., Winter, B., Löser, A., & Gers, F. A. (2019). How does BERT answer questions? A layer-wise analysis of transformer representations. Proceedings of the 28th ACM International Conference on Information and Knowledge Management (pp. 1823–1832). Vilone, G., & Longo, L. (2021). A Quantitative evaluation of global, rule-based explanations of post-hoc, model agnostic methods. Frontiers in Artificial Intelligence, 4. https:// doi. org/ 10. 3389/ frai. 2021. 717899 Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., & Bowman, S. (2018). GLUE: A multi-task benchmark and analysis platform for natural language understanding. EMNLP Workshop BlackboxNLP (pp. 353–355). ACL. https:// doi. org/ 10. 18653/ v1/ W18- 5446 Xu, H., Liu, B., Shu, L., & Yu, P. (2019). BERT post-training for review reading comprehension and aspect-based sentiment analysis. Conference of the North American Chapter of the ACL (pp. 2324–2335). ACL. https:// doi. org/ 10. 18653/ v1/ N19- 1242 Xu, S., Barbosa, S. E., & Hong, D. (2020). BERT feature based model for predicting the helpfulness scores of online customers reviews. In K. Arai, S. Kapoor, & R. Bhatia (Eds.), Advances in Intelligent Systems and Computing. Advances in Information and Communication (Vol. 1130, pp. 270–281). Cham: Springer International Publishing. https:// doi. org/ 10. 1007/ 978-3- 030- 39442-4_ 21 Yan, H., Gui, L., & He, Y. (2022). Hierarchical interpretation of neural text classification. ArXiv Preprint. https:// doi. org/ 10. 48550/ arXiv. 2202. 09792 Yang, Y., Uy, M. C. S., & Huang, A. (2020). FinBERT: A pretrained language model for financial communications. ArXiv Preprint. https:// doi. org/ 10. 48550/ arXiv. 2006. 08097 Yin, D., Bond, S. D., & Zhang, H. (2014). Anxious or angry? Effects of discrete emotions on the perceived helpfulness of online reviews. MIS Quarterly, 38(2), 539–560. https:// doi. org/ 10. 25300/ MISQ/ 2014/ 38.2. 10 Young, T., Hazarika, D., Poria, S., & Cambria, E. (2018). Recent trends in deep learning based natural language processing. IEEE Computational intelligence magazine, 13(3), 55–75. https:// doi. org/ 10. 1109/ MCI. 2018. 28407 38 Zafar, M. B., Schmidt, P., Donini, M., Archambeau, C., Biessmann, F., Das, S. R., & Kenthapadi, K. (2021). More than words: Towards better quality interpretations of text classifiers. ArXiv Preprint. https:// doi. org/ 10. 48550/ arXiv. 2112. 12444 Zhang, R., Yang, W., Lin, L., Tu, Z., Xie, Y., Fu, Z., Xie, Y., Tan, L., Xiong, K., Lin, J. (2020). Rapid adaptation of BERT for information extraction on domain-specific business documents. ArXiv Preprint. https:// doi. org/ 10. 48550/ arXiv. 2002. 01861 Publisher's note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.