Full text
International Journal of Approximate Reasoning 171 (2024) 109128 Available online 17 January 2024 0888-613X/© 2024 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/). Contents lists available at ScienceDirect International Journal of Approximate Reasoning journal homepage: www.elsevier.com/locate/ijar Enriching interactive explanations with fuzzy temporal constraint networks Mariña Canabal-Juanatey a, Jose M. Alonso-Morala,b,∗, Alejandro Catalaa,b, Alberto Bugarín-Diz a,b aCentro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela, 15782 Santiago de Compostela, Spain bDepartamento de Electrónica e Computación, Universidade de Santiago de Compostela, 15782 Santiago de Compostela, Spain A R T I C L E I N F O A B S T R A C T Keywords: Fuzzy temporal constraint networks Fuzzy temporal reasoning Knowledge graphs Language models Conversational agents Humans often use expressions with vague terms which play a fundamental role for effective communication. These expressions are successfully modeled with fuzzy technology, but they are not usually integrated yet with Natural Language Processing models and techniques. Large-scale pre-trained language models yield excellent results in many language tasks, but they have some drawbacks such as their lack of transparency and thorough temporal reasoning capabilities. Therefore, the use of such models may provoke inconsistent or incorrect dialogues in the context of conversational agents which were aimed at providing users of intelligent systems with interactive explanations. In this paper, we propose a model for fuzzy temporal reasoning to overcome some inconsistencies detected in pre-trained language models in a specific application domain of a conversational agent carefully designed for providing users with explanations which are endowed with a good balance between naturalness and fidelity. More precisely, starting from a knowledge graph that provides an intuitive representation of the entities and relations in the application domain, we describe how to map the temporal information onto a fuzzy temporal constraint network. This formalism allows to represent imprecise temporal information and provides mechanisms for checking consistency in conversations. In addition, as a proof of concept, we have developed TimeVersa, a conversational agent which integrates the proposed model into an application domain (i.e., a virtual assistant for tourists) that requires handling imprecise temporal constraints. We illustrate in a use case how the agent can identify temporal inconsistencies and answer queries related to temporal information properly. Results after a user study report that users’ perception of consistency is significantly higher in a conversation with TimeVersa than in a similar conversation using the well-known GPT-3 Large Language Model, when vague temporal information is involved. The proposed approach is a step forward for developing conversational agents operating in application domains that require temporal reasoning under uncertainty. * Corresponding author. E-mail addresses: [email protected] (M. Canabal-Juanatey), [email protected] (J.M. Alonso-Moral), [email protected] (A. Catala), [email protected] (A. Bugarín-Diz). https://doi.org/10.1016/j.ijar.2024.109128 Received 13 January 2023; Received in revised form 8 November 2023; Accepted 14 January 2024
International Journal of Approximate Reasoning 171 (2024) 109128 2 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. 1. Introduction The emergence of disruptive Deep Learning techniques and tools has revolutionized the approach to applications, services and tasks supported by Natural Language Technologies. Currently, the Natural Language Processing (NLP) field is undergoing a paradigm shift with the rise of Neural Language Models, also known as pre-trained Language Models [1], such as those based on the Bidirectional Encoder Representations from Transformers (BERT) [2]and variants like RoBERTa [3]or ALBERT [4], but also Generative Pre-trained Transformers (GPT) [5]. These models are trained on broad data at scale and are adaptable to a wide range of monolingual and multilingual downstream tasks. For instance, GPT-3 [5]was launched in 2021, it has 175,000 million parameters and 96 layers trained on a corpus of 499,000 million tokens of web content. Other models have more than three times the size, with more than 530,000 million parameters [6], and GPT-4 is announced to be with 1.76 trillion parameters. From the data used for training, basically Internet-based corpora, from Wikipedia to the New York Times, primarily in English, the ability of these models to transform natural language specifications into websites, create basic financial reports or solve language puzzles has been very promising so far. This quick paradigm shift means that we have only just started to discover the new possibilities and concerns raised by these models. Compared to the state of the art, the results in many areas are so good that systems are close to human level performance in laboratory benchmarks when testing some difficult language tasks such as machine translation or summarization [7,8]. In addition, recent work has shown that pre-trained language models can robustly perform NLP tasks in a few-shot or even in zero-shot fashion when given an adequate task description in its natural language prompt [5,9]. However, despite their impressive capabilities, these pre-trained language models raise severe concerns. First, training these models requires massive computing capabilities what means a huge energy consumption, and this issue is arguable from a sustainable viewpoint. Second, due to their opaque nature, there is a lack of clear understanding of how they work, how they were trained and validated, as well as what emergent properties they present. In addition, these language models have some limitations. In the case of GPT-3, the following limitations have been identified [10, 11]: output may lack semantic consistency, resulting in text that is gibberish and increasingly nonsensical as it grows longer; outputs embody all the biases that might be found in training data; and outputs may correspond to assertions that are not consonant with the truth. Moreover, the performance exhibited in NLP tasks with low-resources languages (e.g., Vasque, Galician or Gaelic) is much worse than in English, because such languages were underrepresented in the training data. It is worth noting that even if GPT-4 is announced to solve some of the previous limitations, its opaque nature remains. GPT-4 is a black-box model distributed in the form of software as a service and it is impossible to validate its performance from a rigorous scientific viewpoint. These shortcomings are common to other pre-trained language models which could also return inconsistent outputs in some cases. This is certainly a problem, especially when the models are embedded in the core of either question-answering systems or advice-giving systems or any dialog system in general where it is important that the resulting answer is true and correct. Possible inconsistency in the output of pre-trained language models could be explained by their lack of logical reasoning or very simple deductive and inductive abilities. For instance, such models have no clear understanding of temporal reasoning, and their memory sometimes falls short, so they may easily forget even the simplest temporal constraint that could have been introduced a few messages before. They also suffer from the so-called catastrophic forgetting when are fine-tuned for specific tasks [12]. Besides, an added difficulty is the continuous presence of imprecise expressions in the use of natural language by humans. We often use expressions which include vague or imprecise terms such as “more or less”, “a little bit more”, “not so much”, “approximately”, etc. These expressions capture the imprecision of language and summarize the information in a suitable way for human communication. Notice that automatic reasoning and, more particularly, temporal reasoning become more difficult to manage when dealing with this kind of expressions [13–16]. Considering all the above, the main contribution of this work is enriching interactive explanations with a model that can make reasoning with imprecise temporal information. With this aim, we propose a model that combines a Fuzzy Temporal Constraint Network with a knowledge graph, and we describe how it operates into a conversational agent carefully designed for providing users of intelligent systems with interactive explanations within a specific application domain. Namely, we develop an illustrative use case for the tourism sector to show how the agent handles some potential temporal inconsistency problems properly. Moreover, to conduct a thorough validation process of the model described, we have also carried out an empirical user study in which participants can evaluate some aspects of the conversations related to temporal reasoning. The rest of the manuscript is organized as follows. Section 2presents the background on conversational agents until the emergence of pre-trained language models. We also revisit related work with the focus on how to improve the reasoning capabilities of pretrained language models. In Section 3, we define the preliminary concepts needed to understand the rest of the work. In Section 4, we introduce the model proposed for fuzzy temporal reasoning. Section 5describes how such a model operates when it is embedded in a conversational agent. We develop a use case within the specific application domain under consideration, i.e., a virtual assistant for tourists. In Section 6, we report an empirical study with human evaluators to examine the possible effects of the lack of a fuzzy temporal reasoning model in a conversation with a conversational agent. Finally, Section 7summarizes the main conclusions and points out future work. 2. Related work The first implementation of a conversational agent was achieved in 1966 with the development of ELIZA [17], which relied heavily on linguistic rules and pattern matching techniques. Although its scope of knowledge was very limited, ELIZA was a landmark
International Journal of Approximate Reasoning 171 (2024) 109128 3 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. system that stimulated further research in the field. Following ELIZA, PARRY and Racter were introduced in 1972 and 1988 [18], respectively. Later, in 1995, the Artificial Linguistic Internet Computer Entity (ALICE) was introduced in [19]. Its heuristic matching patterns based on the Artificial Intelligence Mark-up Language proved a substantial upgrade with respect to previous conversational agents. Here, each pattern (i.e., each user input) is associated with an output (referred to as a template). The main limitation of rules and pattern matching-based conversational agents is that they are domain dependent. Accordingly, they lack flexibility because they rely on manually written rules for specific domains. With the advances in Artificial Intelligence techniques and NLP tools combined with the availability of computational power, new frameworks and algorithms were created to implement more advanced conversational agents, also called chatbots. In this context, two other chatbots are Jabberwacky and its successor Cleverbot,1released in 1997 and 2006, respectively. They belong to a category that can be referred to as information retrieval-based conversational agents. They generate their responses by selecting a suitable sentence from a large dialogue corpus, i.e., a database of stored conversations. Therefore, their main advantage is that they ensure the quality of the responses, but creating the necessary knowledge can be costly and time consuming. While conversational agents in the two previous categories rely on existing utterances, the so-called generative model-based conversational agents generate their responses word by word using statistical models. When the objective is modelling probability distributions over language (e.g., in terms of which words are more likely to appear before, after or in between some other words), the underlying generative models are called language models. Beyond using only single response generation approaches, some researchers attempt to hybridize the retrieval and rule-based methods with generative-based methods [20]. Originally designed for Machine Translation, one of the most famous models for developing generative-based methods is the neural network model sequence to sequence [21], which is used at the heart of several state-of-the-art conversational agents (e.g., MILABOT [22], Google’s Meena [23]or Microsoft’s XiaoIce [24]). Furthermore, one of the most outstanding innovations in Deep Learning language models has been the introduction of Transformers [25]. The Transformer-based architectures are endowed with training parallelization and this fact permits training on larger datasets than was originally achievable. This led to the development of large-scale pre-trained models such as BERT [2] (Bidirectional Encoder Representations from Transformers), GPT [26](Generative Pre-trained Transformer) and their successors (ALBERT [4], RoBERTa [3], BART [7], GPT-3 [5]or GPT-4 [27]), which were trained with huge datasets, such as Common Crawl,2and may be fine-tuned for downstream task without training a new model from scratch. These pre-trained models have outperformed previous state-of-the-art results on almost all NLP tasks such as machine translation, summarization [7,8], named entity recognition [28]or sentiment analysis [29], some of which have been reported to be close to human-level performance. Despite their powerful capabilities, pre-trained models also present limitations. For instance, in accepting large Internet-based datasets as training data we risk perpetuating dominant viewpoints, increasing power imbalances and further reifying inequality, and introducing stereotypical and derogatory associations among gender, race, ethnicity or disability status that may come with them [30]. In addition, regarding environmental wellbeing, building such models from scratch is very demanding from a computational viewpoint and comes with a very large carbon footprint [31]. The most relevant shortcomings for us, on which we will focus in this paper, are related to the possibility of returning inconsistent or incorrect responses. This was criticized in the case of GPT-3 [10,11,32], but the limitations are common to all pre-trained models mentioned previously and can be explained by a lack of logical reasoning or very simple deductive and inductive abilities in current language models. For instance, some previous publications have shown that pre-trained models lack thorough temporal reasoning [33–35]. However, recognition and understanding of temporal expressions could be crucial for conversational agents, e.g., in scheduling reminders or meetings and booking events, what shapes our motivation behind conducting this work. It is worth noting that much work has already been done on the recognition, extraction, and normalization of temporal expressions [36–39]. Nevertheless, to the best of our knowledge, there has been little work proposing solutions to overcome the possible inconsistencies caused by a lack of temporal reasoning in the context of conversational agents. Existing academic work mostly focuses on improving the mathematical capabilities of pre-trained models [40,41], but not specifically on temporal reasoning, let alone when temporal expressions include imprecise terms. Thus, in this work we will introduce a new model based on a knowledge graph for semantic representation and a fuzzy temporal constraint network for reasoning, that enriches conversational agents with the ability to detect and avoid temporal inconsistencies. There are several approaches with the aim of integrating knowledge graphs into dialog systems for different purposes, for instance tracking user goal during a dialog [42]or incorporating external knowledge and capturing semantics of user inputs [43]. Furthermore, there have been some attempts on modelling temporal information in knowledge graphs with the introduction of Temporal Knowledge Graphs [44–48], which also incorporate the start and end timestamps of each entity. The main difference with our approach is that we intend to use a knowledge graph to represent temporal constraints that restrict the occurrence of events, but they do not necessarily have to have happened. In addition, as we will describe later, we will handle the imprecision present in many temporal expressions in natural language by mapping the knowledge graph onto a fuzzy temporal constraints network. Finally, we would like to conclude this section with the following clarification. Commercial chatbots (e.g., ChatGPT released by OpenAI, BARD from Google, or Bing from Microsoft) are not language models. Indeed, they include several engineering layers which are aimed at fixing undesired behavior of the underlying language models. Unfortunately, there is no public information about how their associated engineering architectures are developed and tested. 1https://www .cleverbot .com/. 2https://commoncrawl .org/.
International Journal of Approximate Reasoning 171 (2024) 109128 4 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. 3. Preliminaries Before presenting the new model in Section 4, it is necessary to define here some preliminary concepts related to fuzzy sets and possibility theory [49]. We consider that time is a discrete set 𝜏={𝑡0, 𝑡1, … , 𝑡𝑖, …}, where 𝑡0is the time origin and 𝑡𝑖represents a precise time instant, for every 𝑖 ∈ℕ. Without losing generality, we assume 𝑡0<𝑡 1<… <𝑡 𝑖<…, with 𝑡𝑖+1 −𝑡𝑖being constant for every 𝑖 ∈ℕ. We denote as the set of possible units of time. Definition 1. Following [50], we define a fuzzy time instant 𝑎(also called date in [14,15]) as a possibility distribution: 𝜋𝑎∶𝜏→[0,1], where, given a precise time instant 𝑡 ∈𝜏, 𝜋𝑎(𝑡) ∈[0, 1] represents the possibility of 𝑎being precisely 𝑡. For simplicity, even if other possibility distributions are allowed, here we assume that any fuzzy time instant 𝑎is given by a trapezoidal distribution, i.e., a function of the form: 𝜋𝑎(𝑡)= ⎧ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎩ 𝑡−𝛼 𝛽−𝛼if 𝛼≤𝑡<𝛽, 1if 𝛽≤𝑡<𝛾, 𝑡−𝛿 𝛾−𝛿if 𝛾≤𝑡<𝛿, 0,otherwise, with 𝛼, 𝛽, 𝛾, 𝛿∈𝜏such that 𝛼≤𝛽≤𝛾≤𝛿. We denote a trapezoidal distribution as the 4-tuple (𝛼, 𝛽, 𝛾, 𝛿). This function becomes a triangular distribution in case that 𝛽=𝛾. Definition 2. Following [14,50], we define a fuzzy time extent 𝑒as a possibility distribution: 𝜋𝑒∶→[0,1], where, given 𝑚 ∈, 𝜋𝑒(𝑚) ∈[0, 1] represents the possibility of 𝑒being precisely a duration of 𝑚units of time. Again, we will consider that any fuzzy time extent is given by a trapezoidal distribution. Let us suppose that 𝑎and 𝑏are two fuzzy time instants whose trapezoidal distributions are determined by the parameters 𝛼, 𝛽, 𝛾, 𝛿 and 𝛼′, 𝛽′, 𝛾′, 𝛿′, respectively. Then, following [15,50], we can represent the fuzzy temporal distance between them, 𝑑(𝑎, 𝑏), by means of the fuzzy time extent: 𝜋𝑑(𝑎,𝑏)(𝑚)∶= max 𝑚=𝑡−𝑠 𝑡,𝑠∈𝜏{min{𝜋𝑎(𝑠),𝜋𝑏(𝑡)}},∀𝑚∈.(1) It is important to remark that 𝜋𝑑(𝑎,𝑏)and 𝜋𝑑(𝑏,𝑎)are symmetric distributions with respect to the y-axis because the distance is directed. Note that a fuzzy time instant 𝑎can always be considered as the fuzzy time extent 𝑒 =𝑑(𝑡0, 𝑎). As an example, Fig. 1shows the graphical representation of two fuzzy time instants: 𝑎 =“around 12 PM” (Fig. 1(a)) and 𝑏 =“a little bit before 12:30 PM” (Fig. 1(b)), given by the trapezoidal distribution 𝜋𝑎= (11.45, 11.55, 12.05, 12.15) and the triangular distribution 𝜋𝑏= (12.10, 12.25, 12.25, 12.30), respectively. We are considering that the time axis 𝜏is formed by the precise time instants from 11.30 AM to 12.30 PM and the unit of time is 5minutes. In Fig. 2, we present the graphical representation of two examples of fuzzy time extents: 𝑒 =“about 35 minutes” (Fig. 2(a)), represented by the trapezoidal distribution 𝜋𝑒= (20, 30, 40, 50), and 𝑓=𝑑(𝑎, 𝑏), with a different shape obtained as the fuzzy temporal distance between fuzzy time instants 𝑎and 𝑏given by the expression (1)(as defined in Fig. 2(b)), with an associated semantics of “a little bit less than approximately 30 min”. Following [50], we now introduce some basic binary operations on fuzzy time extents which will be used later on. Definition 3. We define the intersection of two time extents 𝑒and 𝑒′, noted as 𝑒∩𝑒′, as the fuzzy time extent given by the possibility distribution: (𝜋𝑒∩𝜋𝑒′)(𝑚)∶=min{𝜋𝑒(𝑚),𝜋𝑒′(𝑚)},∀𝑚∈. Definition 4. We define the composition of two time extents 𝑒and 𝑒′, noted as 𝑒◦𝑒′, as the fuzzy time extent given by the possibility distribution: (𝜋𝑒◦𝜋𝑒′)(𝑚)∶= max 𝑚=𝑚1+𝑚2 𝑚1,𝑚2∈{min{𝜋𝑒(𝑚1),𝜋𝑒′(𝑚2)}},∀𝑚∈.
International Journal of Approximate Reasoning 171 (2024) 109128 5 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. Fig. 1. Examples of possibility distributions of fuzzy time instants. Note that the composition operator is nothing but the addition operation of two fuzzy quantities using the extension principle described in [51]. 4. A model for fuzzy temporal reasoning In this section, we present a model for managing fuzzy temporal expressions and detecting inconsistencies between them. Starting from a knowledge graph representative of the specific domain, the model projects the temporal information onto a fuzzy temporal constraint network, which provides reasoning mechanisms for checking consistency and answering explanatory queries. More precisely, we assume a knowledge graph 𝐾𝐺 is known and given beforehand. 𝐾𝐺 employs a graph-based data model to represent knowledge about a specific application domain, providing a concise and intuitive abstraction of data involved [52,53]. It is worth noting that there is no single commonly accepted formal definition and representation of knowledge graphs [54,55]. Accordingly, we simply define 𝐾𝐺 as a directed labeled graph whose nodes are entities of interest in the application domain. These entities represent information which may be temporal or not. If {𝑛1, … , 𝑛𝑚}represents the set of 𝑚nodes in 𝐾𝐺, for each triplet (𝑛𝑖, 𝑟, 𝑛𝑗), 𝑛𝑖is taken as the parent node, which has the relation 𝑟with the child node 𝑛𝑗. We distinguish between entities that include temporal information, which we will denote temporal entities, and those that do not. Thus, we consider precise time instants, fuzzy time instants or fuzzy time extents. Moreover, we consider both temporal relations (e.g., “starts at” or “has a duration of”) and non-temporal relations (e.g., “is a subclass of” or “is located in”). In the case of dealing with temporal information, the relation 𝑟is represented as an arc linking the two given nodes in a Fuzzy Temporal Constraint Network (see formal definitions and further details in Section 4.1). The procedure for mapping 𝐾𝐺 onto is described in Section 4.2. 4.1. Fuzzy temporal constraint networks First, we focus on defining a formalism for knowledge representation of, and reasoning with, temporal information which is not precise but vague or/and uncertain. Let us assume we have a set of 𝑛instantaneous events whose precise time instants of occurrence, 𝑣1, … , 𝑣𝑛∈𝜏, are unknown. If we add some imprecise information about one of the time instants, 𝑣𝑖, we are establishing a constraint over the possible values for 𝑣𝑖, called unary constraint and denoted as 𝐶𝑖. This constraint can be represented as a fuzzy time instant by means of a possibility distribution 𝜋𝐶𝑖defined over 𝜏. If this imprecise information establishes a duration constraint over the possible values of the temporal distance between two precise time instants 𝑣𝑖and 𝑣𝑗, then we call it binary constraint and we denote it as 𝐶𝑖𝑗 . In this case, the
International Journal of Approximate Reasoning 171 (2024) 109128 6 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. Fig. 2. Examples of possibility distributions of fuzzy time extents. constraint can be seen as a fuzzy time extent given by a possibility distribution 𝜋𝐶𝑖𝑗 over . Introducing a precise origin of times, 𝑣0=𝑡0, we can represent every unary constraint 𝐶𝑖as the binary constraint 𝐶0𝑖. For simplicity, we will use the notation 𝜋𝑖𝑗 instead of 𝜋𝐶𝑖𝑗 , ∀𝑖, 𝑗∈{0, … , 𝑛}. Definition 5. Following [14,16,50], we define a Fuzzy Temporal Constraint Network (FTCN) as a pair (, ), where ={𝑣0, … , 𝑣𝑛} is a finite set of nodes, with each node representing an unknown precise time instant 𝑣𝑖∈𝜏, ∀𝑖 ∈{1, … , 𝑛}and 𝑣0=𝑡0the time origin, and a finite set of directed arcs ={𝐶𝑖𝑗| 𝑖, 𝑗∈{0, … , 𝑛}} that represent binary constraints between the nodes. Each constraint 𝐶𝑖𝑗 is a fuzzy time extent that has associated a possibility distribution 𝜋𝑖𝑗 . Definition 6. Given 𝛼∈[0, 1], a solution [16]of an FTCN at least to the degree 𝛼is a 𝑛-tuple 𝑆={𝑡1, … , 𝑡𝑛}such that (𝑆) ≥𝛼, where 𝑡𝑖∈𝜏, ∀𝑖 ∈{0, … , 𝑛}and (𝑆)∶= min 𝑖,𝑗∈{0,…,𝑛}{𝜋𝑖𝑗(𝑡𝑗−𝑡𝑖)}. Definition 7. Given 𝛼∈[0, 1], a precise time instant 𝑡 ∈𝜏is a feasible value [16]at least to the degree 𝛼for a node 𝑣𝑖of an FTCN if there exists a solution 𝑆such that 𝑣𝑖=𝑡and (𝑆) ≥𝛼. Definition 8. The degree of consistency [16]of an FTCN is defined by the following expression: 𝐶𝑜𝑛𝑠()∶=max 𝑆∈𝜏𝑛(𝑆). If 𝐶𝑜𝑛𝑠() =𝛼, then we say that is 𝛼-consistent. is consistent if it is 1-consistent and is inconsistent if there is no solution (𝛼=0).
International Journal of Approximate Reasoning 171 (2024) 109128 7 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. Calculating the feasible values to a certain degree for a node and the consistency of the entire network using the expressions defined above implies first finding all possible solutions (𝑆)at any degree. This process means combining all the values of the axis of precise time instants considered, 𝜏, in the particular case that it is finite. With the aim of making the process more efficient, the computation of queries on FTCNs is based on transforming the network into an equivalent representation in which the query answers can be more directly obtained [16,50]. Two networks defined over the same set of nodes are equivalent if they represent the same solution set with the same degree. We will say that an FTCN =(,)is tighter than another one ′=(,′)with the same set of nodes if for all 𝐶′ 𝑖𝑗 ∈𝐶′there exists a constraint 𝐶𝑖𝑗 ∈𝐶such that their associated possibility distributions satisfy that 𝜋𝑖𝑗 (𝑚) ≤𝜋′ 𝑖𝑗(𝑚), ∀𝑚 ∈. We will look for the tightest equivalent network, called the minimal network. Between each pair of nodes in , 𝑣𝑖and 𝑣𝑗, there is a direct constraint 𝐶𝑖𝑗 . There may be additional induced constraints, each corresponding to a possible path connecting the two nodes. The induced constraint from the node 𝑣𝑖to 𝑣𝑗in a path of length 𝑘, 1 ≤𝑘 ≤𝑛, is given by the composition of the direct constraints between each pair of consecutive nodes belonging to the path. We will use the notation 𝐶𝑘 𝑖𝑗 to represent the intersection of the induced constraints corresponding to all the paths of length 𝑘from 𝑣𝑖to 𝑣𝑗, given by the following possibility distribution: 𝜋𝑘 𝑖𝑗 ∶= ⋂ 𝑖0,𝑖1,…,𝑖𝑘∈𝑉𝑘 𝑛 𝑖0=𝑖 𝑖𝑘=𝑗 𝜋𝑖0𝑖1 ◦𝜋𝑖1𝑖2 ◦⋯◦𝜋𝑖𝑘−1𝑖𝑘,(2) where 𝑉𝑘 𝑛represents the set of 𝑘-element variations of 𝑛elements. Finally, we will denote as 𝐶min 𝑖𝑗 the intersection of all the induced constraints for paths from 𝑣𝑖to 𝑣𝑗of any length 𝑘, given by the following possibility distribution: 𝜋min 𝑖𝑗 ∶= 𝑛 ⋂ 𝑘=1 𝜋𝑘 𝑖𝑗 (3) Note that the expressions above can only be calculated if there exists a single constraint between each pair of nodes in the network. Hence, it is necessary to take into account the following aspects: •The absence of constraints between two nodes 𝑣𝑖and 𝑣𝑗is equivalent to consider a constraint 𝐶𝑖𝑗 whose associated function is the constant function 𝜋𝑈(𝑚) ∶= 1, ∀𝑚 ∈. Therefore, there always exists a constraint between each pair of nodes in the FTCN. •If 𝜋𝑖𝑗 stands for the possibility distribution associated to the constraint 𝐶𝑖𝑗 for two nodes 𝑣𝑖and 𝑣𝑗, we will also consider that there exists a constraint 𝐶𝑗𝑖 between 𝑣𝑗and 𝑣𝑖such that 𝜋𝑗𝑖(𝑚) =𝜋𝑖𝑗(−𝑚), ∀𝑚 ∈. •For each node 𝑣𝑖with itself, we consider a constraint 𝐶𝑖𝑖 such that 𝜋𝑖𝑖(𝑚) =1if 𝑚 =0and 𝜋𝑖𝑖(𝑚) =0otherwise. •If there exists more than one constraint between a pair of nodes in the network, we reduce them into an equivalent single constraint applying the intersection of time extents. The FTCN min =(, min), where min ={𝐶min 𝑖𝑗 |0 ≤𝑖, 𝑗≤𝑛}and each constraint 𝐶min 𝑖𝑗 is associated with 𝜋min 𝑖𝑗 , ∀𝑖, 𝑗∈{0, … , 𝑛}, is the minimal network corresponding to the FTCN . Such minimal network contains 𝑛(𝑛 −1)∕2edges, with 𝑛being the number of nodes. According to [50], the calculation of min by means of the aforementioned expressions is exponential with the number of nodes. Nevertheless, some authors have applied path-consistency algorithms, which are executed with a polynomial increase of time, to different types of precise temporal networks [13]. The same idea can be applied to the minimization of an FTCN [16,50]. We have adapted Floyd-Warshall’s all-pairs-shortest-paths algorithm [56]to compute the minimal network associated to an FTCN in a polynomial time (see the pseudocode in Algorithm 1). Algorithm 1 Pseudocode for computing a minimal network. 1: function MINIMALNETWORK(𝑛, =(, )) 2: for 𝑘 =1, … , 𝑛do 3: for 𝑖 =1, … , 𝑛do 4: for 𝑗=1, … , 𝑛do 5: 𝜋𝑖𝑗 ←𝜋𝑖𝑗 ∩(𝜋𝑖𝑘◦𝜋𝑘𝑗 ) 6: if 𝜋𝑖𝑗 =0then 7: Exit (The FTCN is inconsistent) 8: return min The minimal network allows us to determinate in an easy way the consistency of the network and feasible values for each node. However, edges in this network lack semantic interpretability because they are the result of intersection of the constraints that were initially mapped from the knowledge graph. From now on, we consider that =(, )represents a minimal network. The construction of the minimal network ensures that there exists a solution 𝑆max ={𝑡1, … , 𝑡𝑛}, such that 𝑡𝑖∈𝜏, ∀𝑖 ∈{1, … , 𝑛}and: 𝜋𝑖𝑗(𝑡𝑗−𝑡𝑖)=max 𝑚∈{𝜋𝑖𝑗(𝑚)},∀𝑖,𝑗 ∈{0,…,𝑛}.
International Journal of Approximate Reasoning 171 (2024) 109128 8 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. Hence, following the Definition 8, the degree of consistency can be directly obtained as follows: 𝐶𝑜𝑛𝑠()= min 𝑖,𝑗∈{0,…,𝑛}{max 𝑚∈{𝜋𝑖𝑗(𝑚)}}. Similarly, since we manage the tightest constraint for each pair of nodes, computing the feasibility of a value 𝑡for a node 𝑣𝑖is reduced to calculate the value 𝜋0𝑖(𝑡 −𝑡0). In addition, the minimal network simplifies the introduction of two new concepts: the degree of compatibility of a new constraint with the network and the feasible values for a distance between two nodes. Definition 9. Let 𝐶∗ 𝑖𝑗 be a fuzzy temporal constraint given by a possibility distribution 𝜋∗ 𝑖𝑗 . The degree of compatibility [16]of 𝐶∗ 𝑖𝑗 with is defined by the following expression: 𝐶𝑜𝑚𝑝(𝐶∗ 𝑖𝑗,)∶=max 𝑚∈{min{𝜋∗ 𝑖𝑗(𝑚),𝜋𝑖𝑗 (𝑚)}}. Definition 10. Given 𝛼∈[0, 1], a value of 𝑚 ∈is a feasible value for the distance between two nodes 𝑣𝑖and 𝑣𝑗in an FTCN at least to the degree 𝛼if it is satisfied that: 𝜋𝑖𝑗(𝑚)≥𝛼. To sum up, once we have built the minimal network for an FTCN, we can obtain the following information: (i) consistency of the network; (ii) compatibility of a new constraint with the network; and (iii) feasible values both for a node and for any distance between two nodes. 4.2. Mapping a knowledge graph onto an FTCN An FTCN provides us with a formalism for representing of, and reasoning with, fuzzy temporal information. However, in the context of conversational agents, information is not usually provided in the form of precise time instants and constraints on their temporal distance. Furthermore, some additional constraints may have to be added to the network if we consider certain semantic relations between entities in the domain. Therefore, we propose a knowledge graph (KG) as an initial point to represent in a more natural way the information of a specific application domain. A lot of work has been done on entity extraction from KGs [57], refining the information gathered [53]and completion [58]. Assuming we start from a KG that includes the information of interest from the domain consistently related, the question now is how to project it onto an FTCN in order to allow queries such as those specified in section 4.1. First of all, it is important to point out that not all the entities and relations embedded in the KG need to be transferred to the FTCN. It would be sufficient to map only those parts of the graph that are related to temporal information, thus facilitating the implementation of the mapping procedure (see pseudocode in Algorithm 2). For that reason, we will pay major attention to the procedure to map those nodes in the KG which are connected by a temporal relation. Let us suppose that the KG contains the triplet (𝑛𝑖, 𝑟, 𝑛𝑗), 𝑟being a temporal relation. Then, we consider the following situations: •𝑛𝑖is a non-temporal entity and 𝑛𝑗is a temporal entity. We will suppose that any non-temporal entity has a start time instant and an end time instant. In the absence of information on their duration, we will also assume that the two instants have the same value (possibly unknown). Thus, the existence of this triplet in KG implies the creation of two nodes associated with 𝑛𝑖in the FTCN, one for the precise time instant of start and another one for the precise time instant of end (provided they do not already exist). In addition, a constraint is added on the time elapsed between them, which in the absence of information is assumed to be zero. If the child node is temporal, due to the scope of this work we consider three possible relations between both nodes: –One option is that 𝑛𝑖“starts at” 𝑛𝑗. On the one hand, if 𝑛𝑗is a precise time instant, the node associated with the precise time instant at which the entity 𝑛𝑖starts is set to the value of the precise time instant represented by 𝑛𝑗. As an example, we can figure out that in the application domain we collected the information about the boat trip to go to Cíes Island and we had the triplet (outbound trip, starts at, 11.15 AM). On the other hand, if 𝑛𝑗is a fuzzy time instant, this relation defines a constraint over the temporal distance between the origin of times (the node 𝑣0in the FTCN) and the node in the FTCN associated with the precise time instant of start of entity 𝑛𝑖. This would be the case of a triplet such as (outbound trip, starts at, around 12 PM). –The second option is that 𝑛𝑖“finishes at” 𝑛𝑗and the procedure is analogous to that described above, replacing the node representing the precise time instant at which the entity 𝑛𝑖starts by the node representing its precise time instant of end. –The other alternative is that 𝑛𝑖“has a duration of” 𝑛𝑗, with 𝑛𝑗being a fuzzy time extent, as in the example (outbound trip, has a duration of, about 35 minutes). In this case, 𝑛𝑗is projected into the FTCN by means of a constraint between the nodes associated with its precise time instant of start and its precise time instant of end. This constraint expresses the semantics of the temporal expression described in 𝑛𝑗.
International Journal of Approximate Reasoning 171 (2024) 109128 9 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. Algorithm 2 Pseudocode for mapping a KG onto an FTCN. 1: function GRAPHTONETWORKMAPPING(KG, =(, )) 2: for (𝑛𝑖, r, 𝑛𝑗) in KG do 3: if {𝑣𝑖𝑠𝑡𝑎𝑟𝑡 , 𝑣𝑖𝑒𝑛𝑑 } ⊈then 4: ←∪{𝑣𝑖𝑠𝑡𝑎𝑟𝑡 , 𝑣𝑖𝑒𝑛𝑑 } 5: 𝐶𝑖𝑠𝑡𝑎𝑟𝑡𝑖𝑒𝑛𝑑 ←𝜋𝑈 6: addConstraint(𝐶𝑖𝑠𝑡𝑎𝑟𝑡 𝑖𝑒𝑛𝑑 , ) 7: if isTemporalNode(𝑛𝑗)then 8: if 𝑟 =“has a duration of”then 9: 𝐶𝑖𝑠𝑡𝑎𝑟𝑡𝑖𝑒𝑛𝑑 ←createNodeConstraint(𝑛𝑗) 10: addConstraint(𝐶𝑖𝑠𝑡𝑎𝑟𝑡 𝑖𝑒𝑛𝑑 , ) 11: if 𝑟 =“starts at”then 12: 𝐶0𝑖𝑠𝑡𝑎𝑟𝑡 ←createNodeConstraint(𝑛𝑗) 13: addConstraint(𝐶0𝑖𝑠𝑡𝑎𝑟𝑡 , ) 14: if 𝑟 =“finishes at”then 15: 𝐶0𝑖𝑒𝑛𝑑 ←createNodeConstraint(𝑛𝑗) 16: addConstraint(𝐶0𝑖𝑒𝑛𝑑 , ) 17: else 18: if {𝑣𝑗𝑠𝑡𝑎𝑟𝑡 , 𝑣𝑗𝑒𝑛𝑑 } ⊈then 19: ←∪{𝑣𝑗𝑠𝑡𝑎𝑟𝑡 , 𝑣𝑗𝑒𝑛𝑑 } 20: 𝐶𝑗𝑠𝑡𝑎𝑟𝑡𝑗𝑒𝑛𝑑 ←𝜋𝑈 21: addConstraint(𝐶𝑗𝑠𝑡𝑎𝑟𝑡 𝑗𝑒𝑛𝑑 , ) 22: if 𝑟 =“happens <modifier><m units of time>before”then 23: 𝐶𝑖𝑒𝑛𝑑 𝑗𝑠𝑡𝑎𝑟𝑡 ←createRelationConstraint(𝑟) 24: addConstraint(𝐶𝑖𝑒𝑛𝑑 𝑗𝑠𝑡𝑎𝑟𝑡 , ) 25: if 𝑟 =“happens during”then 26: 𝐶𝑗𝑠𝑡𝑎𝑟𝑡𝑖𝑠𝑡𝑎𝑟𝑡 ←createRelationConstraint(“happens before”) 27: 𝐶𝑖𝑒𝑛𝑑 𝑗𝑒𝑛𝑑 ←createRelationConstraint(“happens before”) 28: addConstraint(𝐶𝑗𝑠𝑡𝑎𝑟𝑡 𝑖𝑠𝑡𝑎𝑟𝑡 , ) 29: addConstraint(𝐶𝑖𝑒𝑛𝑑 𝑗𝑒𝑛𝑑 , ) 30: return •𝑛𝑖and 𝑛𝑗are non-temporal entities. Analogously to the previous case, two nodes and a constraint between them associated to each non-temporary entity are created in the FTCN. Two possible relations are considered: –The first one is the relation “happens <modifier><m units of time>before”, where 𝑚 ∈and the modifier could be an expression (e.g., “approximately”, “more or less”, “a little more than”, etc.) which adds information about the temporal distance between the nodes. Notice that, if 𝑛𝑖happens after 𝑛𝑗instead of before, then the relation can be converted into the above expression just by inverting the parent and the child nodes. This relation between the nodes adds a constraint to the FTCN between the precise time instant of end of 𝑛𝑖and the precise time instant of start of 𝑛𝑗by means of a fuzzy time extent that represents the semantics of the temporal expression described in 𝑟. If no modifier is specified or if the modifier is “exactly” (or a synonym), the associated fuzzy time extent will consist of a possibility distribution taking the value 1at 𝑚 and 0otherwise. In any other case, the possibility distribution associated with the constraint will depend on the definition of the modifier, which could be different according to the application domain. Following on the previous example, this case would represent a triplet such as (outbound trip, happens before, return trip). –The other possibility we consider is that 𝑛𝑖“happens during” 𝑛𝑗. In this case, the relation adds two new constraints to the FTCN: one constraint to express that the precise time instant of start of 𝑛𝑗happens before (or at the same time as) the precise time instant of start of 𝑛𝑖and another one to express that the precise time instant of end of 𝑛𝑖happens before (or at the same time as) the precise time instant of end of 𝑛𝑗. One example would be a triplet such as (bar service, happens during, outbound trip). •Other cases. The case when 𝑛𝑖is a temporal entity and 𝑛𝑗is a non-temporal entity as well as the case when 𝑛𝑖and 𝑛𝑗are both temporal entities do not make sense in the context of this work, so we will not consider them. In the case that it is necessary to add a new constraint between two nodes in the network, 𝑣𝑖and 𝑣𝑗, when one already exists, the constraint 𝐶𝑖𝑗 is modified so that it represents the intersection between the existing one and the new one. Initially, it might seem that projecting only the nodes directly connected through temporal relations is sufficient to transfer all the temporal information to the FTCN. However, there could be non-temporal relations in the graph that caused the inheritance of certain relations, including temporal ones. An illustrative example is the relation “is a subclass of”, for which it is obvious that the node representing the subclass must inherit the relations of the superclass node. Accordingly, before projecting those nodes of interest to the FTCN, it is required to transfer all temporal relations inherited in this way, if necessary. Following an approach based on going through all the nodes by exploring and relocating their relations has exponential complexity on the number of nodes. Therefore, at this point it is advisable to apply some existing graph search algorithm to reduce computational complexity. Moreover, evaluating whether any semantic relation has the property of inheriting temporal
International Journal of Approximate Reasoning 171 (2024) 109128 16 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. Fig. 7. Effects of the type of conversational agent and the FTCN complexity on consistency perceived in conversations. The error bars represent the 95% confidence interval for the mean. (For interpretation of the colors in the figure(s), the reader is referred to the web version of this article.) With the aim of assessing the goodness and significance of reported results, we carried out an independent-means 𝑡-test for subexperiment 1.1 and experiment 2. In addition, in case of the sub-experiment 1.2, we applied a Welch’s 𝑡-test, an alternative to the independent-means 𝑡-test when there is a violation in the assumption of equality of variances. On average, the consistency perceived in conversations with the TimeVersa agent was significantly higher than consistency perceived in conversations with the OpenAI agent, no matter if the resultant FTCN after the conversation was simple (𝑡(128) = 10.43, 𝑝 <0.01) or complex (𝑡(127.701) =3.44, 𝑝 <0.01). Furthermore, on average, the consistency perceived in conversations with TimeVersa incorporating the fuzzy temporal reasoning model was significantly higher than consistency perceived in conversations with TimeVersa incorporating the crisp version of the model (𝑡(128) =2.64, 𝑝 <0.01). 6.3. Discussion The results of the study validate the two research hypotheses formulated above and reproduced here: •𝐻1: “In case of involving vague temporal information, the interaction with TimeVersa is perceived as more consistent than the interaction with the OpenAI API.” •𝐻2: “In case of involving vague temporal information, the interaction with TimeVersa incorporating the temporal reasoning model is perceived as more consistent than the interaction with TimeVersa incorporating the crisp version of the temporal reasoning model.” On the one hand, we found that users’ perception of consistency is significantly higher in conversations with the TimeVersa agent than in conversations with the OpenAI agent. Hence, the fuzzy temporal reasoning model increases the consistency of the conversation, when handling vague temporal information in the context of application. Although the differences are significant for both levels of the FTCN complexity, it is worth noting that in the case of the complex FTCN the average consistency, for both agents, is much closer than for the simple FTCN. In addition, the average consistency for conversations with the OpenAI agent and the complex FTCN exceeds the middle value of the scale considered, even though the agent returns an incorrect answer. It is likely that the complexity of the FTCN has sometimes caused the participant not to be able to reason properly with the temporal constraints involved in the conversation. As a result, this situation can lead to total blind reliability in the agent, even if the agent’s response is incorrect. On the other hand, we found a significantly positive effect of using a fuzzy model compared to using a crisp model on the perception of consistency of a conversation with an agent. However, although the observed differences in the average consistency for both models are statistically significant, they are small in absolute terms. Moreover, both values are surprisingly low, so it seems that the participants do not consider either conversation to be consistent enough. Notice that the highest value for the average consistency in experiment 2, which is obtained for the case that involves the TimeVersa agent with a fuzzy temporal reasoning model, barely exceeds 3, corresponding to “somewhat consistent” on the scale considered. However, in experiment 1 the average consistency reaches a value of 4.35 with TimeVersa and the simple FTCN and a value of 3.89 with TimeVersa and the complex FTCN. 7. Conclusions and future work Large language models achieve great results in many NLP tasks such as machine translation, summarization, or sentiment analysis. Nevertheless, they have some limitations, including the possibility of returning inconsistent or factually incorrect responses, that make them less appropriate for other applications like conversational agents, since they may be misleading for users. Part of these inconsistencies arise from the lack of temporal reasoning capabilities of these models, which is also a key weakness, since
International Journal of Approximate Reasoning 171 (2024) 109128 17 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. understanding and reasoning with temporal expressions is very relevant in the context of a conversational agent. In addition, temporal reasoning becomes more difficult to manage for these models when dealing with vague or imprecise information, which play a fundamental role for effective communication. In this paper we have presented a model for fuzzy temporal reasoning that overcomes some inconsistency issues detected in existing pre-trained language models, in a specific application scenario of a conversational agent. Thanks to the use of an FTCN, this model supports the representation of imprecise temporal knowledge and provides designers with mechanisms to combine the information represented as well as detecting inconsistencies. The temporal model handles temporal constraints on both the time of occurrence of an event (e.g., “around 12 PM”) and the time interval between two events (e.g., “about 3 hours”), which are represented by means of possibility distributions. However, in the context of a conversational agent, we found it difficult to directly transform temporal information from a conversation in natural language into an FTCN. Thus, to provide a more comprehensive and intuitive representation of the application domain, we proposed to start from a knowledge graph, whose nodes represent entities of interest and whose edges can represent any semantic relation between these entities. This knowledge graph must be provided beforehand and validated by an expert. Then, we explained how to map the temporal information stored in a representative knowledge graph of the application domain onto an FTCN that is afterwards in charge of checking consistency and answering queries. In addition, we have integrated our model into the TimeVersa conversational agent, which represents a proof of concept and operates in a practical use case related to tourism where the agent must deal with temporal information under uncertainty. As a result, TimeVersa can detect temporal inconsistencies and return correct answers for the given queries using the model proposed. In contrast, a similar conversation using the OpenAI API includes incorrect responses in relation to temporal consistency. Indeed, an empirical user study demonstrates that the observed differences in users’ perception of consistency between a conversation with the TimeVersa agent and a similar conversation with the OpenAI API are statistically significant. Therefore, an encouraging result and relevant conclusion of this work is that the model proposed and the prototype implemented entail a starting point to endow conversational agents with fuzzy temporal reasoning capabilities. Despite this relevant concluding remark, we have identified some limitations in our model, which leave room for further research. As we mentioned in Section 4.2, a relevant step in the procedure for mapping the knowledge graph onto an FTCN involves an exponential computational complexity. It remains as an open problem exploring strengths and weaknesses of approximate search methods in graphs, because integrating an appropriate search method in our model, we may reduce the computational cost of the mapping procedure. For the sake of generalization, it would be desirable to implement a method that built the knowledge graph with the information of interest automatically, instead of imposing the need of an expert to produce such a graph manually beforehand. Similarly, the projection onto the FTCN should be made automatic too. Namely, one of the hardest steps to make the process fully automatic is the semantic interpretation of possibility distributions which define fuzzy constraints associated to temporal information. Moreover, a future user study should include a greater range of conversations with different use cases, to determine whether the present results generalize well to other situations. Finally, we envision that our conversational agent will be automatically connected to a pre-trained language model, that may be available online as open source and could run in backend, and then the discourse history of every conversation is automatically monitored by a model for fuzzy temporal reasoning like the one described in this paper. Declaration of competing interest The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Mariña Canabal-Juanatey reports financial support was provided by Government of Galicia Department of Culture Education and Universities. Alberto Bugarín-Diz reports financial support was provided by Government of Galicia Department of Culture Education and Universities. Alberto Bugarín-Diz reports was provided by Spain Ministry of Science and Innovation. Jose M. Alonso-Moral reports financial support was provided by Spain Ministry of Science and Innovation. Data availability Data will be made available on request. Acknowledgement Mariña Canabal-Juanatey is a PhD Researcher supported by the Galician Ministry of Culture, Education, Professional Training and University (ED481A 2022/212). All authors recognize the support of the Galician Ministry of Culture, Education, Professional Training and University (grants ED431G2019/04 and ED431C2022/19). This work is also supported by the Spanish Ministry of Science and Innovation (MCIN/AEI/10.13039/501100011033/) with grants PID2021-123152OB-C21, PID2020-112623GB-I00, and TED2021-130295B-C33. All previous grants are co-funded by the European Regional Development Fund (ERDF/FEDER program). References [1] X. Han, et al., Pre-trained models: past, present and future, AI Open 2 (2021) 225–250, https://doi .org /10 .1016 /j .aiopen .2021 .08 .002.
International Journal of Approximate Reasoning 171 (2024) 109128 18 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. [2] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT: pre-training of deep bidirectional transformers for language understanding, in: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), ACL, Minneapolis, Minnesota, 2019, pp. 4171–4186. [3] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov, RoBERTa: a robustly optimized BERT pretraining approach, arXiv preprint, arXiv :1907 .11692, 2019, https://doi .org /10 .48550 /arXiv .1907 .11692. [4] Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, R. Soricut, ALBERT: a lite BERT for self-supervised learning of language representations, arXiv preprint, arXiv :1909 .11942, 2019, https://doi .org /10 .48550 /arXiv .1909 .11942. [5] T. Brown, et al., Language models are few-shot learners, Adv. Neural Inf. Process. Syst. 33 (2020) 1877–1901, https://proceedings .neurips .cc /paper /2020 /file / 1457c0d6bfcb4967418bfb8ac142f64a -Paper .pdf. [6] J. Simon, Large language models: a new Moore’s law?, https://huggingface .co /blog /large -language -models, october 2021. (Accessed 5 September 2022). [7] M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, L. Zettlemoyer, BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, ACL, 2020, pp. 7871–7880. [8] S. Edunov, A. Baevski, M. Auli, Pre-trained language model representations for language generation, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), ACL, Minneapolis, Minnesota, 2019, pp. 4052–4059. [9] N. Ding, S. Hu, W. Zhao, Y. Chen, Z. Liu, H. Zheng, M. Sun, OpenPrompt: an open-source framework for prompt-learning, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, ACL, Dublin, Ireland, 2022, pp. 105–113. [10] R. Dale, GPT-3: what’s it good for?, Nat. Lang. Eng. 27 (1) (2021) 113–118, https://doi .org /10 .1017 /S1351324920000601. [11] L. Floridi, M. Chiriatti, GPT-3: its nature, scope, limits, and consequences, Minds Mach. 30 (2020) 1–14, https://doi .org /10 .1007 /s11023 -020 -09548 -1. [12] Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y.J. Lee, Y. Ma, Investigating the catastrophic forgetting in multimodal large language models, arXiv :2309 .10313, 2023. [13] R. Dechter, I. Meiri, J. Pearl, Temporal constraint networks, Artif. Intell. 49 (1–3) (1991) 61–95, https://doi .org /10 .1016 /0004 -3702(91 )90006 -6. [14] S. Barro, R. Marín, J. Mira, A.R. Patón, A model and a language for the fuzzy representation and handling of time, Fuzzy Sets Syst. 61 (2) (1994) 153–175, https://doi .org /10 .1016 /0165 -0114(94 )90231 -3. [15] D. Dubois, H. Prade, Processing fuzzy temporal knowledge, IEEE Trans. Syst. Man Cybern. 19 (4) (1989) 729–744, https://doi .org /10 .1109 /21 .35337. [16] L. Vila, L. Godo, On fuzzy temporal constraint networks, Mathw. Soft Comput. 1(3) (1994) 315–334. [17] J. Weizenbaum, ELIZA—a computer program for the study of natural language communication between man and machine, Commun. ACM 9(1) (1966) 36–45, https://doi .org /10 .1145 /365153 .365168. [18] K.K. Nirala, N.K. Singh, V.S. Purani, A survey on providing customer and public administration based services using AI: chatbot, Multimed. Tools Appl. 81 (16) (2022) 22215–22246, https://doi .org /10 .1007 /s11042 -021 -11458 -y. [19] R.S. Wallace, The anatomy of A.L.I.C.E., in: Parsing the Turing Test: Philosophical and Methodological Issues in the Quest for the Thinking Computer, Springer Netherlands, Dordrecht, 2009, pp. 181–210. [20] L. Yang, J. Hu, M. Qiu, C. Qu, J. Gao, W.B. Croft, X. Liu, Y. Shen, J. Liu, A hybrid retrieval-generation neural conversation model, in: Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Association for Computing Machinery, 2019, pp. 1341–1350. [21] I. Sutskever, O. Vinyals, Q.V. Le, Sequence to sequence learning with neural networks, in: Proceedings of the 27th International Conference on Neural Information Processing Systems, NIPS’14, vol. 2, MIT Press, Cambridge, MA, USA, 2014, pp. 3104–3112, https://proceedings .neurips .cc /paper /2014 /file / a14ac55a4f27472c5d894ec1c3c743d2 -Paper .pdf. [22] I.V. Serban, et al., A deep reinforcement learning chatbot, arXiv preprint, arXiv :1709 .02349, 2017, https://doi .org /10 .48550 /arXiv .1709 .02349. [23] D. Adiwardana, et al., Towards a human-like open-domain chatbot, arXiv preprint, arXiv :2001 .09977, 2020, https://doi .org /10 .48550 /arXiv .2001 .09977. [24] L. Zhou, J. Gao, D. Li, H.-Y. Shum, The design and implementation of XiaoIce, an empathetic social chatbot, Comput. Linguist. 46 (2020) 1–62, https:// doi .org /10 .1162 /COLI _a _00368. [25] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: I. Guyon, U.V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems, vol. 30, Curran Associates, Inc., 2017, pp. 6000–6010, https://proceedings .neurips .cc /paper /2017 /file /3f5ee243547dee91fbd053c1c4a845aa -Paper .pdf. [26] A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al., Improving language understanding by generative pre-training, Tech. Rep., OpenAI, 2018. [27] OpenAI, GPT-4 technical report, 2023, arXiv :2303 .08774. [28] A. Akbik, D. Blythe, R. Vollgraf, Contextual string embeddings for sequence labeling, in: Proceedings of the 27th International Conference on Computational Linguistics, Association for Computational Linguistics, Santa Fe, New Mexico, USA, 2018, pp. 1638–1649, https://aclanthology .org /C18 -1139. [29] P. Ke, H. Ji, S. Liu, X. Zhu, M. Huang, SentiLARE: sentiment-aware language representation learning with linguistic knowledge, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, Association for Computational Linguistics, 2020, pp. 6975–6988. [30] E.M. Bender, T. Gebru, A. McMillan-Major, S. Shmitchell, On the dangers of stochastic parrots: can language models be too big?, in: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, Association for Computing Machinery, New York, NY, USA, 2021, pp. 610–623. [31] E. Strubell, A. Ganesh, A. Mccallum, Energy and policy considerations for deep learning in nlp, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 3645–3650. [32] M. Zhang, J. Li, A commentary of GPT-3 in MIT technology review 2021, Fundam. Res. 1(6) (2021) 831–833, https://doi .org /10 .1016 /j .fmre .2021 .11 .011. [33] S. Thukral, K. Kukreja, C. Kavouras, Probing language models for understanding of temporal expressions, in: Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, Association for Computational Linguistics, Punta Cana, Dominican Republic, 2021, pp. 396–406. [34] S. Vashishtha, A. Poliak, Y.K. Lal, B. Van Durme, A.S. White, Temporal reasoning in natural language inference, in: Findings of the Association for Computational Linguistics: EMNLP, Online, ACL, 2020, pp. 4070–4078. [35] L. Qin, A. Gupta, S. Upadhyay, L. He, Y. Choi, M. Faruqui, TIMEDIAL: temporal commonsense reasoning in dialog, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online, ACL, 2021, pp. 7066–7076. [36] W. Ding, G. Gao, L. Shi, Y. Qu, A pattern-based approach to recognizing time expressions, Proc. AAAI Conf. Artif. Intell. 33 (01) (2019) 6335–6342, https:// doi .org /10 .1609 /aaai .v33i01 .33016335. [37] S. Chen, G. Wang, B.F. Karlsson, Exploring word representations on time expression recognition, Tech. Rep. MSR-TR-2019-46, Microsoft Research, June 2019, https://www .microsoft .com /en -us /research /publication /exploring -word -representations -on -time -expression -recognition/. [38] L. Derczynski, J. Strötgen, D. Maynard, M.A. Greenwood, M. Jung, GATE-time: extraction of temporal expressions and events, in: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), European Language Resources Association (ELRA), Portorož, Slovenia, 2016, pp. 3702–3708, https://aclanthology .org /L16 -1587. [39] O. Kolomiyets, M.-F. Moens, KUL: recognition and normalization of temporal expressions, in: Proceedings of the 5th International Workshop on Semantic Evaluation, Association for Computational Linguistics, Uppsala, Sweden, 2010, pp. 325–328, https://aclanthology .org /S10 -1072. [40] M. Geva, A. Gupta, J. Berant, Injecting numerical reasoning skills into language models, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, ACL, 2020, pp. 946–958.
International Journal of Approximate Reasoning 171 (2024) 109128 19 M. Canabal-Juanatey, J.M. Alonso-Moral, A. Catala et al. [41] P. Piekos, M. Malinowski, H. Michalewski, Measuring and improving BERT’s mathematical abilities by predicting the order of reasoning, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Online, ACL, 2021, pp. 383–394. [42] Y. Ma, P.A. Crook, R. Sarikaya, E. Fosler-Lussier, Knowledge graph inference for spoken dialog systems, in: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5346–5350. [43] S. Yang, R. Zhang, S. Erfani, GraphDialog: integrating graph knowledge into end-to-end task-oriented dialogue systems, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, Association for Computational Linguistics, 2020, pp. 1878–1888. [44] C. Xu, M. Nayyeri, F. Alkhoury, H. Shariat Yazdi, J. Lehmann, TeRo: a time-aware knowledge graph embedding via temporal rotation, in: Proceedings of the 28th International Conference on Computational Linguistics, International Committee on Computational Linguistics, Barcelona, Spain (Online), 2020, pp. 1583–1593. [45] S. Liao, S. Liang, Z. Meng, Q. Zhang, Learning dynamic embeddings for temporal knowledge graphs, in: Proceedings of the 14th ACM International Conference on Web Search and Data Mining, WSDM ’21, Association for Computing Machinery, New York, NY, USA, 2021, pp. 535–543. [46] Y. Zhao, X. Wang, J. Chen, Y. Wang, W. Tang, X. He, H. Xie, Time-aware path reasoning on knowledge graph for recommendation, ACM Trans. Inf. Syst. (2022), https://doi .org /10 .1145 /3531267. [47] T. Jiang, T. Liu, T. Ge, L. Sha, B. Chang, S. Li, Z. Sui, Towards time-aware knowledge graph completion, in: Proceedings of the 26th International Conference on Computational Linguistics (COLING): Technical Papers, Osaka, Japan, 2016, pp. 1715–1724, https://aclanthology .org /C16 -1161. [48] A. Saxena, S. Chakrabarti, P. Talukdar, Question answering over temporal knowledge graphs, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online, ACL, 2021, pp. 6663–6676. [49] L. Zadeh, Fuzzy sets as a basis for a theory of possibility, Fuzzy Sets Syst. 100 (1999) 9–34, https://doi .org /10 .1016 /S0165 -0114(99 )80004 -9. [50] R. Marín, S. Barro, A. Bosch, J. Mira, Modeling the representation of time from a fuzzy perspective, Cybern. Syst. 25 (2) (1994) 217–231, https://doi .org /10 . 1080 /01969729408902325. [51] L. Zadeh, The concept of a linguistic variable and its application to approximate reasoning—I, Inf. Sci. 8(3) (1975) 199–249, https://doi .org /10 .1016 /0020 - 0255(75 )90036 -5. [52] A. Hogan, et al., Knowledge Graphs, Synthesis Lectures on Data, Semantics, and Knowledge, vol. 22, Morgan & Claypool, 2021. [53] H. Paulheim, Knowledge graph refinement: a survey of approaches and evaluation methods, SemanticWeb 8(3) (2016) 489–508, https://doi .org /10 .3233 /SW - 160218. [54] L. Ehrlinger, W. Wöß, Towards a definition of knowledge graphs, in: Proceedings of the 12th International Conference on Semantic Systems (SEMANTiCS), vol. CEUR-1695, 2016, https://ceur -ws .org /Vol -1695/. [55] G. Buchgeher, D. Gabauer, J. Martinez-Gil, L. Ehrlinger, Knowledge graphs in manufacturing and production: a systematic literature review, IEEE Access 9 (2021) 55537–55554, https://doi .org /10 .1109 /ACCESS .2021 .3070395. [56] C.H. Papadimitriou, K. Steiglitz, Combinatorial Optimization: Algorithms and Complexity, Courier Corporation, 1998. [57] T. Al-Moslmi, M. Gallofré Ocaña, A.L. Opdahl, C. Veres, Named entity extraction for knowledge graphs: a literature overview, IEEE Access 8 (2020) 32862–32881, https://doi .org /10 .1109 /ACCESS .2020 .2973928. [58] Z. Chen, Y. Wang, B. Zhao, J. Cheng, X. Zhao, Z. Duan, Knowledge graph completion: a review, IEEE Access 8 (2020) 192435–192456, https://doi .org /10 .1109 / ACCESS .2020 .3030076. [59] A. Field, G. Hole, How to Design and Report Experiments, Sage, 2002. [60] A. Joshi, S. Kale, S. Chandel, D.K. Pal, Likert scale: explored and explained, Br. J. Appl. Sci. Technol. 7(4) (2015) 396–403, https://doi .org /10 .9734 /BJAST / 2015 /14975.