Charting the Evolution and Future of Conversational Agents: A Research Agenda Along Five Waves and New Frontiers
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Schöbel, Sofia et al. Article — Published Version Charting the Evolution and Future of Conversational Agents: A Research Agenda Along Five Waves and New Frontiers Information Systems Frontiers Provided in Cooperation with: Springer Nature Suggested Citation: Schöbel, Sofia et al. (2023) : Charting the Evolution and Future of Conversational Agents: A Research Agenda Along Five Waves and New Frontiers, Information Systems Frontiers, ISSN 1572-9419, Springer US, New York, NY, Vol. 26, Iss. 2, pp. 729-754, https://doi.org/10.1007/s10796-023-10375-9 This Version is available at: https://hdl.handle.net/10419/318043 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Information Systems Frontiers (2024) 26:729–754 https://doi.org/10.1007/s10796-023-10375-9 Charting the Evolution and Future of Conversational Agents: A Research Agenda Along Five Waves and New Frontiers Sofia Sch¨ obel1·Anuschka Schmitt2·Dennis Benner3·Mohammed Saqr4·Andreas Janson2· Jan Marco Leimeister2,3 Accepted: 15 January 2023 ©The Author(s) 2023 Abstract Conversational agents (CAs) have come a long way from their first appearance in the 1960s to today’s generative models. Continuous technological advancements such as statistical computing and large language models allow for an increasingly natural and effortless interaction, as well as domain-agnostic deployment opportunities. Ultimately, this evolution begs multiple questions: How have technical capabilities developed? How is the nature of work changed through humans’ interaction with conversational agents? How has research framed dominant perceptions and depictions of such agents? And what is the path forward? To address these questions, we conducted a bibliometric study including over 5000 research articles on CAs. Based on a systematic analysis of keywords, topics, and author networks, we derive “five waves of CA research” that describe the past, present, and potential future of research on CAs. Our results highlight fundamental technical evolutions and theoretical paradigms in CA research. Therefore, we discuss the moderating role of big technologies, and novel technological advancements like OpenAI GPT or BLOOM NLU that mark the next frontier of CA research. We contribute to theory by laying out central research streams in CA research, and offer practical implications by highlighting the design and deployment opportunities of CAs. Keywords Bibliometric analysis ·Chatbot ·Conversational agent ·Voice assistant ·ChatGPT ·Large language models · Generative artificial intelligence 1 Introduction Industry and service providers have found interest in conversational agents (CAs), oftentimes colloquially called chatbots. Research on such agents is a rapidly growing field, with an increase in publications over the last few decades, especially fueled through the development of CAs such as Amazon’s Alexa that build upon voice as a modality, so called voice bots (Schmitt et al., 2021; Seaborn et al., 2022; Kendall et al., 2020). Further, recent developments such as the release of the ChatGPT beta from OpenAI pave the way towards a disruptive general AI where conversational partners are general assistants to a Anuschka Schmitt and Dennis Benner have contributed equally to this work. Sofia Sch¨ obel [email protected] Extended author information available on the last page of the article. wide variety of tasks (Haque et al., 2022). Historically, the first chatbot was developed in 1966 by Weizenbaum (1966). Since then, numerous studies have discussed and analyzed the relevance and effectiveness of chatbots and other types of CAs. With the advent of artificial intelligence (AI) and the deployment of more complex models such as neural networks (Sedik et al., 2022; Goudos et al., 2019; Kushwaha &Kar,2021), CAs offer the potential to be sophisticated interactive systems, assisting not only customers to handle their complaints but also acting as learning assistants in tutoring students or employees, for instance (Schmitt et al., 2021; Wambsganss et al., 2021). The growing relevance of CAs is further supported by the global market of CA commercialization, which is expected to grow by 24.3% until 2025. This equals a total market net worth of 1.25 billion USD. Despite technical advances and CAs’ market penetration, current research points towards unexpected challenges and previously unaddressed questions. For example, organizations experience issues integrating chatbots into customer service (Adam et al., 2020; Behera et al., 2021) because of customers’ skepticism / Published online: 20 April 2023
Information Systems Frontiers (2024) 26:729–754 toward chatbots who view such bots as unnatural, impersonal, or deceptive. In addition, CAs’ adoption can be impeded by communication and interaction problems, as well as ethical concerns (Heyselaar & Bosse, 2019). With many decades of CA research and new questions arising, the results of extant research can serve as a foundation to sketch out the frontiers of existing CA research, while such an overview also enables to identify new avenues for CA research. Therefore, with this study, we aim to provide an overview of the past, present, and future of CAs and aim to answer the following research questions: RQ1: How has research on CAs developed over time? RQ2: What are common research streams around CA research? RQ3: What are future research directions and potential novel frontiers for the design and use of CAs? To answer our research questions, we present the results of a bibliometric study that considered over 5,000 studies on CAs. In our study, we provide an overview of leading authors and countries and demonstrate the development of keywords in this research domain. Furthermore, we discuss implications for future research by introducing five historical waves of CAs marked by unique technological advances and research focus each. We contribute to theory and practice by presenting a consolidated overview of the state-of-the-art regarding the usage of CAs. This perspective is additionally supported by implications and a discussion of future research. Practitioners find support in implications about what to care about when designing and deploying CAs. In addition, by presenting different waves of the development of CAs over time, we intend to support practitioners in reacting effectively to future trends in CA design. 2 Related Research 2.1 Conversational Agents as a System Class and their Historical Roots CAs are used as communication tools between humans and technical devices. A CA involves a technological artifact that uses natural language to engage in a dialogue with the user, usually to fulfill a task or provide assistance (Følstad & Skjuve, 2019; Laban & Araujo, 2019). The first agent was developed in 1966 by Joseph Weizenbaum, who created a computer program that could communicate with humans via a text-based interface (Weizenbaum, 1966). These textbased communication programs were followed by voicebased interaction systems and embodied CAs (McTear et al., 2016). In our bibliometric study, we discuss and consider these various types and instantiation of agents. Many different terms are used concerning communicative technology. The term “conversational agent” is oftentimes used interchangeably with the terms “intelligent personal assistant” (Hauswald et al., 2016), “smart personal assistant” (Knote et al., 2021), “chatbot” (Brandtzaeg, 2018), or “conversational agent” (Feine et al., 2019). The relevance of CAs has emerged in many disciplines and tasks, such as in customer service that can be available 24/7, in which a bot assists customers in handling their complaints (Qiu & Benbasat, 2009; Behera et al., 2021). In addition, with the rise of AI, especially natural language processing (NLP), new possibilities exist to work with and change our technology. CAs enable us to find a new and convenient way of accessing services and content by improving interactions between information systems and users (Brandtzaeg, 2018; Behera et al., 2021). Since the first construction of a CA, many advancements have been made, enabling us not only to create Q&A bots but also to develop intelligent bot solutions. Some CAs use natural language to communicate with users (Feine et al., 2019). New forms of such agents can process compound natural language and thereby respond to increasingly complex user requests (Knote et al., 2021), such as Amazon’s Alexa. These assistants operate with a voice interface reacting to individual wake words and questions, like “Alexa, what time is it?” or “Alexa, is it going to rain today?” (Fischer et al., 2019). Research and practice are attempting to make both kinds of bots (textual and acoustic) more human by integrating avatars into text-based solutions (Feine et al., 2019; Purington et al., 2017) or by integrating emotional voices (Knote et al., 2021). Integrating human-like characteristics into agents may cause users to exhibit emotional, cognitive, or behavioral reactions resembling human interactions yet can also cloud users’ understanding of the system (Kr¨ amer et al., 2005). 2.2 Existing Review Approaches to Study Conversational Agents To summarize the results of extant research efforts in a particular field of study, researchers typically use literature reviews. As a point of departure for our study, we identified and studied nine literature reviews on CAs (see Table 3). We deemed these reviews detrimental in framing and summarizing research narratives in the CA domain, as well as informing us on extant conceptualizations and overviews of this system class. All of these reviews have been published within the past five years, underlining the current omnipresence of CAs. The majority of reviews focus on studies from 1998 onward, thereby neglecting first papers on CAs such as well-known first text-based agent ELIZA (Weizenbaum, 1966). In addition, we quickly 730
Information Systems Frontiers (2024) 26:729–754 realized that extant reviews mainly take a socio-technical design perspective by exploring specific design themes such as CA team collaboration (Diederich et al., 2022; Poser et al., 2022; Stieglitz et al., 2021; Abedin et al., 2022) or CAs for education (Smutny & Schreiberova, 2020). While these reviews allow for strong theoretical and practical implications regarding the design of such agents for particular contexts, we are missing a holistic overview of the conceptual and empirical foundations of CA research which also considers the underlying technical advances related to CAs. The number of studies analyzed in literature reviews ranges from 29 articles (Rheu et al., 2021) to 262 articles (Diederich et al., 2022). Whereas some of the literature reviews have a more centralized view on topics involving CAs, for example, the role of trust supporting components (Rheu et al., 2021;V ¨ ossing et al., 2022), others discuss the meaning of disembodied chatbots in human-computer interaction (Chaves & Gerosa, 2021) or characteristics of the adoption of AI-based CAs (Rzepka & Berger, 2018; Elshan et al., 2022). More and more questions involve a discussion of how to make CAs more useful by experiencing how to make them more human-like (Chaves & Gerosa, 2021) or about how to make them a better collaborative partner in communication (Diederich et al., 2022; Elshan et al., 2022; Poser et al., 2022; Abedin et al., 2022). Thus, literature reviews often provide a closer look at specific topics surrounding chatbots (e.g., trust) and serve as a valuable contribution to the study of specific areas in CA research. All the literature reviews suggest areas for future research, such as Smutny and Schreiberova (2020), who suggested analyzing how we can support developers to create and offer tools that allow any teacher to integrate chatbots into their classes. Smutny and Schreiberova (2020) analyzed educational chatbots on Meta (formerly Facebook), this perspective could be better supported by including literature to enrich the perspective on the topic. Research has discussed how we can support teachers in using CAs for their teaching (e.g., Winkler and S¨ ollner (2018)). Other studies suggest better analyzing the behavior in the communication of chatbots by differentiating and considering various variables, such as gender or personality, to detect individual designs (Bickmore et al., 2020;Ahmad et al., 2022). Discussed across various literature reviews is the aspect to take a closer look at the role of AI for CAs, such as the fit between users and systems for different AI-enabled systems (Rzepka & Berger, 2018). Compared to other methods of literature screening, reviews often only consider a limited, selected number of studies. After 40 years of analyzing the role and meaning of CAs, we should look back to identify which authors and topics are most impactful and to discuss how and if chatbots could determine future research. Bibliometric analyses support us in obtaining such a perspective by overcoming the limits of literature studies in considering only a limited number of studies. Other than literature reviews, bibliometric studies oftentimes analyze thousands of publications, utilize statistical methods, and, therefore have become a popular method to identify patterns in collected data and work that reveal emerging trends in research as it evolves (Trinidad et al., 2021). A bibliometric approach can thus provide a more exhaustive overview of extant research while simultaneously uncover patterns and details within a research field other research methods are incapable of. To provide an overview of related work, we summarize the results of major CA reviews and the bibliometric study we identified in Appendix B. So far, only one study has used a bibliometric approach to study CAs (Io & Lee, 2017). In their work, the authors analyzed past research on CAs. They presented an overview of the number of publications for each year and calculated a cluster of keywords, research areas, and keyword co-occurrences. However, the presented insights and suggestions for future research (e.g., analyze other CAs, such as mobile chats, embedded gadgets) only remain on a surface level. With our study, we want to expand the work of Io and Lee (2017) and overcome their limitations in different ways. First, we performed a coding procedure to identify relevant studies using rayyan.ai as a tool to determine which study fit our goal. Second, we opened up our study by considering all kinds of CAs, enabling us to detect trends and relationships of a specific group of CAs. 3 Methodology The search protocol followed the preferred reporting items for systematic reviews and meta-analyses literature search extension (PRISMA-S) method (Rethlefsen et al., 2021). We chose the Scopus database because it covers most articles included in Web of Science and includes a larger selection of technical and social science articles (Norris & Oppenheim, 2007). The Scopus database has rigorously maintained metadata and a curated list of journals and conferences that are subjected to a strict inclusion process (Baas et al., 2020; Singh et al., 2021). Several search iterations were performed, and the authors thoroughly examined the results to choose the most appropriate formula. The final search formula was “chatbot*” or “virtual assistant*” or (“conversation* agent*”) or “smart assistant*” or “smart bot*” or “natural language interface”. Only journal articles, conference proceedings, or book chapters of conference proceedings that were published in English prior to 2021 (to include complete years in the trend analysis) were included. Reviews, editorials, and other 731
Information Systems Frontiers (2024) 26:729–754 non-original articles were excluded. The data were downloaded on the 31st of May 2021. A total of 6,397 articles were retrieved. Two reviewers (first author and third author) reviewed 681 articles and agreed on 656, with an interrater agreement Cohen’s Kappa of 0.963 (adjusted = 0.892). The two authors discussed the disagreements and resolved them based on a common understanding. Then, a single reviewer (third author) completed the remaining articles. When the author was in doubt, the article was classified as “maybe” and resolved based on a review of the full text. A total of 5,107 articles were judged as matching the inclusion criteria, and one article was a duplicate and removed. The data were processed, cleaned, and prepared for analysis, so 1) the author names were manually checked and fixed with the aid of Scopus author IDs. Authors’ names with different spellings, name changes, and special characters were therefore fixed. 2) Publication venues were cleaned by inspecting and fixing different venues with the same title and spelling variations (e.g., Computers & Education and Computers and Education). 3) Keywords were compiled and cleaned with the aid of Google Openrefine. Openrefine is an open-source software designed to help clean up messy datasets; it offers several NLP algorithms and clustering methods for identifying similar keywords in terms of spelling or phonetics (e.g., “chatbot,” “chatbots,” “chat bots,” “chat-bots,” “chat bot”). All such keywords with their identical alternatives were checked carefully by the third author and merged. Another round of keyword cleaning was performed to merge keywords that Openrefine could not recognize, such as “NLP” and “Natural Language Processing,” and “AI” and “Artificial Intelligence”. The cleaned dataset was analyzed using the R statistical language Bibliometrix package (Aria & Cuccurullo, 2017). Bibliometrix is an open-source software that offers a wide array of tools for the analysis of bibliometric metadata, such as authors, keywords, citations, countries, and references. Frequencies, plots, and temporal trends were computed and plotted using R statistical language. The top authors were retrieved based on their number of articles within our dataset, and their citations and evolution of citations were computed. Co-authorship networks offer a powerful summarization method for the visualization of collaboration and scientific production that helped build and shape a scientific field. The co-authorship network was constructed based on co-authorship of the same article (i.e., if two authors collaborated on the same article, they were considered connected). To avoid assigning higher weights to manuscripts with a higher number of authors, we chose a fractional counting method to build a weighted co-authorship network (i.e., edge weights are inversely proportional to the number of authors) based on the recom mended methods by Perianes-Rodriguez et al. (2016). Subgroups (i.e., communities) of authors who frequently collaborated were mapped using Louvain modularity (De Meo et al., 2011) to highlight groups who frequently collaborated and their contribution and influence on the field. The resultant network was plotted using Gephi (Bastian et al., 2009) with the Fruchterman Reingold layout algorithm (Fruchterman & Reingold, 1991), and each community was assigned a unique color (Fruchterman & Reingold, 1991). For readability, the network was limited to the top 100 authors, based on a weighted degree threshold of 10. The country collaboration network was constructed in the same way (fractional counting) based on the authors’ country affiliation. Hence, two countries were considered connected if two authors affiliated with the two countries collaborated on the same article. The country network was plotted using the Fruchterman Reingold layout algorithm, and each community was assigned a unique color (Fruchterman & Reingold, 1991), and communities of countries who collaborated were identified using the Louvain modularity (De Meo et al., 2011). A keyword network was constructed using the full counting method, where keywords in the same manuscript were considered connected. The network of keyword co-occurrence was plotted using Gephi and partitioned with Louvain modularity (De Meo et al., 2011)sothat keywords that occurred together were connected and colored similarly. 4 Results 4.1 General Results In the following, we reviewed the ten most cited articles to arrive at an understanding of relevant topics driving CA research (see Table 1). Based on their citation numbers, we assume that these studies are substantially relevant for establishing and driving a common understanding of CA as a concept and making CAs prominent in the fields of Human-Computer Interaction (HCI) and Information Systems. Publication dates for those ten manuscripts are dispersed, ranging from 1994 to 2016. This dispersion points toward the importance of the first technology established in this field and the ongoing interest in research on CAs fueled by ongoing advances in related technology. 4.2 Authors and Countries In a similar vein, we analyzed the most proliferative authors, co-citations and related countries (see Figs. 7,8and 9in Appendix C). 732
Information Systems Frontiers (2024) 26:729–754 Table 1 Top-cited Manuscripts Author (year) Citations Keywords Category Bickmore and Picard (2005) 581 User interfaces-evaluation / methodology, (Embodied) CA graphical user interfaces, interaction styles, natural language, theory and methods, voice Androutsopoulos et al. (1995) 384 N/A Conversational Interface Graesser et al. (2005) 375 CAs, intelligent tutoring systems, CA natural language dialogue Cassell et al. (1994) 340 N/A CA Luger and Sellen (2016) 327 CAs, mental models, evaluation CA Cassell et al. (1994) 294 Conversational characters, multimodal input, Embodied CA intelligent agents, multimodal output Li et al. (2016) 277 N/A CA Vertegaal et al. (2001) 251 Attention-based interfaces, multiparty CA communication, gaze, conversational attention, attentive agents, eye tracking Bickmore and Cassell (2001) 235 Embodied CA, trust, social interface, (Embodied) CA natural language, small talk, personality Graesser et al. (2001) 226 N/A Intelligent Agent With 91 articles, Pelachaud is the most productive author. Pelachaud’s works have been mainly been published in the 2010s and focus on multimodal dialogue systems, conversation and related emotional and nonverbal competencies, and image classification. With 61 published articles, Pelachaud is followed by Bickmore as the second most productive author who published between 2002 and 2020. Bickmore’s works are prominently positioned in two of the most cited articles in CA research (see Table 1). Similar to Pelachaud’s, Brickmore’s research is scattered over more than 20 years, becoming more relevant around the 2010s. His work encompasses domain-specific CA applications, such as a hospital companion agent (Bickmore et al., 2015) and the deployment of CAs for automated substance use screening (Bickmore et al., 2020), and on how to build user trust through the design of CAs (Bickmore & Cassell, 2001). The third most productive author is Griol, with 47 published articles. Griol’s work has mostly been published around the 2010s, focusing on dialogue system design and user analytics (i.e., predicting users’ mental states from CA interaction) (Callejas et al., 2011). Compared with the other two authors, however, Griol focuses much more on speech and language and has thus also been published more heavily in respective outlets (Griol et al., 2008). In addition, Griol’s manuscripts are prevalent from 2010 onwards. In a subsequent step, we explored the co-citation references among the important authors (see Fig. 9). Our co-cited reference network, defined by four key clusters, allows us to better understand how the field was built and which theoretical foundations mainly drove CA research throughout time. One of the key clusters of CA literature dates back to Turing’s (Trinidad et al., 2021) well-known paper proposing his test to assess whether a machine can think or, has a conscience. His paper introduced the ongoing discussion around the ”humanness” of machines and the technical feasibility and future of machines. Less than 20 years later, Weizenbaum (1966) introduced the first computer program based on natural language communication, called ELIZA, and addressed the question of how CAs should communicate by laying out the differences between online and offline language. These publications represent the first wave of research on CAs, which were interactional interfaces built on rule-based models. Importantly, these papers are fundamental to ongoing advances in CA research and development (Dale, 2016). A second dominant cluster comprises interactions among researchers (e.g., Hochreiter and Schmidhuber (1997)) is concerned with the underlying technical functions of CAs. Technical methods, such as storing recent inputs for speech processing using gradient models (Heyselaar & Bosse, 2019) or neural network-based response generation (Shang et al., 2015; Kushwaha & Kar, 2021), can be viewed as key contributions to the technical advancements and thus sophistication of CA. This body of literature is composed of the second wave of CA research around NLP and text-based and statistical methods. The purple-colored research cluster is defined by authors, such as Nass, and is concerned with the personality of CAs. In 1994, Nass et al. claimed through five experiments that computers are social actors (CASA). Following this socalled CASA paradigm, research on anthropomorphism and 733
Information Systems Frontiers (2024) 26:729–754 perceiving technology as a human makes up a predominant literature stream in contemporary CA research. More scattered is a smaller network around the research of Cassell et al. (1994) and Kopp and Wachsmuth (2004), focusing on the structure of communication and nonverbal behavior. In Cassells et al. (1994) work, elements of human conversation are reviewed, which can be inferred from CAs. Kopp and Wachsmuth (2004) synthesized speech-based utterances with nonverbal modalities to achieve congruent multimodality for CAs. We further analyzed publications and research focus across countries (see Fig. 9). In general, authors from the same country published together, with an average collaboration score of 2.46. With almost 10,000 total citations, the United States dominate in terms of prevalent work. A cluster of countries follows with a smaller number of total citations, with Germany with 2,595 total citations as the country with the second-most citations. The United Kingdom, France, China, Italy, and Japan exhibit between 1,000 and 2,000 citations. In a similar vein, the mentioned countries are also the most active in terms of collaboration across countries. As shown in Fig. 9, authors from the United States publish their work with researchers from many other dominant countries, including China, Germany, the UK, Japan, Canada, India, and Australia. We also identified a collaboration network of Western countries, including Germany, France, Italy, Switzerland, the Netherlands, Austria, and Portugal. 4.3 Keyword Analysis The next step of our analysis included a detailed keyword analysis (see Figs. 1and 2). Our top 20 keywords AI, chatbot, neural network, NLP, virtual assistant, and machine learning were only limitedly used until 2015. The use of each of them greatly increased during the last 5–7 years. This is different from the keywords affective computing, embodied CA, emotion, natural language interface, and virtual agent. These keywords have no linear or clear development, but all of them increase over time. Lastly, deep learning has been relevant from the beginning onward, similar to the keyword dialogue system. A comparison of keyword development will demonstrate that some keywords are more dominant than others. The keyword most prominent from 2003 to around 2016 was “embodied conversational agents”. The development of this keyword has decreased since 2017. Other than this, Fig. 1 Distribution of the Top 20 Keywords per Year 734
Information Systems Frontiers (2024) 26:729–754 Fig. 2 Yearly Distribution of Keywords the keyword “conversational agent” has been used more often in studies beginning in 2007. This is similar to the keyword chatbot, which has become more and more relevant beginning in 2015. All other keywords, such as Human- Computer Interaction or NLP, have been relevant over the last couple of years on a continuous level and to a high frequency. 4.4 Keyword Cluster Turning toward the 50 most prominent keywords in CA literature, the definition of the system class represents the most relevant type of keywords. 11 out of the 50 top keywords refer to a particular system class, such as chatbot, embodied CA, or virtual assistant. The most prominent are text-based CAs, with the chatbot(s) being most frequent, followed by embodied CAs. These different definitions and connotations of CAs point toward the multitude of types and diversity of CAs, also dependent on the modality of the CA or the publishing outlet of a paper. This diversity also illustrates the fragmentation of CA research and lack of agreement on the terminology. Another prominent keyword cluster is the reference to the technical model of the CA discussed. Approximately 26% of the top keywords provide insights into the statistical models, pointing toward the importance of technical advances for the proliferation of CAs. AI, machine learning, and natural language (processing) dominate the keywords, which also illustrates how this technical interplay makes the architecture of CAs unique. Approximately 30% of the 50 most prominent keywords illustrate the research focus. Some keywords include related technology relied upon, such as virtual or augmented reality and the Internet of Things. The most relevant research focuses are trust and anthropomorphism, and agent personality and emotions. Interaction-related questions around the multimodality of CAs and the dialogue structure also dominate the keywords in the research focus. Oftentimes, the research domain of a publication is indicated in the keywords. Domains dominating the CA literature include HCI, affective computing, and ontology-related matters. Lastly, keywords indicate the interaction context of a publication (12 out of the 50 top keywords). CA studies seem to be deployed mostly in question answering, information retrieval, and evaluation interactions. E-learning and mental health represent relevant interaction contexts, which have also seen great 735
Information Systems Frontiers (2024) 26:729–754 Fig. 3 Relative Yearly Distribution of Corporate Affiliations prominence in recent publications (Androutsopoulos et al., 1995; Chaves & Gerosa, 2021; Følstad & Skjuve, 2019; Hauswald et al., 2016). Turning toward keyword clusters (Fig. 10), we see that certain tasks and CAs are commonly associated with specific terminology. The term “conversational agent” is mostly linked to terms concerned with the conversational nature and dialogue of the interaction, such as natural language and dialogue management. Tasks such as question answering and information retrieval are associated with CAs. Lastly, smart personal assistants, such as Amazons Alexa, are linked to the term as well. The term “voice user interface” appears within the CA cluster, yet is not linked to any further words. An equally extensive keyword cluster exists around the term “chatbot”. Chatbots are predominantly studied in interaction contexts around education, e-learning, and deception detection. Technical models, such as NLP, deep learning, and machine learning, appear here. The chatbot cluster is closely linked with a smaller cluster around the term AI, where other technical keywords, such as automation, sentiment analysis, and big data, are found. The term “embodied conversational agent”, is more concerned with additional modalities, including the avatar or facial expression of the agent. In this context, the personality and emotions of agents are commonly discussed. Most technical advancements seem to have been empirically investigated in the context of text-based CAs. Voice-based CAs find little appearance in our keyword cluster, which may be the result of a certain choice of dominant keywords or voice-based CAs not having been researched as extensively as text-based ones. The term “conversational agent” does not necessarily seem to act as an overarching term for the different types of agents. Moreover, certain tasks and contexts are connoted with certain terminology. 4.5 Corporate Affiliations Given the commercial availability of CAs and the increasing research focus within large technical companies, we deem corporate affiliations a noteworthy aspect to be mentioned when discussing the past, present, and future of CAs (see Fig. 3). Publications by corporate affiliations were dominated by IBM and Microsoft until 2008, followed by Google from 2009 until 2011. From 2012 onwards, publications by commercially affiliated researchers have become increasingly fragmented (i.e., through a greater number of players and Asian companies, such as Baidu, becoming apparent). Turning toward the temporal development of corporate-related keywords, we identified a steady increase in the frequency of these keywords since 2000 and a steep increase since 2017 (see Fig. 4). More specifically, Google is most frequently listed as a keyword in 2020 (N = 78), followed by Met1(N = 49) and Microsoft (N = 41). Whereas Eastern players (i.e., Baidu, ASAPP) have increased in appearance, western players still dominate. Regardless, corporate 1https://ai.facebook.com/blog/democratizing-access-to-large-scale- language-models-with-opt-175b/ 736
Information Systems Frontiers (2024) 26:729–754 to current state of the art CAs like Alexa. We have described the past and present CA research (RQ1 and RQ2). Further, we have highlighted potential future research streams according to our keyword analysis, pointing out what topics are currently state of the art and what will be the focus of research in the future (i.e., next developmental generation of CAs), thus answering RQ3. Here, we have pointed out that future research may likely see a strong focus on data-driven and AI-driven approaches, as highlighted by our keyword analysis and five-wave view on CA research. In this context, we have highlighted the developmental waves of CA research from the beginning (i.e., the 1960s) to our current time and what could be the next big wave for CA research (i.e., data and AI). In addition, we have pointed out some limitations in the existing body of research on CAs and the importance of addressing these potential gaps. However, these research gaps should be seen as opportunities, as highlighted by the development of keywords leaning toward AI-driven methods. Moreover, we have briefly highlighted some comparative studies with a similar yet different focus and contribution to provide a more holistic overview of the past, present, and future of CA research. Although we conducted our research according to established research guidelines and used established statistical methods that we followed as rigorously as possible, our research may not be without limitations. First, although we did our best to include as many relevant articles as possible by choosing the appropriate keywords and outlets and not using limitations like publication year restrictions, we may have missed some niche domain or research stream. Moreover, we may not have included literature that uses more specific keywords exclusively. An example of this would be research that uses ”smart personal assistant/agent” as keywords. Second, only two of the authors reviewed the literature in detail and made decisions to include or exclude the literature. Reviewing the articles with more or other reviewers may have led to different results and, thus, a somewhat divergent body of literature. Overall, our research highlights the past, present, and potential future development of CA research covering implications for research and practice. Both partial contributions come together in our presented five-wave view on CA research, thus highlighting the origins of CAs, their development until today, and potential future developments with the next big wave (i.e., data- and AI-driven CAs). On the one hand, our research contributes to theory by highlighting this development and pointing out our potential future research venues based on our keyword analysis, derived five waves of CA research and the implications of these findings. On the other hand, our research contributes to practice by highlighting the technical developments based on our keyword analysis and the indicated affiliations with the bigtech corporate world that currently drives CA research in various areas, particularly AI-related topics. Moreover, our five waves of CA research highlight the current status quo and ongoing development future that can form the basis for future CA developments and directions. Here, we find AI and relevant topics related to it (e.g., trust in AI, explainable AI) to be very likely to continue the throughout the next wave. In general, we hope that our research provides both researchers and practitioners with an exhaustive and holistic overview of CA research with an emphasis on future developments and research opportunities that should and most likely will be addressed within the current decade when considering the fifth wave of CA research. Appendix A: Descriptive Results In our study, we present articles published from 1982 to 2021. A summary of the main information is given in Table 2. Table 2 Summary of General Statistics of Bibliometric Analysis General Data Sources (Journals, Books, etc.) 1,920 Documents 5,106 Average years from publication 5.81 Average citation per document 8.047 Average citation per year per document 1.237 Document Types Articles 1,290 Book chapters 130 Conference papers 3,686 Document Contents Keywords plus (ID) 14.801 Authors’ keywords (DE) 8.327 An overview of the annual scientific production over the years and the average article citations per year is shown in Fig. 6. 743
Information Systems Frontiers (2024) 26:729–754 Fig. 6 Annual Scientific Production and Citations. (The size of the circle is proportional to the number of citations) Appendix B: Related Work Table 3 Summary of Related Literature Reviews and Bibliometric Studies Author Method Review Focus of Study Research Implications (Year) (Sample Size) Period Chaves and Gerosa (2021) Literature review Up to Analysis of social characteristics Analyze if and how moral agency improves (56 articles) 2018 of disembodied chatbots communications with bots Analyze social in the HCI context characteristics of embodied chatbots Diederich et al. (2019) Literature review Up to Discussion of status quo in CA in team collaboration settings (36 articles) 2018 research involving CAs Conduct more design-oriented studies Explore embodied agents Study CA introduction and use in the field Diederich et al. (2022) Organizing and Up to Analysis of CA design and 6 research avenues for CA in IS assessing review 2019 research streams 16 actionable directions for research (262 articles) Call for longitudinal studies Elshan et al. (2022) Literature review No limits Understanding of User Acceptance Conflict of adaptability and standardization (107 articles) of CAs and future research agenda Investigation in freedom and autonomy needed Investigating effects of CA components Rheu et al. (2021) Literature Review Up to Discussion and analysis of Research should focus on optimization (29 articles) 2019 trust supporting components Analyze role of trust and user expectations Optimization and trust in AI by design Rzepka and Berger (2018) Literature Review Up to Characteristics of adoption Beneficial degree of anthropomorphism (96 articles) 2018 of AI-based CAs Differences in user behavior towards AI Effect of transparency and AI on interaction Negative interactions between user and AI Smutny and Schreiberova (2020) Field Analysis No Analysis of educational Support CA integration in education (89 CAs) limits CAs on Meta Content analysis of CAs with users Van Pinxteren et al. (2020) Literature review From 1999 Effect of CA behavior on Analyze user needs and group preferences (61 articles) to 2018 outcomes for service encounters Analyze categories of communicative behavior Analyze (non-)verbal behavior relations 744
Information Systems Frontiers (2024) 26:729–754 Table 3 (continued) Author Method Review Focus of Study Research Implications (Year) (Sample Size) Period Io and Lee (2017) Bibliometric study From 1998 Analyze the past of CA research New trends to analyze role of CAs (61 articles) to 2017 Change way of CA development Analyze business roles and use of CAs Appendix C: Network Graphs Fig. 7 Author Network 745
Information Systems Frontiers (2024) 26:729–754 Fig. 8 Co-citation References 746
Information Systems Frontiers (2024) 26:729–754 Fig. 9 Country Network 747
Information Systems Frontiers (2024) 26:729–754 Fig. 10 Keyword Cluster 748
Information Systems Frontiers (2024) 26:729–754 Appendix D: Future Frontiers for CA Research Table 4 Proposed Frontiers for the Future of CA Research Frontier Current Status and Directions for Future Theory Status quo: Many past and current studies rely on old theories not native to CA research (e.g., social presence (Short et al., 1976), computers as social actors (Nass et al., 1994)) Directions: We call for native CA research theories that are based on and adapted to the boundaries of human-agent interaction. Researchers may draw on the existing theories but should advance CA research theory towards a stand-alone theoretical base. Technical Status quo: Advanced natural language models and machine learning algorithms are being used but no “true” AI Directions: Research should focus on establishing a “true” conversational AI. Fundamental technical advances have been made, especially with large language models (e.g., BLOOM, GPT) but need to be developed towards true human natural language and especially human thinking including social features e.g., social cues, humor, sarcasm (Allouch et al., 2021). Design Status quo: Most current CAs do not include personalized features, trust or explainable AI features or feature believeable humanization Directions: Researchers and developers need to find a way to personalize CAs but at the same time respect ethical boundaries and establish trust in the CA and the underlying mechanisms. In this context persuasive design features may provide a solution, however their boundaries must be respected (Benner et al., 2022) General Status quo: Most CA software, programs, libraries, models, etc. are free and open source. This fosters CA research and enables independant researchers to contribute. Directions: While monetarization is important for corporate players, restriction of access may have serious adverse effects and hinder development and contributions by independant researchers and developers. Therefore, corporate players must find a balance between monetarization and open access Managerial Status quo: Many companies use CAs but most being simple clickbots or less advances CAs that may even hurt the interaction. Directions: Managers should transform their customer interaction and invest on CAs with actual conversational capability to improve user experience and their market position (Abedin et al., 2022). Society Status quo: CAs already have a significant impact on society which is expected to extend. Directions: Emergent fields of application for CAs as well as transforming existing fields and pushing the boundaries may become even more significant with increased societal impact (Wahde & Virgolin, 2022) 749
Information Systems Frontiers (2024) 26:729–754 Acknowledgements The authors acknowledge partial funding of the following funding bodies: “Stiftung Innovation in der Hochschullehre” within the project “Universit¨ at Kassel digital: Universit¨ are Lehre neu gestalten”, Swiss National Science Foundation (project ID: 100013 192718), Academy of Finland (Suomen Akatemia) Research Council for Natural Sciences and Engineering for the project Towards precision education: Idiographic learning analytics (TOPEILA), Decision Number 350560, and the fifth author acknowledges personal funding from the Basic Research Fund (GFF) of the University of St. Gallen. Funding Open Access funding enabled and organized by Projekt DEAL. Declarations Authors two and three have contributed equally. The authors have no conflicts of interest to declare that are relevant to the content of this article. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons. org/licenses/by/4.0/ References Abedin, B., Meske, C., Junglas, I., Rabhi, F., & Motahari- Nezhad, H. R. (2022). Designing and managing human-ai interactions. Information Systems Frontiers.https://doi.org/10. 1007/s10796-022-10313-1. Adam, M., Wessel, M., & Benlian, A. (2020). Ai-based chatbots in customer service and their effects on user compliance. Electronic Markets, pp. 1–19. Ahmad, R., Siemon, D., Gnewuch, U., & Robra-Bissantz, S. (2022). Designing personality-adaptive conversational agents for mental health care. Information Systems Frontiers, pp. 1–21. https://doi.org/10.1007/s10796-022-10254-9. Allouch, M., Azaria, A., & Azoulay, R. (2021). Conversational agents: goals, technologies, vision and challenges. Sensors (Basel Switzerland), vol 21(24). https://doi.org/10.3390/s21248448. Androutsopoulos, I., Ritchie, G. D., & Thanisch, P. (1995). Natural language interfaces to databases–an introduction. Natural Language Engineering,1(1), 29–81. Aria, M., & Cuccurullo, C. (2017). Bibliometrix: an r-tool for comprehensive science mapping analysis. Journal of informetrics, 11(4), 959–975. Baas, J., Schotten, M., Plume, A., Cˆ ot´ e, G., & Karimi, R. (2020). Scopus as a curated, high-quality bibliometric data source for academic research in quantitative science studies. Quantitative Science Studies,1(1), 377–386. Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. Behera, R. K., Bala, P. K., & Ray, A. (2021). Cognitive chatbot for personalised contextual customer service: behind the scene and beyond the hype. Information Systems Frontiers:1–21. https://doi.org/10.1007/s10796-021-10168-y. Benner, D., Sch¨ obel, S., Janson, A., & Leimeister, J. M. (2022). Propositions for ethical persuasive design in information systems. AIS Transactions on Human-Computer Interaction, vol 14(4). Bickmore, T., Asadi, R., Ehyaei, A., Fell, H., Henault, L., Intille, S., Quintiliani, L., Shamekhi, A., Trinh, H., & Waite, K. (2015). Context-awareness in a persistent hospital companion agent. International Conference on Intelligent Virtual Agents, pp. 332– 342. Bickmore, T., & Cassell, J. (2001). Relational agents: a model and implementation of building user trust Proceedings of the SIGCHI conference on Human factors in computing systems, pp. 396–403. Bickmore, T. W., & Picard, R. W. (2005). Establishing and maintaining long-term human-computer relationships. ACM Transactions on Computer-Human Interaction (TOCHI),12(2), 293– 327. Bickmore, T., Rubin, A., & Simon, S. (2020). Substance use screening using virtual agents: towards automated screening, brief intervention, and referral to treatment (sbirt). Proceedings of the 20th ACM International Conference on Intelligent Virtual Agents, pp. 1–7. BigScience (2022). BigScience research workshop. https://bigscience. huggingface.co/. Accessed 22 Dec 2022. BigScience (2022b). BLOOM: a 176B-parameter open-access multilingual language model. Bocklisch, T., Faulkner, J., Pawlowski, N., & Nichol, A. (2017). Rasa: open source language understanding and dialogue management. NIPS 2017 Conversational AI workshop. Brandtzaeg, P. B. (2018). Chatbots: changing user needs and motivations. Interactions,25(5), 38–43. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ...Amodei, D. (2020b). Language models are few-shot learners. arXiv:2005.14165. Callejas, Z., Griol, D., & L´ opez-C´ ozar, R. (2011). Predicting user mental states in spoken dialogue systems. EURASIP Journal on Advances in Signal Processing,2011(1), 1–21. Cassell, J., Pelachaud, C., Badler, N., Steedman, M., Achorn, B., Becket, T., Douville, B., Prevost, S., & Stone, M. (1994). Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents. Proceedings of the 21st Annual Conference on Computer Graphics and Interactive Techniques, pp. 413–420. Castelvecchi, D. (2022). Are chatgpt and alphacode going to replace programmers? Nature. https://doi.org/10.1038/d41586-022- 04383-z. Chaves, A. P., & Gerosa, M. A. (2021). How should my chatbot interact? a survey on social characteristics in human–chatbot interaction design. International Journal of Human–Computer Interaction,37(8), 729–758. Dale, R. (2016). The return of the chatbots. Natural language engineering,22(5), 811–817. Davenport, T., & Mittal, N. (2022). How generative ai is changing creative work. Harvard Business Review. De Meo, P., Ferrara, E., Fiumara, G., & Provetti, A. (2011). Generalized louvain method for community detection in large networks. International conference on intelligent systems design and applications, pp. 88–93. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. 750
Information Systems Frontiers (2024) 26:729–754 Diederich, S., Brendel, A. B., & Kolbe, L. M. (2019). On conversational agents in information systems research: analyzing the past to guide future work. International Conference on Wirtschaftsinformatik, pp. 1550–1564. Diederich, S., Brendel, A. B., Morana, S., & Kolbe, L. (2022). On the design of and interaction with conversational agents: an organizing and assessing review of human-computer interaction research. Journal of the Association for Information System (JAIS). Elshan, E., Zierau, N., Engel, C., Janson, A., & Leimeister, J. M. (2022). Understanding the design elements affecting user acceptance of intelligent agents: past, present and future. Information Systems Frontiers.https://doi.org/10.1007/s10796-021-10230-9. Elshan, E., Ebel, P., S¨ ollner, M., & Leimeister, J. M. (2023). Leveraging Low Code Development of Smart Personal Assistants: An Integrated Design Approach with the SPADE Method. In: Journal of Management Information Systems. https://doi.org/10. 1080/07421222.2023.2172776. Feine, J., Gnewuch, U., Morana, S., & Maedche, A. (2019). A taxonomy of social cues for conversational agents. International Journal of human-computer studies,132, 138–161. Fischer, J. E., Reeves, S., Porcheron, M., & Sikveland, R. O. (2019). Progressivity for voice interface design. Proceedings of the 1st international conference on conversational user interfaces, pp. 1– 8. Fruchterman, T. M. J., & Reingold, E. M. (1991). Graph drawing by force–directed placement. Software: Practice and Experience, 21(11), 1129–1164. Følstad, A., & Skjuve, M. (2019). Chatbots for customer service: user experience and motivation. Proceedings of the 1st international conference on conversational user interfaces, pp. 1–9. Goudos, S. K., Tsoulos, G. V., Athanasiadou, G. E., Batistatos, M. C., Zarbouti, D. A., & Psannis, K. E. (2019). Artificial neural network optimal modeling and optimization of uav measurements for mobile communications using the l-shade algorithm. IEEE Transactions on Antennas and Propagation,67, 4022– 4031. Graesser, A. C., Chipman, P., Haynes, B. C., & Olney, A. (2005). Autotutor: an intelligent tutoring system with mixed-initiative dialogue. IEEE Transactions on Education,48(4), 612–618. Graesser, A. C., VanLehn, K., Ros´ e, C. P., Jordan, P. W., & Harter, D. (2001). Intelligent tutoring systems with conversational dialogue. AI Magazine,22(4), 39. Griol, D., Hurtado, L. F., Segarra, E., & Sanchis, E. (2008). A statistical approach to spoken dialog systems design and evaluation. Speech Communication,50(8-9), 666–682. Hauswald, J., Laurenzano, M. A., Zhang, Y., Yang, H., Kang, Y., Li, C., Rovinski, A., Khurana, A., Dreslinski, R. G., & Mudge, T. (2016). Designing future warehouse-scale computers for sirius, an end-to-end voice and vision personal assistant. ACM Transactions on Computer Systems (TOCS),34(1), 1–32. Haque, M. U., Dharmadasa, I., Sworna, Z. T., Rajapakse, R. N., & Ahmad, H. (2022). I think this is the most disruptive technology: exploring sentiments of ChatGPT early adopters using twitter data. arXiv:2212.05856. Heyselaar, E., & Bosse, T. (2019). Using theory of mind to assess users’ sense of agency in social chatbots. International Workshop on Chatbot Research and Design, pp. 158–169. Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation,9(8), 1735–1780. Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A. (2020). Spacy: industrial-strength natural language processing in python. https://doi.org/10.5281/zenodo.1212303. Huang, M.-H., & Rust, R. T. (2021). Engaged to a robot? the role of ai in service. Journal of Service Research,24(1), 30–41. Io, H. N., & Lee, C. B. (2017). Chatbots and conversational agents: a bibliometric analysis. IEEE International Conference on Industrial Engineering and Engineering Management (IEEM), pp. 215–219. Kendall, L., Chaudhuri, B., & Bhalla, A. (2020). Understanding technology as situated practice: Everyday use of voice user interfaces among diverse groups of users in urban india. Information Systems Frontiers,22(3), 585–605. https://doi.org/10. 1007/s10796-020-10015-6. Knote, R., Janson, A., S¨ ollner, M., & Leimeister, J. M. (2021). Value co-creation in smart services: a functional affordances perspective on smart personal assistants. Journal of the Association for Information Systems,22(2),5. Kopp, S., & Wachsmuth, I. (2004). Synthesizing multimodal utterances for conversational agents. Computer Animation and Virtual Worlds,15(1), 39–52. Kr¨ amer, N. C., Iurgel, I., & Bente, G. (2005). Emotion and motivation in embodied conversational agents. Proceedings of the symposium agents that want and like, artificial intelligence and the simulation of behavior, pp. 55–61. Kushwaha, A. K., & Kar, A. K. (2021). Markbot – a language model-driven chatbot for interactive marketing in post-modern world. Information Systems Frontiers.https://doi.org/10.1007/ s10796-021-10184-y. Laban, G., & Araujo, T. (2019). Working together with conversational agents: the relationship of perceived cooperation with service performance evaluations. International Workshop on Chatbot Research and Design, pp. 215–228. Lee, K., Lee, K. Y., & Sheehan, L. (2020). Hey alexa! a magic spell of social glue?: sharing a smart voice assistant speaker and its impact on users’ perception of group harmony. Information Systems Frontiers,22(3), 563–583. https://doi.org/10. 1007/s10796-019-09975-1. Li, J., Monroe, W., Ritter, A., Galley, M., Gao, J., & Jurafsky, D. (2016). Deep reinforcement learning for dialogue generation. arXiv:1606.01541. Luger, E., & Sellen, A. (2016). Like having a really bad pa the gulf between user expectation and experience of conversational agents. CHI Conference on Human Factors in Computing Systems, pp. 5286–5297. Mattioli, D., Herrera, S., & Toonkel, J. (2022). Amazon, in broad costcutting review, weighs changes at alexa and other unprofitable units. The Wall Street. Journal (11.10.2022). Accessed 22 Dec 2022. McTear, M., Callejas, Z., & Griol, D. (2016). The dawn of the conversational interface. In The conversational interface. Springer (pp. 11–24). Nass, C., Steuer, J., & Tauber, E. R. (1994). Computers are social actors. Human Factors in Computing Systems, pp. 72–78. Nguyen, T. H., Waizenegger, L., & Techatassanasoontorn, A. A. (2021). “don’t neglect the user!” – identifying types of human-chatbot interactions and their associated characteristics. Information Systems Frontiers.https://doi.org/10.1007/ s10796-021-10212-x. Norris, M., & Oppenheim, C. (2007). Comparing alternatives to the web of science for coverage of the social sciences’ literature. Journal of informetrics,1(2), 161–169. Perianes-Rodriguez, A., Waltman, L., & van Eck, N. J. (2016). Constructing bibliometric networks: a comparison between full and fractional counting. Journal of informetrics,10(4), 1178– 1195. Porra, J., Lacity, M., & Parks, M. S. (2020). Can computer based human-likeness endanger humanness? – a philosophical and ethical perspective on digital assistants expressing feelings they can’t have. Information Systems Frontiers,22(3), 533–547. https://doi.org/10.1007/s10796-019-09969-z. Poser, M., K¨ustermann, G. C., Tavanapour, N., & Bittner, E. A. C. (2022). Design and evaluation of a conversational agent 751
Information Systems Frontiers (2024) 26:729–754 for facilitating idea generation in organizational innovation processes. Information Systems Frontiers.https://doi.org/10.1007/ s10796-022-10265-6. Purington, A., Taft, J. G., Sannon, S., Bazarova, N. N., & Taylor, S. H. (2017). Alexa is my new bff social roles, user satisfaction, and personification of the amazon echo. Proceedings of the 2017 CHI conference extended abstracts on human factors in computing systems, pp. 2853–2859. Qiu, L., & Benbasat, I. (2009). Evaluating anthropomorphic product recommendation agents: a social relationship perspective to designing information systems. Journal of Management Information Systems,25(4), 145–182. Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with CLIP Latents. 2204.06125. Rethlefsen, M. L., Kirtley, S., Waffenschmidt, S., Ayala, A. P., Moher, D., Page, M. J., & Koffel, J. B. (2021). Prisma-s: an extension to the prisma statement for reporting literature searches in systematic reviews. Systematic reviews,10(1), 1–19. Revang, M., Mullen, A., & Elliot, B. (2022). Magic quadrant for enterprise conversational ai platforms. Gartner Reports. Accessed 09 Jul 2022. Rheu, M., Shin, J. Y., Peng, W., & Huh-Yoo, J. (2021). Systematic review: trust-building factors and implications for conversational agent design. International Journal of Human–Computer Interaction,37(1), 81–96. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2021). High-resolution image synthesis with latent diffusion models. 2112.10752. Rzepka, C., & Berger, B. (2018). User interaction with ai-enabled systems: a systematic review of is research. International Conference on Information Systems. Rzepka, C., Berger, B., & Hess, T. (2021). Voice assistant vs. chatbot – examining the fit between conversational agents’ interaction modalities and information search tasks. Information Systems Frontiers.https://doi.org/10.1007/s10796-021-10226-5. Schmitt, A., Wambsganss, T., Janson, A., & S¨ ollner, M. (2021). Towards a trust reliance paradox ? exploring the gap between perceived trust in and reliance on algorithmic advice. In Fortysecond international conference on information systems, Austin, Texas (pp. 1–17). Schmitt, A., Zierau, N., Janson, A., & Leimeister, J. M. (2021). Voice as a contemporary frontier of interaction design. European Conference on Information Systems. Schwede, M., Zierau, N., Janson, A., Hammerschmidt, M., & Leimeister, J. M. (2022). I will follow you! - how recommendation modality impacts processing fluency and purchase intention. International Conference on Information Systems. Seaborn, K., Miyake, N. P., Pennefather, P., & Otake-Matsuura, M. (2022). Voice in human–agent interaction. ACM Computing Surveys,54(4), 1–43. https://doi.org/10.1145/3386867. Sedik, A., Maleh, Y., Banby, G. M. E., Khalaf, A. A. M., El- Samie, F. E. A., Gupta, B. B., Psannis, K., & El-Latif, A. A. A. (2022). Ai-enabled digital forgery analysis and crucial interactions monitoring in smart communities, technological forecasting and social change. IEEE Transactions on Antennas and Propagation, vol. 177. Shang, L., Lu, Z., & Li, H. (2015). Neural responding machine for short-text conversation. arXiv:1503.02364. Short, J., Williams, E., & Christie, B. (1976). The social psychology of telecommunications. London: Wiley. Singh, V. K., Singh, P., Karmakar, M., Leta, J., & Mayr, P. (2021). The journal coverage of web of science, scopus and dimensions: a comparative analysis. Scientometrics,126(6), 5113–5142. Smutny, P., & Schreiberova, P. (2020). Chatbots for learning: a review of educational chatbots for the facebook messenger. Computers & Education,151, 103862. Stieglitz, S., Mirbabaie, M., M¨ ollmann, N. R. J., & Rzyski, J. (2021). Collaborating with virtual assistants in organizations: analyzing social loafing tendencies and responsibility attribution. Information Systems Frontiers, pp. 1–26. https://doi.org/10.1007/ s10796-021-10201-0. Stokel-Walker, C. (2022). Ai bot chatgpt writes smart essays - should professors worry? Nature. https://doi.org/10.1038/ d41586-022-04397-7. Sun, J., Liao, Q. V., Muller, M., Agarwal, M., Houde, S., Talamadupula, K., & Weisz, J. D. (2022). Investigating explainability of generative ai for code through scenario-based design. In 27th International conference on intelligent user interfaces. ACM digital library. Association for computing machinery (pp. 212–228). https://doi.org/10.1145/3490099.3511119. Suta, P., Lan, X., Wu, B., Mongkolnam, P., & Chan, J. H. (2020). An overview of machine learning in chatbots. International Journal of Mechanical Engineering and Robotics Research, pp. 502–510. https://doi.org/10.18178/ijmerr.9.4.502-510. Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., & Stojnic, R. (2022). Galactica: a large language model for science. arXiv:2211.09085. Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Yu, D., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., ...Chi, E. (2022). Le Quoc: LaMDA: language models for dialog applications. arXiv:2201.08239. Trinidad, M., Ruiz, M., & Calder´ on, A. (2021). A bibliometric analysis of gamification research. IEEE Access,9, 46505–46544. Van Pinxteren, M. M. E., Pluymaekers, M., & Lemmink, J. G. (2020). Human-like communication in conversational agents: a literature review and research agenda. Journal of Service Management. Vertegaal, R., Slagter, R., Van Der Veer, G., & Nijholt, A. (2001). Eye gaze patterns in conversations: there is more to conversational agents than meets the eyes. SIGCHI Conference on Human Factors in Computing Systems, pp. 301–308. V¨ ossing, M., K¨uhl, N., Lind, M., & Satzger, G. (2022). Designing transparency for effective human-ai collaboration. Information Systems Frontiers.https://doi.org/10.1007/s10796-022-10284-3. Wahde, M., & Virgolin, M. (2022). Conversational agents: theory and applications. arXiv:2202.03164.https://doi.org/10.1142/ 9789811246050. Wallace, R. S. (2009). The anatomy of alice. In Parsing the turing test. Springer (pp. 181–210). Wambsganss, T., Kueng, T., Soellner, M., & Leimeister, J. M. (2021). Arguetutor: an adaptive dialog-based learning system for argumentation skills. In Proceedings of the 2021 CHI conference on human factors in computing systems. Association for computing machinery.https://doi.org/10.1145/3411764.3445781. Weizenbaum, J. (1966). Eliza—a computer program for the study of natural language communication between man and machine. Communications of the ACM,9(1), 36–45. Winkler, R., & S¨ ollner, M. (2018). Unleashing the potential of chatbots in education: a state-of-the-art analysis. Annual Academy of Management Meeting. Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. 752