Personal Medication Advisor
Full text
F ACULDADE DE E NGENHARIA DA U NIVERSIDADE DO P ORTO Personal Medication Advisor Joana Polónia Lobo M ASTER IN B IOENGINEERING Supervisor at FhP: Liliana Ferrreira (PhD) Supervisor at FEUP: Aníbal Ferreira (PhD) July, 2013
© Joana Polónia Lobo, 2013 Personal Medication Advisor Joana Polónia Lobo Master in Bioengineering Approved in oral examination by the committee: Chair: Artur Cardoso (PhD) External Examiner: Fernando Perdigão (PhD) Supervisor at FhP: Liliana Ferrreira (PhD) Supervisor at FEUP: Aníbal Ferreira (PhD) ____________________________________________________ July, 2013
Abstract Healthcare faces new challenges with automated medical support systems and artificial intelligence applications for personal care. The incidence of chronic diseases is increasing and monitoring patients in a home environment is inevitable. Heart failure is a chronic syndrome and a leading cause of death and hospital readmission on developed countries. As it has no cure, patients will have to follow strict medication plans for the rest of their lives. Non-compliance with prescribed medication regimens is a major concern, especially among older people. Thus, conversational systems will be very helpful in managing medication and increasing adherence. The Personal Medication Advisor aimed to be a conversational assistant, capable of interacting with the user through spoken natural language to help him manage information about his prescribed medicines. This patient centred approach to personal healthcare will improve treatment quality and efficacy. System architecture encompassed the development of three modules: a language parser, a dialog manager and a language generator, integrated with already existing tools for speech recognition and synthesis. All these modules work together and interact with the user through an Android application. System evaluation was performed through a usability test to assess feasibility, coherence and naturalness of the Personal Medication Advisor. It was concluded that the developed assistant could successfully make users achieve their objectives. In addition, the system was evaluated as easy to use, coherent and with a pleasant interface. Further improvements can still be done to extend the flexibility and convenience of the Personal Medication Advisor so that an easy integration on senior patient’s daily life can be achieved.
Resumo Os cuidados de Saúde estão a ser revolucionados pela introdução de sistemas automáticos de suporte à decisão médica e outras aplicações de inteligência artificial. A incidência de doenças crónicas está a aumentar e torna-se inevitável monitorizar os pacientes em casa. A insuficiência cardíaca é uma síndrome crónica, sendo a principal causa de morte e readmissão hospitalar nos países desenvolvidos. Como não tem cura, os pacientes terão de seguir exigentes regimes de medição para o resto das suas vidas. O não cumprimento dos planos de medicação receitados é também uma preocupação, especialmente entre pacientes mais idosos. Assim, assistentes conversacionais poder-se-ão tornar muito úteis para gerir a medicação e aumentar a adesão aos planos prescritos. O Assistente Pessoal de Medicação procura ser um assistente conversacional capaz de interagir através de linguagem natural falada com o utilizador para o ajudar a gerir os seus medicamentos. Esta visão centrada no paciente pretende aumentar a eficácia e qualidade dos cuidados de Saúde. A arquitetura do sistema passou pelo desenvolvimento de três módulos: um processador de linguagem, um gestor de diálogo e um gerador de linguagem, integrados com ferramentas já existentes para reconhecimento e síntese de voz. Todos estes módulos atuam em conjunto e interagem com o utilizador através duma aplicação para Android. A avaliação do sistema recorreu a um teste de usabilidade para qualificar a performance, coerência e naturalidade do Assistente Pessoal de Medicação. O assistente desenvolvido conseguiu fazer com que os utilizadores atingissem os objetivos pretendidos. A aplicação foi também classificada como fácil de usar, coerente e com uma interface agradável. Futuras melhorias no sistema devem ser realizadas para aumentar a flexibilidade e conveniência do Assistente Pessoal de Medicação. Deste modo, será conseguida uma melhor e mais facilitada integração do assistente no dia-a-dia de pacientes idosos.
Acknowledgments My thanks to Dr. Lilliana Ferreira, supervisor at Fraunhofer Portugal, and Prof. Aníbal Ferreira, supervisor at FEUP, for providing the required support and guidance during this project. To Ana Barros for her good advices and to Bruno Aguiar for letting me use his PharmInx ontology. To Fraunhofer Portugal AICOS for the opportunity of developing this project and to those who volunteered for testing the Personal Medication Advisor or simply gave their suggestions for the project. A special thanks to all my friends and family. To my mom and boyfriend who support me daily: “Gratitude is the memory of the heart.” Joana Polónia Lobo
List of Tables Table 2.1: Evidence-based doses of effective pharmacological drugs used for HF-REF (European Society of Cardiology, 2012). ........................................................................................................... 12 Table A. I: Aetiology of heart failure. Source: (European Society of Cardiology, 2012) ......................... 57 Table A.II: Participant information and PMA performance data for the two tasks executed during the usability testing. ..................................................................................................................................... 60 Table A.III: Average scores and standard deviations for each of the 16 statements of the usability questionnaire with the system property they evaluate. .......................................................... 60
Abbreviations Abbreviation list (by alphabetic order) ACE Angiotensin-converting enzyme ADT Android development tools ARB Angiotensin receptor blocker ASR Automated speech recognition API Application programmatic interface CHF Chronic heart failure EF Ejection fraction ESC European society of Cardiology HF Heart failure HF-PEF Heart failure with preserved ejection fraction HF-REF Heart failure with reduced ejection fraction H-ISDN Hydralazine and isosorbide dinitrate JSGF Java speech grammar format LKB Linguistic knowledge database MEMS Medication event monitoring system MRA Mineralocorticoid receptor antagonist MTT Meaning text theory NLG Natural language generation NLU Natural language understanding OWL Web ontology language PMA Personal medication advisor PUFA 3-n polyunsaturated fatty acid RDF Resource description framework SPARQL SPARQL protocol and RDF query language SUS System usability scale TTS Text-to-speech VPL Voice prescription labels
3 Chapter 1 Introduction Health care technologies seem to be evolving towards interactive and personal health management assistants. Interaction with conversational assistants should be seen as an ongoing dialogue between user and system. Therefore, user’s intentions and goals must be understood by the system so that proactive help can be provided. Such technologies require a high level of system architecture as well as a complex coordination between its different components. This work intended to take another step on the direction of a Personal Medication Advisor (PMA), a conversational system capable of interacting with the user through natural language, mostly voice, to help him managing his prescribed medication regimens. This development was placed under the scope of a Dissertation project for the Integrated Master in Bioengineering at Faculty of Engineering of University of Porto and in association with Fraunhofer Portugal Research Center for Assistive Information and Communication Solutions (AICOS). This chapter firstly introduces the context, scope and needs behind conversational assistants for healthcare. Then, the motivation and objectives that supported the Personal Medication Advisor project are presented, as well as the requirements envisioned for the system and the anticipated benefits and innovative contributions of this research. Finally, the structure of this thesis is outlined. 1.1 The problem As modern Medicine is developing, a large amount of pharmacological drugs continues to appear. The amount of money annually spent on medicines is growing exponentially (Sullivan, Behncke, & Purushotham, 2010). With the existence of so many drugs and possible medication regimens, Pharmacology brings a large amount of information to be managed by healthcare staff and patients.
Introduction 4 While developed world population is aging, there are more demanding health care requirements so that elders can have quality of life. As age increases, the number of patients suffering from chronic illnesses is also raising (Giles, 2004). Heart failure (HF) is a common chronic disease with high prevalence and significant health expenditure due to frequent patient re-hospitalizations (Michalsen, König, & Thimme, 1998). Such diseases usually have numerous pharmacological prescriptions that require strict medication management. With complex medication regimens, patients spend a lot of effort gathering the information required to successfully manage their disease. Most of them even have trouble keeping track of this information. (Ferguson, et al., 2010) Due to unaffordable costs and insufficient medical personnel, healthcare staff cannot offer daily home-assistance. Therefore, there is a need of automated systems for health communication so that intelligence assistance can be provided to patients within a home health environment. Compliance is an important predictor of outcome among patients with chronic diseases like HF. As slight improvements in functional capacity are of the utmost importance to these patients, prescribed treatments should be taken seriously. Studies have shown that, in general, medication adherence is remarkably lower than self-reported adherence (Nieuwenhuis, Jaarsma, van Veldhuisen, & van der Wal, 2012). With this, it is also essential to maintain patient motivation and promote adherence to the pharmacological regimens in chronic patients. 1.2 Motivation and objectives On the last decade, healthcare has arrived to the next stage of evolution with automated medical support systems and personal care applications (Topol, 2010) (Haux, 2006). The role of this work was to encompass this next technological step with a Personal Medication Advisor in harmony with Fraunhofer Portugal AICOS mission of Information and Communication Technologies for ‘care, well-being and inclusion’. Numerous pharmacological treatments rely on strict medication programs that should be complemented with lifestyle measures like dietary changes and appropriate exercise protocols. Therefore, conversational medical assistance may be very helpful, especially with chronic patients and elder healthcare. This work aimed to provide information and knowledge-based advice to help the user managing complex prescription regimens. Personalized intelligence assistance on posology, interactions, indications and adverse reactions will allow in-home medication monitoring. Since many patients do not recognize the value of strictly following medication regimens (Nieuwenhuis, Jaarsma, van Veldhuisen, & van der Wal, 2012), another objective includes motivating the user through interaction and dialogue, trying to increase medication adherence. The scope of the conversational assistant developed was restricted to medication regimens for heart failure. This syndrome was chosen due to its well defined pharmacological practices which greatly depend on self-care and self-management to lengthen patient survival. Since there is no cure for HF, patients will likely have to follow a medication plan for the rest of their lives.
Introduction 5 Although expensive, existing HF home-monitoring programs have proved to be very helpful (Simon & Horowitz, 2002). A personal health assistant to interactively manage medication and prescriptions would help expanding monitoring access for HF patients by reducing healthcare costs, re-hospitalizations and number of face-to-face consultations. Dialogue systems for health communication generally aim to provide intelligence assistance to patients through spoken natural language conversation within a home health environment. Looking for this final achievement, the main goal of this project was to develop a modular system that would understand what the patient says and produce valid sentences, on a context of heart failure medication regimens, answering the user concerns. The integration of language understanding and generator modules with existing voice recognition and synthesis systems allows spoken interaction with the user. To complement this communication, information through images or text can also be available, but never as the main focus of the system-user interaction. 1.3 Requirements As an agent-based spoken dialogue system, the Personal Medication Advisor mission tried to accomplish a ‘Remarkable Technology, Easy to Use’ motto. Under this thought, the PMA was built on the paradigms of consistency and simplicity where the generated sentences are intended to be short, direct and easy to understand. The goal was to make the system reliable, intuitive and portable. The system should not be too constraining, being at the same time easy to use. Because a decisive factor for adoption of technology is convenience (Jimison, et al., 2008), the system tried to require no special training for its utilization. The ultimate delivery to the consumers should be made using an everyday platform like the Android technology. Real-time generation is essential on a dialogue framework. With this requirement in mind, all the system modules had to be reasonably fast so that the PMA can be used on a real healthcare environment. Moreover, the generation had to be psycho-linguistically realistic and stylistically appropriated. Generated phrases must give useful information and advice on prescribed medication posology for heart failure scenarios. The outputted speech must be wellformed, using formal but short and direct sentences. Taking these requirements in consideration, the medication assistant is settled to provide information on: • Posology - Information on drug administration methods and quantities. Subcategories of posology data include administration dosage, route, frequency, drug name and time references. This information on the drug-posology concept was classified as the most important and it should be incorporated during the first step of system development. • Drug indications - Specification of common targets and uses for a certain pharmacological drug.
Introduction 6 • Adverse reactions - Information on the side-effects and harms of the usage of a pharmacological drug under the recommended administration conditions. A later and advanced approach to the Personal Medication Assistant should also manage contraindications, cautions and drug to drug interactions. The Personal Medication Advisor was developed in European Portuguese language, opening an unexplored path on the integration of intelligence assistance in Portuguese healthcare. 1.4 Benefits The Personal Medication Advisor partially shares its beliefs with ‘Information therapy’. This methodology defends that providing specific evidence-based medical information to certain patients at the right time can help them making positive behaviour changes and improve self-consciousness of healthcare (Mettler & Kemper, 2003). The patient centred approach to health services seems to be a key to improve disease outcomes. Tailored health information and personalized health education are thus essential for effective healthcare (Jones, Cawsey, & Bentala, 1999). It is expected that the conversational assistant developed can bring together the advantages of information therapies and automated technologies, improving the effectiveness of health services, especially for people with chronic diseases or advanced age. Since chronic diseases, like HF, require strict and complex medications regimens, they are usually a potential menace to patient life and it was suggested that using automated intelligence assistance to manage medication could help in patient survival (Ferguson, et al., 2010). Therefore, it is anticipated that efficiency of healthcare services is enhanced by increasing patient survival, reducing re-hospitalization and mortality rates, diminishing complications and allowing a better responsiveness and distribution of work time for healthcare staff. Patients will also be able to live more independently with greater compliance and satisfaction with therapeutic regimens. The contribution of new technological solutions will surely lead to a pro-active attitude towards disease monitoring and management. 1.5 Outline After presenting the aim, context and intents of the Personal Medication Advisor project, this thesis continues with Chapter 2, containing some background on heart failure, healthcare monitoring and conversational systems. The information presented is needed so that project context and design can be fully understood. The subsequent chapter, Chapter 3, identifies, characterizes and justifies the resources and tools used, opening the path to Chapter 4, where the PMA architecture, along with each module and feature, is explained in detail. Consequently, the evaluation performed and a discussion around the final prototype of PMA can be seen in Chapter 5, with special focus on the difficulties and challenges encountered. Finally, Chapter 6
Introduction 7 presents a final balance of the five months of project development, concluding on what was achieved and what could still be done in the future.
Background and Literature Review 14 and motivating appropriate lifestyle behaviours for the patient. Even so, the lack of understanding of the disease and treatment as well as the cost of assistive technologies is the most significant barrier to the use of assistant technologies. 2.2.2 Natural Language Generation Tools Tailored hypermedia information can be easily consulted by patients at home. These systems use artificial intelligence and natural language generation (NLG) techniques to generate personalized media with hypertext containing basic factual information about a certain patient medical record, disease and treatment. Depending on the method and the disease, the patient can have simple and cheap access to the information through different media being a leaflet, a webpage, an online questionnaire, an interactive application, etc. Among many others, PIGLIT (Patient Information Generated by Loosely Intelligent Techniques) (Binsted & Jones, 1995), HealthDoc (Hirst, DiMarco, Hovy, & Parsons, 1997), OPADE (Carolis, et al., 1996) and Patient Advocate (Miksch, Cheng, & Hayes-Roth, 1996) are examples of these systems. The objective of these systems is to ensure that patients have the information required to understand and manage their condition. On the other hand, some effort must be spent to design the knowledge database with the general medical information that may be communicated to the patient. Another noticeable problem of these tools is that a recommended drug for a particular disease appears once in the knowledge database but it may be mentioned in textual documents using many different surface forms. Artificial intelligence and NLG tools can be used on healthcare technologies on a more complex level: the conversational assistants. The Personal Medication Advisor fits within this class of systems that is still on an early stage of growth. Already developed or in-development conversational systems for home-monitoring or other medical applications are discussed in section 2.3.5. 2.3 Conversational systems Conversational assistants are computerized systems that communicate and interact with human users to help solving problems. The conversational paradigm implies understanding a verbal input, through speech recognition and context understanding, as well as verbalizing a response, using language generation and speech synthesis (Seneff, 2003). This way the assistant can be engaged in dialogue with the user. A conversational assistant should focus on practical dialogues, meaning that both interlocutors are cooperating to perform a certain task and achieve a certain goal. Practical dialogues can comprise tasks from planning to management, information retrieval or advice-giving (Allen, Ferguson, & Stent, 2001). Different types of conversational assistance differ on the tasks to be performed, their complexity and the degree of initiative taken by both interlocutors. According to the balance of initiative on the conversation, it is possible to have directed, mixed-initiative or multimodal dialogues (Seneff, 2003). While
Background and Literature Review 15 directed dialogue systems are already successfully commercialized, multimodal systems include action interaction and are starting to be developed. Mixed-initiative dialogue assistants are used for specific applications, usually staying as research prototypes. A brief and generalized representation of the structure of a conversational assistant is represented on Figure 2.1. Conversational systems are complex combinations of different modules. A division of dialogue systems can be made according to the role of each group of components: interpretation, behaviour or generation (Ferguson, et al., 2006). Interpretation or natural language understanding (NLU) modules capture and understand what the user said. Behaviour or dialogue management components plan system actions by managing goals, intentions and reasoning. Generation modules construct the dialogue through content planning and surface realization followed by speech output. When trying to assemble conversational technologies, there are four developing stages where attention should be focused, namely speech recognition, language understanding, language generation and speech synthesis (Cole, et al., 1996). Figure 2.1: Schematics of a typical conversational system with the different functional modules. Source: (Seneff, 2003) 2.3.1. Speech Recognition Automated speech recognition (ASR) components capture and process spoken input by converting the acoustic signal to a set of words. Recognized words will then be used as input for analysis. Most ASR systems use statistical methods, neural networks and/or Hidden Markov Models for acoustic modelling (Jelinek, 1997). An ASR system usually consists of front-end, an acoustic model, a language model and a decoder (Figure 2.2). The system will depend on the intended type of recognition that can be, for example, speaker dependent or independent,
Background and Literature Review 16 continuous or by isolated words. Spontaneous natural speech and its intrinsic variability can bring some difficulties to accurately recognize the speech. Variability in ASR lies on phonetic and acoustic dissimilarities as well as intra-speaker characteristics (age, emotional state, health, fluency or socio-linguistic background) (Benzeghiba, et al., 2007). Spoken input also differs significantly due to the false starts and filled pauses that characterize a natural speech. Even so, speech recognition systems have become much more robust in recent years with respect to both speaker and acoustical variability (Cole, et al., 1996). The system should be attentive for ungrammatical constructs, recognizing errors and modelling non-speech events and disfluencies. An extensive lookup on the available tools for recognition is made in Chapter 3. Figure 2.2: Overview of the structure of a typical ASR system. Source: (Neto N. , 2011) 2.3.2. Language Understanding Between recognition and generation there is a NLU module to analyse and understand the acquired speech utterances. Encoded linguistic knowledge, like lexicon, grammar and semantics, helps producing a meaning representation or logical form through parsing and knowledge representation. The representation of the information acquired using adaptive models is of great importance, allowing it to be available afterwards for other applications (Cole, et al., 1996). Parsing establishes the syntactic organization as well as the semantic content of what was said. This is a hard task because subtle differences in phrasing and typical discourse constructions like ellipsis and pronominal or verbal referencing can lead to different interpretations. Ambiguities and ill-constructed sentences should be avoided by the system. Dialogue management tools are then responsible for clarification sub-dialogues, ambiguity dissolution and pragmatic considerations of what was said. The dialogue manager is the central unit for these tasks, generating and evaluating the hypotheses to the intended meaning of the user speech. After choosing the most possible intention of the user, the information is passed on to the next module presenting a hierarchy of active or pending goals and sub-goals. If the intention was not understood, the system should question the user. The dialogue manager
Background and Literature Review 17 should also be responsible to update dialogue context on each iteration so that the system can always be well fitted in the conversation (Traum & Larsson, 2003). 2.3.3. Language Generation Natural language generation (NLG) refers to the process of automatically converting an abstract meaning representation into a string (Baptist, 2000). NLG aims to produce high-quality natural language text from computer internal representations of information. 2.3.3.1 Processes There are several different approaches to language generation. Generation processes can be mainly categorized as non-linguistic (being statistical, canned-text or template-based systems) or as linguistic (being phrase-based, MTT-based or feature-based systems) (Baptist, 2000). The shallow simplicity of non-linguistic generators opposes to the sophistication of linguistic approaches that require much more planning and development. In general, linguistic methods are more robust and present high quality output. Linguistic generators entail simultaneous computational and linguistic knowledge relying on linguistic knowledge databases (LKB), this is, sets of files comprising grammar rules, lexical entries and rewrite rules, which specify the generation domain. A statistical NLG system depends on training and probabilistic models that can be more or less complex (Ratnaparkhi, 2000). Although these systems are fast, they lack in accuracy and flexibility. Canned-text generators are the simplest but most wasteful systems. They only produce a static string of words without any amend, adequate for periodic event-specific generation (Cole, et al., 1996). More sophisticated, template-based systems can introduce minor alterations on generated sentences while a major part of the text is still static. Unlike previous approaches, they can perform multi-sentence generation if the text has a rather constrained structure (Reiter, NLG vs. templates, 1995). Phrase-based systems focus on text planning by using generalized templates and replacing phrasal patterns by selected words. Their robustness works at sentence or discourse level (Hovy, Planning coherent multisentencial text, 1988). However, their weakness lies on the shortage of text plan libraries and on the complex phrasal interrelationships present at longer discourses, not fitting real scenarios. Meaning-text theory (MTT) based generators sequentially map semantic, syntactic, morphological and phonetic levels of representation according to a set of rules. The description of these relationships is made using lexical functions (Goldberg, Driedger, & Kittredge, 1994). With feature-based systems, each possible alternative of a language expression is analysed and categorized in a feature. As consequence, each sentence is characterized by a particular group of features. The final sentence is determined by incrementally match each inputted phrasal segment to the suitable feature. For these systems, the interrelationships are still difficult to manage and the number of features that can be
Background and Literature Review 18 identified must be controlled. There are no multi-sentence featured-based generators (Hovy, Language generation, 1996). NLG was effusively studied between the late 1980s through the 1990s. Initially, simpler approaches like statistical or canned-text systems were preferred. However, over the years, there was acceptance of template-based and, later, of linguistic generators (Baptist, 2000). Integration of template-based and linguistic approaches would be suggested by Ehud Reiter (Reiter, NLG vs. templates, 1995), in 1995, opening a path to hybrid generation systems. Within a hybrid framework, encompassing both linguistic and non-linguistic aspects may present interesting advantages, allowing either to rapidly develop tools for simple domains or to use expert knowledge bases for more challenging domains (Baptist, 2000). The arrival of machine learning techniques after the mid-1990s has allowed further studies on linguistic approaches, mainly on sentence planning (Walker, Rambow, & Rogati, 2001), lexical selection (Bangalore & Rambow, Corpus-based lexical choice in natural language generation, 2000), word order (Bangalore & Rambow, Exploiting a probabilistic hierarchical model for generation, 2000) and grammatical relations (Corston-Oliver, 2000). Stochastic approaches and machine translation are bringing many advances to NLG (Chambers & Allen, 2004). With this, new hybrid methodologies appeared, namely the corpus-based stochastic generation where linguistic models are combined with statistical models derived from annotated corpora (Oh & Rudnicky, 2002). 2.3.3.2 Tasks A modular approach to NLG is helpful so the developer can first be focused in managing and testing each of the three central tasks that comprise the generation process (Hovy, Language generation, 1996). Figure 2.3: Overview of the tasks of a generation module. Source: (Dale, 1995)
Background and Literature Review 19 Then, all the task responsible modules should be assembled and their communication assessed. Different models for the NLG system architecture and interconnection can be found (one example is presented on Figure 2.3). Although not completely independent, each generation task is briefly presented. • Text planning or content determination This first step defines global text structure, deciding what to say. The inputted information is mapped, the text content is established and the information organized accordingly. There are several approaches to text planning as the traditional top-down planners using Rhetorical Structure Theory (Hovy, Planning coherent multisentencial text, 1988) or intentional operators (Moore & Paris, 1993). • Sentence planning The information is structured and ordered into sentences and paragraphs using appropriate reference words as conjunctions and pronouns. Common tasks include pronoun specification, focus signalling, content aggregation to remove redundancies and preliminary phrase ordering. Lexicalization, the transition from concepts to lexical representations, can be included within this task. The lexical choice includes collocations and definition of lexical sub relations (Busemann, 1993). This module is thus seen as a mediator between deep generation (text planning) and surface realization because it limits simultaneously what can be said and how it can be said. Even so, a clear division between this and the other NLG modules is not always very visible (Smedt, Horacek, & Zock, 1996), • Surface realization or surface generation Using grammar knowledge, coherent and grammatically correct sentences are produced in formatted output text. The realization of the planned expressions, determines how to say the information according to morphological and phonological rules. Computational models try to mimic human utterances to approach the output to natural speech. A common approach is the stochastic surface realization: language models are built, candidate utterances are generated and scored, the best utterance is selected and finally the missing slots are filled with the relevant input information (Oh & Rudnicky, 2002). 2.3.3.3 Demands of language generation Developing a NLG system requires good structured lexicons and grammars in the language(s) that the system will use; knowledge bases oriented to the domain-specific application that is being developed and consolidated text planning libraries. Appropriate text-tospeech and speech-to-text systems would be necessary to help integrating the NLG module on a conversational system. Naturalistic speech generation with the inherent disfluencies, interruptions or overlaps is still distant due to the unsolved difficulties related to discourse and interpersonal speech variability.
Background and Literature Review 20 Until now, only general purpose single-sentence or limited purpose multi-sentence generators were developed (Hovy, Language generation, 1996). A multimodal general purpose language generator has not been build due to the unsolved problems that appear when using large lexicons, unclear contexts, unconventional discourse structures, large-sized domains or systems with great expressive potential that can say the same thing in many ways. In this way, there is still a lot of work to be done on developing generators with the desirable coherence and consistency, producing texts of several paragraphs functional on real-world domains. Challenges arise when trying to build multilanguage systems. As the major amount of NLG research is done using English language, generation in Portuguese is still poorly explored. Furthermore, European Portuguese has a complex verbal structure with auxiliary verbs (Baptista, Mamede, & Gomes, 2010). For now, more literature can be found on Portuguese language recognition, parsing and processing (Meinedo, Caseiro, Neto, & Trancoso, 2003) (Silva, Branco, Castro, & Reis, 2010) rather than on Portuguese NLG. 2.3.4. Speech Synthesis Speech synthesis is the automatic generation of speech waveforms. Spoken output of NLG sentences can be produced by text-to-speech (TTS) systems (Figure 2.4). The first speech synthesis systems were hardware-based. With the increase of computational power, today’s synthesizers are mostly software-based, allowing more functionalities, more flexibility and frequent updates (Lemmetty, 1999). In the majority of TTS systems, the first step is text analysis, or high-level synthesis, where a grapheme-to-phoneme conversion is done. Character strings are pre-processed and converted into phonetic representations. These representations are typically a string of phonemes that can also contain information about phasing and intonation. On a second phase, called low-level synthesis, phonemic representations are converted to waveforms that can be outputted as sound. The majority of actual systems are composed of both high and low synthesis components. Figure 2.4: Overview of the structure of a typical TTS system. Source: (Vasavi & Sravya, 2010) There are several methodologies for generating synthetic speech waveforms: from concatenative to articulatory, formant or Hidden Markov Model based synthesis. (Lemmetty,
Background and Literature Review 21 1999) Each TTS system uses one of these methods according to its own strengths and weaknesses and the intended use of the system. Concatenative TTS systems require prerecorded units and sound libraries based on human speech recordings. With this, they provide a synthesized speech with a more natural sound. This naturalness implies the transcription and alignment of large corpora, what can be time-consuming and demand a profound knowledge of phonetics to accurately align temporally the transcription labels (Carvalho, Oliveira, Viana, & Trancoso, 1998). Formant TTS systems are rule-based requiring less resources. Although the voice quality is lower and less natural, the speech is clear and many languages can be provided with a small size system. This may also make these systems faster than concatenative TTS. Modern speech synthesis technologies involve complex methods and algorithms. Among recent approaches to speech synthesis, Hidden Markov Models and Neural networks based methodologies are very used (Schroeder, 1993). TTS engines also differ with different languages, dialects and specialized vocabularies. High quality speech synthesis systems should maximize naturalness and intelligibility (Lemmetty, 1999) (Thakur, Chettr, & Shah, 2012). The current speech synthesis systems can already artificially produce highly intelligible human speech. However, achieving an optimal naturalness with synthesized speech can still be a problem (Cole, et al., 1996). An overview of existent TTS tools is available in Chapter 3. 2.3.5. Current Conversational Systems A majority of the existing conversational assistants is available for profitable applications. They include flight booking assistants, weather informers, restaurant or hotel guides, calendar managers, city navigators, traffic reporters, sports updaters, crisis supervisors, among others (Seneff, 2003). After several years of development in MIT, Genesis II is an example of a hybrid conversational assistant presenting a framework for several languages and several domain applications: from JUPITER in weather domain to MERCURY in flight reservation, ORION in off-line task delegation or VOYAGER in navigation assistance (Baptist, 2000). Amalgam was developed by Microsoft focusing on sentence realization and using machine-learned and knowledge-engineered modules in machine translation contexts (Gamon, Ringger, & CorstonOliver, 2002). Other systems include the IBM Flight Information System, (Axelrod, 2000) the CMU’s statistical generator (Jin, 2003) and a collaborative multimodal system for rescue emergency scenarios TRIPS (Allen, et al., 2001). These systems are usually based in English, German, Japanese and/or French Language generation. Research on European Portuguese NLG and conversational systems is still limited. Meteo (Cardoso, Flores, Langlois, & Neto, 2002), a telephone-based conversational system for weather forecast and the spoken dialogue virtual butler of the Interactive Home of the Future project (Neto & Cassaca, 2004) are examples of the few conversational systems developed in European Portuguese. Medicine related systems started to appear on the last decade. An example of surgery applications for NLG and intelligent assistance is GALEN-IN-USE system for the collaborative construction and maintenance of surgical procedure classifications (Wagner, Rogers, Baud, &
Background and Literature Review 22 Scherrer, 1999). Conversational assistants using NLG for healthcare are still in development like the Chester (Ferguson, et al., 2006) and CARDIAC (Galescu, Allen, Ferguson, Swift, & Quinn, 2009) systems. The CARDIAC is an English spoken dialogue system that conducts health monitoring interviews using natural language for chronic heart failure patient healthcare. Rather than medication assistance, it collects information about the current state of patient’s health from objective (weight) to subjective (pain, fatigue, exercise) conditions. CARDIAC can be useful for patient self-care, for monitoring trends and generating reports for healthcare providers. In a way, it shares part of its motivations with the Personal Medication Advisor. New and more complex approaches have appeared, using NLG and speech along with sensing, acting and gesture recognition to assemble multimode robot systems in what is called an Embodied or Situated Conversational Assistant (Mavridis, 2007). They include robotic lab assistants, household helpers and affective robotic computers. They may portray the future of conversational assistant technologies. 2.4 Summary Over this chapter, the motivation and problem behind this project was further contextualized by characterizing the world-wide Heart Failure healthcare scenario. Heart Failure is a disease with increasing incidence associated with high hospitalization costs, high mortality and morbidity. These patients need to follow strict medication regimens, composed of ACE inhibitors, ARBs, MRAs and beta-blockers, to reassure their survival and reduce rehospitalisation. Daily home-monitoring and medication compliance are extremely important for HF patients. There are several home-monitoring tools they can use. Conversational assistants can be encompassed within a HF home-healthcare environment. The main tasks of a conversational system are to understand what the user says, to plan an appropriate response and to express a response to the user through generated natural language. Thus, these systems usually have a modular structure, with modules for speech recognition, natural language understanding, dialog management, natural language generation and speech synthesis.
23 Chapter 3 Tools and Resources This chapter identifies the resources and tools that were exploited along the different steps and iterations of the Personal Medication Advisor. Although some of these tools are not directly used by the final prototype, they played an important role during the development process. With this, development platforms are pointed up, followed by the resources and tools that were necessary to this project. From speech technologies to the linguistic tools used, an explanation of their characteristics, purpose and the rationale for their choice is given. 3.1. Development platforms During an initial phase, the medical assistant was developed for personal computers, using Java in the Eclipse environment. Java was chosen as the programming foundation of PMA because it is simple, portable, free-ware, allows easy integration of external packages and supports modular programming. Even so, the most important reason was to facilitate the integration of the system with Smartphone technologies. Consequently, at a later stage, PMA was rebuilt as an Android application, using Java and the Eclipse IDE bundled with Android Development Tools (ADT). 3.2. Resources 3.2.1 Ontologies The required medical information regarding medication posology, dosage, indications and adverse reactions, among others, was retrieved from an ontology. An ontology can be defined as ‘a formal, explicit specification of a shared conceptualization’ (Studer, Benjamins, & Fensel, 1998), this is, a model of a certain system that takes the form of a definition of properties of its
Tools and Resources 30 3.5.3 Combined systems There are systems were both TTS and ASR systems are embedded and combined: from the successful Google Now by Google or Siri by Apple to LumeVox speech engine by Lumevox or Loquendo TTS and ASR by Nuance. For example, Zanzibar openIVR by SpokenTech uses FreeTTS and Sphinx4 and goes to the encounter of the tools chosen above. Nevertheless, most of these products are only commercially available for user-end purposes and not for developers. Among them, iSpeech technologies (iSpeech, n.d.) have demo versions that are available for Java and mobile environments and were tested with PMA. The synthesis in European Portuguese has very high quality and is very accurate. Unfortunately, speech recognition on the demo version is limited and synthesized speech sometimes includes a copyright marking in Java developing environments. It was only during the latest phase of development, when PMA moved to the Android implementation, that the best arrangement of speech recognition and synthesis tools was found. Android speech recognition packages, Google Voice (Google, n.d.), were used as they worked better than iSpeech Android recognition. On the other way, iSpeech Android synthesis had more quality than Android speech synthesis and thus was preferred. 3.6. Summary The platforms essential to the development of the project were described in this chapter. In summary, the Eclipse IDE bundle with ADT was the final platform of choice. The resources used included a medical domain specific ontology, obtained through the systematic processing of Portuguese medical leaflets, as well as European Portuguese languages resources such as a dictionary and a domain specific lexicon. The tools used comprise Java frameworks for ontology modelling, visualization and questioning, in addition to language processing packages for part-of-speech tagging and phrase chunking. Among the several speech technologies that were assessed, there was a special focus in CMUSphinx and Google Voice for ASR, as well as FreeTTS, FurbSpeech and iSpeech for TTS.
31 Chapter 4 The Personal Medication Advisor The Personal Medication Advisor was envisioned to be developed within an iterative patient centred approach for personal healthcare. During the whole process of development, the system aimed for consistency, simplicity and full applicability on home health environments, namely for heart failure care. This chapter describes each step of implementation and development, along with an explanation of system architecture. First, a general overview of PMA’s features and performance is given. Then, the final architecture of PMA, along with descriptions of module tasks, is clarified. Main features and guidelines of system architecture are shared by the different PMA prototypes during the different stages of the project. For this reason, the distinction between the different prototypes and stages is only given afterwards. 4.1. General overview Before a detailed analysis of system architecture and of the different stages of development, a brief overview of system flow and functionalities of the system is needed. PMA engages the interaction with the user by generating a greeting according to the time of the day (morning, afternoon or night) and then waits for a user utterance. Due to the complexity of coherent and complete conversational systems, the range of possible user utterances had to be narrowed. Ideally, the user should be free to say whatever he wants independently of the words used or the syntactic structure chosen. However, the nature of this project does not allow PMA to have such a wide flexibility. With this, the user can interact with the system with somewhat rigid and limited sentences. The user should question the system using utterances similar to the following ones: • “Qual é a dose de <nome medicamento> receitada?” (What is the prescribed dosage of <medicine name>?)
The Personal Medication Advisor 32 • “Qual é a frequência de <nome medicamento> receitada?” (What is the prescribed frequency of <medicine name>?) • “Qual é a via para tomar <nome medicamento>?” (What is the route to take <medicine name>?) • “Existem restrições para tomar <nome medicamento>?” (Are there any restrictions to take <medicine name>?) • “Quais são as posologias existentes de <nome medicamento>?” (What are the posologies of <medicine name>?) • “Quais são as reações adversas de <nome medicamento>?” (What are the adverse reactions of <medicine name>?) • “Porque razão estou a tomar <nome medicamento>?” (Why am I taking <medicine name>?) • “Quais são os medicamentos que o médico receitou?” (What are the medicines that the doctor prescribed?) • “Posso tomar <nome medicamento>?” (Can I take <medicine name>?) These questions will be interpreted by the system and an appropriate answer will be given, making use of the information retrieved from PMA resources. For example, if the patients asks: “Quais são os medicamentos que o médico receitou?” (What are the medicines that the doctor prescribed?), PMA would list the names of the medicines that were on the prescription, along with a counting for the total number of medicines prescribed. When the user utterance is invalid or the system does not understand what was said, PMA will retaliate with a random set phrase that requests the user to repeat or reformulate the question. The system can also promptly react to some basic words present in a normal conversation: “sim” (yes), “não” (no), “ajuda-me” (help me), “112”, “adeus” (goodbye) or “obrigado/a” (thanks you). These interactions are not dependent of PMA resources as the previous questions. For example, “sim” (yes) or “não” (no) are useful for the user to answer some questions that PMA uses for confirmation or inquiry. As “ajuda-me” (help me) will trigger a brief explanation on the PMA system, “112” will make the system deal an emergency call to the national number of emergency 112. If the user wants to close the assistant, he should say “adeus” (goodbye). A few courtesy phrases are also triggered if “obrigado/a” (thank you) is said, giving a more natural feeling to the interaction. The PMA has access to user’s medical prescriptions. This way, information can be more reliable and tailored for each patient. The prescription must have the list of prescribed medicines along with the prescribed posology. For testing PMA prototypes, a medical prescription was simulated. Another feature of PMA includes the production of a medical
The Personal Medication Advisor 33 report, resulting from an inquiry to the user, that allows the healthcare staff that is responsible for that patient to be informed when the he does not feel well or has a critical HF symptom. A general dialogue task specification, in the form of the task tree presented in Figure 4.1, encodes the system's behaviour and allows a first insight into system architecture. Each system feature will be explained with more detail in the next section. Figure 4.1: Task tree of PMA. 4.2. Architecture The PMA was developed envisioning a modular architecture, this is, a top-down design. In this design strategy, the system consists of small and autonomous modules, each performing different functions but altogether working as single structure. This makes the system easier for debugging and for latter updating or modification. Having a modular structure is important for a conversational assistant, where different tasks are being carried and a specific module could be reused in other systems (Mellish & Evans, 2004). The three essential modules that comprise the PMA were established as the language parser, the dialog manager and the language generator. Both speech recognition and synthesis modules allow these crucial components to directly interact with the user, using pre-existent speech tools. With this, they are integrated with the remaining components, making a total of five modules that compose the PMA system. A schematics overview of the system is shown on Figure 4.2. Initially this work was meant to have a special focus on the development of the language generation module. However, functional NLU and managing modules are indispensable for a system that tries to come nearer to a conversational system.
The Personal Medication Advisor 34 Figure 4.2: Schematics showing the flow between different modules of PMA and its resources. 4.2.1 Language Parser The language parser is responsible for performing a rudimentary processing and segmentation of what was perceived by the recognition module. The whole process can be seen as a Natural Language Processing and Understanding procedure. The initial phases of project development integrated external tools for language processing, such as the text taggers JTextPro and TreeTagger. However, due to the need of developing PMA on the Android platform and because there are no tools to perform word tagging in Portuguese for Android, it was necessary to rebuild this module. A simple routine that tokenizes and then categorizes each token of the sentence was implemented. The first step is tokenization, this is, breaking the recognized text into meaningful elements (tokens) such as sentences and words. For this, a simple tokenizer routine was employed, making use of the BreakIterator package in Java. Afterwards, token list is parsed by a phrase chunker and a part-of-speech tagger. This syntactic analysis is required so that each sentence is segmented and each word is annotated with the corresponding word class and lemma. The parsing process identifies all the names and verbs on user utterances in order to spot keywords, extract the necessary information and infer the intention of the user. This parsing process relies on a grammar and a lexicon which can have a generalized range of words
The Personal Medication Advisor 35 or can be domain-specific. The final prototype uses of a domain-specific lexicon to make computation faster and lighter. Furthermore, a more complex sentence processing method was not demanded by the current system. In summary, the Language parser or NLU module starts to understand what the user says. The processed information will then be passed to the dialog manager. 4.2.2 Dialog Manager The dialog manager is responsible for goal analysis and information retrieval. When uttered words are tagged and keywords identified, the system tries to bring out the intention of what the user said and, accordingly, delineate a goal for what should be communicated in response to the user. A set of communicative goals were defined for PMA: EXISTENCE, INFORMATION, LIST, PERMISSION, REASON and UNKNOWN. These goals are used in an attempt of question classification. By analysing utterances, the system tries to include each of the user’s questions inside one of the mentioned attributes. As user utterances must always follow a predefined structure, it was settled that the keyword that establishes the goal, and thus reveals what the user requests from the system, is the first word of the utterance, in this case being an interrogative pronoun or a verb. This way, system performance is simplified and it is enough to associate a goal to each of the distinct intentions behind possible user utterances. For example, when the user questions the system “Qual é …?” (What is …), he must want to retrieve information related to a specific topic and the communicative goal is defined as INFORMATION. If the questions goes as “Quais são…?” (What are …), the user would want to see listed the attributes of a certain entity and the goal changes to LIST. If “Porque ...?” (Why …?) is asked, the user wants a justification and reasoning, and the goal is set to REASON. As for a utterance starting with “Posso …?” (Can I …?), the user would like to receive permission for something, being PERMISSION defined as the goal. The similar happens for the goal EXISTENCE. UNKNOWN is used when the goal could not be detected. Finally, when the goal is defined, the system acts according to the type of response that the user is expecting. Afterwards, depending on the goal, the dialog manager tries to identify, on the list of the tokens identified as nouns, a known medicine name and/or one of the featured property names. It also checks if there are verbs in the phrase. If yes, the first verb is also selected as a keyword. If the goal is defined as INFORMATION, LIST, EXISTENCE or REASON, the dialog manager initially looks for the medicine name on the ontology through specific SPARQL queries. After that, it will also check the user’s prescription to know if the uttered medicine was prescribed by the doctor. If that medication is not prescribed or if that medicine name is not valid, the dialog manager performs a Levenshtein Distance iterative routine to assess the similarity between two strings, the uttered medicine name and each name of the prescribed medicines. The Levenshtein Distance algorithm measures the edit distance, this is, the minimum number of single-character edits, such as deletion, insertion, and substitution, necessary to change one word into the other. The medicine name that is on the prescription, and
The Personal Medication Advisor 36 showed more similarity with the uttered medicine name, is then suggested to the user: “Não quererá dizer <nome semelhante de medicamento>?” (Don’t you mean <similar medicine name>?). The user will have to confirm the judgment of the dialog manager by answering yes or no. After checking the prescription and validating the medicine name, the dialog manager again searches for the necessary information on the ontology. This time the retrieved information is the value of the uttered property, such as dosage, frequency, targets, administration route or adverse reactions, for the selected medicine with the posology indicated on the prescription. This information is then sent to the generation module. If the goal was set as PERMISSION, the dialog manager will immediately check if the uttered medicine name is present on the medical prescription. According to its findings, it will give an appropriate positive or negative answer. If positive, it will also give information about the prescribed frequency for that medication. If negative, the system considers strange the fact that the user asked if he could take a medicine that is not prescribed. Patients tend to ask to take unprescribed medication when they fell unwell. Therefore, the dialog manager identifies a secondary intention behind the user’s question and initiates a medical inquiry. The medical inquiry was designed on purpose for HF patients. The objective is to assess if the user has one of the symptoms that characterizes worsen of HF. Some questions are made by the system such as “Tem pés e pernas inchados?” (Do you have swollen feet or legs?) or “Tem falta de ar?” (Do you have shortness of breath?). The user must answer “sim” (yes) or “não” (no) accordingly. If the patient answers “não” (no) to all the questions, PMA will ask him “Porque não se sente bem então?” (Why are you not feeling well then?) so that the user can freely describe its symptoms. Figure 4.3: Example of a medical report generated by PMA . Figure 4.4: Example of a medical report generated by PMA for a patient with two of the critical HF symptoms.
The Personal Medication Advisor 37 If he answered at least two times “sim” (yes), the system will tell the user to contact his doctor. Either way, the dialog manager will generate a report with the date, the patient answers each medical inquiry question and an observation field. The observation field can contain the user description of symptoms or the classification of the symptoms as urgent. An example of this medical report which would be sent by email or by message to the user’s healthcare staff can be seen in Figure 4.3 and Figure 4.4. 4.2.3 Natural Language Generator According to the selected keywords and the goal defined by the dialog manager, the generation module will build appropriate sentences to answer the user. The final objective is not only to output informative text but also to produce direct and suitable statements to clarify the user without causing any ambiguity. A group of short and simple sentences will then be preferred to fewer but large and complex phrases. Natural language generation in PMA is language and application-specific. It exploits grammar rules and lexicons that are specific to healthcare and medical domains. According to the literature, hybrid language generation systems seemed to be more successful and flexible (Galley, Fosler-Lussier, & Potamianos, 2001). When acceptable performance can be achieved without using complex linguist approaches, and only a few types of sentences are being generated, canned text and templates should be sufficient for realization (Reiter & Dale, Building Applied Natural Language Generation Systems, 1997). More complex approaches should make use of MTT, phrase or feature based methodologies. As consequence, PMA tried to implement a hybrid and pratical approach to NLG: simple and typical phrases are generated by canned-text and template-based approaches, while sentences requiring more interaction, initiative and complexity use a combination of phrase and feature-based methodologies. Canned-text was implemented when fixed instructions or direct answers were required. For example, when an error occurs, a statement is randomly selected from a list of fixed sentences that ask the user to repeat, such as “Não percebi o que disse. Pode repetir por favor?” (I did not understand what you said. Can you repeat please?). The template based approach was used for fixed sentences containing ‘gaps’ for the information that can differ but takes always the same place on the phrase. This was used, for instance, with the greetings that depend on the time of the day or while listing the prescribed medicines: “O seu medico receitou <número de medicamentos> medicamentos: <nome medicamento1> <nome medicamento2> ....” (Your doctor prescribed <number of medicines> medicines: <medicine name1> <medicine name2> …). Sentence realization includes linguistic and structure realization, where syntactic and morphological rules must be applied to build sentences. The sentences that were generated using phrase and feature based methods are usually the most complex ones, containing the answers that include information from the ontology. Each sentence is specified according to a certain phrase pattern, being composed by the following phrase elements: verb, subject, object, noun modifier, post-modifier and punctuation. Only the elements verb and punctuation are
The Personal Medication Advisor 38 defined as compulsory components of the phrase. Furthermore, generation is dictated by the characteristics of the sentence desired, according to a feature-based methodology. A unique set of features, such as negative/affirmative and singular/plural, must be defined to allow the sentence to be successfully generated. For getting the appropriate word order, as well as gender and number conjugations, the generator must make use of a lexicon. The XML lexicon contains the necessary information of every word that may be used. Each word element is then characterized by word features like category, number, gender, conjugations, etc… The generator analyses this information and selects the suitable word for the phrase element that it is building, according to the lexicon and the keywords that were recognized. For example, this module automatically chooses the correct word form for a selected noun as well as the determinant article that suits it, in accordance with gender and number. If the noun has an adjective that characterizes it, it will be integrated in conformity. For a verb, the appropriate conjugation is also chosen. Similarly, the object is adapted to the verb used, only appearing if the verb is transitive. Summarizing module division, the activity diagram from Figure 4.5 shows the flow of tasks in the PMA system. 4.3. Stages of development Along the five months of project, the PMA system had three main stages. Each of these phases is characterized by a different system prototype, using different tools and features but with similar architecture. 4.3.1 The first prototype The initial prototype was a stepping stone into exploring available tools, getting to know better the development platforms and making a preliminary design of system modules. This first version of PMA was English based, using CMU Sphinx for speech recognition and FreeTTS for speech synthesis. These two modules worked very well when integrated with the remaining system. Speech interaction in English was successful since FreeTTS presented very good synthesis in English and Sphinx allowed a sufficient recognition based on domain specific Java speech grammar format (JSGF) grammars. This first prototype only had the basic features allowing information retrieval on posology, adverse reactions, targets, frequency and administration route. This preliminary prototype could only perform canned text generation, this is, fixed strings were outputted according to the goal defined. The language parser used JTextPro for tokenization, phrase chunking and part-ofspeech tagging. The linguistic resources included an English dictionary bundled with Sphinx, which was slightly extended by adding the information on some domain specific words that were missing. This prototype was merely experimental as the ontology was in Portuguese and the system only allowed a rigid and very limited Q&A dialog.
The Personal Medication Advisor 39 Figure 4.5: Simplified UML activity diagram of PMA .
Evaluation and Discussion 46 Overall, user engagement with PMA seemed to be successful. However, the majority of the participants said that they felt that PMA was not very intuitive at the beginning. Only after some time of interaction with the system, they got to understand how it worked and then it got easy to use. Some of them even said that PMA should come initially with some embedded instructions. It was also observed that some participants wanted to forcefully talk with the system using non contemplated utterances because they were more natural for them. Moreover, they also considered that asking directly by the symptoms would be better. For example, they wanted to say ‘I have a fever’ or ‘I have a headache’ and PMA would answer with the medicine that is prescribed for that symptom. Figure 5.1: Average score (radius) given by the participants of the usability test to each of the 16 questionnaire’s statements (perimeter) with the respective attribute that each question assesses. Those who had working experience with senior people seemed to think that older patients would not learn how to use PMA quickly, doubting that they would use it for daily-life. In the end they said that some modifications should be introduced to further adapt PMA for an older audience. It was also suggested that it would be interesting to integrate PMA with visual information or even a barcode reader for medicine boxes. As some senior patients are not able to read, they may not know the name of their medication, written on the medicine packages or their prescriptions. Some participants said that some type of information as medication-medication interactions would also make the system more helpful. The usability test was performed with a small sample size, what is not particularly accurate for this type of questionnaire. However, it was enough to have a preliminary idea of how PMA 0 1 2 3 4 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
Evaluation and Discussion 47 can improve its reliability, helpfulness and usability. The conclusions derived from the usability testing will allow refining PMA and should be seen as future work. 5.2. General discussion 5.2.1 System Limitations The developed system presents some limitations that are important to be aware of. Overall, PMA lacks in flexibility. The system should be able to understand a wider variety of questions and user utterances as well as have a more natural flow of conversation. Android implementation brought PMA closer to a mixed-initiative conversational system. Even so, the system should still have more initiative to improve user engagement. In what relates to speech tools, a limitation is the fact that PMA needs to be connected to a network service provider so that speech tools can be used. Unfortunately, this cannot be avoided. In fact, this can be an advantage since web services can make the system faster. For successful recognition, the user must speak loud, clearly and not too fast. In addition, he must only speak when the assistant is silent, otherwise it will not work well. Speech recognition does not need to be perfect in order to be effective (Ferguson, Personal Health Management Assistant, 2007). Natural human-human conversations also have misunderstandings. However, in the current PMA it still occurs too often and affects system performance. The major recognition problem occurs with medicine names. More work could be done in the direction of correcting recognition errors through the introduction of a speech checker to suggest possible replacements for unrecognized words or to complete questions. Currently, PMA only does so when there is a recognition error on a medicine name. PMA is also slow to start because it has to load the whole ontology during application launching. This would be solved by using web services to get ontology information. Other limitation lies in the fact that the information contained in the ontology is extracted from a therapeutic compendium book and thus it is not specific for that patient (not personalized) being sometimes not so suitable for that user. 5.2.2 Other considerations A reasonable question can be raised for discussion: would PMA be used by older people? Some barriers are present in healthcare systems whose target is elderly population. Since older patients do not have easy access, technical literacy or familiarity with technology in general, the adherence to a computerized-assistant could be poor. Senior patients may view the technology as intrusive and not suitable with their lifestyle, offering resistance to the introduction of new equipment or new habits. They are also overly attached to their daily routines and it would take
Evaluation and Discussion 48 a lot of time to feel comfortable with the developed technology. Furthermore, as it was concluded from the usability testing, initial training could be required. At the NLG level, the complexity of PMA has not yet arrived to a stage where stochastic models are needed for generation. Simultaneously, generation is made phrase by phrase, still far away from the desired multi-sentence generation. A spell checker for the generated sentences could also be implemented to recognize possible structure mistakes or misconceptions from users’ queries. It is important that PMA remains a medical information-provider rather than a decisionmaker. A conversational system should never be a physician’s substitute. ‘Technology should be developed to complement and augment human decision-making’ (Ferguson, Personal Health Management Assistant, 2007). Furthermore, it is important to maintain a good granularity for the information given by the intelligent assistant. The type and amount of information that is appropriated to give may not be the same for every patient. For example, the long list of the adverse effects presented on prescribed medicines can pointlessly stress patients, even when most of these events are rare or unworthy of attention. Once again, further personalization of PMA could avoid causing these unnecessary worries or anxiety. The information that is being given to the patient should be adapted and given according to some priorities and to the social context of the patient, so that problems of excess of information can be avoided. This adjustment can also make the assistant more robust (Jones, Cawsey, & Bentala, 1999). 5.3. Summary A semi-formal evaluation of PMA was performed to assess usability and detect its drawbacks so that improvements can be made in the future. Participants were asked to complete two tasks with the developed android application and then fill a usability questionnaire. It was concluded that users felt that the system was easy to use, simple and had a good interface. Participants also gave several suggestions that could help improving PMA. The system still presents limitations such as lack of flexibility, need a wireless connection to the network, or frequent recognition mistakes. System helpfulness for senior users was also discussed.
49 Chapter 6 Conclusions After five months of project development, the Personal Medication Assistant culminated in an accessible android application that interacts with the user through voice to give information about medication. In this last chapter, a balance of the overall project is made: from the evaluation of the initial objectives to the main contributions and the challenges found. Possible futures for PMA refinement are also discussed. 6.1. Final Balance The major objective of this project was to come nearer to a conversational assistant: the Personal Medication Advisor, a system capable of real-time interaction with a user to help him manage his medication through spoken natural language in European Portuguese. Making use of speech tools, ontologies and linguistic resources as grammars and lexicons, the final system resulted of iterations during three stages of development. The last prototype took the form of an Android application where the user can successfully dialog with the system using voice. PMA structure is composed by five modules: a recognizer, a language parser, a dialog manager, a language generator and a synthesizer. Both recognizer and synthesizer modules are external using, in the final PMA version, Google Voice for speech recognition and iSpeech for speech synthesis. Natural language understanding was achieved by the language parser using phrase chunking, part-of-speech tagging and part of keyword spotting. Single-sentence natural language generation was achieved through hybrid methodologies that use simultaneously canned-text, template-based and feature-based approaches. The PMA application was validated through a usability test. This evaluation wanted to assess system’s feasibility, performance and interface, through participants’ opinions after completing two tasks and filling a suitable questionnaire. Task objectives were successfully achieved. It was concluded that users felt that the system was simple and had a good interface.
Conclusions 50 If this project moves forward, further adaptations for senior users must be made to make PMA more helpful and accessible for them. Initially, this work was meant to be focused on the development of a language generator module capable of producing sentences to be synthesized by existent TTS systems. This objective was accomplished and expanded to the incorporation of a language understanding and dialog management module with existing ASR systems. This was because functional language parsing and managing were needed for achieving any type of conversational assistance. With this, all the functionalities planned in the beginning for PMA were achieved. The objective of having successful integration between NLU/NLG modules and speech tools was also accomplished. As proposed, PMA generates phrases that are coherent, short and easy to understand. Strengths & weakness PMA’s good points are inherent in its motivation, its architecture that aimed to enhance modularity, its simplicity and its pleasant interface. The uniqueness of PMA lies on the use of ontologies as source of medical information. On the other hand, the system still lacks in flexibility, initiative and naturalness, compared to users’ expectations of a conversational assistant. The system should require no special training for the patients that are going to use it. However, the usability test showed that for senior users some features need to be slightly modified to facilitate their first interaction with PMA. Contributions The system was envisioned within a patient centred approach to personal healthcare, namely Heart Failure care, so that home-monitoring and treatment quality and efficacy can be enhanced. The specific evidence-based information given by PMA can help patients making positive behaviour changes. For example, the available information will allow users to know possible negative effects of their medication, not being caught completely unaware. The simplicity and interactivity of the system will trigger user engagement and hopefully improve medication adherence. User satisfaction with PMA will raise their self-consciousness of home healthcare and work together with physicians to modify disease outcome. Challenges encountered During the five months of project, work evolved dynamically as the system was being developed. Nevertheless, major challenges appeared during the different stages of PMA. The major obstacle for this project was associated with the lack of freely available speech tools for European Portuguese. Several tools were assessed and only the final version was able to have successful recognition in this language. Additionally, compared to other languages, European Portuguese is still very immature in what is related to natural language research, especially NLG. Herewith, generation presented a great challenge.
Conclusions 51 Android implementation of PMA also required an extra effort because it was necessary to remodel part of what had already been developed on previous stages. This was also a new challenging platform of development, since it was the first contact with it. 6.2. Future work As PMA has limitations, further system development can be made. Due to the complex nature of conversational assistants, there are always new features to add or improvements to make. This is especially true for natural language understanding and generation. Granularity of utterance’s interpretation as well as of the phrase elements that will be part of generated sentences can be iteratively diminished and refined. This will gradually result in a more complex and enhanced system. Step by step, it should also smooth some of the system limitations, such as the lack of flexibility, the lack of initiative or of naturalness. One of these steps could be the implementation of statistical and/or stochastic methodologies. Others include giving the system the ability to learn common user utterances or underlying intentions, building more complete approaches for language generation, or introducing an automated vocabulary growth feature. For example, the ability to learn new words is important for interactive systems and will increase their flexibility. PMA needs to be more tailored to each patient needs. A major improvement at this level would be the incorporation of user’s medical history and past medical records. This could be complemented with a variety of useful data about the patient. Basic information such as name, age, sex, educational level, local background and physiological data (height, weight, blood pressure, ….) could be easily provided. Recognition and synthesis of user’s name would also make the interaction more natural and personalized. With a great increase of complexity, additional intelligent modules could be useful to detect speaker identity, emotional state or even degree of motivation. The system could be programmed to use this information and adjust its decisions and dialogue accordingly. This attempt to personalize the assistance would be very advantageous, although it would present a bigger challenge. Further adaption of PMA to senior users is still necessary. From modification of the type of utterances allowed to the integration of visual information on the medicine appearance, more work can be done to increase acceptance and convenience of PMA for older patients. Incorporation with already existent medication reminders or alerts would also boost system helpfulness. Moreover, as stated in the introduction, a second and advanced approach to the Personal Medication Assistant will have to include contraindications, cautions and drug to drug interactions.
52 References Aguiar, B. L. (2012). Information Extraction From Medication Leaflets. Master Thesis, FEUP, Porto. Allen, J. F., Byron, D. K., Dzikovska, M., Ferguson, G., Galescu, L., & Stent, A. (2001). Toward conversational human-computer interaction. AI Magazine, 22(4), 27-38. Allen, J., Ferguson, G., & Stent, A. (2001). An Architecture For More Realistic Conversational Systems. 6th international conference on Intelligent user interfaces (pp. 1-8). New Mexico: ACM New York. Allnatt, R., Engelhardt, J. B., Gao, J., & Mariano, A. (2001). An Evaluation of the Functionality and Acceptability of the Voice Prescription Label. Journal of Visual Impairment and Blindness, 95(11), 702-706. Axelrod, S. (2000). Natural language generation in the IBM Flight Information System. ANLP/NAACL, Workshop on Conversational Systems, 3, pp. 21-26. Seattle. Bangalore, S., & Rambow, O. (2000). Corpus-based lexical choice in natural language generation. 38th Annual Meeting of the Association for Computational Linguistics, (pp. 464-471). Hong Kong. Bangalore, S., & Rambow, O. (2000). Exploiting a probabilistic hierarchical model for generation. 18th International Conference on Computational, (pp. 42-48). Saarbrücken, Germany. Baptist, L. M. (2000). Genesis II: a language generation module for conversational systems. Master Thesis, MIT, Department of Electrical Engineering and Computer Science, B.A. Baptista, J., Mamede, N., & Gomes, F. (2010). Auxiliary Verbs and Verbal Chains in European Portuguese. 9th International Conference of Computational Processing of the Portuguese Language, (pp. 110-119). Porto Alegre. Benzeghiba, M., Mori, R. D., iDeroo, O., Dupont, S., Erbes, T., Jouvet, D., . . . Wellekens, C. (2007). Automatic speech recognition and speech variability: A review. Speech Communication, 49(10-11), 763–786. Brooke, J. (1986). SUS – A quick and dirty usability scale. United Kingdom: Redhatch Consulting. Cardoso, P., Flores, L., Langlois, T., & Neto, J. (2002). Meteo: A Telephone-Based Portuguese Conversation System in Weather Domain. Third International Conference of Advances in Natural Language Processing, PorTAL, (pp. 175-178). Faro. Carolis, B. D., Rosis, F. d., Grass, F., Rossiello, A., Berry, D. C., & Gillie, T. (1996). Generating recipientcentered explanations about drug prescriptions. Artificial lntelligence in Medicine, 8, 123145. Carvalho, P., Oliveira, L. C., Viana, M. C., & Trancoso, I. (1998). Concatenative speech synthesis for European Portuguese. Lisbon: INESC/ICT.
References 53 Ceia, F., Fonseca, C., Mota, T., Morais, H., Matias, F., Costa, C., & Gouveia-Oliveira, A. (2004). Epidemiology of heart failure in mainland Portugal: new data from the EPICA study. Revista Portuguesa de Cardiologia, 23(3), 15-22. Chambers, N., & Allen, J. (2004). Stochastic Language Generation in a Dialogue System: Toward a Domain Independent Generator. 5th SIGdial Workshop on Discourse and Dialogue. Boston. Cole, R., Mariani, J., Uszkoreit, H., Varile, G. B., Zaenen, A., & Zampolli, A. (1996). Survey of the State of the Art in Human Language. Cambridge University Press. Corston-Oliver, S. (2000). Using Decision Trees to Select the Grammatical Relation of a Noun Phrase. 1st SIGDial workshop on discourse and dialogue, 10, pp. 66-73. Hong Kong. Dale, R. (1995). An Introduction to Natural Language Generation. European Summer School in Logic, Language and Information. Barcelona. Danieli, M., & Gerbino, E. (1995). Metrics for evaluating dialogue strategies in a spoken language system. AAAI Spring Symposium on Empirical Methods in Discourse Interpretation and Generation, 16. Davis, H. K., Biddulph, R., & Balashek, S. (1952). Automatic Recognition of Spoken Digits. Journal of the Acoustical Society of America, 24(6), 627-642. Dudley, H. (1939). The Vocoder. Bell Labs Record, 17, 122-126. Ferguson, G. (2007). Personal Health Management Assistant. PhD Thesis, University of Rochester. Ferguson, G., Allen, J., Blaylock, N., Byron, D., Chambers, N., Dzikovska, M., . . . Swift, M. (2006). Chester: towards a personal medication advisor. Journal of biomedical informatics, 39(5), 500-513. Ferguson, G., Quinn, J., Horwitz, C., Swift, M., Allen, J., & Galescu, L. (2010). Towards a Personal Health Management Assistant. Journal of biomedical informatics, 43(5 Suppl), S13-16. Galescu, L., Allen, J., Ferguson, G., Swift, M., & Quinn, J. (2009). Speech Recognition in a Dialog System for Patient Health Monitoring. IEEE International Conference on Bioinformatics and Biomedicine - Workshop on NLP Approaches for Unmet Information Needs in Health Care. Washington D.C. Galley, M., Fosler-Lussier, E., & Potamianos, A. (2001). Hybrid Natural Language Generation for Spoken Dialogue Systems. European Speech Processing Conference, (pp. 1735–1738). Aalborg. Gatt, A., & Port, F. (2009). Text content and task performance in the evaluation of a Natural Language System. International Conference on Recent Advances in Natural Language Processing. Giles, T. D. (2004). New perspectives in the prevention and treatment of chronic heart failure with reduced systolic function. American Journal of Hypertension, 17(1S), 249A. Goldberg, E., Driedger, N., & Kittredge, R. (1994). Using natural-language processing to produce weather forecasts. IEEE Expert, 9(2), 45-53.
References 54 Grancelli H, V. S. (2003). Randomized trial of telephone intervention in chronic heart failure (DIAL): study design and preliminary observations. Journal of Cardiac Failure, 9(3), 172-179. Haux, R. (2006). Health information systems – past, present, future. International Journal of Medical Informatics, 75(3-4), 268–281. Heart Failure Association of the ESC. (2010). EURObservational Research Programme: The Heart Failure Pilot Survey. European Journal of Heart Failure, 12, 1076–1084. Jelinek, F. (1997). Statistical Methods for Speech Recognition. MIT Press. Jimison, H., Gorman, P., Woods, S., Nygren, P., Walker, M., Norris, S., & Hersh, W. (2008). Barriers and drivers of health information technology use for the elderly, chronically ill, and underserved. Rockville: Agency for Healthcare Research and Quality. Jones, R., Cawsey, A., & Bentala, D. S. (1999). Patient information systems that tailor to the individual. Patient Education and Counseling, 36, 171–180. Kamm, C., Walker, M., & Litman, D. (2000). Evaluating Spoken Language Systems. AT&T Labs Research. Krumholz, H. M., Chen, Y. T., Wang, Y., Vaccarino, V., Radford, M. J., & Horwitz, R. I. (2000). Predictors of readmission among elderly survivors of admission with heart failure. American Heart Journal, 139(1), 72-77. Lemmetty, S. (1999). Review of Speech Synthesis Technology. Master Thesis, Helsinki University of Technology, Helsinki. Mavridis, N. (2007). Grounded Situation Models for Situated Conversational Assistants. PhD thesis, Massachusetts Institute of Technology, Massachusetts. Meinedo, H., Caseiro, D., Neto, J. P., & Trancoso, I. (2003). AUDIMUS.MEDIA: A Broadcast News Speech Recognition System for the European Portuguese Language. 6th International Workshop of Computational Processing of the Portuguese Language. Faro. Mellish, C., & Evans, R. (2004). Implementation architectures for natural language generation. Natural Language Engineering, 10(3-4), 261 - 282. Mettler, M., & Kemper, D. W. (2003). Information therapy: health education one person at a time. Health promotion practice, 4(3), 214-217. Michalsen, A., König, G., & Thimme, W. (1998). Preventable causative factors leading to hospital admission with decompensated heart failure. Heart, 80, 437–441. Moore, J. D., & Paris, C. L. (1993). Planning text for advisory dialogues: Capturing intentional and rhetorical information. Computational Linguistics, 19, 651-694. Naylor, M. D., Brooten, D. A., Campbell, R. L., Maislin, G., McCauley, K. M., & Schwartz, J. S. (2004). Transitional care of older adults hospitalized with heart failure: A randomized control trial. Journal of American Geriatric Society, 52, 675-684. Neto, J., & Cassaca, R. (2004). A Robust Input Interface in the scope of the Project Interactive Home of the Future. ROBUST 04. Norwich.
References 55 Nieuwenhuis, M. W., Jaarsma, T., van Veldhuisen, D. J., & van der Wal, M. L. (2012). Self-reported versus ‘true’ adherence in heart failure patients: a study using the Medication Event Monitoring System. Netherlands Heart Journal, 20(7-8), 313–319. Oh, A. H., & Rudnicky, A. I. (2002). Stochastic natural language generation for spoken dialog systems. Computer Speech and Language, 16, 387-407. Olson, F. H., & Belar, H. (1956). Phonetic Typewriter. Journal of the Acoustical Society of America, 28(6), 1072-1081. Phan, X.-H. (2006). Retrieved March 2013, from JTextPro: A Java-based Text Processing Toolkit: http://jtextpro.sourceforge.net Raistrick, Phillip C; Raistrick, David; En-Vision America. (2006). Audible Prescription Reading Devices: Solving Issues In Privacy, Safety And Health. International Conference on Aging, Disability and Independence. St. Petersburg. Reiter, E. (1995). NLG vs. templates. Fifth European Workshop on Natural Language Generation. The Netherlands. Schroeder, M. R. (1993). A Brief History of Synthetic Speech. Speech Communication, 13(1-2), 231– 237. Seneff, S. (2003). Spoken Computer conversational systems. JHU Summer Language Workshop. Massachusetts: CSAIL, MIT. Silva, J., Branco, A., Castro, S., & Reis, R. (2010). Out-of-the-Box Robust Parsing of Portuguese. 9th International Conference of Computaional Processing of the Portuguese Language, (pp. 7585). Porto Alegre. Simon, S., & Horowitz, J. (2002). Home-based intervention in congestive heart failure: long-term implications on readmission and survival. Circulation, 105(24), 2861-2866. Smedt, K. D., Horacek, H., & Zock, M. (1996). Architectures for Natural Language Generation: Problems and Perspectives. Trends in natural language generation: an artificial intelligence perspective (pp. 17-46). Springer-Verlag. Smith, R. W. (1997). Performance Measures for the Next Generation of Spoken Natural Language Dialog Systems. Workshop On Interactive Spoken Dialog Systems: Bringing Speech And NLP Together In Real Applications. Madrid. Staab, S., & Studer, R. (Eds.). (2009). Handbook on Ontologies. Berlin: Springer. Stewarta, S., Jenkins, A., Buchan, S., McGuire, A., Capewell, S., & McMurray, J. J. (2001). The current cost of heart failure to the National Health Service in the UK. European Journal of Heart Failure, 4(3), 361-371. Studer, R., Benjamins, R., & Fensel, D. (1998). Knowledge engineering: Principles and methods. Data & Knowledge Engineering, 25(1-2), 161–198. Sullivan, R., Behncke, I., & Purushotham, A. (2010). Why do we love medicines so much? Science and Society, 11(8), 572–578.
62 Do you usually work for or with senior people? If YES, how long have you been working for or with them? 8. The information provided by the system is easy to understand. Strongly disagree 1 2 3 4 Strongly agree 9. The system gives error messages that clearly tell me how to fix problems. Strongly disagree 1 2 3 4 Strongly agree 10. I was able to get useful information using this system. Strongly disagree 1 2 3 4 Strongly agree 11. I quickly could get the information I needed. Strongly disagree 1 2 3 4 Strongly agree 12. I felt very confident using the system. Strongly disagree 1 2 3 4 Strongly agree Please answer the next questions having a senior user in mind: 13. The interface of this system is pleasant. Strongly disagree 1 2 3 4 Strongly agree 14. Older people would learn to use this system very quickly. Strongly disagree 1 2 3 4 Strongly agree 15. This system would help senior patients to understand and manage their medication regimen. Strongly disagree 1 2 3 4 Strongly agree 16. For those who follow strict medications regimens, I think they would like to use this system frequently. Strongly disagree 1 2 3 4 Strongly agree