scieee AI-readable full text Open interactive document viewer

Contribuciones al análisis y desarrollo de robots sociables

Déniz Suárez, Oscar

Abstract

Programa de doctorado: Sistemas Inteligentes y Aplicaciones Numéricas en Ingeniería

Full text

' UNIVERSIDAD DE LAS PALMAS DE GRAN CANARIA Departamento de Informática y Sistemas TESIS DOCTORAL Contribuciones al análisis y desarrollo de robots sociables (An Engineering Approach to Sociable Robots) Óscar Déniz Suárez Las Palmas de Gran Canaria, Abril 2006 UNIVERSIDAD DE LAS PALMAS DE GRAN CANARIA Departamento de Informática y Sistemas Tesis titulada Contribuciones al análisis y desarrollo de robots sociables, que presenta D. Óscar Déniz Suárez, dentro del programa de doctorado Sistemas Inteligentes y Aplicaciones Numéricas en Ingeniería, realizada bajo la dirección del Catedrático de Universidad D. Fco. Mario Hernández Tejera y la codirección del Doctor D. José Javier Lorenzo Navarro. Las Palmas de Gran Canaria, Abril de 2006 El director Fdo. Fco. Mario Hernández Tejera El codirector Fdo. José Javier Lorenzo Navarro El doctorando Fdo. Óscar Déniz Suárez To my family and Gloria Acknowledgments The work described in this document would not have been posible without the disinterested help of many people. First of all, 1 want to acknowledge the continuous support of my tutor Mario Hernández. Besides being very knowledgeable about the topics covered in my work, Mario is one of the most cheerful guys I've ever met. My second tutor Javier Lorenzo also supported me on many occasions, impressing me with his incredibly fast thought. Many people contributed sporadically: Jorge Cabrera, Daniel Hernández, Antonio Carlos Dominguez, Cayetano Guerra, Josep Isern, Antonio Falcón, Juan Méndez, Irael Pérez and David Hernández, al1 from my research group. Modesto Castrillón was also a good companion, from back when 1 started working on face recognition. Now 1 know he is the person 1 should resort to if 1 need to know something about facial analysis or travelling around the world. Modesto kindly provided me with the ENCARA face detection module. David let me use its ZagaZ module for action selection. Also, the French mechanical engineers of IFMA who visited our laboratory, Yann, Patrick, Celine, Vincent and Olivier, were decisive in the mechanical design of CASIMIRO. Other people did not contribute so specifically, but they did in other subtle ways that 1 can't articulate: Yeray, Sanyu, José Luis, Claudio, Carlos Moisés, Oscar, Fayna, Atamán, Jaime, Carmelo, Cristóbal, Patricia, Antonio and Marilola. Much to my surprise, instant messaging turns out to be a good concept when you spend so many hours working in front of a computer screen. 1 also want to thank my pupils and fellow teachers at the School and Faculty of Computer Science, especially Miguel Angel Pérez, Alexis Quesada, Carmelo Rubén Garcia, José Antonio Muñoz, Francisca Quintana and Santiago Candela, who were very understanding and helpful with respect to my teaching duties. 1 want to thank my parents and my sisters María Jesús (Susi) and Lorena. Hopefully, now that 1 have finished this work Susi and Lorena will have more access to our home Internet connection. 1 am grateful for the funding granted by the entities that supported our research group: Gobierno de Canarias, Ministerio de Ciencia y Tecnología and Universidad de Las Palmas de Gran Canaria. 1 am personally indebted to Universidad de Las Palmas de Gran Canaria, and specially to the research and academic staff vice-chancellors, who have supported me since 1 got a research grant back in 2000. Finally, 1 want to thank the person who's holding this document right now for showing interest in my work. 1 hope to be back soon. It is my hope that CASIMIRO's abilities will be expanded in the future. I've always thought that the most important thing is not the destination, but the journey itself. In itinere sumus. Contents Abstract xi 1 Introduction 1 1.1 Previous Work ................................. 4 1.2 Analysis .................................... 8 2 The Case of Social Robotics 11 2.1 The Savant Syndrome ............................. 12 2.2 Unconscious Mental Processes ........................ 15 2.3 The Reason why Social Robots may not be Robust ............. 19 3 Approach and Architecture 25 3.1 Analysis vs Synthesis ............................. 25 3.2 Niches ..................................... 28 3.3 Design and Niche Spaces ........................... 3 1 3.4 Design Principles ............................... 33 3.5 Architecture .................................. 36 4 Hardware Oveniew 45 4.1 Hardware ................................... 47 4.2 Software Overview .............................. 52 5 Perception 55 5.1 Omnidirectional Vision ............................ 56 5.2 Sound Localization .............................. 60 5.2.1 Previous Work ............................ 6 1 5.2.2 A Study on Feature Extraction for Sound Localization ....... 62 5.2.3 Implementation ............................ 68 5.3 Audio-Visual Attention ............................ 71 5.3.1 Attention Model ........................... 72 5.3.2 Implementation and Experiments .................. 74 5.4 Face Detection ................................ 78 5.4.1 The ENCARA Face Detector ..................... 78 5.4.2 Performance and Implementation .................. 8 1 5.4.3 Use of Stereo Information and Shiri Elimination .......... 82 5.5 Head Nod and Shake Detection ........................ 83 r. D 5.6 Memory and Forgetting 88 D N ............................ 6 .- L 5.7 Owner Identification ............................. 94 L" m c > .- 5.8 Habituation 96 c 3 .................................. m a - .- - 0 107 m 6 Action o C1 6.1 Facial Expression 107 5 3 ............................... L 0 * 6.1.1 Functional Model 108 m n ........................... N .- - 6.1.2 Transitions Between Expressions 109 m ? .................. c :o 6.1.3 Implementation ............................ 11 1 .- m - m c m .- 6.2 Neck ...................................... 112 O 6.3 Voice Generation ..................... 6.3.1 Expressive Talk .................. 6.3.2 Local Accent ................... 7 Behaviour .......................... 7.1 Reflexes 7.2 Action Selection ...................... ............... 7.2.1 Behaviour Networks 7.2.2 ZagaZ ....................... ............. 7.2.3 Implemented Behaviours 7.3 Emotions ......................... 8 Evaluation and Discussion 137 8.1 Evaluating Social Robots, a Review ..................... 138 8.2 Quantitative/QualitativeMeasures ...................... 139 ............................. 8.3 Interaction Examples 140 8.4 Interviews. What to Evaluate? ........................ 143 ................................. 8.5 Questionnaire 147 ............................ 8.6 Results and Discussion 148 9 Conclusions and Future Work 155 ........................... 9.1 Conceptual Contributions 156 ............................ 9.2 Technical Contributions 157 .................................. 9.3 Future work 160 A CASIMIRO's Phrase File Abstract A relatively new area for robotics research is the design of robots that can engage with humans in socially interactive situations. These robots have expressive power (i.e. they al1 have an expressive face, voice, etc.) as well as abilities to locate, pay attention to, and address people. In humans, these abilities fa11 within the ambit of what has been called "social intelligence". For this class of robots, the dominant design approach has been that of following models taken from human sciences like developmental psychology, ethology and even neurophysiology. We show that the reproduction of social intelligence, as opposed to other types of human abilities, may lead to fragile performance, in the sense of having very different performance~ between trainingltesting and future (unseen) conditions. This limitation stems from the fact that the abilities of the social spectrum, which appear earlier in life, are mainly unconscious to us. This is in contrast with other human tasks that we carry out using conscious effort, and for which we can easily conceive algorithms. Thus, a coherent explanation is also given for the truism that says that anything that is easy for us is hard for robots and vice versa. For some types of robots like manipulators one can extract a set of equations (or algorithms, representations, ...) that are known to be valid for solving the task. Once that these equations are stored in the control computer the manipulator will always move to desired points. Sociable robots, however, will require a much more inductive development effort. That is, the designer tests implementations in a set of cases and hopes that the performance will be equally good for unseen (future) cases. Inductive processes crucially depend on a priori knowledge: if there is little available one can have good performance in test cases but poor performance in unseen cases (overfitting). In Machine Learning, complexity penalization is often used as a principled means to avoid overfitting. Thus, we propose to develop sociable robots starting from simple algorithms and representations. Implementations should evolve mainly through extensive testing in the robot niche (the particular environment and restrictions imposed on the robot tasks, physical body, etc.). Such approach places more emphasis in the engineering decisions taken throughout the robot development process, which depend very much on the niche. This work describes the ideas and techniques involved in the design and development of CASIMIRO, a robot with a set of basic interaction abilities. The robot has been built following the mentioned approach. In particular, the main difficulties lay in parsimoniously exploiting the characteristics of the robot niche in order to obtain better performances. Publications This thesis is based on research that was previously reported in the following publications: "Face Recognition Using Independent Component Analysis and Support Vector Machines': 0. Déniz, M. Castrillón, M. Hernández. IX Spanish Symposium on Pattern Recognition and Image Analysis, Castellón (Spain), May 2001. "Face Recognition Using Independent Component Analysis and Support Vector Machines': 0. Déniz, M. Castrillón, M. Hernández. Procs. of the Third International Conference on Audioand Video-Based Person Authentication. Lecture Notes in Computer Science 2091, pp. 59-64. Halmstad, Sweden, June 2001. 'Estudio Experimental sobre la Combinación Temporal de Resultados en el Reconocimiento de Caras con Secuencias de Video': O. Déniz, M. Castrillón, J. Lorenzo, M. Hernández. IX Conferencia de la Asociación Española para la Inteligencia Artificial CAEPIA-2001, Vol. 1 pp. 273-282, Gijón, November, 2001. 'El Método IRDB: Aprendizaje Incremental para el Reconocimiento de Caras': O. Déniz, M. Castrillón, J. Lorenzo, M. Hernández. IX Conferencia de la Asociación Española para la Inteligencia Artificial CAEPIA-2001, Vol. 1 pp. 273-282, Gijón, November, 200 1. 'IAprendizaje Incremental para la Identificación Facial a Partir de Secuencias de Video": O. Déniz, M. Castrillón, J. Lorenzo, M. Hernández. Revista Española de Visión por Computador. vol. 6, March, 2002. 'Modelado de expresiones para una cara robótica': 0. Déniz, A. Falcón. Revista Buran, vol. 10, no 18, April2002. 'Xn Incremental Learning Algorithm for Face Recognition': O. Déniz, M. Castrillón, J. Lorenzo, M. Hernández. Lecture Notes in Computer Science, vol. 2359, pp. 1-9, ... Xlll Springer-Verlag, PostECCV'02 Workshop on Biometric Authentication, Copenhagen, Denmark June 2002. "CASIMIRO: A Robot Head for Hurnan-Cornputer Interaction': O. Déniz, M. Castrillón, J. Lorenzo, C. Guerra, D. Hernández, M. Hernández. 11th IEEE International Workshop on Robot and Human Interactive Communication, pp. 319-324, Berlin, Germany, 25-27 September, 2002. "A Computational Mechanism for Habituation in Perceptual User Interfaces", O. Déniz, J. Lorenzo and M. Hernández. Int. Conference on Computational Intelligence for Modelling, Vienna, Austria, February 12-14, 2003. 'Building a Sound Localization Systern for a Robot Head': O. Déniz, J. Cabrera and M. Hernández. Revista Iberoamericana de Inteligencia Artificial, 18 - Winter 2003. 'Multirnodal Attention Systern for an Interactive Robot': O. Déniz, M. Castrillón, J. Lorenzo, M. Hernández and J. Méndez. Lectures Notes in Computer Science, vol. 2652. 1st Iberian Conference on Pattern Recognition and Image Analysis. 4-6 June 2003, Pto. Andratx, (Mallorca, Spain). "Face Recognition using Independent Cornponent Analysis and Support Vector Machines': 0. Déniz, M. Castrillón and M. Hernández. Pattern Recognition Letters, vol 24, issue 13, pp. 2153-2157, September 2003. 'BDIE: A BDI like architecture with ernotional capabilities': D. J. Hernández, O. Déniz, J. Lorenzo and M. Hernández. AAAI Spring Symposium, 22-24 March 2004. '2 simple habituation rnechanisrn forperceptual user intefaces': O. Déniz, J. Lorenzo and M. Hernández. Revista Iberoamericana de Inteligencia Artificial, vol. VIII, 33, pp. 7-16, 2004. "Useful Cornputer Vision Techniques for Hurnan-Robot Interaction': O. Déniz, A. Falcón, J. Méndez, M. Castrillón, ICIAR 2004, International Conference on Image Analysis and Recognition, September 2004, Porto, Portugal. 'Expressive Robotic Face for Interaction': O. Déniz, L. Antón-Canalís, M. Castrillón, M. Hernández. VI1 Workshop de Agentes Físicos (WAF'Oó), April2006, Las Palmas de Gran Canaria, Spain. Chapter 1 Introduction 'Mari is by nature a social animal" Aristotle, Politics. In recent years there has been a surge of interest in a topic called social robotics. As used here, social robotics does not refer to groups of robots that cooperate and interact with each other. For a group of robots, communication is relatively simple from a technological point of view, they can use whatever complex binary protocol to "socialize" with their partners. Back in the 40's [Holland, 19971 robotic tortoises already interacted in a "social" manner using headlamps attached to the robots and by means of phototaxis. This field of Artificial Intelligence evolved and produced concepts like "swarm" robots, "ant-like" robots, self-organization, etc. For us, the adjective social refers to humans. In principle, the implications of this are much wider than in the case of groups of robots. Although the distinction may be much fuzzier, social robotics aims at building robots that do not seem to have a specific task (like playing chess with ahuman) but to simply interact with people. Socializing with humans is definitely much harder, not least because robots and humans do not share a common language nor perceive the world (and each other) in the same way. Many researchers working on this topic use other names like human-robot interaction, perceptual interfaces or multimodal interfaces. However, as pointed out in [Fong et al., 20031 we have to distinguish between conventional human-robot interaction (such as that used in teleoperation scenarios or in friendly user interfaces) and socially interactive robots. In the latter, the common underlying assumption is that humans prefer to interact with robots in the same way that they interact with other people. CHAPTER f. INTRODUCTION Traditional robots can carry out tasks that are well beyond human capabilities, with greater precision, no risk and a large number of times. Then why so much interest in social robots? Admittedly robots are built with the aim of imitating or reproducing human intelligence. "Before" social robotics, most robots excelled at certain tasks, although they were incapable of doing other "simple" things, not even in a partial way. Almost in an unconscious fashion, researchers realized that those robots were extremely practical, but were not considered more intelligent (at least by the general public). In a sense, the more practical and precise the robot, the less human it was considered. The fact is that humans have a wide range of abilities. Comparatively, robots tend to have far fewer abilities, although with high proficiency levels, in some cases even higher than in humans (see Figure 1.1). Such performance unbalance may be key to building social robots. It suggests that two necessary conditions for achieving a more "human" robot may be: 1. to replicate a large number of human abilities, and 2. that these abilities have similar development levels. obilities obilities Figure 1.1: Abilities vs. leve1 of development. On the left: human performance. On the right: typical robot performance. Human intelligence does not seem to be restricted to certain abilities. According to Gardner's theory [Gardner, 19831, now widely accepted, human intelligence is not a unitary capacity that can be measured by IQ tests. He proposes eight classes of intelligence (see Table l. l ): linguistic, musical, logical-mathematical, spatial, bodily-kinesthetic, interpersonal, intrapersonal and naturalist. Intelligence has two meanings. First, it is a species-specific CHAPTER f. INTRODUCTION characteristic (humans can exhibit those eight intelligences). Second, it is also an individualspecific characteristic (each individual shows his particular blend of intelligences). Gardner claims that the eight intelligences rarely operate independently. Rather, the intelligences are used concurrently and typically complement each other as individuals develop skills or solve problems. For example, a dancer can excel in his art only if he has 1) strong musical intelligence to understand the rhythm and variations of the music, 2) interpersonal intelligence to understand how he can inspire or emotionally move his audience through his movements, as well as 3) bodily-kinesthetic intelligence to provide him with the agility and coordination to complete the movements successfully. The theory states that al1 eight intelligences are needed to productively function in society. 1 Intelligence 1 Operations 1 Lincuistic 1 svntax. ohonolom semantics. oracmatics 1 Musical 1 pitch, rhythm, timbre Lopical-mathematical number. cateporization. relations Table 1.1 : The eight human intelligences proposed by Gardner. Spatial Bodily-kinesthetic Interpersonal Gardner's proposal of multiple intelligences was prompted by severa1 signs, like isolation by brain damage, in which certain forms of intelligence are impaired while others remain relatively intact. Damage to the prefrontal lobes of the cerebral cortex can impair personal and social intelligence, while other abilities remain normal [Damasio, 19941. Gardner mentions exceptional cases like autistic savants, and he notes that both the Down syndrome and the Alzheimer's disease are cognitive impairments that do not severely hinder a person's ability to get along with other people. By contrast, Pick's disease seriously diminishes a person's interaction abilities, though some intellectual capacities are not affected [Kihlstrom and Cantor, 20001. accurate mental visualization, mental transfomation of images control of one's own body, control in handling objects awareness of others' feelings, emotions, goals, motivations Not only theory favours the fact that there are different types of intelligence. In a series of empirical studies, [Sternberg et al., 19811 found that people generally identify intelligence with three core abilities: problem-solving, verbal and social competence. Hudson's theory of convergent and divergent thinking is also related with the idea emphasized above [Hudson, 19671. Convergent thoughts are those that lead to the solution of a concrete problem, while divergent thought are those that are prompted by stimuli. Convergent thoughts are related to mathematics and sciences. Divergent thoughts are related to creativity, arts and CHAPTER f. INTRODUCTION humanities. Every individual has a little amount of each component, and some individuals show preference for (or are more inclined to) a particular component. Problems appear when one of the components is not well developed. Al1 these findings suggest that social intelligence aspects must be addressed if we are to build robots that imitate human intelligence. What is more, social intelligence could be even more important than other capacities. There is evidence that in primates social intelligence is one important prerrequisite for the evolution of non-social, domain-independent intelligence [Dautenhahn, 19951. As an example, the highly elaborated ability of symbolization is suggested to be a social act of agreeing [Bullock, 19831. Some authors contend that social intelligence is also necessary for the development of generic intelligence in humans, see [Lindblom and Ziemke, 20031. This 'Social Situatedness Hypothesis' (also known as Machiavellian Intelligence Hypothesis) emphasizes the importance of designing and building robots that can interact with other agents. In this respect, the work of Brooks already showed the importance and benefits of this approach [Brooks, 19911. This document describes the experiences, ideas and techniques involved in building CASIMIR0 ', a robot with basic social abilities. CASIMIR0 is a robotic head that includes facial features and neck movements. The robot was not designed for performing a certain precise task. If any, its task would be to interact with humans (note the vagueness of that). At some point in the development process the option of restricting the robot's behaviour to a game or entertaining task was considered, although that was always discarded. For some members of our group (including me) this produced at times a feeling that the robot would end up doing anything. When 1 have that sort of horror vacui feeling, 1 think in babies. They do not seem to be doing anything, although they already have some physical and mental capabilities that are unattainable for a machine. 1.1 Previous Work In order to provide the reader with a more practica1 viewpoint on social robotics, in this section a brief description of the most influential social robots built is given. Not al1 of such robots appear here. Being an emergent field, their number seem to increase on a monthly basis. This section is intentionally short, as other robots will be referenced and described in the rest of the document. The following three major application areas can now be distinguished for interaction 'The name is an Spanish acronym of "expressive face and basic visual processing for an interactive robot" CHAPTER f. INTRODUCTION robots: + "Robot as a persuasive machine" [Fong et al., 20031: the robot is used to change the behaviour, attitudes or feelings of people. Also, any social robot that is used as a (bidirectional) user interface may also fa11 into this category, as would robots that are social for entertainment or demonstration purposes. + Robots that are designed and used to test theories of human social development or communication. This and the first type overlap and complement each other. Many social robots are inspired by theories of social development, and that effort is obviously in the direction of making the robot more persuasive, as observed by humans. + Robots that must interact/collaborate with people in order to accomplish the assigned tasks. Here the main task is not that of interacting or socializing with people. Collaborationlinteraction is just used to get the main task done. An example of the first type of robots can be found in the AURORA project, which uses robots in autism therapy [Dautenhahn and Weny, 20011. Potential groups of application are children and the elderly. Another such robot is the sea1 robot Paro, which has been tested at nursing homes and with autistic and handicapped children, and has been recognized by the Guinness Book of Records as the "world's most therapeutic robot" [Shibata and Tanie, 20011. Social robots that fa11 into the first category explained above find very fitting working scenarios in places like museums and exhibition centres. These guide robots may also fa11 into the third category (this depends on the importance of the robot itself as an attraction as compared to the importance of the museum tour task). Minerva [Schulte et al., 19991 is an interactive tour-guide mobile robot that displays four basic expressions, namely happy, neutral, sad and angry. A state machine is used to transition between those emotional states, in that specific order. The robot is happy when it can move freely and sad when people block its way. Perhaps one of the most attractive aspects of Minerva is its capability to learn how to attract people. It performs series of actions (facial expression, gaze direction and sound) and then evaluates them based on a reinforcement signal. The signal is positive when there is an increase in closeness and density of people around the robot. This shows the importance of adapting to the user's behaviour and, particularly, to maintain a memory that allows the robot to change its own behaviour. Kismet [Breazeal, 20021 has undoubtedly been the most influential social robot appeared. The most important robot that CASIMIR0 relates to is Kismet, and it was taken from the beginning as a model and inspiration (CASIMIRO's externa1 appearance is in fact CHAPTER f. INTRODUCTION very similar to that of Kismet, albeit this was not achieved intentionally). It is an animal-like robotic head with facial expressions. Developed in the context of the Social Machines Project at MIT, it can engage people in natural and expressive face-to-face interaction. Kismet was conceived as a baby robot, its abilities were designed to produce caregiver-infant exchanges that would eventually make it more dexterous. An overview of Kismet is available at [MIT A1 lab, Humanoid Robotics Group, 20031. Kismet is equipped with visual, auditory and proprioceptive inputs. The vision system includes four colour CCD cameras. Two of them have a larger field of view, and the other two are foveal cameras that allow higher-resolution processing. Kismet's gaze is controlled by three degrees of freedom in the eyes and another three in the neck. An attentional system based on basic visual features like skin tone, colour and motion allows it to direct its attention to relevant stimuli and gaze toward them. The auditory system processes signals gathered by a wireless microphone worn by the caregiver. The auditory system can recognize the affective intent of the speaker, i.e. it can recognize praise, prohibition, attention, and comfort. Kismet's face has 15 DOF that allows it to display facial expressions like fear, accepting, tired, unhappy, disgust, surprise, anger, stern and content. The vocalization system is based on a DECtalk v4.5 speech synthesizer in which parameters are adjusted to convey personality (Kismet babbles like a young child) and emotional state. Another MITrobot, Cog [Adams et al., 2000, Brooks et al., 1999, Scassellati, 20001, has a trunk, head and arms, for a total of 22 degrees of freedom. The head has 4 DOF in the neck and 3 DOF in the eyes. Its capabilities include human-like eye movements, head and neck orientation (in the direction of a target), face and eye detection, imitation of head nods, basic visual feature detectors (colour, motion and skin tone), an attentional system that combines them, sound localization, reflex arm withdrawal, shared attention (see below), reaching to visual targets and oscillatory arm movements. Recently, the work of Arsenio [Arsenio, 20041 has endowed the robot with significant high-leve1 capabilities such as object, face and scene recognition, acoustic segmentation and recognition, activity recognition, etc. Both Kismet and Cog were developed with the aim of exploiting scaffolding. Scaffolding is a teaching strategy introduced by Constructivist psychologist Vygotsky's sociocultural theory. Parents modulate and facilitate learning by creating an appropriate environment for the infant. Aspects like the use of simple phrases and achievable goals can facilitate learner's development. In the case of a robot, an instructor would pide interactions so as to foster novel abilities. The instructor may for example mark the critica1 aspects of the task or concept to learn, reduce the degrees of freedom, show the robot the effects of its actions with CHAPTER 2. THE CASE OFSOCIAL ROBOTICS extraordinary skills not exhibited by most people [Treffert, 19891. The most common forms involve mathematical calculations, memory feats, artistic and musical abilities. Autistic savants behave very much like computers (this should not be taken as a negative trait). From [Treffert, 19891 : "The savant develops into a well-trained 'robot' with little ability to change with changing conditions; in the face of change or challenge, the savant holds to obsessive, stereotyped, concrete responses." One of the best known cases of autistic savants is that of the twin calculators, John and Michael, who were born with the same physical and neurological deficiencies. Despite being autistic, they could perform certain arithmetic calculations at incredible speeds, see [Sacks, 19851. Autistic savants are very rare, the estimated prevalence of savant abilities in autism is 10%. Their skills are very concrete, restricted to rather limited domains, as evidenced in the most widely known account of an autistic savant: the Ruin Man movie (there are other two, Being there with Peter Sellers, and possibly Forrest Gump with Tom Hanks). There are many theories that seek to explain the strange phenomenon of autistic savants. Many researchers find the case extremely interesting because of the fact that abnormalities can help discover what is "normal" in human brains ('Everything we know about health we have learned from the study of disease': [Treffert, 19891). According to a recent theory proposed by Snyder and Mitchell [Snyder and Mitchell, 1999, Snyder and Thomas, 1997~1, autistic savants have not developed in their brains the part that gives meaning and interpretation to the things that we perceive. Autistic savants perceive "everything", they have an eye for detail, they perceive images and sounds with extreme precision. However, they have difficulty in attributing high-leve1 meaning to such detail. They struggle to isolate the important part of the data, the part that conforms a representation of high-leve1 concepts. That inability can render them socially incompetent. Healthy people can isolate useful data, and in fact we do that in order to survive, for we must make useful decisions quickly. We are concept-driven. Only with very hard training can healthy people manage detailed information and use it for example for drawing (our concept-driven thought can be seen as an advantage but also as a serious obstacle [Snyder and Thomas, 1997bl). According to Snyder and Mitchell, healthy people can also perceive with extreme detail, though as that ability is unconscious, they become more conscious of the final product, the concept ("we see what we know"), see Figure 2.2. On the CHAPTER 2. THE CASE OFSOCIAL ROBOTICS makes ahuman more machine-like '. This is not just unbalanced performance (i.e. abilities that are present in humans but not in the robot). The point is that autistic savants can easily accomplish tasks that require much effort for us (calculus, calendars, etc.) whereas they struggle to accomplish tasks that are trivial and almost automatic for us (like social tasks) 3. We try to build robots that look and behave like (healthy) humans. Then why do nonautistic people build machines that turn out to be "autistic"? On the other hand, in Artificial Intelligence it is a well-known fact that certain tasks that are trivial for humans are hard for computers/robots and vice versa. But why is it so? The following sections propose an explanation for these two important questions. 2.2 Unconscious Mental Processes From the discussion in the previous section now it seems clear that if we try to somehow model human intelligence and behaviour we will do it from an unavoidable concept-driven point of view. Some details will remain hidden to our awareness, for it seems that, in healthy adults, an enormous amount of processing is done unconsciously. Nowadays, the existence of unconscious processes in our brain seems to be beyond doubt [Wilson and Keil, 19991. Freud's work already acknowledged that unconscious ideas and processes are critica1 in explaining the behaviour of people in al1 circumstances. Helmholtz [H. von Helmholtz, 19241, studying vision, pointed out that even basic aspects of perception require deep processing by the nervous system. He argued that the brain constructs perceptions by a process of unconscious inference, reasoning without awareness Blindsight [Weiskrantz, 19981 refers to a set of residual visual functions, like the ability to locate a source of light, in cortically ("physically") blind patients. The cortical blindness is a result of the visual cortex's destruction, caused by tumours or other causes. The fact is that patients have those abilities, though they consistently claim not to see the stimuli. Moreover, the abilities are present even in unconscious (comatose) patients. Face recognition, an ability which obviously falls within the ambit of social intelligence, is another example of unconscious processing. We do not know what features in a face te11 us that the face belongs to individual X. Yet, we carry out that process fast and 'An interesting aspect to explore is the "Confederate Effect": in the annual Loebner's Contest (a version of the Turing test in which human judges try to distinguish chatbots from humans in a text conversation) some human interlocutors have been considered machine-like [Shah and Henry, 20051. 3Note that this confirms the two conditions for building a more "human" robot, already suggested in the previous chapter. That is, a) replication of a large number of human abilities and, b) that these abilities have similar performance levels. CHAPTER 2. THE CASE OFSOCIAL ROBOTICS robustly every day. Prosopagnosia is another rare neurological disorder characterized by a specific inability to recognize faces [De Renzi, 19971. As is the case in blindsight, patients with prosopagnosia show an ability to perceive and process the visual input of which they claim not to be aware. They see faces, though they can not recognize them. These patients are further evidence for the existence of unconscious processing. In linguistics something similar has also been observed. There is evidence that speakers unconsciously assign a structure in constituents to sequences of words. Besides, the rules that constitute knowledge of the language, i.e. the rules that enable us to produce grammatically correct sentences, are unconscious [Chomsky, 19801. On the other hand, studying knowledge representation, Bartletí was led to propose the concept of schemata: much of human knowledge consists of unconscious mental structures that capture the generic aspects of the world [Wilson and Keil, 19991. Many other examples of conscious/unconscious dissociation have been encountered. Researchers have demonstrated that this dissociation is present in perception, artificial grammar learning, sequence learning, etc., see [Cleeremans, 20011 for descriptions of experiments. These findings suggest that the presence of unconscious influences on behaviour is pervasive. Some authors argue that newborns, unlike adults, perceive the world with every detail, and only with their growing do concepts appear [Snyder et al., 20041. Infants have eidetic imagery, which tend to disappear as they grow. Experiments show that infants at 6 months can discriminate between monkey faces as well as between human faces, but not after 9 months. They also possess "absolute pitch (the ability to identify or produce any given musical tone without the aid of any reference tone). Absolute pitch is very rare in adults, although al1 musical savants and many individuals with autism possess it. The same authors argue that metaconcepts are also developed. Details (or the concepts that constitute the metaconcepts) tend to fade into our unconscious. In the context of learning, the "conscious competence model" explains the process and stages of learning a new skill (or behaviour, ability, etc.) (see [Kirkpatrick 19711). We move through conscious/unconscious levels of learning: 1. Unconscious incompetence: we are incompetent and ignorant of it. 2. Conscious incompetence: we are incompetent but we can recognize our incompetence. 3. Conscious competence: learnings that develop more and more skill and understanding. 4. Unconscious competence: the skill becomes so practised that it enters the unconscious CHAPTER 2. THE CASE OFSOCIAL ROBOTICS 1 Property Conscious 1 Unconscious 1 Engaged by Novelty, emergencies, danger Repetition, expected events, safety 1 1 Used in New circumstances 1 Routine situations 1 Table 2.1: Properties of the cognitive conscious and the cognitive unconscious (taken from [Raskin, 20001). Capacity Controls Persists for parts of the brain. Common examples are driving, sports activities, typing, manual dexterity tasks, listening and communicating. It becomes posible for certain skills to be performed while doing something else, for example, knitting while reading a book. The person might now be able to teach others in the skill concerned, although after some time of being unconsciously competent the person might actually have difficulty in explaining exactly how they do it (the skill has become largely instinctual). Why are some mental processes unconscious? Unconscious mental processes are fast, allowing us to do things like riding a bicycle without having to think about how to control each of our movements (see Tables 2.1 and 2.2). Some authors contend that it is practice and habituation what makes details go unconscious [Baars, 1988, Mandler, 19841. Only novel, informative inputs trigger conscious processing (or "mental attention", which is a scarce resource). When events are easily predictable, they are no longer conscious [Berlyne, 1960, Sokolov, 19631. This actually constitutes an attractive theory of learning: we learn when we habituate to certain details or stimuli that appear repeatedly. As we habituate we have less and less conscious consideration for the stimuli (i.e. we put less mental attention to those redundancies). Being consciousness a scarce -1imitedresource, this process would perfectly account for the progress of the individual that learning implies. This mechanism is very similar to the "chunking" technique used in the SOAR candidate unified theory of cognition, see [Newell, 19901. Chunking allows an agent to identify and store the mental processes that consistently lead to a goal. Later, the agent does not engage in mental processing for that task but it retrieves the stored pattern automatically. Chunking is meant to produce the "power of law of practice" that characterizes the improvements in human performance during practice. Tiny Volition Tenths of seconds If practice and habituation is what makes details go unconscious, then social abilities should be relatively more unconscious, as they appear first in life and are always present in the acquisition of other abilities (recall that in Chapter 1 we saw that the Social SituHuge Habits Decades ílifelonrl CHAPTER 2. THE CASE OFSOCIAL ROBOTICS Exvlicit inferences 1 Automatic inferences Conscious Explicit cognition Immediate memoty Novel, informative, and significant events Atiended information Foca1 contents Declarative memoty (facts) Supraliminal stimulation Effortful tasks Remembering (recall) Available memories Strategic control Table 2.2: Some widely used euphemisms for conscious and unconscious phenomena (taken from [B.J. Baars, 20041). Unconscious Implicit cognition Long-term memoty Routine, predictable, uninformative events Unatiended information Fringe contents (e.g. familiarity) Procedural memoty (skills) Subliminal stimulation Spontaneous/automatic tasks Knowing (recognition) Unavailable memories Automatic control atedness Hypothesis even proposes that generic intelligence evolves only after social intelligence has developed). In fact, unconscious processes are behind al1 or part of what we cal1 social abilities, like face and language recognition (as opposed to other mental processes like for example solving a differential equation, which require conscious effort) [Bargh and Williams, 20061. From [Wheatley and Wegner, 20011: "Much of our behavior in social life is unconsciously automatic. There is evidence that people can respond automatically and unthinkingly to facial expressions, body gestures, hints about a person's sex, ethnicity, or sexual orientation, information about someone's hostility or cooperativeness, and a variety of other social stimuli. People also have unconscious automatic responses to things they like and dislike, from foods or books to ideas and social groups. ... Many of the automatic behaviors we do every day are things of which we are perfectly aware at the outset. We know we are getting in the car and heading off to work, for instante, or we know we are beginning to take a shower. Yet because we have done the act so often driving to work every day, showering every darn year, whether we need it or not we no longer need to think about the act after we have consciously launched it. These behaviors are often acquired skills, actions that become automatic only after significant repetition. CHAPTER 2. THE CASE OFSOCIAL ROBOTICS When we begin to learn an action, such as driving, we think of the action at a very detailed level. We think 'engage clutch, move gear shifi down into second, lifi left foot off the clutch and right foot onto gas.' Ski11 acquisition starts off as labored, conscious learning and afier consistent, frequent practice becomes more automatic and unconscious. Once the action is well learned, the behavior becomes automatic in the sense that it does not require constant conscious monitoring. " At this point it seems clear that there must be a difference in the process involved in building precision manipulators or calculus machines and that involved in building a robot that can socialize with humans as we do. Note that many other problems in robotics and perception also suffer from our lack of conscious access to the mechanisms involved. However, this effect is comparatively more accentuated in social robotics for, as explained before, social skills appear earlier in life. 2.3 The Reason why Social Robots may not be Robust When building a social robot, for the most advanced abilities (face detection, face recognition, speech recognition, ...) researchers normally resort to machine learning techniques and, particularly, to supervised learning. Machine learning studies computer algorithms that improve automatically through experience [Mitchell, 19971. In supervised learning, given a training set (examples), algorithms can be built for providing outputs for novel samples (i.e. different from those in the training set). This capacity to learn from experience results in a system that can offer increased efficiency and effectiveness. Most importantly, it is the most appropriate option for cases in which we can not think of a concise relationship between input~ and outputs but there are many labelled examples available from which something can be learned. Learning algorithms use the available training samples to pide a search for a solution in a hypothesis space H,. A robust solution has to predict, i.e. produce the correct output for future, unseen samples. Training samples alone are not sufficient for this, for there is a large number of posible hypotheses that fit the training set. This is the essential idea behind Wolpert's No Free Lunch theorem (see [Wolpert, 1996]), which states that, on the criterion of prediction performance, there are no reasons to prefer the hypotheses selected by one learning algorithm over those of another. The perfect fit to a training set does not guarantee low error for future, unseen sam- CHAPTER 2. THE CASE OFSOCIAL ROBOTICS ples. What about the available knowledge about the solution? A priori knowledge about the solution to the task allows to define the hypothesis space. We may know, for example, that the solution is a polynomial function. Better still, we may know that the solution is actually a quadratic. The more knowledge we have of the form of the solution the smaller the search space in which our learner will look for a candidate solution: Ho> Hl > H2 > ... > H' H* being the objective hypothesis (the solution). Virtually al1 "practical" learners employ some sort of complexity penalization technique [Scheffer, 19991, including the method of sieves and Bayesian, Minimum Description Length, Vapnik-Chervonenkis dimension and validation methods [Nowak, 20041. The basic idea of complexity penalization is to minimize the sum of the error in the training set and a complexity measure of the hypothesis space. The quantity that we are ultimately interested in is the risk: a being the parameters of the learner (to be set). The risk is a measure of discrepancy between the values f (x, a) produced by our learner and the objective values y produced by H*. Theoretical work by Vapnik and others has allowed to obtain upper bounds on this risk [Vapnik, 19951. With probability 1 - 7 the following bound holds: where e,,, is the empirical risk (the error on the training set), and h a non-negative integer called VC dimension. The VC dimension of a set of functions H,, a measure of its complexity, is the maximum number of training points that can be shattered4 by H,. Note that we can in principle obtain an e,,, as low as we want, though we would have to do it by making our hypothesis space more complex (i.e. by having a large number of degrees de freedom). Alternatively, we can consider only simple hypotheses, but then we would obtain large e,,,. 4for a 2-class recognition problem and a set of 1 training points, if the set can be labelled in al1 posible 2' ways, and for each labelling a member of the set H, can be found which correctly assigns those labels, then we say that the set of points is shatteved by that set of functions. 20 CHAPTER 2. THE CASE OFSOCIAL ROBOTICS From the definition of VC dimension the hypothesis spaces considered above satisfy: The second term of the right-hand side of Equation (2.3) is monotonically increasing with h (see [Burges, 19981). Thus, if we call: In Equation (2.3), e,,, is the error obtained in the training set for an individual hypothesis. Let eemPM,,(Hi) be the maximum training error that can be obtained in H, (i.e. the maximum training error of al1 the individual hypotheses in H,). Then: Let us assume that the error in the training set for the objective hypothesis H* is O. From (2.1): From (2.8) the following holds: From (2.7), (2.6) and (2.9) we can conclude that the risk decreases as we increase the amount of howledge. With reduced howledge good performances cannot be guaranteed. Therefore, in our case in which available howledge is little, we cannot guarantee good performance for samples different than those tested. Thus, if we obtain low error in a test set we could be actually overfitting to that test test! '. The involved inductive process 'In the discussion we have not mentioned the effect of the samples used. Obviously, if the test samples are "representative" then the error measured is applicable to the future samples that the system will encounter. As an example, let us suppose that the domain is made up of a very frequent set of samples (set A) plus a very infrequent set of samples (set B). Thus P(x) is much larger for values of x inside A. Then if we measure low error CHAPTER 2. THE CASE OFSOCIAL ROBOTICS is valid in predicting future cases only as long as we are capturing howledge about the solution, whether intentionally or not. The conclusion is: in any attempt to reproduce a social ability in a certain domain there is a serious risk of overfitting to the test cases. This means that in practice we will have a situation like that depicted in Figure 2.3. Original domain error Figure 2.3: Situation when there is little howledge about the form of the solution. We can be in one of three possibilities: 1) have low error in a domain different than the original, 2) work in the original domain but with large error, or 3) have low error in the original domain but in that case solving a problem different than the original. The reasoning path taken in this chapter has been the following. First, instead of autism, the condition of autistic savants was seen as more analogous to the social robots that we build, based on the unbalanced intelligences and capacities that are observed. Studies on the origin of that condition show that there are unconscious mental processes that provide healthy adults with the final useful products of reasoning: concepts. Finally, unconscious mental processes, useful as they are, are formally demonstrated to hinder our capacity to reproduce some of our own social capacities. Based on these ideas, namely the unavoidable lack of robustness (specific to tasks that fa11 into the social range of human abilities) and the fact that we are concept-driven, a way to approach the design of social robots is outlined in the next chapter. Note that the ideas developed in this chapter constitute a coherent explanation of the well-hown fact that certain tasks that are trivial for us are hard for computers/robots and in the set A we can obviously expect low error in the future. However, suchreasoning is essentially equivalent to saying that the more (labelled) samples we know of the domain the more reliable the error figure measured on those samples. But knowing many samples of the domain is actually equivalent to knowing somethng about the solution function, only that it would be knowledge by enumeration, instead of by comprehension. Having little knowledge @y comprehension) of the solution is just equivalent to knowing only a very small part of the posible samples in the domain. CHAPTER 3. APPROACH AND ARCHITECTURE Moisture l 1 fundamental niche 1 . Temperature Figure 3.3: Example of fundamental and realized niches. An organism can live under a potential range of moisture and temperature conditions (i.e. they are required for survival). The realized niche is the range of conditions that the organism actually utilizes in its habitat. The robot, however, is always expected to work in a much more restricted environment, its realized niche. Some particular cases that we ourselves encounter in our daily lives will never appear in the niche the robot is to work in. Obviously, throughout the robot development process we must be interested in minimizing error in its realized niche. If we try to reproduce the ability as in the fundamental niche, with simple implementations we may obtain improvements in the realized niche. However, due to our limited knowledge, if we go too far and use too much of the available knowledge or intuitions we could be actually worsening performance in the realized niche. That, in fact, is something relatively frequent. Consider again the problem of face detection. If we reflect on how we detect faces, we would say that by detecting the presence of certain features simultaneously, such as skin colour, eyes, mouth, etc. Sophisticated detection systems exist that use these and other features. However, it is our experience that only skin colour blobs with certain widthheight ratios are sufficient to detect faces in many practica1 cases. Using a more elaborated or intuitive algorithm without extensive testing and tuning in situ usually leads to many typical faces going undetected (i.e. false negatives). Note that the search for implementations that work well for the robot's realized niche is not an easy one. We still have the same lack of conscious knowledge about the posible solution to the task. However, the more narrow the realized niche the more relative descriptive value have the available cases used for testing and tuning the system. In the limit, if the realized niche was just a single case, knowing the output value for that case would be obviously sufficient. CHAPTER 3. APPROACH AND ARCHITECTURE Note also that, as long as the available knowledge is still too little, the more of it we utilize the better the performance in parts of the fundamental niche, but also the more "patchy" the performance inside the realized niche. This is due to the fact that in our case there rnay be overfitting to the samples used for testing (see Section 2.3). Such patchy performance is obviously undesirable. Ideally, we would like to measure low test error inside the realized niche and at the same time have some guarantee that this measured error applies to the whole niche. The more knowledge we use the more likely it is to measure low test error inside the niche, but there is still little guarantee that such error applies to the whole of it. Thus, large errors rnay be experienced in untested cases. This could actually mean that the global error experienced in the niche rnay be very large. Using only basic knowledge that situation will be less frequent. This suggests that basic knowledge will generally be applicable to most of niches, whereas the use of more of the available knowledge could lead to major improvements only in few specific niches. This, in turn, suggests that it would be advisable to make the system as simple as possible. Note that this is not the same as saying that simpler techniques should perform better than more complex ones. The point is that complex implementations should be carefully adapted to the realized niche. The best way to achieve this is to start from simple implementations. Using an analogy, it would be like a custom suit. The tailor starts by using a standard pattern of sizes (which fits everybody relatively well) and then, through successive fittings with the client, adjusts different parameters to produce a better final product. In the end the suit will fit the client very well, although it will probably not be adequate for other clients. It is known that, on an evolutionary scale, organisms tend to adapt to its realized niche (see for example [Kassen and Bell, 19981). If we concede that robot designers rnay at times take a role similar to evolution, we will have to achieve this too. Therefore, for the kind of robots that we pursue, we see that the role of the robot niche is crucial. The robot designer has to make a significant effort in discovering the opportunities in the robot environment that allow to obtain useful implementations for the desired abilities. This rnay involve resorting to assumptions and algorithms that rnay seem even unintuitive. In other words, it rnay involve discarding intuitive ideas that seem to be useful. Obviously, if such opportunities are not found, then the designer will have to define the minimum set of restrictions to the problem that allows to obtain useful implementations. For the reasons given above, in this work we have adopted a parsimonious and opportunistic synthetic approach. That is, our interest is in developing an application with whatever means. The main source of insight will still be the findings of other disciplines like psychol- CHAPTER 3. APPROACH AND ARCHITECTURE ogy or ethology. However, we will use that knowledge only as long as it allows to obtain practica1 results. As we have seen, using too much knowledge can be counter-productive. Therefore, in building the robot, we start by considering simple techniques, simple features, simple algorithms, even though they may seem unintuitive. Then, if needed, more elaborated versions are tested. As soon as the results are satisfactory in the working scenario the technique is implemented as final. Evolution has allowed some species to have organs perfectly adapted to their environment. Actually, there has been no adaptation; only those individuals who were adapted survived. What evolution has made through natural selection, the robot builder has to achieve through a teleological, and also iterative, design perspective. In such approach low-leve1 design aspects become more important. Issues like the robot niche and how the design is adapted to it in order to achieve a desired function are now fundamental, to the point of defining the value of the robot. 3.3 Design and Niche Spaces Design and niche spaces are two concepts introduced by Sloman [Sloman, 19951 that also fit very well in our discussion, see Figure 3.4. The design space is the space of al1 the posible designs for the robot. Obviously, design spaces are huge and generally very complex, for they may include aspects like part sizes, types of motors, process communications, architectures, etc. The niche space implicitly defines the requirements. design space niche space Figure 3.4: Design and niche spaces. Sloman sees A1 as the study of design and niche spaces and their interrelations. One design, for example, may fit a niche in different ways. It could be simply good or bad, or CHAPTER 3. APPROACH AND ARCHITECTURE we could say that it is good with respect to battery charge, or with respect to cost. This is indicated in Figure 3.4 with different styles of arrows. Thus, a design can be considered good for more than one niche. Besides, both spaces can be seen at different abstraction levels (the lowest level of the design space, for example, could be the implementation). The opportunistic synthetic approach would consist of turning the robot building process into a design space exploration. It would also be an exploration in the niche space, as long as some requirements may be discarded along the process for being too difficult to fulfil. In that case, of course, the less stringent the new requirements the better. Note that we are not referring here to the exploration being made by many authors, through hundreds of papers and built robots. The scale of the exploration we refer to is that of a specific robot, from the early design stages to its final implementation details. In such exploration we should look for a tight coupling between the robot design and implementation and its niche. The search for such coupling must start from simple versions, which are not associated to a particular niche (or, in other words, that are relatively valid for a large number of niches). Then, the exploration should develop and test new versions that may use knowledge accumulated through analysis but that, above all, exploit the characteristics of the specific niche in which the robot is to work. In order to take into account both analysis and synthesis, we could use in the discussion both a robot design space and ahuman design space, see Figure 3.5. The latter would consist of whatever space one could imagine to represent the knowledge acquired about ourselves through analysis (neurophysiological level, cognitive level, etc.). Our human design for accomplishing a certain task (whatever it may be) is such that we achieve good performance under a broad range of circumstances. It is so good a design that we of course can achieve good performance on a narrower range of circumstances, like that of the specific niche of the robot. The design for the robot, however, has to be coupled (in a sense, overfitted) to its niche. This will give good results in that niche, although the solution may seem at times unintuitive or contrary to the knowledge acquired through analysis. We want to emphasize that this tight coupling is necessary because, in our context, our knowledge about the solution to the desired task remains relatively poor. Should our knowledge be richer, we could provide solutions that do not rely so much on specific features of the niche. In a sense, we have no option but to overfit the system to its niche, which may involve discarding intuitive approaches, algorithms and representations. Much of the designer effort will revolve around this overfitting. CHAPTER 3. APPROACH AND ARCHITECTURE designer O Robot design space /Tx ANALYSIS w Human design space m---+ bad performance, possibly unintuitive solution + goodpeifonnance specific niche , ofthe robot (rdied niche) ' fundamental niche of the human task to reproduce Figure 3.5: Tight coupling to the specific niche of the robot. 3.4 Design Principles In the previous sections we have argued for a kind of opportunistic synthetic approach. The characteristics ofthe social robotics problem, described in the first chapters, led us to propose that methodology. It can be summarized as shown in Figure 3.6. Such approach obviously involves the use of many heuristic ideas, especially those obtained after extensive experience in the hands-on design and testing of social robots. To us, that experience, which includes many engineering insights, is one of the most valuable aspects of the research literature on the topic. In fact, the writing of this document has been heavily influenced by that idea. Besides the guidelines of Figure 3.6, we consider particularly appropriate for our context the set of principles introduced by Pfeifer and Scheier for designing autonomous agents [Pfeifer and Scheier, 19991. Those authors, who also advocate for a synthetic approach, have extensive experience in building robots that try to reproduce interesting human capacities. From that experience that they have accumulated, they extracted a set of useful principles, CHAPTER 3. APPROACH AND ARCHITECTURE Make an effort in discovering opportunities for improving performance in the niche of the robot. This can lead to unintuitive implementations (or, in other words, can cal1 for originality). Proceed from simple to complex algorithms and representations. The final implementation should be as simple as possible. Perform extensive testing of the algorithms and representations in the niche of the robot. Adjustments to the (or selection of new) algorithms and representations should be guided by the the results of these tests. Treat available human knowledge very cautiously, and always following the two previous guidelines. Basic knowledge will almost always be applicable r. D D N to most niches. Detailed knowledge may be appropriate only for few specific .- 6 L m niches. c L" > .- c 7 - m Figure 3.6: Summary of the opportunistic synthetic approach. a - .- - e m which we summarize in Table 3.1. These principles, in turn, overlap in many aspects with the design principles established by other authors like Brooks [Brooks and Stein, 19941 and Maes [Maes, 19891. Principle 2 indicates that the robot should be: autonomous: it does not require human intervention while it is working (independence of control). self-sufficient: it can perform multiple tasks, can exhibit multiple behaviours in the real world over extended periods of time; that is, they do not incur an irrecoverable deficit in any of their resources. embodied: it has a physical body embedded in a real environment. situated: it has sensors that allow it to acquire information about the environment CASIMIR0 fulfils to different degrees these four requirements. In particular, its physical body is one of the most important values, for it has a clearly visible impact on visitors. This implies that the whole process should be a hardware-software codesign. The fourth requirement is in turn related to the principle of sensory-motor coordination (4), which argues that interesting behaviour requires coupling actions to sensory input. This, in turn, involves extensive sensory processing (that, following principle 6, comes from CHAPTER 3. APPROACH AND ARCHITECTURE Name The three-constituents principle The complete-agent principle The principle of parallel, loosely couple processes The principle of sensoty-motor coordination The principle of cheap designs The redundancy principle The principle of ecological balance The value principle Summay Designing autonomous agents always involves three constituents: 1) definition of ecological niche, 2) definition of desired behaviours and tasks, and 3) design of the agent. The agents of interest are the complete agents, i.e., agents that are autonomous, self-sufficient, embodied, and situated. Intelligence is emergent from an agent-environment interaction based on a large number of parallel, loosely coupled processes that nin asynchronously and are connected to the agent's sensoty-motor apparatus. Al1 intelligent behaviour is to be conceived as a sensoty-motor coordination that semes to stnicture the sensoty input. Designs must be parsimonious and exploit the physics and constraints of the ecological niche. Sensoty systems must be designed based on different sensoty channels with potential information overlap. The complexity of the agent has to match the complexity of the task environment. The agent has to be equipped with a value system and with mechanisms for self-supemised learning employing principles of self-organization. Table 3.1: Design principles of autonomous agents proposed by Pfeifer and Scheier. many channels, both visual and auditive, some of them redundant). The paradigm of active vision [Aloimonos et al., 1987, Ballard, 19911, where movement is considered to be an integral aspect of the perceptual process, is thus very useful here. In the previous section reasons were given for the importance of carefully adapting the design to the robot niche. The principle of cheap design (5) embodies this aspiration. Although the main justification for the principle seems to be the application of the Occam's razor paradigm, we have already shown that in our context such approach is more than justified and it has been, in fact, one of the most important guidelines in the building of our robot CASIMIRO. The principle of ecological balance (7) is closely related with the principle of cheap design. CHAPTER 3. APPROACH AND ARCHITECTURE The modules that constitute the robot's software are a manifestation of the principle of parallel, loosely coupled processes (3). In this case, it is the simultaneous, asynchronous performance of a number of processes (perception of people, face detection, sound perception, ...) that give the robot its current externa1 appearance and behaviour. There is still a "central" module that decides which actions the robot executes, though many of these processes accomplish high-leve1 tasks on their own. As an example, the audio-visual attention system, itself implemented in anumber of modules, is able to modulate the robot's behaviour in a way directly observable by the individuals interacting with the robot. While it is not used for learning, the value principle (8) is also present in the robot. In fact, it is very important that this robot, whose main function is to interact with individuals, has a way to evaluate the current status of the interaction session. In CASIMIR0 the value system is represented mainly in the emotional module, the effect of which is directly visible to the observer. This system allows the robot to have a means of judging what it is good and what it is not in the interaction. Thus it can be considered the basic capacity to develop further abilities. 3.5 Architecture In robotics, an architecture is the high-leve1 description of the main robot components and their interactions. The choice of an architecture already defines some of the most significant properties of the system. After introducing the design approach and basic principles, in this section we start to take design decisions in earnest. This part of the document will also serve the reader to get an overall impression of the robot software suite, with details left for subsequent chapters. Social robot architectures generally make use of anthropomorphical models taken from human sciences. Independent of the validity or explanatory power of those models, they are always a good starting point to divide the (often complex) design process into manageable parts. In CASIMIR0 we have taken advantage of that aspect and thus we have made extensive use of abstractions like emotions, memory (as an analogy of human memory), habituation, etc. In this respect, it is important to consider the ideas developed in the previous chapter which showed that healthy people mainly work with high-leve1 concepts. "Robotic" traits are present in individuals who, for one reason or another, are unable to form and work with such high-leve1 concepts or abstractions. Thus, an additional objective would be to endow the robot with enough expertise to recognize and use high-leve1 concepts, generalizations, CHAPTER 3. APPROACH AND ARCHITECTURE interpretations. Note that by high-leve1 concepts we mean those concepts normally used by humans. The abstraction is relative to the concepts that humans have. Figure 3.7 shows the high-leve1 block diagram of CASIMIRO's architecture. Figure 3.7: Robot architecture. Note that, muchas a "standard suit", the different elements ofthe architecture are very common in robotic architectures, especially the triad (perception, action , behaviour control). The following paragraphs describe and justify the presence of these and other elements. Perception As mentioned in Section 3.4, one of the principles that should guide the robot design is that of sensory-motor coordination. The principle emphasises the use of advanced perceptual and motor capabilities. If we think in terms of high-leve1 concepts, it is straightfonvard to see that useful high-leve1 concepts can only be identified if the appropriate sensors are used. Poor sensory data can not lead to rich high-leve1 concepts. Therefore, the robot has been given the capacity to process data from two rich perception channels: visual and auditive. Figure 3.8 shows a block diagram of the perception abilities in CASIMIRO. An important module is devoted to the detection of individuals in the interaction area. This is CHAPTER 3. APPROACH AND ARCHITECTURE obviously a fundamental capacity for the robot, which is expected to engage in interactions with people around it. Individuals may move around the room, and so they are also tracked during the interaction session. Both capacities will be described in Sections 5.1 and 5.4. To Memory, Behaviow Contml, Emotim and Actim Fmm Behaviour Control Attention 6 - FmmEmotion n Auditive Pmeptim Figure 3.8: Perception. Face processing is itself divided into two functional modules. One of them simply detects faces in the images taken by the robot "eyes". The other module is devoted to processing the face zone of the input images to detect head nods and shakes. Head nods and shakes are interpreted as yeslno responses to questions made by the robot. Although minimal, this is a valuable capacity, for it allows the robot to have a direct way to know if the interaction is going well. Moreover, it is a capacity that can be fulfilled in practice with good performance. This design decision was actually taken after weighing different niches for the robot. The ideal niche would be that of a normal conversation. Speech recognition, however, did not fulfil our performance requirements in practice and so, following the approach introduced in Chapter 3, we opted for detecting affirmativelnegative face gestures. Head nodlshake detection will be covered in detail in Section 5.5. Note that person and face detection are in a sense redundant perceptions. However, given the fact that people are the most important environment object for the robot, this is Chapter 4 Hardware Overview "Your face is a book, where men muy read strange matters" William Shakespeare. Paraphrasing words of Kofi Annan: intelligence is the only thing that is equally distributed al1 over the world. Robotics is nowadays a topic more or less known by the general public. People al1 over the world get captivated by its appeal. They want to build their own robots, fascinated by Kismet or Cog. DYI books are also beginning to appear [Williams, 20041. The ingenuity is in most of the cases in the way they manage to fulfil the hardware or software requirements: "The main objective of this project is to put al1 of the mentioned components together and show the possibility of developing such a complex platform from scratch with very elementary and low-cost components with aid of basic tools", H. Mobahi (builder of the robot Aryan [Mobahi, 20031). In some cases it is quite an achievement: "Well, after seeing al1 of this, this (my) robot is not so bad, if we bear in mind that it has not been sponsored by any firm, government or university. It is the result of work performed by a Spanish technician keen on robotics and with limited economical resources.", J.L. Martinez (builder of the robot SEGURITRON [Martinez, 20031) CHAPTER 4. HARDWARE OVERVIEW CASIMIR0 also shares that spirit. The implementation details are the most difficult part of any robotic project. In the abstract everyone can devise fancy architectures or capabilities, though the implementation is always the ultimate effort. This in fact is one of the ideas of the approach outlined in the previous chapter: fitting the implementation to the robot niche. It is in the implementation where one has to choose between options and make a decision. For someone involved in developing a project, implementation details are the most rewarding information of the available bibliography or web resources. Often, that information is what keeps them excited: "Years ago when 1 was a wee laddie fresh out of college 1 would read academic papers on robotics projects procured at some expense and often taking weeks or months to arrive as crude photocopies from the British Library. Papers such as Brooks' Elephants don'tplay chess, together with others by Luc Steels, Inman Harvey, Cynthia Ferrel and many more were al1 fascinating to read, but there was always an absence of detailed hardware and software information which meant that getting started presented a big obstacle to overcome. It wasn't possible to reproduce the experiments and find out how other people's robots worked, because the details of particular implementations were usually kept secret, or were simply not described out of laziness. To me this lack of detailed description seemed very unscientific, opening up the possibility of researchers making exaggerated or misleading claims about the results of their experiments. So for my own robotics projects I'm determined to make the whole process as transparent as possible, making both source code and details of electronics and physical construction available wherever possible." - Bob Mottram (builder of the robot Rodney [Mottram, 20031). Again, CASIMIR0 and particularly this document, were designed with those ideas in mind. The hope is that some day they can be useful to other people who want to go a little farther. The following sections describe the hardware of the robot. Details will be in general left out as the information is mainly technical data available elsewhere. It is important to introduce the hardware at this point of the document (instead of considering it an appendix). That helps to define in part the robot niche, to which the rest of the work will have to adhere. CHAPTER 4. HARDWARE OVERVIEW 4.1 Hardware If we are to build an anthropomorphic face, would not it be easier to use a graphical face? The Vikia robot [Bruce et al., 20011, for example, has a flat screen monitor in which an animated face is displayed. Kidd and Breazeal [Kidd and Breazeal, 20031 compared people's reactions to a robot and to an animated (flat) character. The results showed that the robot consistently scored higher on measures of social presence and engagement than the animated character. Subjects also rated the robot as more convincing, compelling and entertaining. On the other hand, Bartneck [Bartneck, 20031 found that embodiment had no significant influence on enjoyability. However, in the robotic character (as compared with a screen character) a social facilitation effect and a high forgiveness for speech recognition errors was observed. In any case, a physical robot can be viewed from different angles, it can be touched and it is part of the space occupied by people. People expect that moving 3D objects require intelligent control, while flat images likely result from the playback of a stored sequence as in film or television [Mobahi, 2003, King and Ohya, 19961. CASIMIR0 is a physical face: a set of motors move a number of facial features placed on a metal skeleton. It also has a neck that moves the head. The current aspect of the robot is shown in Figure 4.1. The study in [DiSalvo et al., 20021 tries to look for certain externa1 features of robot heads that contribute to people's perception of humanness. Humanness is desirable, though a certain degree of "robot-ness" is also needed to avoid false expectations. The study analysed 48 robot heads and conducted surveys to measure people's perception of each robot's humanness. The authors of the study give the following 6 suggestions for a humanoid robotic head: + To retain a certain amount of "robot-ness" the head should be slightly wider than it is tall. + The set of facial features (nose, mouth ...) should dominate the face. Less space should be given to forehead, hair, jaw or chin. + To project humanness the eyes should have a certain complexity: surface detail, shape, eyeball, iris and pupil. + Four or more facial features should be present. The most important features are nose, mouth and eyelids. CHAPTER 4. HARDWARE OVERVIEW Figure 4.6: Wireframe design of the head. Figure 4.7: Tilting carousel used as a neck. Courtesy of Rhino Robotics Ltd. Stereo Cameras Most of the visual tasks accomplished by the robot depend on a pair of cameras placed just above the nose. It is a STH-MDl-C FireWire stereo head fiom Videre Design. The device was selected because of its high image quality and speed. Also, it includes a library for efficiently obtaining depth maps. Microphones, Amplifiers and Sound Card CASIMIR0 has two omnidirectional microphones, placed on both sides of the head. Sound signals feed two amplifiers. An EWS88 MT audio system from Terratec is used to capture CHAPTER 4. HARDWARE OVERVIEW Each module can communicate with other modules through TCPIIP sockets. A configuration file stores the association between module, IP address of the machine in which it runs, and listening socket. A special module called Launcher performs al1 the necessary initialization steps and then runs al1 the modules in both computers. It can also stop al1 the modules. CHAPTER 4. HARDWARE OVERVIEW CHAPTER 5. PERCEPTION we could distinguish different sounds with one ear alone, pinpointing where the sounds are coming from requires at least two ears. Reliably localizing a sound source in 3-D space requires even more hearing sensors. Sound differences between the signals gathered in our two ears account for much of our sound localization abilities. In particular, the most important cues used are Interaural Leve1 Difference (ILD) and Interaural Time Difference (ITD). ILD cues are based on the intensity difference between the two signals. This intensity difference, which can be of up to 20dB, is caused mostly by the shading effect of the head. ITD cues are based on the fact that sound coming from a source will be picked up earlier by the ear nearest to the sound source. This difference will be maximum when the sound source is directly from one side, and minimum when it is in front of the head. Both ILD and ITD cues are dependent on the sound frequency. ITD cues are reliable for relatively low frequencies (up to 1 Khz, approximately), while ILD cues are better for higher frequencies (see [GCAT, 19991 for an explanation of this). Humans use additional cues for sound localization. The shape of our head and outer ears affect received sounds in a manner dependent on arrival angle and frequency. A model of this process referred to in the literature is the Head Related Transfer Function (HRTF). HRTF-based cues allows us to obtain an estimate of the sound source elevation and also to distinguish between sound originating in front of and behind the listener. A more detailed description of sound localization mechanisms can be found in [Blauert, 1983, Yost and Gourevitch, 1987, Hartmann, 1999, GCAT, 19991. These and other physiological findings have been emulated in computer-microphone systems with relative success. Sound localization can play an important role in humanmachine interaction and robot interaction with its environment. 5.2.1 Previous Work The first important work on a computer sound localization system is [Irie, 19951. With a combination of hardware and software the system aims to learn to localize sounds in complex environments. The output of the system can be one three values: frontal, right and left. Both ILD and ITD cues are extracted from signals gathered from two microphones and a preamplifier circuit. Signals were previously high-pass filtered to remove background noise and then they were divided into segments. For each segment, the cues extracted are: difference of the two maximum positive values, difference in the positions of these maxima, delay between signals (computed by performing a cross-correlation of both signals), difference in the sum of magnitudes of the signals and filterbank-based cues. Filterbank-based cues are computed CHAPTER 5. PERCEPTION by dividing the spectrum of the signals in a number of equally spaced banks, computing the sum of magnitudes in each bank. The cue itself is the difference between the sums of the two signals. 4 banks were used, so the complete feature set used had 8 cues. These cues were fed into a feedfonvard multi-layer perceptron with three outputs. This network was trained and tested using three sounds (hand clap, spoken "ahh and door slam). This system is currently working on the Cog humanoid robot at MIT l. In [Alexander, 19951 a similar system is introduced. The input signals were divided into segments. The extracted cues were: difference between maximum values, difference in the positions of the maxima, correlation and difference in the sum of magnitudes. A classifier was not used, the output of the system was programmed (basically by means of comparing values). This can be a disadvantage in certain settings, because many thresholds have to be manually found (think for example that the difference in intensities could not be exactly zero for a perfectly frontal source, because the two microphones andlor pre-amplifier circuits could have different gains). A workthat used only one ITD cue is described in [Reid and Milios, 19991. Afterperforming high-pass and low-pass filtering, a signal leve1 test was performed to discriminate between sound and silence. After that, correlation was performed to obtain the ITD estimate and another threshold test was performed on its result (based on the ratio peakíaverage correlation values). Finally, in order to discard outliers, many estimates were gathered before giving their median value as a response. Correlation was only computed for the possible range of temporal displacement values (as the sound speed is finite, there is a maximum delay possible in the ITD cue, and it depends on the distance between microphones), and this in turn allowed for a faster response. The output of the system was an angle, and it was tested with two types of sound (impulsive sound and speech). For examples of simulated auditory models or systems that use more than two microphones or special-purpose hardware see [Rabinkin et al., 1996, Harma and Palomaki, 19991. For the use of sound localization for robot positioning see [J.Huang et al., 1999, Ryu, 20011. 5.2.2 A Study on Feature Extraction for Sound Localization In this section the system described in [Irie, 19951 has been used as a base line for comparison, as it uses both ITD and ILD cues and has found practica1 use. We describe here a new cue extraction procedure that can eliminate some minor errors. The extracted cues for a computer sound localization system are always subject to error because of background noise, 'Prof. Rodney Brooks, personal communication CHAPTER 5. PERCEPTION the maximum is searched for MZ MZ here L L the maximum is searched for here Figure 5.5: a) M1 does not fa11 in the initial or final "dangerous" zones, b) M1 falls in the "dangerous" zone, c) both M1 and Mr fa11 in "dangerous" zones. In the last case the sample is discarded. electrical equipment noise, and specially echoes and reverberation. Echoes originate when sound signals reflect off planar surfaces. Often the effect of multiple reflective sound paths can be as loud or even louder than the sound travelling a direct path from the source. An important fact to consider is the effect of using segments of the input signals. Al1 systems described in Section 5.2.1 divide the input signal in segments, and extract features from these. However, none of the systems described consider problems that could arise at boundaries. If we consider for example the first extracted cue, difference of maximum positive values, the maximum of signal L (left) could be just at the beginning of the segment. If the source is on the right side, signal L will be delayed with respect to signal R (right). Thus the maximum of signal R is not associated with the maximum in signal L. This in turn affects the second extracted cue, the difference in maximum positions. We propose to extract the first cue as follows. The maximum of signal L is found, be it Ml. Then we search in signal R for the maximum in a zone around the position of Ml. The zone has a length of 2W, where W is the maximum possible interaural delay. The value of W depends on the distance between microphones and the sound speed. Any (correct) ITD estimate must be equal or lower than W (in absolute value). If M1 falls in the initial zone of the segment, of length W, or in the final zone of the segment, also of length W, it is discarded and we repeat the procedure beginning with signal R. If the maximum of signal R, Mr, also falls in one of these "dangerous" zones, and the zone in which it falls is different from that of Ml, the segment is discarded (no cues are extracted from it). Figure 5.5 shows the three possible cases. This way, some segments are not used for localization, though the first (and second) cues extracted for other segments should be more reliable. CHAPTER 5. PERCEPTION explained in Section 5.2.1, an additional test was used in [Reid and Milios, 19991 in order to reduce such (and other) errors in the ITD estimate. If the ratio between peak and average values of the correlation result was lower than a threshold, the sample was rejected. In our system that test is used too, though the sample is never rejected. If the obtained ratio is lower than the threshold, the value of the ITD cue for the sample is substituted by the last higher-than-the-ratio value obtained. As the sample is not discarded, this allow us to take advantage of the useful ILD information in the sample. The same mechanism was used for the second cue (difference in the positions of the maxima): if the correlation ratio is lower than the threshold, the value of the second cue is substituted by the last second cue value obtained in which the correlation ratio was higher than the threshold. In order to test the method, sounds were recorded using two Philips Lavalier omnidirectional microphones, pre-amplifier circuits and a professional sound card (see Section 4.1). A DirectX application was developed to integrate al1 the processing stages: low-pass filtering, sound source detection, feature extraction, data recording and playing (for off-line analysis), and classifying (see Figure 5.8). In the experiments, the two microphones were placed 28 cm apart on both sides of a custom-made plastic head (see Figure 5.9). Four different sounds were used in the experiments: hand claps, a cal1 tone from a mobile phone, a maraca and a whistle, see Figure 5.10. The objective was to detect if the sound was coming from the left, right or frontal side. Sounds were recorded in front of the head and at between 35 and 45 degrees on the left and right sides, at a distance of at least one meter to the head. As indicated before, we have compared our feature extraction method with that used in [Irie, 19951, which will be referred to as 'Cog'. In order to study the reliability of the extracted cues, the ratio between inter-class to intra-class variances will be used as a measure of overlap between samples: This is actually the Fisher criterion for feature evaluation. A classifier was not used because our interest is only in the error present in the individual extracted features. The larger the ratio the better the separation of the samples for a given feature. On the other hand, a number F of consecutive cue vectors was extracted and the mean of them was given as features. The results obtained for F=250 are shown in Table 5.1. The results obtained with the proposed method achieve in general a higher separation CHAPTER 5. PERCEPTION Extract cues GClassify cues Localkation Store cues vector and associated localkation label Figure 5.8: Steps performed by the developed sound localization module. The work described in this section focuses on the cue extraction stage. ratio for the four features used. Note that this results are achieved with a high value for F. In Table 5.2 the same values are shown, for F=O (cue values are not averaged). In this case, the ratio values for the normalized cues (1 and 4) are worse in the first and second sounds. As F is low, the error is higher and it could be amplified. This reflects negatively in the two first sounds because these sounds contain no significant changes in volume. The other two sounds still give a better ratio because they contain significant changes in volume, as can be seen in Figure 5.10. The results using the four sounds together appear in Table 5.3, for both F=O and F=250. Again, there is a significant improvement with the proposed method. To summarize, this section describes a new method for feature extraction in the context of sound localization. Using the proposed procedure, extracted cues are more reliable, though a reject possibility is introduced. In typical environments changes in the volume of the sound signals are commonplace. Such changes are due to variations in the volume of the signal itself and changes in the distance to the sound source. In the proposed procedure CHAPTER 5. PERCEPTION Sound 1 Cuel Cue2 Cue3 Cue4 1 clavs 1 0.002 0.004 0.001 2e-4 1 maraca 0.119 0.015 0.014 0.239 whistling 0.251 0.007 0.030 0.247 Table 5.2: Results obtained for F=O, 1 Sound Cuel Cue2 Cue3 Cue4 1 Table 5.3: Results obtained considering the four sounds together. ton. + It gives an angle that can be used for an attentional mechanism (see Section 5.3) A serious problem for sound localization is the fact that servomotors make noise. Actually, servomotors make noise even when not moving, and noise is also generated by computer fans nearby. In the SIG Humanoid this is solved by using four microphones, two inside the head cover and two outside, and cancelling components of the interna1 signals in the externa1 signals [Nakadai et al., 20001. This ingenuous solution depends on the head cover attenuating the sounds. In earlier versions of CASIMIR0 the microphones were located in the head, right below the eyes, and the sound localization module only produced results when no servomotor was moving. This is obviously the simplest option. Currently, the two microphones are located ata distance of approximately 80 cm from the servomotors. There is a wooden box (the head is on this box) between the head and the microphones. In their present position the microphones do not capture the noise of the motors, and still capture voice and other sounds coming from the interaction space. CHAPTER 5. PERCEPTION 5.4 Face Detection Omnidirectional vision allows the robot to detect people in the scene, just to make the neck turn toward them. When the neck turns, there is no guarantee that omnidirectional vision has detected a person, it can be a coat stand, a wheelchair, etc. A face detection module that tries to detect people is described in this section. It uses colour images taken by the stereo camera described in Section 4.1. The module was originally developed for face detection and normalization integrated in a platform designed for general purpose artificial vision applications known as DESEO [Hernández Tejera et al., 19991. 5.4.1 The ENCARA Face Detector Face detection systems described in the literature can be classified by attending different criteria. One of them is based on the use of knowledge employed by these systems: implicit or explicit. The first group focuses on learning a classifier from a set of training samples, providing robust detection for restricted scales and orientations at low rates. These techniques perform with brute force, without attending to some evidences or stimuli that could launch the face processing modules, similar to the way some authors consider that the human system works [Young, 19981. On the other hand, the second group exploits the explicit knowledge of structural and appearance face characteristics that could be provided from human experience, offering fast processing for restricted scenarios. ENCARA [Castrillón, 20031, the face detection module integrated in CASIMIRO, merges both orientations in order to make use opportunistically of their advantages and conditioned by the need of getting a real-time system with standard general purpose hardware. ENCARA selects candidates using explicit knowledge for later applying a fast implicit knowledge based approach. Classification is the crucial process in face detection. There are multiple posible solutions that, roughly speaking, can be divided into two groups: Individual and Multiple classifiers. The complex nature of the face detection problem is easily addressed by means of an approach based on multiple classifiers. The architecture for combination of classifiers used in ENCARA follows [Viola and Jones, 20011 and is sketched in Figure 5.13. However, there is a main difference in relation to that work where the classifiers are based only on rectangle features [Viola and Jones, 20011, in this model the different nature of the classifiers used is assumed and promoted. Initially, evidence about the presence of a face in the image is obtained and the face CHAPTER 5. PERCEPTION with fluid filtering in a filtering cascade. In this case, each filtering stage rejects a fraction of impurity. The more stages with a rejection rate, the more pure fluid is obtained at the output. How to select the individual classifier modules? Different options are possible. In ENCARA, an opportunistic criterion is employed to extract cues and to use, in a convenient fashion, explicit and implicit knowledge to restrict the solutions to a solution space fraction which can comply with real-time restrictions and have a flexible framework to test different solutions, adding modules or deleting others, allowing each module in the cascade to be also a combined classifier. ENCARA is briefly described in terms of the following main modules, organized as a cascade of hypothesis confirmations/rejections: M0.- Tracking: If there is a recent detection, the next frame is analysed first searching for facial elements detected in the previous frame: eyes and mouth corners. If the tracked positions are similar to the one in the previous frame and the appearance test is passed, ENCARA considers that a face has been detected. MI.- Face Candidate Selection: The current implementation makes use of a skin colour approach to select rectangular areas in the image which could contain a face. Once the normalized red and green image has been calculated, a simple method based on defining a rectangular discrimination area on that colour space is employed for skin colour classification. Dilation is applied to the resulting blob image using a 3 x 3 structuring element. M2.- Facial Features Detection: Frontal faces would verify some restrictions for severa1 salient facial features. In the candidates areas selected by the MI module, the system removes heuristically elements that are not part of the face, i.e. neck and fits an ellipse to obtain the vertical position of the blob. Later, this module searches for a first frontal detection based on facial features and its restrictions: geometric interrelations and appearance. This approach would first search potential eyes in selected areas taking into consideration that for Caucasian faces, the eyes are dark areas on the face. After the first detection of an individual, the detection process will be adapted to the individual's dimensions and appearance as a consequence of temporal coherence enforcement. M3.- Normalization: In any case, the development of a general system capable of detecting faces at different scales must include a size normalization process in order to allow for a posterior face analysis and recognition reducing the problem dimensionality. CHAPTER 5. PERCEPTION M4.- Pattern Matching Confirmation: A final confirmation step of the resulting normalized image is necessary to reduce the number of false positives. This step is based on an implicit howledge technique. For eye appearance, a certain area (11 x 11) around both eyes is projected to a Principal Component Analysis (PCA) eigenspace and reconstructed. The reconstruction error provides a measure of its eye appearance [Hjelmas and Farup, 20011, and could be used to identify incorrect eye detections. If this test is passed, a final appearance test applied to the whole normalized image in order to reduce false positives makes use of a PCA representation that is classified using Support Vector Machines [Burges, 19981. If the tests are passed, the mouth and nose are located in relation to eye pair position and their dark appearance in a face. In any other case, when no frontal face is detected, the system computes if there was a recent face detection in which at least one facial feature was not lost according to tracking process, and the posible face location is estimated with high likelihood. 5.4.2 Performance and Implementation The main features of ENCARA are: + The resulting system integrates and coordinates different techniques, heuristics and common sense ideas adapted from the literature, or conceived during its development. + The system is based on a hypothesis verificationlrejection scheme applied opportunistically in cascade, making use of spatial and temporal coherence. + The system uses implicit and explicit howledge + The system was designed in a modular fashion to be updated, modified and improved according to ideas andlor techniques that could be integrated. ENCARA detects an average of 84% of the faces detected using the well-hown Rowley-Kanade's detector [Rowley et al., 19981, but 22 times faster using standard acquisition and processing hardware. ENCARA provides also the added value of detecting facial features for each detected face. More details of experiments carried out with ENCARA can be found in [Castrillón, 20031. The ENCARA system was fully integrated in earlier versions of CASIMIRO. Currently, only the first filter of skin blob detection and a simple blob ratio filter are being used from ENCARA. The rest of the filters were too restrictive for our environment, with a significant amount of frontal faces being discarded by the system. Note that this is in line with CHAPTER 5. PERCEPTION Figure 5.17: Face rectangles obtained without (left) and with (right) shirt elimination, + Symbolic selection: identifying one of severa1 alternatives by non-spatial means. Head nods and shakes fa11 into the fourth category. They are very simple in the sense that they only provide yeslno, understandingldisbelief, approval/disapproval meanings. However, their importance must not be underestimated: + The meaning of head nods and shakes is almost universal '. + They can be detected in a relatively simple and robust way. + They can be used as the minimum feedback for learning new capabilities. + YES and NO have been shown to be by far the two most common human inputs in conversations between humans and chatbots [Wallace, 20051. The head nodlshake detection system described in [Tang and Nakatsu, 20001 achieved a recognition rate of around 90%. It used a Kanade, Luca and Tomasi tracker to follow facial features. Then, a vector is formed with the evolution of those features. The vector feeds a (previously trained) neural network that produces a decision. The system for nodlshake detection described in [Kapoor and Picard, 20011 achieves a recognition accuracy of 78.46%, in real-time. However, the system uses complex hardware and software. An infrared sensitive camera synchronized with infrared LEDs is used to track pupils, and a HMM based pattern analyzer is used to the detect nods and shakes. The system had problems with people wearing glasses, and could have problems with earrings too. The same pupil-detection technique was used in [Davis and Vaks, 20011. That work emphasized the importance of the timing and periodicity of head nods and shakes. However, in our view that information is not robust enough to be used. Innatural human-human interaction, head 'In Bulgaria, a head nod signals a NO, while a head shake signals a YES. CHAPTER 5. PERCEPTION with an incredible number of terabytes in which to store facts and events indefinitely. Behavioural changes occur when we memorize but also when we forget. Thus, a forgetting mechanism can also be helpful in our effort, especially if we take into account the fact that actions chosen by the action-selection module do not always produce the same visible outcome (i.e. the Talk actions). Suppose that a behaviour is triggered by a certain state. As long as that state is present, the behaviour will execute the associated action over and over again, with a frequency imposed by the cycle of the action-selection implementation. With memory, the forgetting mechanism would have the control of the repetitions. Basically, actions should repeat only when: + The robot forgets that it has executed them + The robot do not forgets but a certain time has passed (in which it is reasonable to try again) The first controlled studies of forgetting mechanisms were carried out by Ebbinghaus [Ebbinghaus, 19131. Those experiments, replicated many times, concluded that the forgetting process is more accelerated (we tend to forget more information) in the first minutes and hours after memorization. This can be characterized by a power function (of the form y = at - b, where a and b are positive real numbers), as demonstrated by Wixted and colleagues [Wixted and Ebbesen, 1991, Wixted and Ebbesen, 1997, Kahana and Adler, 20021. In [Rubin and Wenzel, 19961 over a hundred forgetting functions were compared and it was found that the power function was one of only four that provided a good fit to a wide range of forgetting data. In CASIMIRO, forgetting is modelled in the following way. Let f (t) be a forget function, which we use as a measure of the probability of forgetting something: where k and 1 are constants. We apply the f function to the set of Boolean predicates that the robot retains in memory (both global and associated to individuals). There is evidence that some facts are forgotten earlier than others. Some facts are never forgotten. Interference effects are thought to be one of the factors that account for those differences. For simplicity, we do not model those aspects and consider them represented in the stochastic nature of the CHAPTER 5. PERCEPTION However, we are more interested in the role of the person who cares the robot and uses it more frequently (this person is generally the owner too). This person is the one who switches the robot on, which is usually done from a certain part of the robot or from a certain computer. That cue may be exploited to recognize the robot owner or caregiver. Amazing Amanda [Playmates Toys Inc., 20051, a do11 released in autumn of 2005, is able to recognize the girl that takes the mother role. Once the do11 is activated, it starts asking questions. That way, the do11 can "learn" the girl's voice patterns. From that moment on, the do11 is able to recognize the utterances of its "mommy". Other voices can lead Amanda to say "You don't sound like Mommy". Such technique may seem rather ad hoc. However, the approach finds striking examples in nature. Lorenz, one of the founders of ethology, found that, upon coming out of their eggs, geese follow and become attached to the first moving object that they encounter. He showed this by rearing the geese from hatching. From that moment on the geese would follow him. Such phenomenon, which also appears in mammals, is known as imprinting [Lorenz, 19811. In the case of CASIMIRO, the main computer (from which the robot is switched on) is situated behind the robot, on the same table, see Figure 5.27. A camera was placed on top of that computer. The owner detection module uses that camera to search for a skin coloured blob in the image. When the robot is switched on this module will detect a skin coloured blob. The camera has a wide-angle lens, and a relatively low resolution of 160x120 is used. When no blob is encountered in the image the module notifies the Attention module of that event. At that moment the owner detection module exits in order to free CPU resources. Once it has been notified by the owner detection module, the Attention module considers the owner as the first blob that "enters" the omnidirectional camera image from the left. The "owner" property is stored along with the individual in the tracking process. This simple procedure is a form of imprinting. In a sense, the robot finds its ownercaregiver in the first human it sees. It does not stores any biometric features to recognize the owner after being switched on, only its position. Face recognition researchers tend to measure performance in terms of the number of individuals that the system can recognize and measured error rate. A measured error rate of 5-10% can be considered very good under restricted conditions. The approach presented here recognizes a single individual with guaranteed zero error. No face recognition method would recognize the owner with such low error. Note that this is the result of following the approach introduced in Chapter 3: we have been able to devise the simplest algorithm (or CHAPTER 5. PERCEPTION Figure 5.27: The computer from where CASIMIR0 is started. The interaction space is on the left. one of the simplest) that allows to recognize the owner. For that purpose we have fitted the solution to the robot niche. Habituation Habituation is a filtering mechanism that has received a lot of attention in physiology and psychology. In particular, some researchers have investigated the mechanisms of habituation in animals, being one of the most known works the study of the Aplysia's gill-withdrawal reflex [Castellucci et al., 19701. When the animal's siphon is touched, its gil1 contracts for a few seconds. If the siphon is stimulated repeatedly, the gill-withdrawal effect tends to disappear. Crook and Hayes [Crook and Hayes, 20011 comment on a study carried out on two monkeys by Xiang and Brown who identified neurons that exhibit a habituation mechanism since their activity decreases as the stimulus is shown repeatedly. Stanley's model [Stanley, 19761 of habituation, proposed to simulate habituation data obtained from the cat spinal cord, has been widely used in the literature. This model describes the decrease efficacy y of a synapsis by the first-order differential equation: CHAPTER 5. PERCEPTION where yo is the normal, initial value of y, S(t) represents the external stimulation, T is a time constant that governs the rate of habituation and a regulates the rate of recovery. Equation (5.20) ensures that the synaptic efficacy decreases when the input signal S(t) increases and returns to its maximum yo in the absence of an input signal. The model given by (5.20) can only explain short-term habituation, so Wang introduced a model to incorporate both short-term and long-term habituation using an inverse S-shaped curve [Wang, 19951, where a, yo an have the same meaning than in (5.20), 0 regulates the habituation and z(t) decreases monotonically with each activation of the external stimulation S(t), and models the long-term habituation. Due to this effect of z(t) after a large number of activations, the recovery rate is slower. Note that novelty detection is a concept related to habituation. Novelty detection is the discovery of stimuli not perceived before and so habituation serves as a novelty filter [Stiles and Ghosh, 19951. From a engineering viewpoint, perceptual user interfaces, like human-like robots, should be endowed with a habituation mechanism. The interest is twofold. First, it would be a filtering mechanism, discarding (or minimizing the importance of) repetitive information while paying attention to new experiences. This is in part motivated by the desire to distinguish between artificial and human signals. Artificial signals are often static or repeat with a fixed frequency. We do not want our robot to pay much attention to the hands of a wall-mounted clock. Instead, it would be more interesting to detect non-repetitive stimuli, such as a conversation or a sudden loud noise. Note that we generally consider monotonous signals as those having a fixed frequency or frequencies (which can be zero, that is, the signal does not change) but signals whose frequency changes in a periodic pattern could also be considered monotonous. Higher scales are also posible but we do not consider them in this work because they are very hard to visualize and real examples of them are not so common. Second, habituation would lead to a more human-like behaviour, as perceived by users of the interface. As an example of this, consider Kismet. Someone can catch the eye of the system while waving a hand in its visual field of view, but if the stimulus is repetitive for a long time the system can show a lack of interest in it. Many aspects of Kismet's mental CHAPTER 5. PERCEPTION architecture are directly or indirectly influenced by the detection of monotonous sensory signals: stimulation and fatigue drives and the arousal dimension of its affect space (and in turn some emotional states, like surprise, boredom or interest). Although we focus our work on the abilities described above, many other applications are also imaginable. In the robotics field, habituation mechanisms have been used to reduce oscillations caused by collision-avoidance behaviours when navigating through anarrow corridor [Chang, 20001. Marsland [Marsland et al., 20001 uses a SOM neural network as a memory for novelty detection. To add short-term habituation to the original network, each neuron of the SOM is connected to an output neuron with habituable synapses based on the model (5.20). Habituation is also used in [Stoytchev and Arkin, 20031 for controlling reactivity strength, visual attention [Peters and Sowmya, 1998, Breazeal and Scassellati, 19991, and general learning [Damper et al., 19991. On the other hand, there is considerable interest in the field of musicology in Beat Tracking Systems (BTS) [Goto and Muraoka, 19971. BTS systems aim to find the tempo of an audio signal, which is basically the rate of repetitions. The main applications of BTS systems are audiolvideo editing, synchronization of computer graphics with music, stage lighting control and audio content searching. If we use the model of Equation (5.20) we can obtain undesired effects with certain stimuli. A periodic input signal (with frequency greater than zero) can produce a response that does not exhibit habituation. This is due to the fact that the model does not account for changing stimuli, but for continuous ones. In order to include this fact in the model, we propose to use an auxiliary signal which will be zero when the stimulus is stationary or with a fixed frequency, and one otherwise, and use this signal as an input to the habituation model (5.20). The auxiliary signal, which basically detects monotonous stimuli, is obtained from the spectrogram of the stimulus itself. The spectrogram is a time-frequency distribution of a signal, and it is based on the Fourier Transform with a sliding window [Holland et al., 20001. The equation gives the definition of a spectrogram with a Gaussian window function of half-width T, and it is the power spectrum of a signal which corresponds to the squared magnitude of the Fourier transform of the windowed signal. The window can have other forms apart from the Gaussian one. In Figure 5.28 we show an audio signal and its corresponding spectrogram, where brighter areas correspond to higher power. Two well defined frequency spectra can CHAPTER 5. PERCEPTION be distinguished, for there is a change in the input signal at time 0.5 s. Temporal patterns of the stimulus signal have a specific pattern in the spectrogram. A fixed frequency signal corresponds to a straight line parallel to the time axis in the spectrogram, and the length of this line indicates how long has been the stimulus present. -0811111111111 O O1 O2 O3 O4 O5 O6 O7 O8 O9 1 Time (sec) o O O1 O2 O3 O4 O5 O6 O7 O8 O9 Time Figure 5.28: Audio signal (left) and its corresponding spectrogram (right). Spectrograms are computed from windows of the input signal. These windows, of length 1, overlap by 1 - 1 samples. Let each spectrogram be represented as a matrix M, in which rows represent frequencies and columns represent time. We calculate the variance of each row of M, which produces a column vector v. The norm of this vector v is a measure of how monotonous the input signal is. The norm will be high when the signal is changing, and low otherwise. Thus, the auxiliary signal needed is simply the thresholded norm of v. The amplitude of the input signal affects the power content of the spectrograms, and in turn the norm of v. Thus, prior to calculating the FFT the input signal must be normalized dividing each input window by the sum of its absolute values. A value of 1 for the auxiliary signal will mean that there are changes in the input signal, while a value of O indicates that the input signal is monotonous. Once the auxiliary signal is available, the model (5.20) is used to get the desired habituation behaviour, as controlled by parameters r and a. Formally, let N and 1 be the number of rows and columns of M, respectively, and let mi,j represent the element in row i and column j of M. Vector v is calculated as: where: CHAPTER 5. PERCEPTION The auxiliary signal is then, for a given threshold T: With this method both static an fixed frequency stimuli can be detected. However, there are stimuli that change their frequency according to a periodic pattern. These stimuli should also be considered as monotonous. The hissing sound of a siren, for example, is a signal whose frequency changes in a repeated pattern. After few repetitions the signal will be considered monotonous. One way to detect these kind of stimuli is to use the same method with the auxiliary signal. If the input signal changes its frequency content in a repeated pattern, the auxiliary signal will be periodic with a fixed frequency, and that can be detected as explained in the previous paragraph. Thus, two thresholds will be needed, one for the "first level" and one for the "second level". Higher levels could conceivably be used, but we have not considered them because they are very difficult to visualize and encounter in the physical world. Note that the second-leve1 auxiliary signal will be 1 when there are changes in the first-leve1 auxiliary signal, and thus when there are changes in the input signal, and O otherwise. Thus, the final input to the habituation model (5.20) will be the secondlevel auxiliary signal. Note that this second level introduces additional computation, and in some cases we could consider it unnecessary, if we decide to detect only simple monotonous signals. There is only one detail left. If the first-leve1 auxiliary signal is 1 (meaning that the input signal is changing), and this remains for a while, the second-leve1 auxiliary signal will be O (because the second-leve1 norm of the variance vector will be O) which is not the correct value. In order to correct this, the second level must detect when the norm is O and, if so, use the value of the first-leve1 auxiliary signal, instead of the second-leve1 auxiliary signal. Note that if the first-leve1 auxiliary signal is periodic the second-leve1 variances obtained should theoretically be O, which would prevent the use of this correction. However, in al1 the experiments carried out this never happened, because there is always an unavoidable amount of fluctuations in the input signal, which makes the variances larger than O. A previous version of the method proposed here has been already published elsewhere [Lorenzo and Hernández, 2002b, Lorenzo and Hernández, 2002~1. That version used only the frequency associated to the maximum power. Habituation should be present when the plot of that frequency versus time is a straight line. Changes are detected by fitting a line to the last k values of the frequency and computing the difference between the current value and the predicted value with the fitted line. That approach, however is too simplistic CHAPTER 5. PERCEPTION in the sense that it assumes that the input signal is entirely represented by the frequency of maximum power. The algorithm described above was implemented to test it with different input signals. The first experiments that we present use only the first level mentioned above. In order to gather signals from the visual domain, we recorded video containing a yellow bright stimulus (a yellow card) that was moved in a repetitive fashion, see Figure 5.29-a). Using simple segmentation techniques we extracted the centroid of the card on each frame (384x288) and summed the x and .y pixel coordinates to form the one-dimensional signal of Figure 5.29b). The sequence of card movements throughout the recording was: horizontal movement, random (aperiodic) movement, vertical movement and vertical movement at a different frequency than the previous one. l 1 O 20 30 40 50 60 Time (sec) Figure 5.29: a) Video recording used for the visual habituation experiment, b) onedimensional signal extracted from it. The results appear in Figure 5.30. Windows of 128 samples were used, and the variance threshold was set at 1000. As for the audio domain, we recorded signals with a standard PC microphone, at a 22050 Hz sample rate, 8 bits. Figure 5.31 shows the results obtained for an audio signal that contains three sequential parts: silence (0-O.%), people speaking (0.5-1s) and a tone played from an electric piano (1-1.4s). Note that there is an initial delay due to the need to fill the input window, here of length 1 = 5120. The habituation level, obtained using the model of (5.20), shows a satisfactory response. Figure 5.32 shows the results obtained for an audio signal that contains another three sequential parts: a tone played from an electric piano (0-O.%), silence (0.5-1s) and another tone (1-1.4s). The same window length 1 = 5120 was used, and again the habituation level CHAPTER 5. PERCEPTION l 20 30 40 50 60 Time (sec) Time (sec) b) Figure 5.30: a) Evolution of the (12) norm of the variance vector v, b) habituation level, using r = 5,a = 1. shows a satisfactory behaviour. In order to test both the first and second levels of the method, we built an audio signal containing three sequential parts: a beep repetitive sound from a mobile phone, people speaking and a tone played from an electric piano. This signal was accelerated to reduce computation time, which does not alter the qualitative results of the experiments. Results are shown in Figure 5.33. The window length was 1 = 5120 for the first level and 1 = 2148 for the second. In this case the repetitive beeps (clearly observed as a repetitive pattern in the first part of the spectrogram) are correctly considered as monotonous. This would not have occurred if we had used the first-leve1 auxiliary signal alone, for numerous changes are detected (see Figure 5.33-d). Next, we discuss a few aspects of practica1 interest. Particularly, we will comment on the effect of the values of the different parameters to use: + Length of the input window, 1: It should be the largest possible, in order to detect stimuli with large period. However it cannot be too large because that would introduce an unacceptable delay in the response to stimuli with smaller period. Thus, it depends on the type of stimuli. A flexible solution would be to implement multiple instances of the problem, each one with a different size for this parameter, in a multiscale fashion. + Tau, r: It controls the rate of habituation. + Alpha, a: It controls the rate or recovery. + Number of discrete frequency levels, N: Dependent on the type of input stimulus, it should normally be the largest possible. For the case of auditive signals, the minimum CHAPTER 6. ACTION Expression: "Surprise" Group: Mouth Group: Right eyebrow Group: Lefi eyebrow Group: Right ear Group: Lefi ear Group: Right eyelid Group: Lefi evelid Pose: Open Pose: Raised Pose: Raised Pose: Raised Pose: Raised Pose: Raised Pose: Raised Degree: 90 Degree: 90 Degree: 90 Degree: 100 Degree: 100 Degree: 80 Degree: 80 Table 6.1: Typical definition of an expression. The modeller is given the opportunity to modify these transitions, as some of them could appear unnatural. A number of intermediate points can be put in al1 along the transition trajectory. Additionally, velocity can be set between any two consecutive points in the trajectory. The possibility of using non-linear interpolation (splines) was considered, although eventually it was not necessary to obtain an acceptable behaviour. The first pose that the modeller must define is the neutral pose. Al1 the defined poses refer to a maximum degree for that pose, 100. Each pose can appear in a certain degree between O and 100. The degree is specified when the system is running, along with the pose itself. It is used to linearly interpolate the points in the trajectory with respect to the neutral pose. In another level, facial expressions refer to poses of the different groups, each with a certain degree. "Surprise", for example, could be represented by Table 6.1. The facial expression is specified while the system is running along with a degree that allows, by multiplication, to obtain the degree to apply to the poses of the different groups. For more control, the modeller can also specify a time of start for each group. This way, "surprise" could be achieved by raising first the eyebrows and then opening the mouth. Now it is easy to see that the level hierarchy allows to move individual features, like for example winking, blinking or opening the mouth for talking. With respect to this particular case, the motion of the mouth can be combined with poses of other facial features, producing combinations like for example "talking with expression of surprise". 6.1.2 Transitions Between Expressions Earlier it was mentioned that the degree of a pose can be specified at run time, while the defined poses referred to the maximum degree. How to obtain the trajectory from a pose A with a degree Gj to a pose B with a degree Gf? In other words, at a given time a group CHAPTER 6. ACTION is at pose A with degree Gi and we want the system to adopt pose B with degree Gf. For a two-motor group, the procedure is depicted in Figure 6.1, where N represents the neutral pose, X the point corresponding to the initial degree and Y the point corresponding to the final degree. To obtain the expression that gives the transition trajectory the following limit conditions are imposed: The trajectory that fits these restrictions is: Symbols ? and : correspond to the IF-THEN-ELSE instruction (as in the C language). The trajectory equation is continuous in the values of Gi and Gf. The same relationship is used to obtain velocities along the trajectory. m2 Figure 6.1 Transitions between expressions in motor space. CHAPTER 6. ACTION 6.1.3 Implementation The modelling framework described was implemented in the form of a pose editor. The editor gives the modeller a simple and interactive way to define and test poses, transitions, etc. This editor uses a low-leve1 library for control of the servomotors. The specifications of poses and transitions can be saved to afile. Later. this file will be the only thing necessary to reproduce (and generate) the different movements and trajectories. The pose editor can thus work both at design and run time, and it can also be controlled by other modules. The modeller has to follow a few simple steps: connect the ASC 16 board (see Section 4.1), establish the physical motion limits and speciSl the poses and transitions. The physical limits represent the minimum and maximum values that can be given to a motor, and they are also saved in the file. The central part of the editor allows to define groups of motors, poses and transitions, see the following figure. On the top left side there are scrollbars to move motors. On the central left side motor groups are defined. On the central right side poses and transitions can be defined for a selected group. The big STOP button stops al1 motors immediately. Seivo 1 A Seivo 2 A Seivo 3 '_I Saivo 4 4 1 Seivo 5 4 Seivo 6 i] S21~0 7 Seivo 8 Añadir 1 Quitar 1 Poses de segundo nivel - Nombra. -1 1841 Seivo 13 '_U I] 0 II d 2024 Seivo 14 do Siempre visible b d 2414 Seivo9 m 110 d 2146 Seivo10 u do @ 1800 S,,,, 11 ,JJ do Parar I] 1031 Seivo 12 '_U I( 0 motores Poses de prmcr nvtl y sus transiciones 4 Guardar Posición en ASE16 , ~n~c~al ~ntermedm f~nal Velocidad incial Velocidad entre el oto 1 1 I nterrned~o y el slgulente Añadir ] Añadir 1 1'3 Quitar 1 Quitar 1 Qu~tai 1 Giado inicial E1 m << 1 < 1 '>' 1 > 1 >> 1 Giadof~na K] m Grupos Congiado[%] m m Pose de primer nivel lorela 1 Añadir 1 Tiempo de comienzo [ms] Ir a la Pose Qu~tar 1 Figure 6.2: Main window of the pose editor. 11 1 CHAPTER 6. ACTION Two parameters control motion continuity. Each part of the trajectory can have a different velocity (recall that intermediate points can be defined) and, moreover, al1 the velocities are adjusted so that the motors stop at the same instant. The motor control board sends a signal when the motor has reached the final destination (or an intermediate point). However, there is a slight delay between the sending of this signal, its processing in the PC, and the next motion command, which causes the motion to be discontinuous. There is a "softness" parameter that makes the control board send the final of motion signal ahead of time. There is also a distance threshold to eliminate from trajectories too close consecutive points. The pose editor uses a buffer of requests that momentarily stores "Go to pose X with degree Y" petitions coming from other modules. The buffer is necessary, as requests take a certain time to complete. The buffer check al1 the time the pending requests. Those compatible with the request currently being served (i.e. they do not have any motor in common) are also served, always in the order in which the requests were made. Blinking has great importance for the robot. Blinking human forms have been appraised to have significantly higher accuracy and to be more intelligent than other humanlike forms [King and Ohya, 19961. In our implementation, blinking is treated as a special case. In principle, with the framework described it would have to be implemented using to consecutive commands sent to the pose editor, one for each eyelid. However, the eyelids must return to the original position, no matter what, and the velocity should always be set at maximum. Therefore, the system automatically converts two motion commands into four, and maximum velocity is imposed. In the case of winking, one motion command produces two commands internally. Also, to obtain maximum velocity and continuity in the blinking motion, the robot only blink when no other motor on the face is working. In conclusion, the pose editor is the module that allows a modeller to define how the facial features will move to adopt the different expressions. The pose repertoire can be upgraded (if motors are added) easily. More flexibility could have been given to the program. That was discarded because the main objective is to get the face to move well, as observed by humans, and because in practice the number of motors is relatively reduced. Figure 6.3 shows the facial expressions modelled in CASIMIR0 using the pose editor. 6.2 Neck As commented in Section 4.1 there is a 2-DOF neckunder the robot's face. The neck module of CASIMIRO's software is very simple. A velocity PID controller is used for small move- CHAPTER 6. ACTION Figure 6.4: Position of the camera and the pan axis of the neck. , and Once an angle a has been measured by the omnidirectional vision module, there must be a correction of that angle to get a value for the neck angle 0, so that the robot can look toward the detected person. The exact value for 0 depends on the distance c to the person, which is unknown. As an approximation, we may use the typical interaction distance value of 100cm in Equation 6.2 and obtain an expression for 0 that does not depend on c. Besides, we approximated the trigonometric function by a (much faster) quadratic. The quadratic fitting to the trigonometric function was obtained using the least-squares method. These approximations are good enough, as can be seen from Table 6.2. 6.3 Voice Generation Innatural language generation four basic schemes are used: canned te&, templates, cascaded items and features [School of Computing, University of Leeds, 2003, FLUIDS Project, 20031 In the canned text approach, specific fixed phrases are used, without modifications whatsoever. This is very easy to implement, although it is rather inflexible. To avoid repeatability a large number of phrases would have to be introduced beforehand. Templates are phrases that CHAPTER 6. ACTION Table 6.2: Theoretical angle errors using the implemented approximation. include empty slots that are filled with data, like a mail merge. Cascaded items and features are even more complex models. CASIMIR0 uses canned text for language generation, not least because of its simplicity. A text file contains a list of labels. Under each label, a list of phrases appear. Those are the phrases that will be pronounced by the robot. They can include annotations for the text-to-speech module (a commercially available TTS system was used. Annotations allow to change parameters like word intonation, speed, volume, etc.). Labels are what the robot wants to say, for example "greet", "something humorous", "something sad", etc. Examples of phrases for the label "greet" could be: "hi!': "good morning!': "greetings earthling". The Talk module, which manages the TTS system, reads the text file when it starts (Appendix A shows the phrase file. The text is in Spanish). It keeps a register of the phrases that haven been pronounced for each label, so that they will not be repeated. Given a label, it selects a phrase not pronounced before, randomly. If al1 the phrases for that label have been pronounced, there is the option of not saying anything or start again. The text file include annotations that had to be added to each particular phrase to enhance the naturalness of the utterance. In [Koku et al., 20001 the humanoid ISAC is used to investigate an interesting interaction technique. When the robot detects a person, it tries to initiate a conversation with himlher by using a simple phrase based on daily facts. Typical phrases are "Did you know that Tutu calls for abolition of death penalty" or 'Sisten to this, Fed in focus on Wall Street". This system was based on the idea that when people share the same environment for a while, they do not start a conversation by greeting each other, but instead they generally initiate CHAPTER 6. ACTION the conversation based on some daily facts that the other side might be interested in. In CASIMIRO, a somewhat similar technique has been used. It was considered that it would be very useful that the robot could say things that are very popular at present, like for example something about Prince Felipe's fiancée (which was one of the main pieces of gossip in Spain at the time of writing). That would be interpreted positively by the observer. As an example, consider the phrase 'Prince Felipe 'sjancée is dejnitely beyond my expectations! ". The phrase is fixed, though it will only pronounced if the thing, event or person is popular at present. Otherwise, it would not make sense to pronounce the phrase. These cases are implemented by preprocessing the text file mentioned above. The file contains those X's inserted in the phrases. The substitutions are made when the system starts, producing afile ready to be used by the Talk module. Another example: if "Beckham" is popular at present, then the phrase "í'm more famous than Beckham!" can be pronounced. As another example, in our region of Spain, it is humorous and relatively common to greet somebody with the name of a person who is extremely popular at the time. How to determine if something or someone is popular? The solution implemented looks in electronic newspapers of the last K days for mentions of a term (or phrase). If the term has a frequency of apparition that exceeds a threshold then it is considered popular. In the text file, "Popular" is a special label that gives the list of terms to look for. In the rest of the file those terms appear as labels, with their phrases to pronounce. If many terms are popular they al1 will be used, in a random order. Once a popular term is used it will not be used again. If the threshold is carefully chosen and the phrases are not too specific then it would seem to be in context. For each newspaper, terms are counted only once (i.e. whether they appear in it or not). Note that the robot should know when it is appropriate to speak. Turn-taking is a basic form of organization for any conversation. Before speaking, the Talk module interrogates the sound localization module (Section 5.2) in order to know if any sound has been detected in the last K milliseconds (currently K=2000). If so, the robot assumes that someone is speaking and postpones its action. 6.3.1 Expressive Talk The Talk module pronounces phrases with an intonation that depends on the current facial expression. Before pronouncing a phrase, the Talk module interrogates the pose editor to obtain the current facial expression of the robot. The mapping from expression to voice parameters is based on Table 6.3. CHAPTER 6. ACTION Table 6.3: Effect of emotions on human speech [Breazeal, 20021 1 1 1 1 1 Happiness 50 70 70 60 O Sleep 1 35 NDb 30 45 50 Facial expression Surprise Sadness Aneer Table 6.4: Values for NM used. s=speed, b=pitch baseline, f=pitch fluctuation, v=volume, h=breathiness. With the TTS system used four voice parameters can be controlled: speech rate, pitch average, pitch range and intensity. Besides, the rising contour for the surprise expression can be achieved. Are these four parameters sufficient to convey the emotion represented in the facial expression? In [Schroder, 20011 it is stated that some emotions can be recognized reasonably well through average pitch and speed. s 50 35 50 Each of the four voice parameters have a value N associated that indicates how strong is the effect or characteristic. As the facial expression has a degree (between O and 100, O=neutral), the following value for N was used: where g is the degree of the facial expression, ND is the default value for N (supposedly corresponding to a neutral intonation) and NM is the maximum (or minimum) value that N can have. Values for NM for each expression and parameter were established manually, starting from the values of Table 6.3 and the indications in [Montero et al., 19981, see Figure 6.4. An additional parameter set was included for the expression "Sleep". In this case, the intonation is monotonous and slow, and breathiness is high. b 70 50 75 f NDf 40 70 60 45 60 vh O O O CHAPTER 6. ACTION 6.3.2 Local Accent CASIMIR0 has been built in the Canary Islands, a region of Spain where people has a softer accent that in the mainland. For local people, that accent sounds more familiar and promotes empathy. The most significant variation in the Canarian accent is that the "c" (as in "hacer") is pronounced as a "S" (see [Hernández, 19991 for more details). Mexican Spanish, also available in our TTS system, is very similar, although phrase intonation profiles are different. The voice generated by the TTS system has a standard Spanish accent. Simple substitutions were made in order to get a local accent, see Table 6.5. Not al1 the cases are covered, though the effect was easily noticed. Mad0"i "áo" Medo"* "éo" "idO"* M "io" "úa" Table 6.5: Substitutions for getting the local accent (i=only when it appears at the end of a word). This section has described the simple -yet functionallanguage generation system implemented in CASIMIRO. Sophisticated language generation entails a number of requirements. First, the robot has to have something meaningful to say. It could certainly say many things, though they should make sense. Second, there should be variety in the vocabulary and phrases. Othenvise the robot will appear repetitive and dumb. Third, we believe that this effort should be accompanied by speech recognition capabilities, so that conversation can be rich and lasting. Some techniques have been proposed for compensating speech recognition deficiencies in dialog systems, such as reducing the vocabulary. In principle, it would be desirable that voice synthesis produced a more human voice. CHAPTER 7. BEHAVIOUR step environment inputs can make the system select any behaviour. Therefore, both cases of planning and reactivity are dealt with. The weight given to one or other capacity depends on the ratio of 4 to y. 7.2.2 ZagaZ CASIMIRO'S action selection module is based on Zaga [Hernández-Cerpa, 200 11. Zaga is an implementation of PHISH-Nets [Rhodes, 19961, an enhanced version of Maes' Behaviour Networks. It has a graphical interface that allows to execute and debug specifications of PHISH-Nets. Specifications have to be compiled before they can be executed. There are two compilation modes: Release and Debug. Figure 7.1 shows the main windows of the application. Name - Argurnent Type -1 Action Irnplernenter I Nearness Esuatim Irnplernenter -1 Preconditions Ordered? Figure 7.1 : Main windows of the Zaga application. (Courtesy of D. Hernández) Zaga capabilities are not fully exploited in CASIMIRO, although it is thought they will be used in the future as more complex behaviours are added. Currently, the system works as a priority-guided rule system. The behaviour that is finally selected for execution is chosen among those executable according to preset priorities (see below). Also, the Add and Delete lists are not used, for they are actually implemented through the memory system. This way a better control can be exercised over the memorized and forgotten items. Only CHAPTER 7. BEHAVIOUR one goal was used, and it is present in the Add list of each behaviour. The implementation guarantees that the goal is never achieved. 7.2.3 Implemented Behaviours CASIMIRO's behaviour was made intentionally simple. It engages in interaction with the subjects that enter its interaction space (roughly a 180" zone in front of it). It tries to te11 the subjects poems, jokes and data about itself, in that order. Basic interaction stability is maintained: each person is greeted, and at times the robot also makes "continuity", "funny" comments. It can also detect when the person is distracted, which makes it change topic (between poems, jokes and data about itself). It can also ask the subject whether helshe wants to hear more phrases of the current topic. When the subject does not collaborate (i.e. when the robot can not accomplish its task as stated above) the robot gets angry. Table 7.2 shows a list of the perceptions involved in the specification of CASIMIRO's behaviours. Al1 of the perceptions are Boolean predicates. Some predicates are just perceptions of the world at a given moment. Others take their values from facts stored in memory (either in the individual's or in global memory), see Section 5.6. That is, they are true or false depending on the predicate being present in memory or not. For example, the perception Greeted is true only when the Greeted fact is present in memory (the Greet action inserts it). Memory is initially empty, the execution of behaviours adds predicates to it. From time to time a predicate is forgotten (i.e. removed from memory). The forgetting parameters associated to these predicates are shown in Table 7.3. These forgetting parameters are very important for the correct sequence of behaviours. Since some predicates take their values from facts stores in memory, perceptions like "LikesMe" and "NotLikesMe" are not quite the opposite of each other. "LikesMe" is true when the robot has asked a "Do you like me" question and the answer was Yes, and false otherwise. "NotLikesMe" is true when the robot has asked a "Do you like me" question and the answer was No, and false otherwise. Table 7.4 shows a list of the available actions. Note that, as commented in Section 3.5, actions are high-level. The TalkAbout[X] actions can turn into different phrases to say (see Section 6.3). Other actions like Greet can also produce different final actions. Instead of using the energy value to select a behaviour, a priority scheme was used. Each behaviour action has an assigned priority (Table 7.5) that allows to pick only one behaviour when more than one is executable. Table 7.6 shows a summary of the behaviours defined in ZagaZ. CHAPTER 7. BEHAVIOUR Predicate 1 Descripfion 1 Fmm momory? 1 FrontalFace 1 truewhen arouehlvfrontal face of m individual has been detected 1 1 1 .l1111I. ~.llll. 11, 1. . 111111.1.111.11. 11.1 -1111 . 111.11 1 l. I1.L I l.. 1% 1, Il ll ll., ll..ll\l.lllll I.I. .l.L 1. II., II8.I \.,l.. '1 ... 1, Il ll ll., ll..ll\l.lll,I ,.Ir l.. .l. .L I II., I I . I \l \,l.? 1, Il ll ll., ll..ll\l.lllll 1.11..\11.1 I1.111.I \.~\I.\,I.? 1, Il ll ll., ll..ll\l.lllll ,.Ir 11..\11.1 11. 1,. I I111111II 1, Il ll ll., ll..ll\l.lllll l.,. IILCI. .II.LI\L.~ 111.111111111. 11111.1 li) ,l., I ., 1 iI.c.\l. 1, Il ll ll., lI8.I ..l..lilcl.II.~l ll., ll..ll\l IIIII 111.c. 11 ).. 1,. Ii\i.lii,l \.r ,I.c.\lc 1, Il ll ll., lI8.I ..l..lilcl.II.~l ll., ll..ll\l IIIII ).. 1,. Ii\i.lii,l 1 does not like it l Owner l triie if the individiial is the rohot's mner TooMmyPeople 1 true if more thm 2 individuals are being detected xound 1 1 FrontalSound WavingHmd WavingHmdl (the sensor is described in Section 4.1) 1 true if local temperature exceeds a threshold the robot true if the robot has heard a sound coming from in front true if the Omnidirectional vision module has detected awaving hmd (see Section 5.1) true if the face detection module has detected awaving EmotionHighArousal EmotionLmArousal EmotionHighValence EmotionLmValence hmd true if average arousal experiencedwith the individual exceeds a threshold (see Section 7.3) true if average arousal experiencedwith the individual is below a threshold (see Section 7.3) true if average valence experienced with the individual exceeds a threshold (see Section 7.3) true if averaee valence exoerienced with the individual Table 7.2: List of available high-leve1 perceptions WmtRoboWata NotWmtRoboWata MoreRoboWataAvailable NotMoreRoboWataAvailable NotTalkedAboutw] Queslionedw] NotQuestionedw] IAmAlone , true if the individual wmts to hear data about the robot true if the individual does not wmt to hex (more) data about the robot true if there are more robot datafor the robot to say true if there are not more robot datafor the robot to say true if the robot has not talked about [X] true if the robot has asked the individual about w] true if the robot has not asked the individual about w] true if there is no me is the interaction xea ,, , yes, individual yes, individual yes, individual yes, individual yes, individual CHAPTER 7. BEHAVIOUR 1 OuestionedDoYouWantRobotData 1 O 1 O Table 7.3: k and 1 forgetting parameters for predicates stored in memory. CHAPTER 7. BEHAVIOUR Actiori Greet TalkAboutILikeYou TalkAboutHotDay TalkAboutIDontLikeYou TalkAboutSthingToAttractPeople U IncreaseValence DecreaseValence LookAtOther IncreaseArousal DecreaseArousal TalkAboutPoem TalkAboutRobotData TalkAboutOwner TalkAboutIAm Alone QuestionDoYouLikeMe QuestionDoYouWantPoem QuestionDoYouWantJoke QuestionDoYouWantRobotData QuestionDoYouStillWantPoem QuestionDoYouStillWantJoke OuestionDoYouStillWantRobotData Descnption The robot greets the individual The robot savs the individual that it likes himher The robot says sthing when the room - is too cold The robot savs sthim. when the dav is hot U The robot says the individual that it does not like himher The robot says sthing. when - there are too many people around The robot says sthing to attract people The robot says sthing funny ("continuity" comment) The robot reurimands the individual for being too close The robot increases its valence level other individual The robot increases its arousal level The robot decreases its arousal level The robot tells a poem The robot tells a joke The robot te11 sth'ig. about itself The robot tells sthng. to its owner The robot complains of its loneliness The robot asks if the user likes it. This action activates either the LkesMe or NotLikesMe perceptions The robot asks the subiect if helshe wants to hear poems. Ths action activates either the WantPoem or NotWantPoem perceptions The robot asks the subject if helshe wants to hear jokes. This action about itself. This action activates either the WantRobotData or NotWantRobotData uerceutions Table 7.4: List of available high-leve1 actions. CHAPTER 7. BEHAVIOUR Action Greet WantPoem WantJoke WantRobotData TalkAboutIDontLikeYou TalkAboutILikeYou LookAtOther TalkAboutOwner QuestionDoYouStillWantPoem QuestionDoYouStillWantJoke QuestionDoYouStillWantRobotData QuestionDoYouWantJoke QuestionDoYouWantRobotData TalkAboutWavingHand TalkAboutPoem TalkAboutJoke TalkAboutRobotData TalkAboutSpeaking TalkAboutMoving TalkAboutTooClose TTalkAboutIAmAlone Priority 1 1 1 1 1 1 2 3 3 3 3 4 4 5 5 5 5 6 7 7 8 Table 7.5: Priorities assigned to actions. Actions that do not appear in the table have al1 an equal priority value of O. CHAPTER 7. BEHAVIOUR 7.3 Emotions Many emotional models have been proposed both within the A1 community and in psychology (see the Emotion Home Page [E. Hudlicka and J.M. Fellous, 20041). The most wellknown model is perhaps that of Russell [Russell, 19801, which considers that emotions fa11 in a bidimensional space, with orthogonal valence and arousal components, see Figure 7.2. ANGRY SAD JOYFUL - Valence RELAXED Figure 7.2: Arousal and valence emotional space. This bidimensional space (also called circumplex structure) has received wide support in the related literature [Carney and Colvin, 20051. Many forms of human emotional experience Cjudgement of the similarity between pairs of affect terms, self-reports of current emotion and from perceptions of similarity between static photographs of expressed emotion) point to an ordering of basic emotions around the perimeter of a circle with arousal and valence axes. The central zone of that space would correspond to "no emotion". It is a sort of neutral state, where there is no feeling of being well or bad, excited or calmed. In this unemotional state, it is like emotions are nonexistent. They do not influence behaviour, attention or perception. This state is much like that of being "a machine", in which behaviour tends to detailed calculi and deliberation, without time restrictions. On the other hand, zones that are far from the centre of the emotional space correspond to normal emotional states in humans, though they are rarely contemplated in machines. For Sloman there are only three types of emotions: basic, secondary and tertiary [Sloman, 20011. Picard [Picard, 19971 and Damasio see only two types. Basic emotions come directly from certain stimuli. Other emotions arise after a cognitive appraisal. These two types of emotions are present in CASIMIRO: CHAPTER 7. BEHAVIOUR + Basic emotions: Direct influence from sensors: If the robot hears a loud sound, it increases its arousal and valence, especially arousal. + Secondary emotions: Influence in the Emotions module from ZagaZ (see Section 7.2). The Emotions module maintains a position in a 2D valence and arousal space. The module receives messages to shifi the current position in one or the two dimensions. The 2D space is divided into zones that correspond to a facial expressions. In order to simplify the module, it is assumed that the expression is given by the angle in the 2D space (with respect to the valence axis), and the degree is given by the distance to the origin. The circular central zone corresponds to the neutral facial expression. When the current position enters a different zone a message is sent to the pose editor so that it can move the face, and to the Talk module so that intonation can be adjusted. The facial expressions are assigned to the 2D space as shown in Figure 7.3. Values of arousal and valence are not always inside the exterior circle, though the expression degree is maximum for values that lie outside the circle. Figure 7.3: Assignment of facial expression according to the emotional state 0 Anger . f 1 ' I 1 1 \ 1 .- \ \ Sad , .--,pRelative displacements in the arousal and valence axes need a correction. Consider the case depicted in Figure 7.4 in which the current position in emotional space is P. If we want to lead the robot to Anger, we increase arousal and decrease valence with a displace- + ment d. However, the resulting position will be Q, which is associated to Surprise. Obviously, the effect of changes in arousal and valence depends on the current position, which is undesirable. Surprise . . ' Happiness '\ -\ 5 1 I 1 b 1 v I / / . 0 Sleep The correction, which we have not seen previously in the literature, is as follows. + + Given a displacement d = (v, a), the components of the new displacement vector d' are CHAPTER 7. BEHAVIOUR Figure 7.4: Effect of an increase in arousal and decrease in valence. given by: , where 1 = sqrt((v - vp)' + (a - ap)') and m = sqrt(v2 + a'). Figure 7.5 illustrates the effect of the correction. Figure 7.5: Effect of the correction in emotional space when three (arousal-increase,valencedecrease) displacements are submitted to the system. Note that the position in the emotional space tends to the desired expression. When the current position is at the angle of the desired expression only the distance to the centre increases, which in turn increases the degree of the expression. A very simple decay is implemented: every once in a while arousal and valence are divided by a factor. This does not change the angle in the 2D space, and thus the facial expression does not change, only the degree. This procedure is in accordance with the fact that emotions seem to decay more slowly when the intensity is lower [Bui et al., 20021. In Capítulo 5 Percepción "He visto cosas ... cosas que vosotros pequeñas personas no creeriais. .. naves de ataque en llamas frente al brazo de Orión brillantes como el magnesio. Subi en las cubiertas posteriores de un carguero; he visto rayos C relucir en la oscuridad cerca de la puerta de Tanhauser.. Todos esos momentos ... se irán." - Blade Runner; guión de Hampton Fancher y David Peoples, 1981. E STE capítulo describe las capacidades perceptoras de CASIMIRO. La percepción es uno de los aspectos más importantes de un robot. CASIMIR0 tiene una serie de módulos que le permiten captas algunas características de su entorno, especialmente las relacionadas con personas. La Sección 5.1 describe el módulo de visión omnidireccional. La localización de sonido se explica en la Sección 5.2. Las salidas de estos dos módulos se combinan en un sistema de atención audio-visual, que se describe en la Sección 5.3. En la Sección 5.4 se muestra el módulo de detección facial, en la Sección 5.6 los mecanismos de memoria y olvido. El capítulo finaliza con un estudio de mecanismos de habituación. 5.1 Visión Omnidireccional La mayoría de los robots de interacción usan dos tipos de cámara: una cámara de campo de visión ancho (alrededor de 70°), y10 una cámara foveal. En los últimos tiempos ha crecido el interés en la visión omnidireccional, que permite capturar imágenes que abarcan 360". Cuatro técnicas se usan para lograr esto [Fraunhofer Institut AIS, 2004, Nayar, 19981: + Cámaras con lentes de ojo de pez. + Cámaras con espejos curvos (convexos). Esta variante, llamada catadióptrica, es la más común, ver Figura 5.1. + Conjuntos de cámaras montadas en configuración de anillo o esfera + Una cámara que rota alrededor de un eje y toma una secuencia de imágenes que abarca 360". En cámaras normales la resolución es uniforme, mientras que en sistemas catadióptricos la resolución es mayor en el centro y menor en las zonas exteriores. No obstante esto, las ventajas de tener un campo de visión ancho son obvias, especialmente para determinadas aplicaciones como navegación [Gaspar, 2002, Winters, 20011, vigilancia [Boult et al., 1999, Haritaolu et al., 20001 o grabación de reuniones [Stiefelhagen et al., 20031. Figura 5.1: Montaje típico de visión omnidireccional. CASIMIR0 puede localizar personas que entran en la habitación usando visión omnidireccional. La cámara omnidireccional (un montaje catadióptrico) mostrada en la Figura 4.8 le da al robot un campo de visión de 180°, similar al de los humanos. El software asociado implementado está basado en substracción adaptativa del fondo. El primer paso es descartar parte de la imagen, pues solo nos interesa la zona frontal, cubriendo 180" de lado a lado. Así, la imagen de entrada se enmascara para usar solo la mitad superior de una elipse, que es la forma del espejo tal como se ve desde la cámara. El modelo de fondo se obtiene como el valor medio de un cierto número de frames tomados cuando no hay nadie presente en la sala. Ya en funcionamiento, las imágenes de entrada substraídas son umbralizadas y se aplica el operador de cierre. De la imagen así obtenida se localizan las componentes conectadas y se estima su área. Además, para cada componente conectada, se estima la distancia Euclídea que hay desde el punto más cercano de la componente al centro de la elipse (Figura 5.2), así como el ángulo del centro de masas de la componente con respecto al centro de la elipse y su eje mayor. Figura 5.2: Medida de distancia aproximada tomada con la cámara omnidireccional. En este caso una persona se estaba acercando al robot, desde una distancia de 260cm a 60cm. El módulo de visión omnidireccional también detecta zonas de movimiento, que se usan como un mapa de características en el sistema de atención (ver Sección 5.3). La Figura 5.3 muestra al módulo de visión omnidireccional funcionando. En pruebas iniciales se detectó que la gente tendía a mover el brazo frente al robot, como intentando llamar su atención. Se decidió detectar esto para hacer al robot responder con alguna frase graciosa. 5.2 Localización de Sonido La localización de sonido juegaun papel muy importante en muchos seres vivos, siendo vital para la supervivencia de muchas especies. Los búhos, por ejemplo, pueden cazar en la oscuridad total gracias a sus extraordinarias habilidades de localización de sonido [HHMI, 19971. En humanos, la capacidad de detectar de dónde viene un sonido nos avisa de peligro potencial. La localización de sonido es además un importante mecanismo de fijación de atención, especialmente en un escenario de comunicación verbal. Nuestras habilidades de localización de sonido provienen del hecho de tener dos orejas. Si bien podríamos distinguir diferentes sonidos con solo una oreja, localizar de dónde provienen requiere al menos dos. Las diferencias de las señales recogidas en una y otra oreja + El método da un ángulo que puede usarse directamente por un mecanismo atencional (ver Sección 5.3). Un serio problema para la localización de sonido es el hecho de que los servomotores de la cabeza hacer ruido cuando están en funcionamiento. Realmente, los servomotores hacen ruido incluso cuando no se están moviendo, y también se genera ruido por los ordenadores cercanos. En el humanoide SIG este problema se resolvió usando cuatro micrófonos, dos dentro de la cubierta de la cabeza y dos fuera, cancelando en las señales externas componentes de ruido captadas por los micrófonos internos [Nakadai et al., 20001. Esta ingeniosa solución depende de que la cubierta de la cabeza atenúe el sonido. En las primeras versiones de CASIMIR0 los micrófonos estaban ubicados en la misma cabeza, a la altura de los ojos, y el módulo de localización de sonido solo producía resultados cuando ningún servomotor se estaba moviendo. Esta es obviamente la solución más simple. Actualmente, los dos micrófonos se han ubicado a una distancia de aproximadamente 80cm de los servomotores. Hay una caja de madera (la cabeza descansa sobre esta caja) entre la cabeza y los micrófonos. En su actual posición los micrófonos no capturan mucho ruido de los motores, y sí captan la voz y otros sonidos provenientes del espacio de interacción. 5.3 Atención Audio-Visual El objetivo más importante en los robots sociales es la capacidad de interacción. En este sentido un sistema de atención es crucial, tanto como filtro para centrar los recursos perceptuales del robot como medio de hacer saber al observador que el robot tiene intencionalidad. En esta sección se describe un simple pero funcional sistema de atención. El sistema, implementado en CASIMIRO, fusiona información auditiva y visual obtenida del entorno, y puede incorporar en la atención influencias de más alto nivel basadas en conocimiento. La atención es un proceso de selección por el cual solo una pequeña parte de la enorme cantidad de información sensorial llega a centros de proceso superiores. En interpretación visual, por ejemplo, la atención permite dividir el problema en una rápida sucesión de problemas locales, computacionalmente menos costosos. La literatura relacionada divide la atención humana en dos etapas funcionalmente independientes: una etapa preatentiva, que opera en paralelo sobre todo el campo visual, y una etapa atentiva, de capacidad limitada, que solo procesa un ítem a la vez. La etapa preatentiva detecta estímulos intrínsecamente destacados, mientras que la etapa atentiva lleva a cabo un proceso más costoso y detallado con cada estímulo detectado. Los valores de importancia de la etapa atentiva dependen de la tarea actual, conocimiento adquirido, etc. [Heinke and Humphreys, 20011. El robot Kismet incluía un sistema de atención basado en el modelo "Guided Search 2.0 JGS2)" de Wolfe [Wolfe, 19941. GS2 se basa en extraer de la imagen caracteristicas básicas (color, movimiento, etc.) que son combinadas linealmente en un mapa de activación. Se localiza la región de máxima activación en el mapa y el foco de atención (FOA) se dirige a esa región. Es un hecho ampliamente aceptado que la atención es controlada tanto por ítems de importancia sensorial como por factores cognitivos (conocimiento que se tenga, tarea actual, ...) [Corbetta and Shulman, 20021. En [Scassellati, 20011 se usa un mapa de caracteristicas adicional con el propósito de asignar mayor activación a las zonas de atención compartida entre el robot y una persona. En el modelo implementado para CASIMIRO, los valores de activación son controlados por la siguiente ecuación: donde F y G son funciones que se aplican a los mapa de caracteristicas visuales (f,) y auditivos (gj), con el fin de agrupar zonas de actividad y10 tener en cuenta el error en la posición de las zonas de actividad detectadas. Las posiciones espaciales y temporales se representan en los mapas con las variables p y t. vi, sj y K son constantes. C es una función que da más activación a las zonas cercanas al FOA actual: C(p, t) = e-~p-~~~(~-~) . T( P, t) representa el efecto de módulos de alto nivel, que pueden actuar sobre todo el campo de atención. El máximo del mapa de activación define el nuevo FOA, siempre que sea mayor que un umbral U: maxpA(p,t) simaxpA(p,t)>O FOA(t) = FOA(t - 1) enotrocaso El modelo se muestra en la Figura 5.4, usando sonido y visión para extraer mapas de caracteristicas. Nótese que un mecanismo de atención compartida usaría la componente T de la ecuación 5.2. La implementación en CASIMIR0 usa un mapa auditivo: la localización de una única fuente de sonido. El mapa visual se extrae de imágenes tomadas con la cámara omnidireccional. para la interacción. En esta sección, se ha descrito un sencillo aunque funcional modelo de atención. El modelo se implementó usando características tanto auditivas como visuales extraídas de la zona que rodea al robot. identificación del propietario del robot podría ser factible, pues el problema de reconocimiento se reduce a dos clases (propietario y no-propietario). Aun así, el uso de reconocimiento de caras o de la voz para el reconocimiento del propietario llevaría esencialmente a la misma falta de robustez que en el caso de reconocimiento de caras y voz para varios individuos. ¿Quién es el propietario del robot? La respuesta correcta es: la persona que lo compra. Sin embargo, estamos más interesados en el rol de la persona que cuida del robot y lo usa con más asiduidad (esta persona es normalmente el propietario). Esta persona es la que conecta el robot, lo que se hace generalmente desde una parte concreta del robot o desde un determinado ordenador. En CASIMIRO, una cámara está situada encima del ordenador principal de control (desde el cual se enciende el robot). El módulo de detección de propietario usa la cámara para buscar un blob de color de piel en la imagen. La cámara tiene una lente de ángulo ancho, y una relativamente baja resolución de 160x120. Cuando no se encuentra dicho blob en la imagen el módulo notifica al módulo de atención. En ese momento el módulo termina para no seguir usando CPU. El ordenador principal de control está situado detrás del robot, sobre la misma mesa, ver Figura 5.12. Una vez que ha sido notificado por el módulo detector de propietario el módulo de atención considera como propietario al primer blob que entra en la imagen de la cámara omnidireccional por la izquierda. La propiedad "propietarioMse almacena asociada al individuo durante el seguimiento 1D. Figura 5.12: Ordenador desde donde se enciende CASIMIRO. El espacio de interacción queda a la izquierda. Los investigadores en reconocimiento de caras tienden a evaluar los algoritmos en función del número de individuos que el sistema puede reconocer y a la tasa de error medida. Una tasa de error medida de hasta un 10% puede considerarse como muy buena. El enfoque presentado aquí reconoce un único individuo con un error cero garantizado. 5.8 Habituación Los seres vivos poseen mecanismos de habituación que les permiten ignorar estímulos repetitivo~. Si tales estímulos no fueran gradualmente ignorados, la respuesta continuada llevaría al ser vivo al completo agotamiento. Una gran cantidad de información se recibe continuamente del entorno, y debe ser filtrada de alguna forma para que el agente pueda fijarse en los datos interesantes. Muchos animales, incluidos los humanos, tienen algún tipo de mecanismo para filtrar estímulos no interesantes. El modelo de habituación de Stanley [Stanley, 19761, propuesto para simular los datos de habituación obtenidos de la médula espina1 del gato, se ha usado ampliamente en la bibliografía. El modelo describe la disminución de eficiencia y de una sinapsis por una ecuación diferencial de primer orden: donde yo es el valor normal, inicial de y, S(t) representa la estimulación externa, T es una constante de tiempo que controla la tasa de habituación y a regula la tasa de recuperación. La Ecuación (5.5) asegura que la eficacia sináptica decrece cuando la señal de entrada S(t) aumenta, y retorna a su máximo yo en ausencia de señal de entrada. El modelo de (5.5) solo puede explicar la habituación a corto plazo. Wang introdujo un modelo para incorporar habituación tanto corto como a largo plazo, usando una curva en forma de S invertida [Wang, 19951, donde a, yo y y tienen el mismo significado que en (5.5), 0 regula la habituación y z(t) decrece monótonamente con cada activación del estímulo externo S(t), y modela la habituación a largo plazo. Debido a este efecto de z(t) tras un gran número de activaciones la tasa de recuperación es más lenta. Si usamos el modelo de la Ecuación (5.5) podemos obtener efectos no deseados con determinados estimulos. Una señal de entrada periódica (con frecuencia mayor que cero) puede producir una respuesta que no muestra habituación. Esto se debe al hecho de que el modelo no tiene en cuenta estímulos cambiantes. Para incluir este aspecto en el modelo, proponemos usar una señal auxiliar que será cero cuando el estimulo sea estacionario o de frecuencia fija, y uno en otro caso, y usar esta señal como entrada al modelo de habituación (5.5). La señal auxiliar, que básicamente detecta estimulos monótonos, se obtiene del espectrograma del propio estimulo. El espectrograma es la distribución tiempo-frecuencia de una señal. Está basad en la Transformada de Fourier con una ventana móvil [Holland et al., 20001. La ecuación: da la definición de espectrograma con una ventana Gaussiana de ancho 2T, y es el espectro de potencia de una señal que corresponde a la magnitud cuadrada de la Transformada de Fourier de la señal "ventaneada". La ventana puede tener otras formas además de la Gaussiana. En la Figura 5.13 se muestra una señal de audio y su espectrograma correspondiente, donde las regiones más brillantes corresponden a mayor potencia. Se pueden distinguir dos espectros de frecuencia bien definidos, pues hay un cambio en la señal de entrada en 0.5s. Los patrones temporales de una señal estimulo tienen un patrón especifico en el espectrograma. Una señal de frecuencia fija corresponde a una linea recta paralela al eje temporal en el espectrograma. La longitud de esta linea indica cuánto tiempo ha estado presente el estimulo. O O1 O2 O3 O4 O5 O6 O7 O8 O9 1 Time (sec) O O1 O2 O3 O4 O5 O6 O7 O8 O9 Time Figura 5.13: Señal de audio (izquierda) y su correspondiente espectrograma (derecha). Los espectrogramas se calculan a partir de ventanas de la señal de entrada. Estas ventanas, de longitud 1, se solapan en 1 - 1 muestras. Representemos cada espectrograma como una matriz M, en la cual las filas representan frecuencias y las columnas representan tiempo. Calculamos la varianza de cada fila de M, lo que produce un vector columna v. La norma de este vector es una medida de cuan monótona es la señal de entrada. La norma será alta cuando la señal es cambiante, y baja en caso contrario. Así, la señal auxiliar que necesitamos es simplemente la norma umbralizada de v. La amplitud de la señal de entrada afecta al contenido de potencia de los espectrogramas, y a su vez a la norma de v. Por ello, antes de calcular la FFT la señal de entrada debe ser normalizada dividiendo cada ventana de entrada por la suma de sus valores absolutos. Un valor de 1 para la señal auxiliar significará que hay cambios en la señal de entrada, mientras que un valor de O indica que la señal de entrada es monótona. Una vez la señal auxiliar está disponible, se usa el modelo (5.5) para obtener el comportamiento de habituación deseado, controlando los parámetros T y a. El algoritmo descrito en la sección anterior fue implementado para probarlo con diferentes señales de entrada. Se grabó un vídeo conteniendo un estímulo amarillo brillante (una tarjeta amarilla) que se movía de forma repetitiva, ver Figura 5.14-a). Usando técnicas sencillas de segmentación se obtuvo el centroide de la tarjeta en cada frame, y se sumaron sus coordenadas x e 9 para formar la señal unidimensional de la Figura 5.14-b). LA secuencia de movimientos de la tarjeta durante la grabación fue: movimiento horizontal, movimiento aleatorio, movimiento vertical y movimiento vertical a una frecuencia diferente del anterior. l 1 O 20 30 40 50 60 Time (sec) Figura 5.14: a) Vídeo usado para el experimento de habituación visual, b) señal unidimensional extraída de él. Los resultados aparecen en la Figura 5.15. Se usaron ventanas de 128 muestras, y el umbral de varianza se situó en 1000. l 20 30 40 50 60 Time (sec) Time (sec) b) Figura 5.15: a) Evolucón de la norma l2 del vector de varianzas v, b) nivel de habituación, usando r = 5, a = 1. El mecanismo de habituación descrito se implementó en CASIMRO para señales del dominio visual. Se calcula la diferencia entre el frame actual y el anterior. Esta diferencia es umbralizada y filtrada con operadores Open y Close. Los blobs resultantes de área muy pequeña se descartan. A continuación se obtiene el centro de masas de la imagen resultante. La señal que entra al algoritmo de habituación es la suma de las componentes z e y del centro de masas. De esta forma cuando la imagen no cambia significativamente o hay presentes movimientos repetitivos durante un cierto tiempo la señal de habituación crece. Cuando crece por encima de un umbral se envía una señal de inhibición al módulo de atención, el cual cambia entonces el foco de atención. Los movimientos de cuello producen cambios en la imagen, pero se observó que no eran periódicos, y por tanto no hacían crecer la habituación. Para señales de audio el algoritmo funciona demasiado lento. Podría aplicarse a señales de audio si las señales se comprimieran en tiempo real. Capítulo 6 Acción CHRISTOF: Nos hemos aburrido de ver actores que transmiten falsas emociones. Nos hemos cansado de la pirotécnia y los efectos especiales. Si bien el mundo en que él habita es en algunos aspectos falso, no hay nada falso en Truman. No hay guiones, no hay cartas marcadas ... No siempre es Shakespeare, pero es genuino. Es una vida. - El Show de Truman, Guión de Andrew Niccol. E STE capitulo describe las acciones que CASIMIR0 puede realizar. La Sección 6.1 explica cómo se han conseguido las expresiones faciales. La generación de voz se describe en la Sección 6.3, junto con breves descripciones del habla expresiva y aspectos de humor en el lenguaje. 6.1 Expresiones Faciales La cara de Kismet tiene cejas (cada una con dos grados de libertad), párpados (un grado de libertad) y boca (un grado de libertad). La boca ha ganado más grados de libertad en las últimas versiones. Kismet puede adoptar expresiones de enfado, fatiga, temor, disgusto, excitación, felicidad, interés, tristeza y sorpresa, todas ellas fácilmente reconocibles por un observador humano. El sistema de movimiento facial en Kismet se dividió en tres niveles. En un primer nivel hay procesos que controlan cada motor. En el siguiente nivel hay procesos que coordinan el movimiento de los rasgos faciales, como por ejemplo una ceja. En el tercer nivel hay procesos que coordinan los rasgos faciales para formar las expresiones. Este Expresión: "Sorpresa" Grupo: Boca Pose: Abierta Grado: 90 Grupo: Ceja derecha Pose: Levantado Grado: 90 Grupo: Ceja izquierda Pose: Levantada Grado: 90 Grupo: Oreja derecha Pose: Levantada Grado: 100 Grupo: Oreja izquierda Pose: Levantada Grado: 100 Grupo: Párpado derecho Pose: Levantado Grado: 80 Grupo: Párpado izquierdo Pose: Levantado Grado: 80 Tabla 6.1 : Típica definición de una expresión facial. esquema de tres niveles permite descomponer el trabajo de modelado de una forma natural y escalable. Kismet también usa una intensidad para cada expresión, que es un grado con respecto a una pose considerada neutral. Las expresiones faciales en CASIMIR0 usan la jerarquía de tres niveles comentada. Primero se definieron grupos de motores que controlan un rasgo facial concreto. Por ejemplo, dos motores se agrupan para controlar una ceja. Para cada uno de los grupos de motores se define también las poses que el rasgo facial puede adoptar, como 'ceja derecha neutral','ceja derecha levantada',etc. Las transiciones por defecto entre las diferentes poses usan una línea recta en el espacio de valores de control de motor. El modelador tiene la oportunidad de modificar estas transiciones por defecto, si es que alguna de ellas parece poco natural. La primera pose que el modelador debe definir es la pose neutral. Todas las poses que se definan se refieren a un grado máximo, 100. Ya en funcionamiento, cada pose puede adoptarse en un cierto grado entre O y 100. El grado se usa para interpolar linealmente los puntos en la trayectoria con respecto a la pose neutral. En un nivel superior, las expresiones faciales hacen referencia a poses de los diferentes grupos, cada uno con un grado diferente. "Sorpresa", por ejemplo, podría representarse como en la Tabla 6.1. La expresión facial se especifica cuando el sistema está funcionado, junto con un grado que permite, por multiplicación, obtener el grado a aplicar a las poses de los diferentes grupos. Para un mayor control, el modelador puede también especificar un tiempo de comienzo para cada grupo. De esta forma, "sorpresa"podría lograrse elevando primero las cejas y abriendo después la boca. El funcionamiento de la cara nos lleva preguntarnos sobre las transiciones entre las expresiones: ¿Cómo obtener la trayectoria desde una pose A con grado Gj a una pose B con grado Gf! En un momento dado un grupo está en la pose A con grado Gi y queremos que adopte la pose B con grado Gf. Para un grupo de dos motores el procedimiento se muestra en la Figura 6.1, donde N representa a pose neutral, X el punto correspondiente al grado inicial e Y el punto correspondiente al grado final. Para obtener la expresión que da la trayectoria de transición se imponen las siguientes condiciones límite: La trayectoria que se ajusta a estas restricciones es: Los símbolos ? y : corresponden a la instrucción IF-THEN-ELSE (como en el lenguaje C). La ecuación de la trayectoria es continua en los valores de Gi y Gf. La misma relación se usa para obtener velocidades a lo largo de la trayectoria. m1 Figura 6.1 : Transiciones entre expresiones en espacio motor. El esquema de modelado descrito se implementó como un editor de poses (Figura 6.2). El editor da al modelador una forma simple e interactiva de definir y probar poses, transiciones, etc. Usa una librería de bajo nivel para el control de los servomotores. Las especificaciones de poses y transiciones pueden salvarse en un fichero. Más tarde, este fichero será lo único necesario para reproducir (y generar) los diferentes movimientos y trayectorias. Seivo 1 fi ' 2414 Seivo3 u Seivo 2 A 2146 Seivo10 u Seivo 3 A d1800 ~eivo11u Seivo 4 A 1031 Seivo 12 Seivo 5 d 1841 Seivo 13 U Seivo 6 A A 2024 Seivo 14 Siempre visible r Seivo 7 4 A 1845 Seivo 15 Seivo 8 fi do Seivo16U d 0 Paso ScrollBari 150 Suav~dad: 110 Umbral de distancia 1100 Boca(1) Cela Derechal2.31 Cela lzquierda[4.5! Orela Derecha[6] Orela lrquitrda[7) r, Añadir 1 1 Porer de segundo nivel - Nombra. r Poser de prmer nvel y sus tiansiciones I 1 rd "termedios: Velocidad incial Velocidad entre el pto nlerrned~a y el slgulente Añadir 1 insertar 1 Añadir 1 l-3 Quitar Quitar 1 Qu~tai Giado in~ciaI[%l m 1 <<, ;, '>', >[ >>,' Gmdof~nalR)m Grupos Cangiada[%j. m Pose de prmer nivel 10rela *lrb Giado de la pose m de le, nivel [%! Tiempo de comienzo [ms] 10 Ir a lo Pois - 1 Figura 6.2: Ventana principal del editor de poses. El editor de poses puede por tanto funcionar tanto para tiempo de diseño como para tiempo de ejecución, y también puede ser controlado por otros módulos. El parpadeo tiene gran importancia en el robot. En ciertos experimentos formas humanas parpadeantes fueron consideradas como poseedoras de mayor precisión e inteligencia que otras formas humanas [King and Ohya, 19961. En nuestra implementación. el parpadeo se trata como un caso especial en el que el movimiento de párpados se realiza a la máxima velocidad. El repertorio de poses puede actualizarse fácilmente (si se añaden motores). La Figura 6.3 muestra las expresiones faciales modeladas en CASIMIR0 usando el editor de poses. Cuello Como se comentó en la Sección 4.1 un cuello de 2 grados de libertad mueve la cabeza del robot. Para pequeños movimientos se usa un controlador PID. Para grandes movimientos del cuello solo se usan comandos de posicionamiento. De esta forma el cuello es capaz de hacer Beber: Precondiciones: En-sitio-A Add list: Sin-sed Delete list: Comer: Precondiciones: En-sitio-B Add list: Sin-hambre Aproximarse a comida: Precondiciones: Add list: En-sitio-B Delete list: En-sitio-A Aproximarse a agua: Precondiciones: Add list: En-sitio-A Delete list: En-sitio-B Tabla 7.1: Comportamientos ejemplo. Las metas son Sin-hambre y Sin-sed. aunque no del todo, puesto que en cada paso las entradas del entorno pueden hacer que el sistema elija cualquier comportamiento. Por tanto, se contemplan los dos casos: reactividad y planificación. El módulo de selección de acción de CASIMIR0 está basado en la herramienta ZagaZ [Hernández-Cerpa, 20011. ZagaZ es una implementación de las redes de comportamientos. Actualmente, las capacidades de ZagaZ no están totalmente explotadas en CASIMIRO, aunque se espera mejorar esto en el futuro. El sistema de selección de acción está funcionando como un sistema de reglas con prioridad, por lo que no se usan las listas Add y Delete. Dichas listas están en realidad implementadas a través del módulo de memoria, lo que permite ejercen un mayor control sobre los datos que se memorizan y que se olvidan. El comportamiento de CASIMIR0 se hizo intencionalmente simple. El robot entra en interacción con los individuos que entran en su espacio de interacción (una zona de 180" frente a él). Intenta decir poemas, bromas y datos sobre sí mismo a los individuos, en ese orden. La estabilidad básica de la interacción se mantiene: cada persona es saludada, y a veces el robot hace también comentarios de "continuidad, comentarios jocosos. Puede también detectar cuando la persona está distraída, lo que le hace cambiar de tema (entre poemas, bromas y datos sobre sí mismo). También puede preguntar al individuo si quiere oír más frases del tema actual o cambiar de tema. Cuando el individuo no colabora el robot se enfurece. 7.3 Emociones Se han propuesto muchos modelos emocionales tanto dentro de la comunidad en IA como en psicología. El modelo más conocido es el de Russel [Russell, 19801, que considera que las emociones caen en un espacio bidimensional, con componentes ortogonales de excitación y valencia, ver la Figura 7.1. ANGRY JOYFUL SAD Valence RELAXED Figura 7.1 : Espacio emocional de excitación y valencia. La zona central del espacio correspondería a la "no emoción". Es por tanto un estado neutral, donde no habría sensación de estar bien o mal, excitado o calmado. En este estado las emociones no influencian el comportamiento, la atención o la percepción. Para Sloman hay tres tipos de emociones: básicas, secundarias y terciarias. Las emociones básicas se generan directamente a partir de determinados estímulos. Las otras emociones sobrevienen tras una valoración afectiva. El módulo emocional mantiene una posición en el espacio 2D de excitación y valencia. El módulo recibe mensajes para desplazar la posición actual en una o las dos dimensiones. El espacio 2D se divide en zonas que corresponden a expresiones faciales. Se asume que la expresión está dada por el ángulo en el espacio 2D, y el grado de la expresión está dado por la distancia al origen del espacio. Las emociones que el robot ha sentido durante la interacción con un individuo se almacenan en una memoria asociada a ese individuo particular (véase la Sección 5.6). Por otra parte, la distancia desde la posición actual al centro del espacio emocional se usa para enviar al módulo de atención un mensaje para reducir la fijación en el foco de atención actual. Capítulo 8 Evaluación y Discusión 'So importante es no parar de hacerse preguntas." Albert Einstein. L OS módulos que constituyen CASIMIR0 han sido evaluados independientemente en los capítulos previos. La evaluación de esos subsistemas es relativamente directa. Básicamente, cada técnica o subsistema se compara con otras técnicas de referencia. El rendimiento de los subsistemas se puede evaluar a través de resultados numéricos, pudiéndose aplicar técnicas estadísticas y extraer conclusiones relativamente sólidas. Pero ¿cómo evaluar el rendimiento de un robot social completo? En una situación ideal tendríamos que tener dos robots diferentes disponibles, y evaluarlos en las mismas tareas y bajo las mismas condiciones. Eso permitiría obtener medidas comparativas. Sin embargo, esa opción es claramente inviable debido a la complejidad y costo de los robots y su soporte físico y humano asociado. Se requieren técnicas más subjetivas. Este capítulo se centra en observar y analizar cómo funciona CASIMIR0 globalmente. La Sección 8.1 hace una revisión de la literatura sobre la cuestión. Luego se describen ejemplos de interacciones breves. Estos ejemplos permiten al lector conocer como se desarrolla una sesión de interacción típica. En las Secciones 8.3 a 8.4 se muestran los resultados de una serie de entrevistas en las que personas manifiestan sus impresiones sobre el robot. 8.1 Evaluación de Robots Sociables Con robots industriales o de navegación la evaluación del rendimiento debe incluir aspectos como la precisión, repetibilidad, autonomía, velocidad, consumo de batería, etc. No es extraño que estas medidas resulten inapropiadas en nuestro contexto de robots sociables. De hecho, parecería que son contrarias a lo que se se debería medir. Por encima de todo, los robots interactivos no tienen aún una tarea definida como en el caso de los robots industriales. Scholtz y Bahrami [Scholtz and Bahrami, 20031 argumentan que este hecho hace que medidas típicas de evaluación en interacción hombre-máquina como eficiencia, efectividad y satisfacción del usuario resulten inapropriadas. El enfoque natural es intentar sistematizar la inherentemente subjetiva evaluación del robot. En este sentido el método preferido parece ser el de pasar cuestionarios a los individuos que han interactuado con el robot. Sin embargo, muchos aspectos son críticos, como las preguntas exactas a realizar y las escalas en las respuestas. Siempre existe el riesgo de que la pregunta no cubra exactamente el aspecto que se quiere analizar. Duffy y colegas [Duffy et al., 20021 argumentan que el campo de la robótica social hereda un problema metodológico de la sociología, a saber, la cuestión de si métodos cuantitativos capturan fielmente la complejidad del dominio social. Ellos abogan por una aproximación cualitativa, por dos razones: 1) se pierde información cuando las situaciones y escenarios se reducen a números, y 2) los métodos cuantitativos tienen a menudo bases cualitativas. A pesar de las limitaciones inherentes, el uso de cuestionarios y escalas numéricas para las respuestas parece ser el método preferido. La observación y análisis de sesiones de interacción ha sido ampliamente utilizada también, especialmente en los trabajos de Dautenhahn. Ambas técnicas han sido empleadas en nuestro trabajo. 8.2 Ejemplos de Interacción La Tabla 8.1 muestra un ejemplo de sesión de interacción con el robot. En este caso a persona entra en el área de interacción y toma una actitud poco cooperativa. La Figura 8.1 muestra la evolución del estado emocional del robot en esa sesión y en una sesión con una persona más cooperativa. - Tiempo (S) o 18 23 29 3 1 35 39 40 47 50 5 1 57 59 67 70 74 75 80 86 93 97 98 1 O6 108 114 115 120 - Observaciones robot arranca robot saluda la persona habla, el robot le reprende la persona se acerca mucho, el robot le reprende la persona responde Sí la persona habla la persona se mueve la persona agita su brazo el robot intenta cambiar de tema. Respuesta=% la persona responde No Comportamientos considerados QuestionDoYouWantPoem TalkAboutSpeaking,TalkAboutPoem TalkAboutPoem TalkAboutSthingFunny,QuestionDoYouLikeMe TalkAboutPoem,QuestionDoYouLikeMe QuestionDoYouLikeMe TalkAboutSthingFunny,DecreaseValence, DecreaseArousal TalkAboutPoem,DecreaseValence,DecreaseArousal Estado emocional Neutral 100% Enfado 70% Anger 35% Enfado 17% Neutral 100% Enfado 70% Enfado 35% Enfado 17% Enfado 88% Enfado 100% Enfado 57% Enfado 28% Enfado 14% Tabla 8.1: Interacción ejemplo una persona interactuando con el robot. Los comportamientos en negrita fueron ejecutados por el robot. 8.3 Entrevistas Muchos investigadores tienden a pensar que la mejor evaluación para un robot social es el efecto que tiene sobre la gente. Una medida típica es el tiempo de interacción: cuando más largas las sesiones de interacción mejor el robot. En las entrevistas se pregunta a la gente sobre aspectos como entretenimiento o empatía con el robot. Pero todo esto está asumiendo que el cometido fundamental del robot es producir un efecto positivo en la gente (al menos, eso es lo que se que pretende medir). Desde nuestro punto de vista tales aspectos de empatía no deberían ser centrales en la evaluación. Una cosa así sería equivalente a establecer si una película ha sido rodada time (S) Figura S. 1: Valores de valencia en una sesión con una persona cooperativa (izquierda) y con una persona poco cooperativa (derecha). correctamente solo en base a las opiniones de un grupo de espectadores. Además, es extremadamente difícil, si no imposible, aislar aspectos como empatía o duración de la interacción de características de las personas que interactúan, como edad, profesión, estado de ánimo, nivel de conocimientos, etc. Es un hecho que incluso un sencillo programa de ajedrez puede dar lugar a tiempos de interacción muy largos dependiendo de la persona. Cuando se compara interacciones con adultos y con niños, se ha demostrado que éstos parecen ser más receptivos con el robot, dándole una valoración más positiva [Walters et al., 20051. En este trabajo nuestro objetivo no es tanto producir un robot que entretenga personas como imitar habilidades sociales básicas. Por otro lado, hay un aspecto que a menudo se pasa por alto. Se espera que el lector de este documento tenga amplios conocimientos en robótica o inteligencia artificial. Si queremos dar al lector una visión fiel del robot una buena opción sería reproducir aquí las impresiones de otras personas con parecido nivel de conocimientos. Por estas razones, todos los sujetos a los que se entrevistó (19) tienen una base de conocimientos en informática. Todos los entrevistados eran al menos estudiantes de informática, con algunos ya teniendo el grado de Doctor en informática o alguna ingeniería (el 37.5%). Todos tenían cierta familiaridad previa con el robot, y algunos incluso contribuyeron con software creado para otros proyectos. En todos los casos la interacción se estructuraba de la siguiente forma. Cuando el sujeto está en las inmediaciones del robot, éste se pone en marcha. Entonces se desarrolla la interacción. En algún momento el controlador (la persona que conectó el robot) entra en el espacio de interacción. Desde ese momento, dos personas están en el espacio de interacción del robot. La sesión de interacción finaliza cuando el controlador ve que el robot se está aburriendo (lo cual tiende a ocurrir cuando el robot no tiene nada más que decir a la persona). Durante la interacción el controlador trata de atraer la atención del robot al menos una vez. Tras la sesión de interacción los sujetos completaron un cuestionario. El cuestionario se divide en tres secciones más una pregunta adicional. Las tres secciones son: 1) "Entiendo al robot", 2) "El robot me entiendeMy 3) "Impresión global". Las primeras dos secciones permitirían inferir hasta qué punto el robot tiene mínimas habilidades social bidireccionales. Las preguntas del cuestionario se muestran en la Tabla 8.2. Sección 1 : 1.He entendido todo lo que el robot me ha dicho 2.E1 robot ha transmitido sus emociones a través de expresiones faciales 3.E1 robot ha transmitido sus emociones a través del tono de voz 4.El robot me presta atención 5.E1 robot es consciente de mi presencia 6.E1 robot es consciente de mis movimientos 7.El robot reconoce personas 8.Este robot es un buen punto de partida para mantenerme informado Sección 2: 9.E1 robot entiende lo que la gente dice 10.E1 robot entiende las expresiones faciales de la gente 11.El robot sabe a donde dirijo mi atención 12.Tal como está, este robot podría usarse para que aprenda nuevas cosas de las personas Sección 3: 13.No he tenido que hacer mucho esfuerzo para adaptarme al robot 14.E1 robot ha tenido muchos fallos (cosas que es obvio debía hacer y no hizo) 15.¿Qué crees que debería tener el robot para que fuera usado con frecuencia? Tabla 8.2: Cuestionario usado para evaluar al robot. Los entrevistados tenían que dar entre 1 y 5 puntos a cada pregunta (1 significa no o totalmente en desacuerdo, 5 significa sí o totalmente de acuerdo). La última pregunta permitía respuesta libre. Nótese que no preguntamos a los sujetos por cuestiones como empatía o entretenimiento. Estamos fundamentalmente interesados en dotar al robot de habilidades básicas de interacción, por lo que la palabra "transmitir"es muy frecuente en las preguntas. Habituación a estímulos repetitivos Animación de los rasgos faciales Detección del propietario del robot 9.3 Trabajo Futuro A nivel conceptual deberemos indagar más en las implicaciones del hecho de que las habilidades sociales son fundamentalmente subconscientes. Creemos que esta especificidad del problema merece un análisis adicional. En particular, sería muy importante establecer hasta qué punto pueden la introspección y la observación indirecta hacer visible los entresijos de los procesos subconscientes. Desde un punto de vista práctico, podemos identificar una cuestión que merece la pena ser desarrollada en el futuro: el reconocimiento de personas. Generalmente, se piensa que la mejor forma de probar algoritmos de reconocimiento de caras es mostrar el rendimiento para un gran número de individuos. Desde nuestro punto de vista, sería mejor garantizar un buen rendimiento aunque sea para un pequeño número de individuos. Para reconocer personas de entre un pequeño grupo (una familia, por ejemplo) podrian ser apropiadas otras técnicas, como la altura del individuo. Reconocer personas y mantener algún tipo de memoria sobre las mismas es tan importante para la interacción que debería explorarse también técnicas más novedosas. La identificación por radiofrecuencia (RFID, por sus siglas en inglés) es una técnica que se está abriendo camino en numerosos campos, especialmente en la identificación de artículos de consumo. Estos sistemas se componen de un aparato emisor y una etiqueta receptora. El aparato emisor emite una señal de radio que las etiquetas pueden detectar y a las que pueden responder. El aparato emisor puede de esta forma identificar la etiqueta (que emite un código único). Si las etiquetas llegan a fabricarse en formato pegatina, podrian ser transportadas fácilmente por las personas a ser identificadas por el robot (en el móvil, llaves, etc.). Hay ya etiquetas disponibles comercialmente de 0.16mm2, más finas que una hoja de papel, si bien los alcances de la lectura son todavia pequeños (10mm a 1 metro). Por otra parte, el reconocimiento del habla es un tema prometedor que se estudiará en el futuro. Los resultados del cuestionario de evaluación muestran claramente que la gente todavia siente que el robot no les entiende. En CASIMIR0 no se implementó reconocimiento del habla porque en ningún momento se consiguió un rendimiento mínimamente aceptable sin llevar el micrófono cerca de la boca del hablante. El reconocimiento del habla a distancia es probablemente uno de los temas de investigación mas interesantes en interacción hombre-máquina. La utilidad es clara, sobre todo porque permitiría a los usuarios liberarse de los micrófonos. Dos técnicas parecen especialmente atractivas. Por una parte, los arrays de micrófonos pueden filtrar selectivamente señales provenientes de determinadas direcciones. Por otra parte, el reconocimiento de habla audio visual usaría no solo las señales de audio, sino las imágenes de los labios moviéndose al hablar. Bibliografía [Adams et al., 20001 B. Adams et al.. Humanoid Robots: A New Kind of Tool. IEEE Intelligent Systems, vol. 15(4), 25-31, 2000. URL citeseer . nj . nec . com/ adamsO0humanoid.html. [Adini et al., 19971 Y. Adini et al.. Face Recognition: The Problem of Compensating for Changes in Illumination Direction. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 19(7), July 1997. [Baars, 19881 B. J. Baars. A cognitive theory of consciousness. Cambridge University Press, NY, 1988. [Berlyne, 19601 D. Berlyne. ConJict, arousal and curiosity. McGraw Hill, N.Y., 1960. [Beymer, 19931 D. Beymer. Face Recognition Under Varying Pose. AIM-1461, p. 14, 1993. URLciteseer.ist.psu.edu/beymer94face.html. [Binsted, 19951 K. Binsted. Using humour to make natural language interfaces more friendly, 1995. URL citeseer .nj .nec . com/binsted95using.html. [B.J. Baars, 20041 B.J. Baars. Recovering consciousness: A timeline, 2004. URL http : //www.sci-con.org/news/articles/20020904.html. [Blauert, 19831 J. Blauert. Spatial hearing. MIT press, Cambridge, MA, 1983 [Blumberg, 19971 B. Blumberg. Old Tricks, New Dogs: Ethology andInteractive Creatures. Ph.D. thesis, MIT Media Lab, 1997. [Bouguet, 19991 J. Bouguet. Pyramidal Implementation ofthe Lucas Kanade Feature Tracker. Tech. rep., Intel Corporation, Microprocessor Research Labs, OpenCV documents, 1999. [Boult et al., 19991 T. Boult et al.. Frame-rate omnidirectional surveillance and tracking of camouflaged and occluded targets. pp. 48-58, 1999. [Breazeal et al., 20041 C. Breazeal et al.. Humanoid Robots as Cooperative Partners for People. International Journal ofHumanoidRobots jsubmitted), 2004. [Breazeal, 20021 C. L. Breazeal. Designing social robots. MIT Press, Cambridge, MA, 2002. [Brooks, 19861 R. Brooks. Arobust layered control system for amobile robot. IEEE Journal ofRobotics and Automation, (2), 1 &23, 1986. [Brooks et al., 19991 R. Brooks et al.. The Cog Project: Building a Humanoid Robot. Lecture Notes in Computer Science, vol. 1562,52-87, 1999. URL ci teseer . nj . nec . com/brooks99cog.html. [Bruce et al., 20011 A. Bruce et al.. The role of expressiveness and attention in humanrobot interaction, 2001. URL citeseer .nj .nec . com/bruceOlrole .html. [Bullock, 19831 D. Bullock. Seeking relations between cognitive and social-interactive transitions. In Levels and transitions in children's development: new directions for child development (edited by K. Fischer). Jossey-Bass Inc., 1983. [Burges, 19981 C. Burges. A Tutorial on Support Vector Machines for Pattern Recognition. Data Mining andKnowledge Discovery, vol. 2(2), 121-167, 1998. [Castrillón, 20031 M. Castrillón. On Real-Time Face Detection in Video Streams. An Opportunistic Approach.. Ph.D. thesis, Universidad de Las Palmas de Gran Canaria, March 2003. [Chomsky, 19801 N. Chomsky. Rules and representations. Behavioral and Brain Sciences, (3), 1-21, 1980. [Corbetta and Shulman, 20021 M. Corbetta and G. Shulman. Control of goal-directed and stimulus-driven attention in the brain. Nature reviews neuroscience, vol. 3, 201-215, March 2002. [Damasio, 19941 A. Damasio. Descartes ' Error: Emotion, Reason and the Human Brain. Avon Books, New York 1994. [Darrell et al., 19981 T. Darrell et al.. Integrated Person Tracking Using Stereo, color, and Pattern Detection. In Procs. ofIEEE Computer Society Conference on Computer Vision andpattern Recognition, pp. 601-608. Santa Barbara, CA, 1998. [Dautenhahn, 19951 K. Dautenhahn. Getting to know each other - Artificial social intelligence for social robots. Robotics andAutonomous Systems, (6), 333-356, 1995. [Dautenhahn and Weny, 20011 K. Dautenhahn and 1. Weny. The AURORA Project: Using Mobile Robots in Autism Therapy. Learning Technology online newsletter; publication ofIEEE Computer Society Learning Technology TaskForce JLTTF), vol. 3(1), 2001. [Duffy et al., 20021 B. Duffy et al.. Issues in Assessing Performance of Social Robots. In Procs. of the 2nd WSEAS International Conference on Robotics, Distance Learning and Intelligent Communication Systems. Skiathos Island, Greece, September 2002. [Fong et al., 20031 T. Fong et al.. A survey of socially interactive robots. Robotics and Autonomous Systems, vol. 42(3-4), March 2003. [Fraunhofer Institut AIS, 20041 Fraunhofer Institut AIS. Omnidirectional Imaging for Robotic Applications, 2004. URL http : //ais . gmd . de/services/ OmniVision/omni-intro.htm1. [Gardner, 19831 H. Gardner. Frames of mind: The theory of multiple intelligences. Basic Books, New York 1983. [Gaspar, 20021 J. Gaspar. Ornnidirectional vision for mobile robot navigation. Ph.D. thesis, Instituto Superior Técnico, Universidade Técnica de Lisboa, December 2002. [GCAT, 19991 GCAT. Perception of Direction, 1999. Available at http : //www . gcat . clara.net/Hearing/perception - of - direction.htm. [Georghiades etal., 19981 A. Georghiades et al.. Illumination cones for recognition under variable lighting: Faces, 1998. URL citeseer . ist .psu . edu/ georghiades98illumination.html. [Goetz, 19971 P. Goetz. Attractors in recurrent behavior networks. Ph.D. thesis, Graduate School of the State University of New York at Buffalo, August 1997. [Grange et al., 20021 S. Grange et al.. Vision based Sensor Fusion for Human-Computer Interaction, 2002. URL citeseer .nj .nec . com/grangeO2visionbased. html. [Gross et al., 20021 R. Gross et al.. Fisher light-fields for face recognition across pose and illumination, 2002. URL citeseer. ist .psu. edu/gross02fisher. html. [Gross et al., 20041 R. Gross et al.. Face Recognition Across Pose and Illumination. In Handbook of Face Recognition (edited by S. Z. Li and A. K. Jain). Springer-Verlag, June 2004. [H. von Helmholtz, 19241 H. von Helmholtz. Physiological Optics. Optical Society of America 1924. [Haritaolu et al., 20001 1. Haritaolu et al.. W-4: Real-time surveillance of people and their activities. IEEE Transactions Pattern Analysis, and Machine Intelligence, vol. 22(8), 809-830,2000. [Hartmann, 19991 W. Hartmann. How we localize sound. Physics Today, vol. 52(1 l), 2429, 1999. [Heinke and Humphreys, 20011 D. Heinke and G. Humphreys. Computational Models of Visual Selective Attention: A Review. In Connectionist models in psychology (edited by G. Houghton), 2001. [Hernández-Cerpa, 20011 D. Hernández-Cerpa. ZagaZ: Entorno Experimental para el Tratamiento de Conductas en Caracteres Sintéticos. Master's thesis, Universidad de Las Palmas de Gran Canaria, 2001. [Hernández Tejera et al., 19991 F. Hernández Tejera et al.. DESEO: An Active Vision System for Detection, Tracking and Recognition. In Lectures Notes in Computer Science, International Conference on Vision Systems jICVS'99) (edited by H. 1. Christensen), vol. 1542, pp. 379-391, 1999. [HHMI, 19971 HHMI. Seeing, Hearing and Smelling the World. Tech. rep., Howard Hughes Medical Institute, 1997. Available at http : //www . hhmi . org/senses/c210 . html. [Holland et al., 20001 S. Holland et al.. Determination of plate source, detector separation fron one signal. Ultrasonics, vol. 38, 620423, 2000. [Hutchinson, 19581 G. Hutchinson. Concluding Remarks. In Cold Spring Harbor Symp. Quant. Biol., 22, pp. 415-427, 1958. [Kidd and Breazeal, 20031 C. Kidd and C. Breazeal. Comparison of social presence in robots and animated characters, 2003. HCI Journal, special issue on Human-Robot Interaction (Submitted). [King and Ohya, 19961 W. King and J. Ohya. The representation of agents: Anthropomorphism, agency and intelligence. In Proc. of CHI-96, 1996. [Kozima, 20021 H. Kozima. Infanoid: An experimental tool for developmental psychorobotics. In Procs. Int. Workshop on Developmental Study, 2002. [Kozima and Yano, 20011 H. Kozima and H. Yano. A robot that learns to communicate withhumancaregivers, 2001. URL citeseer.nj .nec .com/kozimaOlrobot. html. [Lindblom and Ziemke, 20031 J. Lindblom and T. Ziemke. Social Situatedness of Natural and Artificial Intelligence: Vygotsky and Beyond. Adaptive Behavior, vol. 11(2), 7996. 2003. [Lungarella et al., 20041 M. Lungarella et al.. Developmental Robotics: a Survey. Connection Science, vol. 0(0), 1-40,2004. [Maes, 19891 P. Maes. The dynamics of action selection. In Procs. of the 1 lth International Joint Conference on Artificial Intelligence (IJCAI-89), pp. 991-997, 1989. [Maes, 19901 P. Maes. How to do the Right Thing. Connection Science Journal, Special Issue on Hybrid Systems, vol. 1, 1990. URL citeseer.nj .nec .com/ maes89how.html. [Mandler, 19841 G. Mandler. Mind and body: Psychology of emotion and stress. Norton, N.Y., 1984. [Metta et al., 2000al G. Metta et al.. Babybot: A biologically inspired developing robotic agent. In Proceedings of SPIE. Boston, USA, November 2000a. [Metta et al., 2000bl G. Metta et al.. Babybot: an artificial developing robotic agent. In Proceedings ofSAB 2000. Paris, France, September 2000b. [Mobahi, 20031 H. Mobahi. Building un Interactive Robot Face from Scratch. Master's thesis, Azad University, Tehran, Iran, May 2003. [Moreno etal., 20011 F. Moreno et al.. Localization of human faces fusing color segmentation and depth from stereo, 2001. URL citeseer . nj . nec . com/ morenoOllocalization.html. [Nakadai et al., 20001 ~.Ñakadai et al.. Humanoid Active Audition System Improved by the Cover Acoustics. In PRICAI-2000 Topics in Artlficial Intelligence (Sixth Paclfic Rim International Conference on Artlficial Intelligence), vol. 1886 of Lecture Notes in Artljicial Intelligence, pp. 544-554. Springer Verlag, Melbourne, Australia, August 2000. [Nayar, 19981 s.Ñayar. Omnidirectional Vision. In Procs. of the British Machine Hsion Conference. Southampton, UK, 1998. [Oviatt, 20001 S. Oviatt. Taming recognition errors with a multimodal interface. Communications of the ACM, vol. 43(9), 45-5 1, 2000. [Picard, 19971 R. Picard. Affective Computing. MIT Press, Cambridge, MA, 1997. [Raskin, 20001 J. Raskin. The Humane Interface: New Directions for Designing Interactive Systems. Addison-Wesley Professional, 2000. [Rowley et al., 19981 H. Rowley et al.. Neural Network-Based Face Detection. IEEE Trans. on Pattern Analysis andMachine Intelligence, vol. 20(1), 23-38, 1998. [Russell, 19801 J. Russell. A circumplex model of affect. Journal ofPersonality and Social Psychiology, vol. 39(6), 1161-1 178, 1980. [Scassellati, 20001 B. Scassellati. How Robotics and Developmental Psychology Complement Each Other. In NSF/DARPA Workshop on Development and Learning. Lansing, MI, 2000. [Scassellati, 20011 B. Scassellati. Foundations for a Theory ofMind for a HumanoidRobot. Ph.D. thesis, MIT Department of Computer Science and Electrical Engineering, May 2001. [Scholtz and Bahrami, 20031 J. Scholtz and S. Bahrami. Human-Robot Interaction: Development of an Evaluation Methodology for the Bystander Role of Interaction. In Procs. of the Systems, Man, and Cybernetics Conference. Washington DC, 2003. [School of Computing, University of Leeds, 20031 School of Computing, University of Leeds. Cognitive Science Learning Resource. Natural Language Generation, 2003. URL http: //www.comp. leeds .ac .uk/ugadmit/cogsci/ spchlan/nlgen.htm. [Schulte et al., 19991 J. Schulte et al.. Spontaneous short-term interaction with mobile robots in public places. In Procs. of the IEEE Int. Conference on Robotics and Autornation, 1999. [Selfe, 19771 L. Selfe. Nadia: A case of extraordinary drawing ability in children. Academic Press, London, 1977. [Snyder and Mitchell, 19991 A. Snyder and D. J. Mitchell. 1s Integer Arithmetic Fundamental to Mental Processing?: The Mind's Secret Arithmetic. Proceedings of the Roya1 Society oflondon, (266), 587-592, 1999. [Snyder and Thomas, 19971 A. Snyder and M. Thomas. Autistic artists give clues to cognition. Perception, (23), 93-96, 1997. [Snyder et al., 20041 A. Snyder et al.. Concept formation: 'Object' attributes dynamically inhibited from conscious awareness. Journal of Integrative Neuroscience, vol. 3(1), 3 1-46,2004, [Sokolov, 19631 E. Sokolov. Perception and the conditioned repex. MacMillan, N.Y., 1963. [Stanley, 19761 J. Stanley. Computer Simulation of a Model of Habituation. Nature, vol. 261, 146-148, 1976. [Stiefelhagen et al., 20031 R. Stiefelhagen et al.. Capturing Interactions in Meetings with Omnidirectional Cameras. Nice, France, 2003. [Terratec Electronic GmbH, 20031 Terratec Electronic GmbH. EWS88 MT Audio System, 2003. URL http : //www . terratec . com. [Treffert, 19891 D. Treffert. Extraordinary people: Understanding the savant syndrorne. Harper and Row, New York, 1989. [Tyrrell, 19931 T. Tyrrell. Cornputational rnechanisrns for action selection. Ph.D. thesis, Centre for Cognitive Science, University of Edinburgh, 1993. [Wallace, 20051 R. Wallace. Zipf's Law, 2005. URL http : //www . alicebot . org/ articles/wallace/zipf.html. [Walters et al., 20051 M. Walters et al.. Close encounters: Spatial distances between people and a robot of mechanistic appearance. In Proc. IEEE-RAS International Conference on HurnanoidRobots jHurnanoids2005), December 2005. [Wang, 19951 D. Wang. Habituation. In The Handbook of Brain Theory and Neural Networks (edited by M. A. Arbib), pp. 441-444. MIT Press, 1995. [Wang et al., 20041 J. Wang et al.. Face Image Resolution versus Face Recognition Performance Based on Two Global Methods. In Procs. of the Asia Conference on Cornputer Hsion jACCV'2004). Jeju island, Korea, January 2004. [Wenger, 20031 M. Wenger. Noise rejection, the essence of good speech recognition. Tech. rep., Emkay Innovative Products, 2003. [Wilson and Keil, 19991 R. Wilson and F. Keil, eds. The MITEncyclopedia of the Cognitive Sciences. The MIT Press, Cambridge, Massachusetts, 1999. [Winters, 20011 N. Winters. A Holistic Approach to Mobile Robot Navigation using Ornnidirectional Vision. Ph.D. thesis, Instituto Superior Técnico, Universidade Técnica de Lisboa and Trinity College, University of Dublin, October 2001. [Wolfe, 19941 J. Wolfe. Guided search 2.0: A revised model of visual search. Psychonornic Bulletin andReview, vol. 1(2), 202-238, 1994. [Yamato et al., 20041 J. Yamato et al.. Effect of Shared-attention on Human-Robot Communication. In Procs. of the CHI2004 Workshop Shaping Hurnan-Robot Interaction, Understanding the Social Aspects of Intelligent Robotic Products. Vienna, Austria, April 2004. [Yost and Gourevitch, 19871 W. A. Yost and G. Gourevitch. Directional hearing. SpringerVerlag, New York 1987.