Development of crowd simulation models using unity for immerssive VR applications
Abstract
We develop a virtual reality application that simulates a lecture-type talk in a crowded room and an emergency evacuation taking place in it. We perform an experiment where we immerse five participants in our application, and we analyze their behavior.
Full text
Master in Innovation and Research in Informatics Facultat d’Inform` atica de Barcelona (FIB) Universitat Polit` ecnica de Catalunya (UPC) Development of Crowd Simulation Models using Unity for Immerssive VR Applications Master’s Thesis Author: Alicia Nic´as Miquel Director: Julien Pettr´e Co-Director: Nuria Pelechano January 21, 2019
Abstract This project aims to research human behavior in crowds. We use virtual reality in our research as a tool that allows us to perform crowd experiments with great control over the experiment conditions. We develop a virtual reality application that simulates a lecture-type talk in a crowded room and an emergency evacuation taking place in that environment. We perform an experiment where we immerse five participants in our application. We analyze their behavior when they are immersed in both the emergency evacuation and a normal exiting of the room after the talk.
Contents 1 Introduction 6 2 Related work 8 2.1 Human Behavior in Crowds . . . . . . . . . . . . . . . . . . . . . 8 2.1.1 Real World Observations . . . . . . . . . . . . . . . . . . 9 2.1.2 Controlled Experiments . . . . . . . . . . . . . . . . . . . 11 2.2 VirtualReality ............................ 14 2.3 Evacuation Behaviors . . . . . . . . . . . . . . . . . . . . . . . . 16 2.4 CrowdSimulation .......................... 17 2.4.1 RVO.............................. 20 3 Overview 23 4 Development 25 4.1 TheTools............................... 25 4.1.1 The Game Engine . . . . . . . . . . . . . . . . . . . . . . 25 4.1.2 The Experimental Platform . . . . . . . . . . . . . . . . . 27 4.2 TechnicalAspects........................... 29 4.2.1 Sound............................. 29 4.2.2 Lighting............................ 29 4.3 CrowdSimulation .......................... 30 4.3.1 RVO.............................. 31 4.3.2 Pathfinding.......................... 33 4.3.3 FlowControl ......................... 34 4.3.4 Staggered Leaving . . . . . . . . . . . . . . . . . . . . . . 35 4.3.5 Animation........................... 36 4.4 Code.................................. 37 5 Experiments 40 1
6 Results 45 6.1 Physiological data . . . . . . . . . . . . . . . . . . . . . . . . . . 46 6.1.1 Galvanic Skin Response . . . . . . . . . . . . . . . . . . . 46 6.1.2 Heartrate........................... 46 6.1.3 Discussion........................... 46 6.2 SpatialData ............................. 48 6.2.1 Navigation Speed . . . . . . . . . . . . . . . . . . . . . . . 48 6.2.2 Time to Start Moving . . . . . . . . . . . . . . . . . . . . 50 6.2.3 Exit Door Choice . . . . . . . . . . . . . . . . . . . . . . . 51 6.3 Gaze.................................. 51 6.3.1 Angle between Fixations . . . . . . . . . . . . . . . . . . . 52 6.3.2 Gaze vs Trajectory and Camera Rotation . . . . . . . . . 52 6.3.3 Discussion........................... 52 7 Conclusions and Future Work 56 7.1 Conclusions.............................. 56 7.2 FutureWork ............................. 57 2
List of Figures 2.1 Pattern analysis in observation experiments. (a) Trajectories extracted from the observation of a train station [68]. (b) Average walking pattern for groups of typical sizes [45]. . . . . . . . . . . 9 2.2 Collective motions detected in crowd videos in [67] . . . . . . . . 10 2.3 Observed trajectories from [11] (a) trajectories when an S-turn is forced, showing a constraint on the turning radius. (b) Trajectories with no turn often exhibit curvature . . . . . . . . . . . . . . 12 2.4 Collision avoidance experiment from [47] (a) Experimental setup. (b) Picture taken during experiment. (c) tcross is the time when the distance between the walkers is minimal. . . . . . . . . . . . 12 2.5 Snapshot of the experiment from [37]. Note the formation of lanes in the flow through the bottleneck. . . . . . . . . . . . . . . 13 2.6 Goalkeeper facing a virtual thrower in a handball case study from [9].................................... 14 2.7 A burn patient undergoing wound cleaning is immersed in an icy virtual world to distract him from his pain in [30]. . . . . . . . . 14 2.8 Trajectories followed by participants in the virtual environment in [14] wrt the different variables studied (the interpersonal distances of the group, the direction of the movement of the group, and their appearance) . . . . . . . . . . . . . . . . . . . . . . . . 16 2.9 A still from Game of Thrones using agents simulated with the Massive crowd simulation software. . . . . . . . . . . . . . . . . . 18 2.10 Emergent vortex as four groups simulated by [60] cross paths. . . 18 2.11 The path-following and separation rules described in [54] . . . . . 19 2.12 The formation of lanes in a crowded hallway modelled by [27]. . 19 2.13 Reciprocal velocity obstacles as computed in [61] . . . . . . . . . 20 2.14 RVO [61] simulation where 250 agents move to the diametrically opposite positions in a circle. . . . . . . . . . . . . . . . . . . . . 21 3
4.1 (a) Render with direct lighting only. (b) Render with global illumination. Indirect lighting computation is needed for effects such as color bleeding. . . . . . . . . . . . . . . . . . . . . . . . . 26 4.2 A* pathfinding on a regular grid. Colored nodes are the ones that have been explored . . . . . . . . . . . . . . . . . . . . . . . 27 4.3 The bone structure of a humanoid in mecanim. . . . . . . . . . . 28 4.4 Agents walking away from their goal. . . . . . . . . . . . . . . . . 32 4.5 (a) An arch pattern at a doorway. (b) An agent (red) reducing its radius to attempt to squeeze through a crowded area . . . . . 32 4.6 (a) Agents stuck against walls. (b) Invisible walls. . . . . . . . . 33 4.7 Navigation mesh computed for our scene. Different colors represent different areas. The green arrows indicate the agent’s current goal................................... 34 4.8 Lines showing the ray cast to check whether the agent can start to move. Red means can’t move, green means can move. . . . . . 35 4.9 The animation state graph of the agents. . . . . . . . . . . . . . . 36 4.10 The overall structure of the code of our application . . . . . . . . 38 5.1 Equipment .............................. 41 5.2 Experiment training scenes . . . . . . . . . . . . . . . . . . . . . 42 5.3 Initial distribution of the crowd. The red circles indicate the radius of the crowd agents. The blue crosshairs indicates the initial position of the agent. . . . . . . . . . . . . . . . . . . . . . 43 5.4 The robot giving its speech on the stage . . . . . . . . . . . . . . 43 5.5 The crowd at various points in the scenarios . . . . . . . . . . . . 43 6.1 Derivatives of the galvanic skin response over time for the five participants.............................. 47 6.2 Pulse frequency over time for the five participants . . . . . . . . 47 6.3 (a) Speed aggregated over all the participants for all trials. (b-f) Speed per participant for all trials. . . . . . . . . . . . . . . . . . 49 6.4 (a) Angle between successive fixations aggregated over all the participants for all trials. (b-f) Angle changes per participant for alltrials................................. 53 6.5 (a) Difference between gaze and camera orientation aggregated over all the participants for all trials. (b-f) Angle difference per participant for all trials. . . . . . . . . . . . . . . . . . . . . . . . 54 6.6 (a) Difference between gaze and and trajectory orientation aggregated over all the participants for all trials. (b-f) Angle difference per participant for all trials. . . . . . . . . . . . . . . . . . . . . . 55 4
List of Tables 6.1 Table of the time between the first agent moving to leave the room and the participant starting to move . . . . . . . . . . . . . 50 6.2 Door the participant used to exit the room . . . . . . . . . . . . 51 5
Chapter 1 Introduction I carried out this project at INRIA in Rennes. The French National Institute for Research in Computer Science and Automation, INRIA is a public research institute for informatics. The Rainbow team is a joint Inria/IRISA Project-Team in partnership with CNRS, Universit´e de Rennes 1, and Insa of Rennes. Its research is in robotics, targeting applications such as medical robotics, assistive mobility devices, and coordination of multiple robots for spatial navigation tasks. Julien Pettre is a researcher with the Rainbow team working on crowd simulation and the study of crowd behaviors. Simulated crowds are required in a wide variety of application areas, from videogames and movies to architectural simulations. Achieving realism in crowd simulations requires in-depth knowledge about the behavior of humans in crowds. The development of crowd simulation algorithms at the local interaction level requires extending knowledge on individual behaviors in crowds and their repercussion at the collective scales. Knowledge about the behavior of crowds may be acquired by analyzing observations of real-world crowds or by doing controlled experiments. Hovever, doing crowd experiments with real crowds at even a modest level of complexity is hard to do with real-life crowds, due to problems with standardizing conditions and ethics problems with putting people in certain situations. Virtual reality allows us to set real humans in interaction with their digital counterparts. Doing so, we can study more detailed situations of local interactions, perfectly control our experimental situations, and more easily acquire experimental data. This project is part of a longer project aiming to study the behavior of people in a crowd during stressful situations. 6
In the frame of this masters’ thesis, we develop a virtual reality application that simulates a crowded environment and a stressful event taking place in that environment. We run a small pilot experiment where we immerse five participants in our application. We analyze the data from our experiment in order to validate the usefulness of our application in the larger project. 7
2.2 Virtual Reality Steuer [59] defines virtual reality as ”a real or simulated environment in which a perceiver experiences telepresence”. The term telepresence, or presence, refers in the virtual reality community to ”the subjective experience of being in one place or environment, even when one is physically situated in another” [64][58]. The sense of presence is often evaluated in VR studies using the subjective self-assesment of the user [64][56]. A more objective way of evaluating presence is through measurements of the user’s physiological response to events happening in the virtual environment, such as the heart rate response to a stressful situation [41]. The sense of presence achieved by an application has been shown to influence the user’s behavior in experiments. Various studies show [57][24] that the sense of presence in the application correlates positively with the subject behaving more like they would in the real world. Figure 2.6: Goalkeeper facing a virtual thrower in a handball case study from [9]. Figure 2.7: A burn patient undergoing wound cleaning is immersed in an icy virtual world to distract him from his pain in [30]. The best-known applications of virtual reality are in entertainment, but the properties of presence have enabled VR to be used in the research of human behaviors [9][39][44] (see Figure 2.6 for an example experiment), or as a healthcare 14
tool [29][36][30] (See Figure 2.7 for an example application). If we could use virtual reality in our crowd behavior experiments, instead of real-world controlled experiments, it would have some big advantages. First, in a virtual reality experiment, we have full control over the stimuli that we expose the users to. We can make the stimuli interactive and responsive to the user’s behavior. We can perform repetitions of the same experiment conditions with a high degree of accuracy. We can also more easily control and isolate the factors that influence our experiment. Finally, virtual reality experiments allow us to study the behavior of the user in the midst of large crowds without the inherent difficulty of designing and coordinating an experiment with a real-world crowd [13]. However, for us to be able to use virtual reality to study human behaviors, we need to know if people behave in virtual environments like they do in reality. There are in fact several studies that show differences in the subject’s perception of their environment in VR and the real world. Banton et al. [8] find that users feel that the virtual environment is moving slower than it actually is when they’re walking on a treadmill while immersed in the virtual environment. G´erin-Lajoie et al. [25] find that the size of the personal space bubble of a person is slightly increased in the virtual environment, while its general shape is preserved. Hollman et al. [31] find that walking on a treadmill while immersed in a virtual environment caused instability in the walk of participants. Several studies also find that the perception of distances is compressed in the virtual environment [34][63][53]. Mohler et al. [43] however find that this distance underestimation can be reduced after the participant spends 5 minutes walking to targets in a virtual hallway with continuous visual feedback. On the other hand, there have been studies that show that, while there are quantitative differences in behavior between real-world experiments and their virtual reality counterparts, the qualitative characteristics of the behavior does remain. [20] compare the obstacle-avoidance trajectories of subjects in the real world and in the virtual environment, and find quantitative differences but qualitatively similar trajectories. [14] is a study that uses virtual reality for the study of crowd behaviors. They use VR to study how people avoid a large incoming group of people (do they snake through it, or do they go around it) depending on the group’s size. The trajectories they observe can be seen in Figure 2.8. Considering the difficulty and limitations of real-world experimentation and the limitations of learning from uncontrolled observations, we consider that crowd experiments in VR are a useful alternative that can be used to study complex crowd behaviors. The perceptual differences that users experience in 15
Figure 2.8: Trajectories followed by participants in the virtual environment in [14] wrt the different variables studied (the interpersonal distances of the group, the direction of the movement of the group, and their appearance) VR experiments should however be carefully accounted for. The need to ensure that presence is achieved in the application if we want the user’s behavior to be similar to their real-world behavior should also be considered. 2.3 Evacuation Behaviors We use the word ’stress’ to refer to a state of physical of psychical tension that result from factors —stressors— that tend to alter an existent equilibrium [42]. Stress is a wide concept, encompassing acute and chronic stress. There doesn’t exist a full consensus on its definition [33]. Some definitions of stress are: ”A particular relationship between the person and the environment that is appraised by the person as taxing or exceeding his or her resources and endan- 16
gering his or her well-being.” [22] ”A substantial imbalance between demand (physical or psychological) and response capability, under conditions where failure to meet the demand has important consequences.” [40] In our project, we consider only acute stress due to an active emergency situation, such as a fire or an earthquake. These kinds of catastrophes necessitating an emergency evacuation of large numbers of people happen regularly around the world. In studying these evacuations, we regularly find inappropriate behaviors on the part of the evacuees. People may take a while to start evacuating, they may not use the closest exits, or they may push on the people around them, which is counterproductive, as it reduces the throughput of exits, and may make people fall and be trampled by the crowd [12]. It’s important to study, and be able to accurately reproduce, the behaviors of people facing an evacuation. Knowing more about these behaviors can help us design buildings that are better prepared to meet the building occupants’ needs in an emergency evacuation [35]. During a fire, someone’s psychological stress levels may rise because their capacity for processing information is exceeded [52], or they are confronted with an unfamiliar situation [35]. Too much psychic stress can impair cognitive processes and how an individual responds to a given situation [52]. On the other hand, the commonly held idea that there’s panic —defined as irrational behavior that may be counter-productive, like pushing people to get through exits— has been found to not hold up. Multiple studies have found no proof of the presence of panic in major disasters [32][26]. In fact, [32] finds that what outside observers perceive from the outside as irrational panic behaviors are described by the person involved in terms of rational decisions, and, additionally, that altruistic behavior is the norm. 2.4 Crowd Simulation Simulated crowds are required by a wide range of applications. See Figure 2.9 for an example application in TV filming. These applications don’t all have the same requirements for the properties of the simulations. Videogames and movies require their crowd simulations to be believable while also allowing high control over them. Videogames and interactive experiences additionally require real-time performance and interactable agents. In applications related to event security or urban planning, for example, the main concern is the statistical accuracy of the simulation and its ability to predict the outcomes of the simulated 17
Figure 2.9: A still from Game of Thrones using agents simulated with the Massive crowd simulation software. situation [65]. The range of models and algorithms of crowd simulation is, accordingly, wide and varied. To a first approximation, they can be divided into two main approaches: microscopic and macroscopic algorithms [13]. Figure 2.10: Emergent vortex as four groups simulated by [60] cross paths. Macroscopic algorithms model the behavior of the crowd as a whole, with models being derived from fluid dynamics [55][60] (see Figure 2.10 for some of their results). This kind of algorithms is useful for the simulation of very dense crowds with a small computational cost. However, some issues arise because these algorithms model the agents as particles carried by the flow. The agents can have erratic trajectories and the individual agents can’t have the kind of complex interactions that can be modeled with microscopic approaches. This type of algorithm therefore breaks down in the simulation of less dense crowds. 18
Microscopic algorithms consider each individual member of the crowd independently. They model the behaviors of the individuals, and model a crowd by simulating multiple independent individuals. Microscopic algorithms focus on modeling local interactions. They model the interactions of the agent with the elements of its environment —mainly obstacles and other agents. These interactions can be of many types. The most common and therefore most studied interaction is collision avoidance. (a) Path-following rule (b) Separation rule Figure 2.11: The path-following and separation rules described in [54] One of the first crowd simulation models, that due to Reynolds [54], models interactions as rules, such as those shown in Figure 2.11. It works by giving each agent a steering vector that determines its direction and speed of movement. This steering vector is modulated by the application of ”steering behaviors” that result from the evaluation of the rules of the model. For example, if the agent is too close to a nearby agent, a steering force will be added to its steering vector that steers it in the direction away from the nearby agent. Figure 2.12: The formation of lanes in a crowded hallway modelled by [27]. The social forces model [27] models individuals as particles subjected to various forces. The motion of the agents is then defined by a sum of forces, where each of the forces corresponds to an interaction. The model can reproduce emergent patterns such as the lane formation shown in Figure 2.12. These are position-based algorithms, where the next position is computed by adding forces to a steering vector. They are mainly reactive, they only 19
modify the behavior of the agents when there’s a risk of collision. Another set of microscopic algorithms try to predict future risk of collision and base the behavior of the agents on it. Velocity-based algorithms, such as RVO [61] compute the agent’s next velocity by adding geometric constraints to the space of possible velocities. They then choose the best velocity, according to some metric, among the remaining ones. In Section 2.4.1, we describe RVO, the velocity-based microscopic algorithm that we use in our crowd simulation. 2.4.1 RVO (a) Velocity obstacle (V O) (b) Reciprocal velocity obstacle (RV O) (c) Union of all the reciprocal velocity obstacles the agent will avoid Figure 2.13: Reciprocal velocity obstacles as computed in [61] In our crowd simulation, we use RVO to model the collision avoidance behavior of the agents in the crowd. Reciprocal Velocity Obstacles or RVO [61] is a microscopic, velocity-based, collision avoidance algorithm. For every simulation step, RVO computes the ve- 20
locity of an agent in the next simulation step by applying geometric constraints to the space of possible 2D velocities. The set of discarded velocities is known as the agent’s velocity obstacle. Once the constraints are applied, it chooses the new velocity as the remaining velocity closest to the agent’s preferred velocity. If no permissible velocities remain, a velocity inside the velocity obstacle that has a high time to collision may be chosen. RVO considers the velocities of surrounding agents when computing the velocity constraints. This allows the algorithm to act on the future risk of collision and better handle collision avoidance with agents or obstacles moving at high speed. RVO is based on an earlier algorithm, called Velocity Obstacles [21]. As an example, let’s consider the computation by the VO algorithm of the velocity obstacle of an agent Ainduced by the presence of an agent B. Figure 2.13 shows a graphical representation of this computation. VO computes the velocity obstacle V OB A(vb) as the set of velocities vathat will lead to a collision with the agent Bmoving at a velocity vbat some point in time. The velocity obstacle is computed in the VO algorithm as the set of velocities vasuch that the relative velocity between the two agents va−vbintersects the Minkowski sum of Band −A. Figure 2.14: RVO [61] simulation where 250 agents move to the diametrically opposite positions in a circle. RVO is an extension of VO that is ”reciprocal” because each agent assumes that other agents will try to avoid it as much at it will try to avoid them, and therefore does only half of the necessary collision avoidance. This removes the oscillating velocity artifacts caused by the VO algorithm. Figure 2.14 shows some results obtained with RVO. There are some situations where RVO leads to artifacts. In particular, in dense situations, several agents may get stuck in a symmetrical pattern, where none of them is able to get the right of way. [BRaVO: Biased Reciprocal Velocity 21
Obstacles Break Symmetry in Dense Robot Populations] 22
Chapter 3 Overview This project is part of a longer project. The ultimate goal of the larger project is to study the behavior of people in a crowd during stressful situations. For the reasons explained in Section 2.2, we consider that virtual reality is a useful tool to reach that goal. In the frame of this masters’ thesis, we do two main tasks. First, we develop a virtual reality application. The aim of this task is to create the virtual environment where we will later immerse the participants in our experiments. The application must then simulate a crowded environment and a stressful event taking place in that environment. We choose to simulate a lecture-type talk (we use a TED Talk as the talk audio) in a small venue with a crowd of 100 people attending. We script and develop two possible endings to the talk. In one of them —our baseline scenario— the talk ends normally and the crowd leaves the room calmly. In the second scenario —our emergency scenario— the talk is interrupted by an explosion and the subsequent emergency evacuation of the venue. In Chapter 4 we describe the development process of the application. In Section 4.1 we describe the pre-existing tools we built our application on top of. Section 4.2 describes some technical aspects of the development of our application qua game-like application. Section 4.3 describes some challenges we encountered regarding crowd simulation, and the solutions we arrived at.Finally, Section 4.4 talks about the structure of the code developed for the application. In the second part of the project, we do some preliminary validation of the application we built. Our goal with this validation is to figure out if it’s suited to our research goals. Since we want to study stressful situations, we want to know if the application (specifically, the emergency scenario) is able to induce stress on the user. Because we want to study behavior, we want to check whether our 23
most frames. Some requirements due to the nature of our scene create constraints on the lighting that preclude using less expensive solutions. We need the lights to change when the performance ends and when the explosion happens. This makes baked GI not a possibility for the main lights, which means that we need to use realtime lights for the main lights in our scene. We use Unity’s precomputed realtime GI to achieve greater realism than regular direct lighting. Dynamic objects (i.e. the crowd) are lit using light probes. Light probes are sample points situated in the scene that store the light hitting that point from all directions. The lights that light the rooms outside the main room are baked lights. Mixing precomputed realtime GI and baked lights adds some overhead, as does adding more realtime lights to a scene. We don’t see a performance difference between the two approaches for the outside lights. Another issue has to do with per-pixel vs. per-vertex lights. Per-pixel lights compute lighting at each lit pixel, while per-vertex lights do it only at each vertex. Per-pixel lights cause a big performance overhead and impede certain optimizations. In our scene, the lights on the stage should be rendered per-pixel because they are simulating spotlights on a performer. To reduce the number of lights on the stage to two, while maintaining the illusion of a multi-light stage lighting rig, we use a single spotlight with a cookie (a texture added to a light to change its shape) to fake a multi-spotlight wash. However, two per-pixel lights are still somewhat problematic, especially because they illuminate the crowd too. We don’t try to optimize this further, because the computer we use to run our experiment is powerful enough. The use of shadows would be a big improvement on the realism of the scene, but they imply a big performance overhead. The appearance of the emergency exit signs glowing in the dark room is simulated with the bloom post-processing effect. 4.3 Crowd Simulation In this section, we talk about some aspects of the simulated crowd in our application. Our simulated crowd needs to look like a crowd at a performance. During the performance, they should stay in place and look at the stage. When the performance ends, they should applaud, and then leave in an ordered way that looks as realistic as possible. In our emergency scenario, when the explosion happens, the crowd should look startled and scared, and then leave in a hurry. 30
Our crowd simulation is a microscopic crowd simulation, i.e. each agent is simulated independently, and a crowd is made up of a group of independent agents. The locomotion of each agent of our crowd is made up of two parts: a control law, which handles pathfinding, and a simulation, which handles local movement (in this case, the only local movement behavior we handle is collision avoidance). Each agent additionally has animations. The animations are controlled using Unity’s system of humanoid animation. We use RVO as a local movement algorithm. An overview of RVO can be found in Section 2.4.1. During development, we encountered some situations in which the way that RVO does collision avoidance led to unrealistic behaviors. These issues and the solutions we found are described in Section 4.3.1. Section 4.3.2 describes the way we handle pathfinding in our simulation as well as some issues we had with the exit doors as crowd flow bottlenecks and how we solved them. Section 4.3.5 describes how the animation of the crowd agents adapts to their movement. 4.3.1 RVO In order to use RVO to create a realistic crowd simulation in our application, we needed to address some issues with the algorithm. Because of the geometric nature of the algorithm, agents sometimes get stuck in a symmetric pattern where all of them yield to the others and nobody moves. In our scene, this happens at doors, where the agents get stuck in an arch pattern (see Figure 4.5a). When many agents want to go through a door at once, some of them start walking away from the door (see Figure 4.4), since that’s the best way to avoid collision. In an evacuation, this looks unrealistic, since what we’d expect is queueing behaviors. In the next paragraphs, we describe the tweaks we did to RVO2 to prevent these behaviors. Forbid walking backwards by adding an ORCA line that removes velocities with a backward component from the possibilities. In spite of this fix, the agent may still end up with a small backwards vector because of RVO2 needing to relax some constraints or because of numerical issues. To deal with this, when the simulation asks the agent to move in a direction away from its goal with a very small speed, the agent does move but keeps its previous rotation. 31
Figure 4.4: Agents walking away from their goal. Add a random tie-breaking component. Each agent is given a random value that represents how willing they are to squeeze through in a tie situation. (a) (b) Figure 4.5: (a) An arch pattern at a doorway. (b) An agent (red) reducing its radius to attempt to squeeze through a crowded area When an agent computes its best feasible velocity, it uses its radius and its neighbors’ radii. Our modification allows the agent to consider its own radius to be smaller in tie conditions. This simulates the behavior that can be observed in a crowded area —such as a crowded subway— where people may try to squeeze through the crowd to be able to move. 32
To detect a tie condition, the agent checks the time tstuck since it’s moved (only if its preferred speed is larger than zero) and reduces the radius it considers itself to have once tstuck reaches a threshold tsqueeze.tsqueeze is an attribute set randomly for each agent and therefore acts as a tie-breaker. Figure 4.5b shows the results. Agents get stuck at walls. When agents simulated with RVO2 encounter an obstacle that runs perpendicular to their preferred velocity, they may be unable to overcome it. This is because the vector closest to their preferred velocity in this case is the null vector. In our scene, agents would get stuck at walls when they tried to exit the room (see Figure 4.6a). (a) (b) Figure 4.6: (a) Agents stuck against walls. (b) Invisible walls. We address this limitation by adding invisible walls to the scene that add a bevel to the inside of the jambs of doorways, as shown in Figure 4.6b. This breaks the perpendicularity of the wall and the agent’s preferred velocity and makes it so that the agent can find a valid velocity at which to move on. Using smarter pathfinding also helped agents not get cornered against walls in the first place. 4.3.2 Pathfinding Our initial pathfinding solution was a basic control law which aimed the agent towards specific waypoints. Setting a waypoint at a doorway and another outside the room gave us reasonable behavior. However, this, as can be expected, led to a few artifacts. We set a tolerance to the waypoints, such that agents could move on to the next waypoint even if they couldn’t reach the exact way- 33
point position. However, when large jams happened at one of the doorway, some agent would exit through an unclogged door, then try to re-enter through its assigned doorway. Aside from unrealistic movements, this caused agents to meet head-on while trying to move in opposite directions, which is a situation that may cause issues with RVO2. Figure 4.7: Navigation mesh computed for our scene. Different colors represent different areas. The green arrows indicate the agent’s current goal. In order to fix this we switched to Unity’s implementation of the navigation mesh algorithm for pathfinding. Unity’s NavMesh navigation system allows you to create characters that can intelligently move around the game world, using navigation meshes that are created automatically from your Scene geometry. It allows baking a NavMesh, which we use to create the general navigation areas, and it also allows dynamic obstacles, that can move at runtime. We use dynamic obstacles in the emergency evacuation scene, to have casualties lying on the ground, which aren’t agents or static obstacles, since they have animations that may cause them to move around a bit. The NavMesh also offers the possibility to set a per-agent mask to block some walkable areas, which we use to solve the problem of uneven flow in the next section. Figure 4.7 shows our navigation mesh and the distribution of walkable areas we use. 4.3.3 Flow Control Another problem we encountered, is that static ways of assigning an exit door to an agent ended up with large accumulations of people trying to go through a doorway while other doorways were empty. While this is a behavior that is found sometimes in crowds, it led to unrealistic locomotion in the agents waiting 34
to exit. We fixed this by setting up a system of subscription to doors. Each agent, when it wants to leave the building, has to find a door with a sufficiently low number of subscribers, and can only exit when it finds one. When the agent has exited the building, it unsubscribes from its door. When a door has a low subscriber count, it sends an event, that all the agents are listening to. The listening agents may then choose to switch their subscription to the lowsubscribers door, if its current door has a high amount of subscribers and it isn’t too close to leaving already. To make this work with the NavMesh, we setup different NavMesh areas on each doorway, and block the agent (with the walkable area mask) from walking on areas that contain doorways it isn’t subscribed to. 4.3.4 Staggered Leaving The agents in a crowd don’t leave all at the same time. Figure 4.8: Lines showing the ray cast to check whether the agent can start to move. Red means can’t move, green means can move. They cast a ray between their position and the door they want to exit through and compute a score based on the number of obstacles that the ray intersects, their distance to the agent, and the amount of time that they’ve been waiting to exit. Given a set of ray intersection distances r={r1...rn}, and the distance to the doorway dthe leaving score at time tsince the agent starting trying to leave the room is computed, heuristically, as: 35
leavingScore = 5 −X ri∈r (2 r2 i )−min(0.1d, 1) + min(0.2t, 1) Figure 4.8 shows the resulting value for a variety of situations. 4.3.5 Animation Unity offers an animation system called Mecanim. The Mecanim system is a powerful way of setting up animation of characters in games. It offers features like humanoid animation retargetting and the ability to set up the animation of a character with a visual state machine editor that can be extended with C# code to create complex transitions and behaviors. Figure 4.9 shows the state machine for our crowd agents. An issue I encountered often when working with Mecanim, on the other hand, is the lack of documentation of some parts of the Mecanim scripting API. Figure 4.9: The animation state graph of the agents. Walking and Running Unity offers the possibility of blending several animation clips into one animation via a feature called blend trees. Blend trees are a good way to set up walking and running animations that change according to the angular velocity at which the agent is moving. We have animations of people walking forward and walking in wide and narrow circles, right and left. Using a blend tree allows us to create a continuum of animations between our existing animations by linearly blending them and 36
using some parameter set programatically to set the blending weight. The blending weight in our walking and running blend trees is set as the angular velocity of the agents. The agents choose to use walking or running animations depending on their speed. This could also be set as a blend tree parameter, creating a 2D blend tree with axes speed and angular velocity, but the results weren’t visually convincing. Shuffling When their speed is low enough, normal walking animation slowed down starts looking unrealistic. Instead, they use an animation where they do one step, then stand for a bit, then do another step. Turning Around Sometimes, the agents have to turn around in place. We use animations where the agents turn around in place. We wanted to keep the animation behavior separate from the locomotion behavior. The turning around animation is triggered when the agent turns around (its forward vector changes direction by more than 45ofrom the previous frame) while its speed is zero. When this happens, the orientation of the root bone of the agent is decoupled from the orientation of the agent object (which is used for locomotion computations) such that the agent appears to continue to look in the same direction. Then, the root bone is gradually rotated until it matches the agent object orientation, while the turn in place animation plays. 4.4 Code Figure 4.10 shows a diagram of the overall code structure of the project. The code is divided into two main blocks, coordinated by the IndoorSceneTrial class, which implements the TrialManager class from CrowdMP. One of the blocks, which we’ll call the Crowd block, handles the behavior of the crowd agents. The second block, the Scenario block, handles the ”script” of our two scenarios (baseline and emergency). MultiControlLawAgent is the main class in the Crowd block. It behaves as a state machine. IndoorSceneTrial communicates the changes in the scenario state to the MultiControlLawAgent. The state of MultiControlLawAgent mainly follows the scenario state changes with some delays to avoid the unrealistic situation where everyone in the crowd does something at the exact same time. MultiControlLawAgent implements the Agent class from CrowdMP. A CrowdMP Agent has a ControlLaw component (i.e. a pathfinding algorithm) that it queries at each frame to find its next position and rotation. The position and 37
Figure 4.10: The overall structure of the code of our application 38
rotation given by the Control Law can be overriden by the control simulation (in our case RVO) to handle local movement. MultiControlLawAgent extends Agent by allowing an agent to have multiple control laws. This is needed so that our agents can start the simulation standing in the same spot to watch the performance, and then start to move towards a goal (the outside). Each agent object has an AnimationController component that manages the animation of the agent independently of its locomotion. At each frame, the AnimationController computes the velocity and angular velocity of the agent, plus gets some state information from MultiControlLawAgent (e.g. is the agent applauding?) and updates the animation accordingly. The Scenario part of the code is tasked with coordinating what we could call the ”script” of the scene. The ScenarioManager class receives from IndoorSceneTrial (which in turn is interpreting the instructions from the trial configuration file) the signal to trigger an event (which can be the normal end of the performance or the emergency ending). The ScenarioManager class then coordinates the behavior of the elements of the scene not related to the crowd (the sounds, the lighting, the robot performer, and the casualties in the emergency scenario). After the inputs from TrialManager have triggered an ending ”script”, the ScenarioManager coordinates the subsequent stages of the script using a timer and constants coded into the class that indicate how long each stage should last. IndoorSceneTrial also reads the scenario state from ScenarioManager and passes it on to the MultiControlLawAgent instances, which use it as part of the input for their own state transitions. 39
6.1 Physiological data We want to be able to study the behavior of humans in crowds in situations of stress. In order to do this, our application should be able to cause a stress response in the experimental subjects. We used physiological sensors to collect the users’ galvanic skin response and heart rate during their participation in the experiment. In this section, we present and analyze the collected data. In Section 6.1.1, we present our skin conductivity results. In Section 6.1.2, we present our heart rate results. In Section 6.1.3 we analyze our results and discuss our conclusions. 6.1.1 Galvanic Skin Response The theory of GSR holds that the conductance of human skin increases with the activity of the sweat glands in it. Sweating increases in situations of fear and stress. Therefore, we can use measurements of the conductance of a subject’s skin as a measure of their stress or psychological arousal [10]. In our experiment, we expect to see a rise in the GSR of the subject when the alarm sounds, and no such rise in the baseline trials. If this happens, it means that our emergency scenario can cause a physiological reaction in the experimental subjects. As seen in Figure 6.1, the GSR value of all but one of the participants rose sharply in the emergency trial when the explosion occurred. 6.1.2 Heart rate Heart rate increases with stress and fear [46]. We can then use a measure of heart rate to check whether our experimental subjects experience a stress response to the emergency situation we expose them to in our experiment. As seen in Figure 6.2, we find that the heart rate of all the participants stays mostly stable during all the trials. Unlike what we expected, we don’t find an increase of the heart rate in our emergency situation. 6.1.3 Discussion In our GSR data, our results mostly coincide with what we expected. Four out of five participants experience a peak in GSR when the alarm sounds. The measurement additionally stays mostly constant for participants one and two in the baseline trials. However, for one of the participants, there isn’t a peak in the emergency-scenario GSR, only a continuation of a constant rising trend 46
(a) (b) (c) (d) (e) Figure 6.1: Derivatives of the galvanic skin response over time for the five participants (a) (b) (c) (d) (e) Figure 6.2: Pulse frequency over time for the five participants that starts before the alarm. Participants 3 and 5 also have additional peaks in the baseline and emergency trials, which makes it harder to be sure that the alarm scenario peak is caused by the emergency. 47
Our pulse frequency results, on the other hand, don’t show a rise in heart rate for any of the participants. This contrasts with the GSR results, which do seem to show a stress response. This may indicate that our application doesn’t induce stress in the participants. In the feedback given after their participation, some of the participants indicated that they hadn’t felt scared because they had a good view of the exits, and the evacuation was fluid, without any clogging at the doors. Only participant 4 reported feeling scared by the emergency situation, which contrasts with their GSR measurements. Additionally, participant 4 reported developing motion sickness during immersion, which may alter their results, since motion sickness is correlated with lower feeling of presence. The inconsistent results may be due to human variability in stress responses. They also may be a sign, especially given the heart rate results, that our application doesn’t cause enough stress in the participant for the signal to be clear. It may also be the case that the participants only feel startled by the unexpected explosion, but don’t feel the situation as stressful, maybe because they lack the feeling of presence, or because the emergency doesn’t feel dangerous enough. 6.2 Spatial Data If the participant in a VR experiment is immersed with a feeling of presence in the virtual scene, we expect them to behave —qualitatively— like a person would behave in a real-world situation. In the situation of the normal ending of a performance, we expect to see social behaviors like queueing and maintaning one’s personal space, avoiding physical contact with others [48]. In an emergency situation, we expect to see those social behaviors preserved, since our scenario doesn’t simulate the kind of emergency with a clear and immediate danger (like an out-of control fire inside the room) that could lead to panic-like behaviors [18]. In this section, do some exploratory analysis the locomotion patterns of the participants and the choices they make when exiting the room. In Section 6.2.1, we analyze the speed at which the participants move when exiting the room. In Section 6.2.2 we look at when the participants start moving with relation to the timeline of events of the scenario. In Section 6.2.3 we look at the door the participants choose to exit the room through. 6.2.1 Navigation Speed We analyze the speed at which the user moves when exiting the building. We expect to find that this speed increases in the emergency situation. We find 48
that the speed does seem higher, as seen in Figure 6.3. The average speeds for the four trials are 0.29m/s, 0.32m/s, 0.34m/s, and 0.74m/s, where the last one corresponds to the emergency trial. We consider the speeds starting when the crowd begins to move. Discussion A higher navigation speed doesn’t neessarily imply a higher arousal. We know that individuals in a crowd tend to follow the actions of the people around them [48]. Even if the subject doesn’t experience the emergency situation as stressful and dangerous, they may move at an increased speed to match the people around them. (a) (b) (c) (d) (e) (f) Figure 6.3: (a) Speed aggregated over all the participants for all trials. (b-f) Speed per participant for all trials. 49
∆t1∆t2∆t3∆t4 Participant 1 12.86 7.13 3.06 3.42 Participant 2 1.33 1.77 3.23 -9.01 Participant 3 1.53 5.50 10.61 2.25 Participant 4 7.83 3.00 1.98 8.58 Participant 5 2.93 4.74 5.25 -2.82 Mean time 5.30 4.43 4.82 0.48 std of time 4.45 1.88 3.08 5.97 Table 6.1: Table of the time between the first agent moving to leave the room and the participant starting to move 6.2.2 Time to Start Moving We analyze the time that it takes for the subjects to start moving, in relation to the time the rest of the crowd starts moving. In the baseline scenario, the crowd starts moving after they have finished applauding, at the same time that the lighting of the room changes from dark performance lighting to bright end-of-show lighting. This happens 33 seconds into the scenario. The end of the speech happens in second 25 of the scenario. In the emergency scenario, when the explosion sounds the performance lights go out, and flicker back on after a few seconds. Meanwhile, an alarm starts sounding, and the sounds of sirens and screams start to sound outside the room. the crowd starts to run to evacuate after the lights come back on, on second 27. The explosion happens 18 seconds into the scenario. As shown in Table 6.1, the experimental subjects generally started to move a few seconds after the crowd started to move. Nobody started to leave while the crowd was applauding. In the emergency scenario, some of the participants started moving before the crowd started moving. Discussion Participants moving before the crowd in the emergency scenario may indicate that they were anxious to evacuate the room. On the other hand, some participants reported being confused about whether the explosion and the sudden darkness were intended. This may have prompted them to test the application for a response. With the two early movers removed, the average time to leave is similar to those in the other trials (4.75). This is expected, since previous studies have shown that, even if a fire alarm is sounding, people tend to wait and see how the people around them respond, rather than be the first to take action [52]. 50
Baseline 1 Baseline 2 Baseline 3 Emergency Participant 1 1 None 1 3 Participant 2 4 4 2 1 Participant 3 1 4 4 1 Participant 4 3 2 3 2 Participant 5 4 4 4 3 Table 6.2: Door the participant used to exit the room 6.2.3 Exit Door Choice As shown in Table 6.2, our subjects chose a door they hadn’t previously used in all the emergency situations, and there’s no obvious pattern to their door choices. Participant 1 exited through the wall in the Baseline 2 trial. Discussion Previous studies have shown that people usually choose to exit a building through an exit that is familiar to them, such as the one they used to enter the building. If an emergency exit is usually barred from circulation, it’s usually ignored in an emergency [49]. Since the subjects experience first the baseline non-emergency situation, and have time to familiarize themselves with the exits, we expect that they will tend to choose the same door in all the experiments, or at least choose a familiar exit in the emergency situation. Our participants didn’t conform to this pattern. This may be because all exits are close together in the room and clearly visible, making them interchangeable. The tendency to keep to familiar routes tends to increase in an emergency [52]. That the users chose a new exit in the emergency scenario may also be a sign that they weren’t feeling stressed by the emergency. 6.3 Gaze Using the FOVE headset allows us to track the gaze direction of the experimental subject. The ultimate goal of building our scenario, which falls beyond the scope of this project, is to use analysis of gaze data to study the behavior of individuals in dense crowds. In this section, we analyze our preliminary data to see whether we can find differences in the gaze behavior between the baseline scenario and the emergency scenario. In Section 6.3.1, we look at the difference between gaze fixations. In Section 6.3.2, we compare the direction of the subject’s gaze to their trajectory and the 51
orientation of their body. In Section 6.3.3 we discuss our findings. 6.3.1 Angle between Fixations We use MATLAB scripts to analyze the gaze data of the subjects and group the raw gaze data into a series of fixations on different points. The code for the computation of the fixations was developed by Florian Berton. Following [38], we consider that a temporal series of gaze directions constitute a fixation when, for a temporal window of 80ms or more, all the gaze directions are within a range of 3.0 degrees from the initial direction. We analyze the horizontal angles between each two successive fixations to measure how much the subject is looking around. We analyze the data after the speech has ended or the explosion has occurred, since that’s when the scenarios start to differ. As shown in Figure 6.4, we don’t find a visible difference in the angles between fixations in the two scenarios. 6.3.2 Gaze vs Trajectory and Camera Rotation In Section 6.3.1 we compare the directions of the subject’s gaze among themselves. Here, we compare the direction of the gaze with both the trajectory of the navigation of the subject and the rotation of the camera, i.e. the orientation in space of the virtual body of the subject. Our aim here is also to detect a difference in gaze behavior between our two scenarios. As shown in Figures 6.5 and 6.6, we don’t find a visible difference in the angles between the gaze and the camera angle or the trajectory in the two scenarios. 6.3.3 Discussion We didn’t find a difference in the amount people looked around in our two scenarios by doing our analysis. As far as we know, there’s no scientific literature on the subject of gaze behavior in emergency situations. We therefore didn’t know what we should expect for these results. If there’s a difference to be found, a reason that we may not have found it is our chosen method of navigation. Our users used a joystick for navigation. This is an unnatural way to represent locomotion, and may influence how much the person looks around, e.g. depending on the relative difficulty of turning around with a joystick vs. turning around while walking vs. looking around. 52
(a) (b) (c) (d) (e) (f) Figure 6.4: (a) Angle between successive fixations aggregated over all the participants for all trials. (b-f) Angle changes per participant for all trials. It’s also possible that the difference in gaze behavior is not in how much the person looks around, but what the person looks at. We didn’t collect data on gaze allocation, which requires non-trivial setup to do. 53
(a) (b) (c) (d) (e) (f) Figure 6.5: (a) Difference between gaze and camera orientation aggregated over all the participants for all trials. (b-f) Angle difference per participant for all trials. 54
(a) (b) (c) (d) (e) (f) Figure 6.6: (a) Difference between gaze and and trajectory orientation aggregated over all the participants for all trials. (b-f) Angle difference per participant for all trials. 55
[46] V Mustonen and M Pantzar. Tracking social rhythms of the heart: from dataism to art. 2013. [47] Anne H´el`ene Olivier, Antoine Marin, Armel Cr´etual, and Julien Pettr´e. Minimal predicted distance: A common metric for collision avoidance during pairwise interactions between walkers. Gait and Posture, 36(3):399–404, 2012. [48] Xiaoshan Pan, Charles S. Han, Ken Dauber, and Kincho H. Law. A multiagent based framework for the simulation of human and social behaviors during emergency evacuations. AI & SOCIETY, 22(2):113–132, oct 2007. [49] Nuria Pelechano and Norman Badler. Modeling Crowd and Trained Leader Behavior during Building Evacuation. IEEE Computer Graphics and Applications, 26(6):80–86, nov 2006. [50] Christopher Peters and Cathy Ennis. Modeling groups of plausible virtual pedestrians. IEEE Computer Graphics and Applications, 29(4):54–63, 2009. [51] Julien Pettr´e, Jan Ondˇrej, Anne-H´el`ene Olivier, Armel Cretual, and St´ephane Donikian. Experiment-based modeling, simulation and validation of interactions between virtual walkers. In Proceedings of the 2009 ACM SIGGRAPH/Eurographics Symposium on Computer Animation - SCA ’09, page 189, 2009. [52] G. Proulx. A Stress Modell for People Facing Fire. Journal of Environmental Psychology, 13:137–147, 1993. [53] Rebekka S. Renner, Boris M. Velichkovsky, and Jens R. Helmert. The perception of egocentric distances in virtual environments - A review. ACM Computing Surveys, 46(2):1–40, 2013. [54] Craig W Reynolds. Steering Behaviors For Autonomous Characters. In Proceedings of Game Developers Conference, 1999. [55] Roger L. Huges. The Flow of Human Crowds. Annual Review of Fluid Mechanics, 35:169–82, 2003. [56] Martijn J. Schuemie, Peter van der Straaten, Merel Krijn, and Charles A.P.G. van der Mast. Research on Presence in Virtual Reality: A Survey. CyberPsychology & Behavior, 4(2):183–201, apr 2001. [57] Mel Slater, Amela Sadagic, Martin Usoh, and Ralph Schroeder. Small- Group Behavior in a Virtual and Real Environment: A Comparative Study. Presence: Teleoperators and Virtual Environments, 9(1):37–51, feb 2000. 62
[58] Mel Slater, Martin Usoh, and Anthony Steed. Depth of Presence in Virtual Environments. Presence: Teleoperators and Virtual Environments, 3(2):130–144, jan 1994. [59] Jonathan Steuer. Defining Virtual Reality: Dimensions Determining Telepresence. Journal of Communication, 42(4):73–93, dec 1992. [60] Adrien Treuille, Seth Cooper, Zoran Popovi´c, Adrien Treuille, Seth Cooper, and Zoran Popovi´c. Continuum crowds. In ACM SIGGRAPH 2006 Papers on - SIGGRAPH ’06, volume 25, page 1160, New York, New York, USA, 2006. ACM Press. [61] Jur van den Berg, Ming Lin, and Dinesh Manocha. Reciprocal Velocity Obstacles for real-time multi-agent navigation. In 2008 IEEE International Conference on Robotics and Automation, pages 1928–1935. IEEE, may 2008. [62] Wouter van Toll, Roy Triesscheijn, Marcelo Kallmann, Ramon Oliva, Nuria Pelechano, Julien Pettr´e, and Roland Geraerts. A comparative study of navigation meshes. In Proceedings of the 9th International Conference on Motion in Games - MIG ’16, pages 91–100, 2016. [63] Peter Willemsen, Mark B. Colton, Sarah H. Creem-Regehr, and William B. Thompson. The effects of head-mounted display mechanics on distance judgments in virtual environments. In Proceedings of the 1st Symposium on Applied perception in graphics and visualization - APGV ’04, 2004. [64] Bob G. Witmer and Michael J. Singer. Measuring Presence in Virtual Environments: A Presence Questionnaire. Presence: Teleoperators and Virtual Environments, 7(3):225–240, jun 1998. [65] David Wolinski. Microscopic crowd simulation : evaluation and development of algorithms. PhD thesis, jan 2016. [66] Beibei Zhan, Dorothy N. Monekosso, Paolo Remagnino, Sergio A. Velastin, and Li Qun Xu. Crowd analysis: A survey. Machine Vision and Applications, 19(5-6):345–357, oct 2008. [67] B Zhou, X Tang, and X Wang. Measuring Crowd Collectiveness. In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, pages 3049–3056, 2013. [68] Bolei Zhou, Xiaogang Wang, and Xiaoou Tang. Understanding collective crowd behaviors: Learning a Mixture model of Dynamic pedestrian-Agents. 63
In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 2871–2878, 2012. 64