scieee AI-readable full text Open interactive document viewer

Vision and Crowdsensing Technology for an Optimal Response in Security: Project results

Álvarez García, Juan Antonio; Salazar González, Jose Luis; Deniz Suárez, Óscar; Ruiz Santaquiteria Alegre, Jesús

Abstract

This paper describes the progress and work carried out during the execution of the VICTORY project (Vision and Crowdsensing Technology for an Optimal Response in Security). It describes both its motivation and the progress made in the field of managing violent attacks in buildings. It is assumed that these buildings have a WiFi infrastructure that allows indoor localisation through mobile applications and a closed-circuit television system that allows real-time images to be captured. The advances of this project have allowed 1) the design of a crowdsensing application that allows indoor localisation in a reliable way and with a low battery consumption as well as the detection of stampedes and falls; 2) the development of a multi-agent model that simulates human behaviour in the face of a violent attack and 3) new approaches to improve the state of the art in computer vision for the detection of weapons

Full text

Juan A. ´ Alvarez-Garc´ ıa, Jose L. Salazar-Gonz´ alez Dpto. de Lenguajes y Sistemas Informaticos Universidad de Sevilla Sevilla, Spain jaalv[email protected], [email protected] Oscar Deniz Suarez, Jesus Ruiz Santaquiteria-Alegre VISILAB - E.T.S.I.Industrial Ciudad Real Universidad de Castilla-La Mancha Ciudad Real, Spain oscar[email protected], [email protected] Abstract—This paper describes the progress and work carried out during the execution of the VICTORY project (Vision and Crowdsensing Technology for an Optimal Response in Security). It describes both its motivation and the progress made in the field of managing violent attacks in buildings. It is assumed that these buildings have a WiFi infrastructure that allows indoor localisation through mobile applications and a closed-circuit television system that allows real-time images to be captured. The advances of this project have allowed 1) the design of a crowdsensing application that allows indoor localisation in a reliable way and with a low battery consumption as well as the detection of stampedes and falls; 2) the development of a multi-agent model that simulates human behaviour in the face of a violent attack and 3) new approaches to improve the state of the art in computer vision for the detection of weapons. Index Terms—Crowdsensing, computer vision, deep learning, multi-agents model. I. INTRODUCTION All cities in the world suffer events of diverse nature that endanger the life of its citizens. The private security sector protects public and private institutions with professionalism, but the great diversity of possible events makes it very difficult to prevent them, and above all, to plan beforehand the optimal response to each threat. Unfortunately, the current global situation due to the proliferation of terrorist acts, directly oriented towards citizenship, has only emphasized the need to evolve current security systems to deal with threats in the best possible way. The objective of this paper is to describe the motivation and results of research project ”Vision and crowdsensing for an optimal response in physical security” (VICTORY). The basic goals of the project are as follows: 1) Design a global solution to detect, analyze and classify security threats using new technologies and generating progress in the state of the art, complementing and improving current solutions. 2) Provide quick and personalized information to potential victims within the security zone, as well as to security personnel to facilitate an optimized management of the threat. Proj. TIN2017-82113 funded by Ministerio de Econom´ ıa y Competitividad, Plan Estatal 2013-2016 Retos. To achieve the basic objectives, the following systems are designed within the project and depicted in Figure I: •An efficient crowdsensing system with geolocation for the recognition of personal and group physical activities of the occupants of the security zone (right side of Figure I). It will allow detecting anomalous situations, falls and/or stampedes among other relevant indications using the inertial units of smartphones and machine learning techniques. •A computer vision system that automatically analyzes the images captured by the security cameras installed in the security zone using Deep Learning techniques (left side of Figure I). The analysis contemplates two different and complementary approaches that will provide robustness to the system. On the one hand, the generation of detectors specific to relevant patterns (such as different types of weapons), and on the other hand the use of autoencoders that alert when anomalous situations deviate from the usual scenes captured by the cameras. •An integration framework that analyzes the indications that are received from the other systems, initially vision and crowdsensing, and that processes them to establish levels of priority, probability percentages and actions to be carried out (central part of Figure I). The results will be shown to the users according to their role: –Security personnel: a central console displaying all information relating to the indications received, giving the possibility of confirming the threat or discarding it and offering a list of possible actions to be taken in response. The main objective in this case will be to optimize the precision in all the components involved so that the system is of the maximum utility for the specialized personnel. –Rest of users of the security zone: they will receive indications to follow through a mobile application (for example an escape route or a recommendation to hide) based on their position relative to the threat, the nature of the threat and the user profile (e.g. reduced mobility). The biggest challenge is to provide methods capable of adapting and reacting quickly in a changing scenario considering a large number of Vision and Crowdsensing Technology for an Optimal Response in Security: Project results parameters. Figure I shows a diagram of the project’s major building blocks. Fig. 1. System architecture. II. RESULTS A. Crowdsensing To manage a threat and inform potential victims in real time, a mobile application capable of locating the occupants of a building at the University of Seville indoors was designed and tested. The area of the building is 14325.59 m2. In addition, this application allowed the detection of falls and runs of users using the accelerometer and the gyroscope of the mobile phone so that this information could be provided to security guards in a geolocated manner to determine the threat or false positives. Figure 2 shows some screens of the application. Fig. 2. Screens showing the interface of the mobile application. The indoor location system is described in [1] where an extensive comparison have been carried out between 1) regression KNN (the one that gives better results using 11 neighbours, Manhattan distance and distance-based fingerprints weights), 2) classification KNN and 3) Naive Bayes applied to indoor localization using WiFi fingerprints. In order to prevent the application’s lifetime from being shortened, a combination of techniques to significantly reduce the battery consumption caused by the localization system have been used: accelerometer measurement to detect displacements and duty cycle reduction when access points compatible are not accessible. Results were promising, having an energy consumption of only 0.8 Wh (having a 3.7 V battery) during a 24-hour cycle and an average localization error of 4.51 meters. Worth to mention that as computation is done on server side the system can be expanded to multiple buildings and floors. Furthermore, the training dataset used, 7175 fingerprints collected from 489 different locations (almost 15 fingerprints per location) has been published1making possible to test new algorithms in the same environment. In the figure 3 the map of the building can be seen coloured as follows: green indicates an error of less than 5 meters, orange between 5 and 15 meters and red an error of more than 15 meters. The majority of places with little accuracy are more distant from the access points (APs), the zones that are surrounded by APs obtain greater accuracy. Fig. 3. Localization results - Heatmap floor 0 using regression KNN. Taken from [1] Although the proposed system supports more than 1200 active users with, in the worst case, an average response time of 243 ms, fortunately in Spain, the application did not add value for students and lecturers working in the building and the application was quickly uninstalled. B. Simulation Given the aforementioned circumstances, it was decided to simulate the behavior of potential victims and attackers by developing a model for Netlogo, a multi agent-based modeling software. An approach that combines aspects from discrete and continuous models (through aggregated components) for effi1https://dx.doi.org/10.21227/49yg-5d21 ciently simulating emergency situations under complex scenarios has been implemented including a specific agent behavior model to simulate hazards under several attack situations and pilot response protocols. Then a set of experiments has been carried out to test whether a mobile app providing realtime information can reduce the number of casualties under different conditions. A map inspired by one of the pavilions of a high school has been designed to carry out the experiments. This pavilion has an approximate size of 20x50m2(see Fig. 4, containing 6 classrooms (with sizes between 3x6m2and 6x7m2), 26 offices (with sizes between 2x2m2and 3x3m2) and 6 seminar rooms (of about 4x5m2). These rooms are connected by several corridors of about 3m wide, while the width of the doors is about 0.8m. Some rooms are marked as exits (in red) and some others as safe (in green), that agents can lock down from inside to be safe from danger until law enforcement arrives. In some experiments the safe rooms have been changed in order to test some different scenarios. Fig. 4. Example pavilion designed for the experiments and node positions. Taken from [2]. Fig. 5. Graph representation. Taken from [2]. Once the division of areas have been decided, a graph that provides the discrete representation of the space that our model use was obtained (see Fig. 5. By using as reference values used in similar spaces, or obtained as mean values in experiments with social forces models, the max-flow value for the doors was set up to 2.5, and 20 for corridors. All the implementation has been described in [2] and the software is publicly available in [3]. Figure 6 shows that, the results when using an app such as the implemented in Section II-A are really good, where the average of casualties is much lower than those when the app is not in use. Having many safe rooms reduces the distance of the agents’ movements, thus reducing the number of accidents and increasing the number of agents who manage to reach a safe room. In the figure, y-axis shows the percentage of experiments reaching the number of individuals in the x-axis for the different cases. Fig. 6. Distribution of number of agents for different states (Killed, Rescued, Accident, and in Secure Room) with different secure rooms. Blue - App on, Red - App off. Taken from [2]. C. Computer Vision From a computer vision point of view, the aim is to be able to detect a weapon visible from a CCTV camera using three approaches: 1) Analysing the state of the art of weapon detectors against close-to-reality attacks (attackers’ weapons are not in the foreground and there are significant occlusions); 2) Using an autoencoder to reduce false positives and detect anomalous situations in the images; 3) Using the pose of the attacker in addition to the detectors to improve the accuracy of the system. All the approaches are explained in the following subsections. 1) Mock attack and baseline: Since the number of public datasets related to weapons is small and generally focuses on images of people holding a gun in the foreground, it was decided to record a simulated attack on our building. Figure 7 shows a pair of images of this attack. Fig. 7. Images applying the weapon detector developed. This allowed on the one hand to generate a fairly realistic public dataset and on the other hand to analyse the state of the art in weapon detectors. In addition, to train the system, another dataset was generated with synthetic images from a video game designed using the Unity 3D engine, which allowed for the automatic generation and labelling of images in which a weapon appeared. Both datasets are publicly available2. In order to obtain a baseline for the dataset using the state of the art detectors, one camera of our CCTV was used as testing set, and the rest of datasets were used for training 2https://github.com/Deepknowledge-US/US-Real-time-gun-detection-inCCTV-An-open-problem-dataset process. Faster R-CNN was applied using Feature Pyramid Network [4] with ResNet-50 resulting in a weapon detection model able to be used in quasi real-time CCTV (90 ms of inference time with an NVIDIA GeForce GTX-1080Ti card) improving the state of the art [5] in mean average precision (mAP) on weapon detection in a two stages training. In this work [6], the impact of synthetic datasets on the training of weapon detection systems was demonstrated, but, as expected, the inference results for datasets under ideal conditions, using foreground images, are far from the real situations, as can be seen in Figure 8. Fig. 8. Plot Precision-Recall curve and iso-f1 curves for evaluation on unrealistic conditions, such as close-ups of weapons (left) and realistic ones, such as the ones in Figure 7 (right). Taken from [6]. 2) Autoencoder for anomalous situations: Even if a powerful weapon detection is available, the chances are that, when put to work in a specific end-user scenario, it will produce several false positives. This is very serious, as a high rate of false positives will make the human operators deactivate the system. That almost infinite pool of false positives (which can be obtained by operating the system for some time while human monitoring is still in place) can be leveraged for improving the detector. Given the strict dependence of deep learning methods on the training samples used, this can help fine tune models that were trained with data from elsewhere. The current detector may have been trained with data that is no longer available once the system has been deployed, so retraining may not be an option. Our proposal is to add a filtering step to the existing detector. This step will explicitly discard false positives that are generated in the end scenario, see Figure 9. This corresponds to the scene modeling module in Figure I. We have focused on handguns as it is a mortal weapon (more than knives) and is small (which represents a harder task for vision techniques). The filter has been modeled as a deep autoencoder. In experiments with a weapon detector, the results obtained show a reduction in the number of false positives between 22.5% and 87.2% and an increase of the system’s precision of 1.2%- 47% when the autoencoder is applied. 3) Pose inclusion: On the computer vision side, the major contribution in the project is the use of body pose to recognize hand guns. Body pose represents valuable information for Fig. 9. Proposed system. Taken from [7] cases in which the appearance-based handgun detector may fail (long distance to the camera, low ambient light, etc.). The methods developed are based on learning typical poses of people brandishing handguns and discriminating them from other poses [8]. This classification is then combined with the output of an appearance-based handgun detector such as Yolo. Figure 10 shows some of the results obtained with the method developed. It can be seen that the pose information allows to detect the handgun in adverse conditions (i.e. when the weapon is far from the camera and appears very small). Not only that, the proposed method runs approximately twice as fast as Yolo v3 [9]. III. CONCLUSIONS AND FUTURE WORK The VICTORY project has opened up new lines related to the development of crowdsensing applications, multi-agent simulation systems and computer vision for weapons detection. We believe it is possible to further improve results in all these fields. Specifically in computer vision we are testing Transformers techniques and at the multi-agent simulation level a simulator has been developed from scratch in the LUA language that will allow parallelisation of tasks and improve efficiency. Also noteworthy is the difficulty of using a mobile application if users do not perceive a daily added value, something that is complicated given that the application designed has a very specific functionality for situations that are fortunately infrequent. In that sense, we consider that extending the CCTV to avoid blind spots and modifying the infrastructure so that the images are not only used for a posteriori viewing but also for real-time analysis would allow a control of these events without the need to involve every single occupant of the building. Future work has been noted in [10]–[12] where pose and video analysis is used to improve the accuracy of detection of weapons and violent acts. ACKNOWLEDGMENT This research was partially supported by The Spanish Ministry of Economy and Competitiveness MINECO/FEDER R&D, UE through project VICTORY (grant no.: TIN201782113). REFERENCES [1] J. L. Salazar-Gonz´ alez, L. M. Soria-Morillo, J. A. ´ Alvarez-Garc´ ıa, F. Enr´ ıquez, and A. R. Jim´ enez, “Energy-efficient indoor localization wifi-fingerprint system: An experimental study,” IEEE Access, vol. 7, pp. 162 664–162 682, 2019. [2] E. J. L´ opez-Ortiz, F. Sancho-Caparrini, M. M. del Amor, L. M. Soria-Morillo, and J. A. ´ Alvarez Garc´ ıa, “Hybrid agent-based methodology for testing response protocols,” Knowledge-Based Systems, p. 107005, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0950705121002689 Fig. 10. A Yolo-based handgun detector vs. the proposed detector that leverages body pose. Detections are marked with a yellow box. In the four cases the weapon is detected by the proposed method but not by Yolo. Best viewed in color. [3] F. Sancho-Caparrini, “Hybrid agent-based methodology for testing response protocols,” https://www.comses.net/codebases/0478941a-af4f49c5-b08f-f21956c0a58e/releases/1.0.0/, 2021, online; CoMSES Computational Model Library; accessed 03-February-2021. [4] T.-Y. Lin, P. Doll´ ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2117–2125. [5] R. Olmos, S. Tabik, and F. Herrera, “Automatic handgun detection alarm in videos using deep learning,” Neurocomputing, vol. 275, pp. 66–72, 2018. [6] J. L. Salazar-Gonz´ alez, C. Zaccaro, J. A. ´ Alvarez-Garc´ ıa, L. M. SoriaMorillo, and F. Sancho-Caparrini, “Real-time gun detection in cctv: An open problem,” Neural networks, vol. 132, pp. 297–308, 2020. [7] N. Vallez, A. Velasco-Mata, and O. Deniz, “Deep autoencoder for false positive reduction in handgun detection,” Neural Computing and Applications, pp. 1–11, 2020. [8] J. Ruiz-Santaquiteria, A. Velasco-Mata, N. Vallez, G. Bueno, J. A. ´ Alvarez-Garc´ ıa, and O. Deniz, “Handgun detection using combined human pose and weapon appearance,” arXiv preprint arXiv:2010.13753, 2020. [9] J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” arXiv preprint arXiv:1804.02767, 2018. [10] J. Salido, V. Lomas, J. Ruiz-Santaquiteria, and O. Deniz, “Automatic handgun detection with deep learning in video surveillance images,” Applied Sciences, vol. 11, no. 13, 2021. [Online]. Available: https://www.mdpi.com/2076-3417/11/13/6085 [11] F. J. Rend´ on-Segador, J. A. ´ Alvarez Garc´ ıa, F. Enr´ ıquez, and O. Deniz, “Violencenet: Dense multi-head self-attention with bidirectional convolutional lstm for detecting violence,” Electronics, vol. 10, no. 13, 2021. [Online]. Available: https://www.mdpi.com/20799292/10/13/1601 [12] A. Velasco, J. Ruiz, N. V´ allez, and O. D´ eniz, “Using human pose information for handgun detection,” Neural Computing and Applications, vol. in-press, 2021.