scieee AI-readable full text Open interactive document viewer

Técnicas de Deep Learning aplicadas a la detección de comportamientos agresivos

Rendón Segador, Fernando José

Abstract

La detección y comprensión de comportamientos agresivos son áreas críticas en diversos campos, como la seguridad, la psicología, la educación y las redes sociales. En esta tesis, presentada como conjunto de artículos de investigación, se exploran las técnicas de aprendizaje profundo (Deep Learning) como herramientas para abordar este desafío. El desarrollo de sistemas capaces de detectar comportamientos agresivos presenta múltiples ventajas en diversos ámbitos, incluyendo la prevención de la violencia, la mejora de la seguridad pública, la atención médica y el análisis de redes sociales. Sin embargo, este avance enfrenta desafíos significativos. La precisión en la identificación de comportamientos agresivos es crucial, pero complicada debido a la variabilidad cultural y contextual. La escalabilidad es otro desafío importante, especialmente en entornos de alta escala como las redes sociales. Las preocupaciones sobre privacidad y ética también deben abordarse cuidadosamente para evitar la vigilancia excesiva y el sesgo algorítmico. Finalmente, la capacidad de generalización de los modelos a diferentes contextos y poblaciones es fundamental para su efectividad y aplicabilidad en la práctica. Estos desafíos deben ser considerados y abordados en el desarrollo de sistemas de detección de comportamientos agresivos efectivos y éticos. En este trabajo, se presenta un sistema de detección de comportamientos violentos en videos mediante técnicas de Deep Learning. El modelo final de red neuronal profunda propuesta es un vision transformer (ViT) junto al paradigma de aprendizaje neuronal estructurado con regularización adversaria (NSL). Este modelo ha sido entrenado y evaluado con los conjuntos datos NTU CCTV Fights, UBI Fights, XD Violence, UCF Crime, Hockey Fights, Violent Flows y Surveillance Camera Fights. El modelo y el paradigma de aprendizaje propuestos obtienen entre un 99 y un 100% de precisión en los conjuntos de datos mencionados, mejorando significativamente todos los resultados previos del estado del arte. Además, el modelo propuesto consigue una mayor eficiencia en términos de recursos computacionales y en tiempo de inferencia pudiendo aplicarse a entornos reales. Por otro lado, surge la problemática de un incremento de falsos positivos en validaciones cruzadas entre distintos conjuntos de datos. Para solventar este problema se optó finalmente por el desarrollo de un modelo de red neuronal profunda de ventana deslizante basado en transformer con umbral adaptativo. La solución propuesta logra eliminar la cantidad de falsos positivos, mejorando la precisión entre un 5% y un 10 % en AUC ROC en los distintos conjuntos de validación, consiguiendo así una mejora en la generalización de la detección de comportamientos violentos.

Full text

UNIVERSIDAD DE SEVILLA DEPARTAMENTO DE LENGUAJES Y SISTEMAS INFORMÁTICOS Técnicas de Deep Learning aplicadas a la detección de comportamientos agresivos. Tesis Doctoral presentada por D. Fernando José Rendón Segador dirigida por Dr. Juan Antonio Álvarez García yDr. Luis Miguel Soria Morillo. Septiembre 2024 A mi familia. Por hacer siempre todo lo posible para hacerme feliz. Agradecimientos Quiero expresar mi más sincero agradecimiento a mis directores de tesis, el Dr. D. Juan Antonio Álvarez García y el Dr. D. Luis Miguel Soria Morillo, por su constante apoyo, dedicación y esfuerzo en la realización de este trabajo. Vuestra guía ha sido fundamental para mi crecimiento, tanto personal como profesional, a lo largo de estos años. Gracias por vuestra confianza en mí. En especial, quiero agradecer a Juan Antonio, a quien le estaré eternamente agradecido por su dedicación inquebrantable. Siempre has estado pendiente de mí, preocupado por mi bienestar y por todo lo que me rodeaba, ofreciéndome tu apoyo incondicional en cada paso de este camino. A mi familia y amigos, tanto los presentes como aquellos que ya no están, gracias por haberme acompañado siempre en todas mis decisiones sin poner jamás objeción alguna. Los valores que me habéis inculcado —trabajo, esfuerzo y superación— han sido el motor que me ha permitido alcanzar grandes metas en mi vida. Os llevo siempre en mi corazón. A todos vosotros, muchas gracias de corazón. ÍNDICE GENERAL I Prefacio 1 Introducción 15 1.1 Motivación de la investigación . . . . . . . . . . . . . . . . . . . . . . 15 1.2 Metodología de investigación . . . . . . . . . . . . . . . . . . . . . . . 17 1.3 Pregunta de investigación . . . . . . . . . . . . . . . . . . . . . . . . 19 1.4 Criteriosdeéxito ............................. 20 1.5 Propiedades analizadas y discutidas . . . . . . . . . . . . . . . . . . . 22 1.6 Esquemadelatesis............................ 23 2 Estado Del Arte 27 2.1 Principales Enfoques en la Detección de Violencia . . . . . . . . . . . 28 2.2 DatasetsUtilizados............................ 32 2.3 Comparación de Métodos . . . . . . . . . . . . . . . . . . . . . . . . . 35 2.4 Falsos Positivos y Negativos en Cross-Datasets, Técnicas Basadas en VentanasTemporales........................... 37 página vii ÍNDICE GENERAL II Trabajos de investigación seleccionados 3 ViolenceNet: Dense Multi-Head Self-Attention with Bidirectional Convolutional LSTM for Detecting Violence 43 4 CrimeNet: Neural Structured Learning using Vision Transformer for violence detection 63 5 Transformer and Adaptive Threshold Sliding Window for Improving Violence Detection in Videos 79 III Observaciones finales 6 Conclusiones y trabajo futuro 103 6.1 Conclusiones................................103 6.2 Trabajofuturo ..............................105 A Curriculum 109 A.1 Revistas indexadas JCR . . . . . . . . . . . . . . . . . . . . . . . . . 109 A.2 OtrasRevistas...............................110 A.3 ProyectosI+D+i .............................111 Bibliografía 112 página viii ÍNDICE DE FIGURAS 2.1 Esquema de la arquitectura Vision Transformer (ViT) [12] . . . . . . 30 2.2 Esquema de aprendizaje neuronal estructurado para la clasificación de imágenes (2)............................... 31 ÍNDICE DE TABLAS 1.1 Resumen de artículos publicados en revistas JCR indexadas. . . . . . 24 2.1 Comparativa de los conjuntos de datos para la detección de acciones violentasenvideos............................. 34 2.2 Estado del arte del los métodos aplicados de la detección de acciones violentasenvideos............................. 36 página ix calidad de vida de las personas. Desde la seguridad en las calles hasta el ambiente en las escuelas y el bienestar en las redes sociales. La capacidad de detectar y prevenir la agresividad puede salvar vidas, reducir el estrés y mejorar la convivencia (1). Por ejemplo, en entornos educativos, la identificación temprana de comportamientos agresivos puede prevenir el acoso escolar y fomentar un ambiente de aprendizaje seguro. A nivel científico, esta investigación contribuye al campo del aprendizaje profundo y su aplicación en el análisis de comportamiento humano. Aunque ya existen estudios sobre la detección de violencia, poseen limitaciones en términos de precisión, escalabilidad y generalización. Esta tesis aborda estas lagunas al proponer un modelo avanzado que combina Vision Transformers y aprendizaje neuronal estructurado con regularización adversaria, logrando resultados superiores a los existentes. Las aplicaciones prácticas de esta investigación son diversas y de gran alcance. En el ámbito de la seguridad pública, los sistemas de vigilancia pueden integrar el modelo propuesto para identificar incidentes violentos en tiempo real, permitiendo una respuesta rápida por parte de las autoridades. En el sector de la salud, los profesionales pueden utilizar estas herramientas para monitorizar el comportamiento de pacientes con trastornos agresivos. Por otro lado, en las redes sociales, las plataformas pueden emplear estos sistemas para detectar y mitigar el contenido violento, creando un entorno más seguro para los usuarios y controlando la propagación de comportamientos agresivos en línea. A pesar de los avances tecnológicos, la detección de comportamientos agresivos presenta múltiples desafíos. La variabilidad cultural y contextual puede dificultar la identificación precisa de la agresividad, mientras que la necesidad de escalabilidad y eficiencia computacional es crucial para su implementación en tiempo real. La generalización del concepto de violencia es otro desafío significativo, ya que lo que se considera agresivo puede variar ampliamente entre diferentes culturas y contextos. Además, las preocupaciones éticas sobre la privacidad y el sesgo algorítmico deben (1)https://elpais.com/podcasts/hoy-en-el-pais/2024-02-01/podcast-moderadores-de-contenidoen-redes-protectores-desprotegidos.html ser cuidadosamente gestionadas para evitar consecuencias no deseadas. Personalmente, esta investigación refleja mi compromiso con la utilización de la tecnología para resolver problemas sociales críticos. Profesionalmente, busco contribuir al desarrollo de sistemas inteligentes que no sólo sean tecnológicamente avanzados, sino también éticamente responsables y socialmente beneficiosos. Esta tesis es un paso hacia el logro de estos objetivos, combinando mi pasión por el aprendizaje profundo con mi deseo de impactar positivamente en la sociedad. La motivación detrás de esta investigación es proporcionar una solución innovadora y eficaz para la detección de comportamientos agresivos, con el potencial de transformar diversas áreas críticas de nuestra sociedad. 1.2. Metodología de investigación Este trabajo sigue una técnica de investigación científica estándar [24] que incluye las siguientes fases: 1. Definir el problema de investigación: La detección y clasificación de comportamientos agresivos en videos es crucial para diversas aplicaciones como la seguridad pública, la atención médica y el control de contenidos en redes sociales. En esta tesis se propone un sistema de clasificación de acciones violentas en videos en tiempo real utilizando técnicas de Deep Learning. Este sistema supera a los enfoques tradicionales y permite una mayor precisión y eficiencia, abordando tanto la identificación inicial de comportamientos agresivos como la corrección de falsos positivos y negativos. 2. Revisión de la literatura: Durante el desarrollo de esta tesis doctoral se llevó a cabo una revisión exhaustiva de la literatura existente sobre la detección de comportamientos agresivos utilizando técnicas de Deep Learning. Esto incluyó el análisis de estudios previos sobre Redes Neuronales Convolucionales, Vision Transformers, y distintos tipos de aprendizaje además de otras metodologías relevantes, tal como se muestra en las referencias incluidas en cada artículo. 3. Formular hipótesis: El grupo de investigación discutió la aplicación de técnicas avanzadas de inteligencia artificial, específicamente Vision Transformers y redes neuronales convolucionales, para mejorar la precisión y eficiencia en la detección de comportamientos agresivos. Se planteó la hipótesis de que la combinación de estos enfoques con un modelo de ventana temporal adaptativa podría reducir significativamente los falsos positivos y mejorar la generalización de los modelos a diferentes contextos. 4. Diseño de la investigación: Se analizaron diversos modelos de Deep Learning para identificar puntos clave susceptibles de mejora. Esto incluyó la evaluación de diferentes arquitecturas de Transformers y la integración de mecanismos de regularización adversaria. Además, se diseñó un modelo complementario basado también en Transformers para la corrección de falsos positivos y negativos. 5. Recolectar datos: Se emplearon varios conjuntos de datos ampliamente utilizados y validados por la comunidad científica para la tarea de reconocimiento de comportamientos agresivos, como NTU CCTV Fights, UBI Fights, XD Violence, UCF Crime, Hockey Fights, Violent Flows y Surveillance Camera Fights. Estos datos fueron utilizados para entrenar y evaluar los modelos propuestos. Además, se generó un dataset específico de situaciones violentas mediante sistemas de detección y clasificación desarrollados durante esta investigación. 6. Ejecución del proyecto: Las arquitecturas de redes neuronales diseñadas, incluyendo el Vision Transformer y el modelo de ventana temporal adaptativa, fueron entrenadas y validadas utilizando los datos recolectados. El proceso de entrenamiento involucró la optimización de hiperparámetros y la implementación de técnicas de regularización para mejorar la generalización. 7. Análisis de datos: Los resultados obtenidos por los sistemas desarrollados fueron analizados utilizando métricas estándar de rendimiento, como precisión, recall, AUC-PR y AUC-ROC. La precisión mide la proporción de verdaderos positivos, es decir, las detecciones correctas, sobre el total de elementos clasificados como positivos (verdaderos positivos más falsos positivos). Esta métrica es crucial en situaciones donde es necesario minimizar los falsos positivos. Por otro lado, el recall, también conocido como sensibilidad, evalúa la capacidad del modelo para identificar todos los casos positivos reales. Se calcula como la proporción de verdaderos positivos sobre el total de elementos que realmente son positivos (verdaderos positivos más falsos negativos). Esta métrica es especialmente importante en contextos donde no detectar un caso positivo puede tener consecuencias graves. Además, el AUC-PR (Área bajo la curva de precisión-recall) proporciona una medida del rendimiento del modelo en términos de precisión y recall, siendo particularmente útil en conjuntos de datos desbalanceados, ya que se centra en el rendimiento de la clase positiva. En cuanto al AUC-ROC (Área bajo la curva de la característica operativa del receptor), esta métrica representa la capacidad del modelo para distinguir entre clases. Mide la tasa de verdaderos positivos frente a la tasa de falsos positivos en diferentes umbrales de clasificación, donde un AUC cercano a 1 indica un excelente desempeño del modelo. 8. Interpretar e informar: Una vez que los sistemas propuestos fueron analizados y sus resultados interpretados, se publicaron varios artículos de investigación. Estos artículos documentaron las hipótesis iniciales, los métodos utilizados, los resultados obtenidos y las conclusiones derivadas del estudio, contribuyendo al conocimiento en el área de la detección de comportamientos agresivos mediante técnicas de Deep Learning. 1.3. Pregunta de investigación La pregunta de investigación que conduce esta tesis doctoral es: ¿Podemos lograr un modelo que generalice el concepto de violencia, mejorando la precisión de los sistemas de detección de comportamientos agresivos en videos, sin disminuir la eficiencia en términos de recursos computacionales en espacio y tiempo, utilizando técnicas de Deep Learning? Dentro del contexto de la detección de comportamientos agresivos, nos centramos en desarrollar varias arquitecturas de redes neuronales que combinen Vision Transformers y técnicas de regularización adversaria. Los Vision Transformers permiten procesar secuencias de imágenes (videos) de manera eficiente, capturando relaciones espaciales y temporales cruciales para la identificación de comportamientos agresivos. Estos modelos se complementan con una ventana temporal deslizante adaptativa que corrige falsos positivos y negativos, mejorando así la precisión general del sistema. El objetivo final es analizar cómo estos elementos se comportan al ser entrenados y evaluados con diferentes conjuntos de datos de comportamientos agresivos, aplicando optimizadores basados en algoritmos de descenso de gradientes. Además, buscamos entender cómo la combinación de Vision Transformers y la ventana temporal adaptativa puede mejorar la capacidad del sistema para generalizar el concepto de violencia a distintos contextos y poblaciones, manteniendo al mismo tiempo la eficiencia en términos de recursos computacionales (espacio y tiempo). Dentro del contexto de la detección de comportamientos agresivos, nos centramos en analizar y comparar exhaustivamente el comportamiento de las distintas redes neuronales propuestas en la literatura, adaptándolas específicamente al contexto de la detección de comportamientos violentos en videos. La investigación también aborda la necesidad de mantener la eficiencia en términos de requisitos de memoria y tiempo de inferencia, asegurando que los sistemas desarrollados sean aplicables en entornos reales. 1.4. Criterios de éxito El éxito se logrará si la pregunta de investigación se resuelve. Esto significa que debemos comprobar, por un lado, que el sistema de detección es capaz de identificar correctamente comportamientos agresivos en videos reales. Por otro lado, debemos verificar que el modelo propuesto generaliza bien el concepto de violencia y mantiene la eficiencia en términos de recursos computacionales (espacio y tiempo). Los resultados mostrados en esta tesis, en forma de artículos de investigación, demuestran que coinciden con la predicción formulada en la hipótesis de partida. Para ello, hemos definido los siguientes criterios de éxito específicos: Precisión en la detección: El sistema de clasificación debe ser capaz de identificar comportamientos agresivos en videos con alta precisión, minimizando tanto los falsos positivos como los falsos negativos. Esto se evaluará mediante métricas estándar como precisión, recall, F1-score, y AUC-ROC en varios conjuntos de datos de comportamiento agresivo (NTU CCTV Fights, UBI Fights, XD Violence, UCF Crime, Hockey Fights, Violent Flows y Surveillance Camera Fights). Generalización del concepto de violencia: El modelo debe demostrar su capacidad para generalizar el concepto de violencia a distintos contextos y poblaciones. Esto implica que el sistema debe mantener una alta precisión incluso cuando se evalúa en datos que no se utilizaron durante el entrenamiento. Se realizará una validación cruzada entre diferentes conjuntos de datos para verificar esta capacidad de generalización. Eficiencia en recursos computacionales: El sistema debe ser eficiente en términos de uso de memoria y tiempo de inferencia. Esto se evaluará midiendo el consumo de recursos computacionales durante la ejecución y comparando estos resultados con otros sistemas de detección de comportamientos agresivos. La solución debe ser viable para su implementación en tiempo real en entornos prácticos, como sistemas de vigilancia y plataformas de redes sociales. Corrección de falsos positivos y negativos: El modelo complementario de ventana temporal deslizante con umbral adaptativo debe demostrar su eficacia en la corrección de falsos positivos y negativos, mejorando así la precisión general del sistema. La evaluación de este criterio se realizará mediante la comparación de los resultados del modelo principal con y sin la aplicación de la ventana temporal. Esta tesis establece nuevos estándares en la precisión de la detección de comportamientos agresivos en videos, logrando una mejor generalización del concepto de violencia y manteniendo la eficiencia computacional. Los modelos desarrollados están preparados para ser utilizados en entornos reales y ofrecen una solución robusta y eficaz para la detección de comportamientos agresivos. 1.5. Propiedades analizadas y discutidas La detección y clasificación de comportamientos agresivos en videos se abordan desde varios puntos de vista: Precisión: En relación a la precisión de los modelos de redes neuronales entrenados y evaluados utilizando conjuntos de datos públicos de comportamientos agresivos. Esta propiedad es crucial para asegurar que el sistema sea robusto y confiable en la identificación de comportamientos violentos. La precisión se mide a través de métricas estándar como precisión, recall, AUC-PR, y AUCROC, lo que permite evaluar la efectividad del modelo en diversos escenarios y contextos. Generalización: En relación a la capacidad del modelo para generalizar el concepto de violencia a diferentes contextos y poblaciones. Esta propiedad se evalúa mediante validaciones cruzadas entre distintos conjuntos de datos, asegurando que el modelo mantenga su precisión incluso cuando se enfrenta a datos no vistos previamente. La generalización es esencial para que el sistema sea aplicable en una amplia gama de situaciones reales. Eficiencia: En relación al coste computacional, consumo de memoria, y tiempo de ejecución que requieren los modelos propuestos. Esta propiedad permite seleccionar el modelo que mejor se adapte a las necesidades prácticas, asegurando que el sistema sea viable para implementaciones en tiempo real. La eficiencia se mide comparando el uso de recursos computacionales y el tiempo de inferencia de los modelos con otros enfoques existentes en la literatura. El trabajo presentado en este documento propone una solución para mejorar la robustez y la precisión de los sistemas de detección de comportamientos agresivos, al mismo tiempo que se mejora la eficiencia en términos de recursos computacionales. Además, se enfatiza la importancia de la generalización del concepto de violencia, asegurando que el modelo sea aplicable a diferentes contextos y poblaciones sin perder precisión ni eficiencia. 1.6. Esquema de la tesis Este documento está estructurado de la siguiente forma. En la Parte I se presenta la introducción y el estado del arte correspondientes al capítulo 1 y al capítulo 2. En la Parte II se muestran los tres artículos de investigación derivados de esta tesis doctoral, divididos en 3 capítulos: Capítulo 3 - “ViolenceNet: Dense Multi-Head SelfAttention with Bidirectional Convolutional LSTM for Detecting Violence”; Capítulo 4 - “CrimeNet: Neural Structured Learning using Vision Transformer for violence detection”; Capítulo 5 - “Transformer and Adaptive Threshold Sliding Window for improving violence detection in videos”. Las revistas donde se han publicado estos artículos están incluidas en el ranking JCR de Thomson Reuters y todos ellos están relacionados con el problema de la detección y clasificación de acciones violentas en videos: ViolenceNet: Dense Multi-Head Self-Attention with Bidirectional Convolutional LSTM for Detecting Violence. Rendón-Segador, F. J., Álvarez-García, J. A., Enríquez, F., & Deniz, O. Publicado en Electronics, MDPI, ISSN: 0957-4174, Fecha de Publicación: 3 Julio 2021, Volumen: 10(13), En Páginas: 1601, DOI: https://doi.org/10.3390/electronics10131601, [JCR-2023 2.9] [Q2 en Computer Science, Information Systems (123/251)]. CrimeNet: Neural Structured Learning using Vision Transformer for violence detection. Rendón-Segador, F. J., Álvarez-García, J. A., SalazarGonzález, J. L., & Tommasi, T. Publicado en Neural Networks, Elsevier, Fecha de Publicación: Abril 2023, Volumen: 161, En Páginas: 318-329, DOI: https://doi.org/10.1016/j.neunet.2023.01.048, [JCR-2023 7.8] [Q1 en Computer Science, Artificial Intelligence (28/145)]. Transformer and Adaptive Threshold Sliding Window for improving violence detection in videos. Rendón Segador, Fernando José, Álvarez García, Juan Antonio, Soria Morillo, Luis Miguel. Publicado en Sensors, MDPI, Fecha de Publicación: 22 Agosto 2024, Volumen: 24(16), En Páginas: 5429, DOI: https://doi.org/10.3390/s24165429, [JCR-2023 3.4] [Q2 en Computer Science, Artificial Intelligence (119/354)]. Un resumen del ranking de estos artículos de investigación se puede encontrar en la Tabla 1.1. Título Revista F.I. Ranking ViolenceNet: Dense Multi-Head Self-Attention with Bidirectional Convolutional LSTM for Detecting Violence Electronics 2021 2.9 Q2 CrimeNet: Neural Structured Learning using Vision Transformer for violence detection Neural Networks 2023 7.8 Q1 Transformer and Adaptive Threshold Sliding Window for improving violence detection in videos Sensors 2024 3.4 Q2 Tabla 1.1: Resumen de artículos publicados en revistas JCR indexadas. Finalmente, en la Parte III, se exponen comentarios finales, conclusiones y se discute el trabajo futuro. pueden generalizar mejor ante ejemplos no vistos. Esto es particularmente útil en tareas donde los datos no siguen un patrón homogéneo, como en la detección de violencia en videos, donde los contextos pueden variar significativamente. NSL puede hacer los modelos más robustos frente a datos ruidosos o perturbaciones adversarias. Durante el entrenamiento, puede aplicar técnicas como el entrenamiento adversarial, donde el modelo es entrenado para manejar ejemplos perturbados que intentan inducir errores. Además de los datos de entrada, NSL puede aprovechar información adicional, como etiquetas suaves o representaciones de grafos que describen relaciones implícitas entre los datos, mejorando la capacidad del modelo para aprender de manera más rica y estructurada. 2.2. Datasets Utilizados La elección del conjunto de datos es crítica para evaluar la eficacia de los modelos de detección de violencia en videos. En esta tesis, se han utilizado diversos datasets públicos que abarcan distintos tipos de escenarios violentos, desde entornos de videovigilancia hasta contenido multimedia general. Cada conjunto de datos ofrece diferentes características en cuanto a tamaño, tipos de violencia y la complejidad de las escenas, permitiendo evaluar la capacidad de generalización de los modelos propuestos. A continuación se describen los datasets utilizados en esta tesis: XD-Violence [39]: Conjunto de datos extenso y diverso, compuesto por clips de vigilancia y contenido multimedia con múltiples categorías de violencia. Su gran variabilidad en tipos de violencia y la presencia de eventos ambiguos lo hacen un desafío para los modelos de detección. UCF-Crime [35]: Este dataset está orientado a la videovigilancia en escenarios reales y contiene una amplia gama de incidentes delictivos, incluyendo violencia física. La diversidad de contextos y la longitud de los clips presentan un reto significativo para los sistemas de detección. RWF-2000 [9]: Conjunto de videos que contiene principalmente peleas grabadas con cámaras de seguridad en entornos controlados. Se utiliza frecuentemente para evaluar modelos en escenarios de videovigilancia donde los eventos son más directos. Hockey Fights [5]: Dataset pequeño, centrado en peleas durante partidos de hockey sobre hielo. Aunque es más limitado en términos de diversidad, es útil como referencia para evaluar modelos en situaciones deportivas controladas. Real Life Violence Situations [34]: Contiene videos de peleas y violencia grabados en situaciones cotidianas, con cámaras de teléfonos móviles o de seguridad. La calidad variable de los videos y los escenarios no controlados añaden una capa de complejidad. NTU-CCTV Fights [29]: Un conjunto de videos capturados por cámaras de seguridad, centrado en peleas y disturbios. Es útil para evaluar el rendimiento de los modelos en entornos de vigilancia reales. Surveillance Camera Fights [3]: Este dataset incluye videos de peleas grabados en entornos de videovigilancia, ofreciendo una amplia gama de situaciones de violencia en tiempo real. UBI-Fights [11]: Este conjunto de datos contiene videos de peleas de vigilancia y eventos violentos en espacios públicos. Es especialmente relevante para la evaluación de modelos en escenarios del mundo real donde las peleas no son escenificadas. Violent Flows [19]: Se compone de videos de violencia en flujo continuo, es decir, situaciones donde los eventos violentos no están claramente delimitados en el tiempo. Esto lo convierte en un desafío para modelos que dependen de la detección temporal. MediaEval 2013 Violent Scene Detection (VSD) [32]: Un conjunto de datos multimedia que contiene clips de películas con escenas violentas. Aunque los escenarios son más controlados y menos caóticos que en videos de vigilancia, es útil para evaluar la detección de violencia en contenido audiovisual más general. La Tabla 2.1 ofrece una comparación de los datasets en términos de tamaño, duración de los videos, cantidad de categorías (binarias o multiclase) y tipo de escenarios violentos representados. Dataset NºClips Duración Promedio (s) Categorías Escenarios Tipos de Violencia XD-Violence [39] 4754 10-60 Multiclase Videovigilancia y multimedia Variada UCF-Crime [35] 1900+ 50-300 Multiclase Videovigilancia Variada RWF-2000 [9] 2000 5-30 Binaria Cámaras de vigilancia Peleas Hockey Fights [5] 1000 1-10 Binaria Deporte (hockey) Peleas en deporte Real Life Violence Situations [34] 1000+ 10-30 Binaria Grabaciones móviles Peleas en público NTU CCTV Fights [29] 1500+ 20-40 Binaria Cámaras de seguridad Peleas y disturbios Surveillance Camera Fights [3] 500+ 15-45 Binaria Videovigilancia Peleas UBI-Fights [11] 1200+ 5-25 Binaria Cámaras públicas Variada en público Crowd Violence/Violent Flows [19] 800+ 10-40 Binaria Flujos de video continuos Variada MediEval 2013 VSD [32] 2000+ 10-30 Binaria Peliculas y series Escenas violentas en el cine Tabla 2.1: Comparativa de los conjuntos de datos para la detección de acciones violentas en videos. Cada uno de estos conjuntos de datos ofrece características únicas que los hacen adecuados para diferentes tipos de evaluación. Por ejemplo, XD-Violence [39] y UCFCrime [35] son particularmente desafiantes debido a su diversidad de escenarios y su combinación de tipos de violencia, lo que permite evaluar la robustez de los modelos en situaciones complejas. En cambio, RWF-2000 [9] y Surveillance Camera Fights [3] se centran más en incidentes de peleas en entornos controlados de videovigilancia, lo que los convierte en buenos benchmarks para aplicaciones de seguridad en tiempo real. Violent Flows [19] añade una dimensión adicional al problema al introducir videos con violencia que ocurre de manera fluida, sin delimitaciones claras de inicio o fin, desafiando así las capacidades de los modelos para detectar eventos violentos a lo largo del tiempo. Finalmente, MediEval VSD [32] se destaca por su enfoque en escenas de violencia dentro de películas y series, proporcionando un escenario de evaluación más orientado al análisis de contenido multimedia general, lo que es útil para modelos aplicados a plataformas de monitoreo audiovisual. 2.3. Comparación de Métodos En la Tabla 2.2, se presenta una comparación de los principales métodos de detección de violencia en videos, evaluados en términos de exactitud, área bajo la curva ROC (AUC ROC) y precisión promedia (AP) en diferentes conjuntos de datos. Esta tabla resume los resultados de estudios recientes y proporciona una visión integral del rendimiento de los enfoques más destacados en el campo. Los modelos basados en redes neuronales convolucionales (CNN) han sido fundamentales en la evolución de la detección de violencia. Entre los enfoques más destacados se encuentran EfficientNet CNN con el bloque TSE [22], que muestra un excelente rendimiento en varios conjuntos de datos, alcanzando hasta un 99.60% de precisión. Este modelo es especialmente eficaz cuando se complementa con técnicas de atención temporal y espacial, lo que permite capturar mejor las dinámicas de eventos violentos. De manera similar, los modelos CNN + LSTM [4, 37, 3, 28], que integran procesamiento temporal, también han obtenido resultados sólidos, como el 99 % de precisión en el conjunto de datos Hockey Fights. Estos métodos demuestran la capacidad de las CNN para extraer características espaciales y temporales relevantes en entornos complejos. Las arquitecturas basadas en convoluciones 3D, como ConvNet 3D y C3D (Convolutional 3D Networks) [14, 20, 2, 29], son ampliamente utilizadas para capturar patrones espacio-temporales en videos, especialmente en la detección de violencia. Estos modelos extienden las capacidades de las CNN tradicionales al operar directamente en el dominio temporal. Por ejemplo, C3D combinado con SVM [2] ha logrado un notable 99.29 % de precisión en el conjunto Crowd Violence. Este enfoque demuestra la eficacia de las convoluciones 3D en la captura de información dinámica en escenarios violentos. Dataset Method ROC AUC AP Accuracy RWF-2000 EfficientNet CNN + TSE Block [22] - - 87.25 % ConvNet 3D [9] - - 92.00 % MediEval 2013 VSD C3D [27] - - 68.3 % CNN-LSTM [28] - - 61.00 % Surveillance Camera Fights CNN-LSTM + Attention [3] - - 72.00 % EfficientNet CNN + TSE Block [22] - - 92.00 % Real Life Violence Situations CNN + Transformer [1] - - 98.25 % ViT [33] 98.00 % 98.00 % 98.00 % EfficientNet CNN + TSE Block [22] - - 97.80 % Crowd Violence/Violent Flows C3D + SVM [2] 99.00 % - 99.29 % Bi-CNN-LSTM [18] - - 96.32 % Hockey Fights ConvNet 3D [20] - - 98.00 % CNN-LSTM + Attention [4] - - 98.00 % CNN-LSTM [37] - - 99.00 % EfficientNet CNN + TSE Block [22] - - 99.60 % NTU CCTV Fights 3D CNN VGG16 [29] 79.50% 75.90% - CNN + Transformer [1] - - 96.00 % UBI Fights MIL + C3D + BayesianNet [11] 90.60% - - Weakly Supervised Two-Stage [30] 93.10% - - CNN + Transformer [1] - - 91.80 % XD Violence CNN Contrastive + Attention [8] - 76.9% - MIL + Multi-scale temporal network [36] - 77.81 % - CNN + Relation modeling module [40] - 78.64 % - Contrastive Instance + Weakly-Supervised Audio-Visual [41] - 83.40 % - CNN + Attention [26] - 81.69 % - UCF Crime ConvNet 3D + MIL [14] 76.67% - - MIL + Multi-scale temporal network [36] 84.03% - - MIL + C3D + Attention [16] 82.30 % - - DMIL + MRM [13] 81.91 % - - MIL + CNN AutoEncoder [21] 80.10% - - Tabla 2.2: Estado del arte del los métodos aplicados de la detección de acciones violentas en videos. En los últimos años, los modelos basados en Transformers, como Vision Transformer (ViT) y CNN + Transformer, han emergido como un avance clave en la detección de violencia en videos. Estos métodos destacan por su capacidad para modelar relaciones espaciales y temporales de manera más eficiente que las arquitecturas tradicionales. Por ejemplo, ViT ha alcanzado hasta un 98 % de precisión y AUC ROC en el conjunto Real Life Violence Situations, superando a muchos otros enfoques en términos de precisión y robustez. Además, los Transformers son particularmente efectivos en escenarios complejos, donde la violencia puede involucrar interacciones humanas intrincadas que requieren una modelización más detallada. Los enfoques basados en MIL (Multiple Instance Learning) han sido especialmente efectivos en conjuntos de datos grandes y desafiantes, como UCF Crime y XD Violence. Estos métodos se destacan en situaciones donde la anotación precisa de cada frame es difícil, permitiendo que el modelo aprenda de etiquetas globales para secuencias de video completas. MIL combinado con redes 3D y módulos de atención ha logrado ROC AUC de hasta 84.03%, lo que resalta su potencial para manejar videos con violencia distribuida de manera desigual a lo largo del tiempo. En resumen, los avances recientes en la detección de violencia en videos han sido impulsados por una combinación de modelos convolucionales, técnicas de atención y arquitecturas basadas en Transformers. Mientras que los enfoques clásicos basados en CNN y LSTM continúan siendo eficaces, los modelos más recientes que integran Transformers y aprendizaje por instancias múltiples (MIL) han elevado significativamente el estado del arte, logrando mejoras considerables en precisión y capacidad de generalización. Sin embargo, sigue existiendo un reto en cuanto a la robustez y la capacidad de los modelos para generalizar entre distintos conjuntos de datos, especialmente en escenarios de videovigilancia y situaciones reales complejas. 2.4. Falsos Positivos y Negativos en Cross-Datasets, Técnicas Basadas en Ventanas Temporales La experimentación en cross-datasets implica entrenar un modelo utilizando un conjunto de datos y evaluarlo con otro conjunto de datos diferente. Este enfoque es fundamental para medir la capacidad de generalización del modelo, es decir, su habilidad para reconocer patrones en datos no vistos previamente. La generalización adecuada es crucial en la detección de acciones violentas, ya que un modelo que no generaliza bien puede sufrir problemas de sobreajuste (overfitting), lo que conduce a un rendimiento deficiente en nuevos datos. Esto genera falsos positivos y negativos no deseados, afectando la efectividad del sistema en situaciones reales. Un desafío adicional en la detección de eventos violentos es el uso de métodos que analizan los videos frame a frame o en pequeños conjuntos de fotogramas, los cuales tienden a introducir un nivel considerable de ruido. Este ruido puede surgir cuando la transición entre cuadros con violencia y aquellos sin violencia es difusa, lo que resulta en detecciones inconsistentes. En consecuencia, los modelos pueden alternar entre detectar violencia y no violencia de manera errónea, afectando la fiabilidad del sistema. Uno de los enfoques más prometedores para mitigar tanto los problemas de ruido como los falsos positivos y negativos, especialmente en cross-datasets, es el uso de modelos de ventanas temporales deslizantes con umbrales adaptativos. Estos modelos permiten que el sistema procese secuencias de video en pequeñas “ventanas” de tiempo, analizando la información de manera más granular y capturando eventos violentos que pueden ser de corta duración pero que tienen un gran impacto en el contexto de seguridad. Los métodos tradicionales, como los basados en CNN o ConvNet 3D, suelen operar frame a frame o utilizando una ventana temporal fija. Sin embargo, estos enfoques a menudo resultan insuficientes en situaciones donde la violencia se desarrolla de forma más sutil o en secuencias más largas. Los modelos basados en ventanas temporales deslizantes, en cambio, ajustan dinámicamente la longitud de la ventana y los umbrales de detección, lo que permite una mayor sensibilidad a los cambios contextuales en los videos. Modelos de Umbrales Adaptativos Dentro de esta categoría, los modelos que implementan un ajuste dinámico de umbrales han demostrado ser eficaces en la mejora de la precisión en la detección de eventos violentos. Al modificar los umbrales de detección en función del contexto de la escena o la secuencia temporal analizada, estos sistemas logran reducir falsos positivos, que suelen surgir en entornos caóticos o con movimientos abruptos no violentos. De manera similar, la capacidad de ajustar los umbrales permite que el modelo sea más robusto frente a falsos negativos, asegurando que los eventos de violencia no sean ignorados, incluso cuando son menos evidentes o de corta duración. Por ejemplo, en varios trabajos recientes que combinan modelos CNN o ViT con un enfoque de ventana deslizante [6, 40, 25, 31], se ha reportado una mejora en la precisión general del sistema, tanto en términos de área bajo la curva ROC (AUC ROC) como de reducción de falsos positivos y negativos. Este tipo de mecanismos son especialmente útiles en escenarios de videovigilancia continua, donde la sobrecarga de alarmas falsas puede llevar a que los operadores humanos desactiven el sistema, reduciendo su efectividad en la práctica. Estos avances han sido clave para reducir el número de falsos positivos y negativos, manteniendo al mismo tiempo una alta sensibilidad a eventos violentos reales. Esto es particularmente crucial en sistemas que operan de manera autónoma o semiautónoma, donde la confianza en la detección automática es fundamental para evitar la intervención humana constante. PARTE II Trabajos de investigación seleccionados página 41 Electronics 2021,10, 1601 4 of 17 formats and more as inputs. These are convolutional multi-stream models in which each stream analyzes a type of video feature. In [ 38 ] the authors presented a multi-stream model called FightNet with three types of input modes, that is, RGB images, optical flow images, and acceleration images, each one with an associated network. The input video is divided into segments and for each segment the feature maps of each stream are obtained. Finally, all these feature maps are merged. The final output of the network is the average score of each segment of the video. Our architecture is based on the recurrent convolutional architecture with an attention mechanism. The most similar work to our proposal is that of [ 36 ], but we improve it on the following relevant points: (1) the optical flow of the video as input to the network instead of the RGB format, (2) the inclusion of the DenseNet architecture adapted to three dimensions, using 3D convolutional layers instead of 2D ones and (3) the use of multi-head self-attention layer [ 39 ] instead of attention mechanisms [ 37 ], creating a novel architecture for the detection of violence. Once the new model is obtained, a cross-dataset experiment is carried out to check whether it is capable of generalizing violent actions. The following section describes it in detail. 3. Model Architecture To correctly classify violence in videos, the generation of robust video encoding is fundamental to later classify it using a fully connected network. To achieve it, each video is transformed from RGB to optical flow. Then a Dense network is used to encode the optical flow as a sequence of feature maps. These feature maps are passed through a multi-head self-attention layer and then through a bidirectional ConvLSTM layer to apply the attention mechanism in both temporal directions of the video (forward and backward pass). This spatio-temporal encoder with an attention mechanism extracts relevant spatial and temporal features for each video. Finally, the encoded features are fed into a four-layer classifier that classifies the video into two categories (violence and non-violence). The architecture of the model, called ViolenceNet, is shown in Figure 1. Electronics 2021,10, 1601 5 of 17 Dense Block Fully Connected Convolutional 3D MaxPooling 3D Convolutional 3D AveragePooling 3D AveragePooling 3D Bidirectional Convolution2DLSTM x6 x12 x24 x16 Bidirectional Convolutional LSTM 2D Classifier Batch Normalization Convolutional 3D (1 x 1 x 1) Batch Normalization Convolutional 3D (3 x 3 x 3) Batch Normalization Convolutional 3D (1 x 1 x 1) Batch Normalization Convolutional 3D (3 x 3 x 3) Batch Normalization Convolutional 3D (1 x 1 x 1) Batch Normalization Convolutional 3D (3 x 3 x 3) Batch Normalization Convolutional 3D (1 x 1 x 1) Batch Normalization Convolutional 3D (3 x 3 x 3) Batch Normalization Convolutional 3D (1 x 1 x 1) Batch Normalization Convolutional 3D (3 x 3 x 3) Video RGB Video Optical Flow Section B Dense Net Multi-head Self-Attention h = 6 Feature Map Q Q K K V V Attention Feature Map Backward/Fordward Attention Feature Map Section A Dense Block x5 Figure 1. In Section Athe architecture of the ViolenceNet model, which takes the optical flow as input, is shown. It is composed of four parts: a DenseNet-121 network spatio-temporal encoder, a multi-head self-attention layer [ 39 ], a bidirectional convolution 2D LSTM (BiConvLSTM2D) layer and a classifier. Below each Dense Block, its number of components is indicated. The variable h corresponds to the number of heads used in parallel by the multi-head self-attention layer and the variables Q,K,Vtheir inputs. Section Bshows the internal architecture of a five-component Dense Block (X5). 3.1. Model Justification The blocks that compose the architecture of the model have been successfully tested in the field of human action recognition and, specifically, in the field of violent actions, as shown in Section 2.2. In addition, the 3D DenseNet variant has been used for video classification [ 40 ], bidirectional recurrent convolutional block convolutional improved the efficiency in detecting violent actions [ 34 ], as it allows analyzing features in both temporal Electronics 2021,10, 1601 6 of 17 directions. Attention mechanisms have been used in recent years with good results in the task of recognizing human actions [ 41 , 42 ] and the combination of convolutional networks and bidirectional convolutional recurrent blocks has already shown efficiency in learning spatial and temporal information in videos [ 43 ]. These facts prove that the model is based on blocks that are useful to recognize human actions in videos and have led us to use them to develop our proposal. The following subsections describe the architecture of the ViolenceNet model in detail. 3.2. Optical Flow One of the inputs of our network is the dense optical flow [ 44 ]. This algorithm generates a sequence of frames where those pixels that move the most between consecutive frames are represented with greater intensity. This information is a key element in violent scenes since the most important components are contact and speed: pixels tend to move more during that segment of the video than in the rest and tend to cluster in one area of the scene. Once the algorithm, a 2-channel matrix with optical flow vectors that include magnitude and direction is obtained. The direction corresponds to the hue value of the image, while the magnitude corresponds to the value plane. The hue value is used for visualization only. We selected dense optical flow over sparse optical flow because the former provides the flow vectors for the full-frame, up to one flow vector per pixel, while sparse flow only provides the flow vectors some “interesting features”, such as some pixels that represent the edges or corners of an object within the frame. In deep learning models, like the proposed one, the selection of features is unsupervised, therefore it is better to have a wide range of features than a limited one. Mainly for this reason dense optical flow is used as input to the model. 3.3. DenseNet Convolutional 3D The DenseNet architecture [ 45 ] was designed to be used with images so it is composed of 2D convolutional layers. However, it can be adapted so that it can process videos. Two modifications to the original have been made: the first one is to replace the 2D convolutional layers with 3D ones; the second, is to replace the 2D reduction layers with 3D reduction ones. DenseNet uses the reduction layers MaxPool2D and AveragePool2D with a pool size of (2, 2) and (7, 7). The reduction layers MaxPool3D and AveragePool3D were used with a pool size of (2, 2, 2) and (7, 7, 7). DenseNet takes its name from the dense blocks that are the foundation of the network architecture. These blocks concatenate the feature maps of a layer with all of its descendants. In our proposal, four dense blocks have been used in total, each one with a different size. Each dense block is made up of a series of layers that follow the sequence: batch normalization-convolutional 3D-batch normalization-convolutional 3D as it can be seen in Figure 1(Section B). The DenseNet model has been chosen because of the way in which it concatenates the feature maps, which is simpler than models like Inception [ 46 ] or ResNet [ 47 ]. Its architecture is more robust and requires a very low number of filters and parameters to achieve high efficiency, unlike other models [ 45 ]. It has been used successfully for the treatment of biomedical images [ 48 , 49 ], showing better results than other architectures. From the point of view of detecting violence in videos, the DenseNet model extracts the necessary features to perform the detection task more efficiently than other models in terms of the number of trainable parameters and training time and inference time. 3.4. Multi-Head Self-Attention The multi-head self-attention [ 39 ] is an attention mechanism that links different positions of a single sequence and thus generates a representation of it focusing on the most relevant parts of the sequence. It is based on the attention mechanism that was first in- Electronics 2021,10, 1601 7 of 17 troduced in 2014 [ 37 ]. Self-attention has had remarkable success being used in natural processing language tasks and text analysis [ 50 , 51 ], to determine which other words are relevant while the current one is being processed. Multi-head self-attention is a layer that essentially applies multiple self-attention mechanisms in parallel. The procedure is based on projecting the input data applying different linear projections learned from the same data. Then the attention mechanisms are applied to each one of them concatenated. We use the multi-head self-attention layer in combination with the recurrent convolutional bidirectional layer, to determine which relevant elements are common in both temporal directions, generating a weighted matrix that holds more relevant past and future information simultaneously. The parameters of the multi-head self-attention layer are the next: number of heads h= 6, dimension of questions d_q= 32, dimension of values d_v= 32 and dimension of keys d_k= 32. An ablation study is carried out in Section 5to test the improvements provided by this layer. In the task of detecting violent actions in videos, multi-head self-attention mechanisms establish new relationships between features, determining which of them are the most important in determining whether or not it is a violent action. 3.5. Bidirectional Convolutional LSTM 2D A bidirectional recurrent cell is a recurrent cell with two states. The two states are the past state (backwards) and the future state (forward). In this way, the output layer to which the bidirectional recurrent layer is connected can obtain information on both states simultaneously. The principle of bidirectional recurrent layers is to divide the neurons of a regular recurrent layer in two directions, positive and negative time directions. This is especially useful in the context of detecting violent actions in videos, as performance improvement can be gained by having the ability to look back. In a standard recurrent neural network (RNN), temporal features are extracted but spatial ones are lost. To avoid this problem, fully connected layers are replaced with convolutional ones. This is how the ConvLSTM layer can learn the spatio-temporal features of a video, allowing us to take full advantage of the spatio-temporal information that arises from the correlation between convolution and recurrent operations. BiConvLSTM is an enhancement on ConvLSTM that allows to analysis of sequences forward and backward in time simultaneously. A BiConvLSTM layer can access information in both directions of a video’s timeline. In this way, a better overall understanding of the video is achieved. The bidirectional convolutional LSTM 2D module is known in the field of video and image classification to extract spatio-temporal features, being used successfully in other proposals. Zhang et al. [ 52 ] used it to recognize gestures in videos and classify them by learning long-term spatio-temporal features. In [ 53 ] it is used to classify hyperspectral images, but instead of learning spatial and temporal features, the latter are replaced by spectral features. 3.6. Classifier The classifier part is made up of four fully connected layers. The number of nodes in each layer, ordered sequentially, is 1024, 128, 16 and 2. Hidden layers use the ReLu activation function. The output of the last layer output is a binary predictor that employs the Sigmoid activation function classifying the input into Violence and Non-Violence categories. 4. Data In the experiments, the four datasets that appear the most in studies on the detection of violent actions were selected. They are widely accepted and used to compare approaches to detecting violent behaviour. These datasets are as follows: • Hockey Fights (HFs) [ 17 ] a collection of hockey games from the USA’s National Hockey League (NHL) that includes fights between players. Electronics 2021,10, 1601 8 of 17 • Movies Fights (MFs) [ 17 ] a 200-clip collection of scenes from action movies that includes fight and non-fight events. • Violent Flows (VFs) [ 25 ] a collection of videos that include violence in crowds. It differs from the previous ones in that it is focused on crowds and not on person-to-person violence but it is interesting to check the versatility of the model. • Real Life Violence Situations (RLVSs) [ 54 ] a collection of 1000 violence and 1000 nonviolence videos collected from YouTube, violence videos contain many real street fight situations in several environments and conditions. Additionally, non-violence videos are collected from many different human actions like sports, eating, walking, etc. All four datasets had the same labels, were balanced and were split in a 80–20% ratio for training and testing respectively. Table 1shows the information of each dataset. The datasets cover indoor and outdoor scenarios as well as different weather conditions. The Hockey Fights dataset only shows indoor scenarios, specifically an ice hockey arena. The Movies Fights dataset scenes vary between indoor and outdoor scenes, but none of them show adverse weather conditions. The Violent Flows dataset focuses on mass violence that always occurs outdoors. Some scenes contain adverse weather conditions such as rain, fog and snow. The Real Life Violence Situations dataset shows a great variability of indoor and outdoor scenarios ranging from the street to different venues for sporting events, different rooms in a house, stages for music shows, etc. It also shows different adverse weather conditions, although the most frequent is rain. Table 1. Hockey Fights, Movie Fights, Violent Flows and Real Life Violence Situations dataset features. Dataset Number of Clips Average Frames Hockey Fights [17] 1000 50 Movies Fights [17] 200 50 Violent Flows [25] 246 100 Real Life Violence Situations [54] 2000 100 5. Experiments This section summarizes the training methodology and proposes an ablation study to test the importance of the self-attention mechanism. In addition, cross-dataset experimentation is proposed to evaluate the level of generalization of violent acts. All the experiments are available through Github (https://github.com/FernandoJRS/violencedetection-deeplearning) (accessed on 2 July 2021). 5.1. Training Methodology For the model, the weights of all neurons were randomly initialized. The pixel values of each frame were normalized to be in the range of 0 to 1. The number of frames of the input video was the average of the frames of all the videos in the dataset. If an input video had more frames than the average, the excess frames were eliminated, if there were fewer frames than the average, the last frame was repeated until the average was reached [ 55 , 56 ]. Frames were resized to 224 ×224 ×3, the standard size for Keras pre-trained models. A base learning rate of 10 −4 , a batch size of 12 videos and 100 epochs were selected. Weight decay was initiated to 0.1. Furthermore, the default configuration of the Adam optimizer was used. The Binary Crossentropy function was chosen as the loss function and the Sigmoid function as the activation function for the last layer of the classifier. To perform the experiments, the CUDA toolbox was used to extract deep features on Nvidia RTX 2070 Super GPU. The operating system was Windows 10 using Intel Core i7. The performances with the datasets were carried out using a random permutation cross-validator method. A 5-fold cross-validation was chosen. The experiments were carried out with two kinds of inputs, the first batch with optical flow and a second batch with adjacent frames subtraction, which we called pseudo-optical Electronics 2021,10, 1601 9 of 17 flow. Both entries implicitly represented the temporal dimension but in different ways. This was to find out which kind of input obtained the best results for our model. The pseudo-optical flow was obtained by subtracting two adjacent frames [ 57 ]. Given a sequence of frames, (f0. . . fk) a matrix subtraction was applied to each pair of adjacent frames, ∀n<k∈[0, k]:sn=fn−fn+1 . With this method, any difference between the pixels of two adjacent frames was represented. Figure 2shows the transformation of three violent scenes into their respective optical flows and pseudo-optical flows. Figure 2. Transformation of sequences of two consecutive frames, of four violent scenes for each dataset (columns 1 and 2), in their dense optical flow (column 3) and adjacent frame subtraction (column 4). The color in (column 3), is for better visualization. The main difference between both methods is that the pseudo-optical flow also represents those pixels that did not move between two consecutive frames. If the same pixel in both frames did not change in value, by subtracting both frames that pixel turned black, regardless of whether it moved. For the optical flow method, the pixels that turn black are those that have not moved between two consecutive frames. 5.2. Metrics To measure the efficiency of our model, the following set of metrics was used: • Train accuracy: The number of correct classifications of the model on examples it was constructed on divided by the total amount of classifications. • Test accuracy: The number of correct classifications of the model on examples it has not seen divided by the total amount of classifications. Electronics 2021,10, 1601 10 of 17 • Test inference time: The average latency time when making predictions atomically on the test dataset. 5.3. Ablation Study For the ablation study, a double experiment was proposed to test the importance of the self-attention mechanism and to determine how the use of optical flow and pseudooptical flow affected the results. Avoiding the self-attention mechanism involved the direct connection of the DenseNet to the bidirectional convolutional LSTM. 5.4. Cross-Dataset Experimentation Cross-dataset experimentation is intended to determine whether a model trained on one dataset can correctly evaluate instances of another dataset. The main objective is to determine if the concept of violence learned by the model is general enough to be able to correctly evaluate other datasets. Two kinds of cross-dataset setups were tested. In the first one the model was trained with one of the datasets and evaluated with the rest. In the second one the model was trained with a combination of three datasets and evaluated with the remaining one. 6. Results In this section results obtained from the ablation study, the comparison with state of the art and the cross-dataset experimentation are shown. 6.1. Ablation Study Results Although a more powerful backbone network was used than in previous work, we considered it interesting to check how the performance improved by changing the network input (optical flow and pseudo-optical flow) and using the attention mechanism. When comparing the two versions of the proposed model (optical flow vs pseudo-optical flow input) with the ones without the self-attention module two main advantages were obtained: better accuracy and shorter inference time. As can be seen in Table 2, both accuracy and inference time were consistently better when the attention module was used. Inference time with attention mechanisms was shorter than without them because bidirectional convolutional recurrent layer operations took longer to apply to featured maps resulting from a convolutional network than to concatenated sequences of attention layers. Luong et al. [58] showed how attention mechanisms could reduce inference time. It can be seen that when the datasets were composed of clips of 50 frames on average (HF and MF), the differences in accuracy were small, but when the clips were 100 frames on average (VF and RLVS), the accuracy of the model with self-attention outperformed the others by 2 points. Inference time was consistently shorter on each dataset and went from 4% (VF) to 16% (HF) less than without attention. Finally, comparing the results obtained in each dataset in the pseudo-optical flow without attention and the optical-flow with attention versions, relevant gains were observed in all cases, except in the Movies Fights dataset (very simple). Specifically, the gains were 2 points in HF, 4.4 points in VF and 3.4 points in RLVS. Table 2. Ablation study of architecture Bi-Dense attention and Bi-Dense without attention. Dataset Input Test Accuracy (with Attention) Test Accuracy (without Att.) Test Inference Time (with Attention) Test Inference Time (without Att.) HF Optical Flow 99.20 ±0.6% 99.00 ±1.0% 0.1397 ±0.0024 s 0.1626 ±0.0034 s HF Pseudo-Optical Flow 97.50 ±1.0% 97.20 ±1.0% MF Optical Flow 100.00 ±0.0% 100.00 ±0.0% 0.1916 ±0.0093 s 0.2019 ±0.0045 s MF Pseudo-Optical Flow 100.00 ±0.0% 100.00 ±0.0% VF Optical Flow 96.90 ±0.5% 94.00 ±1.0% 0.2991 ±0.0030 s 0.3114 ±0.0073 s VF Pseudo-Optical Flow 94.80 ±0.5% 92.50 ±0.5% RLVS Optical Flow 95.60 ±0.6% 93.40 ±1.0% 0.2767 ±0.020 s 0.3019 ±0.0059 s RLVS Pseudo-Optical Flow 94.10 ±0.8% 92.20 ±0.8% Electronics 2021,10, 1601 11 of 17 6.2. State of the Art Comparison After the experiments were carried out, better results were observed for the input of optical flow than with that of pseudo-optical flow. The results of training and testing procedure of one iteration for each dataset and for each type of input to the model are shown in Table 3. Table 3. Performance comparison for one iteration of our model for Hockey Fights, Movies Fights, Violent Flows and Real Life Violence Situations datasets. Dataset Input Training Accuracy Training Loss Test Accuracy Violence Test Accuracy Non-Violence Test Accuracy HF Optical Flow 100% 1.20 ×10−599.00% 100.00% 99.50% HF Pseudo-Optical Flow 99% 1.35 ×10−597.00% 98.00% 97.50% MF Optical Flow 100% 1.18 ×10−5100% 100% 100% MF Pseudo-Optical Flow 100% 1.19 ×10−5100% 100% 100% VF Optical flow 98% 1.50 ×10−497.00% 96.00% 96.50% VF Pseudo-Optical Flow 97% 2.94 ×10−495.00% 94.00% 94.50% RLVS Optical Flow 97% 3.10 ×10−496.00% 95.00% 95.50% RLVS Pseudo-Optical Flow 95% 7.31 ×10−494.00% 93.00% 93.50% As can be seen, the optical flow allowed the spatio-temporal dimension of the videos to be highlighted better than the pseudo-optical flow and favoured the training to achieve a greater decrease of the loss function. When comparing our proposal with the state of the art (Table 4), it is observed that it outperformed previous studies, even those that did not use cross-validation, maintaining a low number of parameters. The best results were obtained for HF and MF where personto-person violence was present. In particular, the MF dataset was the most homogeneous and the least challenging. The model also worked very well with violence in crowds. Table 4. State of the art for HF, MF, VF and RLVS datasets. OF stands for optical flow. Model HF test Accuracy MF Test Accuracy VF Test Accuracy RLVS Test Accuracy Train−Test Validation Params VGG13-BiConvLSTM [34] 96.54 ±1.01% 100 ±0% 92.18 ±3.29% −80 −20% 5 fold cross − Spatial Encoder VGG13 [34] 96.96 ±1.08% 100 ±0% 90.63 ±2.82% −80 −20% 5 fold cross − FightNet [38] 97.00 ±0% 100 ±0% − − 80 −20% hold-out − Three streams + LSTM [31] 93.90 ±0% − − − − − − AlexNet - ConvLSTM [32] 97.10 ±0.55% 100 ±0% 94.57 ±2.34% −80 −20% 5 fold cross 9.6 M Hough Forest + CNN [22] 94.60 ±0.6% 99.00 ±0.5% − − 80 −20% 5 fold cross - FlowGatedNetwork [59] 48.10 ±0% 59.00 ±0% 50.00 ±0% −80 −20% hold-out 5.07 K Fine-Tuning Mobile-Net [60] 87.00 ±0% 99.50 ±0% − − 80 −20% hold-out − Xception BiLSTM Attention 10 [36] 97.50 ±0% 100 ±0% − − 80 −20% hold-out 9 M Xception BiLSTM Attention 5 [36] 98.00 ±0% 100 ±0% − − 80 −20% hold-out 9 M SELayer-C3D [61] 99.00 ±0% − − − 80 −20% hold-out − Conv2D LSTM [62] 94.50 ±0% − − 92.00 ±0% 80 −20% hold-out − ViolenceNet Pseudo-OF 97.50 ±1.0% 100 ±0% 94.80 ±0.5% 94.10 ±0.8% 80 −20% 5 fold cross 4.5 M ViolenceNet OF 99.20 ±0.6% 100 ±0% 96.90 ±0.5% 95.60 ±0.6% 80 −20% 5 fold cross 4.5 M In the HF dataset, our model generalized very well the amount of movement of the elements given that hockey is a sport where the players are in constant movement and sometimes there is physical contact. The temporal features that our model learned were based on energetic movement when it occurred. The difficulty of generalizing the concept of violence in the VF dataset was different from that in the HF dataset. The videos from the VF dataset showed violent acts at mass events such as demonstrations, concerts, etc. In massive events, many actions occurred simultaneously. The viewpoints were far from the scene and thus captured many people appearing in low resolution. Many actions in a single video from a far viewpoint made them seem small and made it harder to distinguish if a touch action was violent or not, even by a person. Furthermore, the context of the mass event included specific situations such as catching a golf ball by a crowd (that could seem like the beginning of a fight). It Electronics 2021,10, 1601 12 of 17 was also difficult to generalize the concept of violence with the RLVS dataset as it is very heterogeneous. Unlike the other three datasets, the RLVS scenes were not topic-specific. The heterogeneity of the dataset was more visible in the non-violence category, where the actions of each scene were very different from each other. The test accuracy obtained for the different datasets reached the state of the art. Our model was in a good position compared to others. Before our proposal, other models were applied for the HF, MF, VF and RLVS datasets. For the MF dataset, several models reached a test accuracy of 100 ± 0% which made them unbeatable. For the VF dataset, our proposal outperformed the closest one, [ 32 ], by more than 2 points. The RLVS dataset was only tested with a model prior to ours, [54]. This model had a test accuracy for the RLVS dataset of 92.00% using hold-out validation. Our model achieved a value of 95.60%, again improving on the state of the art. The test accuracy values were slightly higher with the optical flow input than with the pseudo-optical flow input. This occurred with all datasets except the MF dataset, which had the same test accuracy value for both. Even comparing our model with those who used a hold-out validation methodology, it can be seen that our method improved the state of the art. Another remarkable advantage of our proposal was the number of trainable parameters used, less than for the rest of the models for which these data were available. This is due to the Dense architecture and its feature map concatenation method. The only model with a number of trainable parameters lower than the proposed one was FlowGatedNetwork − 3 DCNN −Flow −RGB [ 59 ], however, this model was not relevant because its test accuracy did not exceed 60% for any of the datasets. 6.3. Cross-Dataset Experimentation Results After performing the cross-dataset experiments, two logical facts were observed: on the one hand, there was a slight correlation between more heterogeneous datasets and a better generalization of the concept of violent actions. On the other hand, the experiments performed by training the model with unions of different datasets showed a better generalization of the concept of violence than in those experiments in which the model was trained with a single dataset. It can also be observed that for pairs of datasets very different in the type of context of violence, the results did not show effective generalization. An example of this was the pair MF and VF, where the contexts were clearly different (MF violence between two people or a very small group and VF focuses on mass violence), obtaining a very low accuracy (52.32%) in the MF -> VF direction and somewhat higher in the opposite direction (60.02%), given that the length and variability of VF was higher. The best result obtained during cross-dataset experimentation was the one where the model was trained with the combination of the HF, RLVS and VF datasets, achieving a test accuracy of 81.51%, but very low compared to cross-validation using the same dataset (100%). Finally, experiments that included the RLVS dataset for model training obtained better results than experiments in which the model was not trained with it. The RLVS dataset was the largest and most heterogeneous of the four. The results in testing for five iterations for each type of input are shown in Table 5. Electronics 2021,10, 1601 13 of 17 Table 5. Cross-dataset experiment results. Dataset Training Dataset Testing Test Accuracy Optical Flow Test Accuracy Pseudo-Optical Flow HF MF 65.18 ±0.34 64.86 ±0.41 HF VF 62.56 ±0.33 61.22 ±0.22 HF RLVS 58.22 ±0.24 57.36 ±0.22 MF HF 54.92 ±0.33 53.50 ±0.12 MF VF 52.32 ±0.34 51.77 ±0.30 MF RLVS 56.72 ±0.19 55.80 ±0.20 VF HF 65.16 ±0.59 64.76 ±0.49 VF MF 60.02 ±0.24 59.48 ±0.16 VF RLVS 58.76 ±0.49 58.32 ±0.27 RLVS HF 69.24 ±0.27 68.86 ±0.14 RLVS MF 75.82 ±0.17 74.64 ±0.22 RVLS VF 67.84 ±0.32 66.68 ±0.22 HF + MF + VF RLVS 70.08 ±0.19 69.84 ±0.14 HF + MF + RLVS VF 76.00 ±0.20 75.68 ±0.14 HF + RLVS + VF MF 81.51 ±0.09 80.49 ±0.05 RLVS + MF + VF HF 79.87 ±0.33 78.63 ±0.01 6.4. Detection Process in CCTV The ViolenceNet model allowed classifying videos between the categories of violence and non-violence. In a CCTV system, our model worked with short video fragments of equal length. Each of these fragments was preprocessed, applying the dense optical flow algorithm that generated the input to the ViolenceNet model that was in charge of classifying these fragments into violence or non-violence, as shown in Figure 3. ViolenceNet Non Violence Violence Camera Input Video Optical Flow Video Figure 3. CCTV scheme for detection of violent scenes. In a CCTV system, our model receives the optical flow of the fragments provided by a camera and classifies each of them. The red box represents violent segments and the blue box represents non-violent segments. 7. Discussion The developed model for the detection of violent actions has been more efficient in improving the state of the art than in generalizing the concept of violent actions. This is demonstrated in the results presented in the previous section. However, some results obtained in cross-dataset experiments show that it is possible to improve the generalization of violence if large and heterogeneous datasets with a large number of instances that contemplate different contexts and scenarios are used. Datasets with few instances like MF or HF with a single type of context are not effective in generalizing the concept of violence regardless of whether the input is optical flow or pseudo-optical flow. to adversarial. CrimeNet no solo supera ampliamente los trabajos previos, sino que también reduce prácticamente a cero los falsos positivos. Las pruebas realizadas en los cuatro conjuntos de datos más desafiantes relacionados con la violencia (tanto binarios como de múltiples clases) demuestran la efectividad de CrimeNet, mejorando el estado del arte entre 9.4 y 22.17 puntos porcentuales en términos de AUC ROC, dependiendo del conjunto de datos. Además, presentamos un estudio de generalización en nuestro modelo, entrenándolo y evaluándolo en diferentes conjuntos de datos. Los resultados obtenidos muestran que CrimeNet supera a los métodos competidores con una ganancia de entre 12.39 y 25.22 puntos porcentuales, demostrando una robustez notable. Este avance ha sido posible gracias al apoyo parcial de los proyectos DISARM (Grant n. PDC2021-121197) y HORUS (Grant n. PID2021-126359OB-I00), financiados por MCIN/AEI/310.13039/501100011033 y por la “Unión Europea NextGenerationEU/PRTR”. Neural Networks 161 (2023) 318–329 Contents lists available at ScienceDirect Neural Networks journal homepage: www.elsevier.com/locate/neunet CrimeNet: Neural Structured Learning using Vision Transformer for violence detection Fernando J. Rendón-Segadora,∗, Juan A. Álvarez-Garcíaa, Jose L. Salazar-Gonzáleza, Tatiana Tommasib aDpto. de Lenguajes y Sistemas Informáticos, Universidad de Sevilla, Spain bPolitecnico di Torino & Italian Institute of Technology, Italy article info Article history: Received 16 June 2022 Received in revised form 20 December 2022 Accepted 30 January 2023 Available online 2 February 2023 Keywords: Deep learning Neural Structured Learning Vision Transformer Violence detection Adversarial Learning abstract The state of the art in violence detection in videos has improved in recent years thanks to deep learning models, but it is still below 90% of average precision in the most complex datasets, which may pose a problem of frequent false alarms in video surveillance environments and may cause security guards to disable the artificial intelligence system. In this study, we propose a new neural network based on Vision Transformer (ViT) and Neural Structured Learning (NSL) with adversarial training. This network, called CrimeNet, outperforms previous works by a large margin and reduces practically to zero the false positives. Our tests on the four most challenging violence-related datasets (binary and multi-class) show the effectiveness of CrimeNet, improving the state of the art from 9.4 to 22.17 percentage points in ROC AUC depending on the dataset. In addition, we present a generalisation study on our model by training and testing it on different datasets. The obtained results show that CrimeNet improves over competing methods with a gain of between 12.39 and 25.22 percentage points, showing remarkable robustness. ©2023 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). 1. Introduction Violence detection is a very important functionality in public or private security. In smart cities, more and more surveillance cameras are being installed that enable several use cases such as traffic management, infraction or weapon detection (Salazar González, Zaccaro, Álvarez-García, Soria Morillo, & Sancho Caparrini,2020). In addition, in schools, hospitals, shopping centres, and other buildings they are also generally used as a dissuasion and in the worst case to identify criminals once a crime has been committed. This growing number of cameras requires sufficient human resources to control the volume of video they generate, however the number of cameras needing attention is greater than the number of staff. Furthermore, after 20 min of monitoring a CCTV system, operators’ attention spans are considerably reduced (Ainsworth,2002;Velastin, Boghossian, & Vicencio-Silva,2006). The current difficulties in tackling this problem become even more evident when considering that security is an issue of international concern which scales up from single cities to the World Wide Web. The amount of violent audiovisual content circulating on the network is excessive. To ∗Corresponding author. E-mail addresses: [email protected] (F.J. Rendón-Segador), [email protected] (J.A. Álvarez-García), [email protected] (J.L. Salazar-González), [email protected] (T. Tommasi). such an extent that the operators in charge of controlling and filtering videos of this type on social networks end up with mental health problems due to over-exposure to violent content.1 Despite this, video surveillance systems are still being manned by humans because the number of false positives is not acceptable in production environments and a human in the loop is still necessary. When analysing the literature, we found that for the most challenging datasets the state of the art does not reach 90% accuracy (Lv et al.,2021). To overcome this issue, we aim at designing a robust and sufficiently accurate model to minimise the number of false positives, making it possible to use a video violence detector in real environments. This work covers the video detection of all types of violent events, both visually intentional actions such as a fight between people, as well as unintentional acts such as an explosion. Moreover, we go beyond standard anomaly detection which differentiates violent from non-violent events: we target a model able to differentiate between different types of violence. This is a complex problem since many types of violent actions can be similar even if their classes are different (there are categories such as abuse, arrest or assault that can be confused with each other) or they can even occur in the same video. To tackle these 1https://www.bloomberg.com/news/articles/2021-12-24/tiktok-sued-bycontent-moderator-traumatized-by-graphic-videos https://doi.org/10.1016/j.neunet.2023.01.048 0893-6080/©2023 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/bync-nd/4.0/). F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 challenges we propose to combine the powerful self-attention learning paradigm with Neural Structured Learning (NSL). The former is obtained by exploiting the most recent Vision Transformer deep architecture (ViT, Dosovitskiy et al.,2021). The latter leverages the relation among neighbouring samples during training. Specifically, it provides a regularisation effect by biasing the network to learn similar hidden representations for close instances. Our key contributions can be summarised as follows: •We introduce CrimeNet: a deep model that combines adversarial NSL with ViT for violent activity recognition in videos. Up to our knowledge this is the first time that NSL is used with Transformers rather than with standard convolutional neural networks, and also the first time that it is applied for video recognition. •CrimeNet improves the state of the art for violence detection on four datasets by reaching an accuracy over 99.98%, with an advantage ranging from 9.4 to 22.17 percentage points in ROC AUC over its competitors. A detailed ablation shows that NSL provides and improvement of 9.55% points in ROC AUC over the use of only ViT as the architecture of the model. •We present an extensive cross-dataset analysis and show the generalisation abilities of CrimeNet. Despite the challenging setting of training and testing on different datasets, CrimeNet advances the state of the art from 12.39 to 25.22 percentage points with respect to existing methods that are affected by the domain shift. This paper is organised as follows: Section 2provides a brief survey of the state of the art of the problem. Section 3describes in detail each of the datasets used in this work. Section 4provides in-depth details on the proposed model architecture. Section 5 summarises the types of experiments proposed in this work. Then, in Sections 6and 7the results obtained from the experimentation are shown and described. Finally Section 8summarises the relevance of the results obtained and the possible lines of progress for future development. 2. Related works In this section, a review of the state of the art on video detection of violence is carried out. Furthermore, several Neural Graph Learning proposals are shown, where NSL is a specific case, applied to computer vision problems. 2.1. Deep learning for violence behaviour detection in videos Crowd behaviours analysis (Li, Chen, Nie & Wang,2017a, 2017b;Sadeghian, Alahi, & Savarese,2017) and detection of violent actions in videos are well-known and well-studied research fields (Bermejo Nievas, Deniz Suarez, Bueno García, & Sukthankar,2011;Deniz, Serrano, Bueno, & Kim,2014), however, since 2014 when the first paper using a deep learning approach (Ding, Fan, Zhu, Feng, & Jia,2014) appeared, violence detection has advanced by leaps and bounds. Specifically, for some datasets (Bermejo Nievas et al.,2011;Hassner, Itcher, & Kliper-Gross,2012), based on short clips and labels for the whole clip, there exist approaches reaching up to 100% accuracy (Rendón-Segador, Álvarez-García, Enríquez, & Deniz,2021). Given that these datasets were not sufficiently challenging, researchers created new testbeds with hours of CCTV camera recordings: the videos are longer and often annotated at framelevel. In addition, some of them have moved from 2 classes (violence and non-violence) to multiple classes distinguishing the typology of violence in the video. The first paper to introduce this new family of datasets was (Sultani, Chen, & Shah,2018) in 2018: the UCF Crime collection incorporates the greatest variety of violence classes (14 types). This work also presented a model called DMIL Ranking, in which anomaly detection is approached as a regression problem considering video segments as instances in ‘‘multiple instances learning’’ (MIL) and using a ranking-based loss function to evaluate a fully connected neural network. The NTU CCTV Fights dataset (Perez, Kot, & Rocha,2019) was presented in 2019: it contains only long violent videos, labelled at frame-level with violent or non-violent classes. This dataset has been evaluated using different known feature extractors such as Two-stream Convolutional Neural Network (CNN) or 3D CNN and different classifiers such as End-to-End CNN, Long–short Time Memory (LSTM) and Support Vector Machine (SVM). Another method combining the spatio-temporal feature extractor of ResNet 3D (Dubey, Boragule & Jeon,2019) and the loss function of Sultani et al. (2018) was also published that year, improving the results for the UCF Crime dataset. Zhong et al. (2019) evaluated the same dataset using a convolutional graph to clean and refine the classifier based on Temporal Segment Networks (Wang et al.,2018). New work focused on anomaly detection emerged in 2020. Degardin (2020) provided a new dataset, UBI Fights, and proposed an architecture based on the Gaussian mixture model (GMM) for the detection of abnormal events in videos applied under the weakly supervised learning paradigm. Also noteworthy is the work of Kamoona, Gosta, BabHadiashar, and Hoseinnezhad (2020) in the weakly supervised setting, using an encoder–decoder architecture (DMIL AutoEncoder). Furthermore in the same year, Wu et al. (2020) released the new dataset XD Violence, with 7 classes, the largest number of videos (4754), and hours (217), also incorporating sound. The authors presented a method that exploits visual and audio feature extractors whose output is mixed and provided to graph neural networks to detect shortand long-range temporal relationships. In 2021 Tian et al. (2021) proposed an approach for anomaly detection in videos where a learning function recognises positive instances on the basis of the feature magnitude learning function and using self-attention mechanisms (Vaswani et al.,2017). In the same year, Lv et al. presented a study in which they propose a weakly supervised anomaly localization (WSLA) method that measures variations in both spatial and temporal contexts. Their results mark the current state of the art for the UCF Crime dataset. In the work by Chang, Li, Shen, Feng, and Zhou (2021), frame anomalies are detected starting from a single binary annotation at the video level. Based on the extracted visual features, attention mechanisms are used to refine the classification of anomalous instances. Dubey, Boragule, Gwak and Jeon (2021) proposed a model that addresses context-dependency by analysing motion and appearance features. They use a 3D ResNet network to extract spatiotemporal and motion feature sets which are then fused and provided as input to a network that learns context-dependencies in a weakly supervised manner using multiple classification measures (MRM). Another noteworthy study is that of Feng, Hong, and Zheng (2021) in which they developed a multi-instance self-training framework (MIST) to efficiently refine task-specific discriminative representations with only video-level annotations. The model is composed of a multi-instance pseudo-label generator and a self-guided attention function encoder that aims to automatically focus on anomalous regions in frames while extracting task-specific representations. Finally, Degardin and Proença (2021), presented an iterative learning framework composed of two expert systems working in 319 F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 the weakly supervised and self-supervised paradigms. This work combines 3D convolutional neural networks for both paradigms with a Bayesian network that is responsible for performing data augmentation. To provide some reference background, the next section reviews state of the art of Neural Graph Learning, and general cases for NSL and ViT. 2.2. Applied neural graph learning Graph-based neural learning has been applied to several tasks related to human action recognition. In 2019 Shi, Zhang, Cheng, and Lu (2019) proposed to represent skeleton data in a directed acyclic graph based on the kinematic dependence between joints and bones. In 2021, Xu and Takano (2021) presented a convolutional graph neural network to estimate 3-D human pose in which the input data were structured using an hourglass graph. In the same year, a model to classify sports videos was presented in Gao, Cai, and Liu (2021): it consists of a convolutional model designed with attention mechanisms for assigning weights to neighbouring nodes, combined with a third-order hourglass graph used to structure the features of the videos. Graph-based learning has been also used in Yin, Shen, Gao, Crandall, and Yang (2021) to detect 3D objects in videos. In this case, the data were encoded through a grid message passing network (GMPNet). Considering each grid as a node, the data were structured using a k-NN network and the model was a spatio-temporal Transformer-GRU. There are other fields of research that have experimented with NSL and adversarial learning with significant results, such as Ren, Wang, Zhang, and Chang (2020) for fake news detection through social networks. It has also been applied to protect against adversarial attacks using perturbed data (Jin et al.,2020), achieving significantly better performance compared to state of the art defense methods. 2.3. Applied vision transformer ViT (Dosovitskiy et al.,2021) profits the Transformer (Vaswani et al.,2017) potential, avoiding inductive bias such as translation invariance and locally restricted receptive field in images. To do it, it splits an image in a sequence of patches, flattens them, produces linear embeddings, adds positional embedding to know where is located each patch in the original image, and feeds this sequence as an input to a standard transformer encoder (composed by a multi-head attention layer (Vaswani et al.,2017) that allows the model to jointly attend to information from different representation subspaces at different positions) as it can be seen in Fig. 4. The success of Transformers, ViT, and their variations (Khan et al.,2022) is beyond doubt, and have improved the state of the art in many areas such as frame synthesis (Liu et al.,2020), action recognition (Girdhar, Carreira, Doersch, & Zisserman,2019), or object detection in videos (Chen, Cao, Hu, & Wang,2020). Our approach differs from previous proposals although being inspired by the use of graph neural networks already leveraged in Wu et al. (2020) and Zhong et al. (2019). As we will describe in the following, we propose a new approach based on supervised NSL (Bui, Ravi, & Ramavajjala,2018;Gopalan et al.,2021) and ViT (Dosovitskiy et al.,2021). To our knowledge, the NSL paradigm is used here for the first time for violence recognition in videos. Table 1 Information on datasets used. Dataset №Items №Classes №Hours NTU CCTV Fights (Perez et al.,2019) 1000 2 17.68 UBI Fights (Degardin,2020) 1000 2 80 XD Violence (Wu et al.,2020) 4754 7 217 UCF Crime (Sultani et al.,2018) 1900 14 128 3. Datasets For our work, we focus on four video datasets recording violent events. They all contain 1000 videos or more, each one ranging from hundreds to thousands of frames. A summary of the datasets’ information is in Table 1, while the following list provides further details: •NTU CCTV Fights (Perez et al.,2019) is a dataset containing 1000 videos depicting real-world fights. Of this dataset, 280 videos are recorded from CCTV and 720 from other sources such as mobile cameras or drones, containing different types of fights, ranging from 5 s to 12 min, with an average duration of 2 min. •UBI Fights (Degardin,2020) is a large-scale dataset of 80 h of video fully labelled at frame-level. It consists of 1000 videos, where 216 videos contain a fight event and 784 are normal everyday situations. All unnecessary video segments (e.g., video introductions, news, etc.) that could disrupt the learning process were removed. The title of the videos contains indicators related to the type of the respective video. The dataset is divided into binary classes: violence and normal. •XD Violence (Wu et al.,2020) is a large-scale dataset with a total duration of 217 h, containing 4754 untrimmed videos with audio signals and video-level tagging. The dataset is divided into the following seven anomalous categories: abuse, car accident, explosion, fight, shooting, and riot. •UCF Crime (Sultani et al.,2018) is a dataset consisting of long untrimmed surveillance videos covering 14 real-world violent classes, including abuse, arrest, arson, assault, traffic accident, burglary, explosion, fight, robbery, burglary, shooting, theft, shoplifting, and vandalism. 4. Model achitecture This section shows the type of model and architecture used to address the problem. An in-depth definition of the model and the adaptations applied for our use case is provided. 4.1. Pre-processing 4.1.1. Optical flow Since our model analyses the video frame by frame, including temporal information in each frame is critical. We use optical flow for this purpose (Farnebäck,2003): it takes two adjacent frames and represents in an image the amount of pixel variation caused by the observed movements. Of course, the parts of a frame that move together will correspond to pixels with the same intensity as in the example of Fig. 1. The state of the art demonstrates that the use of optical flow as input in violence detection typically improves the use of RGB input (Mahmoodi & Salajeghe,2019;Rendón-Segador et al.,2021; Zhou, Ding, Luo, & Hou,2018). 320 F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 Fig. 1. Sequence of frames in RGB format and their corresponding optical flow. In this frame sequence, we go from a normal event to a violent event (explosion). Images from the XD-Violence dataset (Wu et al.,2020). 4.1.2. Adversarial neighbours As it will be seen when describing the model’s details (Section 4.3), NSL is used, a new learning paradigm to train neural networks by using structured signals in a graph. This assumes that the model receives two inputs: the RGB frames, in our case transformed by an optical flow algorithm, and a similarity graph. This graph, or structured signal, is used to represent relationships between samples. The similarity graph regularises the training of a neural network, forcing the model to learn accurate predictions by minimising the supervised loss function while maintaining the input structural similarity by reducing the loss function of the neighbouring node. The similarity graph is generated from the training examples using a graph builder. Each entry is assumed to have an ID and an embedded vector as features. On the one hand, the ID uniquely records and identifies each instance; on the other hand, the embedded vector is assumed to capture the essence of each example by representing it as a list of floating-point values. The graph generator compares the embedded vectors of all input pairs. The degree of similarity between any two samples is calculated as the cosine similarity of their embedded vectors. The brute-force approach to constructing a similarity graph from instance embeddings is O(n2), which does not scale well to large training sets. To mitigate that problem, the graph builder uses a wellknown randomisation technique called locality-sensitive hashing or LSH (Charikar,2002). To describe this approach, it is assumed that we have a bidimensional embedding vector represented by npoints plotted on a cartesian coordinate plane. The first step of the LSH process is to choose some random hyperplanes through the origin as it can be seen in Fig. 2. These hyperplanes divide the space and thus the points into discrete sections that are called LSH buckets. Although there are several buckets, the number of points in each bucket is expected to be much smaller than the complete input set. The quadratic nature of comparing all pairs within each bin has a much smaller impact on performance; comparisons within each bin result in a certain number of graph edges within the bin. This generates a series of connected components each of which corresponds to a bucket, these components are separate and are not part of an overall graph. To complete the similarity graph, the bucketing process is repeated several times, with each round of bucketing choosing to select a different number of random hyperplanes until all connected components are connected. We highlight that the similarity graph is needed only at training time, since during inference it would not be useful to regenerate this graph by adding only the samples to be inferred as the whole LSH process would be necessary again. The inference workflow will not change, simply providing predictions on new testing video frames. Obtaining the exact similarity relationship between input data is difficult: it is not trivial to evaluate the likeness of one instance to another without a predefined feature embedding that differentiates the type of violence or separates crimes from normal actions. To overcome this issue we propose to generate synthetic neighbouring instances via Adversarial Learning (Zhang, Lemoine & Mitchell,2018). In the case of image classification networks, adversarial examples are nothing more than samples to which the pixels have been modified in order to alter the predicted class. Their appearance is similar to that of the original images, but they induce model confusion causing prediction errors. Generally, the adversarial examples are obtained by exploiting the inverse gradient direction of an optimised model. We adopt this logic on the optical flow images. We start from the hypothesis that normal and violent events have different pixel clustering and variability in the amount of motion measured using optical flow, with the latter being greater in violent events. Likewise, between different types of violent events, there will be variations in the amount of motion and type of pixel clustering. Fig. 3 shows the scheme of a structured signal in which the optical flow images in Fig. 1 are grouped into two connected components αand β, assuming a binary classifier. The images and small circles are the new samples produced via adversarial learning. With this procedure, we obtain adversarial instances, one per image with respect to the total training set. The samples in the same connected components are visually similar but were created on purpose to confuse the original model, thus this data makes it more robust and able to generalise better. 4.2. Vision transformer As basic classification model that we expect to be initially fooled by the adversarial examples, we use a deep transformer network. Specifically, we consider the Vision Transformer (ViT) model (Dosovitskiy et al.,2021) which has achieved remarkable results compared to CNN, in addition to reducing the computational cost required for training and exhibiting a generally weaker 321 F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 Fig. 2. A representation of the LSH process for the construction of the similarity graph. The dashed red lines corresponding to the labels H1, H2, and H3 are the set of hyperplanes that divide the different points and group them into buckets. αand βcorrespond to two connected components of the similarity graph. The circles represent instances of the training dataset, in this paper, frames from each training video. Fig. 3. Similarity graph of optical flow frames generated using adversarial perturbation of the original input (Fig. 1). Component αare images of non-violent events and component βare images of violent events. This similarity graph is used as the second input to the model. The images inside small circles correspond to other adversarial samples. inductive bias. ViT is a model based on a Transformer architecture (Vaswani et al.,2017) initially used for image classification but later adapted to other visual tasks such as nextframe prediction (Jahanbakht, Xiang, & Azghadi,2022) or video classification (Arnab et al.,2021). It divides an image into a sequence of fixed-size fragments called patches, correctly embeds each of them, and includes their positional encoding as input to the Transformer encoder. The fragments are used similarly to the series of embedded words 322 F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 Fig. 4. ViT Architecture for frame classification, based on Dosovitskiy et al. (2021) and Vaswani et al. (2017) works. for text Transformers, and the output is a prediction label for the image. Although training a ViT model that has good performance requires from 10 million to more than 100 million images and a huge amount of time and resources (Dosovitskiy et al.,2021), the authors of Steiner et al. (2022) released more than 50000 ViT models trained under diverse settings (including patch size of 8, 16 and 32) on various datasets.2The one selected in this work is called ViT-S or DeiT-S (Touvron et al.,2021) and it is trained using a different strategy than the original ViT (Dosovitskiy et al., 2021), achieving good metrics (83.73% in ImageNet (Russakovsky et al.,2015) top-1) with reduced size (115 MB). It is pre-trained with ImageNet 21K and fine-tuned with ImageNet 1K and the patches used are of dimension 16 ×16 pixels. Given that the input frames are resized to 224 ×224, that means the ViT model ingests 142patches. It is worth considering that smaller patch sizes are computationally more expensive. For our work we keep this standard decomposition cardinality of the video frames in patches since we are mainly interested in how ViT combines with NSL: we are aware that the patch size influences the performance of the ViT models (Dosovitskiy et al.,2021;Steiner et al.,2022; Touvron et al.,2021), but this is orthogonal to our analysis and we expect that any improved choice of the patch size would inevitably further improve the observed results. ViT architecture (Dosovitskiy et al.,2021;Paul & Chen,2022), is shown in Fig. 4. We split the image into fixed-size patches and linearly embed each of them. We add positional embeddings by generating a vector sequence with which a standard Transformer encoder is fed. For classification, an extra classification element is added to the vector sequence. Classification is performed by a multilayer perceptron. 4.3. Neural structured learning In order to take advantage of the power of the similarity graph generated by means of adversarial instances, we follow the scheme proposed in NSL (Bui et al.,2018;Gopalan et al., 2021) (See Fig. 5). In it we can see that we have a pair of inputs, the frames of the training set transformed by optical flow, and a similarity graph or structured signal with instances generated by adversarial modified versions of the sample frame to which it is applied with small perturbations. The generated adversarial neighbours form a similarity graph. In the next step, the original 2https://github.com/google-research/vision_transformer instances are combined with their neighbours and serve as input to the ViT. An encoding of examples and their neighbours is obtained as output. The final regularised graph is the sum of the discriminate loss and the regularisation loss of the graph (Bui et al.,2018). Using this connection, neural networks learn to maintain the similarities between the sample and the adversarial neighbours while avoiding the confusion resulting from misclassifications, thus improving the quality and accuracy of the overall neural network. Being T=t1. . . tnthe training dataset an Y=y1. . . ymthe set of labels associated with the instances of the training dataset, through the adversarial learning process a set of neighbour nodes nis generated for each instance of the training dataset, see Eq. (1). ∀ti∈T→ti: {n1. . . nk}(1) For each subset of neighbours associated with an instance tia subgraph Hti=(V′,E′) is generated, where each of the neighbours will be connected to the instance, see Eqs. (2) and (3). V′= {ti,n1. . . nk}(2) E′= {(ti,n1). . . (ti,nk)}(3) The sum of all subgraphs forms the graph G(V,E) which we call the structured signal (Eq. (4)). G(V,E)=∑ ti∈T Hti(4) From the structured signal S=G(V,E) one can distinguish vertices containing the instances (vertices) of the pure training set Vt= {t1. . . tn}and those of the neighbouring vertices generated by adversarial learning Vn= {n1. . . nl}both of which form the input to the ViT model that finish in a multi-layer perceptron (MLP) generating a feature vector. The model returns an embedding Φ(Vt) from Vtand an embedding Φ(Vn) from Vn, the latter only during training. The feature vector Φ(Fv) is the result of the graph regularisation between the vectors Φ(Vt) and Φ(Vn), see Eq. (5). Φ(Fv)=Φ(Vt)+Φ(Vn) (5) It should be noted that the ViT model takes each of the vertices of the structured signal Sand fragments the image it contains into a sequence of patches P(v)= {pv1. . . pvn}that embeds a 323 F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 Fig. 5. Neural Structured Learning using as processing model ViT (Fig. 4). The neural network minimises two loss functions, the supervised and the adversarial loss function, which is shown in Eq. (9). Figure inspired in Juan et al. (2020) and Gopalan et al. (2021) works. vector Γand then by the process of structured learning is reembedded under the vector Φ(Fv), the actual embedding process is represented mathematically by the Eqs. (6),(7) and (8). Vt=⋃ ∀vt∈V P(vt) (6) Vn=⋃ ∀vn∈V P(vn) (7) Φ(Fv)=Φ(Γ(Vt)) +Φ(Γ(Vn)) (8) Once the vectors are embedded, the final step consists of regularising the vector Φ(Fv) and applying the adversarial loss (sparse categorical cross-entropy) function (Goodfellow, Shlens, & Szegedy,2015). Being yi∈Ythe actual label value of a vertex i and gθ(nj) the prediction of the neighbour node j, the adversarial loss function is shown in Eq. (9). ∑ nj∈V(ti) ϵ(yi,gθ(nj)) (9) 5. Experimental setting This section shows the guidelines followed during the experimentation. What kind of experimentation has been performed, how it has been performed, and the motivations related to such experimentation. All experiments have been conducted using a computer with an Intel(R) Core(TM) i7-9700F CPU 3.00 GHz, 16 GB of RAM and an NVIDIA RTX 2080 Super GPU. Details of experimentation are shown through GitHub.3 5.1. Datasets partitions In the first batch of experiments, the model is trained and tested with each of the datasets following the predefined partitions for each of them. Each of the datasets opts for a different type of labelling. These types of labelling can be grouped into three. Firstly, frame-level labelling, where each frame has a label associated with it depending on whether it is a frame that harbours violent behaviour or not, this is the case of UBI Fights. 3https://github.com/FernandoJRS/CrimeNet-ViT-NSL Secondly, interval labelling, in which for each instance of the dataset the interval in frames or seconds in which the violent action occurs is provided, this is the case of NTU CCTV Fights (intervals per second) and XD Violence (intervals per frame). And the third and last case is video-level labelling, where the entire instance of the dataset has a single label associated with it that classifies the entire video, this is the case of UCF Crime training set (the testing set is labelled using intervals per frame). We consider the finest labelling to be at frame level so we will use it for our model. For interval–labelled datasets, we take each interval and label each frame belonging to that interval based on interval class. In the case of UCF Crime training set, all frames belonging to an instance will have the same label as the instance. The datasets provide pre-defined guidelines for their division into training, validation and testing datasets. Rather than creating our own divisions, we strictly follow the guidelines provided by each of the datasets: NTU CCTV Fights uses three randomly selected partitions: 50% training, 25% validation and 25% testing; UBI Fights three fixed subsets (80%, 5% and 15%); XD Violence use 3954 videos for training and 800 for testing; UCF Crime uses 800 normal and 810 anomalous videos for training and 150 normal and 140 anomalous for testing in 4-fold cross-validation. The divisions are made at the instance level (videos) so first, the partition is made into training, validation and test subsets and then the frames are extracted from each instance and labelled. 5.2. Preparation of ablation study data The second batch of experiments aims to see the contributions of NSL to a deep learning model, in this case, a ViT. For this purpose, an ablation study is performed where NSL is eliminated, including the similarity graph using adversarial learning, and replaced by supervised learning, only standard optical flow frames are used as input of the ViT. In this ablation study, the same type of experimentation is performed as with the first batch (following the predefined partitions), substituting one type of learning (ViT +NSL) for another (only ViT). The previous and this experiments are called In-domain experiments. 5.3. Cross-datasets data preparation In the last batch of experiments, the objective is to measure the generality of the concept of violent action by means of crossdatasets experiments (do not confuse it with cross-validation). 324 F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 Fig. 6. Comparison between two frame sequences of the classes Car Accident XD Violence (Top) and Road Accident UCF Crime (Bottom). Table 2 Matching classes between the UCF Crime and XD Violence datasets. UCF crime classes XD violence classes Match classes Normal Normal ✓ Abuse Abuse ✓ Arrest – × Arson – × Assault – × Burglary – × Explosion Explosion ✓ Fighting Fighting ✓ – Riot × Road Accident Car Accident ✓ Robbery – × Shooting Shooting ✓ Shoplifting – × Stealing – × Vandalism – × For this purpose, the whole of the source dataset is used as the training set and the whole of the target dataset with which the model is to be evaluated as the test dataset. On the one hand, the single-class datasets, NTU CCTV Fights, and UBI Fights are crossed with each other by training the model with one set and validating with the other. On the other hand in the multi-class datasets, XD Violence and UCF Crime, crossdatasets experiment with all the classes cannot be performed since they do not have the same amount of classes and not all of them are coincident, for that reason, the model must be trained only with those classes coincident between both datasets. The matched classes between XD Violence and UCF Crime are shown in the Table 2. Given the similarity of the instances of the classes ‘Road Accident’ in UCF Crime and ‘Car Accident’ in XD Violence, we consider them matched classes. Two sequences of frames from the classes Car Accident from the XD Violence dataset and Road Accident from the UCF Crime dataset are shown in Fig. 6. The frames from the UCF Crime dataset are of lower quality than those from XD Violence, in this case from a movie. However, both classes capture the same concept of traffic accidents and can be compared and considered the same class despite the difference in quality and camera focus. 5.4. Metrics The metrics used in the experiments are as follows: •Confusion Matrix (CM): Allows the display of the number of predictions made by a model that matches the labels. •Receiver Operating Characteristic Area Under the Curve (ROC AUC): Calculates sensitivity versus specificity for a classifier system as the discrimination threshold is varied. •Average Precision (AP): Summarises the area under the precision–recall curve (PR AUC) as the weighted average of the accuracy achieved at each threshold. 6. In-domain: Results from each dataset The results in Table 3 show the effectiveness of CrimeNet with respect to its competitors. More precisely, Fig. 7 presents the nearly perfect confusion matrices of CrimeNet for the multi-class XD Violence and UCF Crime datasets. We highlight that the UCF Crime dataset comes with four different train/test dataset splits: for all of them the CrimeNet results are stable with almost zero standard deviation, confirming the robustness of the model. The inference time for these results is about 40 ms. By reporting the results of ViT we provide an ablation on the role of NSL: indeed CrimeNet builds over ViT and further exploits the sample neighbour graph via NSL. The results indicate that CrimeNet has an advantage over ViT of around 10% points confirming that NSL generates a more robust model by establishing stronger relationships between similar image features. These surprisingly good results are in agreement with those obtained in other works using NSL such as Uddin and Soylu (2021). We note that ViT is already surpassing all the state of the art models except in the case of UBI Fights where results are very close in ROC AUC to Sultani et al. (2018) (89.76% vs 90.60%). The approximate inference time of ViT remains at 40 ms, showing the inclusion of NSL in the model does not cause a prediction delay. Finally, CrimeNet outperforms the current state of the art by far. For the less studied datasets such as NTU CCTV Fights and UBI Fights, perfect results are obtained, which is 20.5% points higher in AP for NTU CCTV Fights and 9.4% points higher ROC AUC in UBI Fights. For the case of the multi-class datasets, the results obtained are again significantly higher than the state of the art. For the XD Violence dataset, the improvement is 22.17% in ROC AUC while for the UCF Crime dataset is 14.6%. 7. Cross-domain: Results across datasets Given the results of our model CrimeNet using ViT and NSL so close to 100% of the metrics, this experiment is essential to check if overfitting is occurring. The results of this experiment show a good performance of the model in generalising the concept of violent action. Figs. 8 and 9show the confusion matrices for the cross-datasets experiment between NTU CCTV Fights - UBI Fights and XD Violence-UCF Crime datasets (retrained using the common classes shown in Table 2). Table 4 shows the accuracy results for each of the crossdatasets experiments. These results show that the single-class datasets generalise the concept of violence better, achieving results of around 80% accuracy. The multi-class datasets show significantly lower performance, but higher than 70% in both cases. 325 F.J. Rendón-Segador, J.A. Álvarez-García, J.L. Salazar-González et al. Neural Networks 161 (2023) 318–329 Table 3 Comparison of CrimeNet results with the state of the art for the case study datasets. Our ablation study using only ViT is also included. Dataset Method ROC AUC AP NTU CCTV Fights (Perez et al.,2019) 3D CNN (Perez et al.,2019) – 79.50% ViT 90.45% 90.40% CrimeNet 100% 100% UBI Fights (Degardin,2020) BayesianNet (Degardin & Proença,2021) 84.60% – ViT 89.76% 89.73% GMM (Degardin,2020) 90.60% – CrimeNet 100% 100% XD Violence (Wu et al.,2020) HL-Net (Wu et al.,2020) – 78.64% Contrastive Attention (Chang et al.,2021) – 76.90% RTFM (Tian et al.,2021) 77.81% – ViT 87.23% 87.20% CrimeNet 99.98% 99.95% UCF Crime (Sultani et al.,2018) DMIL Ranking (Sultani et al.,2018) 75.41% – GMM (Degardin,2020) 75.90% – 3D ResNet (Dubey, Boragule, Jeon,2019) 76.67% – DMIL AutoEncoder (Kamoona et al.,2020) 80.10% – DMIL-MRM (Dubey, Boragule, Gwak, et al.,2021) 81.91% – GCN (Zhong et al.,2019) 82.12% – MIST (Feng et al.,2021) 82.30% – RTFM (Tian et al.,2021) 84.30% – Contrastive Attention (Chang et al.,2021) 84.62% – WSAL (Lv et al.,2021) 85.38% – ViT 87.50% 87.50% CrimeNet 99.98% 99.97% Fig. 7. (a) XD Violence Confusion Matrix, (b) UCF Crime Confusion Matrix. It should be noted that none of the state of the art works studied used this experiment to check the robustness of their models, so we carried out it with the state of the art models that have obtained the best results for each dataset in Table 3, that is, GMM (Degardin,2020) for UBI Fights, RTFM (Tian et al.,2021) for XD Violence and WSAL (Lv et al.,2021) for UCF Crime and whose code is available to reproduce. ViT (without NSL) is also included to check the performance of CrimeNet over the state of the art and ViT. As it can be seen, CrimeNet far outperforms the crossdatasets results of the other models, improving in 25.22% points ROC AUC for UBI Fights, 18.73% for XD Violence and 12.39% for UCF Crime respectively. For the NTU CCTV Fights dataset, there is no reproducible code is available with which to obtain results for comparison. When XD Violence is used as the training dataset, better generalisation can be observed than when UCF Crime is used. We consider that the size of the first dataset, with more than twice as many videos as the second one, allows better training and ROC AUC than the second one (73.5% vs. 70.20%). The ablation study on the role of NSL again shows that CrimeNet has a significant advantage over ViT, around 10% (+9.47%, +11.06%, +10.89%, and +9.47%) in the cross-validation experiments. This difference is in agreement with the results obtained in the In-domain experiments and shows that ViT applied to video overpass state of the art but NSL includes much more robustness. 7.1. Recommendations The practically perfect in-domain results motivated us to carry out the cross-dataset experiment, showing that CrimeNet maintains its advantage over the competitors even in this challenging setting. 326 del rendimiento entre un 20% y un 30 %. Por ejemplo, al entrenar en UCF-Crime y evaluar en XD-Violence, el rendimiento en AUC ROC descendió al 70,20%. Para abordar estas limitaciones, se desarrolló un modelo de ventana deslizante con umbral adaptativo, el cual ajusta automáticamente el umbral de detección de violencia. Esta incorporación ha permitido mejorar significativamente la precisión de detección, incrementándola entre un 10% y un 15 % en experimentos de cruce de conjuntos de datos cuando se aplica como post-procesamiento a los resultados de CrimeNet. Este avance representa una mejora sustancial en la robustez y eficacia de los sistemas de detección de violencia en entornos audiovisuales. Además, se han identificado futuras líneas de investigación que incluyen la mejora de la generalización del modelo, el abordaje del desequilibrio de datos, la exploración de representaciones multimodales, la realización de pruebas en aplicaciones del mundo real y la extensión del enfoque a interacciones humanas más complejas. Este desarrollo ha sido posible gracias al apoyo parcial del proyecto HORUS—Grant n. PID2021-126359OB-I00, financiado por MCIN/AEI/10.13039/501100011033. Citation: Rendón-Segador, F.J.; Álvarez-García, J.A.; Soria-Morillo, L.M. Transformer and Adaptive Threshold Sliding Window for Improving Violence Detection in Videos. Sensors 2024,1, 0. https://doi.org/ Academic Editor: Yun Zhang Received: 10 July 2024 Revised: 13 August 2024 Accepted: 16 August 2024 Published: 22 August 2024 Copyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). sensors Article Transformer and Adaptive Threshold Sliding Window for Improving Violence Detection in Videos Fernando J. Rendón-Segador †,‡ , Juan A. Álvarez-García ‡and Luis M. Soria-Morillo * Departamento de Lenguajes y Sistemas Informáticos, Universidad de Sevilla, Spain; [email protected] (F.J.R.-S.); [email protected] (J.A.Á.-G.) *Correspondence: [email protected] †Current address: Departamento de Lenguajes y Sistemas Informáticos, Universidad de Sevilla, Spain. ‡These authors contributed equally to this work. Abstract: This paper presents a comprehensive approach to detect violent events in videos by combining CrimeNet, a Vision Transformer (ViT) model with structured neural learning and adversarial regularization, with an adaptive threshold sliding window model based on the Transformer architecture. CrimeNet demonstrates exceptional performance on all datasets (XD-Violence, UCF-Crime, NTU-CCTV Fights, UBI-Fights, Real Life Violence Situations, MediEval, RWF-2000, Hockey Fights, Violent Flows, Surveillance Camera Fights, and Movies Fight), achieving high AUC ROC and AUC PR values (up to 99% and 100%, respectively). However, the generalization of CrimeNet to cross-dataset experiments posed some problems, resulting in a 20–30% decrease in performance, for instance, training in UCF-Crime and testing in XD-Violence resulted in 70.20% in AUC ROC. The sliding window model with adaptive thresholding effectively solves these problems by automatically adjusting the violence detection threshold, resulting in a substantial improvement in detection accuracy. By applying the sliding window model as post-processing to CrimeNet results, we were able to improve detection accuracy by 10% to 15% in cross-dataset experiments. Future lines of research include improving generalization, addressing data imbalance, exploring multimodal representations, testing in real-world applications, and extending the approach to complex human interactions. Keywords: deep learning; sliding window; transformer; violence detection; adaptive threshold 1. Introduction Detecting violent events in videos is a critical endeavour within the realms of security, surveillance, and multimedia content analysis (https://english.elpais.com/economy-andbusiness/2024-01-29/the-horrors-experienced-by-meta-moderators-i-didnt-know-what-humansare-capable-of.html) [ 1 – 3 ]. The ability to accurately and efficiently identify violent acts in audiovisual material not only promotes public safety, but also finds applications in fields as diverse as human rights protection, sports event monitoring, and media monitoring [ 4 , 5 ]. However, this task presents a significant challenge due to the presence of false positives and negatives in existing detection systems. This problem is addressed in various studies with different approaches to its solution [6,7]. In this context, the paper proposes a novel approach to mitigate false positives and negatives in the detection of violent events in videos by focusing on its application in the CrimeNet [ 8 ] model, with the aim of reducing the number of false positives and negatives of said model in cross-dataset experiments (training on one dataset and evaluating on a different one). Although CrimeNet performs excellently on all datasets it has trained on, its performance drops significantly in all cross-dataset experiments. Our approach is based on the application of an auxiliary adaptive threshold sliding window deep learning model designed to contextualize and verify the actual presence of violent events in video sequences. The adaptive threshold sliding window, which shifts over time, improved the accuracy of the detection process in cross-dataset experiments by 10% to 15%. Sensors 2024,1, 0. https://doi.org/10.3390/s1010000 https://www.mdpi.com/journal/sensors Sensors 2024,1, 0 2 of 18 The innovation of our approach lies in the integration of an auxiliary adaptive threshold sliding window model into the main detection model. This auxiliary model is trained to provide additional confirmation before issuing a final classification, thereby significantly reducing false positive and negative rates. Through this approach, our goal is to improve the overall accuracy of violent event detection systems in videos. The main contributions are as follows: • Development of a Transformer-based model designed to function as a sliding window, effectively reducing both false positives and false negatives within a binary prediction sequence. • Adaptive learning capability within the model to dynamically assess whether the proportion of violence versus non-violence in a prediction sequence, sized according to the adaptive threshold sliding window, is indicative enough to classify the entire sequence as violent or non-violent. • Examination of class imbalances across various datasets when analyzed frame by frame, exploring the resultant challenges for the CrimeNet model, and implementation of the sliding window model as a strategy to mitigate this issue. • Fusion of the sliding window model with the CrimeNet violent action detection model to create an integrated detection system, coupled with post-processing refinement techniques. This approach aims to enhance violent action detection in videos, thereby advancing the benchmark performance through cross-dataset experiments. This paper is organized as follows: In Section 2, we review related work and datasets, and highlight the need to address the problem of false positives and negatives. In Section 3, we present, in detail, our adaptive threshold sliding window approach and the methodology we followed in its development, highlighting the most important details. Then, in Section 4, we discuss the experiments performed to validate the effectiveness of our model and the expected results. In Section 5, we present and discuss the results obtained by our model and perform a series of comparisons. Finally, in Section 6, we conclude with a summary of the findings and a discussion of the implications and future applications of our approach. 2. Background 2.1. Benchmarks in the Detection of Violent Actions in Videos This section highlights key datasets utilized at the forefront of video violence detection research. Table 1shows a comparative study of the different datasets covered by state-ofthe-art violence detection in videos. Table 1. State of the art datasets for the detection of violence in videos. Dataset Multiclass Nº Classes Time Annotation Nº Items Audio FPS Hockey Fights [9] - 2 27 min Video-Level 1.000 - 30 Movies Fight Detection [10] - 2 6 min Video-Level 200 - 30 Violent Flows—Crowd Violence [11] - 2 15 min Video-Level 246 - - Real Life Violence Situations [12] - 2 - Video-Level 1.000 - - Mediaeval-2013-VSD [13] - 9 35.18 h Video-Level 32.678 ✓- RWF-2000 [14] - 2 - Video Level 2.000 - 5–30 NTU CCTV Fights [15]✓2 17.68 h Frame-Level 1.000 ✓- UBI Fights [16] - 2 80 h Frame-Level 1.000 - 30 Surveillance Camera Fights [17] - 2 - Frame-Level 300 - 30 XD Violence [18]✓7 217 h Frame-Level 4.754 ✓24 UCF Crime [19]✓14 128 h Video-Level 1.900 ✓30 The study delves into multiple datasets concerning the identification of violent occurrences in videos. In particular, the XD-Violence dataset stands out, featuring 4754 untrimmed videos totaling 217 h, characterized by subtle audio cues and seven realistic anomalies. Additionally, the Hockey Fights dataset contributes 1000 violent and 1000 non-violent clips sourced from NHL field hockey games. Furthermore, the Movies Fights Detection dataset offers 200 clips from action movies, encompassing both violent and non-violent sequences. Other datasets such as Violent Flows, Real Life Violence Situations, Mediaeval-2013-VSD Sensors 2024,1, 0 3 of 18 benchmark, RWF 2000, NTU CCTV Fights, UBI Fights, Surveillance Camera Fights, and UCF Crime datasets are also incorporated, each serving specific research objectives within the realm of video violence detection. 2.2. CrimeNet: A Vision Transformer Model for Video Violence Detection CrimeNet is a Vision Transformer (ViT) model specially designed for video violence detection [ 8 ]. It uses a 16-block Transformer architecture to process sequences of video frames represented through optical flow. In addition, CrimeNet incorporates structured neural learning (NSL) with adversarial regularization to generate a structured signal that enhances its violence detection capability. CrimeNet’s architecture is composed of the following key elements: • Video Sequence Input via Optical Flow: Takes the dense optical flow of video frames individually as input. Each frame is fragmented into several equal-sized images called patches, and these patches are embedded into a vector via an Embedding layer. This patch-based representation allows the model to capture local and global details in the video frames. • Sixteen-Block Transformer Encoder: A Transformer encoder consisting of 16 blocks. These blocks are responsible for processing information over time and capturing the complex relationships between video frames. The depth of the architecture allows for a hierarchical representation of visual and motion features. • Neural Structured Learning (NSL) with Adversarial Regularization: Uses NSL [ 20 ] with adversarial regularization to generate a structured signal. This structured signal incorporates prior knowledge into the model, such as motion patterns, textures, and visual features relevant to violence detection. Adversarial regularization improves the quality of the structured signal by training the model to discriminate between real and generated signals. • Output and Classification: The output is a binary classification that determines whether a video fragment contains violence or not. The model learns to perform this classification during training on the datasets mentioned above. CrimeNet is designed to identify violence in videos by examining visual and motion patterns extracted from optical flow data. By analyzing temporal relationships and utilizing NSL with adversarial regularization to create a structured signal, CrimeNet can detect subtle cues associated with violence. These cues include sudden changes in motion, aggressive gestures, and changes in lighting conditions. 2.3. Problem of False Positives and Negatives in the Detection of Violent Events in Videos Detecting violent events in videos stands as a pivotal challenge within the realms of computer vision and deep learning. With the surge in online video-sharing platforms, there arises a heightened demand for automated tools capable of discerning and categorizing pertinent events within multimedia content. However, this effort is fraught with the inherent risk of producing false positives and negatives, potentially undermining the efficacy and practical utility of detection systems. Table 2shows a summary of the state of the art, with an emphasis on the two bestperforming proposals on the most popular datasets in the field of violent action detection in videos. Sensors 2024,1, 0 4 of 18 Table 2. State-of-the-art detection of violent actions in videos. This state-of-the-art table contains the two best-performing developments for each existing benchmark in the field of video violence detection. Cross-dataset experiments are not shown in this table. * The model has been trained and tested with the corresponding dataset in this paper. Dataset Method AUC ROC Accuracy AP Surveillance Camera Fight EfficientNet CNN + TSE Block [21] - 92.00 - Surveillance Camera Fight CNN-BiLSTM + Attention [17] - 72.00 - Crowd Violence/Violent Flows C3D + SVM [22] 99.00 99.29 - Crowd Violence/Violent Flows EfficientNet CNN + TSE Block [21] - 98.00 - RWF-2000 Structured Keypoint Pooling [23] - 93.40 - RWF-2000 Semi-Supervised Hard Attention [24] - 90.04 - RWF-2000 EfficientNet CNN + TSE Block [21] - 92.00 - RWF-2000 ConvNet 3D [25] - 87.25 - Movies Fights EfficientNet CNN + TSE Block [21] - 100.00 - Movies Fights ViolenceNet [26] 100.00 100.00 100.00 Hockey Fights CNN+LSTM [27] - - 98.00 Hockey Fights Structured Keypoint Pooling [23] - - 99.5 Hockey Fights EfficientNet CNN + TSE Block [21] - 99.60 - Hockey Fights ViolenceNet [26] 99.37 99.20 99.11 NTU CCTV Fights DeVTrV2 CNN + Transformer [28] - 96.00 - NTU CCTV Fights CrimeNet [8] 100.00 100.00 100.00 UBI Fights DeVTrV2 CNN + Transformer [28] - 91.80 - UBI Fights CrimeNet [8] 100.00 100.00 100.00 Real Life Violence Situations DeVTr [29] - 96.25 - Real Life Violence Situations DeVTrV2 CNN + Transformer [28] - 98.25 - Real Life Violence Situations DeVTrV2 ViT [28] - 98.00 - UCF Crime Magnitude-Contrastive Glance-and-Focus Network [30]86.98 - - UCF Crime BatchNorm Weakly Supervised [31] 87.24 - - UCF Crime ViolenceNet * [26] 88.31 88.31 88.22 UCF Crime CrimeNet [8] 99.98 99.99 99.97 XD-Violence PEL [32] - - 85.59 XD-Violence HyperVD [33] - - 85.67 XD-Violence ViolenceNet * [26] 91.35 90.81 91.12 XD-Violence CrimeNet [8] 99.98 99.97 99.95 Mediaeval-2013-VSD ViolenceNet * [26] 94.66 95.83 Mediaeval-2013-VSD CNN-ConvNet 3D [34] - 78.50 - Given the state of the art, it is worth highlighting the excellent results obtained by the EfficientNet CNN + TSE Block model [ 21 ], a model based on convolutional neural networks to which temporal attention modules called Temporal Squeeze-and-Excitation, or TSE, blocks are added, obtaining results above 90% in all the datasets on which it is tested. Furthermore, noteworthy are the models that combine convolutional neural networks with Transformer models or with attention mechanisms, the basis of the Transformer models. The DeVTrV2 CNN + Transformer [ 28 ] model achieves over 90% on complex datasets such as NTU CCTV Fights and UBI Fights. It is worth mentioning the models based on semi-supervised or weakly supervised learning are capable of working with small ds or with limited instances. These achieve an accuracy of over 90% for the RWF-2000 dataset [ 24 ] and over 85% for the UCF-Crime dataset [31]. Although CrimeNet is shown in the table as a superior performing model on the NTU CCTV Fights, UBI-Fights, XD-Violence, and UCF-Crime datasets, it faces a lack of accuracy in cross-dataset experiments. The CrimeNet model trains and evaluates datasets at the frame level. That is, each frame is labeled as either violent or non-violent. This results in a significant data imbalance, since, in most of the current datasets, the number of non-violent frames is significantly higher than the number of non-violent frames. The CrimeNet model overtrains on unbalanced datasets and results in high accuracy when entering and evaluating the same dataset. However, when cross-dataset experiments are performed with CrimeNet trained on one dataset and evaluated on a different dataset, the accuracy drops by 20% to 30%. This problem generates high rates of false positives and negatives, see Table 3. From here, we face the challenge of finding a method to reduce the Sensors 2024,1, 0 5 of 18 rate of false positives and negatives and get a model able to better generalize the concept of violent action. Table 3. CrimeNet results in cross-dataset experiments between the different datasets on which it was evaluated. Dataset Traning Dataset Test Method AUC ROC AP NTU CCTV Fights UBI Fights CrimeNet [8] 78.90 78.87 UBI Fights NTU CCTV Fights CrimeNet [8] 81.35 81.35 XD-Violence UCF Crime CrimeNet [8] 73.50 73.44 UCF Crime XD-Violence CrimeNet [8] 70.20 70.10 To address the challenge of false positives and negatives in detecting violent events in videos, several approaches have been developed in the literature [ 35 – 38 ]. These methods focus on improving model accuracy and reducing classification errors. Some traditional strategies include the following: Threshold adjustment: In many detection systems, decision thresholds can be adjusted to classify an instance as positive or negative [ 39 ]. Increasing the threshold may reduce the number of negative results, which, in turn, will reduce false negatives. However, this can also increase false positives, so finding the right balance is crucial. Feature enhancement: Refining the features used in detection algorithms can help reduce false positives and negatives. Identifying the most relevant and significant features can improve the overall accuracy of the system, [40,41]. Semi-supervised learning: Using semi-supervised learning techniques can help reduce false positives and negatives by allowing the system to learn from unlabeled positive and negative examples, which can improve discrimination between positives and negatives, [42]. Sliding Window: The sliding window approach in video detection reduces false positives and negatives by analyzing video frame sequences. By adjusting its size and sliding rate, it captures events of different scales and speeds, thus enhancing the system’s accuracy in reliably identifying violent events [43]. 3. Methodology This section presents the methodology used to address the detection of violent events in videos. We describe the approach that combines the CrimeNet model with an adaptive threshold sliding window model to reduce false positives and negatives. While the CrimeNet model focuses on analyzing frames to determine whether they contain violent acts, the adaptive threshold sliding window model specializes in analyzing temporal sequences of events in a video to determine the amount of violence in that sequence. Firstly, the adaptive threshold sliding window model based on the Transformer architecture was developed to analyze temporal sequences of events in videos and determine the level of violence present in each sequence. This involved designing the model architecture, including the necessary attention mechanisms or specific layers to effectively capture the temporal relationships within the video sequences. Secondly, a synthetic dataset of binary sequences was generated. These sequences represent the labels of each frame in a sequence of frames, where ones represent violent frames and zeros represent non-violent frames. Each binary sequence is associated with a global label that determines whether the sequence is completely violent or not. For the global labeling of the sequence, an Autoencoder combined with the K-Means clustering algorithm was used. The objective of this synthetic dataset is for the model to be able to detect the sequences that contain false positives or negatives, evaluating the distribution of ones and zeros in the sequence and the label of said sequence. For example, if, in a Sensors 2024,1, 0 6 of 18 sequence labeled as non-violent (zero), there is a smaller distribution of violence (ones), it will mean that those ones correspond to false positives in the sequence. Thirdly, the model was trained and evaluated. The synthetic dataset was used as training and evaluation data for the model. Training was carried out for 50 epochs using an NVIDIA 2070 Super GPU. For the evaluation of the model, the AUC ROC and AUC PR metrics were used to measure the performance of detecting false positives and negatives in violent events in videos, thus completing the methodological process. Finally, CrimeNet was combined with the adaptive threshold sliding window model. The model collected sequences of predictions from CrimeNet in binary format and determined whether these predictions contained false positives or negatives for later correction. A series of cross-dataset experiments was performed on the entire system. In addition, two comparisons were performed: one using an empirical sliding window model instead of the adaptive threshold model and another using the K-Means algorithm directly instead of the adaptive threshold model. These experiments will evaluate the performance of the combined approach and compare it with alternative methods to determine the effectiveness of the proposed model in detecting violent events in videos. 4. Resolution Methods 4.1. Adaptive Threshold Sliding Window Model The adaptive threshold sliding window model is built upon attention mechanisms specific to the Transformer architecture Figure 1. This model emphasizes the analysis of temporal sequences of events, using an attention structure to capture enduring relationships among sequence elements, a crucial aspect for discerning pertinent patterns within the data. The chosen loss function, binary cross-entropy, is well-suited for binary classification tasks such as violent event detection. Additionally, a synthetic dataset comprising one million binary sequences was generated for training purposes, with each sequence labeled according to its content. This synthetic dataset mirrors the process of scrutinizing a video through the adaptive threshold sliding window, enhancing the model’s robustness through simulated real-world scenarios. Figure 1. Adaptive threshold sliding window model architecture. It uses an Embedding layer together with a MultiHead Attention layer following the philosophy of the Transformer models. This architecture takes advantage of the strengths of the Transformer models to establish relationships between the different elements of a sequence of data in a one-dimensional vector. 4.2. Synthetic Data Generation, Labeling and Relationship Building To train the adaptive thresholded sliding window model, we create a synthetic dataset representing sequences of video frames classified as violent or non-violent. This synthetic dataset is a fixed-length binary sequence of 1s and 0s representing violent or non-violent frame labels, e.g., let Si= [ 1,0,0,0,0,0,0,0,0,1, 0 ] be a binary sequence such that Si∈D where D is the synthetic dataset. One million binary sequences are generated for the dataset. Each sequence is globally labeled as a completely violent or non-violent sequence (1, 0). One-quarter of the generated sequences are sequences in which all values are zeros Sensors 2024,1, 0 7 of 18 (non-violent), and another quarter are sequences in which all values are ones (violent). These sequences are globally labeled with 1 for sequences in which all values are ones and with 0 for sequences in which all values are zeros. The rest of the dataset is sequences with randomly generated values. To globally label these sequences, we used an unsupervised autoencoder together with the K-Means algorithm (with the number of clusters K = 2). The autoencoder receives binary sequences as input and learns intrinsic representations of them, captures their structure, and detects the distribution of violent elements. K-Means then assigns labels to the sequences by grouping them into two categories. To test the validity of the autoencoder + K-Means combination, an ablation study was performed, in which only the K-Means algorithm was used for sequence labeling. The results can be seen in Section 5. This method dynamically identifies whether there is a “significant percentage of violence” in the binary sequence without relying on manual definitions or fixed thresholds. A significant percentage of violence refers to the distribution between zeros and ones in the binary sequence. This adaptive process ensures accurate labeling and fitting of data features, which are crucial to effectively detect violent events in videos; see Figure 2. Once we have the synthetic dataset generated and labeled, we can move on to training and evaluating the sliding window model with an adaptive threshold. Generate binary sequences Features codification Auto-Encoder K-Means 10011 00000 1 0 Labelling Binary Sequence Figure 2. Synthetic data labeling process using autoencoder plus K-Means. The autoencoder takes the synthetically generated binary sequences as input and encodes them. From these encodings, the K-Means algorithm groups the corresponding sequences into two clusters and labels them according to the cluster to which they belong after the application of K-Means. 4.3. Integration of CrimeNet with the Adaptive Threshold Sliding Window Model In the final phase of our approach, the adaptive threshold sliding window model, designed to enhance the outcomes of the CrimeNet model, was incorporated into a unified system for detecting violent events in videos. This system is structured into two distinct stages. See Figure 3. Sensors 2024,1, 0 8 of 18 Optical Flow Frames Predicted Labels Samples Frames Frames Dense Optical Flow Processing CrimeNet 1110 1001 1111 1111 Predictions Violence Sequence Frames ViolenceViolenceViolence 0100 0010 0000 0000 Non ViolenceNon ViolenceNon Violence Adaptative Threshold Sliding Window Model ... ... ... ... t [0, 100) t (100, 200] t [0, 100) t (100, 200] ... Figure 3. Complete system that integrates CrimeNet with the adaptive threshold sliding window model to correct false positives and negatives that may be generated by CrimeNet. In the initial stage of the system, the CrimeNet model is used. It processes the optical flow of the input video frames, generating a sequence of binary predictions. Subsequently, in the second stage, the dynamic temporal window model processes these prediction sequences. This stage harmonizes the sequences by assessing whether each one should be fully classified as violent or not. This step effectively mitigates any false positives and negatives that may have been generated by CrimeNet. Finally, the frames are labeled based on the results provided by the adaptive threshold sliding window model. 5. Experimental Results In this section, we present the experimental results of our approach to detect violent events in videos using CrimeNet, as well as the Transformer-based adaptive threshold sliding window post-processing model. We evaluated their performance on the datasets described above and analyzed the impact of our approach on reducing false positives in cross-dataset experiments. This section is subdivided into several subsections that are organized as follows: • Section 5.1 CrimeNet Results: this subsection shows the CrimeNet results for all datasets, as well as the results of the cross-dataset experiments and the distribution of false positives and negatives across videos in those experiments. • Section 5.2 Adaptive Threshold Sliding Window Model Results: This subsection shows the results of the adaptive threshold sliding window model with the synthetic dataset. In addition, we show the attention maps exposing how the attention mechanisms of the model work with respect to various instances of the synthetic dataset. • Section 5.3 Results with the Adaptive Threshold Sliding Window Model in CrossDatasets: This subsection shows the results of the cross-dataset experiments combining the adaptive threshold sliding window model with the CrimeNet model. 5.1. CrimeNet Results CrimeNet demonstrates exceptional performance in detecting violent events across all datasets. It was initially trained and evaluated on the NTU CCTV Fights, UBI-Fights, Sensors 2024,1, 0 9 of 18 XD-Violence, and UCF Crime datasets. In our experimentation, in addition to the datasets used in the original [ 8 ] publication, we tested CrimeNet on the Real Life Violence Situations, Hockey Fights, RWF-2000, Violent Flows-Crowd Violence, Surveillance Camera Fights, and Medieval-2013-VSD datasets. Then, we applied CrimeNet to the remaining datasets that achieved near-perfect results, with AUC ROC and AP metrics peaking between 99% and 100%; see Table 4. Table 4. CrimeNet results for all datasets with five-cross-validation. Without adaptive threshold sliding window. Dataset Accuracy% AUC ROC% AP% NTU CCTV Fights 100.00 ±0.00 100.00 ±0.00 100.00 ±0.00 UBI Fights 100.00 ±0.00 100.00 ±0.00 100.00 ±0.00 UCF Crime 99.99 ±0.01 99.98 ±0.01 99.97 ±0.02 XD-Violence 99.97 ±0.02 99.98 ±0.02 99.95 ±0.04 Real Life Violence Situations 100.00 ±0.00 100.00 ±0.00 100.00 ±0.00 Hockey Fights 100.00 ±0.00 100.00 ±0.00 100.00 ±0.00 RWF-2000 99.98 ±0.02 99.97 ±0.02 99.98 ±0.01 Violent Flows - Crowd Violence 99.97 ±0.03 99.95 ±0.03 99.96 ±0.02 Surveillance Camera Fights 100.00 ±0.00 100.00 ±0.00 100.00 ±0.00 Mediaeval-2013-VSD 99.98 ±0.01 99.95 ±0.04 99.98 ±0.02 To assess CrimeNet’s robustness and generalization, cross-dataset experiments were conducted. Although performance remained high when trained and evaluated on the same dataset, there was a notable decrease (20–30%) in AUC ROC and AP scores during cross-dataset experiments. See Figure 4. This suggests that CrimeNet’s performance is influenced by dataset characteristics and distribution. Furthermore, testing on movie clips outside the training datasets revealed challenges in false positive and false negative rates due to data imbalance and variability in violent event representation. Figure 4. CrimeNet cross-dataset comparison with initial datasets (NTU CCTV-Fights, UBI-Fights, XD-Violence, and UCF-Crime). The CrimeNet model trained with NTU CCTV Fights and UBI Fights was evaluated video by video individually with the rest of the datasets, showing that, in cross-dataset experiments, the false positives are significantly lower than the false negatives in each video. The model is biased in classifying violent actions as non-violent, as shown in Figures 5and Sensors 2024,1, 0 16 of 18 influence of the distribution of training data. Cross-dataset experiments will further enrich the understanding of how CrimeNet performs in different contexts. • Addressing Data Imbalance. Continue to develop approaches that mitigate the impact of data imbalance on the detection of violent events, including generating synthetic data and refining post-processing strategies. Evaluating performance in contexts with different proportions of violent events may be an additional step. • Exploring Multimodal Representations. Investigate how combining visual data with audio information and other sensory modes can enrich representations and improve accuracy in detecting violent events. This would open the door to a more comprehensive approach to violent event detection. Real-World Applications Test and adapt the proposed approach in real-world situations, such as public safety and policing, to evaluate its effectiveness in practical scenarios. Studying how CrimeNet and the adaptive threshold sliding window model behave under varying conditions can help determine their practical applicability. • Human Interaction. Explore how this approach could be extended to the detection of violent events in more complex human interactions and scenarios, such as heated arguments and conflict situations. This would broaden the scope of the application of violent event detection. This work represents a significant advance in the field of violent event detection in videos by presenting an effective approach that combines CrimeNet and the adaptive threshold sliding window model. Future directions offer exciting opportunities to further improve the accuracy and applicability of this technology in a variety of contexts. We hope that this work will inspire further research and advances in the detection of violent events in videos, and open the door to its implementation in real-world applications. Author Contributions: Methodology, F.J.R.-S.; software, F.J.R.-S.; validation, F.J.R.-S. and J.A.Á.-G.; investigation, F.J.R.-S., J.A.Á.-G., and L.M.S.-M.; resources, F.J.R.-S. and J.A.Á.-G.; data curation, F.J.R.-S.; writing—original draft, F.J.R.-S.; writing—review and editing, F.J.R.-S., J.A.Á.-G., and L.M.S.-M.; project administration, J.A.Á.-G., and L.M.S.-M.; funding acquisition, J.A.Á.-G. and L.M.S.-M. All authors have read and agreed to the published version of the manuscript. Funding: This research was partially supported by the HORUS project—Grant n. PID2021-126359OBI00 funded by MCIN/AEI/ 10.13039/501100011033. Institutional Review Board Statement: Not applicable. Informed Consent Statement: Not applicable. Data Availability Statement: The data used to support the findings of this study are available from the corresponding author upon request. Conflicts of Interest: All authors declare that the research was conducted in the absence of commercial or financial relationships that could be interpreted as a potential conflict of interest. References 1. Ullah, F.U.M.; Obaidat, M.S.; Ullah, A.; Muhammad, K.; Hijji, M.; Baik, S.W. A comprehensive review on vision-based violence detection in surveillance videos. ACM Comput. Surv. 2023,55, 1–44. 2. Pu, Y.; Wu, X.; Wang, S.; Huang, Y.; Liu, Z.; Gu, C. Semantic multimodal violence detection based on local-to-global embedding. Neurocomputing 2022,514, 148–161. 3. Acar, E.; Hopfgartner, F.; Albayrak, S. Breaking down violence detection: Combining divide-et-impera and coarse-to-fine strategies. Neurocomputing 2016,208, 225–237. 4. Mumtaz, N.; Ejaz, N.; Habib, S.; Mohsin, S.M.; Tiwari, P.; Band, S.S.; Kumar, N. An overview of violence detection techniques: current challenges and future directions. Artif. Intell. Rev. 2023,56, 4641–4666. 5. Huszar, V.D.; Adhikarla, V.K.; Négyesi, I.; Krasznay, C. Toward fast and accurate violence detection for automated video surveillance applications. IEEE Access 2023,11, 18772–18793. 6. Bianculli, M.; Falcionelli, N.; Sernani, P.; Tomassini, S.; Contardo, P.; Lombardi, M.; Dragoni, A.F. A dataset for automatic violence detection in videos. Data Brief 2020,33, 106587. 7. Sernani, P.; Falcionelli, N.; Tomassini, S.; Contardo, P.; Dragoni, A.F. Deep learning for automatic violence detection: Tests on the AIRTLab dataset. IEEE Access 2021,9, 160580–160595. Sensors 2024,1, 0 17 of 18 8. Rendón-Segador, F.J.; Álvarez-García, J.A.; Salazar-González, J.L.; Tommasi, T. Crimenet: Neural structured learning using vision transformer for violence detection. Neural Netw. 2023,161, 318–329. 9. Bermejo Nievas, E.; Deniz Suarez, O.; Bueno García, G.; Sukthankar, R. Violence detection in video using computer vision techniques. In Proceedings of the International Conference on Computer Analysis of Images and Patterns, Seville, Spain, 29–31 August 2011; pp. 332–339. 10. Nievas, E.B.; Suarez, O.D.; Garcia, G.B.; Sukthankar, R. Movies Fight Detection Dataset. In Proceedings of the Computer Analysis of Images and Patterns, Seville, Spain, 29–31 August 2011; pp. 332–339. 11. Hassner, T.; Itcher, Y.; Kliper-Gross, O. Violent flows: Real-time detection of violent crowd behavior. In Proceedings of the 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, Providence, RI, USA, 16–21 June 2012; pp. 1–6. https://doi.org/10.1109/CVPRW.2012.6239348. 12. Soliman, M.M.; Kamal, M.H.; Nashed, M.A.E.M.; Mostafa, Y.M.; Chawky, B.S.; Khattab, D. Violence recognition from videos using deep learning techniques. In Proceedings of the 2019 Ninth International Conference on Intelligent Computing and Information Systems (ICICIS), Cairo, Egypt, 8–10 December 2019; pp. 80–85. 13. Schedi, M.; Sjöberg, M.; Mironic˘a, I.; Ionescu, B.; Quang, V.L.; Jiang, Y.G.; Demarty, C.H. VSD2014: A dataset for violent scenes detection in hollywood movies and web videos. In Proceedings of the 2015 13th International Workshop on Content-Based Multimedia Indexing (CBMI), Prague, Czech Republic, 10–12 June 2015; pp. 1–6. https://doi.org/10.1109/CBMI.2015.7153604. 14. Cheng, M.; Cai, K.; Li, M. RWF-2000: An Open Large Scale Video Database for Violence Detection. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 10–15 January 2021; pp. 4183–4190. https: //doi.org/10.1109/ICPR48806.2021.9412502. 15. Perez, M.; Kot, A.C.; Rocha, A. Detection of real-world fights in surveillance videos. In Proceedings of the ICASSP 2019— 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, 12–17 May 2019; pp. 2662–2666. 16. Degardin, B.M. Weakly and Partially Supervised Learning Frameworks for Anomaly Detection. Ph.D. Thesis, Universidade da Beira Interior (Portugal), Covilha, Portugal, 2020. 17. Aktı, ¸S.; Tataro˘glu, G.A.; Ekenel, H.K. Vision-based fight detection from surveillance cameras. In Proceedings of the 2019 Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA), Istanbul, Turkey, 6–9 November 2019; pp. 1–6. 18. Wu, P.; Liu, J.; Shi, Y.; Sun, Y.; Shao, F.; Wu, Z.; Yang, Z. Not only look, but also listen: Learning multimodal violence detection under weak supervision. In Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020; pp. 322–339. 19. Sultani, W.; Chen, C.; Shah, M. Real-World Anomaly Detection in Surveillance Videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018. 20. Gopalan, A.; Juan, D.C.; Magalhaes, C.I.; Ferng, C.S.; Heydon, A.; Lu, C.T.; Pham, P.; Yu, G.; Fan, Y.; Wang, Y. Neural structured learning: training neural networks with structured signals. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining, Jerusalem, Israel, 8–12 March 2021; pp. 1150–1153. 21. Kang, M.S.; Park, R.H.; Park, H.M. Efficient spatio-temporal modeling methods for real-time violence recognition. IEEE Access 2021,9, 76270–76285. 22. Accattoli, S.; Sernani, P.; Falcionelli, N.; Mekuria, D.N.; Dragoni, A.F. Violence detection in videos by combining 3D convolutional neural networks and support vector machines. Appl. Artif. Intell. 2020,34, 329–344. 23. Hachiuma, R.; Sato, F.; Sekii, T. Unified keypoint-based action recognition framework via structured keypoint pooling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 22962–22971. 24. Mohammadi, H.; Nazerfard, E. SSHA: Video Violence Recognition and Localization Using a Semi-Supervised Hard Attention Model. arXiv 2022, arXiv:2202.02212. 25. Cheng, M.; Cai, K.; Li, M. Rwf-2000: An open large scale video database for violence detection. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 10–15 January 2021; pp. 4183–4190. 26. Rendón-Segador, F.J.; Álvarez-García, J.A.; Enríquez, F.; Deniz, O. Violencenet: Dense multi-head self-attention with bidirectional convolutional lstm for detecting violence. Electronics 2021,10, 1601. 27. Abdali, A.M.R.; Al-Tuma, R.F. Robust real-time violence detection in video using cnn and lstm. In Proceedings of the 2019 2nd Scientific Conference of Computer Sciences (SCCS), Baghdad, Iraq, 27–28 March 2019; pp. 104–108. 28. Abdali, A.R.; Aggar, A.A. DEVTrV2: Enhanced Data-Efficient Video Transformer For Violence Detection. In Proceedings of the 2022 7th International Conference on Image, Vision and Computing (ICIVC), Xi’an, China, 26–28 July 2022; pp. 69–74. 29. Abdali, A.R. Data efficient video transformer for violence detection. In Proceedings of the 2021 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT), Purwokerto, Indonesia, 17–18 July 2021; pp. 195–199. 30. Chen, Y.; Liu, Z.; Zhang, B.; Fok, W.; Qi, X.; Wu, Y.C. Mgfn: Magnitude-contrastive glance-and-focus network for weaklysupervised video anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 387–395. 31. Zhou, Y.; Qu, Y.; Xu, X.; Shen, F.; Song, J.; Shen, H. BatchNorm-based Weakly Supervised Video Anomaly Detection. arXiv 2023, arXiv:2311.15367. Sensors 2024,1, 0 18 of 18 32. Pu, Y.; Wu, X.; Wang, S. Learning Prompt-Enhanced Context Features for Weakly-Supervised Video Anomaly Detection. arXiv 2023, arXiv:2306.14451. 33. Peng, X.; Wen, H.; Luo, Y.; Zhou, X.; Yu, K.; Yang, P.; Wu, Z. Learning weakly supervised audio-visual violence detection in hyperbolic space. arXiv 2023, arXiv:2305.18797. 34. Constantin, M.G.; ¸Stefan, L.D.; Ionescu, B.; Demarty, C.H.; Sjöberg, M.; Schedl, M.; Gravier, G. Affect in multimedia: Benchmarking violent scenes detection. IEEE Trans. Affect. Comput. 2020,13, 347–366. 35. Aloysius, C.; Tamilselvan, P. A Novel Method to Reduce False Positives and Negatives in Sentiment Analysis. Int. J. Intell. Syst. Appl. Eng. 2022,10, 365–373. 36. Saha, A.; Denning, T.; Srikumar, V.; Kasera, S.K. Secrets in source code: Reducing false positives using machine learning. In Proceedings of the 2020 International Conference on COMmunication Systems & NETworkS (COMSNETS), Bengaluru, India, 7–11 January 2020; pp. 168–175. 37. Ma, Y.; Peng, Y.; Wu, T.Y. Transfer learning model for false positive reduction in lymph node detection via sparse coding and deep learning. J. Intell. Fuzzy Syst. 2022,43, 2121–2133. 38. El Kaid, A.; Baïna, K.; Baïna, J. Reduce false positive alerts for elderly person fall video-detection algorithm by convolutional neural network model. Procedia Comput. Sci. 2019,148, 2–11. 39. Wang, L.; Zhao, X.; Liu, Y. Reduce false positives for object detection by a priori probability in videos. Neurocomputing 2016, 208, 325–332. 40. Gite, S.; Tiwari, C.; Chandana, J.; Chanumolu, S.V.; Shrivastava, A.; Kotecha, D.K. Crowd Violence Detection Using Deep Learning Techniques and Explanation Using Xai. Available at SSRN 4524940, 2011. 41. Nourani, M.; Honeycutt, D.R.; Block, J.E.; Roy, C.; Rahman, T.; Ragan, E.D.; Gogate, V. Investigating the importance of first impressions and explainable ai with interactive video analysis. In Proceedings of the Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 25–30 April 2020; pp. 1–8. 42. Kumar, A.; Rawat, Y.S. End-to-end semi-supervised learning for video action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 14700–14710. 43. Bilinski, P.; Bremond, F. Human violence recognition and detection in surveillance videos. In Proceedings of the 2016 13th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Colorado Springs, CO, USA, 23–26 August 2016; pp. 30–36. Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. PARTE III Observaciones finales página 101 CAPÍTULO 6 CONCLUSIONES Y TRABAJO FUTURO Daría todo lo que sé por la mitad de lo que ignoro - René Descartes 6.1. Conclusiones Esta tesis se centra en el desarrollo y mejora de sistemas de detección de violencia en videos, abordando desafíos críticos en términos de precisión, generalización y eficiencia computacional. El objetivo principal fue diseñar y evaluar un modelo robusto capaz de detectar comportamientos violentos en entornos audiovisuales complejos, manteniendo un equilibrio óptimo entre precisión y recursos computacionales. Esto es particularmente relevante para aplicaciones en videovigilancia y monitorización de contenido audiovisual, donde la fiabilidad y la velocidad de respuesta son cruciales. En las etapas iniciales de este trabajo, se desarrolló ViolenceNet, una primera página 103 aproximación que combinaba técnicas de deep learning con capas convolucionales y modelos de atención. ViolenceNet representó un avance significativo al mejorar la precisión en la detección de violencia en videos respecto a los métodos anteriores. Sin embargo, aún presentaba limitaciones, particularmente en términos de generalización y un alto número de falsos positivos en datasets más complejos. Este primer modelo, aunque eficaz en muchos aspectos, evidenció la necesidad de una arquitectura más avanzada que pudiera adaptarse mejor a la variabilidad de los escenarios violentos en diferentes contextos. Para superar estas limitaciones, se evolucionó hacia el desarrollo de CrimeNet, un modelo basado en Vision Transformer (ViT) y Neural Structured Learning (NSL) con entrenamiento adversarial. CrimeNet fue diseñado para abordar específicamente los problemas identificados en ViolenceNet, logrando una reducción drástica de falsos positivos y mejorando significativamente la precisión en los conjuntos de datos más desafiantes. Las pruebas realizadas demostraron que CrimeNet no solo superó ampliamente a ViolenceNet y otros modelos previos en términos de AUC ROC (mejorando el estado del arte entre 9.4 y 22.17 puntos porcentuales), sino que también alcanzó un rendimiento sin precedentes en términos de confiabilidad y eficiencia. Sin embargo, uno de los desafíos persistentes tanto para ViolenceNet como para CrimeNet fue la generalización del modelo en experimentos de cross-dataset, donde se observó una caída en el rendimiento de entre 20 y 30 puntos porcentuales en AUC ROC al entrenar en un conjunto de datos y probar en otro. Este problema fue abordado mediante la implementación de un modelo de ventana deslizante con umbral adaptativo basado en la arquitectura Transformer, que ajusta dinámicamente el umbral de detección de violencia. Esta mejora permitió aumentar la precisión en un 10 % a 15% en experimentos de cross-dataset, mitigando así uno de los mayores desafíos en la detección de violencia en videos. En resumen, esta tesis ha contribuido significativamente al campo de la detección de violencia en videos, desde la primera aproximación con ViolenceNet hasta la evolución más avanzada en CrimeNet. Estos desarrollos no solo proponen modelos más precisos y eficientes, sino que también abordan el desafío crítico de la genera- lización a diferentes contextos y datasets. Los avances presentados en CrimeNet y sus mejoras posteriores sientan las bases para futuras investigaciones orientadas a mejorar aún más la capacidad de generalización del modelo, abordar problemas de desbalanceo de datos, explorar representaciones multimodales y validar el sistema en aplicaciones del mundo real. Finalmente, estos desarrollos han sido posibles gracias al apoyo de los proyectos DISARM y HORUS, financiados por MCIN/AEI y la Unión Europea NextGeneration EU/PRTR. Estos resultados no solo destacan la robustez y eficacia de CrimeNet, sino que también subrayan la importancia de la colaboración continua en la investigación para superar los desafíos actuales y futuros en la detección de violencia en entornos audiovisuales. 6.2. Trabajo futuro Aunque esta tesis ha logrado avances significativos en la detección de violencia en videos mediante el desarrollo de ViolenceNet y su evolución en CrimeNet, aún existen varias áreas que requieren investigación y desarrollo adicionales para mejorar y extender las capacidades del sistema. Mejora de la Generalización del Modelo: Uno de los principales desafíos observados, especialmente en experimentos de cross-dataset, es la capacidad de generalización de CrimeNet. En el futuro, se podrían explorar nuevas estrategias de regularización, técnicas de transfer learning y la integración de datasets más diversos y representativos para mejorar la robustez del modelo en escenarios fuera del conjunto de datos en el que fue entrenado. Balanceo de Datos y Manejo de Sesgos: El desbalance de clases, donde los eventos violentos son menos frecuentes que los no violentos, sigue siendo un desafío importante. En nuestro trabajo, optamos por utilizar Generative Adversarial Networks (GANs) para abordar este problema. Las GANs se componen de dos redes neuronales en competencia: un generador, que produce BIBLIOGRAFÍA [1] A. R. Abdali and A. A. Aggar. Devtrv2: Enhanced data-efficient video transformer for violence detection. In 2022 7th International Conference on Image, Vision and Computing (ICIVC), pages 69–74, 2022. [2] S. Accattoli, P. Sernani, N. Falcionelli, D. N. Mekuria, and A. F. Dragoni. Violence detection in videos by combining 3d convolutional neural networks and support vector machines. Applied Artificial Intelligence, 34(4):329–344, 2020. [3] Ş. Aktı, G. A. Tataroğlu, and H. K. Ekenel. Vision-based fight detection from surveillance cameras. In 2019 Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA), pages 1–6. IEEE, 2019. [4] Aktı, G. A. Tataroğlu, and H. K. Ekenel. Vision-based fight detection from surveillance cameras. In 2019 Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA), pages 1–6, 2019. [5] E. Bermejo Nievas, O. Deniz Suarez, G. Bueno García, and R. Sukthankar. Violence detection in video using computer vision techniques. In International conference on Computer analysis of images and patterns, pages 332–339. Springer, 2011. [6] P. Bilinski and F. Bremond. Human violence recognition and detection in surveillance videos. In 2016 13th IEEE international conference on advanced video and signal based surveillance (AVSS), pages 30–36. IEEE, 2016. página 113 [7] T. D. Bui, S. Ravi, and V. Ramavajjala. Neural graph learning: Training neural networks using graphs. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, pages 64–71, 2018. [8] S. Chang, Y. Li, S. Shen, J. Feng, and Z. Zhou. Contrastive attention for video anomaly detection. IEEE Transactions on Multimedia, 24:4067–4076, 2022. [9] M. Cheng, K. Cai, and M. Li. Rwf-2000: An open large scale video database for violence detection. In 2020 25th International Conference on Pattern Recognition (ICPR), pages 4183–4190, 2021. [10] S. Das, A. Sarker, and T. Mahmud. Violence detection from videos using hog features. In 2019 4th International Conference on Electrical Information and Communication Technology (EICT), pages 1–5. IEEE, 2019. [11] B. M. Degardin. Weakly and partially supervised learning frameworks for anomaly detection. Master’s thesis, Universidade da Beira Interior (Portugal), 2020. [12] A. Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. [13] S. Dubey, A. Boragule, J. Gwak, and M. Jeon. Anomalous event recognition in videos based on joint learning of motion and appearance with multiple ranking measures. Applied Sciences, 11(3):1344, 2021. [14] S. Dubey, A. Boragule, and M. Jeon. 3d resnet with ranking loss function for abnormal activity detection in videos. In 2019 international conference on control, automation and information sciences (ICCAIS), pages 1–6. IEEE, 2019. [15] I. Febin, K. Jayasree, and P. T. Joy. Violence detection in videos for an intelligent surveillance system using mobsift and movement filtering algorithm. Pattern Analysis and Applications, 23(2):611–623, 2020. [16] J.-C. Feng, F.-T. Hong, and W.-S. Zheng. Mist: Multiple instance self-training framework for video anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14009–14018, 2021. [17] Y. Gao, H. Liu, X. Sun, C. Wang, and Y. Liu. Violence detection using oriented violent flows. Image and vision computing, 48:37–41, 2016. [18] A. Hanson, K. Pnvr, S. Krishnagopal, and L. Davis. Bidirectional convolutional lstm for the detection of violence in videos. In Proceedings of the European conference on computer vision (ECCV) workshops, pages 0–0, 2018. [19] T. Hassner, Y. Itcher, and O. Kliper-Gross. Violent flows: Real-time detection of violent crowd behavior. In 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, pages 1–6, 2012. [20] B. Jiang, F. Xu, W. Tu, and C. Yang. Channel-wise attention in 3d convolutional networks for violence detection. In 2019 International Conference on Intelligent Computing and its Emerging Applications (ICEA), pages 59–64, 2019. [21] A. M. Kamoona, A. K. Gostar, A. Bab-Hadiashar, and R. Hoseinnezhad. Multiple instance-based video anomaly detection using deep temporal encoding– decoding. Expert Systems with Applications, 214:119079, 2023. [22] M.-S. Kang, R.-H. Park, and H.-M. Park. Efficient spatio-temporal modeling methods for real-time violence recognition. IEEE Access, 9:76270–76285, 2021. [23] S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah. Transformers in vision: A survey. ACM computing surveys (CSUR), 54(10s):1–41, 2022. [24] C. Kothari. Research methodology: Methods and techniques. New Age International, 2004. [25] S. M. Mohtavipour, M. Saeidi, and A. Arabsorkhi. A multi-stream cnn for deep violence detection in video sequences using handcrafted features. The Visual Computer, 38(6):2057–2072, 2022. [26] W.-F. Pang, Q.-H. He, Y.-j. Hu, and Y.-X. Li. Violence detection in videos based on fusing visual and audio information. In ICASSP 2021-2021 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 2260–2264. IEEE, 2021. [27] B. Peixoto, B. Lavi, P. Bestagini, Z. Dias, and A. Rocha. Multimodal violence detection in videos. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2957–2961. IEEE, 2020. [28] B. Peixoto, B. Lavi, J. P. P. Martin, S. Avila, Z. Dias, and A. Rocha. Toward subjective violence detection in videos. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8276–8280. IEEE, 2019. [29] M. Perez, A. C. Kot, and A. Rocha. Detection of real-world fights in surveillance videos. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2662–2666. IEEE, 2019. [30] Z. Qi, R. Zhu, Z. Fu, W. Chai, and V. Kindratenko. Weakly supervised twostage training scheme for deep video fight detection model. In 2022 IEEE 34th International Conference on Tools with Artificial Intelligence (ICTAI), pages 677–685. IEEE, 2022. [31] P. C. Ribeiro, R. Audigier, and Q. C. Pham. Rimoc, a feature to discriminate unstructured motions: Application to violence detection for video-surveillance. Computer vision and image understanding, 144:121–143, 2016. [32] M. Schedi, M. Sjöberg, I. Mironică, B. Ionescu, V. L. Quang, Y.-G. Jiang, and C.-H. Demarty. Vsd2014: A dataset for violent scenes detection in hollywood movies and web videos. In 2015 13th International Workshop on Content-Based Multimedia Indexing (CBMI), pages 1–6, 2015. [33] A. Shagufta, M. T. Hesham, S. Masood, and A. Abd El-latif. A vision transformer model for violence detection from real-time videos. In Proceedings of the 5th International Conference on Future Networks and Distributed Systems, pages 834–840, 2021. [34] M. M. Soliman, M. H. Kamal, M. A. E.-M. Nashed, Y. M. Mostafa, B. S. Chawky, and D. Khattab. Violence recognition from videos using deep learning techniques. In 2019 Ninth International Conference on Intelligent Computing and Information Systems (ICICIS), pages 80–85. IEEE, 2019. [35] W. Sultani, C. Chen, and M. Shah. Real-world anomaly detection in surveillance videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6 2018. [36] Y. Tian, G. Pang, Y. Chen, R. Singh, J. W. Verjans, and G. Carneiro. Weaklysupervised video anomaly detection with robust temporal feature magnitude learning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4975–4986, 2021. [37] A. Traoré and M. A. Akhloufi. Violence detection in videos using deep recurrent and convolutional neural networks. In 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pages 154–159, 2020. [38] A. Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017. [39] P. Wu, J. Liu, Y. Shi, Y. Sun, F. Shao, Z. Wu, and Z. Yang. Not only look, but also listen: Learning multimodal violence detection under weak supervision. In European conference on computer vision, pages 322–339. Springer, 2020. [40] P. Wu, X. Liu, and J. Liu. Weakly supervised audio-visual violence detection. IEEE Transactions on Multimedia, 25:1674–1685, 2023. [41] J. Yu, J. Liu, Y. Cheng, R. Feng, and Y. Zhang. Modality-aware contrastive instance learning with self-distillation for weakly-supervised audio-visual violence detection. In Proceedings of the 30th ACM international conference on multimedia, pages 6278–6287, 2022. [42] P. Zhou, Q. Ding, H. Luo, and X. Hou. Violence detection in surveillance video using low-level features. PLoS one, 13(10):e0203668, 2018.