scieee AI-readable full text Open interactive document viewer

Study and training of deep learning models for scene graph generation from non-static environments

Rosales Santana, Kevin David

Full text

Final Master’s Degree Thesis Study and Training of Deep Learning Models for Scene Graph Generation from non-static Environments Master’s Degree in Artificial Intelligence Author: Kevin David Rosales Santana Supervisor: Sergio ´ Alvarez Napagao, Barcelona Supercomputing Center Co-supervisor: Dmitry Gnatyshak, Barcelona Supercomputing Center Tutor: Ulises Cort´es Garc´ıa, Department of Computer Science, UPC Barcelona, 25 January 2022 What I cannot create, I do not understand. – Richard Feynman Abstract Recently, several proposals have managed to enable the generation of scene graphs from still images. Even if it may be believed that it is a solved problem, State-of-the-Art results indicate that they are still far from getting an accurate graph-based description of their images. This problem gets harder when complex environments are involved. In this study, data will be obtained in the form of videos, whose temporal aspect is usually dismissed for this task. With the arrival of deep learning, many fields covered by Artificial Intelligence have improved their results by large margins such as Computer Vision. Moreover, context-aware models as Recurrent Neural Networks or Transformers have provided a new way to handle sequential data. The main aim of this thesis is the study and training of this type of models to analyze the impact of the temporal dimension in the generation of scene graphs. Notwithstanding, large cascade or end-to-end models require an in-depth and sensible training process due to their different built-in parts. For example, when still images are being treated, two different tasks must be addressed: the objects and their relationships classification. Furthermore, this proposal includes the aforementioned context-aware models on the top, which must deal with the features extracted by the previous parts correctly. Finally, all results will be analyzed using both quantitative and qualitative metrics. For this purpose, the most utilized ones will be computed so as to compare them with other proposals, which are low and extremely difficult to replicate due to the newness of this field and the poor quantity of datasets available. On the other hand, our Human-Object Relationship LSTM (HORL) Model is able to outperform State-of-the-Art results and, as it is implemented using the official Scene Graph Generation benchmark from Microsoft as baseline, among others, it is able to provide a fairer way to demonstrate its enhancements with respect to already implemented still images approaches. Some visual examples will be included in the results in order to show a straightforward comparison against other studies, specially against the ones that do not take into account the temporal dimension of the sequential data, obtaining a more accurate description of these complex environments. Keywords: Scene Graph Generation ·Computer Vision ·Non-static Environments, Videos ·Recurrent Neural Networks ·Deep Learning ·Artificial Intelligence. Acknowledgments First of all, I would like to thank both of my thesis supervisors, Sergio ´ Alvarez and Dmitry Gnatyshak for all their support and advice during this project. It has been absolutely a pleasure to work with one of the greatest research centers in Spain, the Barcelona Supercomputing Center (BSC). Moreover, I wish to thank my tutor, Ulises Cort´es, for organizing this awesome Master’s Degree in Artificial Intelligence as well as its staff. I have really enjoyed this journey during this year and a half, getting the opportunity to refine my knowledge about artificial intelligence, increase my co-working abilities and inject enthusiasm into the beautiful research it contains. Furthermore, I would like to express my gratitude to my MAI colleagues, who have been providing an amazing support along the different courses as well. This help was completely necessary to success due to the impact of SARS-CoV-2 on our lives. In addition to this, I would like to mention, as I did in my bachelor’s thesis, the enormous work that the health personnel has been carrying out during these difficult times. Finally, I would like to thank my family and friends from Gran Canaria. Although they are really far, their unconditional support has encouraged me to continue on my path and overcome the difficulties that life entails. Specially, I cannot be more grateful to Alejandro Moreno. Thank you so much. Contents 1 Introduction 1 1.1 Motivation..................................... 1 1.2 Contribution.................................... 3 1.3 Overview...................................... 4 1.4 ChapterSummary ................................ 4 2 Background 5 2.1 SceneGraphGeneration ............................. 5 2.2 DeepLearning................................... 6 2.2.1 Multilayer Perceptron (MLP) . . . . . . . . . . . . . . . . . . . . . . 7 2.2.2 Convolutional Neural Network (CNN) . . . . . . . . . . . . . . . . . 7 2.2.3 Long Short-Term Memory (LSTM) . . . . . . . . . . . . . . . . . . . 8 2.2.4 AttentionMechanism........................... 9 2.3 Object Detection: Faster R-CNN . . . . . . . . . . . . . . . . . . . . . . . . 9 2.4 HORLPreview .................................. 10 2.5 ChapterSummary ................................ 11 3 Related Work 13 3.1 StaticEnvironments ............................... 13 3.1.1 SGG from Objects, Phrases and Region Captions . . . . . . . . . . . 13 3.1.2 SGG by Iterative Message Passing . . . . . . . . . . . . . . . . . . . 14 3.1.3 SGG using Neural Motifs . . . . . . . . . . . . . . . . . . . . . . . . . 16 3.1.4 Graph-RCNN for SGG . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.1.5 SGGfromGANs ............................. 18 3.1.6 Graphical Contrastive Losses for SGG . . . . . . . . . . . . . . . . . 20 3.2 Non-static Environments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.2.1 SGGusingHORT............................. 23 3.2.2 SGG from Target Adaptive Context Aggregation . . . . . . . . . . . 29 3.2.3 SGG using Spatial-Temporal Transformer . . . . . . . . . . . . . . . 32 3.3 ChapterSummary ................................ 34 4 Methodology 37 4.1 ExperimentSetup................................. 37 4.2 Baselines...................................... 38 4.2.1 Visual Genome Mapping Baseline . . . . . . . . . . . . . . . . . . . . 38 ix NLP Natural Language Processing. NMS Non-Maximum Suppression. RelDN Relationship Detector Network. RePN Relation Proposal Network. RNN Recurrent Neural Network. RPN Region Proposal Network. SGG Scene Graph Generation. TSV Tab-Separated Values. Acronyms Acronyms xviii Study and Training of DL Models for SGG from non-static Environments Chapter 1 Introduction This thesis aims to study the importance of the temporal dimension in sequential data in order to enhance the Scene Graph Generation (SGG) from non-static environments, which in this case, will refer to videos. This impact will be the result of the utilization of the latest enhancements from Deep Learning (DL), which will learn the required patterns so as to get a precise insight of the analyzed environment. This chapter will describe the main motivation behind the idea of this thesis (see Section 1.1), the contribution that it provides to the scientific community (see Section 1.2) and a brief overview of this document (see Section 1.3). 1.1 Motivation Scene Graph Generation (see Section 2.1) is a very useful way to describe media and transform it into a specific data structure that can be easily handled by computers. This generation of graphs that describe situations can support many real-world problems. For example, given an image, a symbolic data structure like graphs can be obtained instead of solely extracting its features by looking at its pixels. Consequently, an enhanced perception of the situation happening in the image can be acquired. For instance, the extraction of the behavior of agents in complex environments is not a trivial task. Graphs can be utilized in order to use their inherent triplets (i.e., data structure composed of <subject, relationship, object>. For example, <person, drinking, water>) to guide the reasoning or training of these agents. This concentrated representation of situations allows an easier treatment of visual data that can improve several machine learning-based tasks (see Section 2.1 for more details) such as Image Retrieval, Image Captioning, Visual Question Answering, Deep Reinforcement Learning explainability or Image Generation from graphs, among others. For that reason, this symbolic knowledge extracted from images has been gaining traction 1 CHAPTER 1. INTRODUCTION 1.1. MOTIVATION recently in the artificial intelligence community, who has been providing several solutions to improve the quality of the generated graphs in the different designed approaches. In this thesis, a specific set of scenes will be studied: the human-object relationships. It is crucial to detect, recognize and analyze the diverse set of interactions that people have with their environment. For example, in contexts like human-robot interaction, senior care or health care, it is pivotal to make sensible and adequate relations between the different elements involved in the scene. In fact, the human-object relationships are the main scenes that are generated in the aforementioned deep reinforcement learning approaches, where scenes graphs will aim to describe complex environments where both objects and agents are integrated and working cooperatively and/or competitively. Notwithstanding, this proposal will aim to describe more complex environments. In this thesis, we describe non-static environments as those environments that go beyond still images by using videos or sequential data (see Figure 1.1). These environments are much less studied than general ones (i.e., still images) and provide a harder task due to the main video-based Computer Vision issues such as frames quality, object and relationship recognition in different frames or temporal features that are usually dismissed in this type of approaches. Figure 1.1: SGG from a non-static environment [1] Therefore, the most important set of research questions that are studied in this thesis are: •Hypothesis 1:The context provided by the frames of a video can support an enhancement in the SGG of a specific frame from the same video. In other words, a context-aware model can generate more accurate Scene Graphs from a video than a context-agnostic one, demonstrating a better performance from non-static environments approaches. •Hypothesis 2:The object detector plays a very important role in the SGG computation. Therefore, in those metrics where the object detector must provide an inference, 2 Study and Training of DL Models for SGG from non-static Environments CHAPTER 1. INTRODUCTION 1.2. CONTRIBUTION a system that includes a more precise one will result in a high quality generation of graphs. •Hypothesis 3: The importance of a precise relationship classification in non-static environments is provided by the relationship head of the deep learning model, the semantic features and the context of the frame. Consequently, in order to obtain accurate relationships between a pair of objects, it is important to design a system that merges sensibly the semantic features (i.e., makes use of the empirical distribution bias of the relationships between objects in the dataset), the results provided by its relationship head from the deep learning model and the features provided by the surrounding frames. •Hypothesis 4: The relationships from the frames must be thoughtfully represented in order to transfer this context correctly between frames. In this hypothesis, the relationship representation is emphasized since this large quantity of information must be inserted appropriately in the context-aware model, removing noise or irrelevant features in the process. 1.2 Contribution Our main contribution is the utilization of relationships from previous and next frames in the computation of a specific one in order to enhance the current relationship detection. In other words, the creation of a deep learning context-aware model so as to take into account both spatial and temporal features in the resulting scene graph. Moreover, as it will take into consideration triplets from surrounding frames, objects that do not take part of the action in consecutive frames are not likely to appear as part of the generated triplets. Our proposal will include a detailed process of training and study of the different deep learning models that are going to compute the previously mentioned two stages. It must be taken into account that SGG literature usually misses a consensus between the different implemented approaches. For that reason, our implementation will be based on the official Scene Graph Generation Benchmark [2] from Microsoft [3] so as to provide fairness in our results by computing baselines that can be recognized worldwide. Moreover, video-based SGG State-of-the-Art contains several problems due to the newness of this field. For example, some approaches may contain unavailable code or models, which results in evaluations that cannot be replicated or others that transforms unfairly the ground truth (e.g., by removing those ground truth objects smaller than a specific area) in their evaluation. For that reason, another contribution of this thesis will be the possibility of replicating our results both due to the utilization of a general benchmark and to an implementation that can be easily used. Finally, this thesis will show some quantitative results about how our Human-Object Relationship LSTM (i.e., HORL) model is able to outperform State-of-the-art approaches and some qualitative results so as to demonstrate how irrelevant triplets are removed and a Study and Training of DL Models for SGG from non-static Environments 3 CHAPTER 1. INTRODUCTION 1.3. OVERVIEW more precise graph is generated, leading to an accurate description of the events taking place in a non-static environment or video. 1.3 Overview The first chapter has introduced this master thesis including its motivation and contributions to the artificial intelligence scientific community. This thesis contains five more chapters. First, Chapter 2 will introduce some necessary background knowledge in order to understand the methodology that will be explained along the document. After that, Chapter 3 will provide a study about related work regarding SGG approaches, including those that are the State-of-the-art, depicting their different results. Then, Chapter 4 will describe our main proposal (i.e., the HORL model) as well as technical aspects about the training and study of the design and development of the diverse models that were built for this thesis. Moreover, Chapter 5 will analyze the different experiments that were performed along the thesis together with their results and their appropriate in-depth comparison with State-of-the-art models. Finally, Chapter 6 will conclude the thesis and provide some sensible future lines of research. The Appendix A will contain some more qualitative results apart from the ones provided in Chapter 5. 1.4 Chapter Summary In this first chapter, the main motivation behind our research has been explained, describing why non-static (i.e., video) SGG is an important field in artificial intelligence and its different applications. Furthermore, our main contribution has been detailed by introducing why context-aware deep learning models are able to support the non-static SGG, stating the new State-of-the-Art results obtained by our proposal. Finally, after explaining why our results provide some fairness due to the utilization of an official benchmark, a brief overview of the structure of the entire document has been provided. Next chapter will explain some background knowledge that may be required for a full understanding of this thesis. 4 Study and Training of DL Models for SGG from non-static Environments Chapter 2 Background As it was mentioned in Chapter 1, some background knowledge will be briefly introduced so as to understand correctly the methodology followed by this thesis. Please, all the content from this chapter is a gentle description of different concepts that will be discussed along the document, so it is strongly encouraged to revise references for in-depth explanations if necessary. Namely, this chapter will describe more details about Scene Graph Generation (see Section 2.1) and several deep learning architectures and mechanisms that will be utilized in our proposal (see Section 2.2). Furthermore, since the majority of SGG developed models are based on the utilization of Faster R-CNN, it will be briefly explained in this chapter as well (see Section 2.3). Finally, a brief preview of our major contribution, the HORL system (see Section 2.4), will be depicted too. 2.1 Scene Graph Generation As it was mentioned in the previous chapter, the Scene Graph Generation (SGG) objective is the obtainment of a representation of a situation or scene through graphs. A graph is a structure composed of nodes (i.e., objects in this thesis) and edges (i.e., relationships in this thesis). These semantics result in a symbolic method that can be utilized to support many subsymbolic outcomes. For example, in the Image Retrieval task, such as the reverse search from Google Images [4], it provides better results the comparison between graphs describing the images rather than a pixel-by-pixel comparison. Another example related to the description of images can be the Image Captioning task, where a machine must describe in a sentence what is occurring in the image [5]. Furthermore, other type of objectives that can be tackled by Scene Graph Generation are related to Visual Question Answering, where a machine must provide a response about diverse situations that may be happening in the input image or action recognition, where the different actions taking part in the image must be detected and classified. 5 CHAPTER 2. BACKGROUND 2.2. DEEP LEARNING Moreover, due to its potential in describing situations, Scene Graph Generation has become a future line of research in explainability, specially on deep learning models. For instance, deep reinforcement learning [6] approaches usually contain agents that must learn and fulfill several tasks. In this context, the agent may behave in some unintelligible way for the scientific community. If the set of actions performed by this agent (or these agents) can be shown by images, Scene Graph Generation can be used to describe this behavior using graphs, which is more understandable by both humans and computers. Finally, deep learning-based approaches on Image Generation from graphs [7] have become popular due to their qualitative results and the set of possibilities that it may provide to image understanding or techniques like data augmentation. Although it can be believed that Scene Graph Generation is already a solved task, it is still far away from considering it as a closed problem. The recognition of the different objects and the identification of their relationships in a graph contain many different stages from the point of view of artificial intelligence. Namely, the different objects must be recognized. Despite the fact that there are many object detectors that generate accurate results such as YOLO [8], Faster R-CNN [9] or SDD [10], the results get much worse when SGG steps in. Since not every detected object actually takes part in some relation, there are many false positives from the point of view of graph-based representations. In fact, this error may be propagated to future steps that depend on the objects that were detected in the first one. Moreover, the confidence value provided by the object detector is one of the values that is utilized in the computation of the confidence values of the different triplets to construct the graph. On the other hand, the second stage is the relationships detection. Usually, this step is computed by extracting the different features of the image composed by both elements proposed in the triplet. As it was mentioned in Section 1.1, two types of environments are defined in this proposal: static environments (i.e., still or isolated images) and non-static environments (i.e., sequence of images or videos). This thesis will study the treatment of the latter, since its context (i.e., surrounding frames) must be taken into account in its research. 2.2 Deep Learning In this section, some concepts about deep learning will be defined, starting from popular architectures that were used for this thesis (see Sections 2.2.1, 2.2.2 and 2.2.3) and finishing with the attention mechanism (see Section 2.2.4), which has become popular due to the recent appearance of transformers. 6 Study and Training of DL Models for SGG from non-static Environments CHAPTER 2. BACKGROUND 2.2. DEEP LEARNING 2.2.1 Multilayer Perceptron (MLP) An Artificial Neural Network (ANN) is a system based on biological neural networks from animal brains. Namely, it is a set of connected units that send signals to other neurons, which in this case, it is the result of a weighted sum of its inputs together with a bias term activated with a function (see Equation 2.1). f n X i=0 xiwi!(2.1) Where nis the number of inputs, xare the inputs and ware the weights (x0is 1, so that w0 is the bias factor). ANN [11, 12] has become popular due to their capability of learning representations of data by learning their set of weights. A MultiLayer Perceptron (MLP) [13] is a feedforward ANN that combines neurons to build layers of nodes which can be: the input layer, the hidden layers or the output layer. Backpropagation algorithm [14] is used to enhance the weights in a supervised learning approach by using an optimizer and a loss function. This combination of layers allows MLP to distinguish data that may not be linearly separable. Our baselines (see Section 4.2), object detector (see Section 4.4), relationship classifier (see Section 4.5) and HORL system (see Section 4.6) are based on artificial neural networks. 2.2.2 Convolutional Neural Network (CNN) A Convolutional Neural Network (CNN) [15, 16] is a class of ANN that is commonly used for image-based machine learning tasks. Since traditional ANN are not able to extract local patterns of data, a CNN recognizes them using kernels over the image. These kernels are matrix-based filters applied on the data so as to get a better numerical representation of the image that takes into consideration the aforementioned local context (see Figure 2.1). Figure 2.1: Example of CNN [15] Study and Training of DL Models for SGG from non-static Environments 7 CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS Figure 3.1: MSDN overview [24] In short, the MSDN (see Figure 3.1) contains the following stages: 1. The object, phrase and caption region proposals are computed. It must be taken into account that phrase regions are generated by grouping object regions into pairs. The caption region proposal is equivalent to the entire input image. 2. The convolutional layers from VGG-16 are shared with other parts. Then, each branch contains its own features for its specific task in the feature specialization. 3. A dynamic graph is constructed using these different semantic levels. 4. All the features are mixed so as to refine them. This process is named object, phrase and caption refining. 5. The final prediction is computed. If the object iand jare not classified as background and the predicate (i, j) is not irrelevant, then the two objects iand jare connected through the predicate (i, j). For the region caption generation branch, they use a LSTM-based language model to generate natural sentences to describe the region from the region features. This approach is the only one from the well-known literature that tries to tackle the SGG problem using Natural Language Processing techniques such as region captioning, allowing the obtainment of different scene understanding tasks solutions together. 3.1.2 SGG by Iterative Message Passing In 2017, another proposal named Scene Graph Generation by Iterative Message Passing [27] by Danfei Xu et al. was published. In this case, the idea is to use a model that solves the scene graph inference problem using standard RNNs and learns to iteratively improve its prediction via message passing mechanism. Namely, their major contribution is that, instead of inferring each component of a scene graph in isolation, the model passes messages containing contextual information 14 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS between a pair of bipartite sub-graphs of the scene graph, and iteratively refines its predictions using RNNs. The Regions of Interest (RoI) are obtained using a Region Proposal Network from Faster R-CNN (see Section 2.3) to automatically generate a set of object bounding box proposals from an image as the base input to the inference procedure. They formulate the SGG problem as finding the optimal x∗, which is the argmax of the probabilities computed by: Pr(x|I, BI) = Πi∈VΠj=iP r(Xcls i, xbbox i, xi→j|I, BI) (3.1) Where X(i→j) is the relationship predicate between the i-th and the j-th proposal boxes, Iis the image and BIare the box proposals. First, they obtain the RoI. Then, they use generic RNN units (in particular, Gated Recurrent Units (GRU) [28]) to take previous hidden states and an incoming message as input in order to produce a new hidden state as output. Each node and edge in the scene graph maintains its internal state in its corresponding GRU unit, where all nodes share the same GRU weights (i.e., node GRUs) and all edges share the other set of GRU weights (i.e., edge GRUs). This setup allows the model to pass messages among the GRU units along the scene graph topology. They also propose a message pooling function that learns to dynamically aggregate the hidden state of the GRUs into messages. Furthermore, since nodes and edges compose two disjoint sub-graphs that are essentially the dual graph to each other, the authors consider that the primal graph defines channels for messages to pass from edge GRUs to node GRUs and the dual graph defines channels for messages to pass from node GRUs to edge GRUs. Figure 3.2: IMP overview [27] As it was previously explained, they use a hidden state (i.e., a high dimensional vector) of the corresponding GRU to represent the current state of each node and edge. As each GRU type (i.e., node GRU and edge GRU) share the same update rule, each of them share the same set of parameters. Study and Training of DL Models for SGG from non-static Environments 15 CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS They use mean field1to perform approximate inference and its distribution can be formulated as: Q(x|I, BI) = Πn i=1Q(xcls i, xbbox i|hi)Q(hi|fv i)∗Πj=iQ(xi→j|hi→j)Q(hi→j|fe i→j) (3.2) Where fv iis the visual feature of the i-th node, and fe i→jis the visual feature of the edge from the i-th node to the j-th node. The current hidden state of node iis hiand the current hidden state of edge i→jis hi→j. As it was mentioned before, these visual features are extracted by a RoI Pooling Layer from the image. In this scene graph topology (see Figure 3.2), the neighbors of the edge GRUs are node GRUs, and the other way around. In order to pass messages along the previously mentioned two disjoint sub-graphs, they consider the node-centric primal graph, in which each node GRU gets messages from its inbound and outbound edge GRUs and the edge-centric dual graph, where each edge GRU gets messages from its subject node GRU and object node GRU. Consequently, they are able to improve the inference efficiency by iteratively passing messages between these two sub-graphs avoiding the utilization of a densely connected graph. Since each GRU receives multiple incoming messages, they aggregate them using adaptive weights that can modulate the influences of incoming messages and only keep the relevant information. Namely, they use a message pooling function that computes the weight factors for each incoming message and fuse the messages using a weighted sum. Finally, they use a softmax activation to produce the final scores for the object class as well as for the relationship predicate. 3.1.3 SGG using Neural Motifs One year later, in 2018, Rowan Zellers et al. proposed their neural motifs approach in their paper Neural Motifs: Scene Graph Parsing with Global Context [30, 31]. This proposal tries to generate Scene Graphs analyzing “motifs” (i.e., regularly appearing substructures). It states that object labels are able to help predicting the relation labels and that there are several triplets that are repeated through the entire dataset. For that reason, they try to predict the most frequent relation between object pairs with the given labels in their object detections. Furthermore, they propose an end-to-end architecture named Stacked Motif Network or MotifNet. Their architecture tries to use Global Context so as to take into account useful knowledge from other parts of the image. Namely, they use Faster R-CNN (see Section 2.3) to take the different region proposals identifying the different objects. After that, the different regions are inserted into an architecture composed of a Bidirectional LSTM, a LSTM and a Bidirectional LSTM again (see Section 2.2.3 and Figure 3.3 for more information). 1Approximation strategy developed in the statistical physics literature that allows the research of highdimensional stochastic models by studying a simpler model that is similar to the original by averaging over degrees of freedom [29]. 16 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS It must be taken into account that the LSTM is used in order to link the object and the edge context parts. Several orderings of the region proposals into the LSTM are studied, concluding that the best performance is obtained by ordering them from left to right using the xaxis of the image. Figure 3.3: MotifNet overview [30] Moreover, they provide an online demo [32]. The utilization of RNN-based architectures was relevant for the development of this thesis since, in this case, the context is obtained using the different frames from the input video. 3.1.4 Graph-RCNN for SGG In 2018, Jianwei Jang et al. proposed the paper Graph-RCNN for Scene Graph Generation [33, 34] in collaboration with Facebook AI Research [35]. Their proposal contains a Relation Proposal Network (RePN) that efficiently deals with the possible quadratic number of potential relations between objects in an image and an Attentional Graph Convolutional Network (aGCN) that effectively captures contextual information between objects and relations. Figure 3.4: Graph-RCNN overview [33] Study and Training of DL Models for SGG from non-static Environments 17 CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS Namely, their model, which is depicted in Figure 3.4, can be summarized in three different stages: 1. The object nodes are extracted using Faster R-CNN (see Section 2.3). 2. The relationship edges are pruned using the RePN. This network learns to compute relatedness scores between object pairs which are used to intelligently prune unlikely scene graph connections. They consider an asymmetric kernel function composed of projection functions for subjects and objects in the relationships. Specifically, they use two MLP with identical architecture. After obtaining the score matrix for all object pairs, they sort the scores in descending order and choose top Kpairs. The Non- Maximum Suppression (NMS) technique is applied to filter out pairs that have significant overlap with others. 3. The graph context is integrated using the aGCN. This network propagates higher-order context throughout the graph, updating each object and relationship representation based on its neighbors. They extend the conventional Graph Convolutional Network (GCN) [36] to an attentional version (see Section 2.2.4) by adjusting the adjacency matrix values. In order to predict attention from node features, they use a 2-layer MLP over concatenated node features and compute a softmax over the resulting scores: uij =wT hσ(Wa[z(l) i, z(l) j]) αi=softmax(ui)(3.3) Where Wand ware learned parameters and zis the representation of a node. Consequently, this SGG approach can be considered as: P(S|I) = P(V|I)∗P(E|V, I)∗P(R, O|V, E, I) (3.4) Where Sis the Scene Graph, Idenote an image, Vis the set of nodes corresponding to localized object regions in I,Edenote the edges (i.e., the relationships) between objects, and Oand Rdenote object and relationship labels respectively. It has to be taken into account that P(V|I) is generated by the Object Region Proposal, P(E|V, I) is generated by the Relationship Proposal and P(R, O|V, E, I) is generated by the Graph Labeling. 3.1.5 SGG from GANs One year later, in 2019, Matthew Klawonn et al. published in Combining Supervised Machine Learning and Structured Knowledge for Difficult Perceptual Tasks [37] a proposal about the utilization of GANs [38] for SGG purposes. Usually, in SGG research, two types of nodes are distinguished: the objects (e.g.,car or tree) and attributes (e.g.,red). This proposal combines the utilization of RNN, CNN and GANs to generate both types of triplets. The main difference of this approach is that it is a bounding-box free method and, consequently, it does not determine the exact spatial location 18 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS of an object in an image, reducing this error in the entire system. First, the Triplet Producing Model pipeline is run as follows: 1. The feature extractor maps images to visual features using a CNN. The proposal evaluates this phase using two versions: one that pre-trains feature detectors on the task of image classification and another that jointly trains the CNN component of the network together with the RNN part. They also evaluate the insertion of pooling layers to introduce invariance to transformations of pixel data. 2. After that, the recurrent component (i.e., the RNN) produces output triplets using a LSTM (see Section 2.2.3) of three time steps so as to get the triplet size. It must be taken into account that only 1 triplet is generated per input image inserted into the CNN. Therefore, the GAN is trained in order to generate more than 1 triplet. Figure 3.5: GAN-based Scene Graph Generator overview [37] Both CNN and RNN components are combined to form the GAN generator as it can be observed in Figure 3.5. Then, an attention mechanism (see Section 2.2.4) is applied in order to combine the triplets into the proper scene graph. Namely, if two lexemes with the same label were produced from the same spatial region of the input using the aforementioned attention, it is merged into a single node in the Graph Producing Model. The same system is replicated so as to get the attributes too. In this case, they use ground truth triplets that represents attributes instead of relations in the training phase. The original code was written in TensorFlow 1.13 [39]. For this project, we upgraded it to TensorFlow 2.4 in order to take advantage of the latest versions of CUDA [40] and cuDNN [41]. The model is ready to be deployed to the Mininostrum cluster from Barcelona Supercomputing Center (BSC) [42] in order to be trained. Study and Training of DL Models for SGG from non-static Environments 19 CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS 3.1.6 Graphical Contrastive Losses for SGG The last approach for static environments that will be described in this chapter is the paper Graphical Contrastive Losses for Scene Graph Parsing [43, 44] by Ji Zhang et al. in collaboration with NVIDIA Corporation [45] in 2019. Since this proposal is the one that will be used for relationship classification in our experiments, it will be explained thoroughly. In the paper, authors state that usually SGG models suffer from two common errors: •Entity Instance Confusion: occurs when the model confuses multiple instances of the same type of entity. In other words, the subject or object is related to one of many instances of the same class and the model fails to distinguish between the target instance and the others (see Figure 3.6a). •Proximal Relationship Ambiguity: arises when multiple subject-predicate-object triplets appear in close proximity with the same predicate and the models struggle to infer the correct subject-object pairings. In other words, the image contains multiple subjectobject-pairs interacting in the same way and the models fail to identify the correct pairing (see Figure 3.6b). (a) Entity Instance Confusion (b) Proximal Relationship Ambiguity Figure 3.6: Common errors in SGG [43] This proposal includes a set of contrastive loss formulations that explicitly force the model to disambiguate related and unrelated instances through margin constraints specific to each type of confusion. 20 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS The aforementioned losses compensate drawbacks by contrasting positive against negative edges for each node, providing global supervision to the classifier and alleviating both issues. Namely, the graphical contrastive losses are composed of three types of losses: 1. Class Agnostic (L1): contrasts positive/negative entity pairs regardless of their relation and adds contrastive supervision for generic cases. It aims to maximize the affinity of the lowest scoring positive pairing and minimize the affinity of the highest scoring negative pairing. They randomly sample at most K non-related objects (i.e., negative pairings), so the complexity is O(NK). 2. Entity Class Aware (L2): addresses the Entity Instance Confusion by focusing entities with the same class. It is an extension of the Class Agnostic loss where they specify a class when populating the positive and negative sets. As L1, its complexity is O(NK). 3. Predicate Class Aware (L3): addresses the Proximal Relationship Ambiguity by focusing on entity pairs with the same potential predicate. This loss maximizes the margins within groups of instances determined by their associated predicates. As L1 and L2, its complexity is O(NK). The final loss is expressed as: L=L0+λ1L1+λ2L2+λ3L3(3.5) Where L0is the cross-entropy loss over predicate classes. Figure 3.7: RelDN overview [43] Moreover, they create a relationship detector named Relationship Detector Network (RelDN) (see Figure 3.7) that works as follows: 1. First, it identifies a proposal set of likely subject-object relationship pairs. It returns bounding box regions containing every pair. 2. Then, features are extracted from these candidate regions to perform a fine-grained classification into a predicate class. The RelDN contains a CNN branch for predicates to extract its features with the same structure of the entity detector CNN branch. Study and Training of DL Models for SGG from non-static Environments 21 CHAPTER 3. RELATED WORK 3.1. STATIC ENVIRONMENTS They build these two separate branches in order to get visual features for predicates that focus on the interactive areas of subjects and objects as opposed to individual entities. Figure 3.8 illustrates the CNN features over the two branches. It computes three types of features for each relationship proposal: •Semantic features: it conditions the predicate class predictions on subject-object class co-occurrence frequencies, as [30] (see Section 3.1.3) studied. For each training image, they count the occurrences of predicate class given subject and object classes in the Ground Truth annotations, providing an empirical distribution. •Spatial features: it conditions the predicate class predictions on the relative positions of the subject and the object. Therefore, they capture spatial information by encoding the box coordinates of subjects and objects using the box delta and normalized coordinates. This feature vector is fed through an MLP to attain predicate class logit scores. •Visual features: it produces a set of class logits conditioned RoI feature maps. They extract subject and object RoI features from the entity detector’s convolution layers and extract predicate RoI features from the relationship convolution layers. The subject, object and predicate features vectors are concatenated and passed through an MLP to attain the predicate class logits. They include two skipconnections projecting subject-only and object-only RoI features to the predicate class logits due to the observation of many relationships that can be accurately inferred by the appearance of only the subjects or objects. (a) Ground Truth (b) conv body det (c) conv body rel Figure 3.8: Visualization of CNN features from the two RelDN branches [43] The final probability distribution over predicate classes is obtained by adding the three scores followed by a softmax normalization: ppred =softmax(fvisual +fspatial +fsemantic) (3.6) 22 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.2. NON-STATIC ENVIRONMENTS Finally, they rank relationship proposals (i.e., triplets) by multiplying the predicted subject, object and predicate probabilities. The subject and object probabilities are obtained from the entity detector. Figure 3.9 depicts some example results of RelDN with L0as the only loss and together with the proposed losses. (a) Entity Instance Confusion (b) Proximal Relationship Ambiguity Figure 3.9: Example results of RelDN [43] 3.2 Non-static Environments This section will discuss some approaches that were developed to handle the generation of scene graphs from non-static environments (i.e., sequences of images or videos). Since this is a new topic in artificial intelligence research, only a few studies have been proposed. 3.2.1 SGG using HORT In 2021, the paper Detecting Human-Object Relationship in Videos by Jingwei Ji et al. created a new State-of-the-Art in SGG from non-static environments using the dataset they presented the previous year, the Action Genome [1] dataset (see Section 5.1.2). The idea of this proposal is the consideration of the temporal dynamics that can be crucial so as to contextualize the human-object relationships. Moreover, most of the existing SGG models implicitly assume single-class relationships between each pair of objects, while this does not always hold. For example, in Action Genome [1], a triplet can be <person, looking at - holding - eating, food>. In fact, this assumption comes mainly from the popular dataset Visual Genome [46] (see Section 5.1.1), where each pair of objects contains only a single relationship. However, Action Genome usually contains more than one relation between each person and object. Study and Training of DL Models for SGG from non-static Environments 23 CHAPTER 3. RELATED WORK 3.2. NON-STATIC ENVIRONMENTS extracted using a 2D CNN and the temporal ones in the clip are extracted with a 3D CNN. The static object features are combined with word-embedding for the subsequent blocks. Figure 3.16: Hierarchical Relation Tree visualization example [50] After that, as for the HRTree (see Figure 3.16), it first provides an adaptive structure for organizing the possible relation pairs efficiently, guiding the context aggregation module to capture spatio-temporal structure information. The leaf nodes in this tree represent the objects detected in the center frame. On the other hand, the non-leaf nodes are derived from their child nodes and represent their composite relations. Then, regarding the Target-adaptive Context Aggregation block, it obtains a contextualized feature representation for each relation pair in order to build a classification head to recognize its relation category. A temporal attentive module for the fusion of temporal features is designed in order to support a directional spatial aggregation module, which propagates the context information: 1. The temporal fusion module extracts temporal features with a box tube. A 3D CNN is used in order to extract spatio-temporal features to provide motion information for relation candidates. Their experiments describe how this approach is effective to certain types of motion relations like writing on or carrying. However, for some short-term relation classification, its enhancement is too slight. 2. The spatial propagation module adopts a group tree-GRU [28] scheme for context aggregation in a bidirectional propagation way. 30 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.2. NON-STATIC ENVIRONMENTS 3. Finally, the classification head (see Figure 3.17), which consists of four branches (i.e., Visual Branch, Fusion Branch, Subject/Object Branch and Statistical Prior Branch), provides a final score by summing the individual results from them and using a sigmoid activation function. Figure 3.17: Illustration of the classification head [50] This approach can be easily extended to a single graph for the entire video with a temporal association strategy: 1. First, the long video clip is divided into overlapping video segments. 2. After that, the tracking is performed on each segment. •A quarter of the frames of a video segment is sampled with frame-level scene graphs for the linking. 3. Finally, if one triplet appears in only one frame, it is directly counted with its predicted score. For triplets with the same predicted categories in multiple frames, the triplet is counted once with summed scores if their subjects and objects belong to the same trajectory, respectively. As for the entire video, the triplets among two neighboring segments are associated if their predicted categories are the same and their Intersection over Union (IoU) of subject/object trajectories is over a threshold of 0.5. Study and Training of DL Models for SGG from non-static Environments 31 CHAPTER 3. RELATED WORK 3.2. NON-STATIC ENVIRONMENTS It must be explained that, since most objects in Action Genome [1] touch each other, they only predict the relations in pairs with overlapped bounding boxes for the SGDet (see Section 5.2). The idea of using relationships from the context to enhance the quality of the SGG from individual frames will be utilized in our approach (see Section 4.6). 3.2.3 SGG using Spatial-Temporal Transformer In 2021, the proposal Spatial-Temporal Transformer for Dynamic Scene Graph Generation [52, 53] by Yuren Cong et al. described another approach based on Transformers to generate scene graphs from non-static environments. In this case, their Spatial-Temporal Transformer (STTran) is a neural network that consists of two core modules: a spatial encoder that takes an input frame to extract the spatial context and the visual relationships within a frame and a temporal decoder that captures the temporal dependencies between frames and infer the dynamic relationships. Their underlying idea is based on taking these temporal dependencies since they play an important role, causing that static scene graph generation methods are not directly applicable to non-static SGG. Moreover, Action Genome [1] may contain multiple relationships between a pair of entities. For that reason, they apply multi-label classification in relationship prediction and design a new strategy to generate a dynamic scene graph with confident predictions. Furthermore, they are able to verify that these temporal dependencies have a positive effect on relationship prediction. Regarding their methodology, they state that a dynamic scene graph Gdyn(Vt, Et) can be modeled as a static scene graph Gstat(V, E) with an extra index trepresenting the relations over time as an extra temporal axis. Therefore, they use transformers since the architecture is permutation-invariant and the sequence is compatible with positional encoding. As for the relationship representation, their representation vector xk tof the relation rk t between the i-th and j-th object proposals contains visual appearances, spatial information and semantic embeddings: xk t=⟨Wsvi t, Wovj t, Wuφ(uij t⊕fbox(bi t, bj t)), si t, sj t⟩(3.10) Where ⟨,⟩is concatenation operation, φis flattening operation and ⊕is element-wise addition. Ws,Wo∈IR2048x512 and Wu∈IR12544x512 represent the linear matrices for dimension compression. uij t∈IR256x7x7indicates the feature map of the union box computed by RoIAlign while fbox is the function transforming the bounding boxes of subject and object to an entire feature with the same shape as uij t. The semantic embedding vectors si t, sj t∈IR200 are determined by the object categories of subject and object. Furthermore, vi t∈IR2048 and bi t are the visual features and the bounding boxes provided by the detector. 32 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.2. NON-STATIC ENVIRONMENTS Figure 3.18: STTran overview [52] Their Spatio-Temporal Transformer (see Figure 3.18) preserves the original encoderdecoder architecture: 1. The spatial encoder concentrates on the spatial context within a frame whose input is a single Xt={x1 t, x2 t, ..., xK(t) t}where K(t) is the number of object proposals in the frame. Unlike the majority of transformer methods, no additional position encoding is integrated into the inputs since the relationships within a frame are intuitively happening at the same time. 2. However, they introduce frame encodings to inject the temporal position in the relationship representations. The frame encodings Efare constructed with learned embedding parameters, since the amount of embedding vectors depending on the window size ηin the temporal decoder is fixed and relatively short: Ef= [e1, ..., eη], where e1, ..., eη∈IR1936 are the learned vectors with the same length as xk t. The window size ηis fixed and therefore the video length does not affect the length of frame encodings. 3. Finally, the temporal decoder captures the temporal dependencies between frames. They adopt a sliding window to batch the frames so that the message is passed between the adjacent frames in order to avoid interference with distant frames. This sliding window of size ηruns over the sequence of spatial contextualized representations [X1, ..., XT] and the i-th generated input batch is presented as: Zi= [Xi, ..., Xi+η−1], i ∈ {1, ..., T −η+ 1}(3.11) The decoder consists of Nstacked identical self-attention layers Attdec() similar to the encoder structure. Therefore, in the first layer, the computation is as follows: Q=K=Zi+Ef V=Zi ˆ Zi=Attdec(Q, K, V ) (3.12) The output from the last decoder layer is adopted as the final prediction. Since a sliding window is being used, the relationships in a frame have various representations in different batches. In their proposal, they select the earliest representation appearing in the windows. Study and Training of DL Models for SGG from non-static Environments 33 CHAPTER 3. RELATED WORK 3.3. CHAPTER SUMMARY As for the loss function, they introduce the multi-label margin loss function for predicate classification as follows: Lp(r, P +, P−) = X p∈P+X q∈P− max(0,1−∅(r, p) + ∅(r, q)) (3.13) For a subject-object pair r,P+are the annotated predicates while P−is the set of the predicates that are not in the annotation. ∅(r, p) indicates the computed confidence score of the p-th predicate. During training, the object distribution is computed by two fullyconnected layers with a ReLU activation and a batch normalization in between. The standard cross entropy loss L0is utilized, so the total loss is as follows: Ltotal =Lp+Lo(3.14) At test time, as it has been done with other proposals, the score of each relationship triplet <subject, predicate, object>is computed as: srel =ssub ∗sp∗sobj (3.15) Where ssub,spand sobj are the confidence score of subject, predicate and object respectively. Notice that this approach modify the ground truth by only keeping those bounding boxes with short edges larger than 16 pixels. For that reason, a direct comparison cannot be made in Section 5.3. Finally, this paper states that some of their predictions do not match the ground truth relationships because they are annotated wrongly. Moreover, other relationships from the dataset are ambiguous and difficult to be identified even by humans. 3.3 Chapter Summary In this chapter, some relevant related proposals have been described (see Table 3.1 for a brief comparison). Namely, depending on the type of data that is being faced, static or non-static approaches have been developed in this field. Moreover, this chapter has provided some insight into the main elements that a SGG may contain together with the main problems that this task involves. On the other hand, this chapter depicts how the final objective of most of the proposals is the obtainment of an ordered set of triplets that describe what is happening in each scene, resulting in a final graph that describes the situation in a more appropriate way for machines. Next chapter will discuss our different approaches, giving some insight into their different motivations and an in-depth explanation of their underlying algorithms. 34 Study and Training of DL Models for SGG from non-static Environments CHAPTER 3. RELATED WORK 3.3. CHAPTER SUMMARY Environment Proposal Year Key Element Static MSDN [24] 2017 Different semantic levels in a dynamic graph together with their refinement by mixing them. IMP [27] 2017 Message passing mechanism with image context information using GRU (RNN) units. MotifNet [30] 2018 Use of motifs (most frequent relation between objects given their labels) and global context in a BiLSTM. Graph-RCNN [33] 2018 Utilization of a Relation Proposal Network to prune relationship edges and an attentional Graph Convolutional Network. GAN-based SGG [37] 2019 RNNs, CNNs and GANs to generate objects and attributes triplets. An attention mechanism is used to merge graph nodes. RelDN [43] 2019 Contrastive losses to solve Entity Instance Confusion and Proximal Relationship Ambiguity. The system contains a Spatial Module, a Semantic Module and a Visual Module. This proposal is used in our approach. Non-static HORT [47] 2021 Intra- and Inter-Transformers to enable joint spatial and temporal reasoning on multiple visual concepts of objects, relationships and visual poses along the sequence of images. No code or checkpoints have been provided yet. TRACE [50] 2021 Detect-to-track paradigm using a Hierarchical Relation Tree (HRTree) and a Target-adaptive Context Aggretation block. The spatial features are extracted using a 2D CNN and the temporal ones are extracted with a 3D CNN. STTran [52] 2021 Spatial Encoder, which extracts the spatial context and the visual relationships within a frame and Temporal Decoder, which captures the temporal dependencies between frames and infer the dynamic relationships. They modify the ground truth in their experiments. HORL (Ours, see 2.4) 2022 Attentional BiLSTM, which refines a representation of the relationships generated using static approaches by taking into consideration predicted triplets from the surrounding frames. Table 3.1: Summary of the explained related work Study and Training of DL Models for SGG from non-static Environments 35 CHAPTER 3. RELATED WORK 3.3. CHAPTER SUMMARY 36 Study and Training of DL Models for SGG from non-static Environments Chapter 4 Methodology After explaining some necessary background knowledge (see Chapter 2) and related work (see Chapter 3), Chapter 4 will describe the different approaches that were designed and developed during this thesis, emphasizing in the study and the training of the different deep learning models in order to discuss their results in Chapter 5. Namely, Section 4.1 will describe briefly our experiment setup. Then, Section 4.2 will show some initial baselines that were developed in order to measure the difficulty of the problem that was being faced. After explaining the data preprocessing in Section 4.3, Section 4.4 will describe the training process of our object detector so as to learn how to detect and classify the different active objects of each video frame effectively. Afterward, Section 4.5 will provide some insight into the design of the training stage of our relationship backbone and head, which will be able to provide confidence values to the different relationships that may take part in an interaction between a person and an object. Finally, Section 4.6 will explain our major contribution in this thesis: the design of a Human-Object Relationship attentional BiLSTM model to deal with the context of a frame in a complex non-static environment. 4.1 Experiment Setup Before explaining the baselines, our experiment-setup will be briefly described. First of all, Table 4.1 describes the specifications regarding the hardware and operating system. CPU Intel i7-10875H GPU NVIDIA RTX 2060, GDDR6 6GB Memory DDR4 16GB*2 OS Linux Ubuntu 20.04.3 LTS Table 4.1: Specifications of the PC used for the experiments As for the software environment, IntelliJ IDEA Ultimate [54] was used as IDE. Moreover, we make use of Python 3.8.10 [55] and the main libraries that were utilized for the project were numpy [56], scikit-learn [57], pytorch [58], tensorflow [39] and matplotlib [59], although a detailed requirements file can be found in the repository, which was managed by 37 CHAPTER 4. METHODOLOGY 4.2. BASELINES git [60] and GitLab [61]. As it was mentioned in Chapter 1, the majority of the code was implemented over the official SGG Benchmark [2] from Microsoft [3] in order to provide fairness in our experiments. The official metrics, which will be explained in Section 5.2, are collected using the aforementioned benchmark. 4.2 Baselines As it has been discussed along this document, the SGG from non-static environments is a recent topic in artificial intelligence that has not received much attention from the research community yet for two main reasons: the lack of populated datasets and the poor results that static environments usually obtain. Static environments proposals usually compare their performance using the Visual Genome dataset [46] (see Section 5.1.1). In 2020, when the Action Genome dataset [1] (see Section 5.1.2) was published, the scientific community started to make some interesting proposals that were described in Section 3.2. Nevertheless, all these approaches contain problems from the point of view of reproducibility of the results since some of them do not provide code nor checkpoints (see Section 3.2.1) or modify the ground truth labeling by removing those objects smaller than a specific area (see Section 3.2.3). For that reason, two quick baselines were developed to understand the complexity of this scenario. 4.2.1 Visual Genome Mapping Baseline The first approach that was developed was a baseline based on a mapping between the results of a Visual Genome [46] pre-trained model and the Action Genome dataset. The general workflow is depicted in Figure 4.1. Figure 4.1: Overview of Visual Genome Mapping Baseline The idea behind this model is the obtainment of a fast result that does not take into account the context of the non-static environment and that does not require a new training 38 Study and Training of DL Models for SGG from non-static Environments CHAPTER 4. METHODOLOGY 4.2. BASELINES of deep learning models. It works as follows: 1. First, individual frames from videos from the Action Genome [1] dataset are obtained. 2. Scene Graphs are predicted using already pre-trained models in PyTorch [58]. Namely, the pre-trained model from Graph R-CNN (see Section 3.1.4) [33, 34] was utilized for this purpose. Notwithstanding, these models were trained on the Visual Genome dataset. For this reason, a good performance was not expected from this baseline. Moreover, the quality of the frames from the Action Genome dataset is lower than the frames from Visual Genome dataset, so pre-trained models were expecting better input images as well. 3. After that, a label mapping was performed between the results from the VG dataset and the AG dataset as each dataset contains a specific set of class and relationship labels in their triplets. Even if the distance between each label embedding could have been used, a manual mapping was designed. It must be taken into account that, although both datasets describe scenarios, one may contain labels that are not present in the other one and the other way around. Figure 4.2 depict some of these mappings between both datasets. 4. Finally, a triplets filtering was performed. Since all triplets in the Action Genome dataset are composed of <person, relation,entity >, all predicted triplets that do not start with the person entity are removed from the final set of results. Furthermore, since only one person appears in each scene and they do not interact with themselves, if the entity was a person, it got removed too. As it was expected, due to the utilization of pre-trained models from another dataset, the manual labeling and the lack of use of context, the poor metrics that were obtained demonstrated the hard scenario that was being treated. 4.2.2 Visual Genome CNNBiLSTM Baseline The second baseline was developed over the predicted results from Section 4.2.1. The idea behind this baseline was a quick measure of the complexity of this problem by trying to improve the previous results by including a deep learning model on top. The deep learning model, as it is depicted in Figure 4.3, takes as input a matrix composed of predicted triplets features such as both person and entity bounding boxes (i.e., 8 features), their relationship (i.e., 1 feature) and the entity id (i.e., 1 feature). Since both relationships and entities are numbers from 1 to R(i.e., number of total relationships in Action Genome) or E(i.e., number of total entities in Action Genome), they get into an embedding layer so as to get a 8-dimensional vector from both of them in normalized values. After concatenating all 24 features, this matrix composed of Ntriplets proposals by 24 features is treated as an image and inserted into the Convolutional Network (see Section 2.2.2) that is shown in Figure 4.3. Study and Training of DL Models for SGG from non-static Environments 39 CHAPTER 4. METHODOLOGY 4.4. FASTER R-CNN TRAINING Figure 4.7: ResNet-34 architecture example [48] 46 Study and Training of DL Models for SGG from non-static Environments CHAPTER 4. METHODOLOGY 4.4. FASTER R-CNN TRAINING All the information regarding the evolution of the training process was stored in log files. The training and validation losses for both backbones are depicted in Figure 4.8 and Figure 4.9. (a) Training losses (b) Validation losses Figure 4.8: Faster-RCNN Training and validation losses of using ResNet-50 backbone (a) Training losses (b) Validation losses Figure 4.9: Faster-RCNN Training and validation losses using ResNet-101 backbone As it can be observed, the loss of a Faster-RCNN-based model (see Section 2.3) is divided into four different categories: •loss classifier: Loss provided by the classification of the different objects of the training frames. •loss box reg: Loss provided by the regression of the coordinates of the bounding boxes (i.e., location of the different objects in the training frames). Study and Training of DL Models for SGG from non-static Environments 47 CHAPTER 4. METHODOLOGY 4.5. RELDN TRAINING •loss objectness: Loss obtained by the score that measures if there is an object of interest in a region of the training frame. •loss rpn box reg: Loss obtained by the regression of the coordinates of the bounding boxes provided by the Region Proposal Network (RPN). 4.5 RelDN Training Once the object detector was trained, the relationship classifier training stage was started. In this case, and based on its performance [71] on the Open Images Visual Relationship Detection challenge [72] and the Visual Genome [46] dataset, the RelDN approach (see Section 3.1.6) was used. Usually, relationship classifiers work using the same convolutional layers that were used for the object detection. Nevertheless, a recent study [2] has demonstrated that the utilization of a separated backbone is able to perform better. These decoupled modules (see Figure 4.10) are designed by letting the relationship head, which classifies the relationship between two proposed objects, receive as input the features extracted by a separated relationship backbone. It must be explained that the initial weights from this new backbone are started by duplicating the object detector ones. Figure 4.10: Illustration of coupled and decoupled SGG model architecture [2] Moreover, due to the popularity of the Visual Genome dataset [46], most relationship classifiers are prepared to train having a single relationship constraint between objects. In other words, they make a single relationship assumption between the different objects in the image. However, the Action Genome dataset [1] proposes multiple relationships between a pair of objects. For that reason, our SGG implementation [2] was trained using the unconstrained mode, allowing more than one relationship between objects. Nevertheless, as it was explained in [47], 48 Study and Training of DL Models for SGG from non-static Environments CHAPTER 4. METHODOLOGY 4.5. RELDN TRAINING two relevant changes had to be applied on the code: 1. The relationship score activation was changed from softmax to sigmoid. This score is used in the likelihood ordering of the resulting triplets. 2. The loss function was changed from cross entropy to binary cross entropy. Therefore, in the training phase, more than one relationship can appear between a pair of objects. Furthermore, as it was explained in Section 4.3, a frequency prior tensor was obtained so as to provide semantic features from the training set empirical distribution of the different relationships based on the subject and object classes. In Section 5.3.1, a further study about the importance of this frequency prior will be detailed. As it was done with the object detector, the training hyperparameters were organized in a.yaml file to make the training process easier. Once the different triplets were preprocessed (see labels.tsv in Table 4.2), the Action Genome dataset trained the relationship classifier pipeline. It must be taken into account that, since two object detectors were trained (i.e.,ResNet- 50-FPN and ResNet-101-FPN), two relationship classifiers were trained too. In the ResNet-50 case, it was trained 2500 epochs using 4 images per batch without weight decay. On the other hand, in ResNet-101, due to hardware limitations, it was trained 5000 epochs using 2 images per batch with a weight decay of 10−4. In both cases, the initial learning rate was set to 10−3. Since the validation process consist in taking the validation loss together with a quantitative evaluation of the SGG performance, it was applied every time the model trained with 1000 images. The relationship header was modified in order to classify 27 relationship classes (i.e., background/no relationship + 26 classes). Only those objects with a confidence score greater than 0.05 from the object detector were taken into account in the relationship head. After obtaining the different relationship scores between the pairs of objects, the triplets were obtained following the Algorithm 1, which works as follows: •From the line 8 to 16, in the case of a ResNet-50 backbone, a special processing is performed. Taking into consideration that, at most, only one object of each class appears in each frame, those objects that are far away from the most reliable object given the object detector confidence values are substituted for this one. Consequently, if an object cfrom class xappears in a relation < c, b > and it is far away from the confidence value of the most reliable object aof the same class x, the relation is replaced to < a, b >. This distance is regulated using γ, which in our experiments was set to 0.2 due to empirical studies. Study and Training of DL Models for SGG from non-static Environments 49 CHAPTER 4. METHODOLOGY 4.5. RELDN TRAINING Algorithm 1: Generation of Predicted Triplets 1Input:Objectsclasses,Objectsscores,ObjectsboundingBox,Objectspairs, RelationshipsscoresP erP air,γ 2Output:Triplets,T ripletsboundingBoxes,Tripletsscore 3 4Triplets ←List 5TripletsboundingBoxes ←List 6Tripletsscores ←List 7 8if ResNet50backbone then 9for Class in Objectsclasses do 10 MaxScoreclass ←max(ObjectsscoresClass ) 11 MaxScoreclassindex ←argmax(ObjectsscoresClass ) 12 for Pair in Objectspairs do 13 if Pair0score < MaxScoreClassP air0−γthen 14 Pair0←MaxScoreClassP air0Index 15 if Pair1score < MaxScoreClassP air1−γthen 16 Pair1←MaxScoreClassP air1Index 17 18 for Pair in Objectspairs do 19 for Relationshipscore in RelationshipsscoresP air do 20 Triplets ←T riplets ∪< Pair0, Relationshipscore, Pair1> 21 T ripletsboundingBoxes ←T ripletsboundingBoxes ∪{P air0boundingBox , P air1boundingBox } 22 T ripletsscores ←T ripletsscores ∪{P air0score ∗P air1score ∗Relationshipscore} 23 24 for Triplet in T riplets do 25 if Triplet0= 1 or T riplet2= 1 then 26 /* Subject must be a person, Object cannot be a person. */ 27 Remove Triplet from T riplets. Remove it from T ripletsboundingBoxes and Tripletsscores too 28 29 for Tripletconcat in < T riplets, T ripletsboundingBoxes >do 30 Remove repeated T ripletconcat 31 32 Sortedinds ←argsort(Tripletsscores) 33 Sort Triplets,T ripletsboundingBoxes and T ripletsscores according to Sortedinds 34 35 Return Triplets,T ripletsboundingBoxes,Tripletsscores 50 Study and Training of DL Models for SGG from non-static Environments CHAPTER 4. METHODOLOGY 4.5. RELDN TRAINING •From the line 18 to 22, the different triplets are generated using the input data together with the bounding boxes location and the assigned score. As it is usually done in SGG (see Chapter 3), each score is computed as: Score =Scoresubject ∗Scorerelationship ∗Scoreobject (4.2) Where Scoresubject and Scoreobject is computed by Faster-RCNN and Scorerelationship is calculated by the RelDN. •From the line 24 to 30, some filtering steps are performed. Those triplets whose subject is not a person or the object is a person are removed (since the Action Genome dataset [1] is composed of non-directed HOR). Moreover, repeated triplets (i.e., triplets composed of the same objects, relationships and bounding boxes) are eliminated too. •Finally, in lines 32 and 33, the predicted triplets are sorted using the scores. As it was done in Section 4.4, all the information related to the evolution of the training process was stored in log files too. The training and validation losses for both backbones are depicted in Figure 4.11 and Figure 4.12. It must be taken into account that the models have been previously trained in Section 4.4, so the initial epoch is not 0. (a) Training losses (b) Validation losses Figure 4.11: RelDN Training and validation losses using ResNet-50 backbone As it was explained in Equation 3.5, the Graphical Contrastive Losses (see Section 3.1.6) are composed of four losses. In this case, the L2is not used since a single frame from Action Genome does not contain two objects from the same class. Therefore: •L0is the loss pred classifier. •L1is the loss contrastive sbj and loss contrastive obj. •L3is the loss p contrastive sbj and loss p contrastive obj. Study and Training of DL Models for SGG from non-static Environments 51 CHAPTER 4. METHODOLOGY 4.6. HORL (a) Training losses (b) Validation losses Figure 4.12: RelDN Training and validation losses using ResNet-101 backbone 4.6 HORL Once the different triplets were obtained and processed, the Human-Object Relationship LSTM (i.e., HORL) system was used. The main objective of this approach is the utilization of context (i.e., features from other frames from the same video) in order to improve the generated triplets for each frame. This context-aware model is based on the use of an attentional (see Section 2.2.4) Bidirectional LSTM (see Section 2.2.3). The general overview of the model is depicted in Figure 4.13. The idea is the obtainment of a suitable and concentrated representation of the different relationships in a frame from the dataset. Taking into consideration that, in this case, the subject is always the person and that there may be more than one relationship between this person and a specific object, the following matrix is designed for each frame of the video (hereinafter, timestamp):     A1,1A1,2... A1,M A2,1A2,2... A2,M ... ... ... ... AN,1AN,2... AN,M     (4.3) Where Ai,j measures the score (i.e., confidence) of the triplet that contains the relation jbetween the person and the object i. Therefore, the content of the predicted triplets is summarized into this matrix. If there is a lack of triplets regarding a specific element of the matrix, it is set to zero. In the case that multiple triplets refer to the same relation jand object i, the highest value is stored in the matrix. Since Confidencetriplets ∈[0,1], all values from the matrix are already standardized and ready to be inserted into the attentional many-to-many BiLSTM. 52 Study and Training of DL Models for SGG from non-static Environments CHAPTER 4. METHODOLOGY 4.6. HORL Figure 4.13: HORL Overview Regarding the attentional BiLSTM, it is based on an implementation [73] of a research [74] that made use of it in order to perform a Natural Language Processing (NLP) task: relation classification. An overview of this type of LSTM can be observed in Figure 4.14. Figure 4.14: Attention-based BiLSTM Overview [74] However, this attention-based BiLSTM was modified in order to work with the timestamp matrices. First, the embedding layer was removed since the input matrix is not composed of wordlike representations, being unrelated to NLP. As it was mentioned before, the input matrix is composed of triplet scores, which is already a good representation for LSTM-based models. Secondly, the output layer was substituted for our own HORL model, which takes as input the hidden states from the different frames matrices of the video. Lastly, the number of LSTM layers was set to 1 and dropout [75] was not applied. The attentional LSTM computes features that are processed in order to obtain a refined matrix of the same input dimensions in each timestamp (i.e., frame) of the video. The Study and Training of DL Models for SGG from non-static Environments 53 CHAPTER 4. METHODOLOGY 4.6. HORL two specific HORL architectures that were developed along this thesis are explained in Section 4.6.1 and Section 4.6.2. As for the training, the binary cross entropy loss was used by flattening both the refined matrix and the ground truth matrix. The ground truth matrices were built by assigning ones to those values of the matrices that represented the correct relation jwith the object i. Adam [62] was utilized as the optimizer, and the number of epochs and batch size were set to 5 and 1 respectively. In each epoch, the whole set of training videos were inserted in the training stage (i.e., the batches per epoch were set to the total of available training videos). A set of training videos was used for validation purposes. Regarding the bounding boxes data (i.e., location of the different person and objects of the images), they were inserted into a Bounding Boxes dictionary together with their confidence values provided by Faster R-CNN. This dictionary was used in order to assign a location to the different objects taking part in the refined triplets after the computation of the HORL system. Once the refined matrix was obtained for each timestamp, it was added to the original Horlinput and processed in order to obtain the enhanced triplets. The general inference algorithm can be seen in Algorithm 2 and the obtainment of triplets from the refined timestamp can be analyzed in Algorithm 3. Algorithm 2: HORL Inference 1Input:Triplets,T ripletsboundingBoxes,Objectsscores,Tripletsscores 2Output:Triplets,T ripletsboundingBoxes 3 4Horl ←Load HORL model 5BoundingBoxesdict ←Build aforementioned dictionary of Bounding Boxes using Triplets,T ripletsboundingBoxes,Objectsscores 6Horlinput ←Build Timestamp input matrices from T ripletsscores (see Equation 4.3) 7Horloutput ←Prediction using Horl +Horlinput 8Triplets,T ripletsboundingBoxes ←Apply Algorithm 3 using Horloutput and BoundingBoxesdict 9 10 Return Triplets,T ripletsboundingBoxes Namely, Algorithm 3 describes how the bounding boxes dictionary is used in order to get the possible locations for each obtained triplet from the HORL output matrix. Both person and object location are searched in this bounding box dictionary. First, it gets their location in the current frame. If Faster-RCNN was not able to find these elements in the current frame, the algorithm tries to find them in the closest k-frames. In our experiments, as the person location is mandatory to produce triplets, kpwas set to 30 frames. On the other hand, kowas set to 1, meaning that the algorithm only searches for the bounding boxes, at most, in the closest forward and backward frame with respect to the current one. 54 Study and Training of DL Models for SGG from non-static Environments CHAPTER 4. METHODOLOGY 4.6. HORL Algorithm 3: Obtainment of refined triplets from HORL output 1Input:Horloutput,BoundingBoxesdict,kp,ko 2Output:Triplets,T ripletsboundingBoxes 3 4Triplets ←List 5TripletsboundingBoxes ←List 6Tripletsscores ←List 7 8if person in current frame from BoundingBoxesdict then 9Subbb ←BoundingBoxesdict[personF romCurrentF rame] 10 else 11 Subbb ←Find the Bounding Boxes of the person from the temporally kp-closest frames in BoundingBoxesdict 12 13 for objidx in Horloutput do 14 if object in current frame from BoundingBoxesdict then 15 Objbb ←BoundingBoxesdict[objectF romCurrentF rame] 16 else 17 Objbb ←Find the Bounding Boxes of the object from the temporally ko-closest frames in BoundingBoxesdict 18 19 for relidx in Horloutput do 20 for PossibleSubjectboundingBox in Subbb do 21 for PossibleObjectboundingBox in Objbb do 22 Triplets ←T riplets ∪<1, Relidx, Objidx > 23 TripletsboundingBoxes ←T ripletsboundingBoxes ∪ {PossibleSubjectboundingBox, PossibleObjectboundingBox} 24 Tripletsscores ←Tripletsscores ∪Apply Equation 4.4 25 26 Sortedinds ←argsort(Tripletsscores) 27 Sort Triplets and T ripletsboundingBoxes according to Sortedinds 28 29 Return Triplets,T ripletsboundingBoxes Study and Training of DL Models for SGG from non-static Environments 55 CHAPTER 5. EXPERIMENTS AND RESULTS 5.1. SGG DATASETS •1.7M visual question answers. •3.8M object instances. •2.8M attributes. •2.3M relationships. Figure 5.1: Overview of Visual Genome [46] 5.1.2 Action Genome The Action Genome dataset [1] proposes a set of videos in order to design systems oriented to the generation of scene graphs from non-static environments. The underlying idea behind the creation of this dataset is the evidence from Cognitive Science and Neuroscience that people actively encode activities into consistent hierarchical part structures (see Figure 1.1). 62 Study and Training of DL Models for SGG from non-static Environments CHAPTER 5. EXPERIMENTS AND RESULTS 5.1. SGG DATASETS In Action Genome, the videos are decomposed into actions which can be interpreted as spatio-temporal scene graphs. This dataset captures changes between active objects and their pairwise relationships while an action occurs. Some of its characteristics are: •It is built upon the Charades dataset [80]. •It contains 10000 videos, getting 234000 frames, where roughly 25% are devoted to the test set. •It provides 400000 objects of 36 classes (including the person class) annotated with their bounding boxes. Moreover, it contains the annotation of 1.7M visual relationships with their 26 action categories (including the other relationship class). This information is depicted in Figure 5.2, where it can be analyzed how some object and relationship classes do not have much representation, causing a slight imbalance in the dataset that may affect the frequency prior, among others (see Section 5.3.1). It must be taken into account that most recent models have resorted to end-to-end predictions that produce a single label for a long sequence of frames from a video and do not explicitly decompose its events into a series of interactions between active objects, which may provide a more descriptive explanation of the scene. Nevertheless, scene graphs, as a comprehensive structural abstraction of images, has not yet been much studied in any large-scale video database as a potential representation for action recognition. Therefore, Action Genome is the first large-scale dataset to jointly boost research in scene graphs and action understanding. For that reason, it can be used in order to train explainable models. Figure 5.2: Distribution of object and relationship occurrences in Action Genome [1] Finally, Figure 5.3 show some examples of frames from Action Genome. On the other hand, Figure 5.4 demonstrate that this dataset may contain bad quality frames, specially due to their video origin, since when the different frames composing the video are split, some interlaced and blurred frames appear. Therefore, the developed models must deal with this complexity, causing a more difficult SGG. Study and Training of DL Models for SGG from non-static Environments 63 CHAPTER 5. EXPERIMENTS AND RESULTS 5.1. SGG DATASETS Figure 5.3: Action Genome examples [1, 80] Figure 5.4: Action Genome bad quality examples [1, 80] 64 Study and Training of DL Models for SGG from non-static Environments CHAPTER 5. EXPERIMENTS AND RESULTS 5.2. SGG EVALUATION METRICS 5.2 SGG Evaluation Metrics In order to understand the quantitative analysis, a brief overview about the most common evaluation metrics in SGG will be described. Three types of predictions are usually performed and evaluated1: •Predicate Recognition /Predicate Classification (PredCls): Using the ground truth bounding boxes including their object categories, predict the relationships between them. Therefore, it measures the system performance on the classification of the predicates alone. •Phrase Recognition (PhrCls) / Scene Graph Classification (SgCls): Using the ground truth bounding boxes without their object categories, predict both the object categories and the relationships between them. Consequently, it measures the system performance on the classification of the predicates and objects. •Scene Graph Generation (SGGen) / Scene Graph Detection (SgDet): It does not use ground truth information, so all the prediction is based on the system without any external support. Therefore, it detects the locations and classes of objects and their pair-wise relationships. Another one (i.e.,Scene Graph Generation+ or SGGen+) can be observed in [33]. However, since it is not used in other proposals, it will not be considered for this thesis and its quantitative analysis. Once the predictions are performed, the recall is computed (see Equation 5.1). recall =|groundTruthtriplets ∩predictedtriplets| |groundTruthtriplets|(5.1) Recall mainly penalizes the false negatives (i.e., triplets in the ground truth that are not found in the predicted ones). It must be explained that a top kof predicted triplets is given in order to avoid inserting every possible combination in the result. Usually, k={20,50,100}, meaning that only the most reliable 20, 50 or 100 predicted triplets are inserted in Equation 5.1. As it was explained in Algorithm 1, the triplets are generated together with the bounding boxes of the objects that take part in them. An IoU of 0.5 is performed in order to measure that the system is identifying correctly the position of the triplets objects in SGGen. Both triplets and coordinates must match the ground truth simultaneously in order to classify them as correct. In SGG from non-static environments, two types of recall averages are performed: one that computes the mean recall using all the frames together and another one that computes the recall using the frames from each video independently and then computes a final mean over the resulting averages. 1Depending on the paper, the used terminology is different. Study and Training of DL Models for SGG from non-static Environments 65 CHAPTER 5. EXPERIMENTS AND RESULTS 5.3. QUANTITATIVE ANALYSIS 5.3 Quantitative Analysis First, the set of quantitative metrics will be described. The system performance will be evaluated using the three aforementioned evaluation metrics (i.e.,PredCls,SGCls and SGDet) together with the Average Precision (AP) at 0.5 IoU with COCO metrics [69]. In Table 5.1, a comparison of our approaches with the other proposed models that have been trained in Action Genome is shown. In this case, the different recalls are averaged by frame results. It must be explained that both HORT [47] (1) and STTran [52] (2) are highlighted in red since (1) does not provide code nor checkpoints, so their results are not reproducible and (2) modifies the ground truth bounding boxes by keeping only those with short edges larger than 16 pixels. However, they are included in the tables although a fair comparison cannot be carried out. On the other hand, Neural Motifs [30], Graph-RCNN [33] and RelDN [43] are static approaches. Therefore, they are context-agnostic, so they do not take into account features from other frames when the SGG is performed. Object Detector PredCls SGCls SGDet Method Backbone AP50 @20 @50 @100 @20 @50 @100 @20 @50 @100 Neural Motifs (3.1.3) [50, 30] ResNet-101 - 87.95 93.02 - 45.10 48.87 - 34.41 44.34 - G-RCNN (3.1.4) [50, 33] ResNet-101 - 88.73 93.73 - 45.57 49.75 - 34.28 44.47 - RelDN (3.1.6) [50, 43] ResNet-101 - 90.89 96.09 - 46.47 50.31 - 34.92 45.27 - TRACE (3.2.2) [50] ResNet-101 - 91.60 96.35 - 46.66 50.46 - 35.09 45.34 - ResNet-50 22.7 85.56 90.03 90.28 42.10 44.33 44.47 31.00 42.31 48.35 Our RelDN (4.5) [43] ResNet-101 27.5 85.59 90.11 90.26 45.96 48.92 49.05 35.19 47.83 55.59 ResNet-50 22.7 86.28 91.50 91.81 43.73 49.26 50.11 30.79 42.35 49.52 HORL-CNN (4.6.1) ResNet-101 27.5 86.21 91.55 91.79 47.38 53.68 54.86 35.02 47.58 56.18 ResNet-50 22.7 84.95 91.47 91.78 43.09 49.39 50.15 31.25 42.64 49.83 HORL-Linear (4.6.2) ResNet-101 27.5 86.17 91.59 91.80 47.48 54.13 54.93 35.45 47.80 56.64 HORT (3.2.1) [47] ResNet-101 20.7 71.67 76.16 - 47.68 62.56 - 37.19 47.76 - STTran (3.2.3) [52] ResNet-101 24.6 94.20 99.10 - 63.70 66.40 - 36.20 48.80 - Table 5.1: Quantitative comparison of our approaches with the other proposed models in Action Genome [1]. The recall is averaged by frame. As for the object detector, our Faster-RCNN based on a ResNet-101 backbone is able to outperform the approaches that included their AP results by almost a 3%. Consequently, our training process and its elements, which are described in Section 4.4, are able to create a Faster-RCNN model that is able to learn the underlying patterns from Action Genome in order to detect and classify the different active objects that are present in each frame. The pre-training on other datasets or the weight decay that is applied in the epochs support a better learning from our model which, even without discarding those objects with a short edge smaller than a specific length, is able to recognize the different relevant objects from the image. This enhancement creates an object detector that is able to disentangle the possible problems that may emerge with complex frames (see Figure 5.4). Therefore, 66 Study and Training of DL Models for SGG from non-static Environments CHAPTER 5. EXPERIMENTS AND RESULTS 5.3. QUANTITATIVE ANALYSIS the Hypothesis 2 (i.e.,The object detector plays a very important role in the SGG computation) gets verified. On the other hand, the number of layers from the ResNet is crucial, as ResNet-101 is able to outperform the ResNet-50 by a large margin. For that reason, ResNet-101 is usually utilized as backbone in SGG tasks. The importance of a good object detector is mostly depicted in SGCls and SGDet, where the bounding boxes and/or the classes must be predicted. Regarding the PredCls metrics, where only the relationship between a pair of ground truth objects is predicted, the TRACE approach [50] is able to outperform our approaches by a 5%. However, our results are still quite competitive, since all of them are able to outperform the proposal [47] from the original authors of Action Genome by roughly a 15%. Therefore, both our RelDN and the HORL system share appropriate features between the frames of the videos in order to enhance the scene graph relationships. This behavior will be further visualized in Section 5.4. Furthermore, we outperform all the SGCls metrics, meaning that the object detector is able to provide higher quality object classes. Once the object classes and their confidences are obtained, our RelDN and its HORL on top produce more accurate relationships between the objects. Namely, the HORL-Linear based on the ResNet-101 is able to outperform the aforementioned TRACE [50] by almost a 1% in SGCls@20 and by nearly a 4% in SGCls@50. Other elements like the decoupled SGG model (see Figure 4.10) support the improvement of our system with respect to the other proposals. No direct comparison is made with STTran [52] since their modification of the ground truth objects increase deceivingly their SGCls and SGDet. Finally, SGDet metrics measure the quality of the whole model, since it must predict the relationships, the object classes and their location. In this case, we outperform again all the different recalls. This confirm our Hypothesis 1 (i.e.,The context provided by the frames of a video can support an enhancement in the SGG of a specific frame from the same video). In the case of the non-static approach TRACE [50], we are able to increase their results in a 0.5% in case of SGDet@20 and in more than a 2% in the case of SGDet@50. Therefore, the complete model (i.e., Faster-RCNN, RelDN and HORL on top) is able to generate the most accurate scene graphs in the State-of-the-Art. Once again, even though both HORT [47] and STTran [52] cannot be used in a direct comparison, we are able to provide competitive results, being able to even outperform the SGDet@50 from HORT [47]. In general, our HORL based on a fully connected approach (see Section 4.6.2) is able to outperform our HORL based on a CNN one (see Section 4.6.1). Consequently, the relationship matrix is represented better as a regression problem rather than as a segmentation one. Moreover, as it has been explained, both proposals outperform most of the existing fair approaches (even the non-static ones), which support the Hypothesis 4 (i.e.,The relationships from the frames must be thoughtfully represented in order to transfer this context correctly between frames) Lastly, Table 5.2 describes the same comparison averaged by video results. In other words, Study and Training of DL Models for SGG from non-static Environments 67 CHAPTER 5. EXPERIMENTS AND RESULTS 5.3. QUANTITATIVE ANALYSIS the different recalls by frame are averaged by video and then, the final average is computed over all the resulting means. Consequently, in these metrics, all videos contain the same weight in the results, even if they contain more frames than others. The results demonstrate similar hypotheses to the ones that were explained before. Moreover, in this case, the robustness of our HORL system is able to improve its results by reducing some margins where it was outperformed (e.g., in PredCls) or by increasing those where it was already outperforming (e.g.,SGCls or SGDet). For example, our SGDet@50 is 5% higher than the SGDet@50 from TRACE [50], when it was roughly 2% in the case of Table 5.1. Object Detector PredCls SGCls SGDet Method Backbone AP50 @20 @50 @100 @20 @50 @100 @20 @50 @100 Neural Motifs (3.1.3) [50, 30] ResNet-101 - 86.01 88.59 - 44.47 46.39 - 32.50 41.11 - G-RCNN (3.1.4) [50, 33] ResNet-101 - 86.28 88.93 - 45.11 47.22 - 32.60 41.29 - RelDN (3.1.6) [50, 43] ResNet-101 - 88.77 91.43 - 45.87 47.78 - 33.18 42.10 - TRACE (3.2.2) [50] ResNet-101 - 89.31 91.72 - 46.03 47.92 - 33.38 42.18 - ResNet-50 22.7 85.45 89.28 89.48 41.55 43.42 43.54 30.46 41.32 47.11 Our RelDN (4.5) [43] ResNet-101 27.5 85.48 89.35 89.46 45.79 48.33 48.43 34.58 46.72 54.25 ResNet-50 22.7 86.29 90.82 91.07 43.36 48.32 49.04 30.26 41.55 48.43 HORL-CNN (4.6.1) ResNet-101 27.5 86.26 90.86 91.07 47.41 53.19 54.22 34.64 46.87 55.27 ResNet-50 22.7 85.28 90.81 91.05 42.86 48.44 49.07 30.95 41.93 48.73 HORL-Linear (4.6.2) ResNet-101 27.5 86.27 90.89 91.06 47.56 53.61 54.28 35.05 47.13 55.72 HORT (3.2.1) [47] ResNet-101 20.7 72.39 76.66 - 47.11 61.61 - 36.51 46.67 - STTran (3.2.3) [52] ResNet-101 24.6 - - - - - - - - - Table 5.2: Quantitative comparison of our approaches with the other proposed models in Action Genome [1]. The recall is averaged by video. Then, the average is computed over all the resulting means. 5.3.1 Ablation Study The objective of the conducted ablation study was the measurement of the importance of the frequency prior matrix that was explained in Section 4.3. These semantic features based on the empirical distribution of the different relationships given the subject and object classes from the training set result in a bias factor when two objects are identified as a potential pair. For example, given the objects person and glass, a potential relationship might be drinking from since multiple instances from the training set contain this relationship between these objects. As it can be observed from Table 5.3, if this frequency prior matrix is removed, the recall is enormously reduced, as [2] stated. Therefore, it can be concluded that the frequency prior matrix is as important as the RelDN in order to compute the potential relationships between a pair of objects. However, the frequency prior matrix bias may not be advantageous always. For instance, if the pair of objects are person and television, and the relationship is not looking at, the frequency prior matrix may state that is actually looking at due to the training set 68 Study and Training of DL Models for SGG from non-static Environments CHAPTER 5. EXPERIMENTS AND RESULTS 5.4. QUALITATIVE ANALYSIS empirical distribution. Nevertheless, the objective of the RelDN and the HORL systems is the enhancement of the results and the avoidance of this type of wrong bias. Method Object Detector SGDet Backbone AP50 @20 @50 @100 Our RelDN without Frequency Prior (4.5) [43] ResNet-101 27.5 11.44 23.02 35.69 Our RelDN (4.5) [43] ResNet-101 27.5 35.19 47.83 55.59 HORL-CNN (4.6.1) ResNet-101 27.5 35.02 47.58 56.18 HORL-Linear (4.6.2) ResNet-101 27.5 35.45 47.80 56.64 Table 5.3: Frequency Prior ablation study. The recall is averaged by frame. All this ablation study demonstrate the Hypothesis 3 (i.e.,The importance of a precise relationship classification in non-static environments is provided by the relationship head of the deep learning model, the semantic features and the context of the frame). 5.4 Qualitative Analysis This last evaluation section presents a visual comparison between the ground truth scene graph, and the graphs generated by our RelDN (i.e., without HORL) and the ones generated by RelDN with the linear HORL system on top. All these results use ResNet-101 as the backbone. The results using ResNet-50 are provided in Appendix A. Namely, three videos are utilized for this qualitative analysis where three frames have been selected for visualization purposes. The temporal distance between the frames is approximately equal, giving the opportunity to HORL to enhance the scene graphs by using features provided by close frames. The individual recall for each predicted graph has been computed too in order to measure quantitatively the enhancement between the two included approaches. Both systems predicted the 20 most feasible triplets for each frame. Some of the wrong ones were removed in order to improve the visualization of the different graphs. The correct and incorrect triplets are highlighted in green and red respectively. Orange triplets are relationships that, even if they are not included in the ground truth, they could be perfectly accepted as well. In the first case, in Figure 5.5, we can see a person cleaning the floor around a box in the first frame. Then, this person starts making some interaction with the box. In the first frame, since the interaction with the box has not started yet, the HORL system is able to avoid the detection of this object as it is not an active one. In other words, in order to select the box as a possible object, first, it must have been selected somewhere with a relationship like not contacting or touching. Study and Training of DL Models for SGG from non-static Environments 69 CHAPTER 5. EXPERIMENTS AND RESULTS 5.4. QUALITATIVE ANALYSIS Moreover, since it is able to distinguish in a better way the active objects from the inactive ones, the quantity of false positive objects is reduced, resulting in more possible relationships between correct ones. The second frame demonstrates this hypothesis by solving the false negatives from the system without HORL in the HORL one. Furthermore, since the next frames contain relationships like <person, touching, box>as it is depicted in the third frame, our HORL system is able to enhance the proposed triplets by adding the not contacting missing relationship. Finally, in the third frame, the box is mostly occluded by the person hands, causing a poor quantity of triplets devoted to the ground truth relationship between the person and the box. However, the triplets from HORL are once again enhanced by including the relationship touching as more feasible than holding by taking into account the entire video context. 70 Study and Training of DL Models for SGG from non-static Environments CHAPTER 5. EXPERIMENTS AND RESULTS 5.4. QUALITATIVE ANALYSIS Figure 5.5: Comparison of SGG in video 06LBQ Study and Training of DL Models for SGG from non-static Environments 71 CHAPTER 6. CONCLUSIONS AND FUTURE WORK 6.1. MAIN HIGHLIGHTS margins, contributing to the addition of a new State-of-the-Art in Scene Graph Generation from videos (i.e., non-static environments). These metrics are supported by both quantitative and qualitative analysis, where several models with their architectures and backbones are compared. Regarding the initial hypothesis that were explained in Chapter 1: •Hypothesis 1 (i.e.,The context provided by the frames of a video can support an enhancement in the SGG of a specific frame from the same video) was demonstrated in Section 5.3, where our HORL system outperformed the current State-of-the-Art using features provided by surrounding frames. •Hypothesis 2 (i.e.,The object detector plays a very important role in the SGG computation) was verified in Table 5.1 and Table 5.2, where the enhancement of the object detector quality was crucial in order to get the aforementioned outperforming results. •Hypothesis 3 (i.e. The importance of a precise relationship classification in nonstatic environments is provided by the relationship head of the deep learning model, the semantic features and the context of the frame) was confirmed in our Table 5.3, where it was demonstrated that the three components play an important role in the obtainment of a high quality relationship classification. •Hypothesis 4 (i.e.,The relationships from the frames must be thoughtfully represented in order to transfer this context correctly between frames.) was demonstrated in Section 5.3 as well, where our representation of the non-static environment provided better results than the State-of-the-Art. Moreover, two types of data transformations were discussed: a segmentation approach (i.e., CNN-based HORL) and a regression approach (i.e., Linear-based HORL). This concentrated representation of images will support other important tasks such as Image Retrieval, Image Captioning, Visual Question Answering or Deep Reinforcement Learning explainability. Moreover, the advancements in both Scene Graph Generation and Image Generation from Graphs will provide more research on graphs as an appropriate way to represent data in a world where its quantity does not stop growing more and more. Therefore, our HORL is able to provide a better representation of situations in videos, meaning that this system will be able to support tasks like the previously mentioned ones in complex scenarios. For instance, and taking into consideration the motivation explained in Section 1.1, the reasoning of an agent in its environment is a sequential set of choices that can be interpreted using our proposal, which takes into account its whole context. On the other hand, it has been demonstrated that an interaction is usually composed of a sequence of actions, resulting in the necessity of this non-static environment model, which can give insight into tasks that are usually restricted exclusively to image-level such as Visual Question Answering. 78 Study and Training of DL Models for SGG from non-static Environments CHAPTER 6. CONCLUSIONS AND FUTURE WORK 6.2. FUTURE WORK Additionally, since Action Genome studies human-object interactions, this research support a further study on the set of actions that people have with their environment, being helpful in contexts like human-robot interaction, senior care or health care. For example, human-robot interaction must be understood as the sequence of actions that people carry out with these robots. For that reason, the utilization of a context-aware model becomes indispensable to support these sequential tasks, whose results can be clearly enhanced by taking into account the temporal aspect of the interaction. 6.2 Future Work However, even if our results outperform the State-of-the-Art approaches, some future lines of research are proposed in order to study whether they can contribute to a more accurate SGG. For example, our approach trained all its elements separately (i.e., the Faster- RCNN, the RelDN and the HORL). Notwithstanding, end-to-end approaches are becoming more popular due to their capability of making accurate representations between its different modules. For that reason, and with the adequate hardware, it is believed that an end-to-end approach may provide better results. On the other hand, the frames were not preprocessed in order to improve their quality so as to provide the same baseline for the compared approaches. Nevertheless, deblurring or deinterlacing techniques could be applied as well in order to facilitate the object detector task. Regarding the relation detector (i.e., the RelDN), it is demonstrated that its involvement in the results is as important as the frequency prior matrix. Further research on the relation detector should be carried out in order to improve its performance, specially in problems where the number of relationships between two objects may be more than one. As for the person features, in our approach a simple bounding box is extracted from it. Nevertheless, much more features could be extracted like the position of the arms, legs or head. Assuming that the actions performed by a person depend on the position of the limbs, these features may support the relationship head in a better manner. Finally, SGG is a task that is still far from being considered as solved. This task is even harder in the context of non-static environments, where Action Genome has been the first large-scale video database in this field. More datasets would support both the robustness of the different approaches that are being designed to tackle this problem and a higher quality training stage. It is expected that, due to the importance that graphs are acquiring, more resources appear in this artificial intelligence area. Study and Training of DL Models for SGG from non-static Environments 79 CHAPTER 6. CONCLUSIONS AND FUTURE WORK 6.3. FINAL WORDS 6.3 Final Words Personally, I am very pleased with the development of the work carried out in this master thesis. Since I had never made any research on SGG, the discovery of its related work has been absolutely an unforgettable experience. I consider that my knowledge has been expanded in several fields like relationship classification between two objects, treatment of models’ backbones, training techniques for deep learning models, utilization of context in temporal problems, representation of complex data or design and development of complex neural network architectures. The Master in Artificial Intelligence provides the required keys to solve many ongoing problems by having a wide overview of the different techniques that one can apply in order to enhance the quality of people lives. In this case, several courses about Deep Learning, Computer Vision, Supervised Learning, Complex Networks or Machine Learning, among others, have provided me the required insight to tackle this problem. Moreover, the emphasis of the master in topics like data preprocessing or representation has supported an ordered development of this thesis, which has been also encouraged by the extensive research that has been carried out along the entire master in different artificial intelligence areas. Lastly, I cannot be more grateful for obtaining the icing on the cake in this project by creating a new State-of-the-Art in video scene graph generation. Thanks to everyone who made it possible. 80 Study and Training of DL Models for SGG from non-static Environments Bibliography [1] Jingwei Ji et al. “Action genome: Actions as compositions of spatio-temporal scene graphs”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020, pp. 10236–10247. [2] Xiaotian Han et al. “Image Scene Graph Generation (SGG) Benchmark”. In: arXiv preprint arXiv:2107.12604 (2021). [3] Microsoft Corporation. Microsoft Official Website.url:https://www.microsoft. com/en-us. (accessed: 20.12.2021). [4] Google LLC. Google Images.url:https://www.google.com/imghp?hl=en. (accessed: 20.12.2021). [5] Yikang Li et al. “Scene graph generation from objects, phrases and region captions”. In: Proceedings of the IEEE international conference on computer vision. 2017, pp. 1261– 1270. [6] Jordi Torres. A gentle introduction to Deep Reinforcement Learning.url:https:// towardsdatascience.com/drl-01-a-gentle-introduction-to-deep-reinforcement- learning-405b79866bf4. (accessed: 20.12.2021). [7] Justin Johnson, Agrim Gupta, and Li Fei-Fei. “Image generation from scene graphs”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2018, pp. 1219–1228. [8] Joseph Redmon et al. “You only look once: Unified, real-time object detection”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2016, pp. 779–788. [9] Shaoqing Ren et al. “Faster r-cnn: Towards real-time object detection with region proposal networks”. In: Advances in neural information processing systems 28 (2015), pp. 91–99. [10] Wei Liu et al. “Ssd: Single shot multibox detector”. In: European conference on computer vision. Springer. 2016, pp. 21–37. [11] Srivignesh Rajan. An Introduction to Artificial Neural Networks.url:https : / / towardsdatascience.com/an-introduction-to-artificial-neural-networks- 5d2e108ff2c3. (accessed: 22.12.2021). [12] Tony Yiu. Understanding Neural Networks.url:https://towardsdatascience.com/ understanding-neural-networks-19020b758230. (accessed: 22.12.2021). 81 BIBLIOGRAPHY BIBLIOGRAPHY [13] Frank Rosenblatt. “The perceptron: a probabilistic model for information storage and organization in the brain.” In: Psychological review 65.6 (1958), p. 386. [14] David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. “Learning representations by back-propagating errors”. In: nature 323.6088 (1986), pp. 533–536. [15] Maurice Peemen NVIDIA Corporation. Convolutional Neural Network (CNN).url: https : / / developer . nvidia . com / discover / convolutional - neural - network. (accessed: 22.12.2021). [16] Sumit Saha. A Comprehensive Guide to Convolutional Neural Networks — the ELI5 way.url:https : / / towardsdatascience . com / a - comprehensive - guide - to - convolutional-neural-networks-the-eli5-way-3bd2b1164a53. (accessed: 22.12.2021). [17] Sepp Hochreiter and J¨urgen Schmidhuber. “Long short-term memory”. In: Neural computation 9.8 (1997), pp. 1735–1780. [18] Behrooz Mamandipoor et al. “Monitoring and detecting faults in wastewater treatment plants using deep learning”. In: Environmental Monitoring and Assessment 192 (Feb. 2020). doi:10.1007/s10661-020-8064-1. [19] Mahendran Venkatachalam. An introduction to Attention.url:https://towardsdatascience. com/an-introduction-to-attention-transformers-and-bert-part-1-da0e838c7cda. (accessed: 22.12.2021). [20] Ashish Vaswani et al. “Attention is all you need”. In: Advances in neural information processing systems. 2017, pp. 5998–6008. [21] Akash Agnihotri. Attending to Attention.url:https://towardsdatascience.com/ attending-to-attention-eba798f0e940. (accessed: 22.12.2021). [22] Ross Girshick et al. “Rich feature hierarchies for accurate object detection and semantic segmentation”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2014, pp. 580–587. [23] Ross Girshick. “Fast r-cnn”. In: Proceedings of the IEEE international conference on computer vision. 2015, pp. 1440–1448. [24] Yikang Li et al. Scene Graph Generation from Objects, Phrases and Region Captions. 2017. arXiv: 1707.09700 [cs.CV]. [25] Yikang Li. Multi-level Scene Description Network Official Implementation.url:https: //github.com/yikang-li/MSDN. (accessed: 22.12.2021). [26] Karen Simonyan and Andrew Zisserman. “Very deep convolutional networks for largescale image recognition”. In: arXiv preprint arXiv:1409.1556 (2014). [27] Danfei Xu et al. Scene Graph Generation by Iterative Message Passing. 2017. arXiv: 1701.02426 [cs.CV]. [28] Kyunghyun Cho et al. “Learning phrase representations using RNN encoder-decoder for statistical machine translation”. In: arXiv preprint arXiv:1406.1078 (2014). [29] Marylou Gabrie. “Towards an understanding of neural networks : mean-field incursions”. PhD thesis. Sept. 2019. 82 Study and Training of DL Models for SGG from non-static Environments BIBLIOGRAPHY BIBLIOGRAPHY [30] Rowan Zellers et al. Neural Motifs: Scene Graph Parsing with Global Context. 2018. arXiv: 1711.06640 [cs.CV]. [31] Rowan Zellers. Neural Motifs SGG Official Implementation.url:https://github. com/rowanz/neural-motifs. (accessed: 23.12.2021). [32] Rowan Zellers. Neural Motifs SGG Demo.url:https://rowanzellers.com/scenegraph2/. (accessed: 23.12.2021). [33] Jianwei Yang et al. Graph R-CNN for Scene Graph Generation. 2018. arXiv: 1808. 00191 [cs.CV]. [34] Jianwei Yang. Graph-RCNN Official Implementation.url:https://github.com/ jwyang/graph-rcnn.pytorch. (accessed: 23.12.2021). [35] Inc. Meta Platforms. Facebook AI Research Official Website.url:https : / / ai . facebook.com/. (accessed: 23.12.2021). [36] Thomas N. Kipf and Max Welling. Semi-Supervised Classification with Graph Convolutional Networks. 2017. arXiv: 1609.02907 [cs.LG]. [37] Matthew Klawonn. Combining Supervised Machine Learning and Structured Knowledge for Difficult Perceptual Tasks. Rensselaer Polytechnic Institute, 2019. [38] Ian J. Goodfellow et al. Generative Adversarial Networks. 2014. arXiv: 1406. 2661 [stat.ML]. [39] Martin Abadi et al. “TensorFlow: A system for large-scale machine learning”. In: 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 2016, pp. 265–283. url:https://www.usenix.org/system/files/conference/ osdi16/osdi16-abadi.pdf. [40] NVIDIA Corporation. CUDA Zone.url:https://developer.nvidia.com/cudazone. (accessed: 23.12.2021). [41] NVIDIA Corporation. NVIDIA cuDNN.url:https : / / developer . nvidia . com / cudnn. (accessed: 23.12.2021). [42] Barcelona Supercomputing Center. Barcelona Supercomputing Center Official Website. url:https://bsc.es/. (accessed: 23.12.2021). [43] Ji Zhang et al. Graphical Contrastive Losses for Scene Graph Parsing. 2019. arXiv: 1903.02728 [cs.CV]. [44] NVIDIA Corporation. Graphical Contrastive Losses for SGG Official Repository.url: https://github.com/NVIDIA/ContrastiveLosses4VRD/. (accessed: 24.12.2021). [45] NVIDIA Corporation. NVIDIA Corporation Official Website.url:https : / / www . nvidia.com/en-us/. (accessed: 24.12.2021). [46] Ranjay Krishna et al. “Visual genome: Connecting language and vision using crowdsourced dense image annotations”. In: International journal of computer vision 123.1 (2017), pp. 32–73. [47] Jingwei Ji, Rishi Desai, and Juan Carlos Niebles. “Detecting Human-Object Relationships in Videos”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2021, pp. 8106–8116. Study and Training of DL Models for SGG from non-static Environments 83 BIBLIOGRAPHY BIBLIOGRAPHY [48] Kaiming He et al. “Deep residual learning for image recognition”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2016, pp. 770–778. [49] Kaiming He et al. Mask R-CNN. 2018. arXiv: 1703.06870 [cs.CV]. [50] Yao Teng et al. Target Adaptive Context Aggregation for Video Scene Graph Generation. 2021. arXiv: 2108.08121 [cs.CV]. [51] Yao Teng et al. SGG by Context Aggregation Official Implementation.url:https: //github.com/MCG-NJU/TRACE. (accessed: 12.01.2022). [52] Yuren Cong et al. Spatial-Temporal Transformer for Dynamic Scene Graph Generation. 2021. arXiv: 2107.12309 [cs.CV]. [53] Yuren Cong et al. Spatial-Temporal Transformer for Dynamic SGG Official Implementation.url:https://github.com/yrcong/STTran. (accessed: 12.01.2022). [54] JetBrains. IntelliJ IDEA Official Website.url:https://www.jetbrains.com/idea/. (accessed: 16.01.2022). [55] Python Software Foundation. Python Official Website.url:https://www.python. org/. (accessed: 16.01.2022). [56] Charles R. Harris et al. “Array programming with NumPy”. In: Nature 585.7825 (Sept. 2020), pp. 357–362. doi:10.1038/s41586-020-2649-2.url:https://doi.org/10. 1038/s41586-020-2649-2. [57] Fabian Pedregosa et al. “Scikit-learn: Machine learning in Python”. In: the Journal of machine Learning research 12 (2011), pp. 2825–2830. [58] Adam Paszke et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. 2019. arXiv: 1912.01703 [cs.LG]. [59] J. D. Hunter. “Matplotlib: A 2D graphics environment”. In: Computing in Science & Engineering 9.3 (2007), pp. 90–95. doi:10.1109/MCSE.2007.55. [60] Linus Torvalds. Git Website.url:https://git-scm.com/. (accessed: 16.01.2022). [61] GitLab B.V. GitLab Official Website.url:https://about.gitlab.com/. (accessed: 04.01.2022). [62] Diederik P Kingma and Jimmy Ba. “Adam: A method for stochastic optimization”. In: arXiv preprint arXiv:1412.6980 (2014). [63] Jingwei Ji et al. Action Genome Official Website.url:https://www.actiongenome. org/. (accessed: 27.12.2021). [64] Jingwei Ji et al. Action Genome Official GitHub.url:https://github.com/JingweiJ/ ActionGenome. (accessed: 27.12.2021). [65] Jingwei Ji et al. Action Genome Official Annotations.url:https://drive.google. com/drive/folders/1LGGPK_QgGbh9gH9SDFv_9LIhBliZbZys. (accessed: 27.12.2021). [66] Xiaotian Han et al. Adding your own dataset to the SGG Benchmark from Microsoft. url:https://github.com/microsoft/scene_graph_benchmark#addingyour- own-dataset. (accessed: 27.12.2021). 84 Study and Training of DL Models for SGG from non-static Environments BIBLIOGRAPHY BIBLIOGRAPHY [67] Meta Research. Model Zoo and Baselines.url:https://github.com/microsoft/ scene_graph_benchmark/blob/main/MODEL_ZOO.md. (accessed: 27.12.2021). [68] Jia Deng et al. “Imagenet: A large-scale hierarchical image database”. In: 2009 IEEE conference on computer vision and pattern recognition. Ieee. 2009, pp. 248–255. [69] Tsung-Yi Lin et al. “Microsoft coco: Common objects in context”. In: European conference on computer vision. Springer. 2014, pp. 740–755. [70] Sami Bourouis et al. “Color object segmentation and tracking using flexible statistical model and level-set”. In: Multimedia Tools and Applications 80 (Feb. 2021). doi:10. 1007/s11042-020-09809-2. [71] Microsoft. SGG Model Zoo.url:https://github.com/microsoft/scene_graph_ benchmark/blob/main/SCENE_GRAPH_MODEL_ZOO.md. (accessed: 01.01.2022). [72] Google AI. Openimages VRD Challenge.url:https://storage.googleapis.com/ openimages/web/challenge.html. (accessed: 01.01.2022). [73] littleflow3r. Attention-based BiLSTM for Relation Extraction implementation.url: https://github.com/littleflow3r/attention-bilstm-for-relation-classification. (accessed: 02.01.2022). [74] Peng Zhou et al. “Attention-Based Bidirectional Long Short-Term Memory Networks for Relation Classification”. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Berlin, Germany: Association for Computational Linguistics, Aug. 2016, pp. 207–212. doi:10.18653/v1/P16- 2034.url:https://aclanthology.org/P16-2034. [75] Nitish Srivastava et al. “Dropout: a simple way to prevent neural networks from overfitting”. In: The journal of machine learning research 15.1 (2014), pp. 1929–1958. [76] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. “U-net: Convolutional networks for biomedical image segmentation”. In: International Conference on Medical image computing and computer-assisted intervention. Springer. 2015, pp. 234–241. [77] Tony Yiu. The Curse of Dimensionality.url:https://towardsdatascience.com/ the-curse-of-dimensionality-50dc6e49aa1e. (accessed: 03.01.2022). [78] scikit-learn developers. Principal component analysis (PCA) explanation.url:https: //scikit-learn.org/stable/modules/decomposition.html#principal-component- analysis-pca. (accessed: 03.01.2022). [79] Barcelona Supercomputing Center (BSC). High Performance Artificial Intelligence Research Group Official Website.url:https://hpai.bsc.es/. (accessed: 04.01.2022). [80] Gunnar A. Sigurdsson et al. Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. 2016. arXiv: 1604.01753 [cs.CV]. Study and Training of DL Models for SGG from non-static Environments 85 BIBLIOGRAPHY BIBLIOGRAPHY 86 Study and Training of DL Models for SGG from non-static Environments Appendix A ResNet-50 HORL Examples This appendix include some qualitative results using the ResNet-50 backbone. The different elements that can be found in Figure A.1 and Figure A.2 are already explained at the beginning of Section 5.4. 87