scieee AI-readable full text Open interactive document viewer

Addressing Multiple Object Tracking with Segmentation Masks

Bendaña Gómez, Manuel

Abstract

Multiple Object Tracking (MOT) aims to locate all the objects from a video, assigning them the same identities across all frames. Traditionally, this problem was addressed following the Tracking by Detection (TbD) paradigm, using detections represented by bounding boxes. However, bounding boxes can contain information from several objects, something that does not happen with segmentation masks. This work takes the ByteTrack MOT system as a starting point. Our proposal, ByteTrackMask, integrates a class-agnostic segmentation method and a segmentation-based tracker in ByteTrack in order to rescue tracks that would have been lost. Results over validation sets of MOT challenge datasets provide improvements in MOT metrics of interest like MOTA, IDF1 and false negatives.

Full text

Addressing Multiple Object Tracking with Segmentation Masks Manuel Benda˜ na G´ omez Master Thesis — Master in Computer Vision University Of Santiago de Compostela Santiago de Compostela, Spain [email protected] Supervisors: Manuel Mucientes Molina ([email protected]) Victor Manuel Brea S´ anchez (victor[email protected]) University of Santiago de Compostela Abstract—Multiple Object Tracking (MOT) aims to locate all the objects from a video, assigning them the same identities across all frames. Traditionally, this problem was addressed following the Tracking by Detection (TbD) paradigm, using detections represented by bounding boxes. However, bounding boxes can contain information from several objects, something that does not happen with segmentation masks. This work takes the ByteTrack MOT system as a starting point. Our proposal, ByteTrackMask, integrates a class-agnostic segmentation method and a segmentation-based tracker in ByteTrack in order to rescue tracks that would have been lost. Results over validation sets of MOT challenge datasets provide improvements in MOT metrics of interest like MOTA, IDF1 and false negatives. Index Terms—Multiple object tracking, segmentation, deep learning I. INTRODUCTION This Master Thesis is centered on Multiple Object Tracking. This topic consists of the detection of all the objects present in every frame of a video, maintaining their identities along time (Fig. 1) [1]. Nowadays, a problem like MOT is solved using Deep Learning strategies, having special relevance the use of Convolutional Neural Networks (CNNs) [2] and, more recently, Vision Transformers (ViT) [3]. Tracking by Detection (TbD) is a common approach followed by MOT systems. This type of solution is based on the idea shown in Figure 2, where the most relevant components referred to a given frame Ft+1 are the following: 1) A detection network identifies all the objects present in a video frame with bounding boxes —dett+1 in Figure 2—, this is, rectangles that delimit the position in which each object is located. 2) Through data association techniques, detections dett+1 and the position of the tracks that were obtained for frame t—Ptin Figure 2– are combined, obtaining the position of the tracks for the actual frame Pt+1. In this way, we can save the trajectories for all objects in a video. Sometimes, an extra module makes an estimation of the positions Ptonto frame t+ 1 before association. A lack of detection for a given track is marked as a miss, even if it is still present in the scene. In this situation, we would like to keep tracking the object, which would improve tracking metrics. Fig. 1. Example of a MOT20 [1] video frame. Green boxes represent the bounding boxes of each of the objects present in that frame. A solution for addressing this would be to predict the position of an object in the actual frame with a visual object tracker. Visual object trackers solve a different problem, as shown in Figure 3. Given an exemplar object Eand a certain frame Ft, a visual object tracker locates the position of that exemplar in that frame Pt. This approach is normally based on bounding boxes, however, this is a problem in situations like the one shown in Figure 4. Let us consider that object B is being tracked. Its bounding box is mainly influenced by information from object A. If we want to track B using said bounding box, an identity switch can be produced, this is, to start following object A instead of object B. We hypothesize that segmentation masks would be more accurate in such a situation. Figure 4 is an example of how segmentation masks would be able to perfectly distinguish both objects even in the presence of occlusions. This may help to avoid confusion between objects and improve tracking metrics. This type of solution introduces two main challenges. The first one is to generate the segmentation masks themselves. The second one is to track said masks. Our solution to these two challenges in this Master Thesis is to integrate a visual object tracking based on segmentation masks in the well-known MOT system ByteTrack [4]. The final system, which we name ByteTrackMask, includes the following contributions: 1 Fig. 2. Representation of the common structure that a MOT system has following the Tracking by Detection (TbD) paradigm. Fig. 3. Representation of the problem addressed by a Visual Object Tracker. Fig. 4. Example of a situation in a MOT17 dataset video in which two objects have a big overlap, but their precise segmentation masks would still be distinguishable. •We integrate ByteTrack MOT system with a new module that helps to keep tracking objects without a detection associated. •We include the Segment Anything Model (SAM) [5] to generate segmentation masks in a class-agnostic fashion in order to provide masks for those tracks of interest. •We incorporate a visual object tracking method that works with segmentation masks —AOT [6]—, taking as exemplars the masks generated with SAM. Our main hypothesis is that tracking those masks will boost the performance of the base ByteTrack. II. STATE OF THE ART A. Multiple Object Tracking Drawing conclusions and guidelines from state-of-the-art MOT results is not straightforward, as difficulties appear at different levels. Object detection leads to issues like deformations, total or partial occlusions, background confusion and illumination. More tracking-specific issues like identity switches or rapid direction changes are also hard to tackle [7]. In this Master Thesis we have focused on ByteTrack [4], which introduces a two-stage association mechanism under the 2 name of BYTE. BYTE runs a first stage focused on matching with high confidence detections followed by a second stage centered on the low confidence ones. ByteTrack is the result of a strong detector alongside this novel association method. In contrast, other tracking by detection solutions only keep high score detection boxes in a single stage. Besides ByteTrack, other proposals have been emerging. For example, GHOST [8] is introduced as a system that tries to preserve ideas from old TbD methods, analyzing where these models failed and solutions to increase performance by means of appearance features and a simple motion model. StrongSORT [9] is another system, designed with the goal of, revising initially a previous classic tracker as DeepSORT [10], addressing problems of missing association and missing detection. For that, it is introduced: an appearance-free link model for doing association without appearance information, and a Gaussian-smoothed interpolation to address missing detections. StrongSORT tries to serve as a new baseline for comparison between different MOT systems —specifically in the TbD field. We want to improve ByteTrack performance with our ByteTrackMask proposal. This is the point in which segmentation masks and tracking of those masks appears: we want to integrate them in the MOT system. B. Segmentation Another problem inside Computer Vision which has special interest is segmentation. We can define it as the classification of the pixels of an image [11]. If we are interested in classifying pixels with semantic labels, this is, in categorizing each pixel in a certain class, e. g. pedestrians, this is known as semantic segmentation. On the other hand, if we are interested in the distinction of the pixels of individual objects —for example, pixels that belong to each specific person—, then we talk about instance segmentation. Panoptic segmentation refers to a combination of both ideas [11]. As in our case we want to segment each object individually in order to track them with masks, the topic of interest for our work is instance segmentation. There are multiple strategies and networks for performing segmentation, from convolutional neural networks or fully convolutional networks to encoder-decoder based ones. It was common to find segmentation alongside detectors, like in the case of Mask R-CNN [12]. This approach allows to detect objects in an image, and at the same time generate a highquality segmentation mask for each instance. For that, based on Faster R-CNN detector [13], which had two outputs for classification and bounding box regression, Mask R-CNN adds a specific third branch for providing the mask. The proposal that has a lot of interest nowadays, and that is the one that also caught our attention in this work, is Segment Anything Model [5], abbreviated as SAM. Based on ideas from Natural Language Processing, a model for promptable segmentation is created. One of the benefits of this method is its class-agnostic nature, this is, for mask generation the class is not taken into account. Therefore, we are able to use this model with no need of a specific training for MOT challenge datasets. SAM has been applied to many tasks with success, like zero-shot single point mask generation, instance segmentation, edge detection and text prompting. The idea is, after checking the feasibility of using SAM for segmenting objects in images of MOTChallenge datasets, to integrate masks into ByteTrack. We have to take into account that in many cases bounding boxes contain information from other objects. But if we regard to the correct segmentation masks, objects can be distinguishable. Taking those masks for tracking may avoid confusions that ByteTrack could have with boxes. C. Visual Object Tracking Another problem related with object tracking is Visual Object Tracking or VOT [14], which refers to the idea of, given a reference frame and an object present in that frame, predict if it is present in subsequent frames of the video and its position. It is important to distinguish it from MOT, as we have to indicate in this case the position of the object in a first, reference frame. In addition, each object is managed independently, although some proposals that have been appearing in recent years [6], [15], [16] consider the possibility of working with multiple instances in parallel. In this work we focus in segmentation-based approaches for VOT. In that field, there is a tracker based on transformers [3] leading the state of the art actually, AOT, which means Associating Objects With Transformers. AOT essentially includes [6]: •An ID assignment mechanism for joint association and decoding of multiple targets. It includes an identification embedding to assign each target a different identification vector, and an identification decoder that predicts probabilities for all the targets. •A Long Short-Term Transformer (LSTT) framework, which allows hierarchical multi-object matching and backpropagation. Beginning with this idea, DeAOT [15] is also proposed as an evolution. It further improves AOT achievements and, compared to its predecessor: •Rethinks hierarchical propagation from AOT, decoupling it in 2 parallel branches: an object-agnostic visual branch, and an object-specific ID branch. •Includes a more efficient module for propagation, based on Single-Head Attention mechanisms, named Gated Propagation Module (GPM). The interesting idea proposed in this series of trackers and their performance also in scenarios like the ones in MOTChallenge datasets led us to use them for segmentationbased tracking. III. BYTETRACK BACKGROUND We give now some details of ByteTrack, which architecture is shown in Figure 5. It is a MOT system formed by many components, following the structure of Figure 2: 3 •Any object detector can be introduced into ByteTrack architecture. •Before data association, a Kalman Filter [17] is applied for positions of tracks in the previous frame. Kalman Filter does an estimation of the position in which the tracks would be located in the actual frame. •The association is based on a novel method proposed by the same authors [4], called BYTE, which has two association steps —or rounds— as mentioned in section II. We will consider always the new frame to predict as t+ 1, and the previous frame as t. Taking this into account, when we have to make a prediction for frame t+ 1, detections are firstly computed —dett+1 in Figure 5. A confidence score is provided with each of them. Detections with a score over a certain threshold τwill be considered of high confidence and used at the first association step —det hct+1 in Figure 5. The rest of them, will be the ones of lower confidence, used at the second step —det lct+1 in Figure 5. In parallel to the detector, Kalman Filter takes object positions from frame tto estimate their location in t+ 1 —trackt+1 in Figure 5. Here is the point where the association begins: •In a first step, high confidence detections det hct+1 and track estimations trackt+1 are matched by means of the Hungarian algorithm [18]. The detection-track pairs that were matched are associated and therefore we can consider those tracks as solved —P1stt+1 in Figure 5. Unmatched tracks u track 1stt+1 are sent to the second step, and unmatched detections u dett+1 are initialized as new tracks —P newt+1 in Figure 5. •In a second step, low confidence detections det lct+1 and the remaining tracks u track 1stt+1 are matched again using the Hungarian algorithm. Matched pairs are associated and therefore those tracks will be solved — P2ndt+1 in Figure 5. No further action is performed for unmatched detections, and unmatched tracks are marked as lost. If after a number of frames Nthose tracks are not recovered with posterior matches, they are deleted. Hungarian algorithm [18] is a well-known optimization method. As an input to it in ByteTrack, a matrix is provided, representing tracks in its rows and detections in its columns. In each position, the intersection over union (IoU) similarity metric is computed for the pair of bounding boxes (last track position, detection). Optimization process is performed afterwards, providing as a result the best set of possible matches between tracks and detections. It is not necessary that they all match, in fact, unmatched tracks and detections are also provided in separate lists. Intersection over Union, also known as Jaccard Index, is a metric that compares the similarity between two arbitrary shapes [19], being computed as follows: IoU(A, B) = |A∩B| |A∪B|(1) Where A and B represent the areas of each of the objects. IoU can take values between 0 and 1. In the worst case, this is, if A and B are disjoint, intersection would be zero, therefore, IoU would take a result of 0. In the best case, this is, if A and B are exactly equal, then intersection and union would have the same value, therefore IoU would take a result of 1. Also, this metric can be computed both for bounding boxes and segmentation masks. We have used IoU in this work in both situations. The final list of tracks with a position in frame t+1 obtained with ByteTrack will be composed of: tracks matched in a first round P1stt+1, unmatched detections from first round that are initialized as new ones P newt+1, and tracks matched in a second round P2ndt+1. They will be subsequently used in Kalman Filter for the next frame estimations. IV. OUR APPROACH: BYTETRACKMASK Taking as reference ByteTrack, we propose to modify data association, adding a third round after the first two associations as shown in Figure 6. Originally, tracks that were not matched in the second association step, u track 2ndt+1 in Figure 6, were directly marked as lost. As a result, we could not be tracking an object which is present on the video. In that situation, we would like to keep tracking the object, but after the two first rounds we do not have any extra detection that we can use for matching. Here is where segmentation masks-based tracking approach appears: all tracks still lost from the second association step are sent to a new module —3rd round estimation with masks, as shown in Figure 6— that will work with them and their masks. Therefore, with this approach we want to avoid losing tracking of objects which are still present in the video, either permanently or temporally. On posterior frames, we can keep those objects tracked either with our third round or in the first/second rounds from ByteTrack. This new third round of the association using masks is represented in Figure 7, which, at the same time, is divided in sub-parts or steps. Firstly, we need to obtain a segmentation mask for the object —mask extraction block from Figure 7— , which can be generated with SAM or from the previous prediction if a tracking with AOT was already running for that object. Once we get a mask, the prediction step is performed, in which AOT estimates the new position of the object regarding its mask —mask prediction block from same Figure 7. A last control step is necessary to check if the mask is coherent —track termination check in Figure 7. If so, the new position of the bounding box is computed and assigned to the track and therefore it will be rescued from being lost. A. Mask extraction For the mask extraction process, given an unmatched track after second round in frame t+1, we firstly check if there is a previous mask available in frame t. In other words, we check if there is an AOT instance active for tracking the corresponding mask for that object, since we will only have them active when performing tracking in this third round. If so, we take the mask from frame tfor the next block —Get AOT prediction 4 Fig. 5. Architecture of ByteTrack with the main components of interest for our work [4]. Fig. 6. Architecture of our approach —which we name ByteTrackMask— proposed to evaluate. in Figure 7. If not, this is, if we have just lost the track from first/second round, we will need to generate a mask —mask generation in Figure 7. For this mask generation process we need to know how SAM works. It follows an encoder-decoder strategy with multiple components, shown in Figure 8: •An image encoder allows us to create an embedding from the input, that only needs to be run once per image in order to begin prompting. •A prompt encoder allows to process the prompt that is given in order to generate the mask. This prompt, as can be seen in Figure 8, can be simply a point at which we know the object is present —or a set of points—, a bounding box surrounding it, or even text. Apart from this, there is the possibility of providing a mask, which is embedded in a different way —using convolutions, as also shown in Figure 8. •Mask decoder is used to map all the embeddings into a mask, which is given as an output of the model. It is worth noting that SAM allows to generate three masks per prompt, to represent the whole object, or a part or a subpart of the object. Each of the masks is accompanied by a confidence score provided by SAM, as an estimation of the IoU [5]. Our mask generation options are detailed in Figure 9. Figure 10 shows the full process that we propose. We can directly provide the bounding box assigned to the track in frame tas a prompt to SAM. A more refined second alternative consists on extracting a set of masks using only one point as a prompt. As we can see in the example of Figure 9, a bounding 5 Fig. 7. Details of the third round proposed for ByteTrackMask in order to try to achieve improvements. Fig. 8. Architecture of SAM including the possible prompts that can be an input to the model [5]. Fig. 9. Example of a mask generated with SAM using the bounding box as a prompt (left) or a correct point located inside the object (right). box can introduce information from other objects as well, making the segmentation process difficult, something that does not happen when providing a correct point from the object. Moreover, frame trepresents the moment before losing the object in first/second round. In that situation, object can be occluded —like in the scene shown in Figure 9— and it may be hard to extract the correct mask for an architecture like the one from SAM and a bounding box prompt. Centering in the second option of the point prompt, we need to provide a point inside the object to generate a correct mask —like the shown in Figure 9. For that, we consider a point grid, having as reference the bounding box that we would use as a prompt in the first option. All three possible masks are extracted for each point of the grid. After SAM execution, we get a set of candidate masks — for a M×Mgrid, M·M·3masks—, also in case we use the alternative of generating three masks from a single bounding box —although in that option only 3 would be obtained. Therefore, we need a filtering process in order to select the most correct mask for each object. For that, we propose the following filters, that are also shown in Figure 10 6 Fig. 10. Detail of the mask generation process inside our third round proposal for ByteTrackMask. inside Mask Filtering block, in this order: 1) Comparison with neighbors: select neighbor tracks and generate a mask using SAM and their bounding boxes — from frame t— as prompts. This will give an idea of the position of each of the neighboring objects. With that, we can filter out those candidate masks that overlap with any of the neighbors over a threshold α. Overlapping measure will be Intersection over Union (IoU) between each candidate mask and a frame mask representing a concatenation of the masks of each of the neighbor tracks. 2) Inside bbox (bounding box) analysis: we want to ensure that candidates have the most pixels of the mask within the reference bounding box —the one that indicates the position of the track in frame t. Proportion of pixels inside the bounding box must be greater than a certain value β. We are not interested in segmenting objects mostly present outside that rectangle. 3) Rank by Confidence: in case more than one mask passes the previous filtering process, we return the mask with the highest SAM confidence score. We introduce the previous filters since SAM does not take into account those conditions in the calculation of the confidence score, thus eliminating masks that we do not want to be selected. After this process, we can have two situations. In the first one, a mask is finally selected —with the rank by confidence— , case in which we advance in the process of this third round. The second one occurs when no mask is able to pass the two proposed filters —no valid mask, shown in Figure 10—, a case in which we will not continue trying to save that track as we do not have a mask that we consider good enough to associate it. B. Mask prediction We will only achieve this stage if, after mask generation step, a mask has been successfully generated by SAM or if we have a mask predicted with AOT in previous frame —this is, we have an active AOT instance running for that object. If we are on the first situation, this means that we have not began tracking with AOT. Therefore we need to initialize a new instance of the tracker providing the mask obtained in mask generation block as reference —AOT initialization, as we show in Figure 7. After initialization, or having an active instance of the tracker, we will make the prediction for the new frame t+ 1 —AOT prediction, as we show in Figure 7. C. Track termination check Once mask is predicted, we need to have some control in order to avoid letting non reliable AOT predictions go by and stop tracking process in that situation. For this, we propose a third stage in which two filters are introduced to check consistency between two consecutive frames. Both are represented in Figure 7: •Mask similarity: two masks between two consecutive frames, although different, should be similar. Therefore, we can measure the IoU between them and keep masks which IoU is over a certain threshold γ. •Bounding box similarity: we can apply the same process that we do with masks to bounding boxes: they should also be similar between consecutive frames. In particular, the bounding box that was assigned for the track in the previous frame t, and the one that we would generate for the track in t+ 1 taking into account the predicted mask. Similarity is again measured with IoU, that should be above a δthreshold. If track is not able to pass any of the filters, it will be marked as lost and not maintained in this frame. In other case —both filters passed— we would assign the new bounding box as the position of the track in t+ 1 —Bounding box estimation in Figure 7. As the system works with bounding boxes, position needs to be given to the track. This is done taking into account three elements: •The last position —bounding box in frame t. •The mask generated for frame t. •The mask estimated for frame t+ 1. We managed different options for updating the bounding box. They are shown in Figure 11 for a problematic situation in which we have a partial occlusion on the object of interest. Firstly, the most straightforward, would consist on assigning the bounding box that surrounds the mask predicted in frame t+1 —top-right image in Figure 11. However, bounding boxes are expected to represent where the object is located even in their occluded parts. As a result, if the object is partially visible —for example, just a head—, the bounding box would be different with respect to how it should be. An intermediate step that was performed consisted on moving the center of the last position onto the center of the mask estimated for frame t+ 1 —bottom-left image in Figure 11. However, we would have a problem due to a similar reason than in the previous case: if the object has occlusion, we would displace the bounding box to the center of the visible part of the object, which may not coincide with the real center of the object —maybe occluded. 7 Fig. 11. Detail of the bounding box estimation process alternatives. If we want to perform a more realistic displacement, basing it on the difference between the mask centers from previous and actual frames is the option which provides better results. We need to estimate better how the object displaced, and for doing so we propose to move the center of the bounding box at frame tfollowing the displacement between the center of the masks of frames tand t+1 —bottom-right image in Figure 11. In this way, even if the object is occluded, movement of the bounding box would be more realistic than in previous proposals. V. EXPERIMENTS AND RESULTS Regarding the experimentation made for this Master Thesis, we will show our main results for the proposed model, as well as an ablation study performed with different components. At the same time, interesting results that were extracted from a preliminary study stage will be shown. A. Datasets All testings made with our integration in ByteTrack were addressed in MOTChallenge datasets. MOTChallenge is a relevant challenge introduced in order to standardize MOT methods evaluation. In this work, we have used 2 datasets: MOT17, based on previous MOT16 [20] with improved ground-truth annotations; and MOT20 [1], the most difficult as it is based on crowded scenarios. Besides these two, another dataset named Multiple Object Tracking and Segmentation (MOTS) [21] was used, as it contains a ground-truth with segmentation masks. All these datasets have two subsets: one for training and another for testing. However, in ByteTrack implementation [22] we can find scripts that allow us to TABLE I INFORMATION OF MOT17 AND MOT20 VALIDATION SUBSETS VIDEOS: LENGTH —NUMBER OF FRAMES—, RESOLUTION AND DENSITY —AVERAGE NUMBER OF OBJECTS PER FRAME. DATASET VIDEO LENGTH RESOLUTION DENSITY MOT17 MOT17-02 300 1920 x 1080 31 MOT17-04 525 1920 x 1080 45,3 MOT17-05 418 640 x 480 8,3 MOT17-09 262 1920 x 1080 10,1 MOT17-10 327 1920 x 1080 19,6 MOT17-11 450 1920 x 1080 10,5 MOT17-13 375 1920 x 1080 15,5 MOT20 MOT20-01 214 1920 x 1080 46,32 MOT20-02 1391 1920 x 1080 55,62 MOT20-03 1202 1173 x 880 130,42 MOT20-05 1657 1654 x 1080 194,98 split MOT17 and MOT20 training sets into two, one part for training and one part for validation. In our case, most of the experiments were performed in MOT17 [20] validation half, although some extra evaluations were done in MOT20 [1] validation half to have another reference. It is important to mention that for using test set we would need to compute the results and upload them onto an evaluation server with limited amount of submissions per proposal. Take into account that annotations of this set are not publicly available. For this work, our goal was to check the performance of our module and different variants of it. Therefore, submitting results onto MOT servers would be impractical and at the same time not purely correct, as we could start to compare variants over the test set. That comparison would be expected to be done in validation subsets. Table I shows resolutions and lengths of each of the videos of each validation subset. We can see in this table the differences between MOT17 and MOT20 in number of videos, frames and objects per frame, expecting more crowded scenarios in MOT20 dataset. Resolutions vary in each video, but it is common to have 1920 x 1080. For evaluating how well a method performs in both datasets, many metrics are proposed. In this work we will focus on MOTA, IDF1, FP and FN. FP and FN stand for the number of false positives and false negatives. MOTA is the abbreviation of Multi Object Tracking Accuracy, and it is one of the most common metrics used to evaluate how well does a tracker perform [1]. The following equation defines how it is computed: MOT A = 1 −Pt(FNt+FPt+IDSWt) PtGTt (2) Where FN and FP stand for false negatives and false positives, respectively, IDSW represents the number of identity switches that were produced and GT represents the number of ground truth objects. In addition, tindicates the frame. As we can see, MOTA is obtained dividing the sum of false positives, false negatives and identity switches by the total amount of ground truth objects per frame, and subtracting that value from 1. Therefore, this metric can take values as a percentage in the 8 interval (−∞,100], having the possibility of a negative MOTA in case the number of errors is too high [1]. IDF1 metric stands for ID F1 score, and it represents the ratio of correctly identified detections over the average number of ground-truth and computed detections. IDF1 allows to rank trackers in a single scale balancing ID precision and recall [23]. Its equation can be represented as follows: IDF 1 = 2·IDT P 2·IDT P +IDF P +IDF N (3) Where IDTP represents the number of true positive IDs, IDFP is the number of false positive IDs and IDFN is the number of false negative IDs. This metric can take percentage values in the interval [0, 100]. In the worst case —0 true positives— IDF1 would take a result of zero, and in the best case —neither false negatives nor false positives—, a perfect value would be achieved. B. Implementation details Before showing the results, we present some implementation details. All the implementation was developed in Python language, taking into account that we have libraries like PyTorch, optimized for performing Deep Learning operations, or other libraries that we know from many subjects of this master like NumPy, OpenCV, Matplotlib and Torchmetrics. Python was also the language used in the different GitHub repositories that we needed for implementing our module. We include a detailed list of them in this section: •ByteTrack [22]: the code of the MOT system is taken as reference for our work. •AOT Series Frameworks [24]: it contains the implementation of AOT, including DeAOT version. •Segment Anything [25]: the implementation of Segment Anything Model, other of the components that was integrated in our work. •MOTS tools [26]: the implementation of some tools that allow us to evaluate performance in the Multiple Object Tracking and Segmentation (MOTS) challenge. •COCO API [27]: it is a repository containing APIs in different languages for managing COCO dataset annotations [28]. COCO is a very famous dataset inside visual recognition, that can be used for tasks such as object detection, captioning or object segmentation. In particular, with this API we have the possibility of saving and loading segmentation mask annotations to or from a file, thanks to its encoding and decoding tools. •TrackEval [29]: it contains code for, given a groundtruth and corresponding results files from a MOT system, evaluate performance with a wider set of metrics than the one that a direct execution of ByteTrack has. Experimentation is normally expensive, therefore, we needed access to GPUs in order to perform general purpose operations thanks to CUDA [30], the parallel computing platform by NVIDIA. We were able to get access to two GPGPU computation servers from Centro Singular de Investigaci´ on en Tecnolox´ ıas Intelixentes (CiTIUS) at the University of Santiago de Compostela: ctgpgpu6 and ctgpgpu11. Normally, we used the first one, which has the following properties: •2 processors Intel Xeon Silver 4214. •Four GPUs: –One NVIDIA Quadro P6000 with 25 GB. –One NVIDIA Quadro RTX 8000 with 50 GB. This is the GPU we commonly used in all experimentation. In particular, for the results shown in this report, it was the GPU in which those experiments were run. –Two NVIDIA A30 with 25 GB each. •CentOS Linux Operating System, in version 7.7, including CUDA driver in version 12.2. Regarding details from the execution, we have taken the weights of ByteTrack ablation model —available in their GitHub [22]. It was trained using CrowdHuman and half of the training set of MOT17, as the other half was used for validation. YOLOX [31] was selected as the object detector, with YOLOX-X backbone. Regarding ByteTrack specific parameters, they are aligned with the default values used in their proposal, including confidence τfor classifying detections into high and low confidence, which was set to 0,6. The number of frames Nthat a track is preserved after being lost and before setting it as deleted is also included and it is set to 30. For SAM, weights were downloaded from their GitHub as well [25], in particular for the largest backbone size (ViT-H), being trained in the large 11 million images dataset SA-1B —proposed also by their authors [5]. In the case of AOT, weights were taken again from GitHub [24], and they were the ones of the DeAOT model which yielded best results, this is, DeAOT-L with SwinB transformer backbone. It was trained in YouTube-VOS [32] and DAVIS [33] datasets after a pre-training with many static image datasets. Regarding our thresholds, we have tried to remain with the same global value of 0,9 for β,γand δ, being inverted in the case of αas it is used in the comparison with neighboring masks and we want to check if there is a minimum overlap with them —therefore, αis set to 0,1. We have generated SAM masks using the grid approach, following the idea proposed in Section IV for implementing the module. Masks were generated with a 3 x 3 point grid, having a total of 3·3·3 = 27 masks for each case. We have also not considered the inference tricks introduced in ByteTrack. Some later works [8], [9] compared their results with ByteTrack, but they showed different scores for that system with respect to the ones originally presented in [4]. This occurs as this MOT system uses an offline interpolation that fills lost frames of tracks, which leads to an extra increase in the results. Moreover, ByteTrack introduces a custom parameter setting depending on the videos from the MOTChallenge datasets, in order to further improve its results —as an example, value Nwas changed from the default 30 depending on the video. We follow in our experimentation the 9