Full text
Master in Sound and Music Computing Universitat Pompeu Fabra Understanding Audio Source Separation in Carnatic Music with Multimodal Data Théo Fuhrmann Supervisor: Martín Rocamora Co-Supervisor: Gloria Haro August 2025
Contents Abstract Acknowledgement 1 Introduction 1 1.1 Motivation.................................. 2 1.2 Objectives.................................. 2 2 State Of The Art 4 2.1 SourceSeparation.............................. 4 2.1.1 Audio Source Separation . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.1.2 Audio-Visual Source Separation . . . . . . . . . . . . . . . . . . . . . . 6 2.2 Interpretability and Model Analysis . . . . . . . . . . . . . . . . . . . . 7 2.3 Datasets................................... 8 2.4 CarnaticMusicinMIR........................... 9 3 Dataset and Preprocessing 10 3.1 DatasetOverview.............................. 10 3.2 Pose Estimation and Instrument Labeling . . . . . . . . . . . . . . . . 11 3.3 FeatureExtraction ............................. 11 3.3.1 Motion.................................... 11 3.3.2 Audio .................................... 13 3.4 Synchronization and Source Alignment . . . . . . . . . . . . . . . . . . 13 3.4.1 DataFiltering................................ 14
3.5 Limitations ................................. 14 4 Correlation and Timing Analysis 15 4.1 Sliding Window Correlation Analysis . . . . . . . . . . . . . . . . . . . 15 4.1.1 Speed and Acceleration . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 4.1.2 Vocal-Specific Features . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 4.1.3 Violin-Specific Features . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 4.2 Cross-Correlation and Lag Estimation . . . . . . . . . . . . . . . . . . 20 5 Model Analysis and Interpretability 22 5.1 OverviewofModels............................. 22 5.2 VocalistModel ............................... 22 5.2.1 Model Architecture and Training . . . . . . . . . . . . . . . . . . . . . 22 5.2.2 Gradient-Based Interpretability . . . . . . . . . . . . . . . . . . . . . . 24 5.2.3 AblationStudies .............................. 32 5.3 Vocal&ViolinModel ........................... 34 5.3.1 Model Architecture and Training . . . . . . . . . . . . . . . . . . . . . 34 5.3.2 AttentionAnalysis ............................. 35 5.3.3 FiLMAnalysis ............................... 37 6 Discussion and Conclusion 39 6.1 GeneralDiscussion ............................. 39 6.2 Conclusions ................................. 40 6.3 FutureWork................................. 41 Bibliography 43 .1 Temporal Correlation Analysis Plots . . . . . . . . . . . . . . . . . . . 50 .2 Input ×GradientPlots........................... 51 .3 Integrated Gradient Plots . . . . . . . . . . . . . . . . . . . . . . . . . 52 .4 Attention Analysis Plots . . . . . . . . . . . . . . . . . . . . . . . . . . 53
Abstract This thesis investigates the internal mechanisms of audio-visual source separation models, focusing on how performer motion guides source separation in Carnatic music. To move beyond "black box" performance metrics, we employ a dual approach combining model-independent analysis of audio-visual synchrony with model-specific interpretability techniques applied to Voice-Vision Transformer (VoViT) architectures. A cross-correlation and sliding-window analysis on the Saraga dataset first establishes instrument-specific temporal patterns, revealing that targeted instrumentspecific motion features exhibit stronger and more consistent correlations with audio dynamics than general body motion. While global linear audio-visual synchrony is weak, these analyses highlight the importance of localised and instrument-specific motion cues. Subsequently, we apply gradient-based saliency methods to a vocal separation model, demonstrating its primary reliance on facial keypoints. Ablation studies causally confirm that these facial regions are crucial for vocal separation, while body motion contributes minimally. We further analyze the attention mechanisms of a vocal & violin model to understand how it disentangles spectrally overlapping sources, specifically through a revised hybrid fusion architecture. FiLM (Feature-wise Linear Modulation) analysis reveals that visual information causally modulates audio features by consistently amplifying or suppressing them, acting as a dynamic gating mechanism. This research provides a framework for interpreting gesture-based audio-visual systems, offers novel insights into audio-visual learning, and contributes to the development of more transparent and culturally-aware Music Information Retrieval technologies. Code available at: https://github.com/theofuhrmann/masters-thesis Keywords: Model Interpretability; Audio-visual; Source Separation; Carnatic Music;
Acknowledgement I would like to express my gratitude to Adithi for her time and for providing me with the resources to carry out this thesis. I am also grateful to Martín and Gloria for their helpful feedback, and to my partner Isabelle for her constant support throughout the process. Finally, I thank Gabriel, a former coworker, for lending me his GPU during most of the thesis, which allowed me to explore different directions without compromise.
Chapter 1 Introduction Isolating a single sound source from a complex auditory scene, known as audio source separation, is a key task with applications ranging from hearing aids and speech recognition to music production and remixing. Early approaches were based on statistical signal processing, but the rise of deep learning has brought major advances, enabling models to separate complex musical mixtures with impressive accuracy. In parallel to advances in audio-only methods, researchers have explored multimodal approaches that incorporate visual cues (such as lip movements or musical gestures) to help disambiguate sound. These audio-visual models complement audio-only systems and have opened new possibilities for improving separation. However, their growing complexity often comes at the cost of interpretability. As these deep learning models become harder to analyze, it becomes increasingly difficult to understand how they make decisions. This opacity limits our ability to trust, debug, and improve them, and makes it harder to see possible connections between sight and sound. This thesis tackles the challenge by shifting the focus from performance metrics to model interpretability. We explore how state-of-the-art audio-visual systems use visual information to separate sound sources. To ground our analysis, we focus on Carnatic music, a classical tradition from Southern India known for its rich acoustic textures and complex performance practices. Its unique gestures and expressive style 1
8Chapter 2. State Of The Art For instance, Meyes et al. [31] evaluated the contribution of different neural layers by systematically removing specific units and observing the resulting performance changes on the MNIST dataset. In this work, we apply gradient-based saliency methods to structured visual inputs (2D facial and body landmark coordinates) to identify which landmarks most influence the model’s voice separation output. Unlike typical saliency methods applied to raw images, this approach provides direct insight into how motion and pose guide the model’s predictions, offering a novel perspective on the role of visual cues in audio-visual source separation. 2.3 Datasets Music Information Research (MIR) has traditionally focused on Western music, often underrepresenting the diverse range of musical traditions worldwide. This imbalance has motivated recent efforts to incorporate non-Western music styles into MIR studies. As highlighted by Serra [32], adopting a multicultural approach is crucial to developing more inclusive and representative computational methodologies. Notable Western-focused datasets used for multimodal audio source separation include AudioSet [33], a large collection of 10-second YouTube videos manually labeled among 632 audio classes; MUSIC [17], which provides 714 labeled YouTube recordings of musical performances; and URMP [34], a high-quality manually recorded audiovisual dataset of multi-instrument classical music pieces. Building on recent efforts to expand MIR research beyond Western music, we use the Saraga Audiovisual dataset [35], a recently developed large-scale multimodal collection designed for the study of Carnatic music. Built on the principles of the original Saraga dataset [36], which included only audio recordings, this dataset contains 42 recorded Carnatic concerts, totaling over 60 hours of audio-visual data, with multi-track audio recordings and synchronized video footage. Additionally, the Sanidha dataset serves as another valuable audio-visual resource for Carnatic music research [37]. It contains 5 high-definition recordings of Carnatic
2.4. Carnatic Music in MIR 9 concerts, totaling 15.35 hours. Each performer was recorded separately, with simultaneous video recordings conducted in different rooms to ensure clean, isolated audio and synchronized visual data. These datasets support ongoing efforts to diversify MIR and advance methods for non-Western music. 2.4 Carnatic Music in MIR Despite the dominance of Western music in MIR, there has been a growing recognition of the need to study non-Western musical traditions. Carnatic music, with its unique raga-based structure, intricate ornamentation, and complex rhythmic patterns, presents both significant challenges and exciting opportunities for MIR. Early studies explored Carnatic music through different MIR techniques. For example, research on singer identification in Carnatic music [38] leveraged the 22-tone octave system to develop cepstral coefficients for capturing distinct vocal characteristics. Later, a study on computational approaches for understanding melody in Carnatic music [39] emphasized the need for tailored methodologies in music information processing for Carnatic music. Another example is a study on raga classification [40], which proposed a new system to classify music into ragas using different audio features. More recently, advances in source separation have addressed challenges in Carnatic singing. Plaja-Roglans et al. [41] trained a cold diffusion model to mitigate audio bleeding in performance recordings. In another study, they developed a system to generate vocal pitch annotations using the Saraga dataset, creating the SaragaCarnatic-Melody-Synth (SCMS) dataset, which was then used to train a state-ofthe-art pitch extraction model for Carnatic music [42]. Building on this, Shankar et al. [43] leveraged the Saraga audio-visual dataset to adapt the VoViT architecture for gesture-guided source separation in Carnatic concert recordings. They explored different audiovisual fusion strategies and showed that integrating facial and body keypoints improves vocal and violin separation, even in the presence of source bleeding and partial visual occlusion.
Chapter 3 Dataset and Preprocessing 3.1 Dataset Overview We focused on performances with a consistent left-to-right layout (mridangam, vocal, and violin) as this setup provides a more controlled environment for multimodal analysis. In this configuration, side artists typically face inward toward the vocalist, causing one side of their body to be partially or fully occluded. This asymmetry in visibility can affect the reliability of pose estimation, so we limited our selection to concerts with a consistent layout to ensure that the same side of the mridangam and violin players remained visible throughout. For this reason, we used only 26 concert recordings from the Saraga Audiovisual collection, totaling around 28 hours of audio-visual material. A minority of performances followed an alternate layout (violin, vocal, and mridangam), which were excluded to ensure pose estimation consistency. However, due to improved bow visibility, these 3 recordings are revisited in Chapter 4 for a violin-specific case study involving more specialized motion features. Although limited to 3 concerts, this violin subset provides nearly 4 hours of music, offering sufficient within-performance variability to analyze audiovisual correlations, while generalization to other performers remains a caveat. 10
3.2. Pose Estimation and Instrument Labeling 11 3.2 Pose Estimation and Instrument Labeling Pose estimation was performed using the RTMW-X (384x288) model from MMPose, following the setup by Rodrigues [44]. Since performers remain seated in fixed positions throughout each concert, mean centroids were computed across the full duration of each performance. Detections with low temporal presence were discarded as likely false positives, such as standing spectators. To filter the remaining false positives, mainly wall artwork depicting human figures, the estimations above a certain vertical threshold were excluded. Each valid estimation was then assigned to one of the three instruments based on a manually generated left-to-right performer layout annotated for each video. The assignments were verified by visualizing short clips with color-coded skeleton overlays, one color per instrument, confirming both spatial consistency and pose estimation quality (see Figure 1). Figure 1: Color-coded pose estimation visualization. 3.3 Feature Extraction 3.3.1 Motion Keypoints were normalized per performer by centering and rescaling around their spatio-temporal mean to ensure consistent scale and position. Lower body keypoints were discarded due to limited movement (as performers remain seated) and inaccurate pose estimation from clothing and instrument occlusions. For the vocalist, both arms were used; for side instrumentalists (mridangam and violin), only the visible arm was retained due to frequent occlusions.
12 Chapter 3. Dataset and Preprocessing For the initial analysis, speed and acceleration were computed using NumPy’s gradient function over the 2D keypoints extracted using MMPose, and averaged over the full upper body as well as specific regions: left arm, right arm, head, face, left hand, and right hand. In addition to these general motion descriptors, two case studies were conducted focusing on more domain-specific movement cues: Vocal-specific features: To capture gestures directly related to vocal production [15, 45], we focused on mouth and jaw motion. We extracted 3D face keypoints using the 3DDFA face pose estimator, which is used in the vovit-fb model analysed in Chapter 5. The 3D nature of the keypoints allowed us to rotate the face to a frontal position, minimizing the influence of head pose on the measurements. During inspection, we found that 3DDFA’s face tracking sometimes failed when the vocalist’s face was occluded or turned away, producing inaccurate landmarks. To address this, we used the more stable MMPose nose keypoint as a reference, discarding frames where its distance from the 3DDFA face centroid exceeded a set threshold. This reduced noise in the vocal-specific features. From the filtered and aligned keypoints, we computed two features to capture facial movements. The first is mouth area, which combines mouth width and height to approximate the degree of mouth opening. The second is nose-to-jaw distance, which tracks the vertical displacement of the jaw. Violin-specific features: For violinists with non-occluded bowing arms, we computed additional motion descriptors using the 2D keypoints from MMPose to capture bowing gestures more directly. We chose these features based on their clear relevance in prior biomechanical studies [46, 47]. The first, wrist velocity, captures the speed of the bowing motion. The second, elbow angle, is defined by the angle between the shoulder, elbow, and wrist, reflecting articulation changes during bow strokes. Finally, arm extension measures the Euclidean distance between the shoulder and wrist, indicating the degree of arm reach during performance.
3.4. Synchronization and Source Alignment 13 3.3.2 Audio Audio features include onset envelope and Root-Mean-Square (RMS) energy, both extracted using Librosa. The onset envelope highlights sudden changes in energy, often corresponding to note or syllable attacks, while RMS energy provides a measure of overall loudness over time. Prior work has shown that these temporal energy variations are closely coupled with articulatory motion in audiovisual speech [45]. For the mridangam, stereo tracks were mixed into a single mono signal before calculating the features. 3.4 Synchronization and Source Alignment To enable multimodal frame-wise analysis, audio and video data were temporally aligned. Although most Saraga videos were recorded at 30 frames per second (fps), some performances had slightly different frame rates. Videos recorded at 29.99 fps were resampled to exactly 30 fps for consistency. Videos recorded at 24 fps were left unchanged to preserve their original temporal resolution. Each song’s fps was stored in the metadata. Pose estimation was performed at the original (or adjusted) video fps, and motion features were computed accordingly. To enable a framewise audiovisual correlation analysis, the 48 kHz audio was resampled using linear interpolation to match the number of video frames. The video data was recorded concurrently with the performance, while the audio was captured separately in four different tracks (one per instrument, with the mridangam using two tracks). Although these sources originate from the same performance, they were pieced together manually, which could cause a potential desynchronization due to the editing, recording or post-processing artifacts. While no obvious desynchronization was perceived during qualitative inspection, potential misalignment is formally assessed in Chapter 4.
14 Chapter 3. Dataset and Preprocessing 3.4.1 Data Filtering Pose estimation confidence varied across instruments and frames. In total, 0.79% of frames contained NaN values, and 5.5% had low confidence scores 1. Mean confidence scores per instrument were as follows: mridangam: 7.16, violin: 7.11, vocal: 8.00. This distribution is expected, as vocalists typically face the camera directly and are subject to less occlusion. Frames with confidence below the confidence threshold were discarded, and these missing or unreliable values were masked in all subsequent analyses. No interpolation or padding was applied given the relatively small amount of unreliable data. As a result, all subsequent analyses were conducted only on highconfidence, temporally aligned frames2. 3.5 Limitations Despite the care taken during preprocessing and annotation, several limitations affect the dataset, which are considered when interpreting motion–audio relationships in the subsequent chapters: •Occlusions: The vocalist is sometimes partially occluded by the microphone, and the lateral positioning of the violinist and mridangam player leads to frequent self-occlusion or loss of hand detail. •Pose Estimation Detail: The resolution of the original recordings limits the fine-grained accuracy of pose estimations, particularly for fast hand movements or finger articulation. •Audio Bleed: Although the audio was recorded in isolated tracks, all instruments were captured in the same physical space, leading to audio bleeding across tracks. This may reduce the precision of instrument-specific audio features and complicate their correlation with motion data. 1The confidence threshold was set to 3, with scores ranging up to ≈11 2Some instruments, such as the mridangam, naturally exhibit a delay between physical motion and corresponding audio due to the physics of sound production and human performance. No corrective temporal alignment was applied to account for these expressive lags. Their impact is analyzed later in Chapter 4 in the context of audiovisual correlations.
Chapter 4 Correlation and Timing Analysis This chapter explores the temporal relationships between motion and audio features across different instruments. We analyze how facial and body movements correlate with sonic events over time and discuss the implications of observed lags or synchrony between modalities. 4.1 Sliding Window Correlation Analysis 4.1.1 Speed and Acceleration To capture local temporal dependencies between modalities, we apply a sliding window correlation approach between motion and audio features. Pearson correlations are computed over a 0.5-second window with a 0.1-second step size. To focus on meaningful interactions, only correlations with an absolute value above 0.5 are retained1. Strong correlation windows are identified across four feature pairs: motion speed vs. audio onset, motion speed vs. audio RMS, motion acceleration vs. audio onset, and motion acceleration vs. audio RMS. The number of high-correlation windows is then aggregated per body part and per instrument to highlight which regions contribute most consistently to the audio signal across all feature pairs. 1Both strong positive and strong negative correlations between audio and motion features are treated as equally meaningful, as they each reflect consistent relationships. 15
16 Chapter 4. Correlation and Timing Analysis A large portion of the dataset exhibits strong correlations between motion and audio features. Overall, 95.5% of the total duration contains at least one strongly correlated window. By instrument, coverage is 60.8% for mridangam, 76.7% for vocals, and 60.9% for violin. To account for temporal overlap, we also computed coverage using non-overlapping windows (0.5 s hop), finding 64.6% overall, with 26.3% for mridangam, 37.9% for vocals, and 25.6% for violin. These results suggest that a notable fraction of the dataset may exhibit audio-motion coupling, with the effect persisting even when accounting for window overlap, providing a modelindependent perspective before introducing audio-visual source separation analyses. Figure 2: Number of strong windows per body part and instrument. In Figure 2, the vocalist shows the highest number of strong correlations, likely due to longer active performance durations and frontal visibility, which improve keypoint estimation. For the instrumentalists, correlations are slightly higher in the arms/hands than in the face or head, consistent with their role in sound production. The vocalist also shows more hand than face correlations, possibly reflecting expressive gestures. The mridangam performer shows fewer correlations overall, as only the non-occluded arm was visible, representing half of the instrument’s activity. More generally, the relatively uniform counts across body parts, especially for side-facing performers, suggest that occlusions and pose estimation errors dilute motion signals in active regions like the hands, while more consistently estimated areas such as the face contain stable but less informative correlations. These aggregate results should therefore be interpreted with caution given the limitations noted in Section 3.5.
4.1. Sliding Window Correlation Analysis 17 4.1.2 Vocal-Specific Features We applied the sliding-window correlation method to two vocal-specific facial features (mouth area and nose-to-jaw distance) extracted from the 3D face keypoints. To ensure reliability, only frames where the 3DDFA face estimation passed the accuracy check described in Section 3.3.1 were included. Case study. To better illustrate the analysis, we examined a short performance segment from Ameya Karthikeyan – Jalajakshi and plotted the temporal evolution of correlations between the two vocal-specific motion features (mouth area and noseto-jaw distance) and the audio features (RMS energy and onset envelope). Only the first 30 seconds of the 41-second performance were retained after filtering unreliable 3DDFA frames. As shown in Figure 3, both audio features follow a broadly similar temporal trend, although RMS energy consistently reaches higher correlation values than the onset envelope. Figure 3: Temporal evolution of vocal motion features vs audio features correlation throughout the performance of Ameya Karthikeyan - Jalajakshi Both vocal-specific features produced comparable counts of strong correlation windows: 48 for the onset envelope and 209 for RMS energy (combined over the two
24 Chapter 5. Model Analysis and Interpretability raises the question of whether including body keypoints introduces useful cues or merely noise, motivating our interpretability analysis to examine how each visual modality contributes to voice separation. 5.2.2 Gradient-Based Interpretability To understand how our model leverages visual input for voice separation, we apply gradient-based attribution methods. While gradients are typically used during training to update model parameters by computing ∂L ∂θ , in interpretability contexts, we instead compute the gradient of the loss Lwith respect to the input x. This allows us to estimate how sensitive the output is to changes in each input feature, essentially measuring the importance of our audiovisual features. We apply all attribution methods over 4-second audio-visual chunks, using the Mean Squared Error (MSE) between predicted and ground-truth audio as the target loss. Vanilla Gradients The simplest approach computes the gradient of the loss with respect to each input: Saliencyi=∂L ∂xi . These raw gradients indicate which inputs would most influence the output if slightly perturbed. However, this method is often noisy due to gradient saturation or nonlinearity in the model’s activations. In practice, we found that the gradients had very low-magnitude attribution scores, and even after applying a constant scaling factor of 216 (which was later used for all methods for consistency), they were hard to interpret. Input ×Gradient To improve interpretability, we used the Input ×Gradient method, defined as: Saliencyi=xi·∂L ∂xi .
5.2. Vocalist Model 25 This method approximates a first-order Taylor expansion of the model’s output around a zero baseline, combining the direction of influence (the gradient) with the input’s actual magnitude. This makes attributions more meaningful than raw gradients, especially when feature scales differ, such as between subtle facial motions and broad hand gestures. The implicit zero baseline serves as a reference point, interpreting each input’s contribution relative to its absence [50]. As with vanilla gradients, we applied a scaling factor of 216 to amplify the small attribution values and aggregated saliency scores across time and space to obtain body-part-level insights. Compared to vanilla gradients, Input ×Gradient produced clearer and more stable results, highlighting facial regions, particularly the nose and mouth. In the aggregated analysis (using the mean saliency per part), body regions showed lower importance, with the head standing out slightly among them, reinforcing the model’s reliance on facial cues. For the temporal evolution, we used the total saliency per part to reflect not only the relative importance but also the absolute contribution of each region over time. Using a zero baseline for all inputs (especially the face) remained a limitation, later addressed with a more realistic baseline in the Integrated Gradients method. Full visualizations are included in Appendix .2. Integrated Gradients To address limitations of Input ×Gradient, such as incomplete attributions and the sensitivity to arbitrary baselines, we implemented Integrated Gradients (IG) [51]. IG computes the average gradient along a straight-line path from a baseline input ˜ xto the actual input x: IGi(x)=(xi−˜xi)Z1 0 ∂L(˜ x+α(x−˜ x)) ∂xi dα
26 Chapter 5. Model Analysis and Interpretability We approximate the integral using 25 steps1, computing the gradient of the MSE loss at each interpolated input. Baseline selection was adapted to reflect the model’s preprocessing pipeline. The audio waveform and body keypoints are both normalized and centered around zero, making a zero baseline appropriate. However, the model centers face keypoints around a dataset-wide mean during preprocessing, but unlike the other modalities, these coordinates are not normalized or scaled. This mismatch can introduce a bias in the model’s internal representations and potentially affect how attribution methods interpret the role of facial inputs. To mitigate this, we used the datasetmean face as the baseline for facial keypoints in the Integrated Gradients method. While this choice improves alignment with the model’s actual input space, it doesn’t fully resolve the interpretability challenges introduced by the lack of normalization, particularly when comparing visual modalities2. To check the impact of this baseline mismatch, we also compared attributions obtained with Integrated Gradients against those from a simpler input ×gradients method. Using the case study performance, both methods produced highly similar distributions of saliency across regions (e.g., face regions, Spearman ρ≈0.86; body regions, ρ≈0.93), suggesting that the choice of baseline does not substantially alter the relative importance patterns, even if absolute magnitudes differ. Case study. We first illustrate the method in more detail with a single performance. For each 4-second chunk, we compute a saliency score by averaging gradient values across time and spatial dimensions. These scores are then used in two complementary analyses: a global one that averages saliency over the entire performance, and a temporal one that tracks how saliency evolves throughout the song. The resulting saliency scores confirm earlier trends: facial regions (especially the nose, followed by the mouth) dominate the attributions. As shown in Figure 6c, 1The authors recommend using between 20–1000 steps. We selected 25 for computational cost reasons, and verified its adequacy by comparing results with 50 steps: distributions of facial saliencies were nearly identical (symmetric Kullback-Leibler ≈3.9×10−6, Spearman ρ= 1.0). https://github.com/ankurtaly/Integrated-Gradients/blob/master/howto.md 2This limitation should be kept in mind when interpreting results and comparing saliency across face and body regions.
5.2. Vocalist Model 27 (a) Body keypoints weighted by IG scores (b) Face keypoints weighted by IG scores (c) Averaged IG saliency scores for different body and face parts. Figure 6: Spatial and averaged visualization of Integrated Gradients scores for visual input modalities in the performance of Abhiram Bode - Entha Bhagyamu using VoViT-fb. Landmark radius size indicates averaged saliency score. the more accurate baseline made the body and face scores more comparable, yet the overall distribution pattern remains. Among body parts, only the head shows a notable contribution, reinforcing the model’s reliance on facial cues for voice separation. The facial visualization on Figure 6b reveals that the keypoints around the inner and outer lips, particularly those near the center of the mouth, exhibit the highest saliency, which aligns with the more pronounced movement of these points during mouth opening and closing. There is also a noticeable saliency concentration on
28 Chapter 5. Model Analysis and Interpretability the chin, likely reflecting the role of jaw motion during singing. Another interesting find is the strong attribution present in the nose across all its keypoints, especially along the bridge. As for the body (Figure 6a, although head keypoints remain the most salient, some fingers on the right hand also show elevated saliency. This may be linked to rhythmic hand gestures or subtle movements that coincide with vocal fluctuations, suggesting the model captures not only facial articulation but also complementary body motion cues3. A more detailed view of individual keypoint contributions, grouped by body region, is provided in Appendix .3. Figure 7: Temporal evolution of saliency scores for major body regions throughout the performance of Abhiram Bode - Entha Bhagyamu using VoViT-fb Moving onto the temporal analysis, Figure 7 provides a temporal breakdown of saliency scores for broader regions. While the overall importance of regions fluctuates over time, their relative rankings remain mostly stable. Facial regions consistently dominate, followed by the right hand and head. Interestingly, saliency scores across 3Although hand and finger keypoints occasionally exhibit elevated saliency, masking experiments (Table 2) show little causal effect on separation quality. These visualizations may therefore reflect incidental correlations or model noise rather than strong reliance on body cues.
5.2. Vocalist Model 29 all regions tend to rise and fall together, suggesting shared influence from more abstract factors such as vocal activity or model confidence, rather than isolated spikes in importance for specific regions. This interpretation is supported by the moderate positive correlations observed between saliency dynamics and separation quality (using scale-invariant signal-to-distortion ratio, SI-SDR), with values ranging from r≈0.38 for body regions to r≈0.46 for face-dominated aggregates. These results indicate that the temporal saliency modulation is not arbitrary but reflects moments where visual cues contribute more strongly to effective separation. Figure 8: Temporal evolution of saliency scores with strong window overlay for the performance of Abhiram Bode - Entha Bhagyamu using VoViT-fb To bridge the model-independent exploration in Chapter 4 with our current attribution analysis, we plotted the temporal evolution of aggregated face and body saliency scores, overlaying the timestamps of strong correlation windows (|ρ|>0.66 over 0.5s) for both keypoint speed and acceleration (Figure 8). While acceleration events showed weak-to-moderate correlations with visual attention (r= 0.17–0.22, with face regions most responsive), speed correlations were negligible (r≈0). Acceleration aligned with local saliency variations in some cases, but the effect size was small. Vocal-specific cues showed slightly higher associations: jaw-to-nose motion reached r= 0.21–0.24 (p < 0.05) with about 6 events per 4s chunk on average, and mouth area motion r= 0.12–0.22 with 5.6 events per chunk. These differences suggest that articulatory gestures may offer more consistent temporal structure than
30 Chapter 5. Model Analysis and Interpretability general patterns, though the correlations remain modest overall. Dataset-Wide Analysis. While the case study illustrates local trends, it is unclear whether these generalize. To address this, we repeated the IG analysis across the entire dataset, restricted to performances with a vocalist, violinist, and mridangam player. For each 4-second segment, we computed saliency scores for all visual keypoints and aggregated them by body region. Rather than computing IG for every 4-second segment, we focused on two contrasting subsets: •The top 10% of chunks with the highest number of strongly correlated frames between vocal-specific motion features and audio per performance. •The bottom 10% of chunks with the lowest number of strong correlations. By comparing these subsets, we test whether the model’s attributions are more meaningful when audiovisual synchrony is high. Figure 9: Averaged IG saliency scores for different body and face parts. Figure 9 shows the distribution of average saliency across grouped facial and body regions for these two subsets. The results reveal that the proportional distribution of keypoint saliency remains almost identical4across subsets, closely resembling the pattern seen in the case study. However, the absolute scale of saliency differs 4The L2 distance for the normalized face and body distributions is 0.03 and 0.05 respectively
5.2. Vocalist Model 31 substantially: face regions exhibit saliency values an order of magnitude higher than body regions, and within each group, the top 10% subset shows attributions 2 to 3 times stronger than the bottom 10%. This indicates that while body keypoints are used minimally, the model relies heavily on facial features, especially the mouth and nose, for its predictions. This effect was not limited to visual inputs. Audio saliency also followed the same trend, with average values of 0.0467 for the top subset and 0.0168 for the bottom subset. This suggests that in low-correlation windows the model is overall less input-dependent, potentially relying more on learned priors or exhibiting saturated behavior. Importantly, this dataset-wide result mirrors the case study (Figure 8), where peaks and valleys of face and body saliency coincided with frames of high and low audiovisual correlation. Taken together, these findings indicate that audiovisual synchrony acts primarily as a scaling factor of saliency: it modulates the overall sensitivity of the model to its inputs without altering the relative spatial distribution of attributions across keypoints. While the face + body analysis revealed strong reliance on facial landmarks, it remained unclear whether body input dilutes or complements the facial contribution. To understand whether the body input dilutes or complements the face contribution, we extended the IG analysis to the face-only and body-only models. Comparison with face-only and body-only models. The same integrated gradients analysis was applied to the face-only and body-only models under identical conditions to those used for the face+body model. Figure 10 shows the distributions of average saliency across facial keypoints for the face+body and face-only models. While the overall magnitude of face-only saliency is roughly half that of the face+body model, the relative distribution shifts: the mouth emerges as the most salient region, followed by the nose, whereas in the face+body model the nose exhibits the highest attribution. This suggests that when constrained to facial input alone, the model emphasizes regions most directly linked to vocalization, particularly the mouth, whereas the face+body model spreads attribution more diffusely,
32 Chapter 5. Model Analysis and Interpretability including toward body keypoints that contribute relatively little (see Figure 9). This difference may help explain why the face-only model achieves stronger separation performance. Figure 10: Averaged IG saliency across facial keypoints for face+body and face-only. Audio saliency values further support this interpretation. While the face+body and face-only models exhibit nearly identical reliance on audio input (1.16 vs. 1.17 average saliency in the top 10%), the body-only model shows substantially higher audio saliency (1.54). This indicates that when deprived of informative facial cues, the model compensates by over-relying on the audio stream, confirming that body landmarks alone provide limited useful information for vocal separation. For completeness, we also examined the saliency distribution of the body-only model (Figure 18 in Appendix .3). In this case, the head still emerges as the most salient region, despite the absence of facial landmarks. This suggests that even with coarse body keypoints, the model attempts to exploit head motion as a weak proxy for vocal activity, though these signals are insufficient to support effective separation. 5.2.3 Ablation Studies To test causality rather than correlation, we performed a set of inference-time ablations that directly intervene on the visual inputs and measure the resulting change in separation quality. All ablations were executed on the vocalists on the face+body VoViT variant used in the gradient-based analysis. We report changes in SI-SDR
5.2. Vocalist Model 33 computed between the model’s estimate and the ground-truth target. Two complementary studies were implemented: Temporal shuffle. The temporal shuffle script randomly permutes the time dimension of the face and body keypoint tensors within each 4-s chunk while leaving the audio unchanged. Concretely, for each sample the tensors shaped (B, T, C, K) are permuted along T, then forwarded through the frozen model to produce a separated estimate. For every (artist,song) we record the baseline SI-SDR and the shuffled SI-SDR, and compute ∆SI-SDR =SI-SDRbaseline −SI-SDRshuffled as the persong effect. This intervention tests whether the model requires short-term temporal alignment between gestures and sound. Regional masking. The region masking script zeroes or replaces specific subsets of landmarks before forwarding the sample. We implemented three masks: (i) mouth, (ii) nose, and (iii) full body. Masks are filled with zeros for the body and with the respective dataset mean face keypoints for the face regions. For each masked condition we compute per-song SI-SDR and the corresponding ∆SI-SDR relative to the unmasked baseline. This directly probes which spatial regions are necessary for the model’s performance. Results. The temporal shuffle produced essentially no effect on separation: across the dataset the average SI-SDR changed by ∆SI-SDR ≈ −0.03 dB (baseline ≈12.32 dB, shuffled ≈12.35 dB), indicating that destroying framewise synchrony did not meaningfully degrade performance. By contrast, regional masking produced very large, region-specific degradations. Table 2 summarizes the face+body results Interpretation. The tiny effect of temporal shuffling suggests the model doesn’t actually need frame-by-frame synchrony between gestures and sound, at least not at the scale we tested it5. In contrast, masking the mouth or nose causes a significant 5This intervention shuffled keypoints within 4-second chunks, which destroys fine-grained synchrony but still preserves slower-scale co-occurrence. It therefore rules out frame-level dependence but not longer-timescale audiovisual alignment.
40 Chapter 6. Discussion and Conclusion whether human-interpretable motion features (e.g., speed, acceleration) contained any relation to audio descriptors (onset envelope and RMS). The goal was not only to test for correlations but to better understand the kind of information present in the data before moving on to model-level analysis. Although overall correlations were weak, we observed that instrument-specific cues, such as mouth motion for voice and elbow angle for violin, aligned more consistently with audio activity than general motion. This finding highlighted the importance of selecting where to look for visual information depending on the instrument of interest. Building on this context, Chapter 5 turned to the interpretability of audiovisual separation models themselves. We began by analyzing the face, body, and face+body VoViT model variants and then extended the investigation to the fusion module variants, which added multihead attention and FiLM layers to better exploit crossmodal cues. Several complementary methods were employed to test the models: gradient-based saliency analysis to reveal spatial and temporal focus, case studies of audiovisual alignment, comparisons across different input variants, ablations to assess the causal impact of visual cues, and an in-depth analysis of the fusion strategies. Taken together, these three stages reflect a progression from data preparation, to exploratory data understanding, to model interpretability. This layered approach provided both a practical evaluation of audiovisual separation in Carnatic music and a methodological framework for disentangling how visual cues influence multimodal models. 6.2 Conclusions The results of this work clarify how visual information contributes to music source separation in complex, real-world recordings. Dataset exploration showed that instrument-specific motion features, such as mouth movements for vocals, align more closely with audio activity than general motion, highlighting where meaningful visual cues reside.
6.3. Future Work 41 Interpretability analyses of VoViT models revealed that facial landmarks dominate visual contributions for vocal separation, while body motion is largely peripheral. Ablations confirmed the causal relevance of these visual features. For higher-performing vocal and violin models, feature fusion through multihead attention and FiLM layers enabled selective cross-modal modulation, showing that the models can amplify or suppress audio features based on visual context. Overall, visual cues can meaningfully support source separation, but their impact is instrumentand region-specific. The findings underline the importance of targeted visual feature selection and carefully designed fusion mechanisms in multimodal music separation systems. 6.3 Future Work The insights and limitations of this thesis open up several promising paths for future research. First, the analytical framework developed here could be extended to other musical traditions and instruments, such as hindustani music, to evaluate whether similar visual dominance patterns and cross-modal modulation principles hold. Insights from our interpretability analyses suggest that models could be made more efficient by prioritizing the most informative visual features, such as facial landmarks, and by designing FiLM-like modulatory mechanisms tailored to specific separation tasks. Another promising direction is connecting keypoint-based analyses with pixel-level saliency methods, which could reveal whether models trained on raw video naturally attend to the same informative regions. Finally, the understanding of audiovisual relationships gained here could inform generative and cross-modal synthesis, enabling models to predict or synthesize performer gestures from audio, or vice versa, in a musically meaningful way. In summary, this thesis has provided an integrated investigation into how visual information interacts with audio in music source separation, focusing on the challenging case of Carnatic performance. By combining dataset-driven analysis, crossdomain evaluation, and interpretability methods, it has offered a clearer picture of
42 Chapter 6. Discussion and Conclusion the selective but meaningful role of visual cues. While the findings are necessarily bounded by data and model constraints, they demonstrate the value of a multimodal perspective and lay the groundwork for more robust and culturally inclusive approaches to audiovisual music processing.
Bibliography [1] Chery, C. Some experiments on the recognition of speech with one. Journal of the Acoustical Society of America (1953). [2] Lee, D. D. & Seung, H. S. Learning the parts of objects by non-negative matrix factorization. Nature 401, 788–791 (1999). URL https://www.nature.com/ articles/44565. [3] Smaragdis, P. & Brown, J. Non-negative matrix factorization for polyphonic music transcription. In 2003 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (IEEE Cat. No.03TH8684), 177–180 (IEEE, New Paltz, NY, USA, 2003). URL http://ieeexplore.ieee.org/document/ 1285860/. [4] Ronneberger, O., Fischer, P. & Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation (2015). URL http://arxiv.org/abs/1505. 04597. ArXiv:1505.04597 [cs]. [5] Jansson, A. et al. SINGING VOICE SEPARATION WITH DEEP U-NET CONVOLUTIONAL NETWORKS (2017). [6] Stoller, D., Ewert, S. & Dixon, S. Wave-U-Net: A Multi-Scale Neural Network for End-to-End Audio Source Separation (2018). URL http://arxiv.org/ abs/1806.03185. ArXiv:1806.03185 [cs]. [7] Hennequin, R., Khlif, A., Voituret, F. & Moussallam, M. Spleeter: a fast and efficient music source separation tool with pre-trained models. Journal of Open 43
44 BIBLIOGRAPHY Source Software 5, 2154 (2020). URL https://joss.theoj.org/papers/10. 21105/joss.02154. [8] Défossez, A., Usunier, N., Bottou, L. & Bach, F. Music Source Separation in the Waveform Domain (2021). URL http://arxiv.org/abs/1911.13254. ArXiv:1911.13254 [cs]. [9] Luo, Y. & Mesgarani, N. Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 27, 1256–1266 (2019). URL http: //arxiv.org/abs/1809.07454. ArXiv:1809.07454 [cs]. [10] Mitsufuji, Y. et al. Music Demixing Challenge 2021. Frontiers in Signal Processing 1, 808395 (2022). URL https://www.frontiersin.org/articles/ 10.3389/frsip.2021.808395/full. [11] Seetharaman, P., Wichern, G., Venkataramani, S. & Roux, J. L. Classconditional embeddings for music source separation (2018). URL http:// arxiv.org/abs/1811.03076. ArXiv:1811.03076 [cs]. [12] Stöter, F.-R., Uhlich, S., Liutkus, A. & Mitsufuji, Y. Open-Unmix - A Reference Implementation for Music Source Separation. Journal of Open Source Software 4, 1667 (2019). URL https://joss.theoj.org/papers/10.21105/ joss.01667. [13] Wang, Y., Stoller, D., Bittner, R. M. & Pablo Bello, J. Few-Shot Musical Source Separation. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 121–125 (IEEE, Singapore, Singapore, 2022). URL https://ieeexplore.ieee.org/document/9747536/. [14] Tong, W. et al. SCNet: Sparse Compression Network for Music Source Separation. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1276–1280 (IEEE, Seoul, Korea, Republic of, 2024). URL https://ieeexplore.ieee.org/document/10446651/.
BIBLIOGRAPHY 45 [15] Sodoyer, D., Schwartz, J.-L., Girin, L., Klinkisch, J. & Jutten, C. Separation of Audio-Visual Speech Sources: A New Approach Exploiting the AudioVisual Coherence of Speech Stimuli. EURASIP Journal on Advances in Signal Processing 2002, 382823 (2002). URL https://asp-eurasipjournals. springeropen.com/articles/10.1155/S1110865702207015. [16] Lu, R., Duan, Z. & Zhang, C. Listen and Look: Audio–Visual Matching Assisted Speech Source Separation. IEEE Signal Processing Letters 25, 1315–1319 (2018). URL https://ieeexplore.ieee.org/document/8404105/. [17] Zhao, H. et al. The Sound of Pixels. In Ferrari, V., Hebert, M., Sminchisescu, C. & Weiss, Y. (eds.) Computer Vision – ECCV 2018, vol. 11205, 587– 604 (Springer International Publishing, Cham, 2018). URL https://link. springer.com/10.1007/978-3-030-01246-5_35. Series Title: Lecture Notes in Computer Science. [18] Gao, R. & Grauman, K. Co-Separating Sounds of Visual Objects. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 3878–3887 (IEEE, Seoul, Korea (South), 2019). URL https://ieeexplore.ieee.org/ document/9009045/. [19] Zhu, L. & Rahtu, E. Visually Guided Sound Source Separation Using Cascaded Opponent Filter Network. In Ishikawa, H., Liu, C.-L., Pajdla, T. & Shi, J. (eds.) Computer Vision – ACCV 2020, vol. 12627, 409–426 (Springer International Publishing, Cham, 2021). URL http://link.springer.com/10.1007/ 978-3-030-69544-6_25. Series Title: Lecture Notes in Computer Science. [20] Gan, C., Huang, D., Zhao, H., Tenenbaum, J. B. & Torralba, A. Music Gesture for Visual Sound Separation. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10475–10484 (IEEE, Seattle, WA, USA, 2020). URL https://ieeexplore.ieee.org/document/9157677/. [21] Tan, R. et al. Language-Guided Audio-Visual Source Separation via Trimodal Consistency. In 2023 IEEE/CVF Conference on Computer Vision and Pat-
46 BIBLIOGRAPHY tern Recognition (CVPR), 10575–10584 (IEEE, Vancouver, BC, Canada, 2023). URL https://ieeexplore.ieee.org/document/10203040/. [22] Chen, J. et al. iQuery: Instruments as Queries for Audio-Visual Sound Separation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14675–14686 (IEEE, Vancouver, BC, Canada, 2023). URL https://ieeexplore.ieee.org/document/10205441/. [23] Chatterjee, M., Le Roux, J., Ahuja, N. & Cherian, A. Visual Scene Graphs for Audio Source Separation. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 1184–1193 (IEEE, Montreal, QC, Canada, 2021). URL https://ieeexplore.ieee.org/document/9710769/. [24] Montesinos, J. F., Kadandale, V. S. & Haro, G. A cappella: Audio-visual Singing Voice Separation (2021). URL http://arxiv.org/abs/2104.09946. ArXiv:2104.09946 [cs]. [25] Montesinos, J. F., Kadandale, V. S. & Haro, G. VoViT: Low Latency GraphBased Audio-Visual Voice Separation Transformer. In Avidan, S., Brostow, G., Cissé, M., Farinella, G. M. & Hassner, T. (eds.) Computer Vision – ECCV 2022, vol. 13697, 310–326 (Springer Nature Switzerland, Cham, 2022). URL https://link.springer.com/10.1007/978-3-031-19836-6_18. Series Title: Lecture Notes in Computer Science. [26] Erhan, D., Bengio, Y., Courville, A., Vincent, P. & Box, P. O. Visualizing Higher-Layer Features of a Deep Network . [27] Zeiler, M. D. & Fergus, R. Visualizing and Understanding Convolutional Networks (2013). URL http://arxiv.org/abs/1311.2901. ArXiv:1311.2901 [cs]. [28] Baehrens, D. et al. How to Explain Individual Classification Decisions . [29] Simonyan, K., Vedaldi, A. & Zisserman, A. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps (2014). URL http://arxiv.org/abs/1312.6034. ArXiv:1312.6034 [cs].
BIBLIOGRAPHY 47 [30] Selvaraju, R. R. et al. Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization. International Journal of Computer Vision 128, 336–359 (2020). URL http://arxiv.org/abs/1610.02391. ArXiv:1610.02391 [cs]. [31] Meyes, R., Lu, M., Puiseau, C. W. d. & Meisen, T. Ablation Studies in Artificial Neural Networks (2019). URL http://arxiv.org/abs/1901.08644. ArXiv:1901.08644 [cs]. [32] Serra, X. A MULTICULTURAL APPROACH IN MUSIC INFORMATION RESEARCH. Oral Session (2011). [33] Gemmeke, J. F. et al. Audio Set: An ontology and human-labeled dataset for audio events. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 776–780 (IEEE, New Orleans, LA, 2017). URL http://ieeexplore.ieee.org/document/7952261/. [34] Li, B., Liu, X., Dinesh, K., Duan, Z. & Sharma, G. Creating a Multitrack Classical Music Performance Dataset for Multimodal Music Analysis: Challenges, Insights, and Applications. IEEE Transactions on Multimedia 21, 522–535 (2019). URL https://ieeexplore.ieee.org/document/8411155/. [35] Shankar, A., Plaja-Roglans, G., Nuttall, T., Rocamora, M. & Serra, X. Saraga audiovisual: a large multimodal open data collection for the analysis of carnatic music (2024). [36] Srinivasamurthy, A., Gulati, S., Caro Repetto, R. & Serra, X. Saraga: Open Datasets for Research on Indian Art Music. Empirical Musicology Review 16, 85–98 (2021). URL https://emusicology.org/index.php/EMR/article/ view/7641. [37] Krishnan, V. V., Alben, N., Nair, A. & Condit-Schultz, N. Sanidha: A Studio Quality Multi-Modal Dataset for Carnatic Music (2025). URL http://arxiv. org/abs/2501.06959. ArXiv:2501.06959 [cs].
48 BIBLIOGRAPHY [38] Sridhar, R. & Geetha, T. V. Music Information Retrieval of Carnatic Songs Based on Carnatic Music Singer Identification. In 2008 International Conference on Computer and Electrical Engineering, 407–411 (IEEE, Phuket, Thailand, 2008). URL http://ieeexplore.ieee.org/document/4741017/. [39] Koduri, G. K., Miron, M., Serra, J. & Serra, X. COMPUTATIONAL APPROACHES FOR THE UNDERSTANDING OF MELODY IN CARNATIC MUSIC . [40] Kirthika, P. & Chattamvelli, R. A review of raga based music classification and music information retrieval (MIR). In 2012 IEEE International Conference on Engineering Education: Innovative Practices and Future Trends (AICERA), 1–5 (IEEE, Kottayam, India, 2012). URL http://ieeexplore.ieee.org/ document/6306752/. [41] Plaja-Roglans, G., Miron, M., Shankar, A. & Serra, X. CARNATIC SINGING VOICE SEPARATION USING COLD DIFFUSION ON TRAINING DATA WITH BLEEDING (2023). [42] Plaja-Roglans, G., Nuttall, T., Pearson, L., Serra, X. & Miron, M. RepertoireSpecific Vocal Pitch Data Generation for Improved Melodic Analysis of Carnatic Music. Transactions of the International Society for Music Information Retrieval 6, 13–26 (2023). URL http://transactions.ismir.net/articles/ 10.5334/tismir.137/. [43] Shankar, A. Gesture-guided melodic source separation in carnatic music using cross-modal fusion techniques (2025). [44] Rodrigues, M., Haro, G., Rocamora, M. & Sivasankar, A. S. Pose Estimation For Audio-Visual Singing Voice Separation In Indian Carnatic Music . [45] Chandrasekaran, C., Trubanova, A., Stillittano, S., Caplier, A. & Ghazanfar, A. A. The Natural Statistics of Audiovisual Speech. PLoS Computational Biology 5, e1000436 (2009). URL https://dx.plos.org/10.1371/journal. pcbi.1000436.
BIBLIOGRAPHY 49 [46] Turner-Stokes, L. & Reid, K. Three-dimensional motion analysis of upper limb movement in the bowing arm of string-playing musicians. Clinical Biomechanics 14, 426–433 (1999). [47] Ancillao, A., Savastano, B., Galli, M. & Albertini, G. Three dimensional motion capture applied to violin playing: A study on feasibility and characterization of the motor strategy. Computer Methods and Programs in Biomedicine 149, 19– 27 (2017). URL https://www.sciencedirect.com/science/article/pii/ S0169260716312342. [48] Ephrat, A. et al. Looking to Listen at the Cocktail Party: A SpeakerIndependent Audio-Visual Model for Speech Separation. ACM Transactions on Graphics 37, 1–11 (2018). URL http://arxiv.org/abs/1804.03619. ArXiv:1804.03619 [cs]. [49] Yan, S., Xiong, Y. & Lin, D. Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. Proceedings of the AAAI Conference on Artificial Intelligence 32 (2018). URL https://ojs.aaai.org/index.php/ AAAI/article/view/12328. Publisher: Association for the Advancement of Artificial Intelligence (AAAI). [50] Shrikumar, A., Greenside, P., Shcherbina, A. & Kundaje, A. Not Just a Black Box: Learning Important Features Through Propagating Activation Differences (2017). URL http://arxiv.org/abs/1605.01713. ArXiv:1605.01713 [cs]. [51] Sundararajan, M., Taly, A. & Yan, Q. Axiomatic Attribution for Deep Networks (2017). URL http://arxiv.org/abs/1703.01365. ArXiv:1703.01365 [cs].