scieee AI-readable full text Open interactive document viewer

A bimodal deep model to capture emotions from music tracks

Tobolewski, Jan,Sakowicz, Michal,Turmo Borras, Jorge,Kostek, Bozena

Abstract

This work aims to develop a deep model for automatically labeling music tracks in terms of induced emotions. The machine learning architecture consists of two components: one dedicated to lyric processing based on Natural Language Processing (NLP) and another devoted to music processing. These two components are combined at the decision-making level. To achieve this, a range of neural networks are explored for the task of emotion extraction from both lyrics and music. For lyric classification, three architectures are compared, i.e., a 4-layer neural network, FastText, and a transformer-based approach. For music classification, the architectures investigated include InceptionV3, a collection of models from the ResNet family, and a joint architecture combining Inception and ResNet. SVM serves as a baseline in both threads. The study explores three datasets of songs accompanied by lyrics, with MoodyLyrics4Q selected and preprocessed for model training. The bimodal approach, incorporating both lyrics and audio modules, achieves a classification accuracy of 60.7% in identifying emotions evoked by music pieces. The MoodyLyrics4Q dataset used in this study encompasses musical pieces spanning diverse genres, including rock, jazz, electronic, pop, blues, and country. The algorithms demonstrate reliable performance across the dataset, highlighting their robustness in handling a wide variety of musical styles.

Full text

JAISCR, 2025, Vol. 15, No. 3, pp. 215 A BIMODAL DEEP MODEL TO CAPTURE EMOTIONS FROM MUSIC TRACKS Jan Tobolewski1, Michał Sakowicz1, Jordi Turmo2, and Bo˙ zena Kostek3,∗ 1Faculty of Electronics, Telecommunications and Informatics, Gda´ nsk University of Technology, 11/12 Narutowicza St., 80-233 Gda´ nsk, Poland 2Department of Computer Science, Universitat Polit` ecnica de Catalunya, Jordi Girona Salgado, 1-3, 08034 Barcelona, Spain 3Audio Acoustics Laboratory, Faculty of Electronics, Telecommunications, and Informatics, Gda´ nsk University of Technology, 11/12 Narutowicza St., 80-233 Gda´ nsk, Poland ∗E-mail: [email protected] Submitted: 28th November 2024; Accepted: 21st February 2025 Abstract This work aims to develop a deep model for automatically labeling music tracks in terms of induced emotions. The machine learning architecture consists of two components: one dedicated to lyric processing based on Natural Language Processing (NLP) and another devoted to music processing. These two components are combined at the decision-making level. To achieve this, a range of neural networks are explored for the task of emotion extraction from both lyrics and music. For lyric classification, three architectures are compared, i.e., a 4-layer neural network, FastText, and a transformer-based approach. For music classification, the architectures investigated include InceptionV3, a collection of models from the ResNet family, and a joint architecture combining Inception and ResNet. SVM serves as a baseline in both threads. The study explores three datasets of songs accompanied by lyrics, with MoodyLyrics4Q selected and preprocessed for model training. The bimodal approach, incorporating both lyrics and audio modules, achieves a classification accuracy of 60.7% in identifying emotions evoked by music pieces. The MoodyLyrics4Q dataset used in this study encompasses musical pieces spanning diverse genres, including rock, jazz, electronic, pop, blues, and country. The algorithms demonstrate reliable performance across the dataset, highlighting their robustness in handling a wide variety of musical styles. Keywords: automatic labeling, deep model, emotion, music, lyrics, machine learning 1 Introduction Emotion recognition is a process that uses technology to identify human emotions from various sources, such as facial expressions, speech, text, or physiological signals [1]. Music profoundly influences human emotions, enhancing emotional experiences and strongly evoking different feelings [2, 3]. Research has shown that music-evoked pleasure increases dopamine production, whereas dissonant music often leads to a rise in blood oxygen levels [4]. Understanding the emotional impact of music has significant applications in areas such as music recommendation, intelligent music players, therapy [5], and affective computing. The emotion 10.2478/jaiscr-2025-0011 – 238 216 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek recognition system may be, at the same time, capable of identifying human emotional states. However, it should be pointed out that music does not always change the emotions of the person who listens to it. How well it works and how the person feels depends on many factors. One concern is how long the person listens to the music. Song et al. [6] used the Kruskal-Wallis test and found that people took more time to feel sadness or relaxation than happiness or anger when listening to music. The authors also looked at how people’s mood matched the emotions that the music was supposed to induce. They found no significant differences in how fast people felt emotions, but they found a positive relationship between them. This means that people often feel the same emotions as those inscribed in the music. Hence, based on emotion recognition in music, a recommender system may suggest similar songs to the ones the user listens to or ask the users what mood they are in to present them with a suitable song. Both the lyrics and the melody cause the emotions induced by listening to a music track [7]. Moreover, lyrical and audio features provide distinct and complementary information about the emotional content of music. Although audio captures music’s tonal and expressive elements, lyrics contribute to semantic and contextual understanding. Combining these modalities ensures that the system can still produce meaningful predictions when one modality is less reliable [8]. This may happen, for example, for complex instrumental music with a tiny dose of lyrics or rap with a monotonous melodic line overflowing with content in the lyrical layer. Hence, automatic labeling is a process in which a system autonomously tags a music track based on audio content, such as music characteristics (music genre, rhythm, or emotions) or semantic annotations, either manual, collected from social tagging services, or Web content mining [9]. Although previous studies have investigated emotion recognition from lyrics or audio independently, a comprehensive approach integrating both modalities remains underexplored. Hence, combining multiple modalities—such as audio, text, video, and image—to capture the complexity and diversity of human emotions could enhance emotion recognition [10, 11]. Therefore, the motivation behind our work is to first conduct a unimodal study on lyrics and music emotion classification, and then compare it with a bimodal approach. Moreover, we would like to check both traditional and deeplearning approaches that automatically label music tracks based on the emotions they induce. To that end, we classify lyrics into four classes of emotion (happy, angry, relaxed, and sad), and these classes can be directly inferred from the four quadrants of Russell’s Valence-Arousal (each quadrant corresponds to one of our four classes). 1.1 Related work Algorithms have become increasingly involved in music in recent years [12, 13, 14]. They are used for tasks such as recommending music, composing songs, improving sound quality, removing noise, and adjusting song tempo. Many of these algorithms rely on artificial intelligence (AI), particularly machine learning (ML), which enables them to learn from data and experience, identify patterns, and make predictions. Some systems can automatically label music, but they still require further improvement to enhance accuracy, robustness, and reliability. Additionally, a user-data-driven approach is essential, especially for personalization. This involves adapting an application to make users feel it is tailored specifically for them, as explained by Barata and Coelho [15]. Personalization depends on collecting user data and analyzing their interactions with the platform. It also leverages big data analytics to predict user preferences [16]. This information is gathered as users listen to music, allowing the recommendation model to continuously refine itself based on their music preferences. As a result, personalized music playlists can be created to match a user’s daily activities and mood. Various methods and techniques can be used to analyze and interpret emotions, such as signal processing, computer vision, and natural language processing. Feature-based approaches, such as those of Schmidt et al. (2010) [17], have demonstrated the utility of hand-crafted acoustic features for audio classification. Bimodal approaches, as shown in the work of Yang et al. (2008) [18], further improve emotion classification by integrating audio and lyrical features for more accurate predictions. Deep learning methods, including CNN architectures for audio [19, 20, 21, 22], word embedding 217 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek recognition system may be, at the same time, capable of identifying human emotional states. However, it should be pointed out that music does not always change the emotions of the person who listens to it. How well it works and how the person feels depends on many factors. One concern is how long the person listens to the music. Song et al. [6] used the Kruskal-Wallis test and found that people took more time to feel sadness or relaxation than happiness or anger when listening to music. The authors also looked at how people’s mood matched the emotions that the music was supposed to induce. They found no significant differences in how fast people felt emotions, but they found a positive relationship between them. This means that people often feel the same emotions as those inscribed in the music. Hence, based on emotion recognition in music, a recommender system may suggest similar songs to the ones the user listens to or ask the users what mood they are in to present them with a suitable song. Both the lyrics and the melody cause the emotions induced by listening to a music track [7]. Moreover, lyrical and audio features provide distinct and complementary information about the emotional content of music. Although audio captures music’s tonal and expressive elements, lyrics contribute to semantic and contextual understanding. Combining these modalities ensures that the system can still produce meaningful predictions when one modality is less reliable [8]. This may happen, for example, for complex instrumental music with a tiny dose of lyrics or rap with a monotonous melodic line overflowing with content in the lyrical layer. Hence, automatic labeling is a process in which a system autonomously tags a music track based on audio content, such as music characteristics (music genre, rhythm, or emotions) or semantic annotations, either manual, collected from social tagging services, or Web content mining [9]. Although previous studies have investigated emotion recognition from lyrics or audio independently, a comprehensive approach integrating both modalities remains underexplored. Hence, combining multiple modalities—such as audio, text, video, and image—to capture the complexity and diversity of human emotions could enhance emotion recognition [10, 11]. Therefore, the motivation behind our work is to first conduct a unimodal study on lyrics and music emotion classification, and then compare it with a bimodal approach. Moreover, we would like to check both traditional and deeplearning approaches that automatically label music tracks based on the emotions they induce. To that end, we classify lyrics into four classes of emotion (happy, angry, relaxed, and sad), and these classes can be directly inferred from the four quadrants of Russell’s Valence-Arousal (each quadrant corresponds to one of our four classes). 1.1 Related work Algorithms have become increasingly involved in music in recent years [12, 13, 14]. They are used for tasks such as recommending music, composing songs, improving sound quality, removing noise, and adjusting song tempo. Many of these algorithms rely on artificial intelligence (AI), particularly machine learning (ML), which enables them to learn from data and experience, identify patterns, and make predictions. Some systems can automatically label music, but they still require further improvement to enhance accuracy, robustness, and reliability. Additionally, a user-data-driven approach is essential, especially for personalization. This involves adapting an application to make users feel it is tailored specifically for them, as explained by Barata and Coelho [15]. Personalization depends on collecting user data and analyzing their interactions with the platform. It also leverages big data analytics to predict user preferences [16]. This information is gathered as users listen to music, allowing the recommendation model to continuously refine itself based on their music preferences. As a result, personalized music playlists can be created to match a user’s daily activities and mood. Various methods and techniques can be used to analyze and interpret emotions, such as signal processing, computer vision, and natural language processing. Feature-based approaches, such as those of Schmidt et al. (2010) [17], have demonstrated the utility of hand-crafted acoustic features for audio classification. Bimodal approaches, as shown in the work of Yang et al. (2008) [18], further improve emotion classification by integrating audio and lyrical features for more accurate predictions. Deep learning methods, including CNN architectures for audio [19, 20, 21, 22], word embedding A BIMODAL DEEP MODEL TO . . . representations for lyrics [23], and Transformerbased attention mechanisms [24], have been shown to enhance classification accuracy through automatic feature extraction. Agrawal et al. outlined two approaches: the first being a multitask approach incorporating both emotion and Russell’s ValenceArousal dimensions, and the second a single-task approach focused solely on emotion classification [24]. Han et al. (2022) [25] highlight that music emotion recognition systems have evolved from traditional hand-crafted feature-based machine learning to deep learning models that learn feature representations automatically. In addition, previous studies have explored multimodal learning paradigms that integrate audio and lyrical input to improve emotion recognition [26, 27]. Delbouys et al. (2018) [27] highlight the advantages of deep learning over classical models in music mood detection. Their work compares traditional feature-based approaches with end-to-end deep learning models, showing that deep models significantly outperform classical methods in arousal detection, while valence prediction benefits from multimodal midlevel fusion. 1.2 Study outline Comparison of the current standard, i.e., deep learning (DM) models with classical algorithms, such as support vector machines (SVM), in experiments on music, lyrics, and emotion analysis, may be insightful. SVMs in such studies were often used as baselines as they are well understood and have a long history of robust performance on structured and moderately complex data. Another important point is that this may also help to quantify how much improvement (if any) deep models provide and whether the added complexity of deep learning is justified. Moreover, SVM does not require massive computational resources, but it still provides insights into the importance of features and decision boundaries. This is an easily interpretable benchmark against which the performance of deep models can be measured. In the case of a moderately sized dataset, the comparison can reveal whether the deep models leverage their capacity or overfit. Also, SVM and other baseline algorithms rely on pre-processed features, while deep learning models learn features automatically. Comparison of performances of, for example, SVM and DM can highlight whether extracted features add value in the context of music, lyrics, and emotion evaluation. That is why our investigations start with the baseline algorithms for unimodal emotion classification and then create a bimodal system for labeling music tracks using state-of-the-art deep neural networks. Using advanced neural architectures for both lyric and audio processing, the study attempts to improve the accuracy of emotion classification. Furthermore, the use of diverse musical genres ensures the robustness of the proposed approach, making it adaptable to a wide range of musical styles. The architecture of the applied bimodal neural network model consists of a natural language module to process lyrics and an audio module for processing the soundtrack of music excerpts (see Figure 1). Each module allows for the classifier model training from suitable preprocessed training corpus (i.e., lyric corpus and audio corpus for the natural language module and the audio module, respectively), as well as using these learned models to classify new lyrics and audio. Both modules are connected at the decision stage. To evaluate and compare all the approaches, we explored three databases with respect to the stateof-the-art: Music4All [28], MoodyLyrics [29], and MoodyLyrics4Q [30]. Music4All consists of over 15 thousand songs represented by their lyrics, 30second excerpts of their audio clips, and other attributes (artist, genre, popularity, tags, etc.). The tracks were collected from many different sources and their tags were not uniform. The Music4All database offers a continuous emotion representation according to Thayer’s model [31]. Such an approach allows mapping Arousal-Valence quadrants into four emotions, but the boundaries are not fixed: for one emotion, the boundaries may be in line with the Arousal-Valence axes, and for another one, they may be slightly shifted. An alternative approach is the discrete representation of emotions, which is more natural for human evaluation. MoodyLyrics [29] contains 2,595 labeled songs along with their track titles and artists. The creators of this dataset annotated each song with one of the four emotion classes of Russel’s model [32]: happy, angry, relaxed, and sad. MoodyLyrics4Q [30] is a dataset containing 2000 songs from the funk, rock, jazz, electronic, pop, blues, and country music genres. Each song is represented 218 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek Figure 1. Experimental design by the artist’s name, the song’s title, and the emotion class assigned by last.fm [33] users (happy, angry, relaxed, or sad). The whole set of songs is distributed into 500 pieces for each emotion class. However, the dataset does not contain either the lyrics or the audio of the songs, so we had to get them from other services, as described in Section (2.1). When comparing MoodyLyrics and MoodyLyrics4Q, the latter has fewer repeated songs and more accurate assignments of emotional labels to songs [28, 30]. That is why we decided to employ MoodyLyrics4Q [29] in our study. The paper is organized as follows. Section 2 shows how MoodyLyrics4Q is supplemented with lyrics retrieved from Genius [34] and audio extracted from YouTube [35]. Section 3 encompasses unimodal based on lyrics and audio, as well as joint modalities, i.e., bimodal emotion classification. It also contains lyrics and audio preprocessing. Section 3.2 summarizes the results achieved in a unimodal approach by each of the researched algorithms. This is followed by audio preprocessing and classification, to which SVM, Inception V3, ResNet family, VGGNet, and Inception-ResNet are employed, presented along with the results obtained (see Sections 3.3 and 3.4). In Section 3, bimodal emotion classification is also examined based on majority voting ensemble methodology. The results achieved show that combining two modalities improves, to some extent, the metrics of the final model compared to separately operating architectures. This indicates the effectiveness of the bimodal approach and suggests that such an approach may be applied in other systems related to emotion classification based on music tracks. The paper discusses the results from the performed experiments and compares them with similar approaches from the literature (see Section 4). The overall conclusion, along with limitations, is also given. Also, possible ways to enhance the system’s performance are outlined. Finally, our research contributes significantly by identifying and addressing errors and inconsistencies within the MoodyLyrics and MoodyLyrics4Q datasets. Specifically, we highlight these issues and propose preprocessing steps, such as removing duplicates and instrumental songs, while focusing solely on the English dataset for lyric classification. As a result, our final comparison focuses solely on the English dataset. Furthermore, our findings demonstrate that the majority voting ensemble effectively classifies three out of four emotions—happy, angry, and relaxed—while the concatenation ensemble demonstrates greater reliability across all emotional categories. All investigations presented in this paper were carried out on a computational cluster provided by the Gda´ nsk University of Technology (GUT). This device has a 10-core CPU Intel Xeon Silver 4210, clocked at 2.20 GHz and 378 GB of RAM, supported by NVIDIA A40 with 46 GB memory, which enables utilizing GPUs for computation [36]. 2 Data Preparation In this section, the process involved in preparing the data used for the experiments was described. This included exploring the dataset and splitting it into appropriate subsets to ensure balanced and effective model training. As already mentioned 219 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek Figure 1. Experimental design by the artist’s name, the song’s title, and the emotion class assigned by last.fm [33] users (happy, angry, relaxed, or sad). The whole set of songs is distributed into 500 pieces for each emotion class. However, the dataset does not contain either the lyrics or the audio of the songs, so we had to get them from other services, as described in Section (2.1). When comparing MoodyLyrics and MoodyLyrics4Q, the latter has fewer repeated songs and more accurate assignments of emotional labels to songs [28, 30]. That is why we decided to employ MoodyLyrics4Q [29] in our study. The paper is organized as follows. Section 2 shows how MoodyLyrics4Q is supplemented with lyrics retrieved from Genius [34] and audio extracted from YouTube [35]. Section 3 encompasses unimodal based on lyrics and audio, as well as joint modalities, i.e., bimodal emotion classification. It also contains lyrics and audio preprocessing. Section 3.2 summarizes the results achieved in a unimodal approach by each of the researched algorithms. This is followed by audio preprocessing and classification, to which SVM, Inception V3, ResNet family, VGGNet, and Inception-ResNet are employed, presented along with the results obtained (see Sections 3.3 and 3.4). In Section 3, bimodal emotion classification is also examined based on majority voting ensemble methodology. The results achieved show that combining two modalities improves, to some extent, the metrics of the final model compared to separately operating architectures. This indicates the effectiveness of the bimodal approach and suggests that such an approach may be applied in other systems related to emotion classification based on music tracks. The paper discusses the results from the performed experiments and compares them with similar approaches from the literature (see Section 4). The overall conclusion, along with limitations, is also given. Also, possible ways to enhance the system’s performance are outlined. Finally, our research contributes significantly by identifying and addressing errors and inconsistencies within the MoodyLyrics and MoodyLyrics4Q datasets. Specifically, we highlight these issues and propose preprocessing steps, such as removing duplicates and instrumental songs, while focusing solely on the English dataset for lyric classification. As a result, our final comparison focuses solely on the English dataset. Furthermore, our findings demonstrate that the majority voting ensemble effectively classifies three out of four emotions—happy, angry, and relaxed—while the concatenation ensemble demonstrates greater reliability across all emotional categories. All investigations presented in this paper were carried out on a computational cluster provided by the Gda´ nsk University of Technology (GUT). This device has a 10-core CPU Intel Xeon Silver 4210, clocked at 2.20 GHz and 378 GB of RAM, supported by NVIDIA A40 with 46 GB memory, which enables utilizing GPUs for computation [36]. 2 Data Preparation In this section, the process involved in preparing the data used for the experiments was described. This included exploring the dataset and splitting it into appropriate subsets to ensure balanced and effective model training. As already mentioned A BIMODAL DEEP MODEL TO . . . and justified in Section 1, we employed MoodyLyrics4Q [29] in our study. 2.1 Dataset Creation The dataset was created in two phases: a retrieval phase and a filtering phase. The retrieval phase consisted of obtaining the lyrics and audio for each song included in the dataset. We retrieved lyrics from Genius [34] because of its partnership with Spotify [37], one of the largest music streaming services. We used the Genius API for this purpose. For audio retrieval, we used the YouTube platform [35] instead of the Spotify API. This decision was due to two primary reasons: the incomplete availability of all songs in the selected dataset on Spotify and the inability to select specific audio segments for analysis (e.g., beginning, middle, or end). A scraping script was implemented to download all the required tracks from YouTube. The Pafy library [38] was used in conjunction with the MoviePy module [39] to convert YouTube videos into MP3 format for subsequent analysis. Although it is not possible to download multiple high-quality audios belonging to specific albums and artists fully automatically due to variations in recording quality, such as live recordings from concerts or audience participation, in roughly 2% of the dataset, we searched for higher-quality songs and performed manual replacements as needed. The filtering phase consisted of removing redundant records from the dataset. The duplicates were removed (they are not listed here due to the space limit; however, this information is available from the authors on request). After removing duplicates, the database contained 1990 songs. On the other hand, it was found that most of the resulting songs are in English, concretely, 1883 songs, while 102 are distributed in other 21 languages, and 5 are instrumental songs. Due to the fact that the dataset is strongly biased towards English, we decided to consider songs in other languages and instrumental songs as outliers (around 5% of the MoodyLyric4Q records). In total, the final lyrical dataset (English lyric dataset) used for our experiments contains 1883 lyrics. As shown in Table 1, the number of songs assigned to each class indicates that the four classes are well-balanced. Table 1. Emotion assignment to classes in the English lyric dataset. Emotion Label Happy Angry Sad Relaxed Number of Songs 471 491 462 459 2.2 Data Partition The datasets were partitioned into three subsets: training, validation, and test. The training set employed for the teaching model comprised 70% of the total data. The validation set, used to validate the model and hyperparameter tuning, accounted for 15%, while the test set, assessing the final model’s generalization, consisted of the remaining 15%. Table 2 provides a distribution of the songs among the classes across the data partitions. It indicates an even distribution of learning examples within each subset, ensuring equitable representation for model training. Table 2. Distribution of the classes in the dataset partitions. Emotion Label Happy Angry Sad Relaxed Number of Training Songs 350 348 348 347 Number of Validation Songs 75 75 75 74 Number of Test Songs 75 74 74 75 Table 3. Distribution of the classes in the English lyric dataset partitions. Emotion Label Happy Angry Sad Relaxed Number of Training Songs 331 344 327 325 Number of Validation Songs 68 73 65 67 Number of Test Songs 72 74 70 67 For the English lyric dataset, we used the same division method to eventually combine the audio and lyric elements, allowing us to create the final dataset. Table 3 shows the distribution of classes 220 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek for each partition of the English lyric dataset. Despite the minimal observed discrepancy in the number of examples in the subsets, this partitioning was considered balanced. 3 Methods This Section outlines the development of a bimodal deep-learning architecture for classifying emotions in music tracks. The proposed approach integrates both lyrical and audio processing to capture the emotional nuances embedded within musical pieces. To evaluate classification performance, we employed multiple metrics, including Accuracy, Precision, Recall, Support, and the F1 score. 3.1 Lyrics Preprocessing Several text-cleaning operations were performed for each lyric: the song title was removed from the lyrics, as well as punctuation marks, tagging metadata, and extra newline control characters. After that, part-of-speech (PoS) tagging was performed in order to achieve morpho-syntactic features for some of the lyrics classification methods. For such a process, SpaCy tagger, a Python NLP toolkit [40], was used. Table 4 describes some statistics related to the number of tokens of the lyrics from the training set. In addition, Table 5 shows the distribution of PoS tags occurring in the lyrics (absolute frequencies and percentages per PoS tag in parentheses), pairing attention to those PoS tags that can express feelings (nouns, adjectives, verbs, adverbs, and interjections). As seen from both tables, on average, a lyric is around 300 tokens, and the number of words potentially expressing feelings in the dataset is high. Table 4. Statistics of the length in tokens of a lyric of the English lyric dataset. Happy Angry Sad Relaxed Mean 375.96 342.63 228.97 230.21 Median 340.0 273.0 206.0 201.0 Maximum 1,432 1,321 1,199 872 Minimum 11 56 12 10 Standard Deviation 198.84 224.76 134.74 135.04 From Table 5, some trends can be discerned a priori: a) interjections occur significantly more times in lyrics evoking ’Angry’ (45.8%) than in lyrics of the rest of the categories, b) nouns and verbs appear significantly more times in lyrics evoking ’Angry’ (30.2% and 31.9%, respectively) and ’Sad’ (32.9% and 29.8%, respectively) than in the rest of lyrics. Table 5. Statistics of the PoS tags occurring in the English lyric dataset. PoS tag Happy Angry Sad Relaxed adjective 5,122 (20.1%) 7,888 (31.0%) 7,568 (29.8%) 4,858 (19.1%) adverb 6,470 (23.2%) 8,281 (29.7%) 7,041 (25.2%) 6,131 (22.0%) interjection 1,231 (11.9%) 4,729 (45.8%) 2,350 (22.8%) 2,007 (19.5%) noun 14,846 (18.6%) 24,084 (30.2%) 26,287 (32.9%) 14,585 (18.3%) verb 17,229 (19.3%) 28,484 (31.9%) 26,632 (29.8%) 16,914 (18.9%) others 52,873 (20.0%) 82,461 (31.1%) 77,621 (29.3%) 51,955 (19.6%) Knowing the distribution of the PoS tags throughout the dataset may allow checking whether their distribution corresponds to expressing feelings in poetry [41]. However, it is important to note that understanding natural language alone does not account for the emotional impact of poetry. As Johnson-Laird and Oatley [42] suggest, prosody—encompassing meter, rhythm, and rhyme—enhances both emotions and their perception. This highlights the significance of integrating NLP with music processing, as musical features play a crucial role in this context. 3.2 Lyrics Classification Overall, for lyrics classification, four different approaches were compared using our training set: a Support Vector Machine (SVM) as a baseline, a shallow neural network (FastText), and two neural networks, specifically a four-layer Artificial Neural Network (ANN) and a transformer-based approach. The following sections describe each of these approaches. However, the starting point is Section 3.2.1, which outlines the feature extraction process employed by both the SVM and the 4-layer ANN methods. 221 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek for each partition of the English lyric dataset. Despite the minimal observed discrepancy in the number of examples in the subsets, this partitioning was considered balanced. 3 Methods This Section outlines the development of a bimodal deep-learning architecture for classifying emotions in music tracks. The proposed approach integrates both lyrical and audio processing to capture the emotional nuances embedded within musical pieces. To evaluate classification performance, we employed multiple metrics, including Accuracy, Precision, Recall, Support, and the F1 score. 3.1 Lyrics Preprocessing Several text-cleaning operations were performed for each lyric: the song title was removed from the lyrics, as well as punctuation marks, tagging metadata, and extra newline control characters. After that, part-of-speech (PoS) tagging was performed in order to achieve morpho-syntactic features for some of the lyrics classification methods. For such a process, SpaCy tagger, a Python NLP toolkit [40], was used. Table 4 describes some statistics related to the number of tokens of the lyrics from the training set. In addition, Table 5 shows the distribution of PoS tags occurring in the lyrics (absolute frequencies and percentages per PoS tag in parentheses), pairing attention to those PoS tags that can express feelings (nouns, adjectives, verbs, adverbs, and interjections). As seen from both tables, on average, a lyric is around 300 tokens, and the number of words potentially expressing feelings in the dataset is high. Table 4. Statistics of the length in tokens of a lyric of the English lyric dataset. Happy Angry Sad Relaxed Mean 375.96 342.63 228.97 230.21 Median 340.0 273.0 206.0 201.0 Maximum 1,432 1,321 1,199 872 Minimum 11 56 12 10 Standard Deviation 198.84 224.76 134.74 135.04 From Table 5, some trends can be discerned a priori: a) interjections occur significantly more times in lyrics evoking ’Angry’ (45.8%) than in lyrics of the rest of the categories, b) nouns and verbs appear significantly more times in lyrics evoking ’Angry’ (30.2% and 31.9%, respectively) and ’Sad’ (32.9% and 29.8%, respectively) than in the rest of lyrics. Table 5. Statistics of the PoS tags occurring in the English lyric dataset. PoS tag Happy Angry Sad Relaxed adjective 5,122 (20.1%) 7,888 (31.0%) 7,568 (29.8%) 4,858 (19.1%) adverb 6,470 (23.2%) 8,281 (29.7%) 7,041 (25.2%) 6,131 (22.0%) interjection 1,231 (11.9%) 4,729 (45.8%) 2,350 (22.8%) 2,007 (19.5%) noun 14,846 (18.6%) 24,084 (30.2%) 26,287 (32.9%) 14,585 (18.3%) verb 17,229 (19.3%) 28,484 (31.9%) 26,632 (29.8%) 16,914 (18.9%) others 52,873 (20.0%) 82,461 (31.1%) 77,621 (29.3%) 51,955 (19.6%) Knowing the distribution of the PoS tags throughout the dataset may allow checking whether their distribution corresponds to expressing feelings in poetry [41]. However, it is important to note that understanding natural language alone does not account for the emotional impact of poetry. As Johnson-Laird and Oatley [42] suggest, prosody—encompassing meter, rhythm, and rhyme—enhances both emotions and their perception. This highlights the significance of integrating NLP with music processing, as musical features play a crucial role in this context. 3.2 Lyrics Classification Overall, for lyrics classification, four different approaches were compared using our training set: a Support Vector Machine (SVM) as a baseline, a shallow neural network (FastText), and two neural networks, specifically a four-layer Artificial Neural Network (ANN) and a transformer-based approach. The following sections describe each of these approaches. However, the starting point is Section 3.2.1, which outlines the feature extraction process employed by both the SVM and the 4-layer ANN methods. A BIMODAL DEEP MODEL TO . . . 3.2.1 Feature Extraction for SVM and 4-Layer ANN The feature set extracted from the lyrics was based on the methodology proposed by Giammusso et al. [23]. We conducted a Principal Component Analysis (PCA) on the selected lyrics features, and the results demonstrated that variance is distributed across multiple components rather than being concentrated in just a few. Given that no single principal component dominated the variance, we decided to use all the selected features from Giammusso et al. [23]. Therefore, features included: – percentage of present past tense verbs ca.lculated as the ratio of present past tense verbs to the total number of verbs; – percentage of adjectives as the ratio of adjectives to the total number of words; – percentage of punctuation, i.e., the ratio of punctuation marks to the total number of words; – percentage of echoism defined as the ratio of echoisms, either a sequence of two subsequent repeated words or the repetition of a vowel in a word, e.g., ’yeaaaah’ to the total number of words; – percentage of duplicated lines as the ratio of duplicated lines to the total number of lines in the lyric; – presence of the title, i.e., a Boolean value indicating whether the lyric contains the title string; – sentiment polarity quantifying the degree of positivity [close to 1], neutrality [0], or negativity [close to -1] presence in the lyrics; – subjectivity degree measuring the amount of personal opinion versus factual information. These features were combined with a word embedding vector generated using SpaCy’s pre-trained language model, which is based on GloVe [43]. Stopwords were removed during this step to ensure that the embeddings focused on meaningful content. The reason for selecting these features comes from their ability to capture both the linguistic structure and stylistic elements in the song lyrics. The relevance of the selected features was evaluated and the percentage of punctuation marks was found to be less significant, while the other features were considered relevant. This information is contained in Table 6. For parts-of-speech tagging and generating word embedding vectors, SpaCy pipelining for English (“en core web lg” version 3.5.0 [44]) was utilized. Sentiment analysis of the lyrics was conducted using TextBlob for Python (version 0.16.0) [45], which provided both sentiment polarity and subjectivity degree. Table 6. Features extracted for SVM and 4-layer ANN approaches. Feature Description Scope The % of present tense verbs [1, 100] The % of past tense verbs [1, 100] The % of adjectives [1, 100] The % of punctuation marks [1, 100] The % of echoes [1, 100] The % of duplicated lines [1, 100] Title presence [True or False] Sentiment polarity [-1.0, 1.0] Subjectivity degree [0.0, 1.0] Word embedding vector [1 x 300] Total number of features 310 The final feature vector for each lyric was of dimension 1 x 310, making it suitable for input into any machine learning classifier. 3.2.2 SVM Approach The Scikit-Learn [46] implementation of SVM was used in our experiments. Using 10-fold crossvalidation over the merge of the training and the validation sets, we conducted a grid search to find optimal outcomes in terms of Accuracy for two hyperparameters: regularization parameter Cand kernel function. Even though grid search has some limitations, such as inefficiency in high-dimensional or large search spaces and the inability to prioritize promising regions of the search space, for straightforward and smaller problems, grid search remains a solid choice – in contrast to alternative methods like Bayesian optimization, random search, etc., as it systematically explores the entire search space by evaluating all possible combinations of specified hyperparameters. The best results were achieved for the combination of a linear kernel and C= 222 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek 0.01. The SVM reached 57.1% (+/- 2.9%) accuracy on average on the validation partitions; very small standard deviation values indicate that the achieved results are statistically significant. Finally, the SVM resulted in 55.0% accuracy on the test set. 3.2.3 4-Layers ANN Approach A model was designed with four dense layers, and a dropout layer was applied after each layer to avoid overtraining. ReLU activation function was used for the subsequent layers, while a softmax was employed for the final layer. The deep model architecture is shown in Figure 2. Figure 2. 4-layer ANN architecture. Similar to the SVM approach, a grid search was performed to find the optimal combination of hyperparameters in terms of Accuracy. The model was trained for a fixed number of epochs while monitoring the validation loss. Models with higher dropout rates exhibited lower training accuracy but better generalization performance. Increasing the number of neurons in the dense layers showed an improvement in performance. Finally, lower learning rates and larger batch sizes generally led to more stable training and better generalization. The best results were obtained for learning rate = 0.01, batch size = 64, epochs = 20, dense size = 128, and dropout = 0.3. It achieved 54.1% accuracy on the validation set and 53.9% on the test set. 3.2.4 FastText Approach One of the key challenges in lyric classification is capturing the nuanced contextual relationships between words, which often convey emotion, style, and genre-specific patterns. To address this, FastText was used because of its ability to effectively capture substructure at the word level and contextual information. FastText leverages word ngrams to represent text, allowing the model to capture local word dependencies and patterns more effectively, which are crucial in lyrics, where emotion frequently depends on the arrangement of just a few words. This algorithm represents each word in the lyrics as a unit and generates a set of n-grams from the sequence of words within a defined context window. These n-grams allow the model to encode word order and neighborhood patterns that enhance its classification capabilities. For example, considering a lyric segment ”Take me down to the Paradise City” from Guns N’ Roses [47] with five as the size of the context window, the generated word n-grams include: – Unigrams: “Take,” “me,” “down,” “to,” “the,” “paradise,” “city”; – Bigrams: “Take me,” “me down,” “down to,” “to the,” “the paradise,” “paradise city”; – Trigrams: “Take me down,” “me down to,” “down to the,” “to the paradise,” “the paradise city”; FastText v0.9.2 [48] was used in our experiment. The following parameters and hyperparameters of this function were experimented with: ngrams length (maximum length of a character ngram), window size (size of the context window for a given n-gram), learning rate, number of epochs, loss function (softmax - preferred for this kind of classification problem; or hierarchical softmax). The parameters and hyperparameters were initially tuned based on our domain knowledge and intuition. Subsequently, the autotune function provided by FastText was employed to refine these parameters further. Finally, the optimal model was selected by observing the classification metrics on the validation set. The best results were obtained for word n-grams = 3; size of the context window = 5; loss function = softmax;learning rate = 0.8; and the number of epochs = 20. The model achieved 51.9% accuracy on the validation set and 48.2% on the test set. 3.2.5 Transformer-Based Approach Our transformer-based approach is based on the experiment described by Agrawal et al. [24] using a fine-tuned XLNet. However, the implementation differs as the goal of our study - as already mentioned - was to classify lyrics into four classes of emotion (happy, angry, relaxed, and sad) without considering the rest of the outputs achieved by 223 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek 0.01. The SVM reached 57.1% (+/- 2.9%) accuracy on average on the validation partitions; very small standard deviation values indicate that the achieved results are statistically significant. Finally, the SVM resulted in 55.0% accuracy on the test set. 3.2.3 4-Layers ANN Approach A model was designed with four dense layers, and a dropout layer was applied after each layer to avoid overtraining. ReLU activation function was used for the subsequent layers, while a softmax was employed for the final layer. The deep model architecture is shown in Figure 2. Figure 2. 4-layer ANN architecture. Similar to the SVM approach, a grid search was performed to find the optimal combination of hyperparameters in terms of Accuracy. The model was trained for a fixed number of epochs while monitoring the validation loss. Models with higher dropout rates exhibited lower training accuracy but better generalization performance. Increasing the number of neurons in the dense layers showed an improvement in performance. Finally, lower learning rates and larger batch sizes generally led to more stable training and better generalization. The best results were obtained for learning rate = 0.01, batch size = 64, epochs = 20, dense size = 128, and dropout = 0.3. It achieved 54.1% accuracy on the validation set and 53.9% on the test set. 3.2.4 FastText Approach One of the key challenges in lyric classification is capturing the nuanced contextual relationships between words, which often convey emotion, style, and genre-specific patterns. To address this, FastText was used because of its ability to effectively capture substructure at the word level and contextual information. FastText leverages word ngrams to represent text, allowing the model to capture local word dependencies and patterns more effectively, which are crucial in lyrics, where emotion frequently depends on the arrangement of just a few words. This algorithm represents each word in the lyrics as a unit and generates a set of n-grams from the sequence of words within a defined context window. These n-grams allow the model to encode word order and neighborhood patterns that enhance its classification capabilities. For example, considering a lyric segment ”Take me down to the Paradise City” from Guns N’ Roses [47] with five as the size of the context window, the generated word n-grams include: – Unigrams: “Take,” “me,” “down,” “to,” “the,” “paradise,” “city”; – Bigrams: “Take me,” “me down,” “down to,” “to the,” “the paradise,” “paradise city”; – Trigrams: “Take me down,” “me down to,” “down to the,” “to the paradise,” “the paradise city”; FastText v0.9.2 [48] was used in our experiment. The following parameters and hyperparameters of this function were experimented with: ngrams length (maximum length of a character ngram), window size (size of the context window for a given n-gram), learning rate, number of epochs, loss function (softmax - preferred for this kind of classification problem; or hierarchical softmax). The parameters and hyperparameters were initially tuned based on our domain knowledge and intuition. Subsequently, the autotune function provided by FastText was employed to refine these parameters further. Finally, the optimal model was selected by observing the classification metrics on the validation set. The best results were obtained for word n-grams = 3; size of the context window = 5; loss function = softmax;learning rate = 0.8; and the number of epochs = 20. The model achieved 51.9% accuracy on the validation set and 48.2% on the test set. 3.2.5 Transformer-Based Approach Our transformer-based approach is based on the experiment described by Agrawal et al. [24] using a fine-tuned XLNet. However, the implementation differs as the goal of our study - as already mentioned - was to classify lyrics into four classes of emotion (happy, angry, relaxed, and sad) without considering the rest of the outputs achieved by A BIMODAL DEEP MODEL TO . . . Agrawal et al.’s implementation. Hence, we decided to implement only the outputs related to these quadrants. This allows the results obtained to be compared with other approaches dealing with the dataset. In addition, as mentioned by Agrawal et al. [24], the classification into four categories (a singletask approach) performs marginally better, and the computational cost is much lower. The architecture of the model is shown in Figure 3. The transformer-based classifier is fed by the lyrics, and the outputted hidden states should be passed to the Sequence Summary block based on the average values of the hidden states. This vector passes to the fully connected layer, which is directly connected to the output layer. After that, through the softmax activation function, we obtain the probability that a lyric belongs to the given class. For this approach, the fine-tuning of the pretrained models was applied. Figure 3. Transformer-based model architecture. To apply the fine-tuning approach, the Transformers - Python library from Hugging Face [49] was used; specifically, XLNet-base-cased [50], the pre-trained model, was employed. For model training, following the experiment authors [24], the optimizer AdamW with a learning rate of 2e-5 was employed. Cross-entropy was utilized to calculate the loss. The remaining parameters were for finetuning used according to the XLNet authors’ implementation [51]. The models were fine-tuned for four epochs, with batch size 32. By observing learning curves for training and validation sets, the optimal model was found with parameters: number of embeddings = 256, batch size = 64, learning rate = 1e-5, epochs = 4, and without applying lower case during tokenization. The model achieved 62.3% accuracy on the validation set and 58.0% accuracy on the test set. 3.2.6 Best Model for Lyrics Classification From the conducted experiments, it can be concluded that the larger size of the input embedding positively impacts classification performance; nonetheless, it implies a much longer training time and requirements of available operation memory. It should be, however, noted that the size of the batch was selected experimentally; too small or too large worsened the generalization ability of the model. Table 7 summarizes the results achieved by each of the researched approaches. Values of the classification metrics in the table are the weighted average. Comparing the achieved results, XLNet cased shows promising results, surpassing other algorithms in accuracy measure. The feature-based approach and FastText scored similar results, but SVM proved to be the best classifier of the first three ML algorithms. The highest accuracy was achieved by the transformer-based approach, and at the same time, it had the highest precision and recall. When the studied models are considered as a recommendation system, Precision is a more important metric than Recall, as bad recommendation has a direct impact on listener dissatisfaction. Table 7. Overall results in % on the test set. Bold values indicate the best-performing model. Classifier approach Precision Recall F1 score Accuracy Featurebased SVM 54.8 55.0 54.9 55.0 Featurebased ANN 54.0 53.9 53.9 53.9 FastTextbased 50.3 50.4 50.3 48.2 Transformerbased (XLNet cased) 58.2 58.0 58.1 58.0 Therefore, the transformer-based model was considered the best among the studied approaches in terms of classification performance. The detailed results of the classification metrics for the most promising approach (XLNet cased) are shown in Table 8. The results presented were obtained using the entire test set, including song lyrics in other languages and instrumental pieces (treated as single white-space characters). The model achieved a high Recall value for the Angry class and a bit lower for the happy and relaxed classes. In contrast, the model performed most poorly in recognizing the lyrics of music pieces labeled as sad. The detailed performance of the Transformer-based approach across all emo- 230 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek 3.5.2 Concatenated Approach Another tested approach was combining models using the concatenation method. This technique merges diverse outputs from two separate models into a single model, which can subsequently run the training phase, usually at the level of the final dense layers. However, it is important to note that the data dimensions can rapidly become quite large, increasing the risk of overfitting. To overcome this, a standard solution is to employ pre-trained models, remove the final classification layers and activations, and then fine-tune both partial models included within the ensemble model [60]. Figure 11. Concatenated model architecture. Similar to the previous approach, the PyTorch library was used. First of all, a custom transformerbased class needed to be implemented. This class removed the last classification layer along with the softmax activation. Following that, the weights from the previously trained lyrics model were loaded, along with the weights from Sarkar et al.’s architecture (both without the classification layer and softmax function) [22]. Next, the models were trained using the frozen weights of the lyrics and audio models, with the addition of various numbers of dense layers and optional dropout layers. The conducted grid search did not significantly improve the results. Following this, the optimal model comprised three Dense layers (512, 256, and 128), each followed by the ReLU activation method. The model was trained on batch size=16, with learning rate=1e-6 for 100 epochs. The simplified architecture is shown in Figure 11. The training history for the concentration ensemble model is displayed in Figure 12 (accuracy curve) and Figure 13 (loss curve). As can be observed, the best model’s accuracy was obtained between 80th and 100th epoch. Despite various experiments with a higher number of epochs, validation accuracy did not increase. Figure 12. Accuracy training curve for the concatenated model architecture. Figure 13. Loss training curve for the concatenated model architecture. Ultimately, the model achieved 58.9% accuracy on the validation and 58.4% on the test set. The metrics resulting from the combined modalities are presented in Table 13. Compared to the performance measures calculated for the separate audio and lyrics modules, a significant improvement in the F1 score was observed for the sad emotion. Unfortunately, this improvement was accompanied by a slight deterioration in the F1 score for other emotions. The model appears to perform relatively well in classifying relaxed and angry emotions. However, despite the remarkable improvement in recognizing sad emotion, it still struggles with its classification. The relatively high Recall for the relaxed emotion indicates that the model has a low rate of incorrectly classifying instances of the relaxed class as another emotion, which can also be observed in Figure 14. 231 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek 3.5.2 Concatenated Approach Another tested approach was combining models using the concatenation method. This technique merges diverse outputs from two separate models into a single model, which can subsequently run the training phase, usually at the level of the final dense layers. However, it is important to note that the data dimensions can rapidly become quite large, increasing the risk of overfitting. To overcome this, a standard solution is to employ pre-trained models, remove the final classification layers and activations, and then fine-tune both partial models included within the ensemble model [60]. Figure 11. Concatenated model architecture. Similar to the previous approach, the PyTorch library was used. First of all, a custom transformerbased class needed to be implemented. This class removed the last classification layer along with the softmax activation. Following that, the weights from the previously trained lyrics model were loaded, along with the weights from Sarkar et al.’s architecture (both without the classification layer and softmax function) [22]. Next, the models were trained using the frozen weights of the lyrics and audio models, with the addition of various numbers of dense layers and optional dropout layers. The conducted grid search did not significantly improve the results. Following this, the optimal model comprised three Dense layers (512, 256, and 128), each followed by the ReLU activation method. The model was trained on batch size=16, with learning rate=1e-6 for 100 epochs. The simplified architecture is shown in Figure 11. The training history for the concentration ensemble model is displayed in Figure 12 (accuracy curve) and Figure 13 (loss curve). As can be observed, the best model’s accuracy was obtained between 80th and 100th epoch. Despite various experiments with a higher number of epochs, validation accuracy did not increase. Figure 12. Accuracy training curve for the concatenated model architecture. Figure 13. Loss training curve for the concatenated model architecture. Ultimately, the model achieved 58.9% accuracy on the validation and 58.4% on the test set. The metrics resulting from the combined modalities are presented in Table 13. Compared to the performance measures calculated for the separate audio and lyrics modules, a significant improvement in the F1 score was observed for the sad emotion. Unfortunately, this improvement was accompanied by a slight deterioration in the F1 score for other emotions. The model appears to perform relatively well in classifying relaxed and angry emotions. However, despite the remarkable improvement in recognizing sad emotion, it still struggles with its classification. The relatively high Recall for the relaxed emotion indicates that the model has a low rate of incorrectly classifying instances of the relaxed class as another emotion, which can also be observed in Figure 14. A BIMODAL DEEP MODEL TO . . . Table 13. Results in % on the test set for the concatenation ensemble. Precision Recall F1 Score Support Happy 61.9 52.0 56.5 75 Angry 61.1 59.5 60.3 74 Sad 49.3 50.0 50.0 74 Relaxed 61.4 72.0 66.3 75 Accuracy - - 58.4 298 Macro average 58.4 58.4 58.4 298 Weighted average 58.5 58.4 58.4 298 Figure 14. Confusion matrix on the test set for the concatenation ensemble. 3.5.3 The comparison of joint classification This section compares the two ensemble approaches — Majority Voting and Concatenation — for the joint classification of audio and lyric modalities. These methods aim to combine the strengths of both modalities to enhance overall performance. Table 14. Comparison in % of joint classification on the test set. Bold values indicate the best-performing model. Ensemble Approach Precision Recall F1 Score Accuracy Majority Voting 61.5 60.7 61.1 60.7 Concatenation 58.5 58.4 58.4 58.4 Table 14 shows a comparison of both methods. The Majority Voting ensemble slightly outperformed the Concatenation approach in all metrics. Despite extensive fine-tuning and grid searches to optimize hyperparameters, the Concatenation model achieved a slightly lower F1 score of 58.4% and an accuracy of 58.4%. However, it is worth noting that this approach yielded a noticeable improvement in classifying the “sad” emotion. A plausible reason for the overall lower results in the Concatenation approach could be the limited dataset size available for fine-tuning the joint model. This limitation may have restricted the model’s ability to generalize effectively across all classes. The findings highlight the importance of selecting an appropriate ensemble approach for multimodal classification tasks. While the Majority Voting ensemble demonstrated overall better performance, the Concatenation approach showed potential in specific areas, such as improving the recognition of underperforming classes. 4 Discussion Despite not resulting in a significant accuracy improvement from 60.7% to the previous 56.7% based solely on lyrics and from 59.06% exclusively based on audio, emotion recognition now relies on both components, making the system more reliable. However, as already mentioned, Precision plays a more critical role in ensuring user satisfaction in the context of recommendation systems. This is supported by the literature sources, highlighting that multimodal approaches increase precision, which is crucial for recommendation systems to avoid user dissatisfaction from bad recommendations [61]. Moreover, Baltruˇ saitis et al. [26] pointed out that multimodality helps to enhance the performance and robustness of machine learning models, as well as the precision of recommendations in reallife situations and further augment user experience. Hence, by retrieving precision values from Tables 8, 10, and 14, it may be seen that the values of this metric are as follows: 56.95% for the lyrics only while employing XLNet, 54.33% achieved by each of the lyrics-based researched approaches, and 60.43% for audio-only. In contrast, the bimodal approach returned a Precision of 61.5%, so a small improvement was obtained when fusing two modalities. A comparison of this metric can be observed in Figure 15. Unlike the audio or lyrics modalities working independently, it also tended to offer a 232 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek more balanced prediction across different emotions, reducing severe misclassification, especially in the case of angry and sad emotions. Figure 15. The comparison of Precision across modalities and approaches. The experiments conducted in this study have revealed interesting insights into the fusion of models for musical piece classification. Surprisingly, a majority voting approach led to a modest improvement, achieving 60.7% accuracy on the test set. Initially, it was expected that fine-tuning of the combined models would outperform the simplistic voting method. However, the fine-tuned approach achieved 58.4% accuracy, performing worse than using the audio modality alone. Figure 16 provides a comparison of the F1 scores achieved through the best-performing single modality and concatenated approaches, highlighting the strengths and limitations of each in predicting different emotions. It is important to highlight that majority voting showed a higher F1 score for the three emotions (happy, angry, relaxed) and a notably lower F1 score for the sad emotion, compared to employing a concatenation ensemble, where more balanced results were observed. This disparity might be attributed to inconsistencies in predicting emotions within individual modalities. For instance, a melancholic song with relaxed lyrics or positive lyrics accompanied by aggressive metal sounds may not necessarily be assessed as sad. This suggests that while majority voting enhances performance overall, its effectiveness diminishes in cases where individual modalities provide conflicting signals, as evidenced in the Sad emotion category. Figure 16. The comparison of F1 score across modalities and approaches. In summary, the majority voting ensemble showcased proficient classification for three out of four emotions (happy, angry, and relaxed). In contrast, the concatenation ensemble demonstrated enhanced reliability across all emotions. Figure 17 shows a comparison of the average-weighted classification metrics. These findings underscore the complex nature of multimodal emotion classification. Considering the objective of combining modalities, we tested the model trained on the English lyric dataset on the full dataset, including previously indicated outliers. However, the results achieved in this study are not entirely satisfactory. The outcomes obtained using lyric modality differ from what was expected. Table 15 summarizes the results achieved by each approach researched based on lyric classification. Outlined are methods based on text feature extraction, embedding, capturing character-level representation, and applying attention mechanisms using transformer architecture. It should be noted that the provided comparison only provides an overview of the expected results. The SVM and 4-layer ANN were obtained with Englishonly lyrics, utilizing 90% of the examples as a training set and only 10% as a test set, which may lead to the model over-fitting to the training data. At the same time, the transformed-based approach [24] was achieved using MoodyLyric, which is a previous version of MoodyLyric4Q, containing repetitions and conflict labels. This may be one of the reasons that the authors obtained such good results. 233 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek more balanced prediction across different emotions, reducing severe misclassification, especially in the case of angry and sad emotions. Figure 15. The comparison of Precision across modalities and approaches. The experiments conducted in this study have revealed interesting insights into the fusion of models for musical piece classification. Surprisingly, a majority voting approach led to a modest improvement, achieving 60.7% accuracy on the test set. Initially, it was expected that fine-tuning of the combined models would outperform the simplistic voting method. However, the fine-tuned approach achieved 58.4% accuracy, performing worse than using the audio modality alone. Figure 16 provides a comparison of the F1 scores achieved through the best-performing single modality and concatenated approaches, highlighting the strengths and limitations of each in predicting different emotions. It is important to highlight that majority voting showed a higher F1 score for the three emotions (happy, angry, relaxed) and a notably lower F1 score for the sad emotion, compared to employing a concatenation ensemble, where more balanced results were observed. This disparity might be attributed to inconsistencies in predicting emotions within individual modalities. For instance, a melancholic song with relaxed lyrics or positive lyrics accompanied by aggressive metal sounds may not necessarily be assessed as sad. This suggests that while majority voting enhances performance overall, its effectiveness diminishes in cases where individual modalities provide conflicting signals, as evidenced in the Sad emotion category. Figure 16. The comparison of F1 score across modalities and approaches. In summary, the majority voting ensemble showcased proficient classification for three out of four emotions (happy, angry, and relaxed). In contrast, the concatenation ensemble demonstrated enhanced reliability across all emotions. Figure 17 shows a comparison of the average-weighted classification metrics. These findings underscore the complex nature of multimodal emotion classification. Considering the objective of combining modalities, we tested the model trained on the English lyric dataset on the full dataset, including previously indicated outliers. However, the results achieved in this study are not entirely satisfactory. The outcomes obtained using lyric modality differ from what was expected. Table 15 summarizes the results achieved by each approach researched based on lyric classification. Outlined are methods based on text feature extraction, embedding, capturing character-level representation, and applying attention mechanisms using transformer architecture. It should be noted that the provided comparison only provides an overview of the expected results. The SVM and 4-layer ANN were obtained with Englishonly lyrics, utilizing 90% of the examples as a training set and only 10% as a test set, which may lead to the model over-fitting to the training data. At the same time, the transformed-based approach [24] was achieved using MoodyLyric, which is a previous version of MoodyLyric4Q, containing repetitions and conflict labels. This may be one of the reasons that the authors obtained such good results. A BIMODAL DEEP MODEL TO . . . Figure 17. The comparison of weighted average classification metrics across modalities and approaches. Due to the intensive processing of the lyric corpus, the number of examples in the training set was rather limited. Therefore, fine-tuning the transformer-based model was considered the best among the studied approaches in terms of classification performance. Nevertheless, when observing the confusion matrices, which were a part of the assessment, it can be concluded that ’happiness’ and ’anger’ are the most manageable set of emotions to detect. A possible reason for this is that they are two opposing feelings containing a high emotional charge. There were likely words or sentences in the text directly indicating these emotions. The classifiers did much worse in discerning ‘sadness’ from lyrics. Table 15. Comparison in % of different approaches - lyric modality - achieved on English dataset. Bold values indicate the best-performing model. Approach Accuracy seen in the literature Accuracy achieved in the study Feature-based SVM 58.0 [23] 55.0 Feature-based 4Layers ANN 58.5 [23] 53.9 Feature-based (GloVe-BiGRU) 50.73 [8] - Feature-based (GloVe+HAN) 58.8 [8] - Transformerbased (XLNet) 94.78 [24] (full dataset) 58.0 Table 16 compares the results of the emotion recognition task accomplished by the authors from the literature versus our outcomes using the audio modality. Essentially, deep learning models tend to outperform the classical ones. Also, there are noticeable discrepancies between the results achieved in literature sources and the results reported in this study. However, developing such a methodology, starting from emotion labeling through the choice of a dataset, a way of parametrizing lyrics and music signals, to the test stage, is far from, if not at all, reproducible. Some of the referenced studies employed a different approach to the assessment step, in some cases using smaller segments of musical pieces (e.g., 5 seconds) or evaluating results not on a test dataset but using a validation dataset. Moreover, although MoodyLyric [24],[29] and MoodyLyric4Q [30] were employed in the literature-based studies, they were not processed. In contrast, we decided to use only MoodyLyrics4Q, which has fewer weaknesses, i.e., the ambiguity of the labels of the music samples and fewer reported repetitions, as we removed duplicates at the preprocessing stage. It should, however, be noted that the model performs very well in predicting ’happiness’ and ’relaxed’ classes while performing slightly less effectively on angry and sad emotions. Table 16. Comparison of different approaches - audio modality. Bold values indicate the best-performing model. Approach Accuracy seen in the literature Accuracy achieved in the study SVM 50 [17] 31.38 Ravdess based architecture 65.96 [18] 47.99 InceptionV3 70-90 [19, 20] 52.89 ResNet 77.36 [21] 56.04 VGG16 63.79 [21] 53.69 Sarkar et al.’s architecture 68-78 [22] 59.06 InceptionResNet 84.91-87.24 [21] 56.23 The experiments that were conducted revealed that the two models (both for lyric and audio) complement each other. Losses are propagated in the same manner during the training of both models. However, the lyrics model demonstrates higher Pre- 234 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek cision for the happy emotion and higher Recall for ’anger.’ On the other hand, the audio modality is better regarding Recall for happy and relaxed emotions, and it achieves higher Precision for angry emotion. In contrast, both modalities exhibit low F1 scores for the sad emotion, resulting in unsatisfactory results for this particular emotion. As seen in Table 15 and Table 16, several state-ofthe-art papers were cited along with the experiment outcomes (e.g., [8],[62],[63]. When referring to state-of-the-art (SOTA) methods, we have compared our results with those from previous studies that addressed emotion classification in music using the same dataset. However, it is important to note that these comparisons serve primarily to illustrate the uniqueness of each approach. Although all studies began with the same initial dataset, differences in data processing and methodology mean that direct performance comparisons should be interpreted with caution. Pyrovolakis et al. [62] divided each song into a 4-line sample. This way, they obtained more than 18,000 samples from the 2,000 samples in the training set. Shaday et al. [63] study evaluated an algorithm’s performance on two datasets: the Music Lyrics Kaggle dataset (3,890 entries, eight emotional categories) and the MoodyLyrics dataset (2,000 entries, four emotional categories). Key testing parameters included embedding weights, embedding dimension, learning rate, and epoch. Experimental results show an accuracy of 62% on the MoodyLyrics dataset and 42% on the Music Lyrics Kaggle dataset. Sujeeha et al.’s [8] approach used only 680 emotion-labeled samples extracted from MoodyLyrics as they investigated the impact of attention mechanisms on multi-modal architectures for audio mood classification. Experiments compare single-modal and multi-modal models with and without attention. Three attention mechanisms—hierarchical attention network, self-attention, and channel attention—are utilized. Acoustic features are extracted from songs, and the model is trained using BiGRU (Bidirectional Gated Recurrent Unit) and CNN, achieving testing accuracies of 61.01%, 44.15%, and 50.73%. 5 Conclusion Several experiments were conducted to evaluate the possibility of assigning emotion automatically based on lyrics and audio independently, as well as when fusing these modalities into a bimodal approach. The results showed promising outcomes, with relatively high accuracy and precision rates across multiple emotion categories, indicating the algorithm’s stability and ability to recognize and label emotions in musical pieces effectively. The best algorithm successfully extracted and classified emotional features in musical pieces with 60.7% accuracy. However, for some parameter configurations, the trained model classified all musical pieces into one emotion class. This may be due to the too-small size of the batch, in which emotions from only one category were probably drawn, while at the same time, the learning rate was too high, which perhaps resulted in adjusting the model’s weights to the last training data batch. One of the key findings of this study is the algorithm’s robustness in dealing with a variety of music genres, styles, and cultural contexts. The diverse datasets used in this study encompass a wide range of musical pieces from various genres, including rock, jazz, electronic, pop, blues, and country. Despite this diversity, algorithms consistently demonstrated reliable performance across all genres. This suggests that the algorithm’s effectiveness is not limited to specific musical styles and can be applied to a broad spectrum of music. The model’s not entirely satisfactory performance is due to the complexity of the problem and the wide variety of texts (for example, in terms of length or structure). However, this may indicate that utilizing transformer models is an effective strategy for lyric classification in the context of the emotion evoked on the MoodyLyric4Q database. While the results are promising, it is important to acknowledge the limitations of this study. Firstly, the dataset used, although diverse, may not capture the full spectrum of emotions across all cultures and music traditions. Future research could benefit from larger and more culturally diverse datasets to enhance the algorithm’s cross-cultural applicability. The solution to this problem may be to generate datasets using the teacher-student machine learning paradigm automatically. In addition, the lyrics could be taken directly from the music, resulting in a more accurate recording of the words contained in the song. 235 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek cision for the happy emotion and higher Recall for ’anger.’ On the other hand, the audio modality is better regarding Recall for happy and relaxed emotions, and it achieves higher Precision for angry emotion. In contrast, both modalities exhibit low F1 scores for the sad emotion, resulting in unsatisfactory results for this particular emotion. As seen in Table 15 and Table 16, several state-ofthe-art papers were cited along with the experiment outcomes (e.g., [8],[62],[63]. When referring to state-of-the-art (SOTA) methods, we have compared our results with those from previous studies that addressed emotion classification in music using the same dataset. However, it is important to note that these comparisons serve primarily to illustrate the uniqueness of each approach. Although all studies began with the same initial dataset, differences in data processing and methodology mean that direct performance comparisons should be interpreted with caution. Pyrovolakis et al. [62] divided each song into a 4-line sample. This way, they obtained more than 18,000 samples from the 2,000 samples in the training set. Shaday et al. [63] study evaluated an algorithm’s performance on two datasets: the Music Lyrics Kaggle dataset (3,890 entries, eight emotional categories) and the MoodyLyrics dataset (2,000 entries, four emotional categories). Key testing parameters included embedding weights, embedding dimension, learning rate, and epoch. Experimental results show an accuracy of 62% on the MoodyLyrics dataset and 42% on the Music Lyrics Kaggle dataset. Sujeeha et al.’s [8] approach used only 680 emotion-labeled samples extracted from MoodyLyrics as they investigated the impact of attention mechanisms on multi-modal architectures for audio mood classification. Experiments compare single-modal and multi-modal models with and without attention. Three attention mechanisms—hierarchical attention network, self-attention, and channel attention—are utilized. Acoustic features are extracted from songs, and the model is trained using BiGRU (Bidirectional Gated Recurrent Unit) and CNN, achieving testing accuracies of 61.01%, 44.15%, and 50.73%. 5 Conclusion Several experiments were conducted to evaluate the possibility of assigning emotion automatically based on lyrics and audio independently, as well as when fusing these modalities into a bimodal approach. The results showed promising outcomes, with relatively high accuracy and precision rates across multiple emotion categories, indicating the algorithm’s stability and ability to recognize and label emotions in musical pieces effectively. The best algorithm successfully extracted and classified emotional features in musical pieces with 60.7% accuracy. However, for some parameter configurations, the trained model classified all musical pieces into one emotion class. This may be due to the too-small size of the batch, in which emotions from only one category were probably drawn, while at the same time, the learning rate was too high, which perhaps resulted in adjusting the model’s weights to the last training data batch. One of the key findings of this study is the algorithm’s robustness in dealing with a variety of music genres, styles, and cultural contexts. The diverse datasets used in this study encompass a wide range of musical pieces from various genres, including rock, jazz, electronic, pop, blues, and country. Despite this diversity, algorithms consistently demonstrated reliable performance across all genres. This suggests that the algorithm’s effectiveness is not limited to specific musical styles and can be applied to a broad spectrum of music. The model’s not entirely satisfactory performance is due to the complexity of the problem and the wide variety of texts (for example, in terms of length or structure). However, this may indicate that utilizing transformer models is an effective strategy for lyric classification in the context of the emotion evoked on the MoodyLyric4Q database. While the results are promising, it is important to acknowledge the limitations of this study. Firstly, the dataset used, although diverse, may not capture the full spectrum of emotions across all cultures and music traditions. Future research could benefit from larger and more culturally diverse datasets to enhance the algorithm’s cross-cultural applicability. The solution to this problem may be to generate datasets using the teacher-student machine learning paradigm automatically. In addition, the lyrics could be taken directly from the music, resulting in a more accurate recording of the words contained in the song. A BIMODAL DEEP MODEL TO . . . Secondly, we employed a discrete (hard) representation of emotion when dividing and distinguishing emotions. However, emotions present in musical pieces are often not coherent. A better approach might involve utilizing the arousal-valence dimension to determine the song’s true nature. In contrast, sparse dimension is not as easily understood and deciphered. Therefore, an approach that involves multiple labels (soft approach) could be employed when a song expresses more than one emotion. Additionally, the evaluation of emotional labels - both inscribed and perceived feelings - may be subjective to some extent, as emotions are inherently complex and can vary from person to person. Future work could explore more user-based evaluations - not only in the context of creating a database but also in terms of verifying the results to understand how the algorithm aligns with human perceptions of emotion in music. Moreover, since the LRAP (Label ranking average precision) measure corresponds to the quality of identification in the context of a task with multiple classes present in a given sample, employing such a measure may be advantageous when involving multiple emotions assigned, especially when adopting a soft approach rather than a hard one. Finally, one major limitation of this study is the lack of synchronization between audio fragments and lyrics, which can convey different emotions. Future research on musical piece classification should incorporate full audio tracks and automatically extract synchronized lyrics for more accurate analysis. References [1] L. Smietanka and T. Maka, “Interpreting convolutional layers in DNN model based on time–frequency representation of emotional speech,” Journal of Artificial Intelligence and Soft Computing Research, vol. 14, no. 1, pp. 5–23, Jan. 2024, doi: 10.2478/jaiscr-2024-0001. [2] S. Sheykhivand, Z. Mousavi, T. Y. Rezaii, and A. Farzamnia, “Recognizing Emotions Evoked by Music Using CNN-LSTM Networks on EEG Signals,” IEEE Access, vol. 8, pp. 139332-139345, 2020, doi: 10.1109/ACCESS.2020.3011882. [3] Y. Takahashi, T. Hochin, and H. Nomiya, “Relationship between Mental States with Strong Emotion Aroused by Music Pieces and Their Feature Values,” in Proc. 2014 IIAI 3rd International Conference on Advanced Applied Informatics, 2014, pp. 718-725, doi: 10.1109/IIAIAAI.2014.147. [4] P. A. Wood and S. K. Semwal, “On exploring the connection between music classification and evoking emotion,” in Proc. 2015 International Conference on Collaboration Technologies and Systems(CTS), 2015, pp. 474-476, doi: 10.1109/CTS.2015.7210471. [5] M. Agapaki, E. A. Pinkerton, and E. Papatzikis, “Music and neuroscience research for mental health, cognition, and development: Ways forward,” Frontiers in Psychology, vol. 13, 2022, doi: https://doi.org/10.3389/fpsyg.2022.976883. [6] Y. Song, S. Dixon, M. Pearce, and A. Halpern, “Perceived and Induced Emotion Responses to Popular Music: Categorical and Dimensional Models,” Music Perception: An Interdisciplinary Journal, vol. 33, pp. 472-492, Apr. 2016, doi: 10.1525/mp.2016.33.4.472. [7] Y. Yuan, “Emotion of Music: Extraction and Composing,” Journal of Education, Humanities and Social Sciences, vol. 13, pp. 422-428, May 2023, doi: 10.54097/ehss.v13i.8207. [8] S. A. Sujeesha, J. B. Mala, and R. Rajeev, “Automatic music mood classification using multi-modal attention framework,” *Engineering Applications of Artificial Intelligence*, vol. 128, p. 107355, 2024, doi: 10.1016/j.engappai.2023.107355. [9] M. Schedl, P. Knees, B. McFee, D. Bogdanov, and M. Kaminskas, “Music recommender systems,” in Recommender systems handbook, Springer, 2015, pp. 453-492. [10] MorphCast Technology. Available: https://www.morphcast.com. Accessed: November 2024. [11] S. Zhao, G. Jia, J. Yang, G. Ding, and K. Keutzer, “Emotion Recognition From Multiple Modalities: Fundamentals and methodologies,” IEEE Signal Processing Magazine, vol. 38, no. 6, pp. 59-73, Nov. 2021, doi: 10.1109/msp.2021.3106895. [12] T. Li, “Music emotion recognition using deep convolutional neural networks,” Journal of Computational Methods in Science and Engineering, vol. 24, no. 4-5, pp. 3063-3078, 2024, doi: 10.3233/JCM-247551. [13] P. L. Louro, H. Redinho, R. Malheiro, R. P. Paiva, and R. Panda, “A comparison study of deep learning methodologies for music emotion recognition,” Sensors, vol. 24, no. 7, p. 2201, 2024, doi: 10.3390/s24072201. 236 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek [14] M. Blaszke, G. Korvel, and B. Kostek, “Exploring neural networks for musical instrument identification in polyphonic audio,” IEEE Intelligent Systems, pp. 1-11, 2024, doi: 10.1109/mis.2024.3392586. [15] M. Barata and P. Coelho, “Music Streaming Services: Understanding the drivers of customer purchase and intention to recommend,” Heliyon, vol. 7, p. e07783, Aug. 2021, doi: 10.1016/j.heliyon.2021.e07783. [16] J. Webster, “The promise of personalization: Exploring how music streaming platforms are shaping the performance of class identities and distinction,” New Media & Society, p. 146144482110278, Jul. 2021, doi: 10.1177/14614448211027863. [17] E. Schmidt, D. Turnbull, and Y. Kim, “Feature selection for content-based, time-varying musical emotion regression,” in Proc ACM SIGMM Int Conf Multimedia Info Retrieval, Mar. 2010, pp. 267-274, doi: 10.1145/1743384.1743431. [18] Y.-H. Yang, Y.-C. Lin, H.-T. Cheng, I.-B. Liao, Y.- C. Ho, and H. H. Chen, “Toward Multimodal Music Emotion Classification,” in Advances in Multimedia Information Processing - PCM 2008, 2008, pp. 70-79. [19] T. Ciborowski, S. Reginis, D. Weber, A. Kurowski, and B. Kostek, “Classifying Emotions in Film Music—A Deep Learning Approach,” Electronics, vol. 10, no. 23, p. 2955, Nov. 2021, doi: 10.3390/electronics10232955. [20] X. Han, F. Chen, and J. Ban, “Music Emotion Recognition Based on a Neural Network with an Inception-GRU Residual Structure,” Electronics, vol. 12, no. 4, p. 978, Feb. 2023, doi: 10.3390/electronics12040978. [21] Y. J. Liao, W. C. Wang, S.-J. Ruan, Y. H. Lee, and S. C. Chen, “A Music Playback Algorithm Based on Residual-Inception Blocks for Music Emotion Classification and Physiological Information,” Sensors, vol. 22, no. 3, p. 777, Jan. 2022, doi: 10.3390/s22030777. [22] R. Sarkar, S. Choudhury, S. Dutta, A. Roy, and S. K. Saha, “Recognition of emotion in music based on deep convolutional neural network,” Multimedia Tools and Applications, vol. 79, pp. 765-783, 2019, [Online]. Available: https://api.semanticscholar.org/CorpusID: 254866914. [23] S. Giammusso, M. Guerriero, P. Lisena, E. Palumbo, and R. Troncy, “Predicting the emotion of playlists using track lyrics,” International Society for Music Information Retrieval ISMIR, Late Breaking Session, 2017. [24] Y. Agrawal, R. Shanker, and V. Alluri, “Transformer-based approach towards music emotion recognition from lyrics,” Advances in Information Retrieval. ECIR 2021. Lecture Notes in Computer Science, vol 12657. Springer, 2021, doi: 10.1007/978-3-030-72240-1 12. [25] D. Han, Y. Kong, H. Jiayi, and G. Wang, “A survey of music emotion recognition,” Frontiers of Computer Science, vol. 16, Dec. 2022, doi: 10.1007/s11704-021-0569-4. [26] T. Baltruˇ saitis, C. Ahuja, and L. -P. Morency, ”Multimodal Machine Learning: A Survey and Taxonomy,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443, 1 Feb. 2019, doi: 10.1109/TPAMI.2018.2798607. [27] R. Delbouys, R. Hennequin, F. Piccoli, J. Royo-Letelier, and M. Moussallam, “Music Mood Detection Based On Audio And Lyrics With Deep Neural Net,” ISMIR 2018 https://doi.org/10.48550/arXiv.1809.07276 [28] I. A. P. Santana et al., “Music4all: A new music database and its applications,” in Proc. 2020 International Conference on Systems, Signals and Image Processing (IWSSIP), 2020, pp. 399-404, doi: 10.1109/IWSSIP48289.2020.9145170. [29] E. C¸ ano and M. Morisio, “Moodylyrics: A sentiment annotated lyrics dataset,” in Proc. 2017 International conference on intelligent systems, metaheuristics & swarm intelligence, 2017, pp. 118124, doi: 10.1145/3059336.3059340. [30] E. C¸ ano and M. Morisio, “Music mood dataset creation based on last.Fm tags,” in Proc. 2017 International Conference on Artificial Intelligence and Applications, Vienna, Austria, 2017, pp. 15-26, DOI:10.5121/csit.2017.70603. [31] R.E. Thayer: The Biopsychology of Mood and Arousal, Oxford University Press, 1989. [32] J. Russell, “A Circumplex Model of Affect,” Journal of Personality and Social Psychology, vol. 39, pp. 1161-1178, Dec. 1980, doi: 10.1037/h0077714. [33] Social music service - Last.fm. Available: https://www.last.fm/. Accessed: November 2024. [34] Genius - Song Lyrics & Knowledge. Available: https://genius.com/. Accessed: November 2024. [35] YouTube. Available: https://www.youtube.com. Accessed: November 2024. [36] M. Sakowicz and J. Tobolewski, “Development and study of an algorithm for the automatic labeling of musical pieces in the context of emotion 237 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek [14] M. Blaszke, G. Korvel, and B. Kostek, “Exploring neural networks for musical instrument identification in polyphonic audio,” IEEE Intelligent Systems, pp. 1-11, 2024, doi: 10.1109/mis.2024.3392586. [15] M. Barata and P. Coelho, “Music Streaming Services: Understanding the drivers of customer purchase and intention to recommend,” Heliyon, vol. 7, p. e07783, Aug. 2021, doi: 10.1016/j.heliyon.2021.e07783. [16] J. Webster, “The promise of personalization: Exploring how music streaming platforms are shaping the performance of class identities and distinction,” New Media & Society, p. 146144482110278, Jul. 2021, doi: 10.1177/14614448211027863. [17] E. Schmidt, D. Turnbull, and Y. Kim, “Feature selection for content-based, time-varying musical emotion regression,” in Proc ACM SIGMM Int Conf Multimedia Info Retrieval, Mar. 2010, pp. 267-274, doi: 10.1145/1743384.1743431. [18] Y.-H. Yang, Y.-C. Lin, H.-T. Cheng, I.-B. Liao, Y.- C. Ho, and H. H. Chen, “Toward Multimodal Music Emotion Classification,” in Advances in Multimedia Information Processing - PCM 2008, 2008, pp. 70-79. [19] T. Ciborowski, S. Reginis, D. Weber, A. Kurowski, and B. Kostek, “Classifying Emotions in Film Music—A Deep Learning Approach,” Electronics, vol. 10, no. 23, p. 2955, Nov. 2021, doi: 10.3390/electronics10232955. [20] X. Han, F. Chen, and J. Ban, “Music Emotion Recognition Based on a Neural Network with an Inception-GRU Residual Structure,” Electronics, vol. 12, no. 4, p. 978, Feb. 2023, doi: 10.3390/electronics12040978. [21] Y. J. Liao, W. C. Wang, S.-J. Ruan, Y. H. Lee, and S. C. Chen, “A Music Playback Algorithm Based on Residual-Inception Blocks for Music Emotion Classification and Physiological Information,” Sensors, vol. 22, no. 3, p. 777, Jan. 2022, doi: 10.3390/s22030777. [22] R. Sarkar, S. Choudhury, S. Dutta, A. Roy, and S. K. Saha, “Recognition of emotion in music based on deep convolutional neural network,” Multimedia Tools and Applications, vol. 79, pp. 765-783, 2019, [Online]. Available: https://api.semanticscholar.org/CorpusID: 254866914. [23] S. Giammusso, M. Guerriero, P. Lisena, E. Palumbo, and R. Troncy, “Predicting the emotion of playlists using track lyrics,” International Society for Music Information Retrieval ISMIR, Late Breaking Session, 2017. [24] Y. Agrawal, R. Shanker, and V. Alluri, “Transformer-based approach towards music emotion recognition from lyrics,” Advances in Information Retrieval. ECIR 2021. Lecture Notes in Computer Science, vol 12657. Springer, 2021, doi: 10.1007/978-3-030-72240-1 12. [25] D. Han, Y. Kong, H. Jiayi, and G. Wang, “A survey of music emotion recognition,” Frontiers of Computer Science, vol. 16, Dec. 2022, doi: 10.1007/s11704-021-0569-4. [26] T. Baltruˇ saitis, C. Ahuja, and L. -P. Morency, ”Multimodal Machine Learning: A Survey and Taxonomy,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443, 1 Feb. 2019, doi: 10.1109/TPAMI.2018.2798607. [27] R. Delbouys, R. Hennequin, F. Piccoli, J. Royo-Letelier, and M. Moussallam, “Music Mood Detection Based On Audio And Lyrics With Deep Neural Net,” ISMIR 2018 https://doi.org/10.48550/arXiv.1809.07276 [28] I. A. P. Santana et al., “Music4all: A new music database and its applications,” in Proc. 2020 International Conference on Systems, Signals and Image Processing (IWSSIP), 2020, pp. 399-404, doi: 10.1109/IWSSIP48289.2020.9145170. [29] E. C¸ ano and M. Morisio, “Moodylyrics: A sentiment annotated lyrics dataset,” in Proc. 2017 International conference on intelligent systems, metaheuristics & swarm intelligence, 2017, pp. 118124, doi: 10.1145/3059336.3059340. [30] E. C¸ ano and M. Morisio, “Music mood dataset creation based on last.Fm tags,” in Proc. 2017 International Conference on Artificial Intelligence and Applications, Vienna, Austria, 2017, pp. 15-26, DOI:10.5121/csit.2017.70603. [31] R.E. Thayer: The Biopsychology of Mood and Arousal, Oxford University Press, 1989. [32] J. Russell, “A Circumplex Model of Affect,” Journal of Personality and Social Psychology, vol. 39, pp. 1161-1178, Dec. 1980, doi: 10.1037/h0077714. [33] Social music service - Last.fm. Available: https://www.last.fm/. Accessed: November 2024. [34] Genius - Song Lyrics & Knowledge. Available: https://genius.com/. Accessed: November 2024. [35] YouTube. Available: https://www.youtube.com. Accessed: November 2024. [36] M. Sakowicz and J. Tobolewski, “Development and study of an algorithm for the automatic labeling of musical pieces in the context of emotion A BIMODAL DEEP MODEL TO . . . evoked,” M.Sc. thesis, Gdansk University of Technology and Universitat Polit` ecnica de Catalunya (co-supervised by B. Kostek and J. Turmo), 2023. [37] Genius and Spotify partnering. Available: https://genius.com/a/genius-and-spotify-together. Accessed: November 2024. [38] Pafy library. Available: https://pypi.org/project/pafy/. Accessed: November 2024. [39] Moviepy library. Available: https://pypi.org/project/moviepy/. Accessed: November 2024. [40] M. Honnibal and I. Montani, “spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing,” 2017. Available: https://github.com/explosion/spaCy. Accessed: November 2024. [41] P. N. Johnson-Laird and K. Oatley, “Emotions, Simulation, and Abstract Art,” Art & Perception, vol. 9, no. 3, pp. 260-292, 2021, DOI: https://doi.org/10.1163/22134913-bja10029. [42] P. N. Johnson-Laird and K. Oatley, “How poetry evokes emotions,” Acta Psychologica, vol. 224, p. 103506, 2022, doi: https://doi.org/10.1016/j.actpsy.2022.103506. [43] J. Pennington, R. Socher, and C. Manning, “GloVe: Global Vectors for Word Representation,” in Proc. 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Oct. 2014, pp. 1532-1543, doi: 10.3115/v1/D14-1162. [44] SpaCy - pre-trained pipeline for English. Available: https://spacy.io/models/en\#en_ core_web_lg . Accessed: November 2024. [45] S. Loria, “Textblob Documentation,” Release 0.15, vol. 2, 2018. Available: https://textblob.readthedocs.io/en/dev/. Accessed: November 2024. [46] F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, no. 85, pp. 2825-2830, 2011. Available: http://jmlr.org/papers/v12/pedregosa11a.html. Accessed: November 2024. [47] ”Paradise City” Guns N’ Roses https://genius.com/Guns-n-roses-paradise-citylyrics [48] FastText - text classification tutorial. Available: https://fasttext.cc/docs/en/supervisedtutorial.html. Accessed: November 2024. [49] T. Wolf et al., “Transformers: State-of-the-Art Natural Language Processing,” Jan. 2020, pp. 38-45, doi: 10.18653/v1/2020.emnlp-demos.6. [50] XLNet (base-sized model). Available: https://huggingface.co/xlnet-base-cased. Accessed: November 2024. [51] Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” Advances in neural information processing systems, vol. 32, 2019. https://doi.org/10.48550/arXiv.1906.08237 [52] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” in Proc. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 2818-2826 doi: 10.1109/CVPR.2016.308. [53] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proc. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 770-778, doi: 10.1109/CVPR.2016.90. [54] K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. 3rd International Conference on Learning Representations(ICLR 2015), 2015, pp. 1-14. https://doi.org/10.48550/arXiv.1409.1556 [55] Librosa library. Available: https://librosa.org/. Accessed: November 2024. [56] Chollet, F. et al., 2015. Keras. Available: https://github.com/fchollet/keras. Accessed: November 2024. [57] TensorFlow library. Available: https://www.tensorflow.org/?hl=pl. Accessed: November 2024. [58] S. C. Huang, A. Pareek, S. Seyyedi, I. Banerjee, and M. Lungren, “Fusion of medical imaging and electronic health records using deep learning: a systematic review and implementation guidelines,” npj Digital Medicine, vol. 3, 12, 2020. https://doi.org/10.1038/s41746-020-00341-z [59] A. Paszke et al., 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems, 32. Curran Associates, Inc., pp. 80248035. [60] Combining two deep learning models. Available: https://control.com/technical-articles/combiningtwo-deep-learning-models/. Accessed: November 2024. 238 Jan Tobolewski, Michał Sakowicz, Jordi Turmo, Bo˙ zena Kostek [61] Y. Shi, A. Karatzoglou, L. Baltrunas, M. Larson, A.Hanjalic, and N. Oliver, “TFMAP: Optimizing MAP for top-n context-aware recommendation,” in Proc. 35th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 155-164, Portland Oregon USA, August 2012, doi: 10.1145/2348283.2348308. [62] K. Pyrovolakis, P.K. Tzouveli, and G. Stamou, Multi-Modal Song Mood Detection with Deep Learning. Sensors (Basel, Switzerland), 22, 2022, doi:10.3390/s22031065 [63] E. N. Shaday, V. J. L. Engel, and H. Heryanto, ”Application of the Bidirectional Long ShortTerm Memory Method with Comparison of Word2Vec, GloVe, and FastText for Emotion Classification in Song Lyrics”, Procedia Computer Science, vol. 245, pp. 137-146, 2024, https://doi.org/10.1016/j.procs.2024.10.237 Jan Tobolewski completed his studies at the Faculty of Electronics, Telecommunications and Informatics at the Gdańsk University of Technology (GUT). He organized his student exchange as part of the Erasmus+ programme at UPC Barcelona, which resulted in collaboration with Prof Turmo during his Master’s degree. His master’s thesis explored the complex fi eld of music and emotions, focusing on the development and evaluation of an algorithm for the automatic classifi cation of lyrics. He is an Industrial Ph.D. candidate at the Gdańsk University of Technology, working as a researcher at Asseco Solutions. His scientifi c drive is to uncover the complexity of Large Language Models hallucinations. https://orcid.org/0009-0005-1934-6217 Michal Sakowicz earned his Master of Science in Artifi cial Intelligence at Gdansk University of Technology, focusing his studies on emotion-driven music analysis, a subject he explored extensively in his master’s thesis and also presented at the KES 2023 conference in Athens. Alongside his academic pursuits, he contributed to the IAESTE organization, collaborating with local companies to facilitate international student internships and workshops. Currently, as a Software Developer at Wunderman Thompson Technology, he applies his skills in Java, Docker, and Kotlin to develop and optimize internal systems for fi nancial control. Outside of work, Michal is passionate about AI capabilities and music and especially enjoys winter sports. https://orcid.org/ 0009-0009-2019-2074 Jordi Turmo is a professor in the Computer Science Department at the Universitat Politècnica de Catalunya (UPC), Spain. He is a member of the Center for Language and Speech Technologies and Applications (TALP) and the Center for Intelligent Data Science and Artifi cial Intelligence Research (IDEAI). His research focuses on the fi eld of Natural Language Processing and especially on the study of machine learning methods for Advanced and Adaptive Text Mining, including information extraction, question answering, document clustering, or text classifi cation of standard and non-standard text such as speech transcripts, tweets, or informal notes. He has participated in 7 European projects and 11 Spanish projects, being the local coordinator of some of them, as well as several projects for the transfer of technology to companies. Prof. Turmo has published more than 90 papers in journals and proceedings of international conferences and has supervised several master’s and doctoral theses. He has also been a program committee member of various international conferences, workshops, and shared tasks, leading some of them. https://orcid.org/0000-0002-7521-1115 Bozena Kostek is a professor in the Faculty of Electronics, Telecommunications and Informatics at the Gdansk University of Technology (GUT), Poland. She is a corresponding member of the Polish Academy of Sciences and a fellow of the Audio Engineering Society and the Acoustical Society of America. Her main scientifi c interests are acoustics, psychoacoustics, multimedia, music information retrieval, cognitive and behavioral processing, as well as applications of machine learning to the mentioned domains. Prof. Kostek has presented more than 600 scientifi c papers for journals and at international conferences. She has also published three books related to multimedia applications. She is the recipient of many prestigious awards for research, including those of the Prime Minister of Poland (twice), the Ministry of Science, and the Polish Academy of Sciences. She has supervised more than 300 master’s and eng. works and 25 doctoral theses. She has also led a number of research projects. She was the editor-in-chief of the Journal of the Audio Engineering Society and Archives of Acoustics, as well as Associate Editor of IEEE/ACM TASLP and Guest Editor of JASA and JIIS. https://orcid.org/0000-0001-6288-2908