scieee Open visual document viewer

MERGE App: A Prototype Software for Multi-User Emotion-Aware Music Management

Louro, Pedro; Branco, Guilherme; Redinho, Hugo; Santos, Ricardo Correia Nascimento Dos; Malheiro, Ricardo; Panda, Renato; Paiva, Rui Pedro

Abstract

We present a prototype software for multi-user music library management using the perceived emotional content of songs. The tool offers music playback features, song filtering by metadata, and automatic emotion prediction based on arousal and valence, with the possibility of personalizing the predictions by allowing each user to edit these values based on their own emotion assessment. This is an important feature for handling both classification errors and subjectivity issues, which are inherent aspects of emotion perception. A path-based playlist generation function is also implemented. A multi-modal audio-lyrics regression methodology is proposed for emotion prediction, with accompanying validation experiments on the MERGE dataset. The results obtained are promising, showing higher overall performance on train-validate-test splits (73.20% F1-score with the best dataset/split combination).

Full text

MERGE App: A P o o ype So wa e o Mul i-use Emo ion-awa e Music Managemen Ped o Lima Lou o1 a, Guilhe me B anco1 b, Hugo Redinho1 c, Rica do Co eia1 d, Rica do Malhei o1,2 e, Rena o Panda1,3 , Rui Ped o Pai a1 g 1Uni e si y o Coimb a, Cen e o In o ma ics and Sys ems o he Uni e si y o Coimb a (CISUC), Depa men o In o ma ics Enginee ing, and LASI 2Poly echnic Ins i u e o Lei ia School o Technology and Managemen 3Ci2 — Sma Ci ies Resea ch Cen e , Poly echnic Ins i u e o Toma ped olou [email p o ec ed], guilhe me[email p o ec ed], [email p o ec ed], { ica doco eia, smal, panda, uiped o}@dei.uc.p Keywo ds: Music In o ma ion Re ie al; Music Emo ion Recogni ion; Machine Lea ning; Deep Lea ning; So wa e Abs ac : We p esen a p o o ype so wa e o mul i-use music lib a y managemen using he pe cei ed emo ional con en o songs. The ool o e s music playback ea u es, song il e ing by me ada a, and au oma ic emo ion p edic ion based on a ousal and alence, wi h he possibili y o pe sonalizing he p edic ions by allowing each use o edi hese alues based on hei own emo ion assessmen . This is an impo an ea u e o handling bo h classi ica ion e o s and subjec i i y issues, which a e inhe en aspec s o emo ion pe cep ion. A pa h- based playlis gene a ion unc ion is also implemen ed. A mul i-modal audio-ly ics eg ession me hodology is p oposed o emo ion p edic ion, wi h accompanying alida ion expe imen s on he MERGE da ase . The esul s ob ained a e p omising, showing highe o e all pe o mance on ain- alida e- es spli s (73.20% F1- sco e wi h he bes da ase /spli combina ion). 1 INTRODUCTION The digi al e a has b ough an unp eceden ed amoun o music igh a ou inge ips h ough digi al ma - ke places and s eaming se ices. Wi h he sudden a ailabili y o millions o songs o use s, he neces- si y o au oma ically o ganize and ind ele an mu- sic eme ged. Cu en ecommenda ion sys ems p o- ide pe sonalized sugges ions o use s based on lis- ening pa e ns and using ags, such as gen e, s yle, e c. Howe e , op ions a e lacking when we conside ecommenda ions based on he au oma ic analysis o he emo ional con en o songs. The ield o Music Emo ion Recogni ion (MER) has seen conside able ad ances in ecen yea s in e ms o he mo e classical app oaches. Panda e al. ah ps://o cid.o g/0000-0003-3201-6990 bh ps://o cid.o g/0000-0003-4073-1716 ch ps://o cid.o g/0009-0004-1547-2251 dh ps://o cid.o g/0000-0001-5663-7228 eh ps://o cid.o g/0000-0002-3010-2732 h ps://o cid.o g/0000-0003-2539-5590 gh ps://o cid.o g/0000-0003-3215-3960 (2020) p oposed a new se o ea u es ha consid- e ably inc eased he pe o mance o hese sys ems, achie ing a 76.4% F1-sco e wi h he op 100 anked ea u es. Al hough he ea u e e alua ion is limi ed o one da ase , he imp o emen s a e signi ican com- pa ed o he bes esul s om simila sys ems ha eached a glass ceiling Hu e al. (2008). One d awback o audio-only me hodologies is hei sho comings when di e en ia ing alence. Va - ious sys ems ha e been p oposed using a bimodal app oach le e aging bo h audio and ly ics, a aining conside able imp o emen s when compa ed o sys- ems using only one o he o he Delbouys e al. (2018); Py o olakis e al. (2022). Such sys ems ha e also implemen ed Deep Lea ning (DL) a chi ec u es o skip he ime-consuming ea u e enginee ing and ex ac ion s eps om he classical sys ems and con- side ably speed up he in e ence p ocess o he o e all sys em. In his s udy, we p esen he MERGE1applica ion, 1MERGE is he ac onym o ”Music Emo ion Recog- ni ion nEx Gene a ion”, a esea ch p ojec unded by he Po uguese Science Founda ion. Figu e 1: MERGE applica ion in e ace. The AV plo , alongside music playback and song display con ols, is seen in yellow. The able iew is highligh ed in blue wi h ed highligh il e ing sea ch ba and bu on o adding music. Finally, g een highligh s he bu ons o applica ion in o ma ion and use logou . which au oma ically p edic s he a ousal and alence o songs based on Russell’s Ci cumplex Model Rus- sell (1980). Two axes make up his model: a ousal (Y-axis), which depic s whe he he song has high o low ene gy, and alence (X-axis), which ep esen s whe he he emo ion o he song has a nega i e o pos- i i e conno a ion. The in eg a ed model used o p edic ion is also p esen ed in his s udy alongside alida ion expe i- men s, which ecei ed bo h audio and ly ics in o ma- ion o map he song mo e accu a ely in o he abo e men ioned model. The MERGE applica ion is a ollow-up o he MOODe ec o applica ion, p e iously c ea ed by ou eam Ca doso e al. (2011). The new MERGE app was buil om sc a ch, wi h signi ican code e ac o - ing and op imiza ion, while keeping he o e all use in e ace o he MOODe ec o app. In addi ion, sig- ni ican no el ea u es and imp o emen s we e im- plemen ed, namely: i) a bimodal app, which ex- ploi s he combina ion o audio and ly ics da a o im- p o ed classi ica ion (unlike he single audio modali y in he MOODe ec o app); ii) an imp o ed classi ica- ion model, aining wi h he MERGE da ase Lou o e al. (2024b) and ollowing a deep lea ning app oach Lou o e al. (2024a); iii) and a shi om he mono- li hic single-use pa adigm o he web-based mul i- use pa adigm. 2 MERGE APPLICATION The MERGE applica ion is implemen ed using Ja aSc ip , wi h he addi ion o he jQue y lib a y o handle AJAX, o i s on end, while he backend is se ed using he Exp ess lib a y on op o Node.js. The applica ion in e ace is depic ed in Figu e 1. 2.1 Applica ion O e iew The componen s can be b oken down as: i) he Rus- sell’s Ci cumplex model whe e all songs can be seen (highligh ed in yellow); ii) a able iew o he songs wi h a ious op ions o so ing (highligh ed in blue); iii) op ions o il e and add new songs (seen in he ed egion, om le o igh ); i ) in o ma ion abou he applica ion, he cu en use , and an op ion o lo- gou (highligh ed in g een, also om le o igh ). Songs a e placed in he plo desc ibed in i) acco d- ing o he es ima ed a ousal and alence (AV) alues in an in e al o [-1, 1] o each axis. The p ocess o ob aining hese alues is desc ibed in Sec ion 3. Beyond he AV posi ioning, each poin on he plane is also colo -coded depending on he quad an : g een (happy), ed ( ense), blue (sad), and yellow ( elaxed). The iew o he g aph can be swi ched be ween a) ”Uploaded” o show only songs uploaded by he cu - en use , b) ”My Lib a y” o display a use ’s lib a y, i.e., he songs uploaded by he cu en use , plus songs uploaded by o he use s added by he cu en use , and c) ”All songs” o show all he songs a ailable in he da abase. A no e ega ding he la e op ion is he di - e en ia ion o songs no added by he use , appea ing as g ey do s in he plo , also depic ed in Figu e 1. Each use can change he song’s posi ion di ec ly by mo ing he poin in he plane iew, o by edi ing he AV alues h ough he able iew. These new AV alues a e unique o he use . The applica ion can be used as an audio playback so wa e, hus o e ing usual ea u es such as: • Playback con ols o mp3 iles (play, pause, seek); • Volume con ols, including mu e; • Double-clicking a song o be played ei he in he plo o able iew; Figu e 2: MERGE App Backo ice. Use s wi h adminis a i e p i ileges can o e w i e he model used o AV alues’ p edic- ion (highligh ed in g een) and expo a CSV ile wi h in o ma ion ega ding he anno a ions o all use s o each song in he da abase (highligh ed in blue). Figu e 3: En i y ela ion diag am o he applica ion’s da abase. • Fil e ing and so ing by any o he a ailable song p ope ies ( i le, a is , alence, a ousal, emo ion); • Adding and dele ing songs om he use ’s lib a y. The applica ion also p o ides a backo ice o use s wi h adminis a i e p i ileges, pic u ed in Fig- u e 2. A e logging in, he use can pe o m one o wo ac ions: upload and deploy a new model o AV p edic ion, and expo a CSV wi h he exis ing use AV anno a ions. The la e op ion is designed o eas- ily e ie e each use ’s a ailable anno a ions o songs in hei espec i e lib a ies. In his way, he MERGE app can be used as a c owdsou cing da a collec ion and anno a ion ool, p omo ing he c ea ion o size- able and quali y MER da ase s, a cu en key need in MER esea ch Panda e al. (2020). Rega ding he da abase used o s o e all he el- e an da a om use s and songs, he co esponding en i y ela ionship diag am is depic ed in Figu e 3. The use able s o es he use ’s pe sonal in o ma ion, as well as he use ole, o access p i ileges pu - poses. The pa h o he model used o AV p edic- ion is sa ed in he co esponding able and iden i ies he use ha uploaded he cu en ly deployed model. The song able s o es all song- ela ed in o ma ion, in- cluding me ada a, he AV alues i s p edic ed by he p esen ly deployed model, he mapped quad an in he plo iew, he pa h o he uploaded audio clip, and he e e ence o he use who i s added he song. The use -speci ic anno a ions a e s o ed in he an- no a ion able, which s o es he AV alues de ined pe he use ’s pe cep ion, he co esponding quad an , and a e e ence o he song and use o ha anno a ion. Finally, he lib a y able s o es all use lib a ies, con- aining only songs added by he use , ei he h ough uploading o om o he use s’ lib a ies. Figu e 4: A he op, ”Ano he One Bi es The Dus ” by Queen is added o he lib a y and placed in he plane acco ding o he p edic ed AV alues. A he bo om, he poin ep esen ing he song is mo ed o a mo e accu a e posi ion, acco ding o he use . 2.2 Building An Emo ionally-awa e Lib a y A e adding a new song om an a ailable MP3 ile, AV alues a e au oma ically p edic ed, and a poin is added o he plo alongside a new en y on he able. Should he use disag ee wi h he p edic ed alues, hese can be easily changed by edi ing he en y om he able iew o mo ing he poin in he plo . An ex- ample o he ini ially p edic ed posi ion o a newly added song and he inal posi ion a e adjus men can be seen in Figu e 4. This pe sonaliza ion mechanism p o ides use s wi h he abili y o add ess he in in- sic subjec i i y in MER. Howe e , ackling his issue con inues o p esen a signi ican challenge. 2.3 Pa h-based Playlis Gene a ion The abili y o gene a e a playlis based on a use - d awn pa h is cu en ly implemen ed, as depic ed in Figu e 5. This ea u e allows use s o eely c ea e an emo ionally- a ying playlis . This is done by compu ing he dis ance o he use - de ined Ncloses songs o he e e ence poin s ha make up he d awn pa h. The use may also con igu e how a he songs can be om he pa h o be consid- e ed in o he calcula ions. This h eshold is de ined as a decimal numbe be ween he plane in e al ([-1, 1]). 3 SONG EMOTION PREDICTION In his sec ion, we discuss he me hodology used o p edic AV alues o a gi en song. Fi s , he DL model’s a chi ec u e is p esen ed, ollowed by he p e- p ocessing s eps o each modali y, a desc ip ion o he op imiza ion used, and he e alua ion conduc ed. 3.1 Model A chi ec u e The p oposed a chi ec u e, depic ed in Figu e 6, is based on he one by Delbouys e al. Delbouys e al. (2018). Dis inc audio and ly ics b anches ecei e Mel-spec og am ep esen a ions and wo d embed- dings, espec i ely. The lea ned ea u es o each modali y a e hen used and u he p ocessed by a small Dense Neu al Ne wo k (DNN), inally ou - pu ing he AV alues p edic ion. We op ed o a bimodal audio-ly ics app oach conside ing ha bo h modali ies ha e ele an in o - ma ion o he di e en axes o Russell’s Ci cum- plex Model. Audio has been shown o be e p edic Figu e 5: The use can be seen d awing a pa h o gene a e a playlis wi h he desi ed emo ional ajec o y a he op. The esul o he pa h-based playlis gene a ion is p esen ed a he bo om. a ousal, while ly ical in o ma ion is mo e ele an o alence p edic ion Lou o e al. (2024b). S a ing in he audio b anch, Mel-spec og am ep esen a ions o each sample a e ed o he ea u e lea ning po ion o he baseline a chi ec u e p esen ed in Lou o e al. Lou o e al. (2024a). I is composed o ou con olu ional blocks, composed o a 2D Con- olu ional laye , ollowed by a Ba ch No maliza ion, D opou , and Max Pooling laye , inishing wi h ReLU ac i a ion. As o he ly ics b anch, he wo d em- beddings o ly ics a e also ed o ou con olu ional blocks, each comp ising a 1D Con olu ional laye , ollowed by Max Pooling and a ReLU ac i a ion lay- e s. To balance he in o ma ion om each modali y, we signi ican ly educe he o e whelming amoun o lea ned ea u es om ly ics using a Dense laye be- o e me ging he lea ned ea u es om bo h b anches. The classi ica ion po ion o he model is com- posed o al e na ing D oupou and Dense laye s, which educe and u he p ocess he se o ea u es espec i ely, inally ou pu ing one o Russell’s Ci - cumplex model’s ou quad an s. 3.2 P e-p ocessing S eps A se o p e-p ocessing s eps is necessa y o ob ain he da a ep esen a ions used o each b anch o he a chi ec u e de ailed abo e. The lib osa lib a y McFee e al. (2015) is used o ob ain he Mel-spec og am ep esen a ion o he au- dio b anch. The audio samples, p o ided as mp3 iles, a e i s con e ed o wa e o ms (.wa ) and down- sampled om 22.5 o 16kHz. This is done o educe he complexi y o he model, along wi h he compu- a ional cos o op imiza ion. The downsampling has been shown o p o ide simila esul s o highe sam- pling a es, showing he obus ness o DL app oaches Py o olakis e al. (2022). The spec al ep esen a ions a e hen gene a ed using de aul pa ame e s o he leng h o he Fas Fou ie T ans o m window (2048) as well as he hop size (512). As o wo d embeddings, he Sen ence T ans- o me lib a y om Hugging Face was used, speci i- cally, he all- obe a-la ge- 1 p e- ained model. The embedde ecei es a con ex o up o 512 okens and ou pu s a 1024 embedded ec o . Gi en ha he bes esul s we e p o ided by using he ull con ex win- dow, some o he ly ics had o be cu o a some poin . A e some simple okeniza ion s eps, namely emo ing new line cha ac e s and con e ing all ex o lowe case, he embeddings we e ob ained up o he al eady men ioned con ex size. Figu e 6: The mul i-modal audio-ly ics eg ession model. Emo ionally- ele an ea u es a e lea ned o bo h he audio ep e- sen a ion in he Mel-spec og am- ecei ing b anch and he ly ics ep esen a ion in he b anch ecei ing he p e iously gene - a ed wo d embeddings. AV alues a e p edic ed a e conca ena ing and p ocessing he lea ned ea u es om bo h b anches. 3.3 Model Op imiza ion Model op imiza ion was conduc ed using he Bayesian op imiza ion implemen a ion o he Ke as Tune lib a y O’Malley e al. (2019). This me hod inds he bes combina ion o hype pa ame e s in p e- iously de ined in e als o each, ei he maximizing o minimizing an objec i e unc ion de ined by he use . Since ou me hodology is based on a eg ession ask o p edic a ousal and alence o a gi en sam- ple, he objec i e is de ined as minimizing he sum o he mean squa ed e o (MSE) o bo h. This en- su es ha none is p io i ized, le e aging bo h audio’s be e p edic abili y in e ms o a ousal and he same o ly ics’ p edic abili y o alence. The in e als o each conside ed hype pa ame e , namely, ba ch size, op imize , and co esponding lea ning a e, a e p e- sen ed in Table 1. Table 1: Op imal Hype pa ame e s Fo Each Da ase Bes Hype pa ame e s Ba ch Size Op mize Lea ning Ra e 64 SGD 1e-2 The op imiza ion p ocess is un o e en ials, pe he lib a y’s de aul , s a ing a he lowe end o each in e al. Fo each ial, he model is ained o a max- imum o 200 epochs, wi h an ea ly s opping s a egy de ined o check o no imp o emen s o he alida- ion loss o 15 consecu i e epochs. This conside ably educes he ime needed o conduc he ull op imiza- ion phase since less ime is spen on unde pe o ming se s o hype pa ame e s. We used a 70-15-15 ain- alida e- es (TVT) spli as ou alida ion s a egy, as de ined in Lou o e al. (2024b). The esul ing models o each ial a e backed up o la e usage, including he e alua ion phase, which is discussed nex . 3.4 Da a and E alua ion The MERGE Bimodal Comple e da ase was used o alida ing ou app oach. P oposed in Lou o e al. (2024b), i comp ises a se o 2216 bimodal samples (audio clips and co esponding ly ics). Fo each sam- ple, he da ase p o ides a 30-second audio exce p o he mos ep esen a i e pa o he song, links o he ull ly ics, labels co esponding o each o he quad- an s in Russell’s Ci cumplex model, and AV alues, used o ob ain he p e iously men ioned labels, cal- Table 2: TVT 70-15-15 Resul s Fo MERGE Audio Comple e F1-sco e P ecision Recall R2 RMSE (A/V) (A/V) 73.20% 74.53% 73.49% 0.454 0.133 0.506 0.339 cula ed based on he ex ac ed emo ion- ela ed ags a ailable in AllMusic 2. The abo e-men ioned AV alues a e ob ained h ough he ollowing p ocess. Fi s , he a ailable ags o each song in he da ase a e ob ained om he All Music pla o m. Using Wa ine ’s Adjec i e Dic iona y Wa ine e al. (2013), he exis ing ags a e ansla ed o a ousal and alence alues. Finally, The alues a e hen a e aged ac oss all ags co espond- ing o a speci ic song, ob aining i s inal mapping on Russell’s Ci cumplex model. Fo he TVT s a egy, bo h he aining and ali- da ion se s a e used in he op imiza ion unc ion. The se o op imal hype pa ame e s is ound using he la - e . A e aining he model o each da ase , he ol- lowing me ics a e compu ed be ween he ac ual and p edic ed AV alues in he es se o each class as well as o he o e all pe o mance: F1-sco e, P eci- sion, Recall, R2(squa ed Pea son’s co ela ion), and Roo Mean Squa ed E o (RMSE). Be o e compu ing hese me ics, he p edic ed and eal AV alues we e mapped o Russell’s Ci cumplex model o ob ain classes o calcula ing P ecision, Re- call, and F1-sco e. 4 EXPERIMENTAL RESULTS AND DISCUSSION Tables 2 and 3 show he o e all esul s o he dis- cussed me hodology. The a ousal and alence s an- dalone esul s o he R2and RMSE me ics a e p e- sen ed in consecu i e lines in he o de displayed in he ables. The ob ained esul s o bo h da ase s a e lowe han hose ob ained in p e ious s udies ocused on s a ic MER as a ca ego ical p oblem Lou o e al. (2024a). The bes esul a ained is a 73.20% F1- sco e, which is a ound 6% lowe han he esul s ob- ained o he same da ase and e alua ion s a egy in he men ioned a icle. The lowe esul s a e mos ly due o he semi-au oma ic app oach o ob ain AV al- ues (see Sec ion 3.4, conside ing ha he ags a ail- able on All Music a e use -gene a ed and i s cu a ion is unknown. 2h ps://www.allmusic.com/ Table 3: TVT 70-15-15 Resul s Con usion Ma ix Fo MERGE Audio Comple e P edic ed Q1 Q2 Q3 Q4 Ac ual Q1 61.3% 10.4% 6.6% 21.7% Q2 9.8% 82.4% 5.9% 2.0% Q3 1.4% 4.3% 78.3% 15.9% Q4 7.3% 0.0% 18.2% 74.5% As shown in Table 2, he R2me ic o alence ou pe o med he one o a ousal, al hough ha ing a la ge RMSE. This indica es ha he ela i e alence h oughou songs is easonably cap u ed, despi e he la ge RMSE e o . Al hough he a ained esul s show oom o im- p o emen , hey a e a good s a ing poin o he use . Gi en he subjec i e na u e o each use ’s emo ional pe cep ion, we belie e ha he pe sonaliza ion ea u e included in he MERGE app is a aluable mechanism o handling subjec i i y in MER. In e ms o he esul s o sepa a e quad an s (Ta- ble 3), we can see ha some Q1 songs a e con used wi h Q4 songs (21.77% Q1 songs a e inco ec ly clas- si ied as Q4). Mo eo e , he e is also some con usion be ween Q3 and Q4 (15.9% o Q3 songs a e p edic ed as Q4 and 18.2% o Q4 songs a e classi ied as Q3). This is a known di icul y in MER, as discussed in Panda e al. (2020) ha needs u he esea ch. 5 CONCLUSION AND FUTURE WORK We p esen ed he p o o ype o he MERGE applica- ion. Cu en ly, he ini ial e sion has implemen ed music playback ea u es, he abili y o add and il e songs o a sha ed da abase, lis and plane iews, he la e based on Russell’s Ci cumplex model, and use managemen unc ionali ies. Mo eo e , a bimodal audio-ly ics model is inco po a ed in o he backend o he p o o ype o allow o AV alue p edic ion o use -uploaded songs. Pa h-based playlis gene a ion has also been implemen ed, enabling use s o c a a playlis ha ollows a speci ic emo ional ajec o y hey ha e selec ed. S ill, many mo e unc ionali ies a e planned o he applica ion in u u e i e a ions. The highligh ed unc ionali ies include use -gene a ed ags o a mo e cus omized il e ing expe ience ha would be a ail- able o o he use s; au oma ic ly ics o he ull song sc aped om an a ailable API, e.g., Genius; and Mu- sic Emo ion Va ia ion De ec ion (MEVD) p edic ion suppo , including isualiza ion wi h he same colo code used in he plo . A s andalone desk op applica- ion is also planned wi hou he c oss-use ea u es. in addi ion o implemen ing hese upcoming ea u es, We plan o conduc in-dep h use expe ience s udies o gain a mo e comp ehensi e unde s anding o he sys em’s e icacy and use sa is ac ion. Valida ion expe imen s on wo ecen ly p oposed da ase s a e p o ided alongside a ho ough sys em de- sc ip ion, elaying insigh s in o he ob ained esul s. These a e s ill below he ca ego ical app oach p e- sen ed in Lou o e al. (2024b) due o he al eady dis- cussed semi-au oma ic AV mapping app oach in Sec- ion 3.4. Despi e his, he p edic ions a e a good s a - ing poin o be u he adjus ed o he use ’s pe cep- ion. Rega ding he ac ual model, nei he ea u e lea n- ing po ion may be ideal o he p oblem a hand since hey we e o iginally de eloped o a ca ego i- cal p oblem. De eloping mo e sui able a chi ec u es should hus be conside ed u u e wo k. Fu he mo e, he da a ep esen a ions, especially he wo d embed- dings, may also be u he imp o ed, conside ing ha he p e- ained model used is limi ed o a con ex win- dow o 512 okens. To conclude, we belie e he p oposed app migh be use ul o music lis ene s. Al hough he e is oom o imp o emen (as he a ained classi ica ion esul s show), he pe sonaliza ion mechanism is a use ul ea- u e o handling p edic ion e o s and subjec i i y. Finally, he pe sonaliza ion ea u e and he mul i-use en i onmen ha e he po en ial o acqui e quali y use anno a ions, leading o a u u e la ge and mo e obus MER da ase . ACKNOWLEDGEMENTS This wo k is unded by FCT - Founda ion o Sci- ence and Technology, I.P., wi hin he scope o he p ojec s: MERGE - DOI: 10.54499/PTDC/CCI- COM/3171/2021 inanced wi h na ional unds (PID- DAC) ia he Po uguese S a e Budge ; and p ojec CISUC - UID/CEC/00326/2020 wi h unds om he Eu opean Social Fund, h ough he Regional Ope a- ional P og am Cen o 2020. Rena o Panda was sup- po ed by Ci2 - FCT UIDP/05567/2020. We hank all e iewe s o hei aluable sugges- ions, which help o imp o e he a icle. REFERENCES Ca doso, L., Panda, R., and Pai a, R. P. (2011). Moode ec- o : A p o o ype so wa e ool o mood-based playlis gene a ion. In Simp´ osio de In o m´ a ica - INFo um 2011, Coimb a, Po ugal. Delbouys, R., Hennequin, R., Piccoli, F., Royo-Le elie , J., and Moussallam, M. (2018). Music Mood De ec ion Based On Audio And Ly ics Wi h Deep Neu al Ne . In P oceedings o he 19 h In e na ional Socie y o Music In o ma ion Re ie al Con e ence, pages 370–375, Pa is, F ance. Hu, X., Downie, J. S., Lau ie , C., Bay, M., and Ehmann, A. F. (2008). The 2007 Mi ex Audio Mood Classi ica ion Task: Lessons Lea ned. In P oceedings o he 9 h In e - na ional Socie y o Music In o ma ion Re ie al Con- e ence, pages 462–467, D exel Uni e si y, Philadelphia, Pennsyl ania, USA. Lou o, P. L., Redinho, H., Malhei o, R., Pai a, R. P., and Panda, R. (2024a). A Compa ison S udy o Deep Lea n- ing Me hodologies o Music Emo ion Recogni ion. Sen- so s, 24(7):2201. Lou o, P. L., Redinho, H., San os, R., Malhei o, R., Panda, R., and Pai a, R. P. (2024b). MERGE – A Bimodal Da ase o S a ic Music Emo ion Recogni ion. McFee, B., Ra el, C., Liang, D., Ellis, D., McVica , M., Ba enbe g, E., and Nie o, O. (2015). Lib osa: Audio and Music Signal Analysis in Py hon. In Py hon in Science Con e ence, pages 18–24, Aus in, Texas. O’Malley, T., Bu sz ein, E., Long, J., Cholle , F., Jin, H., In e nizzi, L., e al. (2019). Ke as Tune . h ps:// gi hub.com/ke as- eam/ke as- une . Panda, R., Malhei o, R., and Pai a, R. P. (2020). No el Audio Fea u es o Music Emo ion Recogni ion. IEEE T ansac ions on A ec i e Compu ing, 11(4):614–626. Py o olakis, K., Tzou eli, P., and S amou, G. (2022). Mul i-Modal Song Mood De ec ion wi h Deep Lea ning. Senso s, 22(3):1065. Russell, J. A. (1980). A ci cumplex model o a ec . Jou nal o Pe sonali y and Social Psychology, 39(6):1161–1178. Wa ine , A. B., Kupe man, V., and B ysbae , M. (2013). No ms o alence, a ousal, and dominance o 13,915 English lemmas. Beha io Resea ch Me hods, 45(4):1191–1207.