Full text
MERGE App: A P o o ype So wa e o Mul i-use Emo ion-awa e
Music Managemen
Ped o Lima Lou o1 a, Guilhe me B anco1 b, Hugo Redinho1 c, Rica do Co eia1 d,
Rica do Malhei o1,2 e, Rena o Panda1,3 , Rui Ped o Pai a1 g
1Uni e si y o Coimb a, Cen e o In o ma ics and Sys ems o he Uni e si y o Coimb a (CISUC), Depa men o
In o ma ics Enginee ing, and LASI
2Poly echnic Ins i u e o Lei ia School o Technology and Managemen
3Ci2 — Sma Ci ies Resea ch Cen e , Poly echnic Ins i u e o Toma
ped olou [email p o ec ed], guilhe me[email p o ec ed], [email p o ec ed], { ica doco eia, smal, panda,
uiped o}@dei.uc.p
Keywo ds: Music In o ma ion Re ie al; Music Emo ion Recogni ion; Machine Lea ning; Deep Lea ning; So wa e
Abs ac : We p esen a p o o ype so wa e o mul i-use music lib a y managemen using he pe cei ed emo ional
con en o songs. The ool o e s music playback ea u es, song il e ing by me ada a, and au oma ic emo ion
p edic ion based on a ousal and alence, wi h he possibili y o pe sonalizing he p edic ions by allowing each
use o edi hese alues based on hei own emo ion assessmen . This is an impo an ea u e o handling
bo h classi ica ion e o s and subjec i i y issues, which a e inhe en aspec s o emo ion pe cep ion. A pa h-
based playlis gene a ion unc ion is also implemen ed. A mul i-modal audio-ly ics eg ession me hodology
is p oposed o emo ion p edic ion, wi h accompanying alida ion expe imen s on he MERGE da ase . The
esul s ob ained a e p omising, showing highe o e all pe o mance on ain- alida e- es spli s (73.20% F1-
sco e wi h he bes da ase /spli combina ion).
1 INTRODUCTION
The digi al e a has b ough an unp eceden ed amoun
o music igh a ou inge ips h ough digi al ma -
ke places and s eaming se ices. Wi h he sudden
a ailabili y o millions o songs o use s, he neces-
si y o au oma ically o ganize and ind ele an mu-
sic eme ged. Cu en ecommenda ion sys ems p o-
ide pe sonalized sugges ions o use s based on lis-
ening pa e ns and using ags, such as gen e, s yle,
e c. Howe e , op ions a e lacking when we conside
ecommenda ions based on he au oma ic analysis o
he emo ional con en o songs.
The ield o Music Emo ion Recogni ion (MER)
has seen conside able ad ances in ecen yea s in
e ms o he mo e classical app oaches. Panda e al.
ah ps://o cid.o g/0000-0003-3201-6990
bh ps://o cid.o g/0000-0003-4073-1716
ch ps://o cid.o g/0009-0004-1547-2251
dh ps://o cid.o g/0000-0001-5663-7228
eh ps://o cid.o g/0000-0002-3010-2732
h ps://o cid.o g/0000-0003-2539-5590
gh ps://o cid.o g/0000-0003-3215-3960
(2020) p oposed a new se o ea u es ha consid-
e ably inc eased he pe o mance o hese sys ems,
achie ing a 76.4% F1-sco e wi h he op 100 anked
ea u es. Al hough he ea u e e alua ion is limi ed
o one da ase , he imp o emen s a e signi ican com-
pa ed o he bes esul s om simila sys ems ha
eached a glass ceiling Hu e al. (2008).
One d awback o audio-only me hodologies is
hei sho comings when di e en ia ing alence. Va -
ious sys ems ha e been p oposed using a bimodal
app oach le e aging bo h audio and ly ics, a aining
conside able imp o emen s when compa ed o sys-
ems using only one o he o he Delbouys e al.
(2018); Py o olakis e al. (2022). Such sys ems ha e
also implemen ed Deep Lea ning (DL) a chi ec u es
o skip he ime-consuming ea u e enginee ing and
ex ac ion s eps om he classical sys ems and con-
side ably speed up he in e ence p ocess o he o e all
sys em.
In his s udy, we p esen he MERGE1applica ion,
1MERGE is he ac onym o ”Music Emo ion Recog-
ni ion nEx Gene a ion”, a esea ch p ojec unded by he
Po uguese Science Founda ion.
Figu e 1: MERGE applica ion in e ace. The AV plo , alongside music playback and song display con ols, is seen in yellow.
The able iew is highligh ed in blue wi h ed highligh il e ing sea ch ba and bu on o adding music. Finally, g een
highligh s he bu ons o applica ion in o ma ion and use logou .
which au oma ically p edic s he a ousal and alence
o songs based on Russell’s Ci cumplex Model Rus-
sell (1980). Two axes make up his model: a ousal
(Y-axis), which depic s whe he he song has high o
low ene gy, and alence (X-axis), which ep esen s
whe he he emo ion o he song has a nega i e o pos-
i i e conno a ion.
The in eg a ed model used o p edic ion is also
p esen ed in his s udy alongside alida ion expe i-
men s, which ecei ed bo h audio and ly ics in o ma-
ion o map he song mo e accu a ely in o he abo e
men ioned model.
The MERGE applica ion is a ollow-up o he
MOODe ec o applica ion, p e iously c ea ed by ou
eam Ca doso e al. (2011). The new MERGE app
was buil om sc a ch, wi h signi ican code e ac o -
ing and op imiza ion, while keeping he o e all use
in e ace o he MOODe ec o app. In addi ion, sig-
ni ican no el ea u es and imp o emen s we e im-
plemen ed, namely: i) a bimodal app, which ex-
ploi s he combina ion o audio and ly ics da a o im-
p o ed classi ica ion (unlike he single audio modali y
in he MOODe ec o app); ii) an imp o ed classi ica-
ion model, aining wi h he MERGE da ase Lou o
e al. (2024b) and ollowing a deep lea ning app oach
Lou o e al. (2024a); iii) and a shi om he mono-
li hic single-use pa adigm o he web-based mul i-
use pa adigm.
2 MERGE APPLICATION
The MERGE applica ion is implemen ed using
Ja aSc ip , wi h he addi ion o he jQue y lib a y o
handle AJAX, o i s on end, while he backend is
se ed using he Exp ess lib a y on op o Node.js.
The applica ion in e ace is depic ed in Figu e 1.
2.1 Applica ion O e iew
The componen s can be b oken down as: i) he Rus-
sell’s Ci cumplex model whe e all songs can be seen
(highligh ed in yellow); ii) a able iew o he songs
wi h a ious op ions o so ing (highligh ed in blue);
iii) op ions o il e and add new songs (seen in he
ed egion, om le o igh ); i ) in o ma ion abou
he applica ion, he cu en use , and an op ion o lo-
gou (highligh ed in g een, also om le o igh ).
Songs a e placed in he plo desc ibed in i) acco d-
ing o he es ima ed a ousal and alence (AV) alues
in an in e al o [-1, 1] o each axis. The p ocess
o ob aining hese alues is desc ibed in Sec ion 3.
Beyond he AV posi ioning, each poin on he plane
is also colo -coded depending on he quad an : g een
(happy), ed ( ense), blue (sad), and yellow ( elaxed).
The iew o he g aph can be swi ched be ween a)
”Uploaded” o show only songs uploaded by he cu -
en use , b) ”My Lib a y” o display a use ’s lib a y,
i.e., he songs uploaded by he cu en use , plus songs
uploaded by o he use s added by he cu en use , and
c) ”All songs” o show all he songs a ailable in he
da abase. A no e ega ding he la e op ion is he di -
e en ia ion o songs no added by he use , appea ing
as g ey do s in he plo , also depic ed in Figu e 1.
Each use can change he song’s posi ion di ec ly
by mo ing he poin in he plane iew, o by edi ing
he AV alues h ough he able iew. These new AV
alues a e unique o he use .
The applica ion can be used as an audio playback
so wa e, hus o e ing usual ea u es such as:
• Playback con ols o mp3 iles (play, pause,
seek);
• Volume con ols, including mu e;
• Double-clicking a song o be played ei he in he
plo o able iew;
Figu e 2: MERGE App Backo ice. Use s wi h adminis a i e p i ileges can o e w i e he model used o AV alues’ p edic-
ion (highligh ed in g een) and expo a CSV ile wi h in o ma ion ega ding he anno a ions o all use s o each song in he
da abase (highligh ed in blue).
Figu e 3: En i y ela ion diag am o he applica ion’s da abase.
• Fil e ing and so ing by any o he a ailable song
p ope ies ( i le, a is , alence, a ousal, emo ion);
• Adding and dele ing songs om he use ’s lib a y.
The applica ion also p o ides a backo ice o
use s wi h adminis a i e p i ileges, pic u ed in Fig-
u e 2. A e logging in, he use can pe o m one o
wo ac ions: upload and deploy a new model o AV
p edic ion, and expo a CSV wi h he exis ing use
AV anno a ions. The la e op ion is designed o eas-
ily e ie e each use ’s a ailable anno a ions o songs
in hei espec i e lib a ies. In his way, he MERGE
app can be used as a c owdsou cing da a collec ion
and anno a ion ool, p omo ing he c ea ion o size-
able and quali y MER da ase s, a cu en key need in
MER esea ch Panda e al. (2020).
Rega ding he da abase used o s o e all he el-
e an da a om use s and songs, he co esponding
en i y ela ionship diag am is depic ed in Figu e 3.
The use able s o es he use ’s pe sonal in o ma ion,
as well as he use ole, o access p i ileges pu -
poses. The pa h o he model used o AV p edic-
ion is sa ed in he co esponding able and iden i ies
he use ha uploaded he cu en ly deployed model.
The song able s o es all song- ela ed in o ma ion, in-
cluding me ada a, he AV alues i s p edic ed by he
p esen ly deployed model, he mapped quad an in he
plo iew, he pa h o he uploaded audio clip, and he
e e ence o he use who i s added he song.
The use -speci ic anno a ions a e s o ed in he an-
no a ion able, which s o es he AV alues de ined pe
he use ’s pe cep ion, he co esponding quad an , and
a e e ence o he song and use o ha anno a ion.
Finally, he lib a y able s o es all use lib a ies, con-
aining only songs added by he use , ei he h ough
uploading o om o he use s’ lib a ies.
Figu e 4: A he op, ”Ano he One Bi es The Dus ” by Queen is added o he lib a y and placed in he plane acco ding o he
p edic ed AV alues. A he bo om, he poin ep esen ing he song is mo ed o a mo e accu a e posi ion, acco ding o he
use .
2.2 Building An Emo ionally-awa e
Lib a y
A e adding a new song om an a ailable MP3 ile,
AV alues a e au oma ically p edic ed, and a poin is
added o he plo alongside a new en y on he able.
Should he use disag ee wi h he p edic ed alues,
hese can be easily changed by edi ing he en y om
he able iew o mo ing he poin in he plo . An ex-
ample o he ini ially p edic ed posi ion o a newly
added song and he inal posi ion a e adjus men can
be seen in Figu e 4. This pe sonaliza ion mechanism
p o ides use s wi h he abili y o add ess he in in-
sic subjec i i y in MER. Howe e , ackling his issue
con inues o p esen a signi ican challenge.
2.3 Pa h-based Playlis Gene a ion
The abili y o gene a e a playlis based on a use -
d awn pa h is cu en ly implemen ed, as depic ed in
Figu e 5. This ea u e allows use s o eely c ea e an
emo ionally- a ying playlis .
This is done by compu ing he dis ance o he use -
de ined Ncloses songs o he e e ence poin s ha
make up he d awn pa h. The use may also con igu e
how a he songs can be om he pa h o be consid-
e ed in o he calcula ions. This h eshold is de ined as
a decimal numbe be ween he plane in e al ([-1, 1]).
3 SONG EMOTION PREDICTION
In his sec ion, we discuss he me hodology used o
p edic AV alues o a gi en song. Fi s , he DL
model’s a chi ec u e is p esen ed, ollowed by he p e-
p ocessing s eps o each modali y, a desc ip ion o
he op imiza ion used, and he e alua ion conduc ed.
3.1 Model A chi ec u e
The p oposed a chi ec u e, depic ed in Figu e 6, is
based on he one by Delbouys e al. Delbouys e al.
(2018). Dis inc audio and ly ics b anches ecei e
Mel-spec og am ep esen a ions and wo d embed-
dings, espec i ely. The lea ned ea u es o each
modali y a e hen used and u he p ocessed by
a small Dense Neu al Ne wo k (DNN), inally ou -
pu ing he AV alues p edic ion.
We op ed o a bimodal audio-ly ics app oach
conside ing ha bo h modali ies ha e ele an in o -
ma ion o he di e en axes o Russell’s Ci cum-
plex Model. Audio has been shown o be e p edic
Figu e 5: The use can be seen d awing a pa h o gene a e a playlis wi h he desi ed emo ional ajec o y a he op. The esul
o he pa h-based playlis gene a ion is p esen ed a he bo om.
a ousal, while ly ical in o ma ion is mo e ele an o
alence p edic ion Lou o e al. (2024b).
S a ing in he audio b anch, Mel-spec og am
ep esen a ions o each sample a e ed o he ea u e
lea ning po ion o he baseline a chi ec u e p esen ed
in Lou o e al. Lou o e al. (2024a). I is composed
o ou con olu ional blocks, composed o a 2D Con-
olu ional laye , ollowed by a Ba ch No maliza ion,
D opou , and Max Pooling laye , inishing wi h ReLU
ac i a ion. As o he ly ics b anch, he wo d em-
beddings o ly ics a e also ed o ou con olu ional
blocks, each comp ising a 1D Con olu ional laye ,
ollowed by Max Pooling and a ReLU ac i a ion lay-
e s. To balance he in o ma ion om each modali y,
we signi ican ly educe he o e whelming amoun o
lea ned ea u es om ly ics using a Dense laye be-
o e me ging he lea ned ea u es om bo h b anches.
The classi ica ion po ion o he model is com-
posed o al e na ing D oupou and Dense laye s,
which educe and u he p ocess he se o ea u es
espec i ely, inally ou pu ing one o Russell’s Ci -
cumplex model’s ou quad an s.
3.2 P e-p ocessing S eps
A se o p e-p ocessing s eps is necessa y o ob ain
he da a ep esen a ions used o each b anch o he
a chi ec u e de ailed abo e.
The lib osa lib a y McFee e al. (2015) is used o
ob ain he Mel-spec og am ep esen a ion o he au-
dio b anch. The audio samples, p o ided as mp3 iles,
a e i s con e ed o wa e o ms (.wa ) and down-
sampled om 22.5 o 16kHz. This is done o educe
he complexi y o he model, along wi h he compu-
a ional cos o op imiza ion. The downsampling has
been shown o p o ide simila esul s o highe sam-
pling a es, showing he obus ness o DL app oaches
Py o olakis e al. (2022). The spec al ep esen a ions
a e hen gene a ed using de aul pa ame e s o he
leng h o he Fas Fou ie T ans o m window (2048)
as well as he hop size (512).
As o wo d embeddings, he Sen ence T ans-
o me lib a y om Hugging Face was used, speci i-
cally, he all- obe a-la ge- 1 p e- ained model. The
embedde ecei es a con ex o up o 512 okens and
ou pu s a 1024 embedded ec o . Gi en ha he bes
esul s we e p o ided by using he ull con ex win-
dow, some o he ly ics had o be cu o a some
poin . A e some simple okeniza ion s eps, namely
emo ing new line cha ac e s and con e ing all ex
o lowe case, he embeddings we e ob ained up o he
al eady men ioned con ex size.
Figu e 6: The mul i-modal audio-ly ics eg ession model. Emo ionally- ele an ea u es a e lea ned o bo h he audio ep e-
sen a ion in he Mel-spec og am- ecei ing b anch and he ly ics ep esen a ion in he b anch ecei ing he p e iously gene -
a ed wo d embeddings. AV alues a e p edic ed a e conca ena ing and p ocessing he lea ned ea u es om bo h b anches.
3.3 Model Op imiza ion
Model op imiza ion was conduc ed using he
Bayesian op imiza ion implemen a ion o he Ke as
Tune lib a y O’Malley e al. (2019). This me hod
inds he bes combina ion o hype pa ame e s in p e-
iously de ined in e als o each, ei he maximizing
o minimizing an objec i e unc ion de ined by he
use .
Since ou me hodology is based on a eg ession
ask o p edic a ousal and alence o a gi en sam-
ple, he objec i e is de ined as minimizing he sum
o he mean squa ed e o (MSE) o bo h. This en-
su es ha none is p io i ized, le e aging bo h audio’s
be e p edic abili y in e ms o a ousal and he same
o ly ics’ p edic abili y o alence. The in e als o
each conside ed hype pa ame e , namely, ba ch size,
op imize , and co esponding lea ning a e, a e p e-
sen ed in Table 1.
Table 1: Op imal Hype pa ame e s Fo Each Da ase
Bes Hype pa ame e s
Ba ch Size Op mize Lea ning Ra e
64 SGD 1e-2
The op imiza ion p ocess is un o e en ials, pe
he lib a y’s de aul , s a ing a he lowe end o each
in e al. Fo each ial, he model is ained o a max-
imum o 200 epochs, wi h an ea ly s opping s a egy
de ined o check o no imp o emen s o he alida-
ion loss o 15 consecu i e epochs. This conside ably
educes he ime needed o conduc he ull op imiza-
ion phase since less ime is spen on unde pe o ming
se s o hype pa ame e s. We used a 70-15-15 ain-
alida e- es (TVT) spli as ou alida ion s a egy, as
de ined in Lou o e al. (2024b). The esul ing models
o each ial a e backed up o la e usage, including
he e alua ion phase, which is discussed nex .
3.4 Da a and E alua ion
The MERGE Bimodal Comple e da ase was used o
alida ing ou app oach. P oposed in Lou o e al.
(2024b), i comp ises a se o 2216 bimodal samples
(audio clips and co esponding ly ics). Fo each sam-
ple, he da ase p o ides a 30-second audio exce p o
he mos ep esen a i e pa o he song, links o he
ull ly ics, labels co esponding o each o he quad-
an s in Russell’s Ci cumplex model, and AV alues,
used o ob ain he p e iously men ioned labels, cal-
Table 2: TVT 70-15-15 Resul s Fo MERGE Audio Comple e
F1-sco e P ecision Recall R2 RMSE
(A/V) (A/V)
73.20% 74.53% 73.49% 0.454 0.133
0.506 0.339
cula ed based on he ex ac ed emo ion- ela ed ags
a ailable in AllMusic 2.
The abo e-men ioned AV alues a e ob ained
h ough he ollowing p ocess. Fi s , he a ailable
ags o each song in he da ase a e ob ained om
he All Music pla o m. Using Wa ine ’s Adjec i e
Dic iona y Wa ine e al. (2013), he exis ing ags a e
ansla ed o a ousal and alence alues. Finally, The
alues a e hen a e aged ac oss all ags co espond-
ing o a speci ic song, ob aining i s inal mapping on
Russell’s Ci cumplex model.
Fo he TVT s a egy, bo h he aining and ali-
da ion se s a e used in he op imiza ion unc ion. The
se o op imal hype pa ame e s is ound using he la -
e . A e aining he model o each da ase , he ol-
lowing me ics a e compu ed be ween he ac ual and
p edic ed AV alues in he es se o each class as
well as o he o e all pe o mance: F1-sco e, P eci-
sion, Recall, R2(squa ed Pea son’s co ela ion), and
Roo Mean Squa ed E o (RMSE).
Be o e compu ing hese me ics, he p edic ed and
eal AV alues we e mapped o Russell’s Ci cumplex
model o ob ain classes o calcula ing P ecision, Re-
call, and F1-sco e.
4 EXPERIMENTAL RESULTS
AND DISCUSSION
Tables 2 and 3 show he o e all esul s o he dis-
cussed me hodology. The a ousal and alence s an-
dalone esul s o he R2and RMSE me ics a e p e-
sen ed in consecu i e lines in he o de displayed in
he ables.
The ob ained esul s o bo h da ase s a e lowe
han hose ob ained in p e ious s udies ocused on
s a ic MER as a ca ego ical p oblem Lou o e al.
(2024a). The bes esul a ained is a 73.20% F1-
sco e, which is a ound 6% lowe han he esul s ob-
ained o he same da ase and e alua ion s a egy in
he men ioned a icle. The lowe esul s a e mos ly
due o he semi-au oma ic app oach o ob ain AV al-
ues (see Sec ion 3.4, conside ing ha he ags a ail-
able on All Music a e use -gene a ed and i s cu a ion
is unknown.
2h ps://www.allmusic.com/
Table 3: TVT 70-15-15 Resul s Con usion Ma ix Fo
MERGE Audio Comple e
P edic ed
Q1 Q2 Q3 Q4
Ac ual
Q1 61.3% 10.4% 6.6% 21.7%
Q2 9.8% 82.4% 5.9% 2.0%
Q3 1.4% 4.3% 78.3% 15.9%
Q4 7.3% 0.0% 18.2% 74.5%
As shown in Table 2, he R2me ic o alence
ou pe o med he one o a ousal, al hough ha ing a
la ge RMSE. This indica es ha he ela i e alence
h oughou songs is easonably cap u ed, despi e he
la ge RMSE e o .
Al hough he a ained esul s show oom o im-
p o emen , hey a e a good s a ing poin o he use .
Gi en he subjec i e na u e o each use ’s emo ional
pe cep ion, we belie e ha he pe sonaliza ion ea u e
included in he MERGE app is a aluable mechanism
o handling subjec i i y in MER.
In e ms o he esul s o sepa a e quad an s (Ta-
ble 3), we can see ha some Q1 songs a e con used
wi h Q4 songs (21.77% Q1 songs a e inco ec ly clas-
si ied as Q4). Mo eo e , he e is also some con usion
be ween Q3 and Q4 (15.9% o Q3 songs a e p edic ed
as Q4 and 18.2% o Q4 songs a e classi ied as Q3).
This is a known di icul y in MER, as discussed in
Panda e al. (2020) ha needs u he esea ch.
5 CONCLUSION AND FUTURE
WORK
We p esen ed he p o o ype o he MERGE applica-
ion. Cu en ly, he ini ial e sion has implemen ed
music playback ea u es, he abili y o add and il e
songs o a sha ed da abase, lis and plane iews, he
la e based on Russell’s Ci cumplex model, and use
managemen unc ionali ies. Mo eo e , a bimodal
audio-ly ics model is inco po a ed in o he backend
o he p o o ype o allow o AV alue p edic ion o
use -uploaded songs. Pa h-based playlis gene a ion
has also been implemen ed, enabling use s o c a
a playlis ha ollows a speci ic emo ional ajec o y
hey ha e selec ed.
S ill, many mo e unc ionali ies a e planned o
he applica ion in u u e i e a ions. The highligh ed
unc ionali ies include use -gene a ed ags o a mo e
cus omized il e ing expe ience ha would be a ail-
able o o he use s; au oma ic ly ics o he ull song
sc aped om an a ailable API, e.g., Genius; and Mu-
sic Emo ion Va ia ion De ec ion (MEVD) p edic ion
suppo , including isualiza ion wi h he same colo
code used in he plo . A s andalone desk op applica-
ion is also planned wi hou he c oss-use ea u es.
in addi ion o implemen ing hese upcoming ea u es,
We plan o conduc in-dep h use expe ience s udies
o gain a mo e comp ehensi e unde s anding o he
sys em’s e icacy and use sa is ac ion.
Valida ion expe imen s on wo ecen ly p oposed
da ase s a e p o ided alongside a ho ough sys em de-
sc ip ion, elaying insigh s in o he ob ained esul s.
These a e s ill below he ca ego ical app oach p e-
sen ed in Lou o e al. (2024b) due o he al eady dis-
cussed semi-au oma ic AV mapping app oach in Sec-
ion 3.4. Despi e his, he p edic ions a e a good s a -
ing poin o be u he adjus ed o he use ’s pe cep-
ion.
Rega ding he ac ual model, nei he ea u e lea n-
ing po ion may be ideal o he p oblem a hand
since hey we e o iginally de eloped o a ca ego i-
cal p oblem. De eloping mo e sui able a chi ec u es
should hus be conside ed u u e wo k. Fu he mo e,
he da a ep esen a ions, especially he wo d embed-
dings, may also be u he imp o ed, conside ing ha
he p e- ained model used is limi ed o a con ex win-
dow o 512 okens.
To conclude, we belie e he p oposed app migh
be use ul o music lis ene s. Al hough he e is oom
o imp o emen (as he a ained classi ica ion esul s
show), he pe sonaliza ion mechanism is a use ul ea-
u e o handling p edic ion e o s and subjec i i y.
Finally, he pe sonaliza ion ea u e and he mul i-use
en i onmen ha e he po en ial o acqui e quali y use
anno a ions, leading o a u u e la ge and mo e obus
MER da ase .
ACKNOWLEDGEMENTS
This wo k is unded by FCT - Founda ion o Sci-
ence and Technology, I.P., wi hin he scope o
he p ojec s: MERGE - DOI: 10.54499/PTDC/CCI-
COM/3171/2021 inanced wi h na ional unds (PID-
DAC) ia he Po uguese S a e Budge ; and p ojec
CISUC - UID/CEC/00326/2020 wi h unds om he
Eu opean Social Fund, h ough he Regional Ope a-
ional P og am Cen o 2020. Rena o Panda was sup-
po ed by Ci2 - FCT UIDP/05567/2020.
We hank all e iewe s o hei aluable sugges-
ions, which help o imp o e he a icle.
REFERENCES
Ca doso, L., Panda, R., and Pai a, R. P. (2011). Moode ec-
o : A p o o ype so wa e ool o mood-based playlis
gene a ion. In Simp´
osio de In o m´
a ica - INFo um 2011,
Coimb a, Po ugal.
Delbouys, R., Hennequin, R., Piccoli, F., Royo-Le elie ,
J., and Moussallam, M. (2018). Music Mood De ec ion
Based On Audio And Ly ics Wi h Deep Neu al Ne . In
P oceedings o he 19 h In e na ional Socie y o Music
In o ma ion Re ie al Con e ence, pages 370–375, Pa is,
F ance.
Hu, X., Downie, J. S., Lau ie , C., Bay, M., and Ehmann,
A. F. (2008). The 2007 Mi ex Audio Mood Classi ica ion
Task: Lessons Lea ned. In P oceedings o he 9 h In e -
na ional Socie y o Music In o ma ion Re ie al Con-
e ence, pages 462–467, D exel Uni e si y, Philadelphia,
Pennsyl ania, USA.
Lou o, P. L., Redinho, H., Malhei o, R., Pai a, R. P., and
Panda, R. (2024a). A Compa ison S udy o Deep Lea n-
ing Me hodologies o Music Emo ion Recogni ion. Sen-
so s, 24(7):2201.
Lou o, P. L., Redinho, H., San os, R., Malhei o, R., Panda,
R., and Pai a, R. P. (2024b). MERGE – A Bimodal
Da ase o S a ic Music Emo ion Recogni ion.
McFee, B., Ra el, C., Liang, D., Ellis, D., McVica , M.,
Ba enbe g, E., and Nie o, O. (2015). Lib osa: Audio and
Music Signal Analysis in Py hon. In Py hon in Science
Con e ence, pages 18–24, Aus in, Texas.
O’Malley, T., Bu sz ein, E., Long, J., Cholle , F., Jin, H.,
In e nizzi, L., e al. (2019). Ke as Tune . h ps://
gi hub.com/ke as- eam/ke as- une .
Panda, R., Malhei o, R., and Pai a, R. P. (2020). No el
Audio Fea u es o Music Emo ion Recogni ion. IEEE
T ansac ions on A ec i e Compu ing, 11(4):614–626.
Py o olakis, K., Tzou eli, P., and S amou, G. (2022).
Mul i-Modal Song Mood De ec ion wi h Deep Lea ning.
Senso s, 22(3):1065.
Russell, J. A. (1980). A ci cumplex model o a ec . Jou nal
o Pe sonali y and Social Psychology, 39(6):1161–1178.
Wa ine , A. B., Kupe man, V., and B ysbae , M.
(2013). No ms o alence, a ousal, and dominance o
13,915 English lemmas. Beha io Resea ch Me hods,
45(4):1191–1207.