scieee Science in your language
[en] (orig)

MERGE App: A Prototype Software for Multi-User Emotion-Aware Music Management

Abstract

We present a prototype software for multi-user music library management using the perceived emotional content of songs. The tool offers music playback features, song filtering by metadata, and automatic emotion prediction based on arousal and valence, with the possibility of personalizing the predictions by allowing each user to edit these values based on their own emotion assessment. This is an important feature for handling both classification errors and subjectivity issues, which are inherent aspects of emotion perception. A path-based playlist generation function is also implemented. A multi-modal audio-lyrics regression methodology is proposed for emotion prediction, with accompanying validation experiments on the MERGE dataset. The results obtained are promising, showing higher overall performance on train-validate-test splits (73.20% F1-score with the best dataset/split combination).

Read accessible full text

MERGE App: A Prototype Software for Multi-User Emotion-Aware Music Management

Author: Louro, Pedro; Branco, Guilherme; Redinho, Hugo; Santos, Ricardo Correia Nascimento Dos; Malheiro, Ricardo; Panda, Renato; Paiva, Rui Pedro
Publisher: SCITEPRESS - Science and Technology Publications
Year: 2024
Source: https://comum.rcaap.pt/bitstreams/c491063b-f754-4ab8-bdbc-bee7058b62e0/download
MERGE App: A P o o ype So wa e o Mul i-use Emo ion-awa e
Music Managemen
Ped o Lima Lou o1 a, Guilhe me B anco1 b, Hugo Redinho1 c, Rica do Co eia1 d,
Rica do Malhei o1,2 e, Rena o Panda1,3 , Rui Ped o Pai a1 g
1Uni e si y o Coimb a, Cen e o In o ma ics and Sys ems o he Uni e si y o Coimb a (CISUC), Depa men o
In o ma ics Enginee ing, and LASI
2Poly echnic Ins i u e o Lei ia School o Technology and Managemen
3Ci2 — Sma Ci ies Resea ch Cen e , Poly echnic Ins i u e o Toma
ped olou [email p o ec ed], guilhe me[email p o ec ed], [email p o ec ed], { ica doco eia, smal, panda,
uiped o}@dei.uc.p
Keywo ds: Music In o ma ion Re ie al; Music Emo ion Recogni ion; Machine Lea ning; Deep Lea ning; So wa e
Abs ac : We p esen a p o o ype so wa e o mul i-use music lib a y managemen using he pe cei ed emo ional
con en o songs. The ool o e s music playback ea u es, song il e ing by me ada a, and au oma ic emo ion
p edic ion based on a ousal and alence, wi h he possibili y o pe sonalizing he p edic ions by allowing each
use o edi hese alues based on hei own emo ion assessmen . This is an impo an ea u e o handling
bo h classi ica ion e o s and subjec i i y issues, which a e inhe en aspec s o emo ion pe cep ion. A pa h-
based playlis gene a ion unc ion is also implemen ed. A mul i-modal audio-ly ics eg ession me hodology
is p oposed o emo ion p edic ion, wi h accompanying alida ion expe imen s on he MERGE da ase . The
esul s ob ained a e p omising, showing highe o e all pe o mance on ain- alida e- es spli s (73.20% F1-
sco e wi h he bes da ase /spli combina ion).
1 INTRODUCTION
The digi al e a has b ough an unp eceden ed amoun
o music igh a ou inge ips h ough digi al ma -
ke places and s eaming se ices. Wi h he sudden
a ailabili y o millions o songs o use s, he neces-
si y o au oma ically o ganize and ind ele an mu-
sic eme ged. Cu en ecommenda ion sys ems p o-
ide pe sonalized sugges ions o use s based on lis-
ening pa e ns and using ags, such as gen e, s yle,
e c. Howe e , op ions a e lacking when we conside
ecommenda ions based on he au oma ic analysis o
he emo ional con en o songs.
The ield o Music Emo ion Recogni ion (MER)
has seen conside able ad ances in ecen yea s in
e ms o he mo e classical app oaches. Panda e al.
ah ps://o cid.o g/0000-0003-3201-6990
bh ps://o cid.o g/0000-0003-4073-1716
ch ps://o cid.o g/0009-0004-1547-2251
dh ps://o cid.o g/0000-0001-5663-7228
eh ps://o cid.o g/0000-0002-3010-2732
h ps://o cid.o g/0000-0003-2539-5590
gh ps://o cid.o g/0000-0003-3215-3960
(2020) p oposed a new se o ea u es ha consid-
e ably inc eased he pe o mance o hese sys ems,
achie ing a 76.4% F1-sco e wi h he op 100 anked
ea u es. Al hough he ea u e e alua ion is limi ed
o one da ase , he imp o emen s a e signi ican com-
pa ed o he bes esul s om simila sys ems ha
eached a glass ceiling Hu e al. (2008).
One d awback o audio-only me hodologies is
hei sho comings when di e en ia ing alence. Va -
ious sys ems ha e been p oposed using a bimodal
app oach le e aging bo h audio and ly ics, a aining
conside able imp o emen s when compa ed o sys-
ems using only one o he o he Delbouys e al.
(2018); Py o olakis e al. (2022). Such sys ems ha e
also implemen ed Deep Lea ning (DL) a chi ec u es
o skip he ime-consuming ea u e enginee ing and
ex ac ion s eps om he classical sys ems and con-
side ably speed up he in e ence p ocess o he o e all
sys em.
In his s udy, we p esen he MERGE1applica ion,
1MERGE is he ac onym o ”Music Emo ion Recog-
ni ion nEx Gene a ion”, a esea ch p ojec unded by he
Po uguese Science Founda ion.
Figu e 1: MERGE applica ion in e ace. The AV plo , alongside music playback and song display con ols, is seen in yellow.
The able iew is highligh ed in blue wi h ed highligh il e ing sea ch ba and bu on o adding music. Finally, g een
highligh s he bu ons o applica ion in o ma ion and use logou .
which au oma ically p edic s he a ousal and alence
o songs based on Russell’s Ci cumplex Model Rus-
sell (1980). Two axes make up his model: a ousal
(Y-axis), which depic s whe he he song has high o
low ene gy, and alence (X-axis), which ep esen s
whe he he emo ion o he song has a nega i e o pos-
i i e conno a ion.
The in eg a ed model used o p edic ion is also
p esen ed in his s udy alongside alida ion expe i-
men s, which ecei ed bo h audio and ly ics in o ma-
ion o map he song mo e accu a ely in o he abo e
men ioned model.
The MERGE applica ion is a ollow-up o he
MOODe ec o applica ion, p e iously c ea ed by ou
eam Ca doso e al. (2011). The new MERGE app
was buil om sc a ch, wi h signi ican code e ac o -
ing and op imiza ion, while keeping he o e all use
in e ace o he MOODe ec o app. In addi ion, sig-
ni ican no el ea u es and imp o emen s we e im-
plemen ed, namely: i) a bimodal app, which ex-
ploi s he combina ion o audio and ly ics da a o im-
p o ed classi ica ion (unlike he single audio modali y
in he MOODe ec o app); ii) an imp o ed classi ica-
ion model, aining wi h he MERGE da ase Lou o
e al. (2024b) and ollowing a deep lea ning app oach
Lou o e al. (2024a); iii) and a shi om he mono-
li hic single-use pa adigm o he web-based mul i-
use pa adigm.
2 MERGE APPLICATION
The MERGE applica ion is implemen ed using
Ja aSc ip , wi h he addi ion o he jQue y lib a y o
handle AJAX, o i s on end, while he backend is
se ed using he Exp ess lib a y on op o Node.js.
The applica ion in e ace is depic ed in Figu e 1.
2.1 Applica ion O e iew
The componen s can be b oken down as: i) he Rus-
sell’s Ci cumplex model whe e all songs can be seen
(highligh ed in yellow); ii) a able iew o he songs
wi h a ious op ions o so ing (highligh ed in blue);
iii) op ions o il e and add new songs (seen in he
ed egion, om le o igh ); i ) in o ma ion abou
he applica ion, he cu en use , and an op ion o lo-
gou (highligh ed in g een, also om le o igh ).
Songs a e placed in he plo desc ibed in i) acco d-
ing o he es ima ed a ousal and alence (AV) alues
in an in e al o [-1, 1] o each axis. The p ocess
o ob aining hese alues is desc ibed in Sec ion 3.
Beyond he AV posi ioning, each poin on he plane
is also colo -coded depending on he quad an : g een
(happy), ed ( ense), blue (sad), and yellow ( elaxed).
The iew o he g aph can be swi ched be ween a)
”Uploaded” o show only songs uploaded by he cu -
en use , b) ”My Lib a y” o display a use ’s lib a y,
i.e., he songs uploaded by he cu en use , plus songs
uploaded by o he use s added by he cu en use , and
c) ”All songs” o show all he songs a ailable in he
da abase. A no e ega ding he la e op ion is he di -
e en ia ion o songs no added by he use , appea ing
as g ey do s in he plo , also depic ed in Figu e 1.
Each use can change he song’s posi ion di ec ly
by mo ing he poin in he plane iew, o by edi ing
he AV alues h ough he able iew. These new AV
alues a e unique o he use .
The applica ion can be used as an audio playback
so wa e, hus o e ing usual ea u es such as:
• Playback con ols o mp3 iles (play, pause,
seek);
• Volume con ols, including mu e;
• Double-clicking a song o be played ei he in he
plo o able iew;
Figu e 2: MERGE App Backo ice. Use s wi h adminis a i e p i ileges can o e w i e he model used o AV alues’ p edic-
ion (highligh ed in g een) and expo a CSV ile wi h in o ma ion ega ding he anno a ions o all use s o each song in he
da abase (highligh ed in blue).
Figu e 3: En i y ela ion diag am o he applica ion’s da abase.
• Fil e ing and so ing by any o he a ailable song
p ope ies ( i le, a is , alence, a ousal, emo ion);
• Adding and dele ing songs om he use ’s lib a y.
The applica ion also p o ides a backo ice o
use s wi h adminis a i e p i ileges, pic u ed in Fig-
u e 2. A e logging in, he use can pe o m one o
wo ac ions: upload and deploy a new model o AV
p edic ion, and expo a CSV wi h he exis ing use
AV anno a ions. The la e op ion is designed o eas-
ily e ie e each use ’s a ailable anno a ions o songs
in hei espec i e lib a ies. In his way, he MERGE
app can be used as a c owdsou cing da a collec ion
and anno a ion ool, p omo ing he c ea ion o size-
able and quali y MER da ase s, a cu en key need in
MER esea ch Panda e al. (2020).
Rega ding he da abase used o s o e all he el-
e an da a om use s and songs, he co esponding
en i y ela ionship diag am is depic ed in Figu e 3.
The use able s o es he use ’s pe sonal in o ma ion,
as well as he use ole, o access p i ileges pu -
poses. The pa h o he model used o AV p edic-
ion is sa ed in he co esponding able and iden i ies
he use ha uploaded he cu en ly deployed model.
The song able s o es all song- ela ed in o ma ion, in-
cluding me ada a, he AV alues i s p edic ed by he
p esen ly deployed model, he mapped quad an in he
plo iew, he pa h o he uploaded audio clip, and he
e e ence o he use who i s added he song.
The use -speci ic anno a ions a e s o ed in he an-
no a ion able, which s o es he AV alues de ined pe
he use ’s pe cep ion, he co esponding quad an , and
a e e ence o he song and use o ha anno a ion.
Finally, he lib a y able s o es all use lib a ies, con-
aining only songs added by he use , ei he h ough
uploading o om o he use s’ lib a ies.
Figu e 4: A he op, ”Ano he One Bi es The Dus ” by Queen is added o he lib a y and placed in he plane acco ding o he
p edic ed AV alues. A he bo om, he poin ep esen ing he song is mo ed o a mo e accu a e posi ion, acco ding o he
use .
2.2 Building An Emo ionally-awa e
Lib a y
A e adding a new song om an a ailable MP3 ile,
AV alues a e au oma ically p edic ed, and a poin is
added o he plo alongside a new en y on he able.
Should he use disag ee wi h he p edic ed alues,
hese can be easily changed by edi ing he en y om
he able iew o mo ing he poin in he plo . An ex-
ample o he ini ially p edic ed posi ion o a newly
added song and he inal posi ion a e adjus men can
be seen in Figu e 4. This pe sonaliza ion mechanism
p o ides use s wi h he abili y o add ess he in in-
sic subjec i i y in MER. Howe e , ackling his issue
con inues o p esen a signi ican challenge.
2.3 Pa h-based Playlis Gene a ion
The abili y o gene a e a playlis based on a use -
d awn pa h is cu en ly implemen ed, as depic ed in
Figu e 5. This ea u e allows use s o eely c ea e an
emo ionally- a ying playlis .
This is done by compu ing he dis ance o he use -
de ined Ncloses songs o he e e ence poin s ha
make up he d awn pa h. The use may also con igu e
how a he songs can be om he pa h o be consid-
e ed in o he calcula ions. This h eshold is de ined as
a decimal numbe be ween he plane in e al ([-1, 1]).
3 SONG EMOTION PREDICTION
In his sec ion, we discuss he me hodology used o
p edic AV alues o a gi en song. Fi s , he DL
model’s a chi ec u e is p esen ed, ollowed by he p e-
p ocessing s eps o each modali y, a desc ip ion o
he op imiza ion used, and he e alua ion conduc ed.
3.1 Model A chi ec u e
The p oposed a chi ec u e, depic ed in Figu e 6, is
based on he one by Delbouys e al. Delbouys e al.
(2018). Dis inc audio and ly ics b anches ecei e
Mel-spec og am ep esen a ions and wo d embed-
dings, espec i ely. The lea ned ea u es o each
modali y a e hen used and u he p ocessed by
a small Dense Neu al Ne wo k (DNN), inally ou -
pu ing he AV alues p edic ion.
We op ed o a bimodal audio-ly ics app oach
conside ing ha bo h modali ies ha e ele an in o -
ma ion o he di e en axes o Russell’s Ci cum-
plex Model. Audio has been shown o be e p edic
Figu e 5: The use can be seen d awing a pa h o gene a e a playlis wi h he desi ed emo ional ajec o y a he op. The esul
o he pa h-based playlis gene a ion is p esen ed a he bo om.
a ousal, while ly ical in o ma ion is mo e ele an o
alence p edic ion Lou o e al. (2024b).
S a ing in he audio b anch, Mel-spec og am
ep esen a ions o each sample a e ed o he ea u e
lea ning po ion o he baseline a chi ec u e p esen ed
in Lou o e al. Lou o e al. (2024a). I is composed
o ou con olu ional blocks, composed o a 2D Con-
olu ional laye , ollowed by a Ba ch No maliza ion,
D opou , and Max Pooling laye , inishing wi h ReLU
ac i a ion. As o he ly ics b anch, he wo d em-
beddings o ly ics a e also ed o ou con olu ional
blocks, each comp ising a 1D Con olu ional laye ,
ollowed by Max Pooling and a ReLU ac i a ion lay-
e s. To balance he in o ma ion om each modali y,
we signi ican ly educe he o e whelming amoun o
lea ned ea u es om ly ics using a Dense laye be-
o e me ging he lea ned ea u es om bo h b anches.
The classi ica ion po ion o he model is com-
posed o al e na ing D oupou and Dense laye s,
which educe and u he p ocess he se o ea u es
espec i ely, inally ou pu ing one o Russell’s Ci -
cumplex model’s ou quad an s.
3.2 P e-p ocessing S eps
A se o p e-p ocessing s eps is necessa y o ob ain
he da a ep esen a ions used o each b anch o he
a chi ec u e de ailed abo e.
The lib osa lib a y McFee e al. (2015) is used o
ob ain he Mel-spec og am ep esen a ion o he au-
dio b anch. The audio samples, p o ided as mp3 iles,
a e i s con e ed o wa e o ms (.wa ) and down-
sampled om 22.5 o 16kHz. This is done o educe
he complexi y o he model, along wi h he compu-
a ional cos o op imiza ion. The downsampling has
been shown o p o ide simila esul s o highe sam-
pling a es, showing he obus ness o DL app oaches
Py o olakis e al. (2022). The spec al ep esen a ions
a e hen gene a ed using de aul pa ame e s o he
leng h o he Fas Fou ie T ans o m window (2048)
as well as he hop size (512).
As o wo d embeddings, he Sen ence T ans-
o me lib a y om Hugging Face was used, speci i-
cally, he all- obe a-la ge- 1 p e- ained model. The
embedde ecei es a con ex o up o 512 okens and
ou pu s a 1024 embedded ec o . Gi en ha he bes
esul s we e p o ided by using he ull con ex win-
dow, some o he ly ics had o be cu o a some
poin . A e some simple okeniza ion s eps, namely
emo ing new line cha ac e s and con e ing all ex
o lowe case, he embeddings we e ob ained up o he
al eady men ioned con ex size.

Figu e 6: The mul i-modal audio-ly ics eg ession model. Emo ionally- ele an ea u es a e lea ned o bo h he audio ep e-
sen a ion in he Mel-spec og am- ecei ing b anch and he ly ics ep esen a ion in he b anch ecei ing he p e iously gene -
a ed wo d embeddings. AV alues a e p edic ed a e conca ena ing and p ocessing he lea ned ea u es om bo h b anches.
3.3 Model Op imiza ion
Model op imiza ion was conduc ed using he
Bayesian op imiza ion implemen a ion o he Ke as
Tune lib a y O’Malley e al. (2019). This me hod
inds he bes combina ion o hype pa ame e s in p e-
iously de ined in e als o each, ei he maximizing
o minimizing an objec i e unc ion de ined by he
use .
Since ou me hodology is based on a eg ession
ask o p edic a ousal and alence o a gi en sam-
ple, he objec i e is de ined as minimizing he sum
o he mean squa ed e o (MSE) o bo h. This en-
su es ha none is p io i ized, le e aging bo h audio’s
be e p edic abili y in e ms o a ousal and he same
o ly ics’ p edic abili y o alence. The in e als o
each conside ed hype pa ame e , namely, ba ch size,
op imize , and co esponding lea ning a e, a e p e-
sen ed in Table 1.
Table 1: Op imal Hype pa ame e s Fo Each Da ase
Bes Hype pa ame e s
Ba ch Size Op mize Lea ning Ra e
64 SGD 1e-2
The op imiza ion p ocess is un o e en ials, pe
he lib a y’s de aul , s a ing a he lowe end o each
in e al. Fo each ial, he model is ained o a max-
imum o 200 epochs, wi h an ea ly s opping s a egy
de ined o check o no imp o emen s o he alida-
ion loss o 15 consecu i e epochs. This conside ably
educes he ime needed o conduc he ull op imiza-
ion phase since less ime is spen on unde pe o ming
se s o hype pa ame e s. We used a 70-15-15 ain-
alida e- es (TVT) spli as ou alida ion s a egy, as
de ined in Lou o e al. (2024b). The esul ing models
o each ial a e backed up o la e usage, including
he e alua ion phase, which is discussed nex .
3.4 Da a and E alua ion
The MERGE Bimodal Comple e da ase was used o
alida ing ou app oach. P oposed in Lou o e al.
(2024b), i comp ises a se o 2216 bimodal samples
(audio clips and co esponding ly ics). Fo each sam-
ple, he da ase p o ides a 30-second audio exce p o
he mos ep esen a i e pa o he song, links o he
ull ly ics, labels co esponding o each o he quad-
an s in Russell’s Ci cumplex model, and AV alues,
used o ob ain he p e iously men ioned labels, cal-
Table 2: TVT 70-15-15 Resul s Fo MERGE Audio Comple e
F1-sco e P ecision Recall R2 RMSE
(A/V) (A/V)
73.20% 74.53% 73.49% 0.454 0.133
0.506 0.339
cula ed based on he ex ac ed emo ion- ela ed ags
a ailable in AllMusic 2.
The abo e-men ioned AV alues a e ob ained
h ough he ollowing p ocess. Fi s , he a ailable
ags o each song in he da ase a e ob ained om
he All Music pla o m. Using Wa ine ’s Adjec i e
Dic iona y Wa ine e al. (2013), he exis ing ags a e
ansla ed o a ousal and alence alues. Finally, The
alues a e hen a e aged ac oss all ags co espond-
ing o a speci ic song, ob aining i s inal mapping on
Russell’s Ci cumplex model.
Fo he TVT s a egy, bo h he aining and ali-
da ion se s a e used in he op imiza ion unc ion. The
se o op imal hype pa ame e s is ound using he la -
e . A e aining he model o each da ase , he ol-
lowing me ics a e compu ed be ween he ac ual and
p edic ed AV alues in he es se o each class as
well as o he o e all pe o mance: F1-sco e, P eci-
sion, Recall, R2(squa ed Pea son’s co ela ion), and
Roo Mean Squa ed E o (RMSE).
Be o e compu ing hese me ics, he p edic ed and
eal AV alues we e mapped o Russell’s Ci cumplex
model o ob ain classes o calcula ing P ecision, Re-
call, and F1-sco e.
4 EXPERIMENTAL RESULTS
AND DISCUSSION
Tables 2 and 3 show he o e all esul s o he dis-
cussed me hodology. The a ousal and alence s an-
dalone esul s o he R2and RMSE me ics a e p e-
sen ed in consecu i e lines in he o de displayed in
he ables.
The ob ained esul s o bo h da ase s a e lowe
han hose ob ained in p e ious s udies ocused on
s a ic MER as a ca ego ical p oblem Lou o e al.
(2024a). The bes esul a ained is a 73.20% F1-
sco e, which is a ound 6% lowe han he esul s ob-
ained o he same da ase and e alua ion s a egy in
he men ioned a icle. The lowe esul s a e mos ly
due o he semi-au oma ic app oach o ob ain AV al-
ues (see Sec ion 3.4, conside ing ha he ags a ail-
able on All Music a e use -gene a ed and i s cu a ion
is unknown.
2h ps://www.allmusic.com/
Table 3: TVT 70-15-15 Resul s Con usion Ma ix Fo
MERGE Audio Comple e
P edic ed
Q1 Q2 Q3 Q4
Ac ual
Q1 61.3% 10.4% 6.6% 21.7%
Q2 9.8% 82.4% 5.9% 2.0%
Q3 1.4% 4.3% 78.3% 15.9%
Q4 7.3% 0.0% 18.2% 74.5%
As shown in Table 2, he R2me ic o alence
ou pe o med he one o a ousal, al hough ha ing a
la ge RMSE. This indica es ha he ela i e alence
h oughou songs is easonably cap u ed, despi e he
la ge RMSE e o .
Al hough he a ained esul s show oom o im-
p o emen , hey a e a good s a ing poin o he use .
Gi en he subjec i e na u e o each use ’s emo ional
pe cep ion, we belie e ha he pe sonaliza ion ea u e
included in he MERGE app is a aluable mechanism
o handling subjec i i y in MER.
In e ms o he esul s o sepa a e quad an s (Ta-
ble 3), we can see ha some Q1 songs a e con used
wi h Q4 songs (21.77% Q1 songs a e inco ec ly clas-
si ied as Q4). Mo eo e , he e is also some con usion
be ween Q3 and Q4 (15.9% o Q3 songs a e p edic ed
as Q4 and 18.2% o Q4 songs a e classi ied as Q3).
This is a known di icul y in MER, as discussed in
Panda e al. (2020) ha needs u he esea ch.
5 CONCLUSION AND FUTURE
WORK
We p esen ed he p o o ype o he MERGE applica-
ion. Cu en ly, he ini ial e sion has implemen ed
music playback ea u es, he abili y o add and il e
songs o a sha ed da abase, lis and plane iews, he
la e based on Russell’s Ci cumplex model, and use
managemen unc ionali ies. Mo eo e , a bimodal
audio-ly ics model is inco po a ed in o he backend
o he p o o ype o allow o AV alue p edic ion o
use -uploaded songs. Pa h-based playlis gene a ion
has also been implemen ed, enabling use s o c a
a playlis ha ollows a speci ic emo ional ajec o y
hey ha e selec ed.
S ill, many mo e unc ionali ies a e planned o
he applica ion in u u e i e a ions. The highligh ed
unc ionali ies include use -gene a ed ags o a mo e
cus omized il e ing expe ience ha would be a ail-
able o o he use s; au oma ic ly ics o he ull song
sc aped om an a ailable API, e.g., Genius; and Mu-
sic Emo ion Va ia ion De ec ion (MEVD) p edic ion
suppo , including isualiza ion wi h he same colo
code used in he plo . A s andalone desk op applica-
ion is also planned wi hou he c oss-use ea u es.
in addi ion o implemen ing hese upcoming ea u es,
We plan o conduc in-dep h use expe ience s udies
o gain a mo e comp ehensi e unde s anding o he
sys em’s e icacy and use sa is ac ion.
Valida ion expe imen s on wo ecen ly p oposed
da ase s a e p o ided alongside a ho ough sys em de-
sc ip ion, elaying insigh s in o he ob ained esul s.
These a e s ill below he ca ego ical app oach p e-
sen ed in Lou o e al. (2024b) due o he al eady dis-
cussed semi-au oma ic AV mapping app oach in Sec-
ion 3.4. Despi e his, he p edic ions a e a good s a -
ing poin o be u he adjus ed o he use ’s pe cep-
ion.
Rega ding he ac ual model, nei he ea u e lea n-
ing po ion may be ideal o he p oblem a hand
since hey we e o iginally de eloped o a ca ego i-
cal p oblem. De eloping mo e sui able a chi ec u es
should hus be conside ed u u e wo k. Fu he mo e,
he da a ep esen a ions, especially he wo d embed-
dings, may also be u he imp o ed, conside ing ha
he p e- ained model used is limi ed o a con ex win-
dow o 512 okens.
To conclude, we belie e he p oposed app migh
be use ul o music lis ene s. Al hough he e is oom
o imp o emen (as he a ained classi ica ion esul s
show), he pe sonaliza ion mechanism is a use ul ea-
u e o handling p edic ion e o s and subjec i i y.
Finally, he pe sonaliza ion ea u e and he mul i-use
en i onmen ha e he po en ial o acqui e quali y use
anno a ions, leading o a u u e la ge and mo e obus
MER da ase .
ACKNOWLEDGEMENTS
This wo k is unded by FCT - Founda ion o Sci-
ence and Technology, I.P., wi hin he scope o
he p ojec s: MERGE - DOI: 10.54499/PTDC/CCI-
COM/3171/2021 inanced wi h na ional unds (PID-
DAC) ia he Po uguese S a e Budge ; and p ojec
CISUC - UID/CEC/00326/2020 wi h unds om he
Eu opean Social Fund, h ough he Regional Ope a-
ional P og am Cen o 2020. Rena o Panda was sup-
po ed by Ci2 - FCT UIDP/05567/2020.
We hank all e iewe s o hei aluable sugges-
ions, which help o imp o e he a icle.
REFERENCES
Ca doso, L., Panda, R., and Pai a, R. P. (2011). Moode ec-
o : A p o o ype so wa e ool o mood-based playlis
gene a ion. In Simp´
osio de In o m´
a ica - INFo um 2011,
Coimb a, Po ugal.
Delbouys, R., Hennequin, R., Piccoli, F., Royo-Le elie ,
J., and Moussallam, M. (2018). Music Mood De ec ion
Based On Audio And Ly ics Wi h Deep Neu al Ne . In
P oceedings o he 19 h In e na ional Socie y o Music
In o ma ion Re ie al Con e ence, pages 370–375, Pa is,
F ance.
Hu, X., Downie, J. S., Lau ie , C., Bay, M., and Ehmann,
A. F. (2008). The 2007 Mi ex Audio Mood Classi ica ion
Task: Lessons Lea ned. In P oceedings o he 9 h In e -
na ional Socie y o Music In o ma ion Re ie al Con-
e ence, pages 462–467, D exel Uni e si y, Philadelphia,
Pennsyl ania, USA.
Lou o, P. L., Redinho, H., Malhei o, R., Pai a, R. P., and
Panda, R. (2024a). A Compa ison S udy o Deep Lea n-
ing Me hodologies o Music Emo ion Recogni ion. Sen-
so s, 24(7):2201.
Lou o, P. L., Redinho, H., San os, R., Malhei o, R., Panda,
R., and Pai a, R. P. (2024b). MERGE – A Bimodal
Da ase o S a ic Music Emo ion Recogni ion.
McFee, B., Ra el, C., Liang, D., Ellis, D., McVica , M.,
Ba enbe g, E., and Nie o, O. (2015). Lib osa: Audio and
Music Signal Analysis in Py hon. In Py hon in Science
Con e ence, pages 18–24, Aus in, Texas.
O’Malley, T., Bu sz ein, E., Long, J., Cholle , F., Jin, H.,
In e nizzi, L., e al. (2019). Ke as Tune . h ps://
gi hub.com/ke as- eam/ke as- une .
Panda, R., Malhei o, R., and Pai a, R. P. (2020). No el
Audio Fea u es o Music Emo ion Recogni ion. IEEE
T ansac ions on A ec i e Compu ing, 11(4):614–626.
Py o olakis, K., Tzou eli, P., and S amou, G. (2022).
Mul i-Modal Song Mood De ec ion wi h Deep Lea ning.
Senso s, 22(3):1065.
Russell, J. A. (1980). A ci cumplex model o a ec . Jou nal
o Pe sonali y and Social Psychology, 39(6):1161–1178.
Wa ine , A. B., Kupe man, V., and B ysbae , M.
(2013). No ms o alence, a ousal, and dominance o
13,915 English lemmas. Beha io Resea ch Me hods,
45(4):1191–1207.