scieee Science in your language
[en] (orig)

Design of an Unsupervised Machine Learning-Based Movie Recommender System

Abstract

This research aims to determine the similarities in groups of people to build a film recommender system for users. Users often have difficulty in finding suitable movies due to the increasing amount of movie information. The recommender system is very useful for helping customers choose a preferred movie with the existing features. In this study, the recommender system development is established by using several algorithms to obtain groupings, such as the K-Means algorithm, birch algorithm, mini-batch K-Means algorithm, mean-shift algorithm, affinity propagation algorithm, agglomerative clustering algorithm, and spectral clustering algorithm. We~propose methods optimizing K so that each cluster may not significantly increase variance. We~are limited to using groupings based on Genre and Tags for movies. This research can discover better methods for evaluating clustering algorithms. To verify the quality of the recommender system, we adopted the mean square error (MSE), such as the Dunn Matrix and Cluster Validity Indices, and social network analysis (SNA), such as Degree Centrality, Closeness Centrality, and~Betweenness Centrality. We also used average similarity, computational time, association rule with Apriori algorithm, and clustering performance evaluation as evaluation measures to compare method performance of recommender systems using Silhouette Coefficient, Calinski-Harabaz Index, and~Davies--Bouldin Index.

Read accessible full text

Design of an Unsupervised Machine Learning-Based Movie Recommender System

Author: Putri, Debby Cintia Ganesha; Leu, Jenq-Shiou; Šeda, Pavel
Publisher: MDPI
Year: 2020
DOI: 10.3390/sym12020185
Source: https://dspace.vut.cz/bitstreams/6505c296-fb9b-4baa-b3c2-f29560344010/download
symme y
S
S
A icle
Design o an Unsupe ised Machine Lea ning-Based
Mo ie Recommende Sys em
Debby Cin ia Ganesha Pu i 1,* , Jenq-Shiou Leu 1and Pa el Seda 2,3
1Depa men o Elec onic and Compu e Enginee ing, Na ional Taiwan Uni e si y o Science and
Technology, Taipei Ci y 106, Taiwan; [email p o ec ed]
2Depa men o Telecommunica ions, B no Uni e si y o Technology, Technicka 12, 61600 B no,
Czech Republic; xsedap01@ u b .cz
3Ins i u e o Compu e Science, Masa yk Uni e si y, Bo anica 554/68A, 602 00 B no, Czech Republic
*Co espondence: [email p o ec ed]
Recei ed: 25 Decembe 2019; Accep ed: 13 Janua y 2020; Published: 21 Janua y 2020


Abs ac :
This esea ch aims o de e mine he simila i ies in g oups o people o build a ilm
ecommende sys em o use s. Use s o en ha e di icul y in inding sui able mo ies due o he
inc easing amoun o mo ie in o ma ion. The ecommende sys em is e y use ul o helping
cus ome s choose a p e e ed mo ie wi h he exis ing ea u es. In his s udy, he ecommende sys em
de elopmen is es ablished by using se e al algo i hms o ob ain g oupings, such as he
K
-Means
algo i hm, bi ch algo i hm, mini-ba ch
K
-Means algo i hm, mean-shi algo i hm, a ini y p opaga ion
algo i hm, agglome a i e clus e ing algo i hm, and spec al clus e ing algo i hm. We p opose
me hods op imizing
K
so ha each clus e may no signi ican ly inc ease a iance. We a e limi ed o
using g oupings based on Gen e and Tags o mo ies. This esea ch can disco e be e me hods o
e alua ing clus e ing algo i hms. To e i y he quali y o he ecommende sys em, we adop ed he
mean squa e e o (MSE), such as he Dunn Ma ix and Clus e Validi y Indices, and social ne wo k
analysis (SNA), such as Deg ee Cen ali y, Closeness Cen ali y, and Be weenness Cen ali y. We also
used a e age simila i y, compu a ional ime, associa ion ule wi h Ap io i algo i hm, and clus e ing
pe o mance e alua ion as e alua ion measu es o compa e me hod pe o mance o ecommende
sys ems using Silhoue e Coe icien , Calinski-Ha abaz Index, and Da ies–Bouldin Index.
Keywo ds:
a ini y p opaga ion; agglome a i e spec al clus e ing; associa ion ule wi h Ap io i
algo i hm; a e age simila i y; bi ch; clus e ing pe o mance e alua ion; compu a ional ime;
Dunn Ma ix; mean-shi ; mean squa ed e o ; mini-ba ch
K
-Means; ecommenda ions sys em;
K-Means; social ne wo k analysis
1. In oduc ion
The explosion o in o ma ion on he in e ne is de eloping ollowing he apid ad ancemen o
in e ne echnology. The ecommende sys em is a simple mechanism o help use s ind he igh
in o ma ion based on he wishes o in e ne use s by e e ing o he p e e ence pa e ns in he da ase .
The pu pose o he ecommende sys em is o au oma ically gene a e p oposed i ems (web pages, news,
DVDs, music, mo ies, books, CDs) o use s based on his o ical p e e ences and sa e ime sea ching
o hem online by ex ac ing wo hwhile da a. Some websi es using he ecommende sys em me hod
include yahoo.com,ebay.com, and amazon.com [
1
–
6
]. A mo ie ecommende is an applica ion mos
widely used o help cus ome s selec ilms om a la ge capaci y ilm lib a y. This algo i hm can ank
i ems and show use s high-le el i ems and good con en o p o ide a mo ie ecommended based on
cus ome simila i y. Cus ome simila i y means collec ing ilm a ings gi en by indi iduals based on
Symme y 2020,12, 185; doi:10.3390/sym12020185 www.mdpi.com/jou nal/symme y
Symme y 2020,12, 185 2 o 27
gen e o ags and hen ecommending ilms ha p omise o a ge cus ome s based on indi iduals
wi h iden ic as es and p e e ences.
T adi ional ecommende sys ems always su e om se e al inhe en limi s, such as poo
scalabili y and da a spa si y [
7
]. Se e al wo ks ha e e ol ed a model-based app oach o o e come
his p oblem and p o ide he bene i s o he e ec i eness o he exis ing ecommende sys em. In he
li e a u e, many model-based ecommende sys ems we e de eloped by pa i ioning algo i hms,
such as K-Means, and sel -o ganizing maps (SOM) [8–12].
O he me hods ha can be used in he ecommende sys em include he cla i ica ion me hod,
associa ion ules, and da a g ouping. The pu pose o g ouping is o sepa a e use s in o di e en
g oups o o m neighbo s who a e “like-minded” (closes ) subs i u es o sea ching he en i e use
space o inc ease sys em scalabili y [
13
]. In essence, making high-quali y ilm ecommenda ions wi h
good g oupings emains a challenge and explo ing hose ollowing e icien g ouping me hods is an
impo an issue in he ecommended sys em si ua ion. A e y use ul ea u e in he ecommende
sys em becomes he abili y o guess use p e e ences and needs in analyzing use beha io o o he
use beha io s o p oduce a pe sonalized ecommende [14].
To o e come he challenges men ioned abo e, se e al me hods a e used o classi y pe o mance o
he mo ie ecommende sys em, such as
K
-Means algo i hm [
15
–
17
], bi ch algo i hm [
18
], mini-ba ch
K
-Means algo i hm [
19
], mean-shi algo i hm [
20
], a ini y p opaga ion algo i hm [
21
], agglome a i e
clus e ing algo i hm [
22
], and spec al clus e ing algo i hm [
23
]. In his a icle,
we de elop
a g ouping
ha can be op imized wi h se e al algo i hms, hen ob ain he bes algo i hm in g ouping use
simila i ies based on gen e, ags, and a ings on mo ies wi h he Mo ieLens da ase . Then, he
p oposed scheme op imizes
K
o each clus e so ha i can signi ican ly educe a iance. To be e
unde s and his me hod, when we alk abou a iance, we a e e e ing o mis akes. One way o
calcula e his e o is by ex ac ing he cen oids o each g oup and hen squa ing his alue o emo e
nega i e e ms. Then, all o hese alues a e added o ob ain o al e o . To e i y he quali y o he
ecommende sys em, we use mean squa ed e o (MSE), Dunn Ma ix as Clus e Validi y, and social
ne wo k analysis (SNA). I also uses a e age simila i y, compu a ional ime, ules o associa ion wi h
Ap io i algo i hms, and pe o mance e alua ion g ouping as e alua ion measu es o compa e he
pe o mance o ecommende sys ems.
1.1. P io Rela ed Wo ks
Zan Wang, X. Y. (2014) p esen ed esea ch on an imp o ed collabo a i e mo ie ecommende
sys em o de elop CF-based app oaches o hyb id models o p o ide mo ie ecommenda ions
ha combine dimensional educ ion echniques wi h exis ing clus e ing algo i hms. In a spa se
da a en i onmen , “like-minded” selec ion based on he gene al anking is a unc ion o p oducing
high-quali y ecommended ilms. Based on he Mo ieLens da a se , an expe imen al e alua ion
app oach can p o e ha i is capable o p oducing high p edic i e accu acy and mo e eliable
ilm ecommenda ions o exis ing use p e e ences compa ed o exis ing CF-based clus e ing [
13
].
This s udy also applies he clus e ing me hod o ind he nea es clus e and ecommends a lis o
mo ies based on simila i ies among use s. Ou ecommende sys em da ase e e s o his esea ch
by using Mo ieLens da ase o es ablish he expe imen s, including 100,000 a ings by 943 use s on
1682 mo ies, wi h a disc e e scale o 1–5. Each use has a ed a leas 20 mo ies. Then, he da ase
was andomly spli in o aining and es da a a an 80% o 20% a io. Md. Tayeb Himel, M. N. (2017)
esea ched he weigh based mo ie ecommende sys em using
K
-Means algo i hm [
14
]. This esea ch
uses he
K
-Means algo i hm and explains he esul s. This esea ch mo i a es us o use o he me hods
as a compa ison o iden i y he algo i hm wi h be e pe o mance.
1.2. P oblem Fo mula ion
In o ma ion o e load is a p oblem in in o ma ion e ie al, and he ecommenda ion sys em is
one o he main echniques o deal wi h p oblems by ad ising use s wi h app op ia e and ele an
Symme y 2020,12, 185 3 o 27
i ems. A p esen , se e al ecommenda ion sys ems ha e been de eloped o qui e di e en domains;
howe e , his is no app op ia e enough o mee use in o ma ion needs. The e o e, a high-quali y
ecommenda ion sys em needs o be buil . When designing hese ecommenda ions an app op ia e
me hod is needed. This pape in es iga es se e al app op ia e clus e ing me hods o esea ch
in de eloping high-quali y ecommenda ion sys ems wi h a p oposed algo i hm ha de e mines
simila i ies o de ine a people g oup o build a mo ie ecommende sys em o use s. Nex , expe imen s
a e conduc ed o make pe o mance compa isons wi h e alua ion c i e ia on se e al clus e ing
algo i hms using he
K
-Means algo i hm, bi ch algo i hm, mini-ba ch
K
-Means algo i hm, mean-shi
algo i hm, a ini y p opaga ion algo i hm, agglome a i e clus e ing algo i hm, and spec al clus e ing
algo i hm. The bes me hods a e iden i ied o se e as a ounda ion o imp o e and analyze his mo ie
ecommende sys em.
1.3. Main Con ibu ions
This s udy in es iga es se e al app op ia e clus e ing me hods o de elop high-quali y
ecommende sys ems wi h a p oposed algo i hm o inding he simila i ies wi hin g oups o
people. Nex , we conduc expe imen s o make compa isons on se e al clus e ing algo i hms
including
K
-Means algo i hm, bi ch algo i hm, mini-ba ch
K
-Means algo i hm, mean-shi algo i hm,
a ini y p opaga ion algo i hm, agglome a i e clus e ing algo i hm, and spec al clus e ing algo i hm.
A e ha , we ind he bes me hod om hem as a ounda ion o imp o e and analyze his mo ie
ecommende sys em. We limi o using h ee ags and h ee gen es because o analyze pe o mance
and ge good isualiza ion, he mos s able esul s a e h ee ags and h ee gen es, be o e we ha e
ied mo e han h ee bu he isualiza ion esul s ob ained a e no so good wi h some o he me hods
used in his s udy. We s a losing he abili y o isualize co ec ly when analyzing h ee o mo e
dimensions. Then we limi i by using a o i e gen es and ags and mo e de ails on he algo i hm
compa ison. The main con ibu ions o his s udy a e as ollows:
•
Pe o mance compa ison o se e al clus e ing me hods o gene a e a mo ie ecommende sys em.
•
To op imize he
K
alue in
K
-Means, mini-ba ch
K
-Means, bi ch, and agglome a i e
clus e ing algo i hms.
•
To e i y he quali y o he ecommende sys em, we employed social ne wo k analysis (SNA).
We also used he a e age simila i y o compa e pe o mance me hods and associa ion ules wi h
he Ap io i algo i hm o ecommende sys ems.
The emainde o his pape is o ganized as ollows. In Sec ion 2, we p esen an o e iew o he
ecommende sys em and e iew he clus e ing algo i hm and sys em design o he ecommende
sys em. We de ail he algo i hm design o he
K
-Means algo i hm, bi ch algo i hm, mini-ba ch
K
-Means
algo i hm, mean-shi algo i hm, a ini y p opaga ion algo i hm, agglome a i e clus e ing algo i hm,
and spec al clus e ing algo i hm. Addi ionally, we p oposed a me hod o op imize
K
in some o he
me hods. This session also explains he e alua ion c i e ia. The expe imen s, da ase explana ion,
and esul s a e illus a ed in Sec ion 3. E alua ion esul s o algo i hms ia a ious es cases and
discussions a e shown in Sec ion 4. Finally, Sec ion 5concludes his wo k.
2. Recommende Sys ems
2.1. O e iew
A ecommende sys em is a simple algo i hm o p o ide he mos ele an in o ma ion o use s
o ind pa e ns in he da ase . This algo i hm a es he i em and indica es he use who is a ed high.
I a emp s o ecommend i ems ha bes sui cus ome needs (in he o m o p oduc s o se ices).
A e y use ul ea u e in a ecommende sys em is he abili y o guess he p e e ences and use needs in
analyzing use beha io o o he use beha io o gene a e a pe sonalized ecommende [
14
]. The chie
pu pose o ou sys em is o iden i y mo ies based on use s’ iewing his o ies and a ings p o ided in
Symme y 2020,12, 185 4 o 27
he Mo ieLens sys em and da ase s [
24
] and o use a speci ic algo i hm o p edic mo ies. The esul s
a e e u ned o he use as a ecommende i em wi h he pa ame e s o he use . The illus a ion o
he ecommende sys em is p esen ed in Figu e 1, which explains simila i ies wi hin people o build a
mo ie ecommenda ion sys em o use s.
Use 1 Use 2
A
B
B
A
C
Simila
Recommend
Figu e 1. Simila i ies wi hin people o build a mo ie ecommenda ion sys em o use s.
Figu e 1shows how wo use s ha e a simila in e es in i ems A and B. When his occu s,
he simila i y index o bo h use s will be calcula ed. Fu he mo e, he sys em can ecommend i ems C
o o he use s because he sys em can de ec ha bo h use s ha e simila i ies in e ms o he i ems.
2.2. Sys em Design
In his ecommende sys em,
K
-Means algo i hm, bi ch algo i hm, mini-ba ch
K
-Means algo i hm,
mean-shi algo i hm, a ini y p opaga ion algo i hm, agglome a i e clus e ing algo i hm, and spec al
clus e ing algo i hm a e used o de e mine he bes pe o ming algo i hms in mo ie ecommenda ions
based on op imized
K
alues. A e applying se e al algo i hms, all spaces a e sea ched o ob ain he
use nea es neighbo in he same clus e and Top-
N
Lis o ecommende mo ies. Figu e 2shows an
o e iew o he low o se en exis ing algo i hms.
Symme y 2020,12, 185 5 o 27
S a
Selec Da ase
Usedda ase byeachuse a g a ing
wi h a o i egen e/ a o i e ags
Plo  heeach alueo Kwi h he
Silhoue eSco ea  ha  alue
Clus e ingda ase wi h7me hods
(K-Means,bi ch,mini-ba ch
K-Means,mean-shi ,a ini y
p opaga ion,
agglome a i eclus e ing,and
spec alclus e ing)
End
Clus e  heuse s a emo ies
Find henea es nodewi hEuclidean
Dis ance
Top-Nlis o  ecommenda ionmo ie
o simila i yuse
Meansqua ede o ,
socialne wo kanalysis,
DunnMa ix,A e ageSimila i y,
Compu a ionalTime,Associa ion
Rulewi hAp io iAlgo i hm,
Clus e ingPe o manceE alua ion
Algo i hmcompa ison
 ope o manceanalysis
Figu e 2. Flowcha ha con igu es a ecommenda ion sys em o mo ies.
2.3. Clus e ing Algo i hm
Clus e ing is an analy ical me hod ha was used as ea ly as 1939 by T yon, R. C [
25
]. Clus e ing
was i s used in psychology and hen apidly expanded o o he ields. Since he explosion o he
amoun o in o ma ion a ailable on he in e ne , many e o s ha e been made o educe he p oblem
o in o ma ion o e load. This o e load can be esol ed by using a clus e ing me hod. Clus e ing is
a classi ica ion o he same objec s om di e en g oups wi h pa i ions using exis ing da a in o a
new g oup, and each g oup o da a is iden i ied wi h a ce ain deg ee o dis ance. In clus e analysis,
he e is also a con as be ween pa ame ic and nonpa ame ic app oaches [
26
]. Da a clus e ing is a
echnique commonly used in a ious ields, such as da a mining, pa e n ecogni ion, image analyze,
and a i icial in elligence. Da a clus e ing is employed o educe a la ge amoun o da a by p o iding
he ca ego ies o classi ying he da a ha ha e a high deg ee o simila i y.
2.3.1. K-Means Clus e ing Algo i hm and Op imize KNumbe Clus e
K
-Means clus e ing is a me hod ha au oma ically di ides da ase s in o
k
g oups [
27
]. Resul s
selec ed he ini ial cen al clus e
k
o be i e a i ely e ined by being assigned o he nea es cen al
clus e . Each cen e o he Cjclus e is upda ed o an a e age sample o i s cons i uen s.
K
-Means is used o g oup app oaches gi en i s simplici y, e iciency, and lexibili y in calcula ions
especially conside ing a la ge amoun o da a.
K
-Means calcula e he clus e cen e in assigning objec s
o he closes clus e based on dis ance. When he midpoin does no change, he clus e ing algo i hm
seeks con e gence. Howe e ,
K
-Means lacks he abili y o choose he igh ini ial seeds and could
cause classi ica ion inaccu acies. Selec ing a andom s a ing seed can p oduce a locally good solu ion
ha is qui e in e io o inding he di ec
K
alue. The a ious ini ial seeds ha un on he same da ase
migh deli e di e en pa i ion esul s. Gi en a se o objec s
(x1
,
x2
,
. . .
,
xn)
, whe e each objec is
an m-dimensional ec o , he
K
-Means algo i hm aims a sepa a ing hese objec s o o m
k
g oups
au oma ically. Algo i hm 1p o ides he K-Means p ocedu e.

Symme y 2020,12, 185 6 o 27
Algo i hm 1 K-Means algo i hm and op imize knumbe clus e
Inpu : Selec ing da ase and using da ase by each use a g a ing wi h a o i e gen e/ ags;
Ou pu : Finishing wi h elease Top-Nlis o ecommenda ion mo ie o simila i y use ;
1: unc ion K-MEANS()
2: Choosing kini ial clus e cen e s
3: Cj,j=1, 2, 3 . . . , k;
4: Each xiis assigned o i s closes clus e
5: cen e based on he dis ance me ic
6: J=∑k
j=1∑i∈C emp ||xi−Mj||2,
7: whe e Mjdeno es he mean o da a poin s
8: in C emp;
9: Choosing he igh Knumbe o Clus e s;
10: clus e ing he use ’s a e mo ies;
11: Finding he nea es node o sea ch simila i y wi h
12: use wi h Euclidean dis ance;
13: i he e is no change hen
14: The algo i hm has con e ged and clus e ing ask
15: is ended, also ecalcula e he Mo Kclus e s
16: as he new clus e cen e s and go o;
17: end i
18: end unc ion
Op imize KNumbe Clus e
To o e come he limi a ions abo e, we op imized
K
o choose he co ec numbe o
K
clus e s.
Choosing he bes numbe o
K
clus e s is he key poin o he
K
-Means algo i hm. To ind
K
,
we calcula e he clus e ing e o wi h mean squa ed e o (MSE) and silhoue e sco e. Fi s , we selec a
da ase and choose he ange o
k
alues o es . Then, we de ine he unc ion o calcula e he clus e ed
e o s and e o alues o all k alues. Finally, we plo each alue o
k
wi h he silhoue e sco e a ha
alue. A de ailed explana ion o mean squa ed e o (MSE) and silhoue e sco e is p o ided in [
28
,
29
].
(1) Silhoue e Sco e o de e mine he Knumbe clus e .
The silhoue e was i s in oduced by
Pe e J. Rousseeuw in [
30
] in 1986. This is a me hod o in e p e a ion and alida ion o clea da a
clus e s. This echnique p o ides a g aphical ep esen a ion o how well each objec i s inside he
g oup. The silhoue e alue o an a ibu e is gi en by he equa ion below:
Si=a(i)·b(i)
max {a(i),b(i)}, (1)
whe e
a(i)
is he a e age dissimila i y o da a poin
i
wi h o he da a wi hin he one clus e . He e,
b(i)
is he minimum a e age dissimila i y o da a poin
i
wi h any o he clus e in which
i
is no
inside membe a.
2.3.2. Bi ch Clus e ing Algo i hm
Balanced i e a i e educing and clus e ing using hie a chies (bi ch clus e ing) uses a hie a chical
da a s uc u e ha calls CF- ee o inc emen and dynamically clus e s da a poin s [
31
]. The bi ch
algo i hm uses an inpu se o da a poin s
N
, which is ep esen ed as a ec o o eal alue and he
desi ed numbe o clus e
K
. The i s phase builds a CF- ee om a da a poin . This can be de ined
wi h gi en a se o
N
d-dimensional da a poin s, and he clus e ing ea u e (CF) o he se is used o
es ablish he iple- CF = (N,LS,SS), whe e
Symme y 2020,12, 185 7 o 27
−→
LS =
N
∑
i=1
−→
xi, (2)
is he linea sum and
−→
SS =
N
∑
i=1
−−→
(xi)2, (3)
is he sum o he da a poin s.
Clus e ing ea u es a e o ganized on a CF- ee, a heigh -balanced ee wi hin wo pa ame e s,
including b anching ac o B and h eshold T. Each non-lea node con ains mos B en ies inside he
o m [
CFi
, child
i
], in which child
i
is one poin e o i s
i
he child node and
CFi
is he clus e ing ea u e
ep esen ing he associa ed subclus e .
2.3.3. Mini-Ba ch K-Means Clus e ing
The algo i hm Mini-ba ch
K
-Means was de eloped as a modi ica ion o he
K
-Means algo i hm.
I uses mini-ba ch o educe ime in e y complex and la ge-scale calcula ions o da ase s. The e o s
o op imize g ouping esul s could be used wi h his me hod. Mini-ba ch
K
-Means a e andomly used
as inpu , which is a subse o he en i e da ase . Mini-ba ch
K
-Means is as e han
K
-Means and is
usually used o la ge da ase s. Fo da ase
T={x
1,
x
2,
. . .
,
xn}
,
xi∈Rm·n
,
xi
ep esen s a ne wo k
eco d wi h an
n
-dimensional eal ec o . In addi ion,
m
indica es he numbe o eco ds inside he
da ase
T
. The objec i e o he clus e ing p oblem is o unco e he se
C
o clus e cen e s
c∈Rm·n
o minimize he da ase
T
o eco ds
c∈Rm·n
in unc ion [
32
]. In con as o
K
-Means, mini-ba ch
K
-Means andomly selec s a subse o eco ds om he da ase . Mini-ba ch
K
-Means g ea ly educes
he clus e ing ime and he con e gence ime. The sum o squa ed dis ances is compu ed in one clus e
as ollows:
min ∑
x∈T
|| (C,x)−x||2, (4)
whe e
(C
,
x)
e u ns he closes o clus e cen e
c∈C
o eco d
x
, and
|C|=K
and
K
is he numbe
o clus e s o ob ain.
2.3.4. Mean-Shi Clus e ing
The mean-shi algo i hm is p oposed as a me hod o clus e analysis [
33
]. Howe e , gi en ha
he mean-shi de e mines he g adien ascen , he con e gence o he p ocess needs e i ica ion
and i s ela ion wi h simila algo i hms needs cla i ica ion. The mean-shi algo i hm is pa o a
nonpa ame ic g ouping echnique ha does no equi e p io knowledge o he numbe o clus e s
and does no cons ain he shape o he clus e s. Mean-shi clus e ing is used o disco e blobs in
a smoo h densi y o samples, which wo ks by upda ing he candida es o cen oids o be he mean
poin s wi hin a gi en egion. These candida es a e hen il e ed in a pos p ocessing s age o elimina e
nea -duplica es o o m he inal se o cen oids. Gi en a candida e cen oid
xi
ollowed by i e a ion
,
he candida e is upda ed by he ollowing equa ion:
x +1
i=m(x
i), (5)
Symme y 2020,12, 185 8 o 27
whe e
N(xi)
is he neighbo hood o samples wi hin a gi en dis ance a ound
xi
and
m
is he mean-shi
ec o o each cen oid ha poin s agains a egion o he maximum inc ease in he densi y o poin s.
This is compu ed using he ollowing equa ion:
m(xi) = ∑xj∈N(xi)K(xj−xi)xj
∑xj∈N(xi)K(xj−xi). (6)
The algo i hm can au oma ically de e mine he numbe o clus e s, using bandwid h pa ame e s.
This in o ma ion can de e mine he size o he egion o sea ch h ough. This algo i hm is un eachable
because i equi es se e al close neighbo sea ches du ing i s execu ion.
2.3.5. A ini y P opaga ion Clus e ing
In he a ini y p opaga ion me hod all da a poin s a e conside ed as possible exempla s.
I exchanges eal- alues be ween exempla s un il high-quali y exempla s and co esponding clus e s
a e no p o ided. The pa icula messages a e u he upda ed based on a simple o mula ha
a iocina es a sum-p oduc . A any selec ed poin in ime, he magni ude in each message ep esen s
he cu en a ini y ha one poin has o choosing ano he da a poin as i s exempla , hence he name
“a ini y p opaga ion” [
34
]. The messages sen be ween poin s belong o one o he wo ca ego ies.
The accumula ed e idence
(i
,
k)
o sample
k
should be he exempla o sample
i
. In addi ion,
ega ding a ailabili y
a(i
,
k)
, he accumula ed indica es ha sample
i
should choose sample
k
o be an
exempla . The exempla chosen by samples is simila enough o many samples ha a e ep esen a i e
o hemsel es. The esponsibili y o a sample
k
o be he exempla o sample
i
is gi en by he
ollowing o mula:
(i,k)←s(i,k)−max[a(i,k0) + s(i,k0)∀k06=k]. (7)
The simila i y be ween samples
i
and
k
,
s(i
,
k)
is assessed. The a ailabili y o sample
k
o be an
exempla o sample iis gi en by he ollowing o mula:
a(i,k)←min[0, (k,k) + ∑
i0s, ,i0∈{i,k}
(i0,k)]. (8)
We de ine a clus e wi h upda e
(i
,
k)
and
a(i
,
k)
. All he alues o
and
a
we e se o ze o and
each i e a e is calcula ed un il con e gence is ound. To a oid nume ical oscilla ions, he i e a ion
p ocess equi es he damping ac o γas ollows:
+1(i,k) = λ· (i,k) + (1+λ)· +1(i,k), (9)
a +1(i,k) = λ·a (i,k) + (1+λ)·a +1(i,k). (10)
2.3.6. Agglome a i e Clus e ing
Agglome a i e clus e ing can scale la ge numbe s o samples when used oge he wi h he
connec i i y ma ix and all possible me ges a e conside ed a each s ep. Wa d’s me hod is one o he
agglome a i e clus e ing me hods based on a classical sum-o -squa es c i e ion, p oducing g oups
ha minimize wi hin-g oup dispe sion a each bina y usion. This me hod uses Euclidean dis ance as
he dis ance me ic.
||a−b||2= ∑
i
(ai−bi)2. (11)
Symme y 2020,12, 185 9 o 27
2.3.7. Spec al Clus e ing
Spec al clus e ing uses in o ma ion om he eigen alue (spec um) o a special ma ix ha will
be buil om a g aph o da a se . This ma ix will be buil and in e p e i s spec um using eigen ec o s
o assign da a o clus e s. An eigen ec o is an impo an objec o linea algeb a and helps illus a e he
dynamics o he sys em ep esen ed by he ma ix. Speci ic g ouping uses he concep o eigen alues
and eigen ec o s.
L=D−1/2 ·AD−1/2. (12)
I s closes clus e cen e is assigned based on he dis ance me ic. The a ini y ma ix is o med,
and he diagonal ma ix is de ined. The ma ix is o med and he no malized Laplacian ma ix and
eigen ec o s a e compu ed.
2.4. E alua ion C i e ia
We u ilized he aining da a o de elop he o line model, and he emaining da a a e used
o analyze and p o ide he mo ie ecommenda ion. To e i y he quali y o he ecommende
sys em, we employed he mean squa ed e o (MSE), social ne wo k analysis (SNA), Dunn Ma ix
as clus e alidi y indices, and e alua ion measu es wi h a e age simila i y, compu a ional ime,
associa ion ule wi h Ap io i algo i hm, and clus e ing pe o mance e alua ion. The ollowing is
e i ied and e alua ed.
2.4.1. Mean Squa ed E o
Mean squa ed e o (MSE) is used o acili a e aining and some o e all e o measu e is o en
used as a pe o mance me ic o an objec i e unc ion [35].
MSE =1
M·1
N
M
∑
m=1
N
∑
j=1dmj −ymj2, (13)
whe e
dmj
and
ymj
ep esen he desi ed ( a ge ) alue and ou pu a he
m
he node o he
j
he
aining pa e n espec i ely,
M
is he numbe o ou pu nodes, and
N
is he numbe o he aining
pa e ns. The pu pose o aining is o de ec he se o weigh s ha minimize he objec i e unc ion.
Mean squa ed e o (MSE) is e y cle e in p o iding in o ma ion abou his a i icially buil
model. By minimizing he mean squa ed e o (MSE) alue, he a ian model is minimized. This can
p o ide ela i ely consis en esul s as inpu da a compa ed o models wi h la ge a ian s (la ge mean
squa ed e o (MSE)).
2.4.2. Clus e ing Validi y Indices: Dunn Ma ix
The Dunn index (DI) is a me ic o e alua ing clus e ing algo i hms wi h in e nal e alua ion
schemes, wi h esul s being based on he clus e da a i sel . Dunn’s index is he a io o wi hin and
be ween clus e sepa a ion (Malay K. Pakhi a, S. B., 2004). Simila o all o he indices, he pu pose o
he Dunn index is o iden i y a se o compac clus e s wi h small a ian s be ween clus e membe s,
ha a e well sepa a ed, and wi hin which he a e age clus e di e s signi ican ly compa ed o he
clus e a ian s.
The highe he Dunn index alue is, he be e he g ouping. The numbe o clus e s maximizing
he Dunn index will be aken as he op imal numbe o clus e s
k
. I also has se e al sho comings.
As he numbe o clus e s and da a dimensionali y inc ease, he cos o compu ing also inc eases.
The Dunn index o cnumbe o clus e s is de ined as ollows:
Symme y 2020,12, 185 16 o 27
(a)
(b)
Figu e 8.
Visualiza ion o agglome a i e clus e ing algo i hm. Agglome a i e clus e ing gen e (
a
),
agglome a i e clus e ing ag (b).
To ob ain a mo e delimi ed subse o people o s udy, g ouping is pe o med o exclusi ely ob ain
a ings om hose who like ei he omance o science ic ion mo ies. X and Y-axes a e omance and
sci- i a ings, espec i ely. In addi ion, he size o he do ep esen s he a ings o he ad en u e
mo ies. The bigge he do , he highe he ad en u e a ing. The addi ion o he ad en u e gen e
signi ican ly al e s he clus e ing. The Top
N
-Mo ies lis s se e al clus e ing algo i hms wi h
K
-Means
gen e
n
clus e = 12,
K
-Means ag
n
clus e = 7, bi ch gen e
n
clus e = 12, bi ch ags
n
clus e = 12,
MiniBa ch-
K
-Means gen e
n
clus e = 12, MiniBa ch-
K
-Means ags
n
clus e s = 7, mean-shi gen e,
mean-shi ags, a ini y p opaga ion gen e, a ini y p opaga ion ags, agglome a i e clus e ing gen e
n
clus e = 12, agglome a i e clus e ing ag
n
clus e = 7, spec al clus e ing gen e, spec al clus e ing
ags a e epo ed below.
(a)
(b)
Figu e 9.
Visualiza ion o spec al clus e ing algo i hm. Spec al clus e ing gen e (
a
), spec al clus e ing
ag (b).

Symme y 2020,12, 185 17 o 27
Figu e 10. Example o isualiza ion Gen e K-Means o Top lis -No mo ies.
Conside ing a subse o use s and disco e ing wha hei a o i e gen e was, we de ine a
unc ion ha would calcula e each use ’s a e age a ing o all omance mo ies, science ic ion
mo ies, and ad en u e mo ies. To ob ain a mo e delimi ed subse o people o s udy, we biased ou
g ouping o exclusi ely ob ain a ings om hose use s who who like ei he omance o science ic ion
mo ies. We used he
x
and
y
-axes o he omance and sci- i a ings. In addi ion, he size o he do
ep esen s he a ings o he ad en u e mo ies ( he bigge he do , he highe he ad en u e a ing).
The addi ion o he ad en u e gen e signi ican ly a ec s he clus e s. The mo e da a added o ou
model, he mo e simila he p e e ences o each g oup a e. The inal e sion is Top lis o
N
o mo ies
(e.g., shown in Figu e 10). Addi ionally, we conside ed a subse o use s and disco e ed hei a o i e
ags. We de ined a unc ion ha calcula ed each use ’s a e age a ing o all unny ag mo ies, an asy
ag mo ies, and ma ia ag mo ies. To ob ain a mo e delimi ed subse o people o s udy, we biased ou
g ouping o exclusi ely ob ain a ings om hose use s who like ei he unny o an asy ags mo ies.
We also de e mined ha esul s ob ained be o e he compa ison algo i hm a e Top N mo ies o be
gi en o simila use s. The esul s o Top N mo ies be o e he compa ison algo i hm and Top N mo ies
o gi e o simila use s o in e es in a o i e gen e and ags a e p esen ed.
Op imize KNumbe Clus e
F om he esul s ob ained, we can choose he bes choices o he
K
alues. Choosing he igh
numbe o clus e s is one o he key poin s o he
K
-Means algo i hm. We also use mini-ba ch
K
-Means
algo i hm and bi ch algo i hm. We do no apply o mean-shi and a ini y p opaga ion because he
algo i hm au oma ically se s he numbe o clus e s. Inc easing he numbe o clus e s shows he
ange ha esul ed in he wo s clus e s based on he Silhoue e Sco e. Op imize K is ep esen ed by a
silhoue e sco e. The X-axis ep esen s he la ges sco e, so he g oup is mo e a ied o use and he Y
axis ep esen s he numbe o clus e s. This is so ha i can de e mine he igh numbe o clus e s o
be used in displaying isualiza ions. The esul s o op imize
K
in se e al clus e ing algo i hms a e
p esen ed in Figu e 11.
Symme y 2020,12, 185 18 o 27
Figu e 11. Example o isualiza ion op imiza ion o Kin K-Means gen e a ing.
4. E alua ion and Discussion
This sec ion con ains he e i ica ion and e alua ion esul s o he me hodology. The bes
pe o ming me hod is iden i ied, and a discussion is p esen ed.
4.1. E alua ion Resul
The e i ica ion and e alua ion esul s a e p esen ed below.
4.1.1. Mean Squa ed E o
Shown in Figu es 12 and 13, he mean squa ed e o (MSE) agglome a i e me hod se es as an
example among he se en clus e ing algo i hms.
0
0.01
0.02
0.03
0.04
0.05
0.06
0.07
0 2 4 6 8 10 12
Mean Squa ed E o Sco e
Clus e
MSE Agglome a i e Clus e ing: Gen e
Mean Squa ed E o
Figu e 12. MSE agglome a i e clus e ing gen e.
Symme y 2020,12, 185 19 o 27
0
0.05
0.1
0.15
0.2
0.25
0 1 2 3 4 5 6 7
Mean Squa ed E o Sco e
Clus e
MSE Agglome a i e Clus e ing: Tags
Mean Squa ed E o
Figu e 13. MSE agglome a i e clus e ing ags.
4.1.2. Clus e Validi y Indices: Dunn Ma ix
The Dunn ma ix is used as a alidi y measu e o compa e pe o mance me hods o
ecommenda ion sys ems. The ollowing Dunn ma ix esul s a e shown in Table 1.
Table 1. Dunn Ma ix o se en clus e ing algo i hms.
Me hods Amoun o Clus e s Sco e
K-Means algo i hm: gen e a ing 3 0.41
K-Means algo i hm: ags a ing 3 0.41
bi ch algo i hm: gen e a ing 3 0.49
bi ch algo i hm: ags a ing 3 0.63
mini-ba ch K-Means algo i hm: gen e a ing 3 0.38
mini-ba ch K-Means algo i hm: ags a ing 3 0.37
mean-shi algo i hm: gen e a ing – 0.39
mean-shi algo i hm: ags a ing – 0.63
a ini y p opaga ion algo i hm: gen e a ing – 1.06
a ini y p opaga ion algo i hm: ags a ing – 0.47
agglome a i e clus e ing algo i hm: gen e a ing 3 0.43
agglome a i e clus e ing algo i hm: ags a ing 3 0.45
spec al clus e ing algo i hm: gen e a ing – 2.09
spec al clus e ing algo i hm: ags a ing – 4.60
4.1.3. A e age Simila i y
The a e age simila i y bi ch me hod example esul s om se en clus e ing algo i hms a e shown
in Table 2.
Symme y 2020,12, 185 20 o 27
Table 2. A e age simila i y o bi ch me hod.
Me hods Amoun A e age
o Clus e s Simila i y
bi ch algo i hm: gen e a ing
3 0.99
6 0.97
12 0.96
bi ch algo i hm: ags a ing
4 0.97
6 0.93
12 0.94
4.1.4. Social Ne wo k Analysis
Mean-shi esul s om se en clus e ing me hods o social ne wo k analysis (SNA) a e shown in
Table 3.
Table 3. Mean-shi social ne wo k analysis esul .
Me hods Clus e SNA
mean-shi algo i hm: gen e a ing Clus e : (0, 2, 1, 4)
Deg ee, densi y in clus e 0 ( he highes numbe
in esul ), compa e o sequence lis o he clus e
whe e he dis ance esul is 5831.49.
Closeness, he highes in clus e 0 o clus e 1 (bo h
clus e s ha e a high linkage ela ionship) whe e he
dis ance esul is 1.44.
Be weenness, he highes in clus e 0 ( i s ) o
clus e 2 (end), be ween 1 (clus e 0 is he mos
ha e a ela ionship whe e he clus e 1(be ween)
and 2 whe e he dis ance esul is 1.39.
mean-shi algo i hm: ags a ing Clus e : (3, 0, 1, 2)
Deg ee, he densi y in clus e 0 ( he highes numbe
in esul ), compa e o sequence lis o he clus e
whe e he dis ance esul is 148.11.
Closeness, he highes in clus e 0 o clus e 1 (bo h
clus e s ha e a high linkage ela ionship) whe e he
dis ance esul is 0.67.
Be weenness, he highes in clus e 3 ( i s ) o
clus e 1 (end), be ween 0 (clus e 3 is he mos
ha e a ela ionship wi h clus e 0 (be ween) and 1
whe e he dis ance esul is 3.12.
4.1.5. Associa ion Rule: Ap io i Algo i hm
The associa ion ule wi h Ap io i algo i hm is used as an e alua ion measu e o compa e he
me hod pe o mance o he ecommenda ion sys ems. An example o esul s o associa ion ules wi h
he Ap io i algo i hm om se en clus e ing algo i hms is shown in he ollowing sec ion.
Rule: 12 Ang y men (1957) -> Ad en u es o P iscilla, Queen o he Dese , The (1994)
Suppo : 0.25
Con idence: 1.0
Li : 4.0
========================================
Rule: 12 Ang y Men (1957) -> Ai plane! (1980)
Suppo : 0.25
Con idence 1.0
Li : 4.0
========================================
Rule: Amadeus (1984) -> 12 Ang y men (1957)
Symme y 2020,12, 185 21 o 27
Suppo : 0.25
Con idence: 1.0
Li : 4.0
========================================
Rule: Ame ican Beau y (1999) -> 12 Ang y Men (1957)
Suppo : 0.25
Con idence: 1.0
Li : 4.0
========================================
Rule: Aus in Powe s: In e na ional Man o Mys e y (1997) -> 12 Ang y Men (1957)
Suppo : 0.25
Con idence: 1.0
Li : 4.0
4.1.6. Compu a ional Time
The compu a ional ime is used as an e alua ion measu e o compa e pe o mance me hods o
ecommenda ion sys ems. Compu a ional ime esul s a e epo ed below (see Table 4).
Table 4. Compu a ional ime o se en clus e ing algo i hms.
Clus e ing Me hod Compu a ional Time [ms]
K-Means-gen e 31.16
K-Means- ags 14.43
bi ch-gen e 24.49
bi ch- ags 15.34
mini-ba ch K-Means me hod-gen e 23.82
mini-ba ch K-Means me hod- ags 15.79
mean-shi -gen e 13.75
mean-shi - ags 10.15
a ini y p opaga ion-gen e 20.04
a ini y p opaga ion- ags 8.53
agglome a i e clus e ing-gen e 32.00
agglome a i e clus e ing- ags 10.37
spec al clus e ing-gen e 15.55
spec al clus e ing- ags 6.22
4.1.7. Clus e ing Pe o mance E alua ion
Clus e ing pe o mance e alua ion (CPE) esul o
K
-Means and bi ch me hod examples om
se en clus e ing algo i hms a e epo ed below (see Table 5).

Symme y 2020,12, 185 22 o 27
Table 5. Clus e ing pe o mance e alua ion o K-Means me hod and bi ch me hod.
Me hods
Clus e ing
Sco ePe o mance
E alua ion
K-Means algo i hm: gen e a ing
Silhoue e Coe icien 0.29
Calinski Ha abaz Index 59.41
Da ies Bouldin Index 1.13
K-Means algo i hm: ags a ing
Silhoue e Coe icien 0.25
Calinski Ha abaz Index 7.47
Da ies Bouldin Index 0.86
bi ch algo i hm: gen e a ing
Silhoue e Coe icien 0.23
Calinski Ha abaz Index 39.03
Da ies Bouldin Index 1.24
bi ch algo i hm ags a ing
Silhoue e Coe icien 0.25
Calinski Ha abaz Index 5.73
Da ies Bouldin Index 1.16
4.2. Discussion
A de ailed explana ion o he abo e-men ioned expe imen s is discussed in he subsequen
sec ion. In gene al, o all o hese me hods, he highe he alue o he li , suppo , and con idence,
he be e he link is o he ecommende sys em. Fu he , he highe he Dunn index alue, he be e
he g ouping.
4.2.1. K-Means Pe o mance
Mo ie ecommende quali y is e alua ed wi h
K
-Means. MSE esul s om
K
-Means show
di e en esul s o gen e a ing and a ing ags. The a ing ag esul s a e ela i ely smalle wi h
a ing gen e sco es o 0–0.95 and a ing ags sco es o 0–0.28.
K
-Means has a Dunn Ma ix ha ends
o be e enly dis ibu ed o gen e and ags wi h alues o 0.41. The highe he Dunn index alue,
he be e he g ouping. The a e age simila i y in he gen e showed ha he alue inc eases as he
numbe o clus e s dec eases. The a e age simila i y in
K
-Means ags shows ha he high simila i y
alue depends on he numbe o clus e s. The associa ion ule wi h Ap io i algo i hm in
K
-Means
clus e ing app oach 13% suppo o he gen e and 25% o ags o cus ome s who choose mo ies A
and B. Suppo is an indica ion o how o en he i emse appea s in linkages. The con idence is 61% o
he gen e and 100% o ags o he cus ome s who choose mo ie A and mo ie B. Li ep esen s he
a io o 3.3 o gen e and 4.0 o ags o he obse ed suppo alue. This is a condi ional p obabili y.
Clus e ing pe o mance e alua ion showed ha he
K
-Means me hod showed good pe o mance wi h
he Calinski-Ha abaz Index wi h a sco e o 59.41.
4.2.2. Bi ch Pe o mance
To e alua e he mo ie ecommende quali y wi h bi ch, mean squa ed e o (MSE) esul s om
bi ch showed ela i ely small esul s wi h a a ing gen e sco e ange o 0–0.25 and ag sco es o 0–0.17.
Bi ch ag a ings ha e a Dunn Ma ix alue ha ends o be g ea e han 0.64. The a e age simila i y
in he bi ch gen e showed ha he alue inc eases, as he numbe o clus e s dec eases. The a e age
simila i y in bi ch ags e ealed ha he high simila i y alue depended on he numbe o clus e s.
The associa ion ule wi h Ap io i algo i hm in he bi ch clus e ing app oach p o ides 16% suppo o
gen e and 50% suppo o ags o cus ome s who choose mo ies A and B. Suppo is an indica ion
o how o en he i emse appea s in linkages. Con idence is 100% o he gen e and 50% o ags o
Symme y 2020,12, 185 23 o 27
he cus ome s who choose mo ie A and B. Li ep esen s he a io o 4.0 o gen e and 1.0 o ags
o he obse ed suppo alue. The pe o mance e alua ion showed ha his me hod p o ides good
pe o mance wi h a sco e o 1.24 on he Da ies–Bouldin Index.
4.2.3. Mini-Ba ch K-Means Pe o mance
To e alua e mo ie ecommende quali y wi h mini-ba ch
K
-Means, MSE esul s om mini-ba ch
K
-Means showed di e en esul s in gen e a ing and a ing ags. The a ing ag esul s we e ela i ely
smalle wi h a a ing gen e sco e ange o 0–0.69 and a ing ag sco es o 0–0.19. The mini-ba ch
K
-Means wi h gen e a ing has a Dunn Ma ix alue ha ends o be g ea e han 0.38. The a e age
simila i y in he mini-ba ch
K
-Means gen e showed ha he high simila i y alue depended on he
numbe o clus e s. The a e age simila i y in he mini-ba ch
K
-Means ags showed ha he high
simila i y alue depended on he numbe o clus e s. The associa ion ule wi h Ap io i algo i hm in
mini-ba ch
K
-Means clus e ing app oach p o ides 13% suppo o gen e and 14% suppo o ags o
cus ome s who choose mo ies A and B. Suppo is an indica ion o how o en he i emse appea s in
linkages. Con idence is 100% o he gen e and 100% o ags o he cus ome s who choose mo ie A
and mo ie B. Li ep esen s he a io o 3.75 o gen e and 7.0 o ags o he obse ed suppo alue.
The e alua ion showed ha he mini-ba ch
K
-Means me hod pe o ms well wi h Calinski-Ha abaz
Index wi h a sco e o 48.18.
4.2.4. Mean-Shi Pe o mance
To e alua e he mo ie ecommende quali y wi h he mean-shi , he MSE esul s om he
mean-shi showed ela i ely la ge esul s wi h a gen e a ing sco e ange o 0–1 and a ag sco e o
0–1. The mean-shi algo i hm wi h a ing ags has a Dunn Ma ix alue ha ends o be g ea e han
0.63. The a e age simila i y in he mean-shi gen e showed ha he alue inc eases as he numbe o
clus e s dec eases. The a e age simila i y in ags mean-shi showed ha he high simila i y alue
depended on he numbe o clus e s. The mean-shi in he gen e has he bes compu a ional ime a
13.75 ms. The associa ion ule wi h Ap io i algo i hm in mean-shi clus e ing app oach p o ides 12%
suppo o gen e and 9% o ags o cus ome s who choose mo ies A and B. Suppo is an indica ion
o how o en he i emse appea s in linkages. Con idence is 81% o he gen e and 100% o ags o he
cus ome s who choose mo ie A and mo ie B. Li ep esen s a a io o 3.06 o gen e and 5.5 o ags o
he obse ed suppo alue. The abo e-men ioned e alua ion depic s he a ini y p opaga ion me hod
o p o ide su icien pe o mance wi h Calinski-Ha abaz Index wi h a sco e o 20.14.
4.2.5. A ini y P opaga ion Pe o mance
To e alua e he quali y wi h a ini y p opaga ion mo ie, he esul s o he mean squa ed e o
(MSE) om a ini y p opaga ion showed di e en esul s o gen e a ing and a ing ags. The a ing
gen e esul s we e ela i ely smalle wi h a gen e a ing sco e ange o 0–0.17 and a a ing ag sco e o
0–0.89. The a ini y p opaga ion algo i hm wi h gen e a ing has a Dunn Ma ix which ends o be
highe a 1.06. The a e age simila i y in he a ini y p opaga ion gen e showed ha he high simila i y
alue depended on he numbe o clus e s. The a e age simila i y in a ini y p opaga ion ags showed
ha he high simila i y alue depended on he numbe o clus e s. The associa ion ule wi h Ap io i
algo i hm in a ini y p opaga ion clus e ing app oach p o ided 10.5% suppo o he gen e and 20%
suppo o ags o cus ome s who choose mo ies A and B. Suppo is an indica ion o how o en he
i emse appea s in linkages. He e, 66% con idence is no ed o he gen e and 100% o ags o he
cus ome s who choose mo ie A and mo ie B. Li ep esen s he a io o 4.22 o gen e and 5.0 o
ags o he obse ed suppo alue. Clus e ing pe o mance e alua ion also showed ha he a ini y
p opaga ion me hod showed good pe o mance wi h he Calinski-Ha abaz Index wi h a sco e o 53.49.
Symme y 2020,12, 185 24 o 27
4.2.6. Agglome a i e Clus e ing Pe o mance
Now we e alua e he mo ie ecommende quali y wi h agglome a i e clus e ing. MSE esul s
om agglome a i e clus e ing showed di e en esul s in gen e a ing and a ing ags. The a ing gen e
esul s we e ela i ely smalle wi h a ing gen es sco es o 0–0.06 and a ing ags sco es wi h 0–0.23.
Agglome a i e clus e ing algo i hms wi h a ing ags ha e Dunn Ma ix esul s ha end o be g ea e
han 0.45. The a e age simila i y in he agglome a i e clus e ing gen e showed ha he high simila i y
alue depended on he numbe o clus e s. The a e age simila i y in agglome a i e clus e ing ags
showed ha he high simila i y alue depended on he numbe o clus e s. The associa ion ule wi h
he Ap io i algo i hm in agglome a i e clus e ing app oach p o ides 22% suppo o he gen e and
16% o ags o cus ome s who choose mo ies A and B. Suppo is an indica ion o how o en he
i emse appea s in linkages. In addi ion, 22% con idence is no ed o he gen e and 16% o ags o
cus ome s who choose mo ie A and mo ie B. Li ep esen s a a io o 1.0 o gen e and 1.0 o ags
o he obse ed suppo alue. F om he pe o mance e alua ion, we see ha he agglome a i e
clus e ing me hod pe o ms well wi h he Calinski-Ha abaz Index wi h a sco e o 49.34.
4.2.7. Spec al Clus e ing Pe o mance
Now we e alua e he mo ie ecommende quali y wi h spec al clus e ing. Mean squa ed e o
(MSE) esul s om spec al clus e ing showed di e ences in gen e a ing and a ing ags. The a ing
ag esul s we e ela i ely smalle wi h a a ing gen e sco e ange o 0–0.62 and a ing ags sco es
o 0–0.17. Spec al clus e ing algo i hm wi h ag a ing has he bes Dunn Ma ix esul s a 4.61
and he bes spec al clus e ing algo i hm esul s wi h gen e a 2.09. The a e age simila i y in he
spec al clus e ing gen e showed ha he high simila i y alue depended on he numbe o clus e s.
The a e age simila i y in spec al clus e ing ags showed ha he high simila i y alue depended
on he numbe o clus e s. Spec al clus e ing in ags has he bes compu a ional ime a 6.22 ms.
The associa ion ule wi h Ap io i algo i hm in spec al clus e ing app oach p o ides 12% suppo o
he gen e and 33% suppo o ags o cus ome s who choose mo ies A and B. Suppo is an indica ion
o how o en he i emse appea s in linkages. Con idence is 75% o he gen e and 100% o ags o
he cus ome s who choose mo ie A and mo ie B. Li ep esen s he a io o 3.12 o gen e and 3.0
o ags o he obse ed suppo alues. Clus e ing pe o mance e alua ion showed ha he spec al
clus e ing me hod showed good pe o mance wi h he Calinski-Ha abaz Index wi h a sco e o 16.39.
5. Conclusions
In his s udy, se en clus e ings we e used o clus e pe o mance compa ison me hods o mo ie
ecommenda ion sys ems, such as he
K
-Means algo i hm, bi ch algo i hm, mini-ba ch
K
-Means
algo i hm, mean-shi algo i hm, a ini y p opaga ion algo i hm, agglome a i e clus e ing algo i hm,
and spec al clus e ing algo i hm. The de eloped op imized g oupings om se e al algo i hms we e
hen used o compa e he bes algo i hms wi h ega d o he simila i y g oupings o use s on mo ie
gen e, ags, and a ing using he Mo ieLens da ase . Then, op imizing
K
o each clus e did no
signi ican ly inc ease he a iance. To be e unde s and his me hod, a iance e e s o he e o .
One o he ways o calcula e his e o is o ex ac he cen oid o i s espec i e g oups. Then, his alue
is squa ed ( o emo e he nega i e e ms) and all hose alues a e added o ob ain he o al e o .
To e i y he quali y o he ecommende sys em, he mean squa ed e o (MSE), Dunn Ma ix as
clus e alidi y indices and social ne wo k analysis (SNA) we e used. In addi ion, a e age simila i y,
compu a ional ime, associa ion ule wi h Ap io i algo i hm and clus e ing pe o mance e alua ion
measu es we e used o compa e he me hods o pe o mance sys ems.
Using he Mo ieLens da ase , expe imen e alua ion o he se en clus e ing me hods e ealed
he ollowing:
1.
The bes mean squa ed e o (MSE) alue is p oduced by he bi ch me hod wi h a ela i ely small
squa ed e o sco e in he a ing gen e and a ing ags.
Symme y 2020,12, 185 25 o 27
2.
Spec al clus e ing algo i hm wi h ag a ing has he bes Dunn Ma ix esul s a 4.61 and he
spec al clus e ing algo i hm has he bes gen e esul s a 2.09. The highe he Dunn index alue
is, he be e he g ouping.
3.
The closes dis ance o he social ne wo k analysis (SNA) is he mean-shi me hod,
which indica es ha he dis ance be ween clus e s has a high linkage ela ionship in a iance.
4.
The bi ch me hod had a ela i ely high a e age simila ly o inc ease he numbe o clus e s,
which showed a good le el o simila i y in clus e ing.
5.
The bes compu a ional ime is indica ed by he mean-shi o gen e a 13.75 ms and spec al
clus e ing o ags a 6.22 ms.
6.
Visualiza ion o clus e ing and op imizing
k
o mo ie gen e in algo i hms is be e han mo ie
ags because ewe da a a e used o mo ie ags.
7.
Mini-ba ch
K
-Means clus e ing app oach is he bes app oach o he associa ion ule wi h
Ap io i algo i hm wi h a high sco e o suppo , 100% con idence, and 7.0 a io o li o
i em ecommenda ions.
8.
Clus e ing pe o mance e alua ion shows ha he
K
-Means me hod exhibi s good pe o mance
wi h he Calinski-Ha abaz Index wi h a sco e o 59.41, and he bi ch algo i hm wi h a sco e o
1.24, on he Da ies–Bouldin Index.
9. Bi ch is he bes me hod based on a compa ison o se e al pe o mance ma ices.
Au ho Con ibu ions:
W i ing—o iginal d a p epa a ion, D.C.G.P.; w i ing— e iew and edi ing, P.S.;
supe ision, J.-S.L. All au ho s ha e ead and ag eed o he published e sion o he manusc ip .
Funding: This esea ch ecei ed no ex e nal unding.
Acknowledgmen s:
This esea ch was suppo ed by he Minis y o Science and Technology (MOST) unde he
g an MOST-108-2221-E-011-061- and MIT Labo a o y, Na ional Taiwan Uni e si y o Science and Technology. 2.
Fo he esea ch, in as uc u e o he SIX Cen e was used.
Con lic s o In e es :
The au ho s decla e no con lic o in e es .The unde s had no ole in he design o he s udy;
in he collec ion, analyses, o in e p e a ion o da a; in he w i ing o he manusc ip , o in he decision o publish
he esul s.
Re e ences
1.
Isinkaye, F.; Folajimi, Y.; Ojokoh, B. Recommenda ion sys ems: P inciples, me hods and e alua ion.
Egyp . In o m. J. 2015,16, 261–273. [C ossRe ]
2.
Nilashi, M.; Salahshou , M.; Ib ahim, O.; Ma dani, A.; Es ahani, M.D.; Zakuan, N. A new me hod o
collabo a i e il e ing ecommende sys ems: The case o yahoo! mo ies and ipad iso da ase s. J. So
Compu . Decis. Suppo Sys . 2016,3, 44–46.
3.
Smi h, B.; Linden, G. Two decades o ecommende sys ems a Amazon. com. IEEE In . Compu .
2017,21, 12–18. [C ossRe ]
4.
G eens ein-Messica, A.; Rokach, L. Pe sonal p ice awa e mul i-selle ecommende sys em: E idence om
eBay. Knowl. Sys . 2018,150, 14–26. [C ossRe ]
5.
I mazi, J.; Megías, M. Using ecommenda ion sys ems in cou se managemen sys ems o ecommend
lea ning objec s. In . A ab J. In o m. Technol. 2008,5, 234–240.
6.
Kuma , M.; Yada , D.; Singh, A.; Gup a, V.K. A mo ie ecommende sys em: Mo ec. In . J. Compu . Appl.
2015,124, 7–11. [C ossRe ]
7.
Lu, J.; Wu, D.; Mao, M.; Wang, W.; Zhang, G. Recommende sys em applica ion de elopmen s: A su ey.
Decis. Suppo Sys . 2015,74, 12–32. [C ossRe ]
8.
Shah, N.; Mahajan, S. Documen clus e ing: A de ailed e iew. In . J. Appl. In o m. Sys .
2012
,4, 30–38.
[C ossRe ]
9.
Yang, M.S.; Sinaga, K.P. A Fea u e-Reduc ion Mul i-View
k
-Means Clus e ing Algo i hm. IEEE Access
2019,7, 114472–114486. [C ossRe ]
10.
Wu, J.L.; Chang, P.C.; Tsao, C.C.; Fan, C.Y. A pa en quali y analysis and classi ica ion sys em using
sel -o ganizing maps wi h suppo ec o machine. Appl. So Compu . 2016,41, 305–316. [C ossRe ]