scieee Science in your language
[en] (orig)

On Member Labelling in Social Networks

Abstract

Software agents are increasingly used to search for experts, recommend resources, assess opinions, and other similar tasks in the context of social networks, which requires to have accurate information that describes the features of the members of the network. Unfortu-nately, many member profiles are incomplete, which has motivated many authors to work on automatic member labelling, that is, on techniques that can infer the null features of a member from his or her neighbour-hood. Current proposals are based on local or global approaches; the former compute predictors from local neighbourhoods, whereas the lat-ter analyse social networks as a whole. Their main problem is that they tend to be inefficient and their effectiveness degrades significantly as the percentage of null labels increases. In this paper, we present Katz, which is a novel hybrid proposal to solve the member labelling problem using neural networks. Our experiments prove that it outperforms other pro-posals in the literature in terms of both effectiveness and efficiency.

Read accessible full text

On Member Labelling in Social Networks

Author: Corchuelo Gil, Rafael; Reina Quintero, Antonia María; Jiménez Aguirre, Patricia
Publisher: Springer
Year: 2015
DOI: 10.1007/978-3-319-19222-2_41
Source: https://idus.us.es/bitstreams/98170052-b63e-4cdb-9c9c-d52e81d4321d/download
On Membe Labelling in Social Ne wo ks
Ra ael Co chuelo(B), An onia M. Reina Quin e o, and Pa icia Jim´enez
ETSI In o m´a ica, A da. Reina Me cedes s/n, E-41012 Se illa, Spain
{co chu, einaqu,pa iciajimenez}@us.es
Abs ac . So wa e agen s a e inc easingly used o sea ch o expe s, ecommend
esou ces, assess opinions, and o he simila asks in he con ex o social ne wo ks,
which equi es o ha e accu a e in o ma ion ha desc ibes he ea u es o he membe s
o he ne wo k. Un o u-na ely, many membe p ofiles a e incomple e, which has
mo i a ed many au ho s o wo k on au oma ic membe labelling, ha is, on echniques
ha can in e he null ea u es o a membe om his o he neighbou -hood. Cu en
p oposals a e based on local o global app oaches; he o me compu e p edic o s om
local neighbou hoods, whe eas he la - e analyse social ne wo ks as a whole. Thei
main p oblem is ha hey end o be inefficien and hei effec i eness deg ades
significan ly as he pe cen age o null labels inc eases. In his pape , we p esen Ka z,
which is a no el hyb id p oposal o sol e he membe labelling p oblem using neu al
ne wo ks. Ou expe imen s p o e ha i ou pe o ms o he p o-posals in he li e a u e
in e ms o bo h effec i eness and efficiency.
Keywo ds: Social ne wo ks · Membe labelling · Hyb id app oach · Neu al
ne wo ks
1 In oduc ion
On-line social media ha e sp ou ed ou du ing he las decade. They ha e pa ed
he way o on-line social ne wo ks whose membe s ypically in e ac o sha e o
o e ie e in o ma ion om one ano he . Ne e be o e has i been easie o find
in o ma ion abou indi iduals, hei demog aphics, hei likes, hei dislikes, he
ac i i ies in which hey engage, hei opinions, hei hough s, and so on. And
some hing ha is e en mo e impo an : hei ela ionships.
So wa e agen s a e being used in asks such as sea ching o expe s ega d-
ing a gi en opic, ecommending esou ces (pos s, ideos, music, and he like),
assessing opinions, a ge ing ad e isemen s, sociological s udies, and so on. Fo
hese agen s o succeed in p oducing accu a e in o ma ion, i is e y impo -
an ha he in o ma ion in a membe ’s p ofile be as comple e as possible.
Un o una ely, i is no uncommon ha many membe s do no comple e hei
p ofiles [11], which makes i e y difficul o so wa e agen s o wo k well.
Many au ho s ha e paid a en ion o a p oblem ha is commonly e e ed o
as membe labelling (aka. membe classifica ion, node classifica ion, link-based
classifica ion, o collec i e classifica ion). Simply pu , he idea is o in e he
ea u es o a membe o a social ne wo k as accu a ely as possible using solely
he ea u es a ailable om membe s wi h whom he o she has a ela ionship [1].
This has p o en o wo k well because social ne wo ks ha e a p ope y ha is
known as homophily [19], acco ding o which membe s who ha e simila ea-
u es end o ha e s onge ela ionships han membe s ha ha e e y dissimi-
la ea u es. The cu en p oposals in he li e a u e a e based on local o global
me hods. The o me lea n a p edic o om he ea u es o he membe s o a
social ne wo k, including some neighbou s; he la e ackle he p oblem om a
global pe spec i e and a emp o analyse social ne wo ks as a whole. The main
p oblem wi h cu en p oposals is ha hey ha e p o en o be inefficien and
ineffec i e as he size o a social ne wo k o he numbe o null ea u es inc eases.
This mo i a ed us o wo k on Ka z, which is a no el hyb id p oposal o sol e
he membe labelling p oblem. I is based on neu al ne wo ks, which a e used o
in e a p edic o o each membe ea u e using he in o ma ion p o ided by an
unbounded neighbou hood. I s a s analysing each membe ’s p ofile in isola ion,
and hen explo es his o he neighbou hood sea ching o he ela ionships and
ea u es ha con ibu e he mos o p oducing a be e p edic o . I is no a local
me hod since i explo es an unbounded con ex and selec s he mos in e es ing
ea u es and neighbou s o lea n a p edic o ; nei he is i a global me hod because
i does no a emp o analyse social ne wo ks as a whole; ha is he eason why
we e e o Ka z as a hyb id p oposal. Ou expe imen s on qui e a la ge eal-
wo ld social ne wo k p o e ha i ou pe o ms o he p oposals in he li e a u e
in e ms o bo h effec i eness and efficiency.
The es o he pape is o ganised as ollows: Sec ion 2desc ibes ou p oposal;
Sec ion 3 epo s on he esul s o ou expe imen s; Sec ion 4su eys he ela ed
wo k and compa es i o ou s; finally, Sec ion 5p esen s ou conclusions.
2 Ou P oposal
Ka z wo ks on a social ne wo k ha is ep esen ed as a g aph in which a node
ep esen s all o he ea u es o a membe p ofile and an edge ep esen s a ela-
ionship o ano he membe . I analyses he ne wo k and e u ns a map in which
each ea u e is associa ed wi h a se o neu al ne wo ks ha can be used o label a
new membe ega ding ha ea u e. No e ha each ea u e is p edic ed by means
o a se o neu al ne wo ks ha a e lea n om diffe en pa i ions o he social
ne wo k; he goal, which has been confi med empi ically, is o dec ease he e o
a e by using an ensamble-p edic o app oach ins ead o he single-p edic o
app oach ha is common in he li e a u e. In he ollowing subsec ions, we fi s
p esen he main p ocedu e o Ka z and hen an ancilla y p ocedu e ha is used
o ex end a neu al ne wo k o he mos app op ia e neighbou hood.
Main P ocedu e: Figu e 1shows he main p ocedu e o Ka z. I wo ks on a
g aph (N,E) ha ep esen s a social ne wo k. Nis a collec ion o ec o s o
he o m (m, 1, 2,..., n), whe e mis he unique iden ifie o a membe o he
social ne wo k and ia e he alues o i s ea u es (i=1...n); ea u es can be
1: Ka z(N,E)
2: m=∅
3: o each ea u e used in Ndo
4: ns =∅
5: epea β imes
6: = selec nodes in Nwi h a alue o
7: s = c ea e a aining se wi h γ| | nodes om
8: s = s
9: n=null
10: do
11: (n, s, s)=expandNeu alNe wo k(n, s, s,N,E)
12: exi when n=n
13: (n, s, s)=(n, s, s)
14: end
15: w=1/e o (n, s)
16: ns =ns ∪{(n,w)}
17: end
18: m=m∪{( ,ns)}
19: end
20: e u n m
Fig. 1. Main p ocedu e o Ka z
ei he nume ic (e.g., age, sala y, o opinion pola i y abou a opic) o ca ego ical
(e.g., na ionali y, gende , o dislikes). Eis a collec ion o ec o s o he o m
(m1,m2,k,w), whe e m1and m2a e he iden ifie s o wo membe s o he social
ne wo k, kdeno es a kind o ela ionship be ween hem, and wis he weigh o
ha ela ionship. The ela ionships include any kind o in e ac ion be ween any
wo membe s o a social ne wo k (e.g., eplies o pos s, pos o wa ds, iendship
eques s, message exchanges, and so on). Thus, he weigh o edge (m1,m2,k,w)
is compu ed as he numbe o ac ual in e ac ions o ype k ha ha e occu ed
be ween membe s m1and m2.
The esul o he main p ocedu e is compu ed in a iable m, which is a map
ha associa es e e y ea u e in he social ne wo k wi h a collec ion o uples
o he o m (n,w), whe e nis a neu al ne wo k, which ac s as a eg esso o a
classifie o he co esponding ea u e, and wis i s weigh , which is he in e se o
he e o a e; ha is, he smalle he e o a e, he mo e impo an he neu al
ne wo k and he la ge he e o a e, he less impo an he neu al ne wo k. Ka z
e u ns β ules o e e y ea u e, whe e βis a use -p o ided pa ame e . To label a
new membe ega ding a gi en ea u e, he neu al ne wo ks a e applied one a e
he o he . In he case o nume ic ea u es, he alues p edic ed by each ule a e
weigh ed acco ding o hei no malised e o a e and hen a e aged; in he case
o ca ego ic ea u es, he esul s a e weigh ed acco ding o he no malised e o
a e and he mos o ed one is e u ned.
The main p ocedu e basically i e a es o e he se o ea u es in he social
ne wo k; in each i e a ion, i epea s he ollowing p ocedu e β imes: i fi s
selec s he subse o nodes ha ha e a alue o he ea u e being analysed and
hen spli s i in o a aining se and a alida ion se . The size o he aining se
is con olled by means o γ, which is a use -p o ided pa ame e ; he emaining
1: expandNeu alNe wo k(n, s, s,N,E)
2: i n=null hen
3: n= lea n ne wo k om s
4: else
5: c= expand he neighbou hood o s and s using (N,E)
6: o each (u, )in cdo
7: n= lea n a ne wo k om u
8: s=u
9: s=
10: i e o (n, s)<e o (n, s) hen
11: (n, s, s)=(n, s, s)
12: end
13: end
14: end
15: e u n (n, s, s)
Fig. 2. P ocedu e o expand a neu al ne wo k
ŵϭ Ϯϯ ŵĂůĞ ƐƚƵĚĞŶƚ ƉŚLJƐŝĐƐ
ŵϮ ϮϮ ŵĂůĞ ƐƚƵĚĞŶƚ ĂƌƚƐ
ŵϯ Ϯϯ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ƉŚLJƐŝĐƐ
ŵϰ Ϯϰ ĨĞŵĂůĞ ƐƚĂĨĨ ƉŚLJƐŝĐƐ
ŵϭ Ϯϯ ŵĂůĞ ƐƚƵĚĞŶƚ ƉŚLJƐŝĐƐ ŵϱ Ϯϰ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ĂƌƚƐ ϭϮ
ŵϭ Ϯϯ ŵĂůĞ ƐƚƵĚĞŶƚ ƉŚLJƐŝĐƐ ŵϲ ϯϮ ŵĂůĞ ƐƚƵĚĞŶƚ ƉŚLJƐŝĐƐ ϭϴ
ŵϮ ϮϮ ŵĂůĞ ƐƚƵĚĞŶƚ ĂƌƚƐ ŵϳ ϮϮ ŵĂůĞ ƐƚƵĚĞŶƚ ĂƌƚƐ Ϯ
ŵϮ ϮϮ ŵĂůĞ ƐƚƵĚĞŶƚ ĂƌƚƐ ŵϴ Ϯϰ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ĂƌƚƐ ϭ
ŵϮ ϮϮ ŵĂůĞ ƐƚƵĚĞŶƚ ĂƌƚƐ ŵϵ ϯϮ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ƉŚLJƐŝĐƐ Ϯϯ
ŵϯ Ϯϯ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ƉŚLJƐŝĐƐ ŵϰ Ϯϰ ĨĞŵĂůĞ ƐƚĂĨĨ ƉŚLJƐŝĐƐ ϭϴ
ŵϯ Ϯϯ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ƉŚLJƐŝĐƐ ŵϵ ϯϮ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ƉŚLJƐŝĐƐ ϵϬ
ŵϯ Ϯϯ ĨĞŵĂůĞ ůĞĐƚƵƌĞƌ ƉŚLJƐŝĐƐ ŵϭ Ϯϯ ŵĂůĞ ƐƚƵĚĞŶƚ ƉŚLJƐŝĐƐ ϭϮ
ŵϰ Ϯϰ ĨĞŵĂůĞ ƐƚĂĨĨ ƉŚLJƐŝĐƐ ŶƵůů ŶƵůů ŶƵůů ŶƵůů ŶƵůů ŶƵůů
ǁĞŝŐŚƚ;yϭͿŐĞŶĚĞƌ;yϭͿ ŐƌŽƵƉ;yϭͿ ƐĐŚŽŽů;yϭͿ
yϬс
ŵĞŵďĞƌ
yϬс
ŵĞŵďĞƌ
ĂŐĞ;yϬͿ ŐĞŶĚĞƌ;yϬͿ ŐƌŽƵƉ;yϬͿ ƐĐŚŽŽů;yϬͿ
ĂŐĞ;yϬͿ ƐĐŚŽŽů;yϬͿŐĞŶĚĞƌ;yϬͿ ŐƌŽƵƉ;yϬͿ yϭсƐĞŶĚƐͲ
ŵĞƐƐĂŐĞ;yϬͿ
ĂŐĞ;yϭͿ
Fig. 3. Exce p o an expanded aining se
nodes a e used o alida ion pu poses. I hen ini ialises a neu al ne wo k n
o a null ne wo k ha does no hing, and hen epea edly expands i un il no
u he expansion is possible. Expanding a neu al ne wo k consis s o ex ending
i o some neighbou s as long as his helps o educe he e o a e. We p o ide
addi ional de ails in he ollowing subsec ion.
Expanding Neu al Ne wo ks: Figu e 2shows he p ocedu e o expand a
neu al ne wo k. I wo ks on a neu al ne wo k n, a aining se s, a alida ion
se s, and a social ne wo k (N,E); i e u ns a new neu al ne wo k, he aining
se om which i was lea n , and he alida ion se on which i was alida ed.
The p ocedu e fi s checks i he inpu neu al ne wo k is null, in which case i
simply lea ns a neu al ne wo k om he aining se and e u ns i . O he wise,
i fi s expands he neighbou hood o he aining and he alida ion se s using
he in o ma ion p o ided by he social ne wo k. Expanding he neighbou hood
o a da ase means ha i s ec o s a e expanded wi h addi ional componen s
ha ep esen he ea u es o a kind o neighbou . Fo ins ance, Figu e 3shows
Table 1. Expe imen al esul s
ŐĞ EĂƚŝŽŶ
Nulli .
ETETETETET
Nulli .
ETETETETET
5.00% 5.55 7.01 9.70 12.39 8.85 6.75 5.00% 19.12 5.56 19.77 5.60 3.44 11.77 7.94 13.51 5.40 4.10
10.00% 9.39 8.22 14.30 12.74 12.55 6.69 10.00% 26.85 6.79 21.28 6.95 6.81 12.27 8.33 15.06 9.09 3.76
15.00% 16.45 9.18 10.70 15.60 10.88 4.36 15.00% 27.55 7.33 22.01 7.83 9.00 12.64 14.96 18.77 13.55 3.74
20.00% 6.49 10.61 11.16 16.30 11.75 4.71 20.00% 22.76 9.04 22.75 6.96 10.33 15.55 8.94 20.13 17.36 3.23
25.00% 20.43 11.37 20.00 14.55 18.83 4.70 25.00% 34.96 7.25 23.56 7.39 19.24 16.57 8.59 20.62 25.13 4.40
30.00% 24.08 10.94 31.86 16.23 15.62 4.68 30.00% 48.85 7.99 24.25 7.27 7.88 12.54 17.79 22.72 12.34 5.07
35.00% 35.55 13.56 21.17 15.63 10.59 5.05 35.00% 23.34 8.59 25.05 8.13 22.35 13.76 35.32 26.06 20.84 4.76
40.00% 23.39 12.46 15.31 17.88 17.86 4.66 40.00% 22.85 7.19 35.82 8.14 24.67 16.75 27.80 23.97 20.46 2.44
45.00% 50.13 12.92 36.26 21.65 13.18 6.38 45.00% 54.20 7.53 26.50 8.65 19.85 16.79 19.85 23.39 21.07 3.33
50.00% 40.23 15.81 17.44 20.75 10.48 5.59 50.00% 49.13 7.86 44.71 10.52 52.53 17.77 17.79 27.51 44.04 3.57
Mean 23.17 11.21 18.79 16.37 13.06 5.36 Mean 29.96 7.51 25.57 7.74 17.61 14.64 16.73 21.17 17.93 3.84
'
ĞŶĚĞ
ƌ
^ĐŚŽŽů
Nulli .
ETETETETET
Nulli .
ETETETETET
5.00% 5.40 6.35 10.25 6.54 5.62 5.40 9.73 4.65 8.01 4.08 5.00% 21.75 5.50 16.52 7.54 4.06 5.81 9.85 4.55 6.46 5.41
10.00% 5.72 9.33 13.54 6.90 14.50 5.61 16.40 4.86 12.73 4.14 10.00% 23.86 6.59 18.81 9.03 11.93 5.69 15.55 4.98 5.48 4.04
15.00% 6.52 10.71 22.70 6.04 12.84 6.42 11.45 5.97 11.17 4.44 15.00% 27.77 7.92 19.95 9.47 10.85 5.89 10.73 5.32 8.01 5.27
20.00% 14.85 12.21 26.83 6.96 24.15 6.08 17.67 6.67 15.25 4.76 20.00% 30.22 8.19 21.11 7.34 11.14 6.30 8.69 5.33 7.12 5.24
25.00% 15.29 13.68 31.03 7.16 5.59 7.49 23.57 5.72 17.17 3.22 25.00% 35.98 9.64 22.26 7.92 8.18 6.55 22.06 5.59 21.34 5.30
30.00% 7.71 15.19 20.18 8.41 35.44 5.93 32.64 6.49 10.59 4.14 30.00% 24.09 9.59 23.47 9.40 25.70 6.88 19.83 5.69 28.90 2.09
35.00% 25.13 16.59 21.76 9.59 6.35 4.73 23.68 6.95 16.38 3.55 35.00% 40.91 11.79 24.63 8.97 14.68 7.94 35.23 6.84 28.30 2.07
40.00% 18.76 18.08 23.57 7.91 20.12 3.62 24.36 7.03 14.36 4.06 40.00% 53.60 9.75 25.74 9.76 21.84 9.44 24.57 7.36 25.61 1.57
45.00% 13.36 19.68 29.06 7.91 27.63 3.21 14.86 5.82 13.36 4.09 45.00% 30.26 7.60 30.94 9.48 30.13 9.62 31.53 8.74 29.31 2.28
50.00% 28.18 20.87 26.63 9.50 27.29 2.88 42.70 6.43 15.29 3.83 50.00% 32.96 7.78 28.06 11.67 21.03 12.00 19.98 10.05 32.13 2.29
Mean 12.09 14.27 22.16 7.69 17.95 5.14 21.71 6.06 13.43 4.03 Mean 32.14 8.43 22.75 9.06 15.95 7.61 19.80 6.44 17.27 3.56
'ƌŽƵƉ >ŝŬŝƐ
Nulli .
ETETETETET
Nulli .
ETETETETET
5.00% 14.36 5.01 18.56 6.07 8.53 9.52 9.17 4.13 8.33 5.40 5.00% 19.49 5.49 13.50 7.91 10.34 5.19 15.19 10.84 10.75 6.26
10.00% 23.36 9.27 23.09 9.35 14.53 9.32 13.18 9.30 11.34 5.06 10.00% 24.50 9.33 20.25 9.28 18.35 9.26 20.20 9.31 13.75 5.62
15.00% 27.86 10.79 25.28 10.75 17.54 10.82 15.18 10.74 12.83 3.85 15.00% 27.00 10.78 21.12 10.80 17.84 10.83 22.70 10.83 15.25 9.25
20.00% 32.36 12.26 27.57 12.30 20.53 12.23 17.17 12.27 14.33 4.26 20.00% 29.49 12.18 21.95 12.17 20.34 12.25 25.20 12.20 16.75 7.68
25.00% 36.86 13.84 29.75 13.64 23.53 13.76 19.17 13.82 15.84 4.19 25.00% 31.99 13.62 22.85 13.69 22.85 13.74 27.70 13.75 18.25 6.99
30.00% 41.36 15.22 32.08 15.31 26.53 15.22 21.18 15.13 17.33 5.77 30.00% 34.49 15.10 23.63 15.12 25.35 15.06 30.19 15.14 19.75 9.72
35.00% 45.86 16.70 34.43 16.54 29.53 16.56 23.17 16.60 18.84 5.14 35.00% 37.00 16.73 24.50 16.64 27.84 16.63 32.70 16.51 21.25 9.48
40.00% 50.36 18.02 36.70 18.25 32.53 18.33 25.18 18.06 20.33 2.97 40.00% 39.50 17.95 25.42 18.02 30.35 18.12 35.20 18.09 22.76 16.02
45.00% 54.86 19.72 38.72 19.74 35.53 19.76 27.18 19.64 21.83 3.82 45.00% 42.00 19.64 32.22 19.50 32.85 19.51 37.69 19.46 24.26 19.39
50.00% 59.36 21.27 45.05 21.01 38.53 21.32 29.17 21.35 23.33 3.41 50.00% 44.50 21.34 37.04 20.85 35.35 20.91 40.19 21.19 25.75 13.73
Mean 38.66 14.21 30.72 14.29 24.73 14.68 19.98 14.10 16.43 4.39 Mean 33.00 14.22 22.65 14.40 24.15 14.15 28.70 14.73 18.85 10.41
Ka zLGNJ MP B
NJ LG MP B Ka z
NJ LG MP B Ka z
NJ LG MP B Ka z
NJ LG MP B
NJ LG MP B
Ka z
Ka z
an exce p o an ini ial aining se on he le ; on he igh , ha aining se has
been expanded wi h he ea u es o he neighbou s ega ding he ‘sends-message’
ela ionship; no e ha he weigh o he ela ionship is added as an addi ional
ea u e o he ec o .
Then, he p ocedu e i e a es h ough he se o expansions o he aining se
and lea ns a new neu al ne wo k om each one. I e u ns he expanded neu al
ne wo k ha achie es he smalles e o a e oge he wi h he aining se om
which i was lea n and alida ion se on which he e o a e was compu ed.
3 Expe imen al Resul s
We conduc ed a se ies o expe imen s o analyse how Ka z pe o ms in p ac ice.
The expe imen s we e ca ied ou using a Ja a 1.7 implemen a ion ha was un

ϬϬϬ
ϱϬϬ
ϭϬϬϬ
ϭϱϬϬ
ϮϬϬϬ
ϮϱϬϬ
ϯϬϬϬ
ϯϱϬϬ
ϰϬϬϬ
ϰϱϬϬ
ϱϬϬй ϭϬϬϬй ϭϱϬϬй ϮϬϬϬй ϮϱϬϬй ϯϬϬϬй ϯϱϬϬй ϰϬϬϬй ϰϱϬϬй ϱϬϬϬй
ƌƌŽƌƌĂƚĞ
WĞƌĐĞŶƚĂŐĞŽĨŶƵůůŝĨŝĐĂƚŝŽŶ
E: >' DW  <Ăƚnj
ϬϬϬ
ϮϬϬ
ϰϬϬ
ϲϬϬ
ϴϬϬ
ϭϬϬϬ
ϭϮϬϬ
ϭϰϬϬ
ϭϲϬϬ
ϭϴϬϬ
ϮϬϬϬ
ϱϬϬй ϭϬϬϬй ϭϱϬϬй ϮϬϬϬй ϮϱϬϬй ϯϬϬϬй ϯϱϬϬй ϰϬϬϬй ϰϱϬϬй ϱϬϬϬй
WƌŽĐĞƐƐŝŶŐƚŝŵĞ
WĞƌĐĞŶƚĂŐĞŽĨŶƵůůŝĨŝĐĂƚŝŽŶ
E: >' DW  <Ăƚnj
Fig. 4. G aphic summa y
on a ou - h eaded In el Co e i7 compu e ha an a 2.93 GHz, had 16 GiB
o RAM, Windows 7 P o 64-bi , O acle’s Ja a De elopmen Ki 1.7.902, and
Weka 3.6.8.
We implemen ed he gene al amewo k by Ne ille and Jensen [17](NJ)
and hen he specific p oposals by Lu and Ge oo [12] (LG), Macskassy and
P o os [13] (MP), and Bhaga e al. [3] (B). Rega ding he p e ious p oposals,
we conside ed ha a labelling was es able when no mo e han 5% o he ea u es
changed in an i e a ion o he me hod. Rega ding Ka z, we expe imen ed wi h
se e al combina ions o pa ame e s and kinds o neu al ne wo ks. We ound ou
ha he ollowing alues o he pa ame e s wo k qui e well: β= 10, ha is, 10
neu al ne wo ks lea n o each ea u e, and γ=0.25, ha is, 25% o he nodes
a ailable in he social ne wo k a e used o aining pu poses and he emaining
o alida ion pu poses. Rega ding he lea ning echnique, we ound ou ha
RBFN ne wo ks [4] a e he bes pe o ming in his con ex .
The expe imen s we e pe o med on a da ase ha consis ed in a dump o
ou uni e si y social ne wo k. This ne wo k has 56,431 membe s, each o which
is cha ac e ised by a p ofile ha includes he ollowing ea u es: age (a na u-
al alue), gende (male, emale), g oup (s uden , lec u e , s aff), na ionali y
(Spanish, F ench, I alian, and so on), school (Compu e -Science, Ma hema -
ics, Physics, Philology, and so on), likes, and dislikes; o he ea u es like name,
add ess, na ional id o passpo we e disca ded o keep he da a anonymous;
nei he was i e y in e es ing o a emp o p edic hem. The likes and dislikes
a e se s o key wo ds ha a e selec ed by he membe s om a lis ha is com-
pu ed au oma ically om he messages pos by he membe s o he ne wo k; o
deal wi h hem in ou expe imen s, we selec ed he op 50 key wo ds and c ea ed
bina y ea u es o he o m likes Xo dislikes Y, whe e Xo Y ep esen s key wo ds.
The ela ionships be ween he membe s o he ne wo k a e he ollowing: pos s-
o-wall, eplies- o-pos , o wa ds-pos , sends-message, ollows-membe , eques s-
iendship. This da ase was pa icula ly use ul because almos e e y p ofile has
accu a e ea u es ha a e se au oma ically using he s uden s’ egis a ion da a
o he lec u e s’ and s aff’s wo k con ac s, and he likes and dislikes a e also
selec ed om se s o p e-compu ed key wo ds. Tha is, we had qui e a la ge
co ec ly labelled da ase on which could conduc qui e a p ecise alida ion.
To e alua e ou p oposal and compa e i o o he s, we c ea ed se e al da ase s
om he p e ious one. They we e e sions o he o iginal da ase in which we
nullified he ea u es o 5% up o 50% nodes ha we e chosen andomly. This
helped us o e alua e how ou p oposal wo ks and compa e i o o he s in e ms
o e o a e (E) and p ocessing ime (T). The e o a e was compu ed as he
pe cen age o w ong p edic ions; in he case o nume ic ea u es a ±10% ole ance
h eshold was es ablished o conside a p edic ion w ong. The p ocessing ime
was measu ed in CPU plus IO hou s, since hese imings a e a mo e eliable
and s able han use imes.
Table 1shows ou esul s and Figu e 4summa ises hem using a couple o
cha s. (The columns ha co espond o p oposals NJ and LG ega ding ea-
u e ‘age’ a e emp y because hese me hods canno be applied o nume ic ea-
u es.) Rega ding effec i eness, he fi s conclusion is ha he e o a e inc eases
s eadily as he pe cen age o nullifica ion inc eases, bu Ka z keeps he smalles
global mean in he majo i y o cases, whe e global mean e e s o he compu ed
mean o each ea u e o a gi en nullifica ion pe cen age, c . he uppe pa
o Figu e 4. To compa e he esul s mo e p ecisely, we ha e compu ed he en-
dency lines o each p oposal acco ding o he pe cen age o nullifica ion, which
is deno ed as N:
P oposal E o a e endency R2
NJ 2.78N+14.72 0.99
LG 1.88N+15.17 0.93
MP 2.99N+4.16 0.95
B2.15N+9.17 0.88
Ka z 1.69N+7.37 0.92
No e ha he R2coefficien is e y good in e e y case, which means ha
he e is a clea linea endency in he esul s. The smalles slope co esponds
o Ka z, which means ha i is he p oposal whose e o a e inc eases a he
lowes pace as he pe cen age o nullifica ion inc eases; i is ollowed by LG, bu
no e ha he e o a e o his p oposal is oughly double as Ka z’s.
Rega ding efficiency, he fi s conclusion is ha Ka z seems o ha e a
beha iou ha is e y s able, whe eas he o he p oposals seem o equi e mo e
p ocessing ime as he pe cen age o nullifica ion inc eases. To confi m his idea,
we ha e also compu ed he endency lines o each p oposal, namely:
P oposal P ocessing ime endency R2
NJ 1.04N+5.62 0.98
LG 0.66N+6.08 0.98
MP 0.78N+6.93 0.97
B1.00N+7.63 0.98
Ka z 0.08N+4.81 0.95
No e ha he R2coefficien is again e y good in e e y case. The smalles
slope co esponds again o Ka z, which means ha i is he p oposal whose p o-
cessing ime inc eases a a lowe pace as he pe cen age o nullifica ion inc eases.
No e ha i is e y close o 0.00, which means ha he p ocessing ime emains
almos cons an ; he eason is ha he size o he aining se s dec ease as he
pe cen age o nullifica ion inc eases, which makes lea ning neu al ne wo ks eas-
ie ; un o una ely, as he pe cen age o nullifica ion inc eases, he numbe o
neighbou s ha mus be explo ed o keep as a low e o a e as possible inc eases.
Ka z is ollowed by he o he p oposals, which equi e conside ably mo e p o-
cessing ime since hey ha e o i e a e un il he labelling is s able enough, which
is mo e and mo e difficul as he pe cen age o nullifica ion inc eases.
4 Rela ed Wo k
The e a e wo mains eam app oaches o he membe labelling p oblem [9],
namely: local and global me hods. They bo h wo k on a g aph-based ep e-
sen a ion o he social ne wo k being analysed, whe e he nodes s o e membe
ea u es and he edges keep ack o hei in e ac ions, bu diffe in ha he
o me ocus on lea ning local p edic o s om e e y membe and his o he local
neighbou hood, whe eas he la e analyse he social ne wo k as a whole.
Below, we epo on bo h app oaches and discuss on how ou p oposal
imp o es on hem om a concep ual poin o iew.
Local Me hods: These me hods can be u he classified in o ins an ia ions o
he I e a i e Classifica ion Algo i hm by Ne ille and Jensen [17] o ins an ia ions
o he Gibbs Sampling Algo i hm by Geman and Geman [8].
The me hods ha a e based on he I e a i e Classifica ion Algo i hm [17]
ans o m a social ne wo k in o a da ase o ec o s, each o which p o ides he
ea u es o a membe ’s p ofile plus some agg ega ed ea u es ha co espond o
he membe s in his o he neighbou hood. They analyse each ea u e in isola ion
as ollows: hey fi s lea n a local p edic o om he membe s whose p ofiles
p o ide a non-null alue o ha ea u e. (In o mally, his is commonly e e ed
o as “ he membe is labelled”.) The p edic o is ei he a eg esso o a classifie
depending on whe he he ea u e being analysed is nume ic o ca ego ical. I is
hen used o compu e he label o he unlabelled membe s, as long as hey ha e
a leas a labelled neighbou . No e ha labelling a membe will likely change
he alues o he agg ega ed ea u es in he neighbou hood, so he labelling p o-
cess needs o be epea ed i e a i ely un il he labels do no change d ama ically
o do no change a all. The p e ious idea has been ins an ia ed many imes
in he li e a u e, he diffe ence being he kind o p edic o used: Ne ille and
Jensen [17] used Nai e-Bayes p edic o s, Lu and Ge oo [12] used logis ic eg es-
sion, Macskassy and P o os [13] used a o ing app oach, and Bhaga e al. [3]
and McDowell e al. [16]usedk-nea es neighbou s. Recen ly, Ca al epe e al. [6]
ha e used diffe en ypes o p edic o s o membe ea u es and neighbou hood
ea u es, which a e hen combined o p oduce an ensamble p edic o .
The me hods ha a e based on he Gibbs Sampling Algo i hm [8]wo kin
ou phases, namely: boo s apping, bu n-in, collec ing, and labelling. In he
boo s apping phase, hey lea n a p edic o in a way ha is e y simila o
he me hods ha a e based on he I e a i e Classifica ion Algo i hm, and hen
use i o label he unlabelled membe s. Then, he bu n-in phase is epea ed a
numbe o imes; in each epe i ion, he membe s ha we e ini ially unlabelled
a e andomly o de ed and hen new labels a e compu ed using a p edic o , which
can be he same ha was used in he boo s apping phase o a new one [14]. In
he sample collec ion phase, he p ocess is epea ed a p e-defined numbe o imes
and he coun o labels assigned o each membe is compu ed. Finally, in he
labelling phase, he membe s ha we e ini ially unlabelled a e assigned he mos
likely label acco ding o he coun s ha we e compu ed in he p e ious phase.
Bo h McDowell e al. [15] and Macskassy and P o os [14] ha e ins an ia ed his
idea; he o me used Nai e-Bayes and k-nea es neighbou s and he la e used
diffe en combina ions o p edic o s.
Global Me hods: The mos common me hods in his ca ego y a e based on
andom walks and op imisa ion.
A andom walk on a g aph is a e y special case o a Ma ko chain. The co e
idea was in oduced by Zhu e al. [23]: hey ely on a ansi ion ma ix P ha
encodes he p obabili y ha a andom walk p oceeds be ween any wo membe s
o a social ne wo k using hei ela ionships. Gi en an unlabelled membe , he
me hod assigns i he mos common label ou o he membe s ha can be eached
om i using andom walks. Tha is, i equi es o compu e an app oxima ion