Full text
CLIP-based Few-Sho Mul i-Label Classi ica ion
Me hods: A Compa a i e S udy
Ya˘
gmu C¸ i˘
gdem Ak as¸
Vicom ech Founda ion
Basque Resea ch and Technology Alliance (BRTA)
Mikele egi 57, 20009 Donos ia-San Sebas i´
an (Spain)
[email p o ec ed]g
Jo ge Ga c´
ıa Cas a˜
no
Vicom ech Founda ion
Basque Resea ch and Technology Alliance (BRTA)
Mikele egi 57, 20009 Donos ia-San Sebas i´
an (Spain)
[email p o ec ed]g
Abs ac —Ca ego izing da kweb image con en is c i ical o
iden i ying and a e ing po en ial h ea s. Howe e , his emains
a challenge due o he na u e o he da a, which includes
mul iple co-exis ing domains and in a-class a ia ions. While
many me hods ha e been p oposed o classi y his image con en ,
mul i-label mul i-class classi ica ion emains unde explo ed. The
complexi y o da kweb image y, combined wi h he need o
e icien classi ica ion sys ems, demands inno a i e app oaches
ha can handle bo h he echnical challenges and he sensi i e
na u e o he con en . In his pape , we p esen a compa a i e
s udy o ew-sho mul i-label classi ica ion me hods using he
mul imodal model CLIP. Ou esea ch add esses he g owing
need o obus classi ica ion sys ems ha can e ec i ely ca ego-
ize di e se and complex image con en while main aining high
accu acy and compu a ional e iciency. We pa icula ly ocus on
he challenges o handling mul iple labels simul aneously and he
scalabili y o hese sys ems in eal-wo ld applica ions. We analyze
and compa e ou di e en app oaches: CLIP+Label Empowe
Adap e , CLIP Sigmoid, SIGLIP, and CLIP+ML-Decode . Ou
s udy e alua es hese me hods based on hei p ecision, ecall,
and abili y o handle inc easing class numbe s e icien ly. Finally,
ou esea ch con ibu es o he ield by p o iding de ailed insigh s
in o he s eng hs and limi a ions o each me hod.
Index Te ms—Mul i-label image classi ica ion, Mul i-class im-
age classi ica ion, CLIP, SIGLIP, ML-Decode
I. INTRODUCTION
The da k web, a specialized subse o he as global
in e ne , is accessible only h ough dedica ed web b owse s.
I is widely ecognized as a hub o illici ac i i ies due o
i s co e cha ac e is ics: anonymi y, which p o ides use s wi h
a le el o p i acy no ypically a ailable on he su ace web,
and un aceabili y, which makes i ex emely challenging o
ack he o igins and des ina ions o da a ans e s. As a esul ,
he da k web se es as a c i ical sou ce o h ea in elligence,
o e ing insigh s in o cybe -a acks, s olen asse s, illegal ade,
a ms a icking, con iden ial da a leaks, child exploi a ion, and
o he c iminal ac i i ies.
Au oma ically analyzing he isual con en o he da k
web —such as images and ideos- is essen ial o e icien
image ca ego iza ion, enabling he de ec ion and p e en ion
o po en ial h ea s. Howe e , da k web image classi ica ion
emains an open challenge due o he complex na u e o i s
con en .
Fig. 1: Mul i-label image examples om Da k web da ase
Many images su e om low esolu ion, blu , and small
objec sizes, o en con aining objec s ha blend in o he back-
g ound due o colo simila i ies. Addi ionally, he coexis ence
o mul iple o e lapping domains and in a-class a ia ions —
whe e ins ances wi hin he same ca ego y exhibi signi ican
isual di e ences — u he complica e classi ica ion asks, as
illus a ed in Fig. 1.
This complexi y highligh s he impo ance o mul i-label
classi ica ion, whe e an image can be assigned mul iple ele-
an labels a he han a single ca ego y.
A ange o mul i-label classi ica ion echniques ha e been
explo ed o ackle his p oblem. Some app oaches decom-
pose mul i-label classi ica ion in o sepa a e single-label p ob-
lems[1], [2], while o he s e ine loss unc ions o ac i a ion
unc ions o enhance classi ica ion accu acy[3], [4]. Addi ion-
ally, g aph-based ne wo ks u ilize lea ned label embeddings o
loca e key disc imina i e ea u es [5], [6], while ans o me -
based a chi ec u es ha e been adap ed o mul i-label classi i-
ca ion asks [7].
While hese me hods show p omising esul s, hey ely on
la ge, well-balanced da ase s—which a e di icul o cu a e in
eal-wo ld applica ions. Expanding a da ase wi h a balanced
class dis ibu ion is an exponen ially complex p ocess [8],
making la ge-scale mul i-label lea ning in easible in many
scena ios.
To add ess his, ecen ad ances in ision-language models
like CLIP [9] o e new possibili ies. By aligning isual and
ex ual embeddings, CLIP enables ze o-sho and ew-sho
lea ning in open- ocabula y se ings. Howe e , mos CLIP-979-8-3315-0993-4/25/31.002025IEEE
based me hods a e ailo ed o single-label classi ica ion.
In his s udy, we add ess he p oblem o ew-sho mul i-
label image classi ica ion, ocusing on he da k web as a
eal-wo ld, high-s akes applica ion domain. Ou hypo hesis is
ha adap ing s a e-o - he-a ew-sho mul i-label classi ica ion
echniques o CLIP can imp o e gene aliza ion and pe o -
mance in low- esou ce, high- a iabili y en i onmen s.
Ou con ibu ions can be summa ized as ollows:
•We analyze ou s a e-o - he-a ew-sho mul i-label
classi ica ion echniques and p o ide a po able pipeline
applicable o any eal-wo ld da ase equi ing ew-sho
mul i-label classi ica ion.
•We e alua e hese app oaches by implemen ing CLIP-
based me hodologies o da k web mul i-label image
classi ica ion.
•We compa e ou p e ious app oach wi h newe CLIP-
based mul i-label me hodologies, p esen ing a anspa en
compa a i e s udy which con ibu es o he ield by p o-
iding de ailed insigh s in o he s eng hs and limi a ions
o each me hod.
II. RELATED WORK
A a ie y o me hods ha e been p oposed o sol e he mul i-
label classi ica ion p oblem, which can be b oadly ca ego ized
in o (i) p oblem ans o ma ion me hods and (ii) algo i hm
adap a ion echniques.
Ea ly esea ch ocused on decomposing mul i-label classi i-
ca ion in o independen bina y classi ica ion asks[1]. How-
e e , his app oach ails o cap u e in e -label co ela ions
and equi es aining sepa a e classi ie s o each ca e-
go y—in oducing a signi ican compu a ional bu den. To ad-
d ess his, some s udies p opose classi ie chaining[10], whe e
each classi ie ’s ou pu se es as an inpu ea u e o sub-
sequen classi ie s. While his s a egy imp o es in e -label
dependency modeling, i s ill su e s om high compu a ional
cos s due o he g owing numbe o classi ie s equi ed.
Ano he ans o ma ion-based app oach, known as Label
Powe se (LP) [2], con e s mul i-label p oblems in o mul i-
class p oblems by ea ing each unique label combina ion as
a dis inc class. Howe e , LP s uggles wi h he exponen ial
g ow h o label combina ions, making i imp ac ical o la ge
da ase s.
Se e al me hods aim o imp o e mul i-label classi ica ion
by op imizing he loss unc ion[3] o ac i a ion unc ion[4] o
mi iga e class imbalance issues.
In compu e ision, mul iple app oaches ha e been de el-
oped o mul i-label image classi ica ion. Hypo heses-CNN-
Pooling (HCP) [11] gene a es a la ge numbe o p oposals
h ough objec de ec ion echniques, ea ing each p oposal as
a single-label classi ica ion p oblem.
Beyond CNN-based me hods, G aph Con olu ional Ne -
wo ks (GCNs) ha e demons a ed high e icacy ac oss a i-
ous ision asks, including mul i-label classi ica ion [5], [6],
[12]. GCN-based me hods build classi ie s by modeling label
ela ionships wi hin g aph ne wo ks.
O he app oaches employ weakly supe ised lea ning o
mul i-label classi ica ion [13], le e aging knowledge dis illa-
ion echniques om objec de ec ion models.
Mo e ecen ly, ans o me s[14]—o iginally de eloped o
modeling long- ange dependencies in na u al language p o-
cessing (NLP)[15]–[17]—ha e demons a ed s ong pe o -
mance ac oss compu e ision asks, including image classi i-
ca ion[7], [9] and objec de ec ion[18].
Se e al ans o me -based mul i-label classi ica ion ech-
niques ha e eme ged:
Mul i-label T ans o me (MLT) [19], which models pixel-
wise a en ion o enhance ea u e ex ac ion. T ans o me -
based que y lea ning [20], which uses decode que ies o
p edic label exis ence. G aph-enhanced T ans o me s [21],
which combine GCNs wi h a en ion mechanisms. Me ic
Lea ning T ans o me s [22], which in eg a e me ic lea ning
o assess label simila i ies. T ans o me -CNN Hyb ids [23],
which gene a e a en ion maps pe label, cap u ing in a-class
dependencies h ough con olu ional laye s.
Despi e hese ad ancemen s, exis ing mul i-label me hods
s ill equi e la ge da ase s o achie e high accu acy. Howe e ,
da a sca ci y is a p e alen challenge in eal-wo ld applica-
ions, whe e collec ing and labeling ex ensi e da ase s is bo h
esou ce-in ensi e and ime-consuming.
La ge-scale open- ocabula y models ained on as
da ase s, such as CLIP [9], o e a p omising al e na i e by
le e aging ze o-sho single-label classi ica ion and ew-sho
mul i-label classi ica ion. By adap ing hese models, we aim
o explo e hei po en ial o enhancing mul i-label classi i-
ca ion in da k web image y, whe e adi ional me hods ace
signi ican cons ain s.
III. CLIP-BASED MULTI-LABEL CLASSIFICATION
METHODOLOGIES
The e a e ou ecen mul i-label classi ica ion me hods sui -
able o CLIP mul i-modal model adap a ion: (i) CLIP+Label
Empowe (ou p e ious model), (ii) CLIP Sigmoid (iii) SIGLIP
(i ) CLIP + ML-Decode
A. CLIP+Label Empowe
Label Empowe [2] is one o he app oaches ha con e he
mul i-label classi ica ion p oblem in o a single-label classi i-
ca ion, by mul i-labels o bina y ec o s by ho -encoding hem
and assigning a unique label o each di e en bina y ec o .
As demons a ed in [24], i has he highes pe o mance among
he o he app oaches o sol ing he mul i-label classi ica ion
ask by ans o ming he p oblem in o a single-label one. Fig.
2 shows he a chi ec u e o ou p e iously p oposed me hod
CLIP+Label Empowe Adap e .
Al hough ou p e iously p oposed model has he highes
ecall, as shown in Table I, among e en hese ou s a e-
o - he-a CLIP-adap ed mul i-label classi ica ion app oaches,
no only be ween he adi ional ones, i p ese es a isk o
accu acy d op wi hin a la ge numbe o classes due o i s
exponen ially g owing ho -encoded labels. This possible isk
is especially impo an o such a da ase like da kweb, which
Fig. 2: CLIP based Mul i-Label Me hodologies
is expec ed o g ow as in he sense o numbe o classes,
since new ype o images and i les a e expec ed o occu .
•Wo s Case: O(N!) (All class labels exis in all possible
pe mu a ions in he da ase )
•Bes Case: O(N)(The da ase con ains only single
labels)
B. CLIP Sigmoid
CLIP has he capabili y o classi y an ex emely huge num-
be o di e en classes, hanks o i s open- ocabula y na u e.
Ye we migh need o ine- une his s a e-o - he-a model o
ou eal-wo ld da ase s like da kweb da ase . In a single-label
pipeline, he p ope way o ine- uning such a model is o
add a linea laye ollowing he image and ex embeddings,
ha ing ou pu nodes as much as class numbe s in he da ase .
In a single-label, mul i-class classi ica ion scena io, so max
is commonly used o dis ibu e p obabili ies among di e en
classes, ensu ing ha he sum o all class p obabili ies equals
1.
P(yi|x) = ezi
PN
j=1 ezj
(1)
whe e zi ep esen s he logi s o class iand Nis he o al
numbe o classes.
Howe e , in mul i-label classi ica ion, whe e mul iple labels
can be p esen simul aneously, so max is subop imal as i
o ces mu ual exclusi i y among classes. Ins ead, sigmoid
ac i a ion is mo e app op ia e, as i independen ly p edic s
each label’s p obabili y:
P(yi|x) = 1
1 + e−zi(2)
This allows he model o assign a p obabili y close o 1.0
o each ele an label in an image, ensu ing ha all p esen
classes a e p edic ed wi hou compe i ion om o he labels.
Thus, ine- uning CLIP wi h a linea classi ica ion head
using Sigmoid ac i a ion is ano he app oach. Al hough his
me hod does no ha e any isk o exponen ially g owing wi h
he numbe o class numbe s, ou expe imen al esul s show
i s poo accu acy in compa ison wi h ou p e iously p oposed
me hod.
C. SIGLIP
While SIGLIP [25] is a s a e-o - he-a , mul i-modal model,
ha ing a e y simila a chi ec u e o CLIP, i s main di e ence
is using Sigmoid, ins ead o So max while p e- aining. This
ac p o ides us wi h he possibili y o see i he poo pe o -
mance o CLIP + Sigmoid ine- uning me hodology occu ed
due o using a di e en ac i a ion unc ion in he ine- uning
han p e- aining.
Fig. 2 shows he e y simila a chi ec u e o SIGLIP.
D. CLIP+ML-Decode
ML-Decode [26] is a s a e-o - he-a decode implemen-
a ion aiming o o e come he exponen ially g owing inpu
size in classi ica ion pipelines buil wi h de aul ans o me
decode s (Fig. 3) due o he sel -a en ion laye . The me hod
p oposes o emo e he sel -a en ion laye and claims i s
a ec less on he classi ica ion accu acy. I is also claimed o
use ”g oup que ies”, ins ead o ixed que y embeddings pe
class as in adi ional ans o me -decode s. Being he numbe
o g oups a new hype pa ame e in his a chi ec u e e e s
o he amoun o inpu que y embeddings he decode will
ha e, independen ly om he numbe o classes in he da ase .
Being K is he numbe o g oup que ies, he decode lea ns
o c ea e K que ies e e encing N numbe o classes in a
da ase , ins ead o ha ing N ixed que ies o N classes as
in de aul ans o me decode s. CLIP+ML-Decode app oach
uses bo h classi ica ion loss coming om he inal p edic ion
logi s and an alignmen loss coming om he simila i y o ex
and image embeddings o ob ain a inal loss. The ixed ex
embeddings a e used only o his pu pose, whe eas non- ixed
g oup que ies a e lea nable embeddings and hey a e upda ed
du ing he aining.
Fig. 3: T adi ional s ML Decode s
The e o e his app oach enhances scalabili y, wi h wo key
op imiza ions: (1) Sel -a en ion emo al, which educes he
compu a ional complexi y om O(N2) o O(N), and (2)
G oup decoding, which u he op imizes in e ence by shi ing
he complexi y om O(N) o O(K), whe e KK ep esen s he
numbe o meaning ul g oups ins ead o p ocessing all classes
independen ly.
Sel -a en ion emo al:
O(N2)→O(N)
G oup decoding:
O(N)→O(K)
Fig. 2 shows he a chi ec u e o CLIP + ML-Decode .
IV. EXPERIMENTAL STUDY
A. Da k web da ase
The Da k web da ase , sou ced om CFLW’s Da k Web
Moni o 1, includes images collec ed om da k web domains
abou a ious c ime ca ego ies like inancial c ime and o gani-
za ions, d ugs and na co ics, weapons as well as hei endo s
as indi idual classes. Simila o many eal-wo ld da ase s,
he da k web da ase con ains images wi h mul iple labels,
meaning an image can belong o mo e han one class. I also
includes images wi h a single label. Fig. 1 shows some image
samples along wi h hei g ound u h labels om a ious
classes.
1h ps://c lw.com/dwm/
Fig. 4: Da a dis ibu ion o he da k web da ase o bo h ain and es se s.
The da ase con ains 46 classes om a ious ca ego ies
like D ugs, Na co ics, Weapons, and Financial O ganiza ions
which do no ha e a s ong co ela ion be ween hem bu
i includes subca ego ies ha ing a s ong ela ionship and
he possibili y o being exis ing in he same image. Being
he cu en da kweb da ase is an expe imen al one, many
mo e classes a e expec ed o join and hus, he scalabili y
pe o mance is impo an han a egula Fig. 4 shows he
da a dis ibu ion o he da k web da ase o bo h ain and
es se s.
The da kweb da ase p esen s an imbalanced da a p oblem.
Some dominan classes ha e many samples, whe eas some
o he classes su e om a lack o da a. Some classes like
d ugs and na co ics, coming along wi h any ype o d ug ype
o any indi idual endo ha has o be ca ego ized, become
a e y dominan class by collec ing only a ew samples o
i s subg oups. Fig. 1 shows 3 image samples om he da k
web da ase , 1 belonging o a inancial o ganiza ion ca ego y,
2 belonging o he d ugs and na co ics ca ego y con aining
pic u es o weed and he endo names. While endo names
may a y, each image ha ing a endo name in o ma ion also
con ains weed class, doubling he amoun o ”weed” class in
compa ison o class names e e ing o he indi idual endo
names. Which is obus example showing he eason o s ong
class imbalance p oblem in da kweb da ase .
Las ly, ano he challenge is ha some classes, especially he
endo names like ”deep shop”, o ” anda al”, ha e ex ual
in o ma ion a he han isual ea u es which b ings he need
o conside ing he combina ion o ision and language da a.
B. Implemen a ion De ails
The ine- uning s a egy a ies o each o he me hods we
examine: CLIP + Label Empowe was ine uned by eezing
he CLIP p e- ained model, and upda ing he weigh s o only
o be ween 5 o 10 epochs, using Adam op imize wi h de aul
lea ning a e 1×10−3. So max inal ac i a ion unc ion and
C oss En opy Loss.
CLIP Sigmoid model was ine uned by eezing he CLIP
p e- ained model, and upda ing he weigh s o only he ad-
di ional linea classi ica ion head, o 50 epochs un il con e -
gence, using he same op imize and lea ning a e as p e ious
app oach, wi h a Bina y C oss En opy Loss o e he logi s,
as a de aul app oach o sigmoid classi ica ion.
SIGLIP model was ine uned wi hou eezing any pa
o he a chi ec u e, since ou implemen a ion any addi ional
pa and he na u e o he a chi ec u e is al eady compa ible
o mul i-label classi ica ion. I is ine uned 80 epochs un il
con e gence, using Adam op imize wi h a lea ning a e o
5×10−5.
CLIP+ML-Decode me hod was ine- uned by eezing he
CLIP model and only ocusing on upda ing he weigh s o
Decode . A g id sea ch o e all he possible hype -pa ame e s
o he decode block was made and he bes esul s was
ob ained wi h: mul i-head a en ion head amoun 4, d opou
in eed o wa d ne wo k 0.5, numbe o g oups K 8, numbe
o laye s (decode block amoun ) 2. Simila ly o he p e ious
me hods, Adam op imize is used, wi h a lea ning a e o
1×10−4.
C. E alua ion Me ics
We e alua e model pe o mance using s anda d classi ica-
ion me ics: p ecision, ecall, and 1-sco e. We u he mo e
analyze he scalabili y o he model.
•P ecision: Measu es he p opo ion o co ec ly p edic ed
labels among all p edic ed labels. 3
•Recall: Measu es he p opo ion o co ec ly p edic ed
labels among all ue labels.4
•F1-Sco e: The ha monic mean o p ecision and ecall. 5
•Scalabili y: Assesses how well he model adap s o
inc easing class sizes.
P ecision =T P
T P +F P (3)
Recall =T P
T P +F N (4)
F1sco e =2×P ecision ×Recall
P ecision +Recall (5)
The ue posi i es, alse posi i es, alse nega i es a e cal-
cula ed simila ly o single-label classi ica ion. Fig. 5 shows
an image sample ha ing wo g ound u h labels: VISA and
Bankno es. In case he p edic ed classes a e ”VISA and
Paypal” class, a e ex ac ing he single labels as ”VISA” and
”Paypal”, his p edic ion would con ibu e as a ue posi i e
o ”VISA” class, alse posi i e o ”Paypal” class and a alse
nega i e o ”Bankno es” classes.
Fig. 5: An example mul i-labeled image wi h 2 classes.
V. RESULTS AND DISCUSSION
Table I compa es he pe o mance o he ou me hods
in e ms o p ecision, ecall, 1-sco e and scalabili y and
summa izes ou indings. CLIP+Label Empwoe me hod has
a low scalabili y since i has O(N!) as wo s case scena io,
CLIP Sigmoid and SIGLIP has a mode a e scalabili y since
hey p o ide an imp o emen , bu nei he b ing any g ow h,
meaning he wo s and bes case scena io is he same and
O(N). While CLIP+ML-Decode imp o es he scalabili y om
O(N) o O(K), being N is he numbe o classes and K is he
numbe o g oups de ined by he end-use .
Table II compa es he wo models gi ing he bes esul s
o da k web da ase , on he MS-COCO mul i-label da ase ,
causing he CLIP+LE me hod o ha e a huge d op o accu acy,
due o expanding 80 base classes o 234,581 ho -encoded
classes. These esul s show he a o emen ioned po en ial isk
o CLIP+LE me hod on he da ase s ha ing high ela ion
be ween he classes, causing he exis ence o a ious pe mu-
a ions o mul i-label ec o s among he da ase . E en hough
he esul s migh be imp essi e o such amoun o classes, i
is isible ha he Label Empowe me hod has a huge impac on
CLIP mul i-modal model’s classi ica ion capaci y by exploding
he class numbe s unnecessa ily.
TABLE I
COMPARISON OF FEW-SHOT MULTI-LABEL CLASSIFICATION METHODS. LE:
LABEL EMPOWER, MLD: ML-DECODER
Me hod P ecision Recall F1 sco e Scalabili y
CLIP+LE 0.944 0.931 0.937 Low
CLIP Sigmoid 0.493 0.276 0.351 Mode a e
SIGLIP 0.880 0.735 0.801 Mode a e
CLIP+MLD 0.958 0.916 0.936 High
Ou esul s show ha he CLIP+Label Empowe Adap e
achie es he bes ecall, making i s ill a s ong candida e
o ecall-sensi i e applica ions buil o da ase s no ha ing
huge amoun o classes, due o i s low scalabili y and he
po en ial isk o p o iding poo e esul s wi h a huge numbe
o classes. On he o he hand, he esul s show he medium-
le el scalabili y models (CLIP+Sigmoid, SIGLIP), ha a e no
imp o ing no wo sen he a chi ec u e acco ding o he numbe
o classes, ha e poo accu acy in compa ison wi h CLIP
+ adap e solu ions. Howe e , CLIP+ML-Decode pe o ms
well in bo h p ecision and ecall, while also add essing he
scalabili y issue e ec i ely and i should be a de ini e choice
o da ase ha ing huge numbe o classes, whe e o no ha e
alse posi i e p edic ions is mo e impo an han no missing
any ue posi i e p edic ion, ega ding i ’s p ecision bea ing he
CLIP+Label Empowe me hod while p o iding lowe ecall.
TABLE II
COMPARISON OF MULTI-LABEL MS-COCO MAP SCORE. LE: LABEL
EMPOWER AND MLD: MULTI-LABEL DECODER
Me hod P ecision Recall
CLIP+LE 0.592 0.555
CLIP+MLD 0.839 0.809
VI. CONCLUSION AND FUTURE WORKS
In his s udy, we analyzed ou ew-sho mul i-label classi-
ica ion me hods based on CLIP. Ou esul s demons a e ha
CLIP+Label Empowe Adap e excels in he ecall, whe eas
CLIP+ML-Decode p o ides a mo e scalable solu ion by
mi iga ing he exponen ial g ow h p oblem in inpu ea u es,
p o iding also a obus pe o mance on accu acy me ics.
Fu u e wo k will explo e he e icien deploymen pipeline
o CLIP+ML-Decode app oach o eal-wo ld mul i-label
classi ica ion asks.
ACKNOWLEDGEMENTS
The wo k desc ibed in his pape is pe -
o med in he H2020 p ojec STARLIGHT
(”Sus ainable Au onomy and Resilience o
LEAs using AI agains High P io i y
Th ea s”). This p ojec has ecei ed unding
om he Eu opean Union’s Ho izon 2020
esea ch and inno a ion p og am unde g an
ag eemen No 101021797.
REFERENCES
[1] M.-L. Zhang, Y.-K. Li, X.-Y. Liu, and X. Geng, “Bina y ele ance o
mul i-label lea ning: An o e iew,” F on ie s o Compu e Science,
ol. 12, no. 2, pp. 191–202, Ma . 2018, ISSN: 2095-2236. DOI: 10.
1007/s11704-017-7031-7. [Online]. A ailable: h p://dx.doi.o g/10.
1007/s11704-017-7031-7.
[2] G. Tsoumakas, A. Dimou, E. Spy omi os-Xiou is, V. Meza is, I.
Kompa sia is, and I. Vlaha as, “Co ela ion-based p uning o s acked
bina y ele ance models o mul i-label lea ning,” Jan. 2009, pp. 101–
116.
[3] E. Ben-Ba uch, T. Ridnik, N. Zami , e al.,Asymme ic loss o mul i-
label classi ica ion, 2021. a Xi : 2009.14119 [cs.CV].
[4] A. F. T. Ma ins and R. F. As udillo, F om so max o spa semax: A
spa se model o a en ion and mul i-label classi ica ion, 2016. a Xi :
1602.02068 [cs.CL].
[5] Y. Wang, D. He, F. Li, e al.,Mul i-label classi ica ion wi h label
g aph supe imposing, 2019. a Xi : 1911.09243 [cs.CV].
[6] Z.-M. Chen, X.-S. Wei, P. Wang, and Y. Guo, Mul i-label image
ecogni ion wi h g aph con olu ional ne wo ks, 2019. a Xi : 1904.
03582 [cs.CV].
[7] H. Tou on, M. Co d, M. Douze, F. Massa, A. Sablay olles, and H.
J´
egou, T aining da a-e icien image ans o me s dis illa ion h ough
a en ion, 2021. a Xi : 2012.12877 [cs.CV].
[8] W. Zhang, C. Liu, L. Zeng, B. Ooi, S. Tang, and Y. Zhuang, “Lea ning
in impe ec en i onmen : Mul i-label classi ica ion wi h long- ailed
dis ibu ion and pa ial labels,” in P oceedings o he IEEE/CVF
In e na ional Con e ence on Compu e Vision (ICCV), Oc . 2023,
pp. 1423–1432.
[9] A. Rad o d, J. W. Kim, C. Hallacy, e al.,Lea ning ans e able isual
models om na u al language supe ision, 2021. a Xi : 2103.00020
[cs.CV].
[10] J. Read, B. P ah inge , G. Holmes, and E. F ank, “Classi ie chains: A
e iew and pe spec i es,” Jou nal o A i icial In elligence Resea ch,
ol. 70, pp. 683–718, Feb. 2021, ISSN: 1076-9757. DOI: 10.1613/jai .
1.12376. [Online]. A ailable: h p://dx.doi.o g/10.1613/jai .1.12376.
[11] Y. Wei, W. Xia, M. Lin, e al., “Hcp: A lexible cnn amewo k
o mul i-label image classi ica ion,” IEEE T ansac ions on Pa e n
Analysis and Machine In elligence, ol. 38, no. 9, pp. 1901–1907,
Sep. 2016, ISSN: 1939-3539. DOI: 10 . 1109 / pami . 2015 . 2491929.
[Online]. A ailable: h p://dx.doi.o g/10.1109/TPAMI.2015.2491929.
[12] T. Chen, M. Xu, X. Hui, H. Wu, and L. Lin, Lea ning seman ic-
speci ic g aph ep esen a ion o mul i-label image ecogni ion, 2019.
a Xi : 1908.07325 [cs.CV].
[13] Y. Liu, L. Sheng, J. Shao, J. Yan, S. Xiang, and C. Pan, “Mul i-
label image classi ica ion ia knowledge dis illa ion om weakly-
supe ised de ec ion,” in P oceedings o he 26 h ACM in e na ional
con e ence on Mul imedia, ACM, Oc . 2018. DOI: 10.1145/3240508.
3240567. [Online]. A ailable: h p://dx.doi.o g/10.1145/3240508.
3240567.
[14] A. Vaswani, N. Shazee , N. Pa ma , e al.,A en ion is all you need,
2023. a Xi : 1706.03762 [cs.CL].
[15] J. De lin, M.-W. Chang, K. Lee, and K. Tou ano a, Be : P e- aining
o deep bidi ec ional ans o me s o language unde s anding, 2019.
a Xi : 1810.04805 [cs.CL].
[16] A. Rad o d, J. Wu, R. Child, D. Luan, D. Amodei, and I. Su ske e ,
“Language models a e unsupe ised mul i ask lea ne s,” 2019. [On-
line]. A ailable: h ps://api.seman icschola .o g/Co pusID:160025533.
[17] T. B. B own, B. Mann, N. Ryde , e al.,Language models a e ew-
sho lea ne s, 2020. a Xi : 2005.14165 [cs.CL].
[18] X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, De o mable
de : De o mable ans o me s o end- o-end objec de ec ion, 2021.
a Xi : 2010.04159 [cs.CV].
[19] X. Cheng, H. Lin, X. Wu, e al.,Ml : Mul i-label classi ica ion wi h
ans o me , 2021. a Xi : 2106.06195 [cs.CV].
[20] S. Liu, L. Zhang, X. Yang, H. Su, and J. Zhu, Que y2label: A
simple ans o me way o mul i-label classi ica ion, 2021. a Xi :
2107.10834 [cs.CV].
[21] J. Ye, J. He, X. Peng, W. Wu, and Y. Qiao, A en ion-d i en dynamic
g aph con olu ional ne wo k o mul i-label image ecogni ion, 2020.
a Xi : 2012.02994 [cs.CV].
[22] K. P oko ie and V. So aso , Combining me ic lea ning and a en-
ion heads o accu a e and e icien mul ilabel image classi ica ion,
2022. a Xi : 2209.06585 [cs.CV].
[23] F. Zhu, H. Li, W. Ouyang, N. Yu, and X. Wang, Lea ning spa ial
egula iza ion wi h image-le el supe isions o mul i-label image
classi ica ion, 2017. a Xi : 1702.05891 [cs.CV].
[24] Y. C¸ . Ak as¸ and J. G. Cas a˜
no, “Few-sho mul i-label mul i-class
classi ica ion o da k web image ca ego iza ion,” in 2024 12 h
In e na ional Symposium on Digi al Fo ensics and Secu i y (ISDFS),
2024, pp. 1–6. DOI: 10.1109/ISDFS60797.2024.10527297.
[25] X. Zhai, B. Mus a a, A. Kolesniko , and L. Beye , Sigmoid loss o
language image p e- aining, 2023. a Xi : 2303 .15343 [cs.CV].
[Online]. A ailable: h ps://a xi .o g/abs/2303.15343.
[26] T. Ridnik, G. Sha i , A. Ben-Cohen, E. Ben-Ba uch, and A. Noy,
Ml-decode : Scalable and e sa ile classi ica ion head, 2021. a Xi :
2111.12933 [cs.CV]. [Online]. A ailable: h ps://a xi .o g/abs/
2111.12933.