scieee Science in your language
[en] (orig)

CLIP-based Few-Shot Multi-Label Classification Methods: A Comparative Study

Abstract

Categorizing darkweb image content is critical for identifying and averting potential threats. However, this remains a challenge due to the nature of the data, which includes multiple co-existing domains and intra-class variations. While many methods have been proposed to classify this image content, multi-label multi-class classification remains underexplored. The complexity of darkweb imagery, combined with the need for efficient classification systems, demands innovative approaches that can handle both the technical challenges and the sensitive nature of the content. In this paper, we present a comparative study of few-shot multi-label classification methods using the multimodal model CLIP. Our research addresses the growing need for robust classification systems that can effectively catego-rize diverse and complex image content while maintaining high accuracy and computational efficiency. We particularly focus on the challenges of handling multiple labels simultaneously and the scalability of these systems in real-world applications. We analyze and compare four different approaches: CLIP+Label Empower Adapter, CLIP Sigmoid, SIGLIP, and CLIP+ML-Decoder. Our study evaluates these methods based on their precision, recall, and ability to handle increasing class numbers efficiently. Finally, our research contributes to the field by providing detailed insights into the strengths and limitations of each method.

Read accessible full text

CLIP-based Few-Shot Multi-Label Classification Methods: A Comparative Study

Author: Aktas, Yagmur; García, Jorge
Publisher: Zenodo
DOI: 10.1109/ISDFS65363.2025.11012107
Source: https://zenodo.org/records/17227247/files/CLIP_based_Multi_Label_Classification_Methods__A_Comparative_Study.pdf
CLIP-based Few-Sho Mul i-Label Classi ica ion
Me hods: A Compa a i e S udy
Ya˘
gmu C¸ i˘
gdem Ak as¸
Vicom ech Founda ion
Basque Resea ch and Technology Alliance (BRTA)
Mikele egi 57, 20009 Donos ia-San Sebas i´
an (Spain)
[email p o ec ed]g
Jo ge Ga c´
ıa Cas a˜
no
Vicom ech Founda ion
Basque Resea ch and Technology Alliance (BRTA)
Mikele egi 57, 20009 Donos ia-San Sebas i´
an (Spain)
[email p o ec ed]g
Abs ac —Ca ego izing da kweb image con en is c i ical o
iden i ying and a e ing po en ial h ea s. Howe e , his emains
a challenge due o he na u e o he da a, which includes
mul iple co-exis ing domains and in a-class a ia ions. While
many me hods ha e been p oposed o classi y his image con en ,
mul i-label mul i-class classi ica ion emains unde explo ed. The
complexi y o da kweb image y, combined wi h he need o
e icien classi ica ion sys ems, demands inno a i e app oaches
ha can handle bo h he echnical challenges and he sensi i e
na u e o he con en . In his pape , we p esen a compa a i e
s udy o ew-sho mul i-label classi ica ion me hods using he
mul imodal model CLIP. Ou esea ch add esses he g owing
need o obus classi ica ion sys ems ha can e ec i ely ca ego-
ize di e se and complex image con en while main aining high
accu acy and compu a ional e iciency. We pa icula ly ocus on
he challenges o handling mul iple labels simul aneously and he
scalabili y o hese sys ems in eal-wo ld applica ions. We analyze
and compa e ou di e en app oaches: CLIP+Label Empowe
Adap e , CLIP Sigmoid, SIGLIP, and CLIP+ML-Decode . Ou
s udy e alua es hese me hods based on hei p ecision, ecall,
and abili y o handle inc easing class numbe s e icien ly. Finally,
ou esea ch con ibu es o he ield by p o iding de ailed insigh s
in o he s eng hs and limi a ions o each me hod.
Index Te ms—Mul i-label image classi ica ion, Mul i-class im-
age classi ica ion, CLIP, SIGLIP, ML-Decode
I. INTRODUCTION
The da k web, a specialized subse o he as global
in e ne , is accessible only h ough dedica ed web b owse s.
I is widely ecognized as a hub o illici ac i i ies due o
i s co e cha ac e is ics: anonymi y, which p o ides use s wi h
a le el o p i acy no ypically a ailable on he su ace web,
and un aceabili y, which makes i ex emely challenging o
ack he o igins and des ina ions o da a ans e s. As a esul ,
he da k web se es as a c i ical sou ce o h ea in elligence,
o e ing insigh s in o cybe -a acks, s olen asse s, illegal ade,
a ms a icking, con iden ial da a leaks, child exploi a ion, and
o he c iminal ac i i ies.
Au oma ically analyzing he isual con en o he da k
web —such as images and ideos- is essen ial o e icien
image ca ego iza ion, enabling he de ec ion and p e en ion
o po en ial h ea s. Howe e , da k web image classi ica ion
emains an open challenge due o he complex na u e o i s
con en .
Fig. 1: Mul i-label image examples om Da k web da ase
Many images su e om low esolu ion, blu , and small
objec sizes, o en con aining objec s ha blend in o he back-
g ound due o colo simila i ies. Addi ionally, he coexis ence
o mul iple o e lapping domains and in a-class a ia ions —
whe e ins ances wi hin he same ca ego y exhibi signi ican
isual di e ences — u he complica e classi ica ion asks, as
illus a ed in Fig. 1.
This complexi y highligh s he impo ance o mul i-label
classi ica ion, whe e an image can be assigned mul iple ele-
an labels a he han a single ca ego y.
A ange o mul i-label classi ica ion echniques ha e been
explo ed o ackle his p oblem. Some app oaches decom-
pose mul i-label classi ica ion in o sepa a e single-label p ob-
lems[1], [2], while o he s e ine loss unc ions o ac i a ion
unc ions o enhance classi ica ion accu acy[3], [4]. Addi ion-
ally, g aph-based ne wo ks u ilize lea ned label embeddings o
loca e key disc imina i e ea u es [5], [6], while ans o me -
based a chi ec u es ha e been adap ed o mul i-label classi i-
ca ion asks [7].
While hese me hods show p omising esul s, hey ely on
la ge, well-balanced da ase s—which a e di icul o cu a e in
eal-wo ld applica ions. Expanding a da ase wi h a balanced
class dis ibu ion is an exponen ially complex p ocess [8],
making la ge-scale mul i-label lea ning in easible in many
scena ios.
To add ess his, ecen ad ances in ision-language models
like CLIP [9] o e new possibili ies. By aligning isual and
ex ual embeddings, CLIP enables ze o-sho and ew-sho
lea ning in open- ocabula y se ings. Howe e , mos CLIP-979-8-3315-0993-4/25/31.002025IEEE
based me hods a e ailo ed o single-label classi ica ion.
In his s udy, we add ess he p oblem o ew-sho mul i-
label image classi ica ion, ocusing on he da k web as a
eal-wo ld, high-s akes applica ion domain. Ou hypo hesis is
ha adap ing s a e-o - he-a ew-sho mul i-label classi ica ion
echniques o CLIP can imp o e gene aliza ion and pe o -
mance in low- esou ce, high- a iabili y en i onmen s.
Ou con ibu ions can be summa ized as ollows:
•We analyze ou s a e-o - he-a ew-sho mul i-label
classi ica ion echniques and p o ide a po able pipeline
applicable o any eal-wo ld da ase equi ing ew-sho
mul i-label classi ica ion.
•We e alua e hese app oaches by implemen ing CLIP-
based me hodologies o da k web mul i-label image
classi ica ion.
•We compa e ou p e ious app oach wi h newe CLIP-
based mul i-label me hodologies, p esen ing a anspa en
compa a i e s udy which con ibu es o he ield by p o-
iding de ailed insigh s in o he s eng hs and limi a ions
o each me hod.
II. RELATED WORK
A a ie y o me hods ha e been p oposed o sol e he mul i-
label classi ica ion p oblem, which can be b oadly ca ego ized
in o (i) p oblem ans o ma ion me hods and (ii) algo i hm
adap a ion echniques.
Ea ly esea ch ocused on decomposing mul i-label classi i-
ca ion in o independen bina y classi ica ion asks[1]. How-
e e , his app oach ails o cap u e in e -label co ela ions
and equi es aining sepa a e classi ie s o each ca e-
go y—in oducing a signi ican compu a ional bu den. To ad-
d ess his, some s udies p opose classi ie chaining[10], whe e
each classi ie ’s ou pu se es as an inpu ea u e o sub-
sequen classi ie s. While his s a egy imp o es in e -label
dependency modeling, i s ill su e s om high compu a ional
cos s due o he g owing numbe o classi ie s equi ed.
Ano he ans o ma ion-based app oach, known as Label
Powe se (LP) [2], con e s mul i-label p oblems in o mul i-
class p oblems by ea ing each unique label combina ion as
a dis inc class. Howe e , LP s uggles wi h he exponen ial
g ow h o label combina ions, making i imp ac ical o la ge
da ase s.
Se e al me hods aim o imp o e mul i-label classi ica ion
by op imizing he loss unc ion[3] o ac i a ion unc ion[4] o
mi iga e class imbalance issues.
In compu e ision, mul iple app oaches ha e been de el-
oped o mul i-label image classi ica ion. Hypo heses-CNN-
Pooling (HCP) [11] gene a es a la ge numbe o p oposals
h ough objec de ec ion echniques, ea ing each p oposal as
a single-label classi ica ion p oblem.
Beyond CNN-based me hods, G aph Con olu ional Ne -
wo ks (GCNs) ha e demons a ed high e icacy ac oss a i-
ous ision asks, including mul i-label classi ica ion [5], [6],
[12]. GCN-based me hods build classi ie s by modeling label
ela ionships wi hin g aph ne wo ks.
O he app oaches employ weakly supe ised lea ning o
mul i-label classi ica ion [13], le e aging knowledge dis illa-
ion echniques om objec de ec ion models.
Mo e ecen ly, ans o me s[14]—o iginally de eloped o
modeling long- ange dependencies in na u al language p o-
cessing (NLP)[15]–[17]—ha e demons a ed s ong pe o -
mance ac oss compu e ision asks, including image classi i-
ca ion[7], [9] and objec de ec ion[18].
Se e al ans o me -based mul i-label classi ica ion ech-
niques ha e eme ged:
Mul i-label T ans o me (MLT) [19], which models pixel-
wise a en ion o enhance ea u e ex ac ion. T ans o me -
based que y lea ning [20], which uses decode que ies o
p edic label exis ence. G aph-enhanced T ans o me s [21],
which combine GCNs wi h a en ion mechanisms. Me ic
Lea ning T ans o me s [22], which in eg a e me ic lea ning
o assess label simila i ies. T ans o me -CNN Hyb ids [23],
which gene a e a en ion maps pe label, cap u ing in a-class
dependencies h ough con olu ional laye s.
Despi e hese ad ancemen s, exis ing mul i-label me hods
s ill equi e la ge da ase s o achie e high accu acy. Howe e ,
da a sca ci y is a p e alen challenge in eal-wo ld applica-
ions, whe e collec ing and labeling ex ensi e da ase s is bo h
esou ce-in ensi e and ime-consuming.
La ge-scale open- ocabula y models ained on as
da ase s, such as CLIP [9], o e a p omising al e na i e by
le e aging ze o-sho single-label classi ica ion and ew-sho
mul i-label classi ica ion. By adap ing hese models, we aim
o explo e hei po en ial o enhancing mul i-label classi i-
ca ion in da k web image y, whe e adi ional me hods ace
signi ican cons ain s.
III. CLIP-BASED MULTI-LABEL CLASSIFICATION
METHODOLOGIES
The e a e ou ecen mul i-label classi ica ion me hods sui -
able o CLIP mul i-modal model adap a ion: (i) CLIP+Label
Empowe (ou p e ious model), (ii) CLIP Sigmoid (iii) SIGLIP
(i ) CLIP + ML-Decode
A. CLIP+Label Empowe
Label Empowe [2] is one o he app oaches ha con e he
mul i-label classi ica ion p oblem in o a single-label classi i-
ca ion, by mul i-labels o bina y ec o s by ho -encoding hem
and assigning a unique label o each di e en bina y ec o .
As demons a ed in [24], i has he highes pe o mance among
he o he app oaches o sol ing he mul i-label classi ica ion
ask by ans o ming he p oblem in o a single-label one. Fig.
2 shows he a chi ec u e o ou p e iously p oposed me hod
CLIP+Label Empowe Adap e .
Al hough ou p e iously p oposed model has he highes
ecall, as shown in Table I, among e en hese ou s a e-
o - he-a CLIP-adap ed mul i-label classi ica ion app oaches,
no only be ween he adi ional ones, i p ese es a isk o
accu acy d op wi hin a la ge numbe o classes due o i s
exponen ially g owing ho -encoded labels. This possible isk
is especially impo an o such a da ase like da kweb, which
Fig. 2: CLIP based Mul i-Label Me hodologies
is expec ed o g ow as in he sense o numbe o classes,
since new ype o images and i les a e expec ed o occu .
•Wo s Case: O(N!) (All class labels exis in all possible
pe mu a ions in he da ase )
•Bes Case: O(N)(The da ase con ains only single
labels)
B. CLIP Sigmoid
CLIP has he capabili y o classi y an ex emely huge num-
be o di e en classes, hanks o i s open- ocabula y na u e.
Ye we migh need o ine- une his s a e-o - he-a model o
ou eal-wo ld da ase s like da kweb da ase . In a single-label
pipeline, he p ope way o ine- uning such a model is o
add a linea laye ollowing he image and ex embeddings,
ha ing ou pu nodes as much as class numbe s in he da ase .
In a single-label, mul i-class classi ica ion scena io, so max
is commonly used o dis ibu e p obabili ies among di e en
classes, ensu ing ha he sum o all class p obabili ies equals
1.
P(yi|x) = ezi
PN
j=1 ezj
(1)
whe e zi ep esen s he logi s o class iand Nis he o al
numbe o classes.
Howe e , in mul i-label classi ica ion, whe e mul iple labels
can be p esen simul aneously, so max is subop imal as i
o ces mu ual exclusi i y among classes. Ins ead, sigmoid
ac i a ion is mo e app op ia e, as i independen ly p edic s
each label’s p obabili y:
P(yi|x) = 1
1 + e−zi(2)
This allows he model o assign a p obabili y close o 1.0
o each ele an label in an image, ensu ing ha all p esen
classes a e p edic ed wi hou compe i ion om o he labels.
Thus, ine- uning CLIP wi h a linea classi ica ion head
using Sigmoid ac i a ion is ano he app oach. Al hough his
me hod does no ha e any isk o exponen ially g owing wi h
he numbe o class numbe s, ou expe imen al esul s show
i s poo accu acy in compa ison wi h ou p e iously p oposed
me hod.
C. SIGLIP
While SIGLIP [25] is a s a e-o - he-a , mul i-modal model,
ha ing a e y simila a chi ec u e o CLIP, i s main di e ence
is using Sigmoid, ins ead o So max while p e- aining. This
ac p o ides us wi h he possibili y o see i he poo pe o -
mance o CLIP + Sigmoid ine- uning me hodology occu ed
due o using a di e en ac i a ion unc ion in he ine- uning
han p e- aining.
Fig. 2 shows he e y simila a chi ec u e o SIGLIP.
D. CLIP+ML-Decode
ML-Decode [26] is a s a e-o - he-a decode implemen-
a ion aiming o o e come he exponen ially g owing inpu
size in classi ica ion pipelines buil wi h de aul ans o me
decode s (Fig. 3) due o he sel -a en ion laye . The me hod
p oposes o emo e he sel -a en ion laye and claims i s
a ec less on he classi ica ion accu acy. I is also claimed o
use ”g oup que ies”, ins ead o ixed que y embeddings pe
class as in adi ional ans o me -decode s. Being he numbe
o g oups a new hype pa ame e in his a chi ec u e e e s
o he amoun o inpu que y embeddings he decode will
ha e, independen ly om he numbe o classes in he da ase .
Being K is he numbe o g oup que ies, he decode lea ns
o c ea e K que ies e e encing N numbe o classes in a
da ase , ins ead o ha ing N ixed que ies o N classes as
in de aul ans o me decode s. CLIP+ML-Decode app oach
uses bo h classi ica ion loss coming om he inal p edic ion
logi s and an alignmen loss coming om he simila i y o ex
and image embeddings o ob ain a inal loss. The ixed ex
embeddings a e used only o his pu pose, whe eas non- ixed
g oup que ies a e lea nable embeddings and hey a e upda ed
du ing he aining.
Fig. 3: T adi ional s ML Decode s
The e o e his app oach enhances scalabili y, wi h wo key
op imiza ions: (1) Sel -a en ion emo al, which educes he
compu a ional complexi y om O(N2) o O(N), and (2)
G oup decoding, which u he op imizes in e ence by shi ing
he complexi y om O(N) o O(K), whe e KK ep esen s he
numbe o meaning ul g oups ins ead o p ocessing all classes
independen ly.
Sel -a en ion emo al:
O(N2)→O(N)
G oup decoding:
O(N)→O(K)
Fig. 2 shows he a chi ec u e o CLIP + ML-Decode .
IV. EXPERIMENTAL STUDY
A. Da k web da ase
The Da k web da ase , sou ced om CFLW’s Da k Web
Moni o 1, includes images collec ed om da k web domains
abou a ious c ime ca ego ies like inancial c ime and o gani-
za ions, d ugs and na co ics, weapons as well as hei endo s
as indi idual classes. Simila o many eal-wo ld da ase s,
he da k web da ase con ains images wi h mul iple labels,
meaning an image can belong o mo e han one class. I also
includes images wi h a single label. Fig. 1 shows some image
samples along wi h hei g ound u h labels om a ious
classes.
1h ps://c lw.com/dwm/
Fig. 4: Da a dis ibu ion o he da k web da ase o bo h ain and es se s.
The da ase con ains 46 classes om a ious ca ego ies
like D ugs, Na co ics, Weapons, and Financial O ganiza ions
which do no ha e a s ong co ela ion be ween hem bu
i includes subca ego ies ha ing a s ong ela ionship and
he possibili y o being exis ing in he same image. Being
he cu en da kweb da ase is an expe imen al one, many
mo e classes a e expec ed o join and hus, he scalabili y
pe o mance is impo an han a egula Fig. 4 shows he
da a dis ibu ion o he da k web da ase o bo h ain and
es se s.
The da kweb da ase p esen s an imbalanced da a p oblem.
Some dominan classes ha e many samples, whe eas some
o he classes su e om a lack o da a. Some classes like
d ugs and na co ics, coming along wi h any ype o d ug ype
o any indi idual endo ha has o be ca ego ized, become
a e y dominan class by collec ing only a ew samples o
i s subg oups. Fig. 1 shows 3 image samples om he da k
web da ase , 1 belonging o a inancial o ganiza ion ca ego y,
2 belonging o he d ugs and na co ics ca ego y con aining
pic u es o weed and he endo names. While endo names
may a y, each image ha ing a endo name in o ma ion also
con ains weed class, doubling he amoun o ”weed” class in
compa ison o class names e e ing o he indi idual endo
names. Which is obus example showing he eason o s ong
class imbalance p oblem in da kweb da ase .
Las ly, ano he challenge is ha some classes, especially he
endo names like ”deep shop”, o ” anda al”, ha e ex ual
in o ma ion a he han isual ea u es which b ings he need
o conside ing he combina ion o ision and language da a.
B. Implemen a ion De ails
The ine- uning s a egy a ies o each o he me hods we
examine: CLIP + Label Empowe was ine uned by eezing
he CLIP p e- ained model, and upda ing he weigh s o only
o be ween 5 o 10 epochs, using Adam op imize wi h de aul
lea ning a e 1×10−3. So max inal ac i a ion unc ion and
C oss En opy Loss.
CLIP Sigmoid model was ine uned by eezing he CLIP
p e- ained model, and upda ing he weigh s o only he ad-
di ional linea classi ica ion head, o 50 epochs un il con e -
gence, using he same op imize and lea ning a e as p e ious
app oach, wi h a Bina y C oss En opy Loss o e he logi s,
as a de aul app oach o sigmoid classi ica ion.
SIGLIP model was ine uned wi hou eezing any pa
o he a chi ec u e, since ou implemen a ion any addi ional
pa and he na u e o he a chi ec u e is al eady compa ible
o mul i-label classi ica ion. I is ine uned 80 epochs un il
con e gence, using Adam op imize wi h a lea ning a e o
5×10−5.
CLIP+ML-Decode me hod was ine- uned by eezing he
CLIP model and only ocusing on upda ing he weigh s o
Decode . A g id sea ch o e all he possible hype -pa ame e s
o he decode block was made and he bes esul s was
ob ained wi h: mul i-head a en ion head amoun 4, d opou
in eed o wa d ne wo k 0.5, numbe o g oups K 8, numbe
o laye s (decode block amoun ) 2. Simila ly o he p e ious
me hods, Adam op imize is used, wi h a lea ning a e o
1×10−4.
C. E alua ion Me ics
We e alua e model pe o mance using s anda d classi ica-
ion me ics: p ecision, ecall, and 1-sco e. We u he mo e
analyze he scalabili y o he model.
•P ecision: Measu es he p opo ion o co ec ly p edic ed
labels among all p edic ed labels. 3
•Recall: Measu es he p opo ion o co ec ly p edic ed
labels among all ue labels.4
•F1-Sco e: The ha monic mean o p ecision and ecall. 5
•Scalabili y: Assesses how well he model adap s o
inc easing class sizes.
P ecision =T P
T P +F P (3)
Recall =T P
T P +F N (4)
F1sco e =2×P ecision ×Recall
P ecision +Recall (5)
The ue posi i es, alse posi i es, alse nega i es a e cal-
cula ed simila ly o single-label classi ica ion. Fig. 5 shows
an image sample ha ing wo g ound u h labels: VISA and
Bankno es. In case he p edic ed classes a e ”VISA and
Paypal” class, a e ex ac ing he single labels as ”VISA” and
”Paypal”, his p edic ion would con ibu e as a ue posi i e
o ”VISA” class, alse posi i e o ”Paypal” class and a alse
nega i e o ”Bankno es” classes.
Fig. 5: An example mul i-labeled image wi h 2 classes.
V. RESULTS AND DISCUSSION
Table I compa es he pe o mance o he ou me hods
in e ms o p ecision, ecall, 1-sco e and scalabili y and
summa izes ou indings. CLIP+Label Empwoe me hod has
a low scalabili y since i has O(N!) as wo s case scena io,
CLIP Sigmoid and SIGLIP has a mode a e scalabili y since
hey p o ide an imp o emen , bu nei he b ing any g ow h,
meaning he wo s and bes case scena io is he same and
O(N). While CLIP+ML-Decode imp o es he scalabili y om
O(N) o O(K), being N is he numbe o classes and K is he
numbe o g oups de ined by he end-use .
Table II compa es he wo models gi ing he bes esul s
o da k web da ase , on he MS-COCO mul i-label da ase ,
causing he CLIP+LE me hod o ha e a huge d op o accu acy,
due o expanding 80 base classes o 234,581 ho -encoded
classes. These esul s show he a o emen ioned po en ial isk
o CLIP+LE me hod on he da ase s ha ing high ela ion
be ween he classes, causing he exis ence o a ious pe mu-
a ions o mul i-label ec o s among he da ase . E en hough
he esul s migh be imp essi e o such amoun o classes, i
is isible ha he Label Empowe me hod has a huge impac on
CLIP mul i-modal model’s classi ica ion capaci y by exploding
he class numbe s unnecessa ily.
TABLE I
COMPARISON OF FEW-SHOT MULTI-LABEL CLASSIFICATION METHODS. LE:
LABEL EMPOWER, MLD: ML-DECODER
Me hod P ecision Recall F1 sco e Scalabili y
CLIP+LE 0.944 0.931 0.937 Low
CLIP Sigmoid 0.493 0.276 0.351 Mode a e
SIGLIP 0.880 0.735 0.801 Mode a e
CLIP+MLD 0.958 0.916 0.936 High
Ou esul s show ha he CLIP+Label Empowe Adap e
achie es he bes ecall, making i s ill a s ong candida e
o ecall-sensi i e applica ions buil o da ase s no ha ing
huge amoun o classes, due o i s low scalabili y and he
po en ial isk o p o iding poo e esul s wi h a huge numbe
o classes. On he o he hand, he esul s show he medium-
le el scalabili y models (CLIP+Sigmoid, SIGLIP), ha a e no
imp o ing no wo sen he a chi ec u e acco ding o he numbe
o classes, ha e poo accu acy in compa ison wi h CLIP
+ adap e solu ions. Howe e , CLIP+ML-Decode pe o ms
well in bo h p ecision and ecall, while also add essing he
scalabili y issue e ec i ely and i should be a de ini e choice

o da ase ha ing huge numbe o classes, whe e o no ha e
alse posi i e p edic ions is mo e impo an han no missing
any ue posi i e p edic ion, ega ding i ’s p ecision bea ing he
CLIP+Label Empowe me hod while p o iding lowe ecall.
TABLE II
COMPARISON OF MULTI-LABEL MS-COCO MAP SCORE. LE: LABEL
EMPOWER AND MLD: MULTI-LABEL DECODER
Me hod P ecision Recall
CLIP+LE 0.592 0.555
CLIP+MLD 0.839 0.809
VI. CONCLUSION AND FUTURE WORKS
In his s udy, we analyzed ou ew-sho mul i-label classi-
ica ion me hods based on CLIP. Ou esul s demons a e ha
CLIP+Label Empowe Adap e excels in he ecall, whe eas
CLIP+ML-Decode p o ides a mo e scalable solu ion by
mi iga ing he exponen ial g ow h p oblem in inpu ea u es,
p o iding also a obus pe o mance on accu acy me ics.
Fu u e wo k will explo e he e icien deploymen pipeline
o CLIP+ML-Decode app oach o eal-wo ld mul i-label
classi ica ion asks.
ACKNOWLEDGEMENTS
The wo k desc ibed in his pape is pe -
o med in he H2020 p ojec STARLIGHT
(”Sus ainable Au onomy and Resilience o
LEAs using AI agains High P io i y
Th ea s”). This p ojec has ecei ed unding
om he Eu opean Union’s Ho izon 2020
esea ch and inno a ion p og am unde g an
ag eemen No 101021797.
REFERENCES
[1] M.-L. Zhang, Y.-K. Li, X.-Y. Liu, and X. Geng, “Bina y ele ance o
mul i-label lea ning: An o e iew,” F on ie s o Compu e Science,
ol. 12, no. 2, pp. 191–202, Ma . 2018, ISSN: 2095-2236. DOI: 10.
1007/s11704-017-7031-7. [Online]. A ailable: h p://dx.doi.o g/10.
1007/s11704-017-7031-7.
[2] G. Tsoumakas, A. Dimou, E. Spy omi os-Xiou is, V. Meza is, I.
Kompa sia is, and I. Vlaha as, “Co ela ion-based p uning o s acked
bina y ele ance models o mul i-label lea ning,” Jan. 2009, pp. 101–
116.
[3] E. Ben-Ba uch, T. Ridnik, N. Zami , e al.,Asymme ic loss o mul i-
label classi ica ion, 2021. a Xi : 2009.14119 [cs.CV].
[4] A. F. T. Ma ins and R. F. As udillo, F om so max o spa semax: A
spa se model o a en ion and mul i-label classi ica ion, 2016. a Xi :
1602.02068 [cs.CL].
[5] Y. Wang, D. He, F. Li, e al.,Mul i-label classi ica ion wi h label
g aph supe imposing, 2019. a Xi : 1911.09243 [cs.CV].
[6] Z.-M. Chen, X.-S. Wei, P. Wang, and Y. Guo, Mul i-label image
ecogni ion wi h g aph con olu ional ne wo ks, 2019. a Xi : 1904.
03582 [cs.CV].
[7] H. Tou on, M. Co d, M. Douze, F. Massa, A. Sablay olles, and H.
J´
egou, T aining da a-e icien image ans o me s dis illa ion h ough
a en ion, 2021. a Xi : 2012.12877 [cs.CV].
[8] W. Zhang, C. Liu, L. Zeng, B. Ooi, S. Tang, and Y. Zhuang, “Lea ning
in impe ec en i onmen : Mul i-label classi ica ion wi h long- ailed
dis ibu ion and pa ial labels,” in P oceedings o he IEEE/CVF
In e na ional Con e ence on Compu e Vision (ICCV), Oc . 2023,
pp. 1423–1432.
[9] A. Rad o d, J. W. Kim, C. Hallacy, e al.,Lea ning ans e able isual
models om na u al language supe ision, 2021. a Xi : 2103.00020
[cs.CV].
[10] J. Read, B. P ah inge , G. Holmes, and E. F ank, “Classi ie chains: A
e iew and pe spec i es,” Jou nal o A i icial In elligence Resea ch,
ol. 70, pp. 683–718, Feb. 2021, ISSN: 1076-9757. DOI: 10.1613/jai .
1.12376. [Online]. A ailable: h p://dx.doi.o g/10.1613/jai .1.12376.
[11] Y. Wei, W. Xia, M. Lin, e al., “Hcp: A lexible cnn amewo k
o mul i-label image classi ica ion,” IEEE T ansac ions on Pa e n
Analysis and Machine In elligence, ol. 38, no. 9, pp. 1901–1907,
Sep. 2016, ISSN: 1939-3539. DOI: 10 . 1109 / pami . 2015 . 2491929.
[Online]. A ailable: h p://dx.doi.o g/10.1109/TPAMI.2015.2491929.
[12] T. Chen, M. Xu, X. Hui, H. Wu, and L. Lin, Lea ning seman ic-
speci ic g aph ep esen a ion o mul i-label image ecogni ion, 2019.
a Xi : 1908.07325 [cs.CV].
[13] Y. Liu, L. Sheng, J. Shao, J. Yan, S. Xiang, and C. Pan, “Mul i-
label image classi ica ion ia knowledge dis illa ion om weakly-
supe ised de ec ion,” in P oceedings o he 26 h ACM in e na ional
con e ence on Mul imedia, ACM, Oc . 2018. DOI: 10.1145/3240508.
3240567. [Online]. A ailable: h p://dx.doi.o g/10.1145/3240508.
3240567.
[14] A. Vaswani, N. Shazee , N. Pa ma , e al.,A en ion is all you need,
2023. a Xi : 1706.03762 [cs.CL].
[15] J. De lin, M.-W. Chang, K. Lee, and K. Tou ano a, Be : P e- aining
o deep bidi ec ional ans o me s o language unde s anding, 2019.
a Xi : 1810.04805 [cs.CL].
[16] A. Rad o d, J. Wu, R. Child, D. Luan, D. Amodei, and I. Su ske e ,
“Language models a e unsupe ised mul i ask lea ne s,” 2019. [On-
line]. A ailable: h ps://api.seman icschola .o g/Co pusID:160025533.
[17] T. B. B own, B. Mann, N. Ryde , e al.,Language models a e ew-
sho lea ne s, 2020. a Xi : 2005.14165 [cs.CL].
[18] X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, De o mable
de : De o mable ans o me s o end- o-end objec de ec ion, 2021.
a Xi : 2010.04159 [cs.CV].
[19] X. Cheng, H. Lin, X. Wu, e al.,Ml : Mul i-label classi ica ion wi h
ans o me , 2021. a Xi : 2106.06195 [cs.CV].
[20] S. Liu, L. Zhang, X. Yang, H. Su, and J. Zhu, Que y2label: A
simple ans o me way o mul i-label classi ica ion, 2021. a Xi :
2107.10834 [cs.CV].
[21] J. Ye, J. He, X. Peng, W. Wu, and Y. Qiao, A en ion-d i en dynamic
g aph con olu ional ne wo k o mul i-label image ecogni ion, 2020.
a Xi : 2012.02994 [cs.CV].
[22] K. P oko ie and V. So aso , Combining me ic lea ning and a en-
ion heads o accu a e and e icien mul ilabel image classi ica ion,
2022. a Xi : 2209.06585 [cs.CV].
[23] F. Zhu, H. Li, W. Ouyang, N. Yu, and X. Wang, Lea ning spa ial
egula iza ion wi h image-le el supe isions o mul i-label image
classi ica ion, 2017. a Xi : 1702.05891 [cs.CV].
[24] Y. C¸ . Ak as¸ and J. G. Cas a˜
no, “Few-sho mul i-label mul i-class
classi ica ion o da k web image ca ego iza ion,” in 2024 12 h
In e na ional Symposium on Digi al Fo ensics and Secu i y (ISDFS),
2024, pp. 1–6. DOI: 10.1109/ISDFS60797.2024.10527297.
[25] X. Zhai, B. Mus a a, A. Kolesniko , and L. Beye , Sigmoid loss o
language image p e- aining, 2023. a Xi : 2303 .15343 [cs.CV].
[Online]. A ailable: h ps://a xi .o g/abs/2303.15343.
[26] T. Ridnik, G. Sha i , A. Ben-Cohen, E. Ben-Ba uch, and A. Noy,
Ml-decode : Scalable and e sa ile classi ica ion head, 2021. a Xi :
2111.12933 [cs.CV]. [Online]. A ailable: h ps://a xi .o g/abs/
2111.12933.