scieee Open visual document viewer

CLIP-based Few-Shot Multi-Label Classification Methods: A Comparative Study

Aktas, Yagmur; García, Jorge

Abstract

Categorizing darkweb image content is critical for identifying and averting potential threats. However, this remains a challenge due to the nature of the data, which includes multiple co-existing domains and intra-class variations. While many methods have been proposed to classify this image content, multi-label multi-class classification remains underexplored. The complexity of darkweb imagery, combined with the need for efficient classification systems, demands innovative approaches that can handle both the technical challenges and the sensitive nature of the content. In this paper, we present a comparative study of few-shot multi-label classification methods using the multimodal model CLIP. Our research addresses the growing need for robust classification systems that can effectively catego-rize diverse and complex image content while maintaining high accuracy and computational efficiency. We particularly focus on the challenges of handling multiple labels simultaneously and the scalability of these systems in real-world applications. We analyze and compare four different approaches: CLIP+Label Empower Adapter, CLIP Sigmoid, SIGLIP, and CLIP+ML-Decoder. Our study evaluates these methods based on their precision, recall, and ability to handle increasing class numbers efficiently. Finally, our research contributes to the field by providing detailed insights into the strengths and limitations of each method.

Full text

CLIP-based Few-Sho Mul i-Label Classi ica ion Me hods: A Compa a i e S udy Ya˘ gmu C¸ i˘ gdem Ak as¸ Vicom ech Founda ion Basque Resea ch and Technology Alliance (BRTA) Mikele egi 57, 20009 Donos ia-San Sebas i´ an (Spain) [email p o ec ed]g Jo ge Ga c´ ıa Cas a˜ no Vicom ech Founda ion Basque Resea ch and Technology Alliance (BRTA) Mikele egi 57, 20009 Donos ia-San Sebas i´ an (Spain) [email p o ec ed]g Abs ac —Ca ego izing da kweb image con en is c i ical o iden i ying and a e ing po en ial h ea s. Howe e , his emains a challenge due o he na u e o he da a, which includes mul iple co-exis ing domains and in a-class a ia ions. While many me hods ha e been p oposed o classi y his image con en , mul i-label mul i-class classi ica ion emains unde explo ed. The complexi y o da kweb image y, combined wi h he need o e icien classi ica ion sys ems, demands inno a i e app oaches ha can handle bo h he echnical challenges and he sensi i e na u e o he con en . In his pape , we p esen a compa a i e s udy o ew-sho mul i-label classi ica ion me hods using he mul imodal model CLIP. Ou esea ch add esses he g owing need o obus classi ica ion sys ems ha can e ec i ely ca ego- ize di e se and complex image con en while main aining high accu acy and compu a ional e iciency. We pa icula ly ocus on he challenges o handling mul iple labels simul aneously and he scalabili y o hese sys ems in eal-wo ld applica ions. We analyze and compa e ou di e en app oaches: CLIP+Label Empowe Adap e , CLIP Sigmoid, SIGLIP, and CLIP+ML-Decode . Ou s udy e alua es hese me hods based on hei p ecision, ecall, and abili y o handle inc easing class numbe s e icien ly. Finally, ou esea ch con ibu es o he ield by p o iding de ailed insigh s in o he s eng hs and limi a ions o each me hod. Index Te ms—Mul i-label image classi ica ion, Mul i-class im- age classi ica ion, CLIP, SIGLIP, ML-Decode I. INTRODUCTION The da k web, a specialized subse o he as global in e ne , is accessible only h ough dedica ed web b owse s. I is widely ecognized as a hub o illici ac i i ies due o i s co e cha ac e is ics: anonymi y, which p o ides use s wi h a le el o p i acy no ypically a ailable on he su ace web, and un aceabili y, which makes i ex emely challenging o ack he o igins and des ina ions o da a ans e s. As a esul , he da k web se es as a c i ical sou ce o h ea in elligence, o e ing insigh s in o cybe -a acks, s olen asse s, illegal ade, a ms a icking, con iden ial da a leaks, child exploi a ion, and o he c iminal ac i i ies. Au oma ically analyzing he isual con en o he da k web —such as images and ideos- is essen ial o e icien image ca ego iza ion, enabling he de ec ion and p e en ion o po en ial h ea s. Howe e , da k web image classi ica ion emains an open challenge due o he complex na u e o i s con en . Fig. 1: Mul i-label image examples om Da k web da ase Many images su e om low esolu ion, blu , and small objec sizes, o en con aining objec s ha blend in o he back- g ound due o colo simila i ies. Addi ionally, he coexis ence o mul iple o e lapping domains and in a-class a ia ions — whe e ins ances wi hin he same ca ego y exhibi signi ican isual di e ences — u he complica e classi ica ion asks, as illus a ed in Fig. 1. This complexi y highligh s he impo ance o mul i-label classi ica ion, whe e an image can be assigned mul iple ele- an labels a he han a single ca ego y. A ange o mul i-label classi ica ion echniques ha e been explo ed o ackle his p oblem. Some app oaches decom- pose mul i-label classi ica ion in o sepa a e single-label p ob- lems[1], [2], while o he s e ine loss unc ions o ac i a ion unc ions o enhance classi ica ion accu acy[3], [4]. Addi ion- ally, g aph-based ne wo ks u ilize lea ned label embeddings o loca e key disc imina i e ea u es [5], [6], while ans o me - based a chi ec u es ha e been adap ed o mul i-label classi i- ca ion asks [7]. While hese me hods show p omising esul s, hey ely on la ge, well-balanced da ase s—which a e di icul o cu a e in eal-wo ld applica ions. Expanding a da ase wi h a balanced class dis ibu ion is an exponen ially complex p ocess [8], making la ge-scale mul i-label lea ning in easible in many scena ios. To add ess his, ecen ad ances in ision-language models like CLIP [9] o e new possibili ies. By aligning isual and ex ual embeddings, CLIP enables ze o-sho and ew-sho lea ning in open- ocabula y se ings. Howe e , mos CLIP-979-8-3315-0993-4/25/31.002025IEEE based me hods a e ailo ed o single-label classi ica ion. In his s udy, we add ess he p oblem o ew-sho mul i- label image classi ica ion, ocusing on he da k web as a eal-wo ld, high-s akes applica ion domain. Ou hypo hesis is ha adap ing s a e-o - he-a ew-sho mul i-label classi ica ion echniques o CLIP can imp o e gene aliza ion and pe o - mance in low- esou ce, high- a iabili y en i onmen s. Ou con ibu ions can be summa ized as ollows: •We analyze ou s a e-o - he-a ew-sho mul i-label classi ica ion echniques and p o ide a po able pipeline applicable o any eal-wo ld da ase equi ing ew-sho mul i-label classi ica ion. •We e alua e hese app oaches by implemen ing CLIP- based me hodologies o da k web mul i-label image classi ica ion. •We compa e ou p e ious app oach wi h newe CLIP- based mul i-label me hodologies, p esen ing a anspa en compa a i e s udy which con ibu es o he ield by p o- iding de ailed insigh s in o he s eng hs and limi a ions o each me hod. II. RELATED WORK A a ie y o me hods ha e been p oposed o sol e he mul i- label classi ica ion p oblem, which can be b oadly ca ego ized in o (i) p oblem ans o ma ion me hods and (ii) algo i hm adap a ion echniques. Ea ly esea ch ocused on decomposing mul i-label classi i- ca ion in o independen bina y classi ica ion asks[1]. How- e e , his app oach ails o cap u e in e -label co ela ions and equi es aining sepa a e classi ie s o each ca e- go y—in oducing a signi ican compu a ional bu den. To ad- d ess his, some s udies p opose classi ie chaining[10], whe e each classi ie ’s ou pu se es as an inpu ea u e o sub- sequen classi ie s. While his s a egy imp o es in e -label dependency modeling, i s ill su e s om high compu a ional cos s due o he g owing numbe o classi ie s equi ed. Ano he ans o ma ion-based app oach, known as Label Powe se (LP) [2], con e s mul i-label p oblems in o mul i- class p oblems by ea ing each unique label combina ion as a dis inc class. Howe e , LP s uggles wi h he exponen ial g ow h o label combina ions, making i imp ac ical o la ge da ase s. Se e al me hods aim o imp o e mul i-label classi ica ion by op imizing he loss unc ion[3] o ac i a ion unc ion[4] o mi iga e class imbalance issues. In compu e ision, mul iple app oaches ha e been de el- oped o mul i-label image classi ica ion. Hypo heses-CNN- Pooling (HCP) [11] gene a es a la ge numbe o p oposals h ough objec de ec ion echniques, ea ing each p oposal as a single-label classi ica ion p oblem. Beyond CNN-based me hods, G aph Con olu ional Ne - wo ks (GCNs) ha e demons a ed high e icacy ac oss a i- ous ision asks, including mul i-label classi ica ion [5], [6], [12]. GCN-based me hods build classi ie s by modeling label ela ionships wi hin g aph ne wo ks. O he app oaches employ weakly supe ised lea ning o mul i-label classi ica ion [13], le e aging knowledge dis illa- ion echniques om objec de ec ion models. Mo e ecen ly, ans o me s[14]—o iginally de eloped o modeling long- ange dependencies in na u al language p o- cessing (NLP)[15]–[17]—ha e demons a ed s ong pe o - mance ac oss compu e ision asks, including image classi i- ca ion[7], [9] and objec de ec ion[18]. Se e al ans o me -based mul i-label classi ica ion ech- niques ha e eme ged: Mul i-label T ans o me (MLT) [19], which models pixel- wise a en ion o enhance ea u e ex ac ion. T ans o me - based que y lea ning [20], which uses decode que ies o p edic label exis ence. G aph-enhanced T ans o me s [21], which combine GCNs wi h a en ion mechanisms. Me ic Lea ning T ans o me s [22], which in eg a e me ic lea ning o assess label simila i ies. T ans o me -CNN Hyb ids [23], which gene a e a en ion maps pe label, cap u ing in a-class dependencies h ough con olu ional laye s. Despi e hese ad ancemen s, exis ing mul i-label me hods s ill equi e la ge da ase s o achie e high accu acy. Howe e , da a sca ci y is a p e alen challenge in eal-wo ld applica- ions, whe e collec ing and labeling ex ensi e da ase s is bo h esou ce-in ensi e and ime-consuming. La ge-scale open- ocabula y models ained on as da ase s, such as CLIP [9], o e a p omising al e na i e by le e aging ze o-sho single-label classi ica ion and ew-sho mul i-label classi ica ion. By adap ing hese models, we aim o explo e hei po en ial o enhancing mul i-label classi i- ca ion in da k web image y, whe e adi ional me hods ace signi ican cons ain s. III. CLIP-BASED MULTI-LABEL CLASSIFICATION METHODOLOGIES The e a e ou ecen mul i-label classi ica ion me hods sui - able o CLIP mul i-modal model adap a ion: (i) CLIP+Label Empowe (ou p e ious model), (ii) CLIP Sigmoid (iii) SIGLIP (i ) CLIP + ML-Decode A. CLIP+Label Empowe Label Empowe [2] is one o he app oaches ha con e he mul i-label classi ica ion p oblem in o a single-label classi i- ca ion, by mul i-labels o bina y ec o s by ho -encoding hem and assigning a unique label o each di e en bina y ec o . As demons a ed in [24], i has he highes pe o mance among he o he app oaches o sol ing he mul i-label classi ica ion ask by ans o ming he p oblem in o a single-label one. Fig. 2 shows he a chi ec u e o ou p e iously p oposed me hod CLIP+Label Empowe Adap e . Al hough ou p e iously p oposed model has he highes ecall, as shown in Table I, among e en hese ou s a e- o - he-a CLIP-adap ed mul i-label classi ica ion app oaches, no only be ween he adi ional ones, i p ese es a isk o accu acy d op wi hin a la ge numbe o classes due o i s exponen ially g owing ho -encoded labels. This possible isk is especially impo an o such a da ase like da kweb, which Fig. 2: CLIP based Mul i-Label Me hodologies is expec ed o g ow as in he sense o numbe o classes, since new ype o images and i les a e expec ed o occu . •Wo s Case: O(N!) (All class labels exis in all possible pe mu a ions in he da ase ) •Bes Case: O(N)(The da ase con ains only single labels) B. CLIP Sigmoid CLIP has he capabili y o classi y an ex emely huge num- be o di e en classes, hanks o i s open- ocabula y na u e. Ye we migh need o ine- une his s a e-o - he-a model o ou eal-wo ld da ase s like da kweb da ase . In a single-label pipeline, he p ope way o ine- uning such a model is o add a linea laye ollowing he image and ex embeddings, ha ing ou pu nodes as much as class numbe s in he da ase . In a single-label, mul i-class classi ica ion scena io, so max is commonly used o dis ibu e p obabili ies among di e en classes, ensu ing ha he sum o all class p obabili ies equals 1. P(yi|x) = ezi PN j=1 ezj (1) whe e zi ep esen s he logi s o class iand Nis he o al numbe o classes. Howe e , in mul i-label classi ica ion, whe e mul iple labels can be p esen simul aneously, so max is subop imal as i o ces mu ual exclusi i y among classes. Ins ead, sigmoid ac i a ion is mo e app op ia e, as i independen ly p edic s each label’s p obabili y: P(yi|x) = 1 1 + e−zi(2) This allows he model o assign a p obabili y close o 1.0 o each ele an label in an image, ensu ing ha all p esen classes a e p edic ed wi hou compe i ion om o he labels. Thus, ine- uning CLIP wi h a linea classi ica ion head using Sigmoid ac i a ion is ano he app oach. Al hough his me hod does no ha e any isk o exponen ially g owing wi h he numbe o class numbe s, ou expe imen al esul s show i s poo accu acy in compa ison wi h ou p e iously p oposed me hod. C. SIGLIP While SIGLIP [25] is a s a e-o - he-a , mul i-modal model, ha ing a e y simila a chi ec u e o CLIP, i s main di e ence is using Sigmoid, ins ead o So max while p e- aining. This ac p o ides us wi h he possibili y o see i he poo pe o - mance o CLIP + Sigmoid ine- uning me hodology occu ed due o using a di e en ac i a ion unc ion in he ine- uning han p e- aining. Fig. 2 shows he e y simila a chi ec u e o SIGLIP. D. CLIP+ML-Decode ML-Decode [26] is a s a e-o - he-a decode implemen- a ion aiming o o e come he exponen ially g owing inpu size in classi ica ion pipelines buil wi h de aul ans o me decode s (Fig. 3) due o he sel -a en ion laye . The me hod p oposes o emo e he sel -a en ion laye and claims i s a ec less on he classi ica ion accu acy. I is also claimed o use ”g oup que ies”, ins ead o ixed que y embeddings pe class as in adi ional ans o me -decode s. Being he numbe o g oups a new hype pa ame e in his a chi ec u e e e s o he amoun o inpu que y embeddings he decode will ha e, independen ly om he numbe o classes in he da ase . Being K is he numbe o g oup que ies, he decode lea ns o c ea e K que ies e e encing N numbe o classes in a da ase , ins ead o ha ing N ixed que ies o N classes as in de aul ans o me decode s. CLIP+ML-Decode app oach uses bo h classi ica ion loss coming om he inal p edic ion logi s and an alignmen loss coming om he simila i y o ex and image embeddings o ob ain a inal loss. The ixed ex embeddings a e used only o his pu pose, whe eas non- ixed g oup que ies a e lea nable embeddings and hey a e upda ed du ing he aining. Fig. 3: T adi ional s ML Decode s The e o e his app oach enhances scalabili y, wi h wo key op imiza ions: (1) Sel -a en ion emo al, which educes he compu a ional complexi y om O(N2) o O(N), and (2) G oup decoding, which u he op imizes in e ence by shi ing he complexi y om O(N) o O(K), whe e KK ep esen s he numbe o meaning ul g oups ins ead o p ocessing all classes independen ly. Sel -a en ion emo al: O(N2)→O(N) G oup decoding: O(N)→O(K) Fig. 2 shows he a chi ec u e o CLIP + ML-Decode . IV. EXPERIMENTAL STUDY A. Da k web da ase The Da k web da ase , sou ced om CFLW’s Da k Web Moni o 1, includes images collec ed om da k web domains abou a ious c ime ca ego ies like inancial c ime and o gani- za ions, d ugs and na co ics, weapons as well as hei endo s as indi idual classes. Simila o many eal-wo ld da ase s, he da k web da ase con ains images wi h mul iple labels, meaning an image can belong o mo e han one class. I also includes images wi h a single label. Fig. 1 shows some image samples along wi h hei g ound u h labels om a ious classes. 1h ps://c lw.com/dwm/ Fig. 4: Da a dis ibu ion o he da k web da ase o bo h ain and es se s. The da ase con ains 46 classes om a ious ca ego ies like D ugs, Na co ics, Weapons, and Financial O ganiza ions which do no ha e a s ong co ela ion be ween hem bu i includes subca ego ies ha ing a s ong ela ionship and he possibili y o being exis ing in he same image. Being he cu en da kweb da ase is an expe imen al one, many mo e classes a e expec ed o join and hus, he scalabili y pe o mance is impo an han a egula Fig. 4 shows he da a dis ibu ion o he da k web da ase o bo h ain and es se s. The da kweb da ase p esen s an imbalanced da a p oblem. Some dominan classes ha e many samples, whe eas some o he classes su e om a lack o da a. Some classes like d ugs and na co ics, coming along wi h any ype o d ug ype o any indi idual endo ha has o be ca ego ized, become a e y dominan class by collec ing only a ew samples o i s subg oups. Fig. 1 shows 3 image samples om he da k web da ase , 1 belonging o a inancial o ganiza ion ca ego y, 2 belonging o he d ugs and na co ics ca ego y con aining pic u es o weed and he endo names. While endo names may a y, each image ha ing a endo name in o ma ion also con ains weed class, doubling he amoun o ”weed” class in compa ison o class names e e ing o he indi idual endo names. Which is obus example showing he eason o s ong class imbalance p oblem in da kweb da ase . Las ly, ano he challenge is ha some classes, especially he endo names like ”deep shop”, o ” anda al”, ha e ex ual in o ma ion a he han isual ea u es which b ings he need o conside ing he combina ion o ision and language da a. B. Implemen a ion De ails The ine- uning s a egy a ies o each o he me hods we examine: CLIP + Label Empowe was ine uned by eezing he CLIP p e- ained model, and upda ing he weigh s o only o be ween 5 o 10 epochs, using Adam op imize wi h de aul lea ning a e 1×10−3. So max inal ac i a ion unc ion and C oss En opy Loss. CLIP Sigmoid model was ine uned by eezing he CLIP p e- ained model, and upda ing he weigh s o only he ad- di ional linea classi ica ion head, o 50 epochs un il con e - gence, using he same op imize and lea ning a e as p e ious app oach, wi h a Bina y C oss En opy Loss o e he logi s, as a de aul app oach o sigmoid classi ica ion. SIGLIP model was ine uned wi hou eezing any pa o he a chi ec u e, since ou implemen a ion any addi ional pa and he na u e o he a chi ec u e is al eady compa ible o mul i-label classi ica ion. I is ine uned 80 epochs un il con e gence, using Adam op imize wi h a lea ning a e o 5×10−5. CLIP+ML-Decode me hod was ine- uned by eezing he CLIP model and only ocusing on upda ing he weigh s o Decode . A g id sea ch o e all he possible hype -pa ame e s o he decode block was made and he bes esul s was ob ained wi h: mul i-head a en ion head amoun 4, d opou in eed o wa d ne wo k 0.5, numbe o g oups K 8, numbe o laye s (decode block amoun ) 2. Simila ly o he p e ious me hods, Adam op imize is used, wi h a lea ning a e o 1×10−4. C. E alua ion Me ics We e alua e model pe o mance using s anda d classi ica- ion me ics: p ecision, ecall, and 1-sco e. We u he mo e analyze he scalabili y o he model. •P ecision: Measu es he p opo ion o co ec ly p edic ed labels among all p edic ed labels. 3 •Recall: Measu es he p opo ion o co ec ly p edic ed labels among all ue labels.4 •F1-Sco e: The ha monic mean o p ecision and ecall. 5 •Scalabili y: Assesses how well he model adap s o inc easing class sizes. P ecision =T P T P +F P (3) Recall =T P T P +F N (4) F1sco e =2×P ecision ×Recall P ecision +Recall (5) The ue posi i es, alse posi i es, alse nega i es a e cal- cula ed simila ly o single-label classi ica ion. Fig. 5 shows an image sample ha ing wo g ound u h labels: VISA and Bankno es. In case he p edic ed classes a e ”VISA and Paypal” class, a e ex ac ing he single labels as ”VISA” and ”Paypal”, his p edic ion would con ibu e as a ue posi i e o ”VISA” class, alse posi i e o ”Paypal” class and a alse nega i e o ”Bankno es” classes. Fig. 5: An example mul i-labeled image wi h 2 classes. V. RESULTS AND DISCUSSION Table I compa es he pe o mance o he ou me hods in e ms o p ecision, ecall, 1-sco e and scalabili y and summa izes ou indings. CLIP+Label Empwoe me hod has a low scalabili y since i has O(N!) as wo s case scena io, CLIP Sigmoid and SIGLIP has a mode a e scalabili y since hey p o ide an imp o emen , bu nei he b ing any g ow h, meaning he wo s and bes case scena io is he same and O(N). While CLIP+ML-Decode imp o es he scalabili y om O(N) o O(K), being N is he numbe o classes and K is he numbe o g oups de ined by he end-use . Table II compa es he wo models gi ing he bes esul s o da k web da ase , on he MS-COCO mul i-label da ase , causing he CLIP+LE me hod o ha e a huge d op o accu acy, due o expanding 80 base classes o 234,581 ho -encoded classes. These esul s show he a o emen ioned po en ial isk o CLIP+LE me hod on he da ase s ha ing high ela ion be ween he classes, causing he exis ence o a ious pe mu- a ions o mul i-label ec o s among he da ase . E en hough he esul s migh be imp essi e o such amoun o classes, i is isible ha he Label Empowe me hod has a huge impac on CLIP mul i-modal model’s classi ica ion capaci y by exploding he class numbe s unnecessa ily. TABLE I COMPARISON OF FEW-SHOT MULTI-LABEL CLASSIFICATION METHODS. LE: LABEL EMPOWER, MLD: ML-DECODER Me hod P ecision Recall F1 sco e Scalabili y CLIP+LE 0.944 0.931 0.937 Low CLIP Sigmoid 0.493 0.276 0.351 Mode a e SIGLIP 0.880 0.735 0.801 Mode a e CLIP+MLD 0.958 0.916 0.936 High Ou esul s show ha he CLIP+Label Empowe Adap e achie es he bes ecall, making i s ill a s ong candida e o ecall-sensi i e applica ions buil o da ase s no ha ing huge amoun o classes, due o i s low scalabili y and he po en ial isk o p o iding poo e esul s wi h a huge numbe o classes. On he o he hand, he esul s show he medium- le el scalabili y models (CLIP+Sigmoid, SIGLIP), ha a e no imp o ing no wo sen he a chi ec u e acco ding o he numbe o classes, ha e poo accu acy in compa ison wi h CLIP + adap e solu ions. Howe e , CLIP+ML-Decode pe o ms well in bo h p ecision and ecall, while also add essing he scalabili y issue e ec i ely and i should be a de ini e choice o da ase ha ing huge numbe o classes, whe e o no ha e alse posi i e p edic ions is mo e impo an han no missing any ue posi i e p edic ion, ega ding i ’s p ecision bea ing he CLIP+Label Empowe me hod while p o iding lowe ecall. TABLE II COMPARISON OF MULTI-LABEL MS-COCO MAP SCORE. LE: LABEL EMPOWER AND MLD: MULTI-LABEL DECODER Me hod P ecision Recall CLIP+LE 0.592 0.555 CLIP+MLD 0.839 0.809 VI. CONCLUSION AND FUTURE WORKS In his s udy, we analyzed ou ew-sho mul i-label classi- ica ion me hods based on CLIP. Ou esul s demons a e ha CLIP+Label Empowe Adap e excels in he ecall, whe eas CLIP+ML-Decode p o ides a mo e scalable solu ion by mi iga ing he exponen ial g ow h p oblem in inpu ea u es, p o iding also a obus pe o mance on accu acy me ics. Fu u e wo k will explo e he e icien deploymen pipeline o CLIP+ML-Decode app oach o eal-wo ld mul i-label classi ica ion asks. ACKNOWLEDGEMENTS The wo k desc ibed in his pape is pe - o med in he H2020 p ojec STARLIGHT (”Sus ainable Au onomy and Resilience o LEAs using AI agains High P io i y Th ea s”). This p ojec has ecei ed unding om he Eu opean Union’s Ho izon 2020 esea ch and inno a ion p og am unde g an ag eemen No 101021797. REFERENCES [1] M.-L. Zhang, Y.-K. Li, X.-Y. Liu, and X. Geng, “Bina y ele ance o mul i-label lea ning: An o e iew,” F on ie s o Compu e Science, ol. 12, no. 2, pp. 191–202, Ma . 2018, ISSN: 2095-2236. DOI: 10. 1007/s11704-017-7031-7. [Online]. A ailable: h p://dx.doi.o g/10. 1007/s11704-017-7031-7. [2] G. Tsoumakas, A. Dimou, E. Spy omi os-Xiou is, V. Meza is, I. Kompa sia is, and I. Vlaha as, “Co ela ion-based p uning o s acked bina y ele ance models o mul i-label lea ning,” Jan. 2009, pp. 101– 116. [3] E. Ben-Ba uch, T. Ridnik, N. Zami , e al.,Asymme ic loss o mul i- label classi ica ion, 2021. a Xi : 2009.14119 [cs.CV]. [4] A. F. T. Ma ins and R. F. As udillo, F om so max o spa semax: A spa se model o a en ion and mul i-label classi ica ion, 2016. a Xi : 1602.02068 [cs.CL]. [5] Y. Wang, D. He, F. Li, e al.,Mul i-label classi ica ion wi h label g aph supe imposing, 2019. a Xi : 1911.09243 [cs.CV]. [6] Z.-M. Chen, X.-S. Wei, P. Wang, and Y. Guo, Mul i-label image ecogni ion wi h g aph con olu ional ne wo ks, 2019. a Xi : 1904. 03582 [cs.CV]. [7] H. Tou on, M. Co d, M. Douze, F. Massa, A. Sablay olles, and H. J´ egou, T aining da a-e icien image ans o me s dis illa ion h ough a en ion, 2021. a Xi : 2012.12877 [cs.CV]. [8] W. Zhang, C. Liu, L. Zeng, B. Ooi, S. Tang, and Y. Zhuang, “Lea ning in impe ec en i onmen : Mul i-label classi ica ion wi h long- ailed dis ibu ion and pa ial labels,” in P oceedings o he IEEE/CVF In e na ional Con e ence on Compu e Vision (ICCV), Oc . 2023, pp. 1423–1432. [9] A. Rad o d, J. W. Kim, C. Hallacy, e al.,Lea ning ans e able isual models om na u al language supe ision, 2021. a Xi : 2103.00020 [cs.CV]. [10] J. Read, B. P ah inge , G. Holmes, and E. F ank, “Classi ie chains: A e iew and pe spec i es,” Jou nal o A i icial In elligence Resea ch, ol. 70, pp. 683–718, Feb. 2021, ISSN: 1076-9757. DOI: 10.1613/jai . 1.12376. [Online]. A ailable: h p://dx.doi.o g/10.1613/jai .1.12376. [11] Y. Wei, W. Xia, M. Lin, e al., “Hcp: A lexible cnn amewo k o mul i-label image classi ica ion,” IEEE T ansac ions on Pa e n Analysis and Machine In elligence, ol. 38, no. 9, pp. 1901–1907, Sep. 2016, ISSN: 1939-3539. DOI: 10 . 1109 / pami . 2015 . 2491929. [Online]. A ailable: h p://dx.doi.o g/10.1109/TPAMI.2015.2491929. [12] T. Chen, M. Xu, X. Hui, H. Wu, and L. Lin, Lea ning seman ic- speci ic g aph ep esen a ion o mul i-label image ecogni ion, 2019. a Xi : 1908.07325 [cs.CV]. [13] Y. Liu, L. Sheng, J. Shao, J. Yan, S. Xiang, and C. Pan, “Mul i- label image classi ica ion ia knowledge dis illa ion om weakly- supe ised de ec ion,” in P oceedings o he 26 h ACM in e na ional con e ence on Mul imedia, ACM, Oc . 2018. DOI: 10.1145/3240508. 3240567. [Online]. A ailable: h p://dx.doi.o g/10.1145/3240508. 3240567. [14] A. Vaswani, N. Shazee , N. Pa ma , e al.,A en ion is all you need, 2023. a Xi : 1706.03762 [cs.CL]. [15] J. De lin, M.-W. Chang, K. Lee, and K. Tou ano a, Be : P e- aining o deep bidi ec ional ans o me s o language unde s anding, 2019. a Xi : 1810.04805 [cs.CL]. [16] A. Rad o d, J. Wu, R. Child, D. Luan, D. Amodei, and I. Su ske e , “Language models a e unsupe ised mul i ask lea ne s,” 2019. [On- line]. A ailable: h ps://api.seman icschola .o g/Co pusID:160025533. [17] T. B. B own, B. Mann, N. Ryde , e al.,Language models a e ew- sho lea ne s, 2020. a Xi : 2005.14165 [cs.CL]. [18] X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, De o mable de : De o mable ans o me s o end- o-end objec de ec ion, 2021. a Xi : 2010.04159 [cs.CV]. [19] X. Cheng, H. Lin, X. Wu, e al.,Ml : Mul i-label classi ica ion wi h ans o me , 2021. a Xi : 2106.06195 [cs.CV]. [20] S. Liu, L. Zhang, X. Yang, H. Su, and J. Zhu, Que y2label: A simple ans o me way o mul i-label classi ica ion, 2021. a Xi : 2107.10834 [cs.CV]. [21] J. Ye, J. He, X. Peng, W. Wu, and Y. Qiao, A en ion-d i en dynamic g aph con olu ional ne wo k o mul i-label image ecogni ion, 2020. a Xi : 2012.02994 [cs.CV]. [22] K. P oko ie and V. So aso , Combining me ic lea ning and a en- ion heads o accu a e and e icien mul ilabel image classi ica ion, 2022. a Xi : 2209.06585 [cs.CV]. [23] F. Zhu, H. Li, W. Ouyang, N. Yu, and X. Wang, Lea ning spa ial egula iza ion wi h image-le el supe isions o mul i-label image classi ica ion, 2017. a Xi : 1702.05891 [cs.CV]. [24] Y. C¸ . Ak as¸ and J. G. Cas a˜ no, “Few-sho mul i-label mul i-class classi ica ion o da k web image ca ego iza ion,” in 2024 12 h In e na ional Symposium on Digi al Fo ensics and Secu i y (ISDFS), 2024, pp. 1–6. DOI: 10.1109/ISDFS60797.2024.10527297. [25] X. Zhai, B. Mus a a, A. Kolesniko , and L. Beye , Sigmoid loss o language image p e- aining, 2023. a Xi : 2303 .15343 [cs.CV]. [Online]. A ailable: h ps://a xi .o g/abs/2303.15343. [26] T. Ridnik, G. Sha i , A. Ben-Cohen, E. Ben-Ba uch, and A. Noy, Ml-decode : Scalable and e sa ile classi ica ion head, 2021. a Xi : 2111.12933 [cs.CV]. [Online]. A ailable: h ps://a xi .o g/abs/ 2111.12933.