FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO
P oducing Decisions and Explana ions:
A Join App oach Towa ds Explainable
CNNs
Isabel C is ina Rio-To o de Oli ei a
MASTER’SDEGREE IN ELECTRICAL AND COMPUTERS ENGINEERING
Supe iso : P o . Dou o Luís Filipe Pin o de Almeida Teixei a
Co-Supe iso : Dou o Kelwin Alexande Fe nandes Co eia
Oc obe 1, 2019
c
Isabel Rio-To o, 2019
Resumo
As Redes Neu onais Con olucionais, bem como ou os modelos de Deep Lea ning, êm con-
seguido a ingi esul ados no á eis em a e as como classi icação e de eção de obje os. Con udo,
es es modelos pe manecem, em g ande pa e, caixas-neg as. Com o uso gene alizado de ais edes
em cená ios eais e com a c escen e exigência do di ei o à explicação, já incluída nas no as polí i-
cas do Regulamen o Ge al de P o eção de Dados, especialmen e em á eas al amen e eguladas
como medicina e jus iça, ge a apenas decisões o nou-se insu icien e. Pa a se possí el ope a
nes e no o enquad amen o legal, os modelos de Machine Lea ning êm de se explicá eis, is o
é, comp eensí eis pelos se es humanos, o que implica se em capazes de ap esen a as azões po
de ás das suas decisões.
Os dados são o ing edien e p incipal em odos os pipelines de Deep Lea ning. No en an o, é
inc i elmen e di ícil encon a conjun os de dados ano ados e é mui o ca o ecolhe e ano a um.
Além disso, a explicabilidade de um modelo ou de uma ins ância especí ica depende do u ilizado
e do con ex o de ope ação. Como al, é impe a i o que no as soluções sejam capazes de ge a
explicações sem p ecisa de dados ano ados.
Enquan o a maio pa e da li e a u a se oca em mé odos pós-modelo, es e abalho consis e
numa no a a qui e u a in a-modelo baseada numa ede neu onal con olucional, compos a po um
explicado e um classi icado . Es a a qui e u a ge a não só uma class label, mas ambém uma
explicação isual de al decisão, sem a necessidade adicional de dados ano ados pa a eina o
explicado . O modelo é einado end- o-end, endo o classi icado como en ada não só as ima-
gens do conjun o de dados, mas ambém a explicação esul an e do explicado , pe mi indo, assim,
que o classi icado se oque apenas nas á eas ele an es de al explicação. Um p ocesso de eino
decompos o em ês ases, é, ambém, p opos o, bem como unções de pe da especí icas, que eg-
ula izam as explicações e p omo em p op iedades desejá eis, como espa sidade ou con iguidade
espacial. Duas abo dagens al e na i as são desen ol idas, uma não supe isionada e ou a aca-
men e supe isionada, em que ambas pe mi em ao explicado p oduzi explicações sem p ecisa
de ano ações adicionais dos dados.
Pa a es a a a qui e u a p opos a, oi desen ol ida uma e amen a de ge ação de dados sin-
é icos, que pe mi e ano ação au omá ica e a c iação de imagens de ácil comp eensão que não
exigem o conhecimen o de um especialis a pa a as explica , acele ando, assim, o p ocesso de
a aliação humana quali a i a das explicações p oduzidas.
A a qui e u a oi, ambém, alidada num conjun o de dados médicos eal e os esul ados ob i-
dos mos am que, de ac o, es a é capaz de p oduzi explicações isuais pa a as decisões do clas-
si icado sem supe isão. Es es ambém comp o am que es a pode se u ilizada com qualque
classi icado , desde que sejam ei as as conexões necessá ias ao explicado . Além disso, es e
mé odo cons i ui um melho amen o ace a mé odos do es ado da a e, especialmen e no que se
e e e à explicação de ins âncias nega i as em casos sin é icos. Po im, as explicações p oduzidas
num conjun o de dados médico de canc o do colo do ú e o encon am-se alinhadas com conheci-
das ca a e ís icas isuais associadas a casos cance osos, como mudanças na mo ologia, co e
con o nos dos ecidos do colo do ú e o, que são, assim, indica i as de á eas lesionadas.
i
ii
Abs ac
Con olu ional Neu al Ne wo ks (CNNs), as well as o he Deep Lea ning (DL) models, ha e
shown ema kable pe o mance on asks like classi ica ion and objec de ec ion. Howe e , hese
models la gely emain black-boxes. Wi h he widesp ead use o such ne wo ks in eal-wo ld sce-
na ios and wi h he g owing demand o he igh o explana ion al eady included in he new Gene al
Da a P o ec ion Regula ion (GDPR) policies, especially in highly- egula ed a eas like medicine
and c iminal jus ice, gene a ing accu a e p edic ions is no longe enough. In o de o ope a e
wi hin his new legal amewo k, Machine Lea ning (ML) models ha e o be explainable, i.e.,
unde s andable o humans, which en ails being able o p esen he easons behind hei decisions.
Da a is he co e ing edien in e e y DL pipeline. Howe e , i is inc edibly di icul o come
by labelled da ase s and i is highly expensi e o collec and anno a e one. Mo eo e , he explain-
abili y o a model o a pa icula ins ance is use - and domain-dependen . As such, i is impe a i e
ha new solu ions a e able o gene a e explana ions wi hou needing labelled da a.
While mos o he li e a u e on explainable ML ocuses on pos -model me hods, his wo k
comp ises a no el in-model CNN a chi ec u e, composed by an explaine and a classi ie . This
a chi ec u e ou pu s no only a class label, bu also a isual explana ion o such decision, wi hou
he need o addi ional labelled da a o ain he explaine . The model is ained end- o-end, wi h
he classi ie aking as inpu no only images om he da ase bu also he explaine ’s esul ing
explana ion, hus allowing o he classi ie o ocus on he ele an a eas o such explana ion.
We also p opose a h ee phase aining p ocess wi h cus om loss unc ions ha egula ise he
p oduced explana ions and encou age desi ed p ope ies, such as spa si y and spa ial con igui y.
Two al e na i e app oaches a e p oposed, an unsupe ised and a weakly supe ised app oach, bo h
allowing he explaine o p oduce explana ions wi hou he need o addi ional anno a ions o he
da a.
To es he p oposed a chi ec u e, a syn he ic da a gene a ion ool was de eloped, ha allows
o au oma ic anno a ion and c ea ion o easy- o-unde s and images ha do no equi e he knowl-
edge o an expe o be explained, hus expedi ing he human quali a i e e alua ion p ocess o he
p oduced explana ions.
Ou join app oach was also alida ed on a eal medical da ase and he ob ained esul s show
ha , in ac , i is able o p oduce isual explana ions o he ne wo k’s decisions wi hou supe i-
sion. They also show ha his a chi ec u e can be employed wi h any classi ie , p o ided ha he
necessa y connec ions o he explaine a e made. Fu he mo e, his app oach imp o ed on s a e o
he a me hods, especially when ying o explain nega i e ins ances on syn he ic cases. Finally,
he p oduced explana ions on he medical ce ical cance da ase we e aligned wi h known isual
ea u es associa ed wi h cance ous cases, such as mo phology, colou and con ou changes o he
ce ix issue, and hus a e indica i e o inju ed a eas.
iii
i
Ag adecimen os
Aos meus o ien ado es, P o . Dou o Luís Teixei a e Dou o Kelwin Fe nandes, po odo o
apoio, o ien ação e opo unidades que me p opo ciona am ao longo des e pe cu so.
A odos os meus amigos que semp e me acompanha am e i e am paciência pa a me ou i .
Aos amigos de semp e, Ma iana, Nuno, Sa a e Daniel, ob igada pelo apoio. Ao Nuno, à Ri a e ao
Ped o, mui o ob igada e desculpem as secas gigan es que os dei a ala sob e es e abalho.
Aos meus a ós, que embo a já não isicamen e comigo, con inuam a se uma eno me on e de
inspi ação, o gulho e ajuda.
Aos meus pais po se em um excelen e exemplo de p o issionalismo e humildade, pelo apoio
incondicional e inabalá el paciência que i e am, não só du an e a elabo ação des e abalho, mas
ao longo des es á duos 5 anos: o meu mais since o ag adecimen o.
Po im, à Minie.
Isabel Rio-To o
i
“Be o e he ision, comes he ques ion.”
Inspi oBo 1- I am an a i icial in elligence dedica ed o gene a ing unlimi ed amoun s o unique
inspi a ional quo es o endless en ichmen o poin less human exis ence.
1h ps://inspi obo .me
ii
xi LIST OF TABLES
Abb e ia ions and Symbols
AI A i icial In elligence
ALE Accumula ed Local E ec s
ANN A i icial Neu al Ne wo k
AUC A ea Unde he Cu e
CAV Concep Ac i a ion Vec o
CIN Ce ical In aepi helial Neoplasia
CNN Con olu ional Neu al Ne wo k
COMPAS Co ec ional O ende Managemen P o iling o Al e na i e Sanc ions
CV Compu e Vision
DARPA De ense Ad anced Resea ch P ojec s Agency
DL Deep Lea ning
DNN Deep Neu al Ne wo k
DTD Deep Taylo Decomposi ion
EU Eu opean Union
FDA Food and D ug Adminis a ion
FN False Nega i e
FP False Posi i e
GAM Gene alised Addi i e Model
GD G adien Descen
GDPR Gene al Da a P o ec ion Regula ion
GLM Gene alised Linea Model
GPU G aphics P ocessing Uni
IAPR In e na ional Associa ion o Pa e n Recogni ion
IbPRIA Ibe ian Con e ence on Pa e n Recogni ion and Image Analysis
ICE Indi idual Condi ional Expec a ion
ILSVRC ImageNe La ge Scale Visual Recogni ion Challenge
K-NN K-Nea es Neighbou s
LIME Local In e p e able Model-agnos ic Explana ions
LRP Laye -Wise Rele ance P opaga ion
LTU Linea Th eshold Uni
ML Machine Lea ning
MLP Mul i-Laye Pe cep on
MMD Maximum Mean Disc epancy
MSE Mean Squa ed E o
NCI Na ional Cance Ins i u e
NIH Na ional Ins i u e o Heal h
PDP Pa ial Dependence Plo
ReLU Rec i ied Linea Uni
x
x i ABBREVIATIONS AND SYMBOLS
RGB Red, G een and Blue (addi i e colou model)
ROC Recei e Ope a ing Cha ac e is ic
RQ Resea ch Ques ion
SGD S ochas ic G adien Descen
SHAP SHapley Addi i e exPlana ions
SVM Suppo Vec o Machine
TCAV Tes ing wi h Concep Ac i a ion Vec o s
TN T ue Nega i e
TP T ue Posi i e
US Uni ed S a es
VGG Visual Geome y G oup
XAI Explainable A i icial In elligence
Chap e 1
In oduc ion
1.1 Con ex and Mo i a ion
The eme gence o Deep Lea ning (DL) changed he Machine Lea ning (ML) pa adigm in e-
cen yea s, especially in Compu e Vision (CV). Due o hei signi ican pe o mance imp o emen
in asks like classi ica ion, hese models became dominan . They ha e since been applied o ackle
di e en CV p oblems - e.g. de ec ion, segmen a ion, dep h es ima ion - and also o many o he
domains, such as speech ecogni ion, ex analysis and aud de ec ion.
Despi e his o e whelming ubiqui ousness, DL models and in pa icula Con olu ional Neu al
Ne wo ks (CNNs), a e o he mos pa conside ed black-box models, o which i is ha d o un-
de s and he easons o he labels ha hey gene a e. These models pe o m millions o non-linea
ope a ions, ende ing i impossible o humans o ollow his eno mous amoun o compu a ions.
Fu he mo e, con a ily o wha happened wi h adi ional ML me hods, such as Decision T ees o
Suppo Vec o Machines (SVMs), he whole p ocess is done end- o-end, om aw inpu o ou -
pu , implici ly ans o ming he aw inpu in o he mos con enien ea u e space o he model’s
calcula ions, elimina ing he need o ea u e enginee ing. This au oma ic ea u e ex ac ion, al-
hough highly ad an ageous, adds an ex a laye o complexi y and opaqueness o hese models. In
ac , he e is an impo an ade-o be ween p edic i e pe o mance and opaqueness, as depic ed
in igu e 1.1.
The e o e, unde s anding he decisions o DL models in gene al, and CNNs in pa icula , is o
he u mos impo ance o allow he deploymen o obus , anspa en and us wo hy DL-based
sys ems in eal-wo ld scena ios. This is pa icula ly ele an in c i ical high-s akes a eas, whe e
he consequences o a w ong decision could g ea ly a ec i s use s, o e en lead o he loss o
human li es. Examples o such highly egula ed a eas include medicine, c iminal jus ice, inance
and au onomous d i ing.
Fo hese easons, he e is an inc easing in e es in inding solu ions o ackle, o , a leas ,
mi iga e his limi a ion, and ha can comply wi h ecen legisla ion in oduced in he Eu opean
Union (EU) abou he igh o algo i hmic explana ion. In ac , hese challenges a e nowadays
agg ega ed unde wha now is known as Explainable A i icial In elligence (XAI). The e m i s
1
2In oduc ion
su aced in 2004 in he wo k o Van e al. [1] as a desc ip ion o hei AI aining sys em de eloped
o he Uni ed S a es (US) A my, capable o p oducing explana ions and answe ing ques ions
abou he pla oon’s ammuni ion s a us, o example.
Howe e , his is s ill a ela i ely new esea ch a ea, ha lacks uni ied de ini ions o co e con-
cep s, such as in e p e abili y and explainabili y, as well as a uni e sally accep ed axonomy and
quan i a i e e alua ion amewo ks.
Finally, on a sligh ly di e en no e, da a is he co e ing edien in e e y DL pipeline. In ac ,
DL models a e as good as he da a wi h which hey a e ained. Howe e , i is inc edibly di icul
o come by labelled da ase s and i is highly expensi e o collec and anno a e one, ei he ime-
o labou -wise. Mo eo e , p oducing meaning ul explana ions is an use - and domain-dependen
ask. As such, i is impe a i e ha new solu ions a e able o gene a e explana ions wi hou needing
addi ionally labelled da a.
Figu e 1.1: Lea ning echniques and he Accu acy-Explainabili y ade-o . Ex ac ed om [2].
1.2 P oblem Desc ip ion and Objec i es
Conside ing his g owing impo ance in designing sys ems capable o p oducing explana ions,
on a mac o le el, his wo k’s objec i e is o design and implemen a solu ion o y and ackle he
opaqueness p oblem inhe en o CNN a chi ec u es.
The e o e, he objec i es and echnical challenges o his wo k can be o mula ed as a se ies
o esea ch ques ions (RQ) as ollows:
•RQ1 - Would i be possible o gene a e isual explana ions wi hou supe ision, elimina ing
he need o addi ional anno a ions o he da a?
•RQ2 - Would i be possible o design an in-model explainabili y me hod capable o gene -
a ing isual explana ions as an implici pa o he aining p ocess?
•RQ3 - Would i be possible o ul il he a o emen ioned aspec s wi h an app oach capable o
being applied o di e en CNN a chi ec u es?
1.3 Con ibu ions and Publica ions 3
•RQ4 - Would i be possible o alida e his a chi ec u e no only on syn he ic da a, bu also
on a eal medical applica ion?
Mo e speci ically, his wo k con empla es he de elopmen and s udy o an a chi ec u e com-
posed by an explaine and a classi ie , capable o gene a ing no only class p edic ions, bu also
explana ions o such decisions, in an unsupe ised ashion. This wo k also in ol es he de el-
opmen o a syn he ic da a gene a ion ool wi h au oma ic anno a ion, o speed up he e alua ion
p ocess, by elimina ing he need o expe knowledge in p elimina y es scena ios.
1.3 Con ibu ions and Publica ions
The main con ibu ions o his disse a ion a e:
•A e sa ile and highly cus omisable syn he ic da a gene a ion ool wi h au oma ic anno a-
ion.
•A comp ehensi e and c i ical su ey on de ini ions, axonomies and me hods ound in he
e iewed li e a u e, as well as he denomina ion o wo p e iously unnamed axonomies
(Loca i e and Me hodology c i e ia), acco ding o hei scope. Also, an in-dep h s udy o
he di e ences and o e lap be ween he e iewed axonomies, and a classi ica ion o he
su eyed me hods acco ding o he di e en axonomies.
•A no el in-model join app oach ha p oduces isual explana ions o he decisions o di -
e en CNN-based classi ie s, along wi h a cus om aining p ocedu e and loss unc ions,
and i s alida ion on a eal wo ld medical scena io.
•Cus om loss unc ions (unsupe ised and weakly supe ised), ha allow gene a ing expla-
na ions wi hou he need o supe ision and no u he anno a ions o he al eady exis ing
da a.
F om he de elopmen o his wo k o igina ed he ollowing scien i ic pape :
•Isabel Rio-To o, Kelwin Fe nandes, and Luís F. Teixei a, ”Towa ds a Join App oach o
P oduce Decisions and Explana ions Using CNNs”, In 9 h Ibe ian Con e ence on Pa e n
Recogni ion and Image Analysis, Sp inge In e na ional Publishing, pages 3-15, 2019.
This pape was selec ed o o al p esen a ion as one o he bes anked pape s in he ML ca ego y a
he 9 h Ibe ian Con e ence on Pa e n Recogni ion and Image Analysis (IbPRIA), ha ook place
in July 2019 a Uni e sidad Au onoma de Mad id, Spain. I was he winne o an honou able
men ion and was also selec ed o an ex ended e sion on Pa e n Recogni ion Le e s, he
main a chi al jou nal o he In e na ional Associa ion o Pa e n Recogni ion (IAPR).
4In oduc ion
1.4 Documen S uc u e
The emainde o his documen is di ided in o 6 chap e s, desc ibed as ollows. Chap e
2desc ibes some undamen al concep s ela ed o he ML pa adigm, asks and me ics (sec ion
2.1), ollowed by a b ie his o ical backg ound on DL and some impo an aspec s ega ding his
pa icula opic (sec ion 2.2). The chap e ends wi h basic CNN concep s and he mos widely
used CNN a chi ec u es ele an o his wo k.
Chap e 3 e iews he li e a u e, s a ing by de ailing se e al de ini ions ela ed o XAI, in-
e p e abili y and explainabili y in sec ion 3.1. A e wa ds, aspec s ela ed o XAI sys ems a e
p esen ed, such as i s ele ance nowadays (sec ion 3.2), applica ion scena ios and s akeholde s
(sec ion 3.3), and easibili y and co esponding challenges (sec ion 3.4). The li e a u e e iew
con inues by su eying se e al p oposed axonomies in sec ion 3.5 and e alua ion s a egies in sec-
ion 3.6. Then, an ex ensi e e iew o exis ing explainabili y me hods co e ing model-agnos ic,
model-speci ic and in insically in e p e able models is done in sec ions 3.7,3.8 and 3.9, espec-
i ely. To end he chap e , a summa y and discussion o he e iewed me hods is p esen ed in
sec ion 3.10.
Chap e 4desc ibes he p oposed join app oach, s a ing wi h ela ed wo k ha inspi ed he
p oposed solu ion in sec ion 4.1, as well as a b ie o e iew (sec ion 4.2), ollowed by a de ailed
cha ac e isa ion o bo h classi ie (4.3) and explaine (4.4). Finally, he p oposed aining p ocess
and loss unc ions a e p esen ed in sec ion 4.5.
The nex chap e , chap e 5, is dedica ed o he expe imen al me hodology, including he syn-
he ic and eal da ase s used o es he p oposed a chi ec u e (sec ion 5.1). A e wa ds, he chap e
ocuses on hype pa ame e uning s a egies (sec ion 5.2) and a b ie s udy o he connec ion be-
ween explaine and classi ie (sec ion 5.3).
Chap e 6p esen s he ob ained esul s on bo h syn he ic and eal da ase s, and ega ding
he wo app oaches add essed: unsupe ised (sec ion 6.1) and weakly supe ised (sec ion 6.2)
app oaches.
Finally, chap e 7concludes his wo k and add esses u u e esea ch di ec ions.
Chap e 2
Fundamen als
This chap e ocuses on undamen al concep s ha a e he unde lying building blocks o his
wo k and any ML sys em. Fi s , he chap e s a s wi h an o e iew o wha ML is, as well as he
h ee main lea ning pa adigms: Supe ised, Unsupe ised and Rein o cemen Lea ning. Then, he
ocus is shi ed o DL in pa icula , s a ing wi h some his o ical backg ound, and hen p oceeding
o Deep Neu al Ne wo ks (DNNs), hei p incipal componen s and how hey a e ained. Finally,
he chap e ends wi h CNNs, wha dis inguishes hem om ”no mal” DNNs, and some common
CNN a chi ec u es ele an o he emainde o his wo k. Fo a mo e in-dep h analysis and
explana ion, he eade can ind de ailed in o ma ion in [3] and [4].
2.1 Machine Lea ning
2.1.1 O e iew
ML is a subse o AI ha s udies models capable o lea ning wi hou explici ins uc ions,
elying only on da a and i s pa e ns. The e m da es back o 1959 and is a ibu ed o A hu
Samuel, which de ined i as:
”[Machine Lea ning is he] ield o s udy ha gi es compu e s he abili y o lea n wi hou being
explici ly p og ammed.” - A hu Samuel, 1959
The adi ional p og amming pa adigm, which is shown in igu e 2.1 a), akes some da a as
inpu , applies an algo i hm w i en by a p og amme o ha da a and ou pu s he esul s o he
calcula ions included in he compu e p og am. Con e sely, ML sys ems, igu e 2.1 b), ake as
inpu he da a and he desi ed ou pu , he so-called g ound- u h, and lea n a model, ha ep esen s
he lea ned in o ma ion om ha da a.
5
6Fundamen als
Da a
Algo i hm
Ou pu
T adi ionalP og amming
(a)
Da a
Ou pu
Algo i hm
MachineLea ning
(b)
Figu e 2.1: T adi ional P og amming e sus ML.
The e a e di e en ypes o ML sys ems, bu mos o hem ollow one o hese h ee pa adigms:
supe ised, unsupe ised, o ein o cemen lea ning. Figu e 2.2 ep esen s hese h ee pa adigms
and hei main di e ences.
Figu e 2.2: Lea ning Pa adigms in ML. Ex ac ed om [5].
On he one hand, in supe ised lea ning, he ML model has human supe ision, much like
a eache : he model is ed no only inpu da a, bu also labels o each da a ins ance, i.e., he
desi ed/known ou pu o each ins ance. Fo example, a se o images o dogs and ca s a e gi en o
he model, as well as he in o ma ion o each image i i con ains a dog o a ca . Ha ing aining
da a, X, and i s espec i e labels, Y, he goal o he ML model is o lea n he unc ion ha bes
app oxima es he unknown mapping unc ion o Xin o Y,Y= (X). This app oxima ed unc ion
is ob ained i e a i ely: he algo i hm p edic s he ou pu o an ins ance and is co ec ed i he
p edic ion does no co espond o he label; lea ning goes on un il a ce ain deg ee o pe o mance
has been achie ed. Ha ing such an app oxima ion, one is able o p edic Y om X. So, he e he
goal is o minimise he di e ence be ween he p edic ed and he ac ual ou pu .
On he o he hand, in unsupe ised lea ning scena ios he da a is unlabelled, so he e is no
” eache ” o supe ise he lea ning p ocess. In his app oach, he goal is o unco e unde lying
pa e ns in he da a, such as clus e s o simila animals, o example.
2.1 Machine Lea ning 7
The main p oblem wi h supe ised lea ning esides in he need o eno mous amoun s o la-
belled da a, which is cos ly o ob ain. To mi iga e his need, bu s ill main ain some deg ee o
supe ision, be ween he supe ised and unsupe ised pa adigms lies weakly supe ised lea ning,
which is an app oach whe e he sys em is no ully supe ised. Weak supe ision can be done in
di e en ways [6]:
•incomple e supe ision (semi-supe ised lea ning): only pa o he aining da a is labelled
•inexac supe ision: he aining da a only has coa se-g ained labels o he exis ing labels
a e om ano he ask, usually o highe le el
•inaccu a e supe ision: he gi en labels a e no always g ound- u h
The hi d pa adigm, ein o cemen lea ning, is a bi di e en om he o he wo app oaches.
He e he lea ning sys em is called an ”agen ”, ha obse es he en i onmen and pe o ms ac ions,
ge ing ewa ds o penal ies in e u n.
In his wo k, we will app oach supe ised, weakly supe ised and unsupe ised lea ning.
When using weak supe ision, we a e only alking abou inexac supe ision, o which we will
b oadly e e as ou weakly supe ised app oach.
2.1.2 Tasks
ML can be used o ackle di e en p oblems, which will be b ie ly men ioned in his sec ion.
In ac , ML is e y use ul when sol ing p oblems ha a e oo complex o be sol ed by adi-
ional p og amming echniques, ei he because hey equi e a lo o hand- uning o ha e no known
algo i hm [3].
A ypical supe ised lea ning ask is classi ica ion, in which one is simply in e es ed in p e-
dic ing he class o each da a ins ance om a known disc e e se o classes, o example, p edic ing
ypes o dog b eeds. Ano he common p oblem is eg ession. He e he ou pu is con inuous, so
he goal is o p edic a nume ical alue, o example es ima e he p ice o a house.
Unsupe ised echniques a e mos commonly used o pe o m clus e ing, i.e., ind na u al
g oups in he da a and assign a label o each g oup, o dimensionali y educ ion, educing he da a
ea u e space.
These echniques a e usually independen o he ype o inpu da a, ha can be abula (nume -
ical, ca ego ical, e c), sequen ial o isual. In ac , ML has been applied success ully in many CV
asks. CV is an in e disciplina y ield ha aims a ex ac ing meaning ul in o ma ion om isual
da a, such as images and ideos. Typical CV asks include image classi ica ion, objec de ec ion,
seman ic and ins ance segmen a ion and image cap ioning. Image classi ica ion only add esses he
p oblem o iden i ying which objec s appea in an image. Objec de ec ion localises hose objec s
inside he image, usually h ough bounding boxes. Seman ic segmen a ion consis s in spli ing an
image in o segmen s and labelling each o hese segmen s acco ding o he objec hey con ain.
Ins ance segmen a ion ex ends seman ic segmen a ion by dis inguishing objec s wi hin he same
class. Finally, image cap ioning ocuses on gene a ing desc ip ions o he con en s o an image.
14 Fundamen als
Op imise s When aining a DNN i is also impo an o choose he co ec op imisa ion s a egy,
i.e., he i e a i e op imisa ion p ocess o educe he loss unc ion. Vanilla GD is a ely used when
aining DNNs, since i in ol es using he whole aining se a once, which is usually no possible
as i canno i in o he GPU’s a ailable memo y. Mul iple al e na i es do exis , being S ochas ic
GD (SGD) he p e e ed one. These op imise s ha e a s a ic lea ning a e, bu o he op imise s
ha e adap i e lea ning a es. Gene ally, adap i e op imise s each he minimum as e , because he
esul ing upda es a e poin ed mo e di ec ly owa ds he global minimum, and equi e less uning
o he lea ning a e hype pa ame e . Examples o such op imise s include Adam [17], Adadel a
[18] and, mos ecen ly, Adabound [19]. Much like he choice o ac i a ion unc ion, choosing he
op imise and i s lea ning a e is dependen on he ne wo k a chi ec u e and he da a.
Regula isa ion A DNN has housands o pa ame e s and wi h i he possibili y o i ing huge
complex da ase s. Howe e , his can lead o wha is known as o e i ing. O e i ing occu s when
he ne wo k is oo well adap ed o he aining da a and, hus, gene alises poo ly on unseen da a.
A oiding o e i ing can be done h ough egula isa ion, i.e., by cons aining he model and lim-
i ing i s many deg ees o eedom [3]. One way o doing his is by cons aining he ne wo k’s
weigh s, o example by adding he penalised l1o l2no ms o he weigh s o he loss unc ion.
Ano he egula isa ion echnique is Ea ly S opping and, as he name implies, i simply consis s
in s opping aining when he alida ion loss s ops dec easing. This echnique usually wo ks well
in p ac ice when combined wi h o he egula isa ion me hods, such as D opou , o example [3].
D opou was i s p oposed by Hin on e al. [20]: a e e y aining s ep each neu on has a p ob-
abili y, he d opou a e, o being empo a ily igno ed du ing ha aining s ep. Using D opou ,
neu ons become less sensi i e o small changes in he inpu , because hey canno co-adap wi h
hei neighbou s, no ely excessi ely on some o hei inpu neu ons, which leads o a mo e obus
ne wo k ha gene alises be e . D opou can also be hough o as c ea ing an a e aging ensemble
o smalle neu al ne wo ks, as a unique ne wo k is gene a ed a e e y aining s ep [3]. O he
egula isa ion echnique is Da a Augmen a ion, which simply in ol es gene a ing new aining
ins ances om exis ing ones, by pe o ming ans o ma ions such as o a ions and ansla ions,
a i icially inc easing he size o he aining se and, hus, educing he possibili y o o e i ing.
Finally, Ba ch No malisa ion, p oposed by Io e e al. [21], besides being a egula isa ion ech-
nique, also imp o es he speed, s abili y and pe o mance o DNNs. This echnique add esses
he In e nal Co a ia e Shi p oblem, he ac ha he dis ibu ion o each laye ’s inpu s changes
du ing aining, and he anishing/exploding g adien s p oblem. B ie ly, i consis s in adding an
ope a ion jus be o e he ac i a ion unc ion o each laye , ze o-cen e ing and no malising he in-
pu s, and hen scaling and shi ing he esul , i.e., i allows he model o lea n he op imal scale
and mean o he inpu s o each laye . To do so, he algo i hm compu es he mean and s anda d
de ia ion o he inpu s o e each ba ch o he aining da a. Du ing in e ence, i uses he mean
and s anda d de ia ion o he whole aining se . This echnique educes he anishing g adien s
p oblem, lessens he dependence on weigh ini ialisa ion, allows using highe lea ning a es and,
consequen ly, speeds up he lea ning p ocess, and ac s as a egula ise , educing he need o o he
2.2 Deep Lea ning 15
egula isa ion echniques. Howe e , due o he addi ional calcula ions i in oduces, he ne wo k
makes slowe p edic ions [3].
2.2.3 Con olu ional Neu al Ne wo ks
CNNs a e a ype o DNNs ha ake as inpu image da a. These ne wo ks a e mainly used o
isual applica ions, such as image analysis and classi ica ion, image and ideo ecogni ion, objec
de ec ion and localisa ion, o image segmen a ion.
These ne wo ks eme ged om he s udy o he b ain’s isual co ex by Nobel lau ea es Da id
H. Hubel and To s en Wiesel [22,23,24], ha ing inspi ed he neocogni on [25] in 1980, which
e ol ed in o cu en CNNs. The au ho s showed ha neu ons in he isual co ex ha e small
ecep i e ields ha eac o isual s imuli loca ed in a sub egion o he isual ield. Fu he mo e,
hei wo k showed ha some neu ons eac only o ho izon al lines, while o he s eac o lines
wi h di e en o ien a ions, and ha neu ons wi h la ge ecep i e ields eac ed o mo e complex
pa e ns ha cons i u ed a combina ion o lowe -le el pa e ns [3].
Figu e 2.10 depic s a ypical CNN a chi ec u e, which is composed o some con olu ional and
pooling laye s, which a e desc ibed in sec ion 2.2.3.1, ollowed by some ully connec ed laye s,
ending wi h a So max laye .
Figu e 2.10: Typical CNN a chi ec u e. Ex ac ed om [26].
2.2.3.1 Laye s
CNNs use he same laye s as egula DNNs, bu in oduce wo new ypes o laye s: con olu-
ional and pooling laye s.
Con olu ional Laye The con olu ional laye is he co e building block o CNNs. Ins ead o
being connec ed o e e y single pixel o he inpu image, which would happen in egula DNNs,
leading o an exponen ial inc ease in he numbe o pa ame e s, neu ons on he i s con olu ional
laye a e solely connec ed o pixels in hei espec i e ields, simila ly o wha happens in he
neu ons o he b ain’s isual co ex. This is epea ed h oughou he di e en con olu ional laye s,
making his a hie a chical s uc u e ha ocuses on lea ning low-le el ea u es in he i s laye s
and hen joining hem in o highe -le el ea u es as he laye s s ack-up.
16 Fundamen als
As depic ed in igu e 2.10, he ou pu o each laye is a se o s acked ea u e maps o he
same size. A ea u e map esul s om a laye ull o neu ons using he same il e 4, by means
o he con olu ion ope a ion. This ope a ion consis s in compu ing he do p oduc be ween each
ea u e map o he p e ious laye and he espec i e il e . Du ing aining, he CNN lea ns he
mos sui ed il e s o he ask and combines hem in o mo e complex pa e ns [3]. In o he wo ds,
a con olu ional laye simul aneously applies mul iple il e s o i s inpu s, hus de ec ing mul iple
ea u es. Con olu ional laye s, as desc ibed so a , p ese e he inpu ’s heigh and wid h, changing
only i s dep h. Howe e , heigh and wid h can be changed by using wha is called a s ide, g ea e
han 1. The s ide de ines he dis ance be ween consecu i e ecep i e ields. Fo example, a s ide
o 2 indica es ha he il e mo es 2 pixels in he inpu o e e y pixel in he ou pu , so i de ines he
a io be ween inpu and ou pu , hus o igina ing a ea u e map wi h hal he wid h and he heigh o
he o iginal inpu image. Con e sely, he e e se ope a ion is also possible, in an ope a ion called
anspose con olu ion o ac ionally s ided con olu ion, as i equals a con olu ion wi h a s ide
lowe han 1.
Pooling Laye The pooling laye sub samples one o mo e dimensions o i s inpu image, so as
o educe he compu a ional and memo y equi emen s needed. Simila ly o con olu ional laye s,
he pooling laye s a e connec ed o a limi ed numbe o neu ons in he p e ious laye , he ones
in hei ecep i e ield. Howe e , a pooling laye p esen s no weigh s and simply agg ega es he
inpu s using a unc ion like max, in which only he maximum inpu alue is p opaga ed o he nex
laye , o a e age, in which he mean o he inpu alues is p opaga ed o he nex laye , as depic ed
in igu e 2.11. The pooling ope a ion is usually applied o he image’s wid h and heigh , bu i can
also be applied o i s dep h, educing he numbe o channels o he inpu image. While pooling
pe o ms downsampling, upsampling is also possible h ough unpooling. Finally, one can also
employ global pooling, which is an ope a ion ha educes a ma ix om 3D o 1D, by choosing
only one alue (ei he he maximum o he a e age) om each ea u e map, he e o e educing an
inpu o w×h×d o 1×1×d. A global pooling laye is o en used a e he con olu ion-pooling
s ages and be o e ully connec ed laye s, o p epa e he ea u es o hese las laye s, wi hou he
need o la ening.
Figu e 2.11: Max pooling example. In his case, a s ide o 2 is used, which means ha he inpu
is 2 imes he size o he ou pu . A il e o size 2 ×2 is used, so e e y 4 pixels in he inpu a e
mapped o 1 pixel in he ou pu . Ex ac ed om [27].
4A il e o con olu ional ke nel is a ep esen a ion o a neu on’s weigh s as an image he size o i s espec i e
ecep i e ield [3]. In o he wo ds, i is a k×kma ix o lea nable weigh s.
2.2 Deep Lea ning 17
2.2.3.2 A chi ec u es
As p e iously men ioned, a ypical CNN a chi ec u e, depic ed in igu e 2.10, comp ises con-
olu ional laye s ollowed by ReLU and pooling laye s. Usually, as he image p og esses h ough
he ne wo k i ge s smalle and smalle , bu also deepe and deepe (wi h mo e ea u e maps),
depending on he il e s used [3]. A e hese laye s, a gene al eed o wa d ne wo k is added,
composed o some ully connec ed laye s and hei espec i e ac i a ion laye s, ypically ending
wi h a so max laye o ou pu es ima ed class p obabili ies.
Typical CNN a chi ec u es, b ie ly desc ibed below, include he amous LeNe -5, AlexNe ,
VGG16, GoogLeNe and ResNe . Also, a CNN a chi ec u e called U-ne is e iewed, because i
is ele an o he emainde o his wo k.
LeNe -5 LeNe -5 [28] is one o he bes known CNN a chi ec u es. C ea ed by Yann LeCun in
1998 and applied o handw i en digi ecogni ion on he widely used MNIST da ase , i included
wo con olu ion-a e age pooling s ages, ollowed by ano he con olu ional laye and wo ully
connec ed laye s. I used he hype bolic angen unc ion as ac i a ion and a sligh ly di e en
ou pu laye ha ou pu ed he Euclidean dis ance be ween i s inpu and weigh ec o s [3]. I s
inpu s had size 32 ×32.
AlexNe AlexNe [29] was in oduced in 2012 by K izhe sky e al. and achie ed ou s anding
pe o mances on he ImageNe La ge Scale Visual Recogni ion Challenge5(ILSVRC), a classi-
ica ion challenge o o e 14 million images belonging o 1000 classes. This CNN is simila o
LeNe -5, bu la ge and deepe : wi h wo con olu ion-pooling s ages, h ee con olu ional laye s
and h ee dense laye s. I s inpu s a e se en imes la ge han hose o LeNe -5, wi h size 224×224,
and i uses la ge ke nels in he i s laye s, namely 11 ×11 and 5 ×5. Con a ily o LeNe -5, i
uses ReLU ac i a ion unc ions o he hidden laye s and so max o he ou pu laye .
VGG-16 P oposed by Simonyan e al. [30] (membe s o he Visual Geome y G oup hence
he name VGG) in 2014, i was one o he bes pe o ming models in ILSVRC 2014, along wi h
GoogleNe . I imp o ed on AlexNe by using smalle ke nels, o size 3×3, which s acked in o
3 laye s ha e he same ecep i e ield as 7×7 ke nels, bu inco po a ing 3 ins ead o only 1 non-
linea ec i ica ion laye , making he decision unc ion mo e disc imina i e. Also, a single 7 ×7
ke nel has 81% mo e pa ame e s han 3 s acked 3 ×3 ke nels [30]. This a chi ec u e, p esen ed
in igu e 2.12, also akes as inpu 224×224-sized images, ha go h ough 2 con olu ion-pooling
s ages, ollowed by 3 con olu ion-con olu ion-pooling s ages and 3 ully connec ed laye s, making
a o al o 16 weigh laye s. O he VGG a chi ec u es exis , such as VGG-19, wi h 19 weigh laye s.
5h p://image-ne .o g/
18 Fundamen als
Figu e 2.12: VGG-16 a chi ec u e. Ex ac ed om [31].
GoogLeNe GoogLeNe [32] was de eloped by Szegedy e al. om Google Resea ch and won
he ILSVRC 2014 challenge, by building a much deepe ne wo k han p e ious CNNs. This a -
chi ec u e is composed by smalle ne wo ks called incep ion modules, depic ed in igu e 2.13, ha
allow GoogLeNe o use pa ame e s mo e e icien ly, leading o ewe pa ame e s han i s p ede-
cesso s, such as AlexNe . The incep ion module uses 1×1 con olu ions o pe o m dimensionali y
educ ion and has di e en ke nel sizes in he second se o laye s, 3×3 and 5 ×5, o be able o
cap u e pa e ns a di e en scales. The whole a chi ec u e includes 9 incep ion modules wi h 3
laye s each.
Figu e 2.13: Incep ion module wi h dimensionali y educ ion. Ex ac ed om [32].
ResNe ResNe [33] won he ILSVRC 2015 challenge wi h an ex emely deep ne wo k o 152
laye s and was de eloped by Kaiming He e al.. T aining such a deep ne wo k was possible
h ough he in oduc ion o skip connec ions: connec ions ha link one laye wi h ano he one a
bi highe up in he s ack. When aining a DNN, he goal is o model a unc ion h(x); by adding
he inpu x o he ou pu o he ne wo k, i will be o ced o model (x) = h(x)−x, in wha is called
2.2 Deep Lea ning 19
esidual lea ning [3]. Thanks o his echnique, he signal lows easily ac oss he whole ne wo k,
which conside ably speeds up aining. A diag am o he de eloped esidual block can be ound
in igu e 2.14, whe e he skip connec ion is also depic ed. The whole a chi ec u e is composed by
a e y deep s ack o esidual uni s: 2 con olu ional laye s wi h Ba ch No malisa ion and ReLU
ac i a ion, and 3 ×3 ke nels. Typically, ResNe is used wi h 50 esidual laye s, also known as
ResNe -50.
Figu e 2.14: Residual lea ning block wi h a skip connec ion. Ex ac ed om [33].
Figu e 2.15 shows he e olu ion o he op 5% e o on he ILSVRC compe i ion h oughou
he yea s, om 2010 o 2015, wi h a signi ican inc ease in he numbe o laye s used by 2015.
Figu e 2.15: ILSVRC op 5% e o on ImageNe . Ex ac ed om [34].
U-Ne The U-Ne [35] is a di e en CNN a chi ec u e, because i ou pu s an image wi h he
same size as he inpu , a he han class sco es. I was de eloped by Ronnebe ge e al. o ackle
biomedical imaging segmen a ion wi h small da ase s. I is an encode -decode based a chi ec u e:
he encode maps he inpu space o a di e en /la en space, while he decode lea ns he comple-
men a y unc ion ha maps om he la en space o he a ge space. In his case, he a ge space
is he same as he inpu space and he ne wo k is ained end- o-end, pixels- o-pixels. The main
idea is o only use con olu ional laye s and o pe o m anspose con olu ion ope a ions o eco e
he spa ial dimensions o he o iginal inpu , al e ed by he downsampling in oduced by he con-
olu ional laye s. The encode is a ypical ea u e ex ac o based on he VGG-16 ne wo k and
20 Fundamen als
he decode is a shape gene a o ha ou pu s images om he ex ac ed ea u es. This upsampling
pa uses c opped ea u e maps om he downsampling pa o a oid losing global in o ma ion, as
depic ed in igu e 2.16.
Figu e 2.16: U-Ne a chi ec u e, composed by an encoding and a decoding pa h. Ex ac ed om
[35].
Chap e 3
Explainable Machine Lea ning:
Li e a u e Re iew
This chap e is dedica ed o a comp ehensi e and c i ical e iew o he exis ing li e a u e on
in e p e abili y and explainabili y in ML, mo e speci ically in DNNs. Fi s , undamen al concep s
such as XAI, in e p e abili y and explainabili y a e in oduced. Then, i is explo ed why hese
p ope ies a e needed and wha makes hem ele an o ML sys ems, ollowed by he goals and
applica ions o explainabili y, as well as he challenges ha he esea ch communi y has o ace in
o de o c ea e explainable sys ems.
A e his con ex ualisa ion and mo i a ion, di e en explainabili y axonomies a e desc ibed,
along wi h a ious e alua ion s a egies ound in he e iewed li e a u e.
Then, he ocus is shi ed owa ds s a e-o - he-a explainabili y me hods, s a ing wi h model-
agnos ic me hods, i.e., me hods ha can be applied o any kind o ML model; hese a e u he
di ided in o global and local me hods. The emaining me hods all unde he ca ego y o model-
speci ic me hods and, since his wo k ocuses on CNNs, his sec ion e ol es a ound he explain-
abili y o DNNs.
To close he chap e o he explainabili y me hods a e b ie ly men ioned, such as in insically
in e p e able models, succeeded by a summa y o he e iewed me hods and, inally, a discussion
o he e iewed li e a u e.
3.1 Concep s and De ini ions
3.1.1 Explainable A i icial In elligence (XAI)
A i icial In elligence (AI) is nowadays widely adop ed in ou daily li es, i s p esence anging
om mo ie ecommenda ion sys ems o ailo ed ad e ising. As such, i is inc easingly impo an
o be able o unde s and he easons behind he decisions o hese sys ems (see sec ion 3.2). This
is whe e Explainable AI comes in o play, p oposing a pa adigm shi owa ds anspa en AI [36].
Fo mally, he e is no s anda d de ini ion o XAI, bu i gene ally e e s o esea ch and ini-
ia i es owa ds AI anspa ency and us [36]. Acco ding o he De ense Ad anced Resea ch
21
22 Explainable Machine Lea ning: Li e a u e Re iew
P ojec s Agency (DARPA), XAI is an eme ging esea ch ield wi h he goal o de eloping explain-
able models wi hou sac i icing high p edic ion accu acies and enabling humans o unde s and and
us AI pa ne s [2]. The e m i s su aced in 2004 in a pape by Van e al. [1] as a desc ip ion
o hei AI aining sys em de eloped o he U.S. A my capable o p oducing explana ions and
answe ing ques ions like "Wha is he pla oon’s ammo s a us?".
As can be seen i igu e 3.1, he XAI pa adigm di e s om he p esen AI/ML pa adigm by
eplacing classic Lea ning Models by an Explainable Model and an Explana ion In e ace capable
o answe ing he use ’s ques ions abou he decisions made by he sys em, leading o a mo e
anspa en , obus and us wo hy sys em.
Figu e 3.1: Compa ison be ween oday’s ML pa adigm wi h he u u e pa adigm in oduced by
XAI [2].
3.1.2 In e p e abili y and Explainabili y
Despi e he g owing in e es in XAI as a esea ch ield and as he u u e o AI/ML, he ield
s ill su e s om ambiguous e minology. In ac , wo k in his ield usually employs he e ms
in e p e abili y and explainabili y, and o en does so in e changeably [37]. Howe e , he wo d
in e p e abili y is p e e ed o explainabili y in he ML communi y, as obse ed in igu e 3.2.
O he e ms such as unde s andabili y [38] o in elligibili y [39] appea , bu less equen ly.
Figu e 3.2: Google ends o in e p e abili y and explainabili y in scien i ic and non-scien i ic
con ex s [36].
3.1 Concep s and De ini ions 23
As no ed by Lip on [40], he e is no ag eed upon de ini ion o in e p e abili y, bu pape s
s ill use he e m in a ”quasi-ma hema ical way”, which inc eases he need o a clea axonomy.
Acco ding o he Me iam-Webs e English Online Dic iona y1, o in e p e means ” o explain o
ell he meaning o , p esen in unde s andable e ms”. Based on his de ini ion, Doshi-Velez e al.
[41] de ine in e p e abili y as he abili y o p esen in unde s andable e ms o a human being. The
same au ho s also s a e ha : ”In e p e abili y is NOT abou unde s anding all bi s and by es o
he model o all da a poin s (we canno ). I ’s abou knowing enough o you downs eam asks.”
[42].
Howe e , o he de ini ions a ise in li e a u e, sugges ing ha he concep o in e p e abili y
consis s o a my iad o dis inc ideas, as sugges ed by Lip on [40]. Fo example, Mon a on e al.
[43] dis inguish he e ms in e p e a ion and explana ion:
De ini ion 1. An in e p e a ion is he mapping o an abs ac concep (e.g. a p edic ed class)
in o a domain ha he human can make sense o .
De ini ion 2. An explana ion is he collec ion o ea u es o he in e p e able domain, ha ha e
con ibu ed o a gi en example o p oduce a decision (e.g. classi ica ion o eg ession).
The e o e, an example o an in e p e able domain would be an image o ex , whe eas an
explana ion would be a hea map which highligh s he inpu image’s pixels ha con ibu ed he
mos o a ce ain classi ica ion ou pu , o highligh ed ex in he case o na u al language p ocessing
[44].
Gilpin e al. [45] also makes a dis inc ion be ween in e p e abili y and explainabili y. Sim-
ila ly o wha Doshi-Velez e al. [41] s a e and was p e iously men ioned, Gilpin e al. de-
ine in e p e abili y as a desc ip ion o he in e nals o a sys em in a human-unde s andable way.
No wi hs anding, explainabili y encompasses he abili y o summa ise he easons o he model’s
beha iou , o gain use us o o p oduce insigh s abou he causes o he model’s decisions. As a
esul , he au ho s conside in e p e abili y a c ucial bu no single aspec o achie e explainabili y:
comple eness is also needed. In hei iew, a comple e sys em should be able o desc ibe i s ope a-
ion in an accu a e way and, he e o e, a sys em is mo e comple e when i allows o i s beha iou
o be p edic ed in mo e si ua ions. Thus, an explana ion can be e alua ed acco ding o i s in e -
p e abili y and comple eness; explainable models a e in e p e able by de aul , bu he e e se is no
gua an eed. Mo eo e , his di e en ia ion in oduces an in e p e abili y-comple eness ade-o :
he mo e comple e (accu a e) he explana ion, he less in e p e able i is o humans; con e sely,
highly in e p e able desc ip ions usually do no p o ide p edic i e powe . As an example one can
conside ha a pe ec ly comple e explana ion o a DNN would be a lis ing o all he ma hema ical
ope a ions and pa ame e s o he ne wo k, which is no unde s andable o a human and does no
allow one o ep oduce he mapping om inpu s o ou pu s [45].
1In e p e [De . 1]. ( e b). In Me iam Webs e Online, Re ie ed Feb ua y 03, 2019, om h ps://www.
me iam-webs e .com/dic iona y/in e p e abili y
30 Explainable Machine Lea ning: Li e a u e Re iew
3.4 Feasibili y and Challenges
Ha ing discussed he ele ance and applica ions o explainabili y, i is now impo an o ad-
d ess i i is easible, pa icula ly in he legal con ex ; in o he wo ds, i i is possible o ex ac
om AI sys ems he same kind o explana ions ha a e nowadays equi ed o humans. Doshi-
Velez e al. [50] a gue ha i is, indeed, echnically easible, mainly because an explana ion is no
he same as anspa ency, in he sense ha , as men ioned in subsec ion 3.1.2, explaining does no
mean unde s anding e e y bi ha lows in he sys em, no mo e han an explana ion om a human
being en ails knowing e e y signal low h ough e e y neu on.
The e o e, an explana ion as equi ed unde he law can be summa ised by he ollowing p op-
e ies, de i ed om he ques ions also p esen ed in subsec ion 3.1.2:local explana ion and local
coun e ac ual ai h ulness. A local explana ion consis s o explaining a speci ic decision, a he
han he whole sys em’s beha iou . Local coun e ac ual ai h ulness ela es o he ac ha hu-
mans expec an explana ion o be causal. Conside ing a c edi loan p edic ion sys em, an example
o a local explana ion would be: pe son A go hei loan denied because o paymen his o y, while
o pe son B i was insu icien income ha esul ed in he denial o he loan. An example o local
coun e ac ual ai h ulness would be ha i a pe son was old ha hei income was he main ea-
son o he denial o a loan, hen hey migh expec ha i ha income inc eases, he sys em migh
change i s p e ious decision and allow he loan.
Again, i is impo an o no e ha bo h hese p ope ies can be achie ed wi hou needing o
know he p ocess by which he sys em eached i s decision, which add esses conce ns ega ding
ade sec e s, because an explana ion can be p o ided wi hou he need o unco e ing he con en s
o said sys em. An example is also p o ided in [50]: i a legal ques ion consis s in knowing
whe he ace biased a loan decision, hen he AI sys em migh be ed a ia ions o he o iginal
inpu s changing only he a ibu e ace. I i s ou pu s u n ou o be di e en , hen i is easonable
o a gue ha ano he a ibu e, such as gende , in luenced he decision, which cons i u es a legally
su icien explana ion and no mo e in o ma ion is needed unde he law.
Figu e 3.3: F amewo k o XAI sys ems, consis ing o an explana ion sys em, along wi h he
p op ie a y model. Wi h his app oach, he ML model can emain p o ec ed, which add esses
conce ns ega ding ade sec e s. Ex ac ed om [50].
3.4 Feasibili y and Challenges 31
Following his line o hough , in [50] he au ho s also p opose a amewo k o explainable
AI sys ems, as depic ed in igu e 3.3. I is a gued ha he model i sel emain a black-box model,
possibly p op ie a y, and ha explana ions a e included in a sepa a e sys em. Thus, he AI sys em
can be op imized o p oduce p edic ions, ˆy, ha ma ch he eal wo ld, y, while he explana ion
sys em mus p o ide a human-in e p e able ule, ex(x), ha akes in he same inpu as he p edic-
ion model and also ou pu s a p edic ion, ˜y. So, o sa is y local coun e ac ual ai h ulness, ˆyand
˜ymus be he same unde small a ia ions o inpu x. This amewo k allows o he concep s
desc ibed abo e o be quan i iable: o any x, i can be checked i ˆyequals ˜yand i hese p edic-
ions emain consis en o small a ia ions o x(xcould be he ace a ibu e as desc ibed in he
a o emen ioned example), meaning ha i can be measu ed du ing how much ime an explana ion
sys em is ai h ul and he speci ic ins ances in which i is.
Ano he impo an aspec o ake in o accoun when building XAI sys ems, is he challenge ha
he accu acy-explainabili y ade-o poses. Gene ally speaking, he mo e accu a e he model, he
less explainable, and ice e sa, as depic ed in igu e 3.4. Once mo e, inding he adequa e poin
in his spec um is icky and highly dependen on he a ge domain and use s o he sys em.
Explainabili y
Accu acy +
-
+-
Linea
Reg ession DNNs
Decision
T ees
Random
Fo es s SVMs
Figu e 3.4: Accu acy-Explainabili y ade-o . The mo e accu a e he model, he less explainable.
DNNs a e loca ed on he a igh o his spec um, achie ing ou s anding p edic i e accu acies,
bu emaining black-boxes, while simple models, such as Decision T ees o Linea /Logis ic Re-
g ession exhibi lowe accu acies, bu a e easie o unde s and.
As b ie ly men ioned in sec ion 3.2, gi en hese challenges ha a ec he easibili y o XAI
sys ems, a sensible app oach is o p o ide a ious le els o g anula i y o explana ions, which
make he sys em adap able o di e en con ex s o applica ion. P eece e al. [67] sugges h ee
laye s o explana ions:
•Laye 1 - T aceabili y: ep esen a ion o in e nal s a es o he model, capable o showing
ha he sys em pe o med as expec ed, o example a saliency map o inpu laye ea u es
ha led o he classi ica ion ou pu o ”dog”. This i s laye usually in e es s heo is s and
de elope s and no as much use s.
•Laye 2 - Jus i ica ion: ep esen a ions linked o laye 1 ha o e seman ic ela ionships
be ween inpu and ou pu ea u es, showing how he sys em pe o med as expec ed. These
ep esen a ions can be o di e en modali ies, possibly p esen ed simul aneously, comple-
men ing each o he . An example would be a seman ic anno a ion o he salien dog ea u es.
This laye is ailo ed o de elope s and use s.
32 Explainable Machine Lea ning: Li e a u e Re iew
•Laye 3 - Assu ance: ep esen a ions linked o laye 2, designed o gi e hei ecipien s
con idence in he sys em’s decisions, in mo e global e ms han laye 2; o example coun-
e ac ual examples showing ha he sys em does no classi y a ca as a dog.
In addi ion o ha ing di e en le els o explana ions, combining di e en explainabili y me h-
ods is also a sensible app oach, since such a sys em would be able o p o ide di e en modali ies
o explana ions ha can complemen each o he . Also, each use can choose which modali y(ies)
o analyse acco ding o hei p e e ences o expe ise. An example o a possible sys em is depic ed
in igu e 3.5. The p oposed sys em p edic s i a pa ien has ca diomegaly (a medical condi ion in
which he hea p esen s an enla gemen ) based on ches X-Rays. The sys em is highly adap -
able o he a ge audience, since i can ou pu a p edic ion, i s p obabili ies, a isual explana ion,
coun e ac ual examples and a seman ic desc ip ion. This way, he use can ob ain complemen-
a y in o ma ion om di e en sou ces, leading o a be e unde s anding o he model and i s
decisions, as well as a g ea e lexibili y o accommoda e he a ge audience’s p e e ences and
knowledge.
Figu e 3.5: Combining explana ion me hods in a ca diomegaly p edic ion sys em. The sys em is
able o p oduce a class label, ou pu p obabili ies, isual saliency, coun e ac ual examples and a
seman ic explana ion. Ex ac ed om [68].
3.5 Taxonomy 33
3.5 Taxonomy
Since explainable ML is s ill in i s in ancy as a esea ch opic, he ield lacks a clea and
acknowledged axonomy, simila ly o wha happens o he de ini ions o he co e concep s such
as explainabili y i sel . Ha ing such a axonomy is ex emely impo an o allow esea che s o
classi y, e alua e and compa e hei me hods.
This sec ion ocuses on ca ego ising ML explainabili y me hods acco ding o di e en c i e ia
p esen ed in li e a u e and pa ially su eyed in [69]. The mind-map o igu e 3.6 ep esen s he
i e axonomies e iewed, as well as hei subca ego ies.
B ie ly, he loca i e c i e ion classi ies a me hod acco ding o i s posi ion ela i ely o he
model, i i is applied be o e he model (p e-model), while building he model (in-model) o a e
he model (pos -model o pos -hoc). The agnos ici y c i e ion simply di ides me hods as model-
speci ic, ailo ed o a ce ain model, o model-agnos ic, applicable o a ious ypes o models. The
scope c i e ion sepa a es local om global me hods, while he ou pu c i e ion, as he name im-
plies, ca ego ises he me hod acco ding o i s ou pu , which can ange om a ea u e summa y o
an in insically in e p e able model. The anspa ency c i e ion ela es o he in e nal mechanisms
o he models. Finally, he me hodology c i e ion dis inguishes me hods acco ding o hei unc-
ion and pu pose, be i in es iga e he p ocessing o he ep esen a ions o a model, o example.
These axonomies will be u he de ailed in he ollowing sec ions.
Explainabili y
Taxonomies
Loca i e
Pos -Model
In-Model
P e-Model
Me hodology
Explana ion
P oducing
Rep esen a ion
P ocessing
Agnos ici y
Model-Agnos icModel-Specific
Ou pu
In insically
In e p e able Model
Da a Poin
Model In e nals
Fea u e Summa y
Scope
LocalGlobal
T anspa ency Algo i hmic
T anspa ency
Decomposabili ySimula abili y
Figu e 3.6: Mind-map o di e en c i e ia by which explainabili y me hods can be ca ego ised. A
me hod can be desc ibed acco ding o: i s posi ion ela i e o he ML model (loca i e c i e ion),
i s speci ici y owa ds he ML model (agnos ici y c i e ion), i s ange (scope c i e ion), he ype
o ou pu i p oduces (ou pu c i e ion), i s le el o anspa ency ( anspa ency c i e ion), o i s
me hodology (me hodology c i e ion).
34 Explainable Machine Lea ning: Li e a u e Re iew
3.5.1 Loca i e C i e ion
ML explainabili y me hods can be classi ied acco ding o when he me hod is applied ela i ely
o he model, i.e., i he me hod is applied be o e he model, when building he model o a e he
model [42]. Gi en his loca ion p ope y o he me hod in ela ion o he model and a lack o a
uni ying name o hese h ee subca ego ies in he e iewed li e a u e, we call his he Loca i e
C i e ion.
P e-model Examples o p e-model me hods include isualisa ion and explo a o y da a analysis.
The main goal is o ex ac meaning ul in o ma ion om da a be o e building a ML model. Ex-
plo a o y da a analysis can be achie ed h ough clus e ing me hods such as K-Means o K-Nea es
Neighbou s (K-NN) o h ough he use o p o o ypes and c i icisms, as desc ibed in [52].
In-model Al e na i ely, explainabili y can be included wi hin he model i sel , when building i .
This ype o me hods can also be said o be in insically in e p e able [69]. These models can be
ule-, pe - ea u e-, case-, spa si y- o mono onici y-based.
Rule-based models no mally include decision ees, ule se s and ule lis s, all o which include
a se o ules desc ibing he di e en classes and p edic ions. Pe - ea u e-based model examples
include he linea and gene alized linea models, as well as he gene alized addi i e model [42].
Al hough hese models a e said o be explainable, one has o keep in mind ha he size o he
esul ing model migh in alida e i s unde s andabili y, as he decision ee, he ule se o he
ea u es may become oo complex o a human o keep ack o . Ne e heless, a ule-based model
is as in e p e able as i s o iginal ea u es [37].
Case-based models le e age he impo ance o example-based explana ions ound in human
easoning, o example when a doc o p esc ibes a ea men o pa ien X because i wo ked on a
pa ien who had simila symp oms.
Spa si y-based models, as he name indica es, ely on he spa si y p ope y. Fo example,
he lowe he numbe o ac i a ions in a neu al ne wo k, he easie i is o ind ou which e en s
led o he model’s decisions. Howe e , lowe ing he numbe o ac i a ions can ende a decision
impossible o achie e and hus, spa si y can educe he model’s accu acy [37].
Finally, mono onici y-based models gua an ee he lea n unc ion’s mono onici y in ela ion o
some o he inpu s, which can acili a e i s in e p e abili y.
Pos -model Las ly, he e is pos -model explainabili y, also known as pos -hoc explainabili y.
Pos -hoc explainabili y en ails selec ing and aining a black-box model and applying explainabil-
i y me hods only a e he model is ained [69]. Se e al echniques o pos -hoc explainabili y
we e p oposed, such as Sensi i i y Analysis, Saliency Maps, Backwa d P opaga ion, Mimic o
Su oga e Models [42]. As in mos cases his ype o me hod is applied independen ly o he p e-
iously ained model, i can also be a gued ha a pos -model/pos -hoc explainabili y me hod is
model-agnos ic, as desc ibed in subsec ion 3.5.3.
3.5 Taxonomy 35
3.5.2 Me hod Ou pu C i e ion
Ano he c i e ion o classi ying an explainabili y me hod is p oposed by [69] and i s a es
ha a me hod can be classi ied wi h espec o i s ou pu s. The e o e, a me hod can ou pu he
ollowing da a.
Fea u e summa y In his ca ego y, he ou pu o he me hod can consis o a summa y o how
each ea u e a ec s model p edic ions, in e ms o ea u e impo ance o ea u e in e ac ion mea-
su es. The ou pu can also be a isualisa ion o he summa y s a is ics, such as Pa ial Dependence
Plo s (PDPs) [70].
Model in e nals The me hod p oduces a se o i s in e nal pa ame e s, such as he weigh s in
linea models o he lea ned ee s uc u e ( ea u es and h esholds) o decision ees. Howe e , in
linea models his classi ica ion o e laps wi h he p e ious ypology, because he lea ned weigh s
a e bo h summa y s a is ics o he model’s ea u es as well as i s in e nal lea ned pa ame e s. In
CNNs one can conside he isualisa ion o lea ned ea u e de ec o s he me hod’s ou pu .
I is wo h no ing ha his ca ego y supe imposes bo h he a o emen ioned ca ego y o in-
model me hods (subsec ion 3.5.1) and he subsequen ca ego y o model-speci ic me hods (sub-
sec ion 3.5.3).
Da a poin The ou pu can also be a da a poin , ei he an al eady exis en o newly c ea ed
poin . Examples o hese me hods include coun e ac ual explana ions o p o o ype iden i ica ion.
Howe e , in o de o hese me hods o be explainable, he ou pu ed da a poin s hemsel es ha e
o be explainable, which is he case o images and ex s, bu no o abula da a wi h hund eds o
ea u es.
Su oga e in insically in e p e able model The me hod gene a es an in e p e able model,
such as a linea eg ession, which app oxima es, locally o globally, he black box model. As
men ioned be o e, he in e p e able model i sel is explained wi h he help o i s model in e nal
pa ame e s o ea u e summa y s a is ics, o example.
3.5.3 Agnos ici y C i e ion
Acco ding o [69], an explainabili y me hod can also be classi ied acco ding o he agnos ici y
c i e ion, i.e., a me hod can be model-speci ic o model-agnos ic. On one hand, model-speci ic
me hods a e es ic ed o speci ic classes o models, o example, can only be used wi h neu al
ne wo ks. By de ini ion, in insically in e p e able models a e model-speci ic. On he o he hand,
model-agnos ic me hods can be applied o any ML model and, by de ini ion, a e pos -hoc me hods,
since hey a e applied a e he model is ained. Thus, hese me hods do no ha e access o he
model’s in e nal pa ame e s.
36 Explainable Machine Lea ning: Li e a u e Re iew
3.5.4 T anspa ency C i e ion
Ano he c i e ion, p oposed by Lip on [40], ela es o he anspa ency p ope y o in e -
p e abili y me hods, i.e., ”How does he model wo k?”. As s a ed by he au ho , anspa ency,
he opposi e o ”blackbox-ness”, is ela ed o he unde s anding o he mechanisms by which he
model wo ks. The au ho de ines di e en le els o anspa ency: en i e model, indi idual com-
ponen s and aining algo i hm.
Simula abili y Rela es o anspa ency a he le el o he en i e model: a model is anspa en
i a human can conside he whole model a once, which sugges s ha an explainable model is
a simple one. In o he wo ds, a model is comple ely unde s ood i a human can, in easonable
amoun o ime and pu ing oge he inpu da a and model pa ame e s, p oduce a p edic ion by
compu ing e e y necessa y calcula ion. This no ion o explainabili y is consis en wi h common
claims ha spa se linea models a e mo e explainable han dense linea models o he same inpu s.
Howe e , o some models he ime i akes o make a p edic ion g ows much slowe han he size
o he model (e.g. decision ees). This obse a ion led he au ho o conside wo sub ypes o
simula abili y, one based on he compu a ion equi ed o pe o m in e ence and o he based on he
size o he model.
Decomposabili y This le el o anspa ency is ied o he no ion ha each pa o he model is
explainable. In his ca ego y, Lip on [40] conside s model pa s such as inpu s, pa ame e s and
calcula ions. The au ho also highligh s ha his concep equi es ha he inpu s be in e p e able,
which is o en no he case when enginee ed o anonymous ea u es exis .
Algo i hmic T anspa ency As he name implies, algo i hmic anspa ency consis s o ans-
pa ency a he le el o he lea ning algo i hm i sel . Mode n DL me hods lack his ype o ans-
pa ency, because e en hough he heu is ic op imiza ion p ocedu es a e powe ul, hey a e no
ully unde s ood no can one gua an ee a p io i ha hey will con e ge on unseen da a.
3.5.5 Scope C i e ion
Addi ionally, a me hod can be conside ed local o global, in he sense ha i explains single
p edic ions o he whole model [69].
Global This ype o in e p e abili y ollows he same de ini ion as p esen ed by [40] and p e i-
ously desc ibed in sec ion 3.5.4. Howe e , Molna [69] a gues ha global model in e p e abili y
is e y di icul o achie e in p ac ice, mainly because humans a e unable o i in o memo y mo e
han a ew pa ame e s. In ac , a ea u e space wi h mo e han 3 dimensions is inconcei able
o humans. No wi hs anding, some models can be unde s ood a a modula le el, such as linea
models, whe e he in e p e able modules a e i s weigh s, o decision ees, whe e he in e p e able
modules a e he spli s and lea nodes. Ne e heless, one mus be ca e ul because he in e p e a ion
3.5 Taxonomy 37
o a single weigh assumes ha he o he weigh s emain cons an , which does no happen in eal
wo ld scena ios.
Local Local in e p e abilli y can be aken in o accoun conside ing only a single p edic ion o
g oups o p edic ions. In he i s case, only a single ins ance and i s p edic ed ou pu a e consid-
e ed. This is based on he ac ha locally, a p edic ion migh depend linea ly o mono onously
on some ea u es, a he han p esen ing a complex dependence on hose ea u es [69]. This same
concep can be applied o a g oup o p edic ions, which can be achie ed using p e iously desc ibed
echniques such as global model in e p e abili y on a modula le el o in e p e abili y o a single
p edic ion. To apply he global me hods one can ake he g oup o ins ances as i i we e he whole
da ase and use he global me hods on ha da a. The indi idual explana ions can be applied o
each ins ance in he g oup and hen agg ega ed o he whole g oup.
3.5.6 Me hodology C i e ion
Finally, Gilpin e al. [45] p opose a axonomy speci ic o DNNs, which ca ego ises me hods
acco ding o how hey app oach explainabili y and wha hey a e ying o achie e; be i ei he
he explana ion o he ne wo k’s p ocessing o he ne wo k’s ep esen a ions, o an explana ion-
p oducing sys em. Adadi e al. [36] e e o his c i e ion as a me hodological app oach o
e alua ing in e p e abili y, so we call i he Me hodology C i e ion.
P ocessing These me hods aim a explaining he neu al ne wo k’s p ocessing, i.e., he way he
ne wo k diges s in o ma ion and u ns i in o p edic ions. The undamen al p oblem o such me h-
ods consis s in inding ways o educe he complexi y o he millions o ope a ions pe o med by
DNNs. This is usually done by inding a p oxy model ha app oxima es he o iginal model, bu
is easie o unde s and, o by c ea ing a saliency map ha highligh s po ions o he compu a ions
ha a e mos ele an . Examples o me hods ha all unde his ca ego y a e Linea P oxy Mod-
els [71], decomposi ion o DNNs in o decision ees [72], au oma ic ule ex ac ion o saliency
mapping [73,74,75,76,77,78,79].
Rep esen a ions Explaining DNNs’ ep esen a ions in ol es unde s anding he ole and s uc-
u e o he da a ha lows h ough he ne wo k, which can be done by laye , by uni o by ec o .
When examining he in e nal s uc u e o he ne wo k by laye , he in o ma ion lowing h ough
each laye is conside ed all oge he , while by uni , single neu ons o il e s a e conside ed indi-
idually, as p oposed by Bau e al. [80] wi h a me hod known as Ne wo k Dissec ion. Fu he -
mo e, we can conside o he ep esen a ion spaces, which is he case wi h ec o analysis me hods,
whe e o he ec o di ec ions in he ep esen a ion space besides laye s and uni s a e conside ed.
The mos signi ican example o an analysis o ep esen a ion ec o s is he wo k o Kim e al.
known as Concep Ac i a ion Vec o s (CAVs) [81].
38 Explainable Machine Lea ning: Li e a u e Re iew
Explana ion-P oducing Sys ems Gilpin e al. [45] dis inguish a hi d ype o explainabili y
me hods, which he au ho s call ”Explana ion-P oducing Sys ems”. As he name sugges s, hese
sys ems in insically p oduce explana ions o include me hods ha make he ne wo k easie o
unde s and. The mos common examples a e a en ion ne wo ks (ne wo ks ained wi h explici
a en ion) [82], disen angled ep esen a ions (me hods ha un a el ep esen a ions in o sepa a e
dimensions desc ibing meaning ul and independen ac o s) [83] o gene a i e explana ions (neu-
al ne wo ks ha p oduce explana ions explici ly as pa o hei aining, such as Visual Ques ion
Answe ing [84]).
3.5.7 Compa ison and Discussion
As e iewed along his sec ion, he e a e many ways o di ide and ca ego ise explainabili y
me hods; six axonomies we e e iewed and hei subca ego ies explained wi h examples o some
me hods belonging o each o hese subg oups. Some axonomies a e comple ely di e en om
one ano he , while some p esen a ce ain deg ee o simila i y, o e en o e lap. Mo eo e , di e en
axonomies p esen di e en app oaches o he ca ego isa ion o explainabili y me hods, as well
as a ying deg ees o g anula i y.
The scope axonomy simply di ides me hods as global o local. Al hough his is a ele an
p ope y o ake in o accoun , i is a oo wide a ca ego isa ion, because i does no dis inguish
me hods like Smoo hG ad [79] om Shapley Values [85]. The i s me hod is local and model-
speci ic, while he la e is also local, bu model-agnos ic.
Con e sely, PDPs [86] a e agnos ic and global, while Coun e ac ual Explana ions [87] a e
also agnos ic, bu local. This shows ha he agnos ici y c i e ion is also no enough o classi y
explainabili y me hods.
The ou pu c i e ion di ides me hods acco ding o hei ou pu , which cons i u es a coa se
g ained classi ica ion, which may no be enough o sepa a e explainabili y me hods. Fo example,
i classi ies PDPs and Shapley Values [85] as ea u e summa y me hods, al hough hey a e global
and local me hods, espec i ely.
The anspa ency c i e ion sepa a es me hods acco ding o hei le els o anspa ency, ocus-
ing on how he model wo ks. This axonomy lacks simplici y, which makes i ha de o use in
p ac ice. Fu he mo e, one o i s sub-classes, algo i hmic anspa ency, which ela es o ans-
pa ency a he algo i hm le el and how he algo i hm lea ns a model om he da a, does no e e
o he lea ned model o how i s p edic ions a e made. In ac , his kind o anspa ency only
equi es knowledge abou he algo i hm, o example, backp opaga ion. Howe e , when alking
abou explainabili y, we a e usually mo e in e es ed in he gene a ed models hemsel es and no
only on he algo i hms ha p oduce hem [47].
The me hodology c i e ion dis inguishes h ee ypes o models, acco ding o how hey ap-
p oach explainabili y: emula ing he p ocessing, explaining he ep esen a ions o inhe en ly p o-
duce explana ions. Howe e , his axonomy classi ies only me hods designed o neu al ne wo ks,
excluding o he ML models, such as SVMs, o example.
3.6 E alua ion 39
The las e iewed axonomy, and p obably he mos widely known and used, is he Loca i e
C i e ion, which di ides me hods as p e-, in- and pos -model. Despi e being a simple and disc im-
ina i e axonomy, i has a uzzy ba ie be ween in- and pos -model me hods. Al hough one migh
be ini ially inclined o say ha all pos -model me hods a e model-agnos ic, since hey a e applied
a e he model is buil , his is no always ue. The e a e some me hods ha a e pos -model, bu
a e model-speci ic, such as Model Comp ession [88], which is ailo ed o DNNs. By de ini ion all
in insically in e p e able models (in-model) a e model-speci ic, bu no all pos -model/pos -hoc
me hods a e model-agnos ic.
Figu e 3.7 summa ises he o e lap ound be ween some o he axonomies e iewed. In-model
me hods a e model-speci ic and p e-model me hods a e model-agnos ic, bu pos -model me hods
can be ei he model-speci ic o agnos ic. The me hodology c i e ion only di ides me hods de-
signed o neu al ne wo ks, making i a model-speci ic axonomy. Ye wo o i s h ee classes, Rep-
esen a ion and P ocessing, a e included in he pos -model ca ego y, while Explana ion-P oducing
models a e in-model.
Model
Agnos ic
Model
Specific
In-Model
P e-
Model
Pos -Model
P ocessing
Explana ion-
P oducing
Loca i e
C i e ion
Me hodology
C i e ion
Agnos ici y
C i e ion
Rep esen a ion
Figu e 3.7: Di e ences and o e lap be ween axonomies: agnos ici y, loca i e and me hodology.
The colou s used o each axonomy a e consis en wi h he ones ound in igu e 3.6.
As de ailed abo e, each axonomy alone is no enough o di ide mos o he me hods e iewed.
As such, we p opose he usage o wo axonomies oge he . Fi s , we di ide me hods as model-
speci ic (sec ion 3.7) o model-agnos ic (sec ion 3.8). Model-agnos ic me hods a e u he di ided
in o global o local me hods. Then, since his wo k ocuses on DNNs, we subdi ide he model-
speci ic me hods acco ding o he me hodology c i e ion.
3.6 E alua ion
Jus like he e is li le consensus on he de ini ion o explainabili y o i s axonomy, he e is
s ill li le consensus on how i should be e alua ed and measu ed. As Sil a e al. s a e in [37], he
e icacy o an explana ion is closely ela ed o i s abili y o con ince he a ge audience, making
i suscep ible o a my iad o subjec i e and in angible ac o s, such as he audience’s willingness
o us an explana ion o he audience’s backg ound knowledge. Despi e his subjec i e na u e,
46 Explainable Machine Lea ning: Li e a u e Re iew
he con ibu ion o each ea u e o ha p edic ion, simila ly o wha PDPs and ALE plo s do o
all ins ances. This me hod compu es Shapley alues, an algo i hm om coali ional game heo y,
whe e i is assumed ha each ea u e is a ”playe ” in a game whe e he p edic ion is he payou ,
and ha dis ibu es his payou ai ly among ea u es [69].
SHAP allows isualisa ion o ea u e impo ance o a single p edic ion by means o se e al
plo s: explana ion o ce plo s, whe e i is shown how much each ea u e ei he inc eases o de-
c eases he p edic ion; ea u e impo ance plo s, whe e he ea u es a e o de ed by dec easing
impo ance and compu ed ac oss all da a; summa y plo s, ha exhibi he ea u e impo ance ( he
Shapley alue) o a ea u e and a speci ic ins ance; and ea u e dependence plo s, whe e o each
ins ance i associa es he co esponding ea u e and Shapley alues.
Excep o MMD-C i ic, all hese model-agnos ic me hods, ei he global o local, explain mod-
els o p edic ions, espec i ely, by compu ing some kind o ea u e impo ance. They depend on
isualisa ion ools o allow in e p e a ion, which limi s hese models o less han h ee ea u es,
o he wise he in e ac ions s a o become oo di icul o humans o unde s and. The eade is
sugges ed o [69] o mo e model-agnos ic me hods.
Al hough model-agnos ic me hods ha e he ad an age o being used wi h any ML model,
wi hou he need o e ain he model, hey do no le e age in insic p ope ies o speci ic me hods,
which hinde s hei explana ion capabili y.
3.8 Model-Speci ic Me hods: Explainabili y o DNNs
The high p edic i e capaci y and complexi y o DNNs ha e made hem he objec o s udy
o mos explainabili y me hods. The e o e, and since his wo k ocuses on CNNs, his sec ion
is dedica ed o me hods speci ic o DNNs. The e iewed me hods a e di ided acco ding o he
Me hodology axonomy p esen ed in sec ion 3.5.6: ne wo k p ocessing, ne wo k ep esen a ion
and explana ion-p oducing sys ems.
3.8.1 Explaining Ne wo k P ocessing
DNNs ha e millions o ope a ions, which means ha explaining ne wo k p ocessing in ol es
inding ways o educing he complexi y o all hese ope a ions [45]. Saliency mapping, he c e-
a ion o a saliency map ha highligh s small po ions o he compu a ions which a e mos ele an ,
is pe haps he mos common ype o me hod when one alks abou explainable DNNs.
As men ioned in sec ion 3.6.3, he occlusion me hod [73], a b u e o ce sensi i i y analysis
ha wo ks by occluding pa s o he image and eeding hem o a ne wo k, in o de o ind ou
which pa s in luence he mos he ou pu , can be used as a benchma k o o he saliency me hods.
I one has access o pa ame e s o he DNN, saliency maps can be c ea ed by di ec ly compu ing
he inpu g adien [95].
Saliency mapping me hods a e g adien -based and he main idea behind hem is ha he g adi-
en e lec s he impo ance/ ele ance o inpu ea u es o he p edic ions: he da a poin x
x
xis seen as
3.8 Model-Speci ic Me hods: Explainabili y o DNNs 47
a sum o ea u es (xi)d
i=1and each one is assigned a ele ance sco e,Ri, de e mining how ele-
an each one is o explaining he indi idual decision. This is done by mapping back o he inpu
domain a p edic ion o a ce ain image, c ea ing a hea map, which is he explana ion. The a -
ious exis ing me hods simply cons i u e al e na i es o p oducing he a o emen ioned ele ance
sco es [43], which in ol es p opaga ing di e en quan i ies o he han g adien s [45], since such
de i a i es can miss impo an in o ma ion ha lows in he ne wo k.
These me hods we e also implemen ed in he con ex o [96], "iNN es iga e neu al ne wo ks!",
a lib a y de eloped in o de o p o ide a common in e ace and ou -o - he-box implemen a ion o
di e en analysis me hods, allowing o a sys ema ic compa ison be ween hem. Some o he
ob ained esul s a e shown in igu e 3.12, whe e he analysis was conduc ed on ImageNe [97].
Each column shows esul s om di e en me hods and each ow co esponds o one inpu da a
image. Fo each inpu , he g ound u h and p edic ed labels a e also shown on he le , as well as
he model’s so max ou pu (p ob) and i s logi ou pu be o e he so max laye (logi ). I is wo h
no ing ha all analyses ook in o conside a ion only he logi ou pu .
Figu e 3.12: Resul s ob ained wi h di e en me hods ia he iNN es iga e lib a y. Ex ac ed om
h ps://gi hub.com/albe max/inn es iga e (Accessed on 12-02-2019).
Acco ding o Kinde mans e al. [98], saliency me hods can be di ided in o unc ion, signal
and a ibu ion isualisa ion. These g oups all p esen di e en bu complemen ing in o ma ion
abou he ne wo k, and a e e iewed in he ollowing subsec ions.
Func ion This ca ego y aims a explaining he unc ion implemen ed by he model. In DNNs,
he unc ion is ep esen ed by he ou pu neu on ha encodes a ce ain concep [43]. So, ”...ex-
plaining he unc ion in inpu space co esponds o desc ibing he ope a ions he model uses o
ex ac y om x”. Howe e , in DNNs, which a e highly nonlinea , i is only possible o app oxi-
ma e he complex unc ion (x)[98]. The me hods ha will be desc ibed below cons i u e di e en
ways o making ha app oxima ion.
48 Explainable Machine Lea ning: Li e a u e Re iew
Sensi i i y Analysis es ima es how mo ing along he model’s locally e alua ed g adien di-
ec ion in inpu space in luences y[43]. In his case, he mos ele an ea u es a e hose o which
he ou pu is mos sensi i e. The echnique is easy o implemen conside ing ha he g adien can
be compu ed using backp opaga ion du ing he DNN’s aining. Howe e , i s esul s a e no g ea ,
as can be seen in igu e 3.12. This beha iou can be explained i one emembe s ha sensi i i y
analysis only akes in o accoun he local a ia ion o (x)ins ead o i s alue, which means ha
sensi i i y analysis only p o ides an explana ion o he unc ion’s local slope.
Simple Taylo Decomposi ion is based on decomposing (x)as a sum o ele ance sco es
using a Taylo expansion [43]. I one igno es he highe o de e ms o his expansion, which can
be done o some piecewise linea unc ions, he ele ance sco e becomes a p oduc o sensi i i y
and saliency, i.e., a ea u e is ele an i i posi i ely a ec s he model’s p edic ion and i i is
p esen in he inpu da a. Simple Taylo Decomposi ion pe o ms be e han Sensi i i y Analysis,
bu is cha ac e ised by a la ge amoun o nega i e ele ance.
Smoo hG ad was o iginally p oposed by Smilko e al. in [79] as a way o isually sha pen
g adien -based sensi i i y maps. The co e idea can be simply pu as ” emo ing noise by adding
noise”: i s ake an image as inpu da a, hen sample simila images by adding noise o he o iginal
ins ance and inally compu e he a e age o he sensi i i y maps o each sampled ins ance. Since
di ec ly compu ing a local a e age in a high-dimensional inpu space is no easible, he au ho s
use a s ochas ic app oxima ion. As can also be seen in igu e 3.12, Smoo hG ad sligh ly imp o es
he saliency map’s isualisa ion, smoo hing some o i s noise.
Signal Acco ding o [98], a signal de ec ed by a neu al ne wo k is he componen o he inpu
da a ha caused he ne wo k’s ac i a ions. The me hods desc ibed below a emp o isualise his
signal. Speci ically, Guided BackP op and DeCon Ne use he same algo i hm as he saliency
map, bu use di e en app oaches when dealing wi h he ec i ie neu on. Pa e nNe eme ges as
an imp o emen o e he o he wo app oaches.
The DeCon Ne was p oposed by Zeile e al. in [73] and cons i u es a ”...way o map hese
ac i i ies back o he inpu pixel space, showing wha inpu pa e n o iginally caused a gi en
ac i a ion in he ea u e maps.”. In he sen ence, he au ho s e e o ” hese ac i i ies” as he
ea u e ac i i ies in he in e media e laye s o CNNs. As explained by hem, his a chi ec u e
consis s o a e e se CNN, ha uses he same building blocks bu in e e se di ec ion, mapping
ea u es o pixels. The e o e, examining a CNN in ol es a aching a DeCon Ne o each laye .
As a esul , he ob ained econs uc ion om an ac i a ion esembles a small piece o he o iginal
image, weigh ed acco ding o hei con ibu ion o he gi en ea u e ac i a ion. The au ho s also
no e ha ”.. hese p ojec ions a e no samples om he model...” [73]. The esul s o applying
his me hod a e also shown in igu e 3.12. Al hough i is no depic ed, esul s ob ained in [73]
show ha each laye ’s p ojec ions ep esen he hie a chical na u e o he ne wo k ea u es: laye
2 de ec s co ne s and edges, laye 3 cap u es mo e complex s uc u es, such as simila ex u es and
pa e ns, and he highe laye s a e mo e speci ic o each class [73].
3.8 Model-Speci ic Me hods: Explainabili y o DNNs 49
Guided BackP op was p oposed by Sp ingenbe g e al. in [99] and cons i u es a a ia ion
o he DeCon Ne p e iously desc ibed, solely in he way bo h me hods handle backp opaga ion
h ough he ReLUs.
Pa e nNe was p oposed in [98] by Kinde mans e al. and eme ges as an imp o emen upon
he DeCon Ne and GuidedBackP op isualisa ions, a e he au ho s disco e ed ha o hese
app oaches he di ec ion o he il e s did no coincide wi h he di ec ion o he signal. In he pape ,
he inpu da ax
x
xis modelled as he sum o a signal componen ,s
s
s, and a dis ac o componen , d
d
d. The
signal con ibu es o he ou pu , whe eas he dis ac o does no . The e o e, he goal is o es ima e
as bes as possible he signal. Pa e nNe in pa icula makes use o a wo-componen signal
es ima o , ailo ed o ReLUs. Hence, his app oach ”...yields a laye -wise back-p ojec ion o he
es ima ed signal o inpu space.”, whe e he signal es ima o is app oxima ed by supe posi ion o
neu on-wise nonlinea wo-componen es ima o s in each laye . As can be seen in igu e 3.12, his
leads o a isual imp o emen o he ac i a ions compa ed o p e iously desc ibed me hods.
A ibu ion Ano he app oach o isualising neu al ne wo ks in ol es a ibu ion isualisa ion.
As s a ed in [98], an a ibu ion ep esen s ”...how much he signal dimensions con ibu e o he
ou pu h ough he laye s...”. Se e al echniques we e p oposed o allow a ibu ion isualisa-
ion, such as simply mul iplying he Inpu by he G adien [100], Laye -Wise Rele ance P opa-
ga ion (LRP) and Deep Taylo Decomposi ion (DTD). The concep o a ibu ion is e e ed o as
ele ance in LRP. Pa e nA ibu ion eme ges as an imp o emen o e hese app oaches, jus as
Pa e nNe was an imp o emen o e DeCon Ne and GuidedBackP op.
LRP o iginally p oposed in [74] and la e e iewed in [43], is a backwa d p opaga ion ech-
nique ha , unlike DeCon Ne and GuidedBackP op, possesses a conse a ion p ope y, whe e he
sha e o (x) ecei ed by each neu on is equally edis ibu ed o lowe laye neu ons, om he
ne wo k ou pu o he inpu . Jus as in p e ious me hods, he e is a i s phase whe e a s anda d
o wa d pass is applied and he ac i a ions a each laye a e sa ed. The second phase is basically
a backp opaga ion s ep, whe e speci ic p opaga ion ules a e applied. LRP applies one ule ha ,
simply pu , ei he edis ibu es ele ance o ”coun e - ele ance” o lowe -laye neu ons in p opo -
ion o hei exci a o y o inhibi o y e ec , espec i ely, on he ac i a ion o he neu on whose
ele ance is being dis ibu ed. This ele ance edis ibu ion ule can be unde s ood by looking a
igu e 3.13, whe e posi i e ele ance is shown in ed and nega i e ele ance in blue.
Figu e 3.13: Rele ance dis ibu ion p ocess in LRP o one neu on, o di e en αand β alues.
Ex ac ed om [43].
50 Explainable Machine Lea ning: Li e a u e Re iew
DTD When α=1 and β=0, as shown in igu e 3.13, he LRP is educed o a DTD [101]: he
LRP-α1β0 ule applied o a gi en laye can be seen as he compu a ion o a Taylo decomposi ion
o he ele ance a ha laye on o a lowe laye . The me hod’s name hen ”...a ises om he
i e a i e applica ion o Taylo decomposi ion om he op laye down o he inpu laye .” [43].
Va ious esul s o LRP and DTD applica ion a e also shown in igu e 3.12.
Deep-Li P oposed by Sh ikuma e al. [102], p oceeds simila ly o LRP in a backwa d
ashion. In his case, each uni is assigned an a ibu ion ha ep esen s he ela i e e ec o
he uni ac i a ed by an inpu compa ed o he ac i a ion by some e e ence inpu . The e e ence
alues a e compu ed by unning a o wa d pass using as inpu he e e ence image, which is usually
chosen o be ze o [103].
Pa e nA ibu ion was p oposed alongside Pa e nNe in [98]. I is an ex ension o DTD,
add essing one o he la e ’s main p oblems: he choice o oo poin used in he Taylo expansion.
Thus, Pa e nA ibu ion can be seen as a oo poin es ima o ha lea ns om da a.
In eg a ed G adien s was p oposed in [78] and combines echniques employed by G adien -
based me hods and LRP. This me hod simply consis s in accumula ing he compu ed g adien s a
all poin s along a s aigh line pa h om he baseline x
x
x0 o inpu x
x
x. Speci ically, he in eg a ed
g adien s a e de ined as he pa h in eg al o he g adien s along he s aigh line pa h om he
baseline x
x
x0 o he inpu x
x
x. This baseline inpu could be a black image o CNNs. In p ac ice, he
in eg a ed g adien s can be app oxima ed by a summa ion o g adien s a pa h poin s sepa a ed by
su icien ly small in e als, which can be done simply by compu ing he g adien s in a loop o e
he se o inpu s. This me hod can, he e o e, be easily inco po a ed in o any DNN a chi ec u e.
The esul s o his implemen a ion in he con ex o he iNN es iga e amewo k a e shown in
igu e 3.12.
Figu e 3.14 p esen s a di ision o he e iewed me hods, as p oposed by [104]. DeepLIFT,
G adien ×Inpu o In eg a ed G adien s a e conside ed G adien and Decomposi ion-based me h-
ods, while LRP is essen ially a Decomposi ion me hod, om which many o he o he me hods
can be de i ed. DeCon Ne and Guided BackP op a e decon olu ion me hods by na u e, as hey
in e he no mal ope a ions o CNNs, in o de o p oduce isual explana ions.
As demons a ed by he numbe o me hods b ie ly e iewed, saliency me hods cons i u e a
la ge pe cen age o he esea ch in his ield. Many al e na i es we e p oposed, bu mos o hese
me hods sha e a lo o hei co e ideas, as was concluded in [90]. All e iewed me hods seem
o p oduce isual appealing explana ions (see igu e 3.12). Howe e , elying solely on isual
inspec ion can be misleading [105]. In his wo k, ”Sani y Checks o Saliency Maps” he au ho s
ound ou ha some me hods, e.g. GuidedBackP op a e independen o bo h he da a and he model
pa ame e s. The au ho s base hei indings in wo es s, he model pa ame e andomisa ion es
and he da a andomisa ion es .
Wo k by Kinde mans e al. [106], ”The (un) eliabili y o saliency me hods”, also ound ha
adding a cons an shi o he inpu da a, which does no a ec he model, causes se e al me hods
o inco ec ly a ibu e, mainly G adien ×Inpu and In eg a ed G adien s.
3.8 Model-Speci ic Me hods: Explainabili y o DNNs 51
Al hough saliency maps a e powe ul ools o gain in ui ion abou he way he ne wo ks p o-
cess he da a, some ail basic es s like he ones men ioned abo e. The e o e, de e mining whe e
hese me hods ail and ensu ing ha hey consis en ly gua an ee eliabili y o all possible ans-
o ma ions is impe a i e o hese me hods o be applied in c i ical a eas, such as medicine [106].
The e is also a p essing need o de eloping es s ha iden i y hese p oblems and o benchma k-
ing and compa ing he a ious exis ing me hods.
Figu e 3.14: Ca ego isa ion o di e en saliency maps. Ex ac ed om [104].
3.8.2 Explaining Ne wo k Rep esen a ions
Despi e he eno mous amoun o ope a ions in a DNN, hese a e in e nally o ganised in o
smalle subs uc u es. Explaining DNN ep esen a ions ocuses on unde s anding he ole and
s uc u e o da a lowing h ough he ne wo k [45]. This classi ica ion u he subdi ides me hods
acco ding o he g anula i y examined: by laye , by uni o by ec o .
As men ioned in sec ion 3.5.6, when me hods explain DNNs’ ep esen a ions by laye , he
in o ma ion lowing h ough one laye is conside ed all oge he . T ans e lea ning, he echnique
o using laye s om one ne wo k o sol e o he p oblem, i s in o his ca ego y and was p oposed
by Raza ian e al. [107]. I is nowadays one o he mos popula and use ul echniques o
aining DNNs. The au ho s ound ou ha he ou pu o an hidden laye o a ne wo k ained o
classi y objec s on ImageNe p oduced a ea u e ec o ha could be di ec ly eused o sol e o he
p oblems, such as ine-g ained classi ica ion o species o a ce ain animal [45].
Explaining DNN ep esen a ions by uni can be done by c ea ing isualisa ions o he inpu
pa e ns ha maximise he esponse o a single uni (quali a i e app oach) o by es ing he abili y
o a single uni , ei he a neu on o a con olu ional il e , o sol e a ans e p oblem (quan i a i e
app oach) [45]. These isualisa ions can be c ea ed ei he by op imising an image ia g adien
52 Explainable Machine Lea ning: Li e a u e Re iew
descen [95], by sampling images ha maximise ac i a ions [108] o by aining a gene a i e
ne wo k o c ea e such images [109]. An example o such isualisa ions can be ound in igu e
3.15, whe e he class speci ic spa ial suppo o he class Goose is isualised. Rega ding he
quan i a i e app oach, me hods like Ne wo k Dissec ion [80] measu e he abili y o indi idual
uni s o iden i y objec s, pa s, colou s and ex u es no p esen in he aining se , hus allowing he
cha ac e isa ion o he kind o s uc u es p esen in each uni o he ne wo k [45]. Ne 2Vec [110] is
ano he example o a me hod ha quan i ies how concep s a e encoded by il e s in DNNs, showing
ha mul iple il e s a e equi ed o encode a concep and o en il e s a e no concep -speci ic and
help encode mul iple concep s.
Figu e 3.15: Class saliency isualisa ion o he Goose class. Ex ac ed om [95].
P uning o ne wo ks [111] is ano he way o unde s and he ole o neu ons: he au ho s
showed ha la ge ne wo ks con ain smalle ne wo ks wi h ini ialisa ions conduci e o op imi-
sa ion, which means ha he e exis aining s a egies capable o sol ing he same p oblem, bu
wi h much smalle ne wo ks ha may be mo e explainable [45].
A e iew o me hods ha ocus on unde s anding uni ep esen a ions can be ound in [112].
Finally, one can also cha ac e ise o he di ec ions in he ep esen a ion ec o space o med by
linea combina ions o indi idual uni s [45]. The mos amous and ecen app oach in his ca ego y
is Tes ing wi h Concep Ac i a ion Vec o s (TCAV) [81]. Figu e 3.16 shows how TCAV explains
a DNN’s in e nal s a e in e ms o human- iendly concep s. TCAV aims a gi ing quan i a i e
explana ions: how much one concep was impo an o a p edic ion, e en i ha concep was no
pa o he aining. In o de o be applicable, his me hod needs:
•(a) a se o examples o a concep , e.g. ”s iped”, and some andom examples
•(b) labelled ins ances o he class being s udied, e.g. zeb a
•(c) a ained ne wo k
Ha ing his, TCAV quan i ies he model’s sensi i i y o he concep o ha class eso ing o
Concep Ac i a ion Vec o s, he ed a ow in igu e 3.16 (d), ha a e lea ned by aining a linea
classi ie o dis inguish be ween ac i a ions p oduced by he examples o he concep and ac i a-
ions in any laye . The CAV is he ec o o hogonal o he decision bounda y. In o de o quan i y
he concep ual sensi i i y o he class o in e es , TCAV uses he di ec ional de i a i e (e). This
me hod is a pos -model me hod, since i does no equi e he ne wo k o be e ained. Howe e ,
3.8 Model-Speci ic Me hods: Explainabili y o DNNs 53
CAVs can also be used o so images wi h espec o hei ela ion o he concep , making his an
in e es ing explo a o y analysis ool o be e unde s and and spo biases in he aining da a.
Figu e 3.16: TCAV explana ion pipeline. Gi en a use -de ined se o examples o a concep (e.g.
”s iped”) and andom examples (a), plus labelled da a examples o he class being s udied (e.g.
”zeb a”) (b), and a ained ne wo k (c), TCAV quan i ies he sensi i i y o he model o he concep
o he speci ied class. This quan i ica ion is done h ough CAVs, which a e lea ned by a linea
classi ie ha dis inguishes he ac i a ions p oduced by he examples o he concep and ac i a ions
in any laye (d). The CAV is he ec o o hogonal o he decision bounda y ( ed a ow in d). The
concep ual sensi i i y is hen calcula ed by he di ec ional de i a i e (e). Ex ac ed om [81].
3.8.3 Explana ion-P oducing Sys ems
The las ca ego y in he Me hodology axonomy is Explana ion-P oducing Sys ems. The goal
o hese me hods is o c ea e ne wo ks ha a e designed om he beginning o be easie o explain.
Th ee ways o doing his ha e been p esen ed in he e iewed li e a u e: ne wo ks can be ained
wi h explici a en ion, can be ained o lea n disen angled ep esen a ions o ained o gene a e
explana ions [45].
A en ion-based ne wo ks lea n unc ions ha weigh o e inpu s o ea u es o s ee ha in-
o ma ion owa ds o he pa s o he ne wo k. These ne wo ks we e ini ially employed o na u al
language ansla ion o allow he app op ia e p ocessing o wo ds in non-sequen ial o de [113].
A en ion ne wo ks ha e also been used in ine-g ained image classi ica ion [114], isual ques ion
answe ing [115] and image cap ioning [116]. Al hough modules ha con ol a en ion a e no
ained explici ly o p oduce explana ions, hey e eal a map o he in o ma ion ha lows h ough
he ne wo k, which can se e as a o m o explana ion [45].
Disen angled ep esen a ions a e ep esen a ions ha ha e indi idual dimensions, desc ibing
meaning ul and independen ac o s o a ia ion. Lea ning disen angled ep esen a ions, o he
p oblem o sepa a ing la en ac o s, can be used o c ea e in e p e able CNNs, whose indi idual
uni s de ec cohe en meaning ul pa e ns ins ead o he no mal di icul - o-in e p e mix u es o
pa e ns, by means o special loss unc ions [83].
54 Explainable Machine Lea ning: Li e a u e Re iew
Figu e 3.17: In e p e able s O dina y CNNs. The op 4 ows co espond o he isualisa ion o
il e s in op con olu ional laye s in in e p e able CNNs, while he bo om 2 ows co espond o
hose same il e s in o dina y CNNs. In e p e able CNNs ocus mo e on he animals’ heads, while
o dina y CNNs ocus on o he pa s o he images o classi y he animals. The au ho s ound
ha in e p e able CNNs usually encoded head pa e ns o animals in he op con olu ional laye s.
Ex ac ed om [83].
Finally, ne wo ks can be ained o explici ly gene a e explana ions. Figu e 3.18 ep esen s a
ypical block diag am o explana ion gene a ion ne wo ks. The black-box ML model is accompa-
nied by an Explana ion Me hod ha p oduces an explana ion based on he p edic ion. Explana ion
gene a ion is s ill a e y unexplo ed ield, ha ing been applied o isual ques ion answe ing [84]
and ine-g ained classi ica ion [117]. In bo h hese wo ks, he sys ems gene a e a ”because” sen-
ence ha explains he decision in na u al language [45]. Howe e , as a as his li e a u e e iew
wen , he e a e no sys ems ha explici ly gene a e isual explana ions, and ha is whe e his wo k
is inse ed, as will be u he de ailed in chap e 4.
Figu e 3.18: Block diag am o a ypical explana ion-p oducing sys em. The black-box model is
accompanied by an Explana ion Me hod ha p oduces an explana ion alongside he p edic ion.
Ex ac ed om [47].
3.9 In insically In e p e able Models 55
3.9 In insically In e p e able Models
This sec ion b ie ly co e s some models ha he ML communi y conside s in insically in e -
p e able. As men ioned in sec ion 3.5, hese a e in-model and model-speci ic by de ini ion. These
models a e global on a modula le el: hei pa ame e s and ea u es a e unde s andable o humans
and allow one o pe o m manually he same calcula ions hese models do, in o de o each he
same ou pu s.
In e p e able models can be cha ac e ised by he ollowing p ope ies, ha help inc ease hei
unde s andabili y [69]:
•Linea i y - a model is linea i he mapping be ween ea u es and a ge alues is linea .
•Mono onici y - he ela ionship be ween an inpu ea u e and he a ge ollows he same
di ec ion o e he en i e ea u e space. When an inpu ea u e inc eases (o dec eases) i
always leads o an inc ease o a dec ease in he a ge alue. This cons ain can be en o ced
in any model o inc ease i s in e p e abili y, simila ly o wha Sil a e al. did in [37].
•Spa si y - i imp o es unde s andabili y, because he less non-ze o alued ea u es, he easie
i is o a human o assimila e all o hem.
•In e ac ion - he abili y o au oma ically include in e ac ions be ween ea u es. In e ac ion
can also be included manually in any ype o model h ough ea u e enginee ing [47]. Al-
hough i usually imp o es p edic i e pe o mance, i oo many o oo complex in e ac ions
a e p esen , in e p e abili y will su e .
The mos widely known in e p e able models a e shown in able 3.1, whe e hey a e compa ed
wi h ega ds o linea i y, mono onici y and ea u e in e ac ion, as well as o wha asks hey can be
used in.
Concisely, Linea Reg ession models p edic he a ge as a weigh ed sum o he ea u e in-
pu s. These models, as he name sugges s, canno be applied o classi ica ion asks, so Logis ic
Reg ession models ex end on Linea models by using he logis ic unc ion o squeeze he ou pu
o a linea equa ion be ween 0 and 1 [69]. O he ex en ions o linea models exis , being wo h
no ing Gene alised Linea Models (GLMs) and Gene alised Addi i e Models (GAMs).
Linea and Logis ic eg ession models ail when he ela ionship be ween ea u es and ou come
is nonlinea o when he e is ea u e in e ac ion. To o e come hese d awbacks, one can eso o
Decision T ees; he da a is spli mul iple imes acco ding o ce ain h esholds in he ea u es,
c ea ing subse s o said da a, wi h each ins ance belonging o one subse [69].
Finally, he Naï e Bayes classi ie is a p obabilis ic classi ie based on he Bayes’ heo em,
along wi h assump ions on he independence be ween he ea u es.
62 P oposed Join A chi ec u e
egions highligh ed by he explaine and o disca d egions whe e he explana ion o he class
being p edic ed is no p esen , as will be u he de ailed in sec ion 5.3. A e he 4 con olu ion-
con olu ion-pooling s ages a global pooling laye is in oduced o la en he esul ing ea u e
maps, in o de o eed hem o he dense laye s. Finally, he ne wo k ends in wo 1×1×128 ully
connec ed laye s wi h sigmoid ac i a ion unc ions and 30% d opou a e, and 1 ully connec ed
laye wi h a so max ac i a ion, o allow mul iclass classi ica ion.
Howe e , i is wo h men ioning ha any classi ie can be used, p o ided ha he co ec
connec ions a e in oduced. In ac , he expe imen s pe o med on he ce ical cance da ase
( e e o chap e s 5and 6) include he ResNe -50 [33] as a classi ie , modi ied only o in oduce
he a o emen ioned connec ions. As he p oposed a chi ec u e aims a explaining he classi ie ’s
decisions, he classi ie should be chosen i s , depending on he classi ica ion p oblem and he
a ailable da a; he explaine mus adap o he classi ie and no he o he way a ound.
4.4 Explaine
The explaine ( op ow in igu e 4.2) akes as inpu an RGB image om he chosen da ase
and also ou pu s an RGB image, he isual explana ion, ˆz, wi h he same spa ial dimensions as he
inpu image. I consis s o a con olu ional encode -decode ne wo k, based on U-Ne [35].
The downsampling pa h, o encode , is a simple con olu ion-con olu ion-pooling scheme e-
pea ed 3 imes, simila o he classi ie . Each con olu ional laye uses 3×3 ke nels and is ollowed
by a ReLU ac i a ion laye . All pooling laye s use max pooling and a downsampling ac o o 2.
The i s con olu ion-con olu ion-pooling s age has 32 s acked il e s and 224×224 ea u e maps.
A e wa ds, he numbe o il e s inc eases as a powe o 2 acco ding o he s age le el.
The upsampling pa h, o decode , ollows a ” anspose con olu ion-con olu ion-con olu ion”
scheme, whe e he i s con olu ion ope a ion is applied o he pixel-wise sum o he p e ious
laye ’s ou pu (ou pu o he anspose con olu ional laye ) wi h he co esponding con olu ional
laye in he downsampling pa h. These connec ions allow o he subsequen laye o lea n a mo e
p ecise ou pu . The whole decode has 3 ” anspose con olu ion-con olu ion-con olu ion” s ages,
ending in a con olu ional laye wi h a 1 ×1 ke nel, used o keep he spa ial dimension space
(heigh and wid h), bu o educe he il e space dimensionali y, i.e., educe he numbe o il e s.
In his speci ic case, he numbe o il e s is educed om 32 o 1. This las con olu ional laye
has a linea ac i a ion and is ollowed by a ba ch no malisa ion laye wi h an hype bolic angen
ac i a ion. Con e sely o wha is done in he downsampling pa h, he numbe o il e s s a s a
128 and dec eases as a powe o 2 wi h each s age. On he o he hand, he spa ial dimensions
inc ease in he same way.
4.5 T aining 63
4.5 T aining
T aining his a chi ec u e in ol es he h ee phases desc ibed below and depic ed in igu e 4.3.
T aining
Phase 1
whi e
image
Classifie
Explaine
Classifie
Explaine
T aining
Phase 2
whi e
image
T aining
Phase 3
Classifie
Explaine
ainable
non- ainable/ ozen
Figu e 4.3: T aining p ocedu e o he p oposed a chi ec u e. In he i s phase only he classi ie
is ained, keeping he explaine ozen. A e wa ds, he p ocess is in e ed. In he las phase, he
whole a chi ec u e is ine uned end- o-end.
Fi s , in phase 1, only he classi ie is ained, aking a ”whi e image”, i.e., a 2D a ay illed
wi h ones, as explana ion, while he explaine ’s laye s emain ozen, i.e., unable o be upda ed.
This means ha he ini ial explana ion is he whole image and ha he classi ie is no aking he
explaine ’s ou pu in o accoun , as he explaine is no ained ye . As such, eeding he classi ie
a ”whi e image” ende s he mul iplica ions in oduced by he explaine (pu ple a ows in igu e
4.2) i ele an . I is impe a i e ha a he end o phase 1 he classi ie emain somewha uns a-
ble, i.e., ha i s loss does no pla eau, so ha in phase 3 he classi ie can s ill lea n om he
explaine ’s ou pu . O he wise, in phase 3 he classi ie would no upda e i s pa ame e s wi h he
new in o ma ion p o ided by he now ained explaine .
A e wa ds, in phase 2, he p ocess is e e sed: he classi ie is ozen and he explaine is he
only module being upda ed by backp opaga ion. The classi ie s ill akes as inpu a ”whi e image”,
so ha i ou pu s he same p edic ions as i did in phase 1, gua an eeing ha , a his poin , bo h ex-
plaine and classi ie a e s ill ained independen ly. I a his poin he classi ie had al eady made
use o he explaine ’s ou pu , i would be ocusing on in o ma ion p o ided by an explaine du ing
i s aining p ocess. In u n, based on his e oneous in o ma ion, he classi ie would p edic du -
ing he o wa d pass and his p edic ion would al e he loss unc ion acco dingly, in luencing he
e o p opaga ion in he explaine , al hough no in luencing he classi ie ’s pa ame e s. Then, in
he nex o wa d pass, he classi ie would again use he explaine ’s ou pu , p oduced by an inad-
equa ely upda ed explaine , and again p edic acco dingly, backp opaga ing his new e o o bo h
explaine and classi ie , upda ing hem. This p ocess would epea , culmina ing in an inco ec
aining o bo h explaine and classi ie , as a esul om his accumula ion e ec .
Finally, in phase 3, he whole a chi ec u e is ine- uned end o end. The classi ie is lea ning
wi h he new in o ma ion p o ided by he explaine , h ough he mul iplica ions o he explaine ’s
64 P oposed Join A chi ec u e
ou pu a e e y classi ie ’s s age, as explained in sec ion 4.3. As he classi ie is aining, ideally
he global loss unc ion is dec easing, imp o ing he explaine as well. As he explaine is ain-
ing, ideally he global loss unc ion also dec eases, and he classi ie imp o es. So, in he end,
bo h explaine and classi ie imp o e, ui o his dynamic in e ac ion be ween he wo and i s
espec i e loss unc ions.
4.5.1 Loss
The whole aining p ocess desc ibed abo e eso s o he end- o-end loss unc ion de ined in
equa ion 4.1. This loss cons i u es a weigh ed sum o he in luence o bo h classi ie and explaine
(Lclass and Lexpl, espec i ely), as ealised by he αhype pa ame e .
L=α×Lclass +(1−α)×Lexpl (4.1)
In gene al, α>0.5, because he co ec classi ica ion is an indispensable pa o he sys em,
i di ec ly a ec s he classi ie and indi ec ly a ec s he explaine . In ac , h ough his join loss,
he explaine is always in luenced by he classi ie . Howe e , a sligh dec ease in classi ica ion
accu acy is ole able, as long as i means ha he sys em is also able o p o ide meaning ul ex-
plana ions. I α=1, hen he explaine is being ained wi hou any kind o egula isa ion o i s
ou pu s, being only indi ec ly in luenced by he classi ica ion componen ia he global loss.
The classi ica ion loss, Lclass, de ined in equa ion 4.2, is he commonly used ca ego ical c oss-
en opy loss, employed in he majo i y o mul iclass classi ica ion scena ios.
Lclass =−
N
∑
i=1
y>
i×log(ˆyi)(4.2)
whe e Nis he numbe o ins ances in he aining se , yiis he one-ho encoded column a ay o
he a ge class labels o ins ance iand ˆyia e he classi ie ’s so max p edic ions o ins ance i.
Rega ding he explana ion loss, wo di e en app oaches a e p oposed, an unsupe ised and a
weakly supe ised loss, u he de ailed in subsec ions 4.5.1.1 and 4.5.1.2, espec i ely.
4.5.1.1 Unsupe ised Explana ion Loss
As i would be e y ha d o supe ise he explaine ’s aining, no only because explana ions
a e use -, con ex - and domain-dependen , bu also because an explana ion cons i u es a highe
concep ual le el, he explaine is ained wi h an unsupe ised loss. This loss, ep esen ed in
equa ion 4.3, is also a weigh ed sum o ele an p ope ies one wan s o p omo e on he esul ing
explana ions, as ealised by he βhype pa ame e . These p ope ies, spa si y and con igui y, a e
imposed by means o wo egula isa ion e ms: he penalised l1 no m and o al a ia ion (Lspa si y
and Lcon igui y espec i ely) [125].
Lexpl_unsup =β×
N
∑
i=1
Lspa si y(ˆzi)+(1−β)×
N
∑
i=1
Lcon igui y(ˆzi)(4.3)
4.5 T aining 65
Lspa si y(ˆz) = 1
m×n∑
i,j
|ˆzi,j|(4.4)
Lcon igui y(ˆz) = 1
m×n∑
i,j
|ˆzi+1,j−ˆzi,j|+|ˆzi,j+1−ˆzi,j|(4.5)
whe e mand n ep esen he spa ial dimensions o he ea u e map ˆz.
The penalised l1 no m (equa ion 4.4), p omo es spa si y, as i sh inks less impo an ea u es o
ze o, pe o ming a kind o ea u e selec ion. This penal y wo ks as an explana ion budge , limi ing
he pe cen age o he inpu image ha can be conside ed an explana ion. The g ea e he penal y,
he mo e egula isa ion, and he e o e he less explana ion budge is a ailable.
The o al a ia ion ac o (equa ion 4.5) p omo es spa ial con igui y, i.e., encou ages smoo h-
ness and spa ial localisa ion o he ac i a ions o ˆz, by minimising he local spa ial ansi ions o ˆz.
This ac o was in oduced because ini ially ob ained esul s, u he de ailed in chap e 6, hin ed
a he need o spa se explana ions, bu also connec ed, i.e., explana ions ha connec seman ically
ela ed pa s o he images.
4.5.1.2 Weakly Supe ised Explana ion Loss
One way o impose s onge cons ain s on he p oduced explana ions wi hou needing o ully
supe ise hei gene a ion is o use a weakly supe ised s a egy, p o ided ha some ype o anno-
a ed da a is a ailable om ano he ask. In his case, masks om objec de ec ion o segmen a ion
asks a e needed. The aim is o d i e he explana ions no o ocus on egions ha a e ou side he
in e es egions; in he case o objec de ec ion anno a ions, he egions o in e es would be he
a eas inside he g ound- u h bounding boxes. The e o e, he weakly supe ised explana ion loss
is de ined as:
Lexpl_weakly =
∑
i,j
(1−zi,j)׈zi,j
∑
i,j
(1−zi,j)
(4.6)
whe e ˆzi,jis he alue o pixel (i,j)o ea u e map ˆzand zi,jis he alue o pixel (i,j)o mask
z. This mask is an image whe e egions o in e es a e ep esen ed by ones, and o he egions by
ze os.
This loss punishes explana ions ou side he egion o in e es , by compu ing he p oduc be-
ween he in e ed mask (1−zi,j) and he explana ion (ˆzi,j): in he egion o in e es , he expla-
na ion is mul iplied by ze os, and he explaine is no punished; ou side he egion o in e es , he
explana ion is mul iplied by ones, so he explaine will ha e o punish hese pixels, by lowe ing
hei alue om one, in o de o dec ease he explana ion loss.
The di iso minimises he impac o he size o he egions o in e es in he mask.
Al hough his app oach needs masks om o he asks, i elimina es one o he hype pa ame-
e s, g ea ly educing he ime o une he p oposed a chi ec u e.
66 P oposed Join A chi ec u e
Chap e 5
Expe imen al Me hodology
This chap e desc ibes he expe imen al me hodology employed du ing he cou se o his wo k.
The i s sec ion desc ibes he da ase s in which he p oposed a chi ec u e was es ed. The de el-
oped syn he ic da a gene a ion ool is de ailed, ollowed by he cha ac e isa ion o he eal da ase s
used, namely 16-class-ImageNe [126], he Cue Con lic da ase by Gei hos e al. [127] and he
Ce ical Cance da ase o he Na ional Ins i u e o Heal h (NIH) and he Na ional Cance Ins i u e
(NCI) [128].
A e wa ds, da a p ep ocessing s eps and hype pa ame e uning s a egies a e discussed and,
o close he chap e , a b ie Explaine -Classi ie Connec ion S udy is p esen ed.
5.1 Da ase s
5.1.1 Syn he ic Da ase s
Da a is he co e ing edien in e e y DL pipeline. Howe e , i is inc edibly di icul o come by
labelled da ase s and i is highly expensi e o collec and anno a e one. Mo eo e , he explainabil-
i y o a model o a pa icula ins ance is use - and domain-dependen . As such, a syn he ic da a
gene a ion ool wi h au oma ic anno a ion was de eloped1.
This ool gene a es images con aining simple polygons, like ci cles, iangles and squa es.
This way i is possible o analyse he p oposed ne wo k’s p oduced explana ions wi hou he need
o expe knowledge, since i can be conside ed common knowledge, easily unde s andable by
any human. Fu he mo e, a syn he ic da ase en ails o he ad an ages, such as he gene a ion o
as much images as needed, he de ini ion o he numbe o ins ances o each class, au oma ic
anno a ion o di e en p oblems anging om classi ica ion o de ec ion and he inclusion o
cus om cha ac e is ics like o e lap, occlusion, objec ype, objec colou , image dimensions, e c.
Each polygon is de ined by i s pand q ac o s, whe e pis he numbe o e ices and q ep-
esen s how hese e ices a e connec ed: he e ices a e placed along he ci cum e ence ha
ci cumsc ibes he polygon wi h an equally spaced angle gi en by 2π
p. Fo example, i qequals 3
hen he i s e ex connec s o he ou h and so on. On he o he hand, i pequals -1 hen he
1The code is publicly a ailable a h ps://gi hub.com/ic o/xML.gi .
67
68 Expe imen al Me hodology
polygon is a ci cle and he alue o qis igno ed. Each polygon is andomly placed in he image,
conside ing o e lap, occlusion and o a ion cons ain s. Fo each gene a ed image, anno a ions a e
c ea ed in Pascal VOC o ma , con aining in o ma ion ega ding image cha ac e is ics and also he
classi ica ion p oblem being sol ed. The a ge polygon is gi en, along wi h in o ma ion abou i s
p esence in he image. I he polygon is p esen in he image, hen bounding boxes a e p o ided o
each ins ance o he desi ed polygon. This way, one da ase can be used o simple bina y classi i-
ca ion (”Is he e a leas one iangle in his image?”), mul iclass classi ica ion (”Does his image
con ain 0, be ween 1 and 3 o mo e han 3 iangles?”) o de ec ion (”Whe e a e he iangles
loca ed in each image?”).
To make he ool easie o use and he gene a ed images as cus om as possible, a ious con ig-
u a ion pa ame e s can be se in a JSON con igu a ion ile. A lis o hese con igu able pa ame e s
can be ound in A.1 and an example o a gene a ed XML anno a ion ile in Pascal VOC o ma can
be ound in appendix A.2. The pseudo-code desc ip ion o he algo i hm o he syn he ic images’
gene a ion is p esen ed below.
Algo i hm 1: Syn he ic Image Gene a ion
Resul : RGB image and co esponding XML anno a ion ile in PASCAL VOC o ma
c ea e image wi h he de ined backg ound colou
c ea e anno a ion ile
andomly choose numbe o polygons (be ween min_n _shapes and max_n _shapes)
coun ←0
ies ←0
while (coun < n _shapes - 1) o ( ies > n _ ies) do
andomly choose p,q, x_o ig, y_o ig, ad and o a ion angle o he polygon
o e e y polygon in he image do
check i polygons a e o e lapping
end
i polygons a e NOT o e lapping hen
add he cen e coo dina es o he new polygon o he lis o polygons
d aw polygon by placing i s e ices along he ci cumsc ibing ci cum e ence
i d awn polygon is a ge polygon hen
append he polygon’s bounding box coo dina es o he co esponding XML
anno a ion ile
end
coun ←coun + 1
else
ies ← ies + 1
end
end
upda e XML anno a ion ile wi h numbe o a ge polygons d awn
5.1 Da ase s 69
Figu e 5.1 shows some examples o gene a ed images. P elimina y expe imen s showed ha
a da ase composed by images simila o hose o he ou h ow (d) cons i u ed an o e com-
plex classi ica ion p oblem. As such, only he i s h ee da ase s we e chosen o he es o he
expe imen s wi h syn he ic da a.
(a) Images om a simple da ase wi h colou cues
(b) Images om a simple da ase wi hou colou cues
(c) Images om a da ase wi h mul iple a ge polygons
wi hou colou cues
(d) Images om a da ase wi h mul iple a ge polygons and
5-poin ed s a s wi hou colou cues
Figu e 5.1: Example images o 4 gene a ed da ase s. Fo each ow, he le column illus a es
an example o he posi i e class, while he igh column illus a es he nega i e class. The a ge
polygons a e he iangles.
70 Expe imen al Me hodology
Finally, igu e 5.2 depic s how many ins ances he e a e in each class o he wo syn he ic
da ase s used. Despi e being possible o gene a e as much images as one needs, i was chosen o
ix he numbe o aining samples o 1000, as his is a a he simple bina y classi ica ion p oblem
ha should no need mo e da a, and also o educe he du a ion o he aining p ocess. I was also
decided ha he da ase s should p esen some deg ee o unbalance, as o mimic wha happens in
eal wo ld scena ios, mainly in medical con ex s, when usually he e a e many mo e ins ances o
heal hy pa ien s han o sick pa ien s. The e o e, bo h da ase s consis o 1000 224 ×224 RGB
images, each con aining a iangle ( a ge polygon) and a a iable numbe o ci cles. Fo each
image i is only aken in o accoun he p esence (posi i e class) o absence (nega i e class) o he
a ge polygon, hus making hese expe imen s bina y classi ica ion asks.
(a) (b)
Figu e 5.2: Simple da ase wi hou (le ) and wi h ( igh ) colou cues’ class dis ibu ion.
5.1.2 Real Da ase s
5.1.2.1 16-class-ImageNe
16-class-ImageNe was i s in oduced in he wo k o Gei hos e al., in which he au ho s
de elop an expe imen al pa adigm o compa e he gene alisa ion abili ies o human obse e s and
neu al ne wo ks [126]. This da ase is based on he ILSVRC 2012 da abase [129], in which mo e
han 1000 ine-g ained classes a e p esen . Howe e , such le el o g anula i y is no needed in
bo h he o iginal pape no in his wo k. This s ems om he ac ha humans end o ca ego ise
objec s in o ”en y-le el” ca ego ies, such as dog ins ead o Lab ado [126]. The e o e, he au ho s
mapped he ini ial 1000 classes in o 16 ”en y-le el” classes, using he Wo dNe hie a chy [130]:
ai plane, bicycle, boa , ca , chai , dog, keyboa d, o en, bea , bi d, bo le, ca , clock, elephan ,
kni e, uck. Figu e 5.3 shows he numbe o ins ances pe class. In o al, his da ase is composed
by 213555 RGB images, o which some examples a e shown in igu e 5.4. In ac , his da ase was
no di ec ly used in his wo k, bu i s desc ip ion is needed o he nex subsec ion.
5.1 Da ase s 71
Figu e 5.3: 16-class-ImageNe da ase class dis ibu ion.
(a) (b)
Figu e 5.4: Example images om 16-class-ImageNe .
5.1.2.2 Cue Con lic Da ase
This da ase was i s in oduced in [127], whe e he au ho s alida e ha ImageNe - ained
CNNs a e biased owa ds ex u e. In o de o es his hypo hesis, he au ho s p opose a cue
con lic expe imen in which s yle ans e is employed, in oducing ex u e in images. Figu e 5.5
ep esen s he esul s o one o he conduc ed expe imen s wi h a s anda d ResNe -50 a chi ec u e.
I can be seen ha he ne wo k is able o co ec ly classi y he ex u e and o iginal images, bu i
ails o iden i y he image in which he e is a ex u e-shape cue con lic .
This da ase 2con ains he same 16 classes men ioned in sec ion 5.1.2.1, wi h 80 images each,
making a o al o 1280 images a ailable in he da ase . Examples o such images a e shown in
igu e 5.6.
2publicly a ailable a h ps://gi hub.com/ gei hos/ ex u e- s-shape
78 Expe imen al Me hodology
(a)
(b)
(c)
Figu e 5.10: Resul ing explana ions ob ained on a syn he ic da ase wi h colou cues when con-
nec ing he explaine o he classi ie ’s (a) las , (b) i s and (c) e e y con olu ion-con olu ion-pool
s age. Fo each ow he o iginal inpu image is p esen ed, as well as he explaine ’s ou pu ed ex-
plana ion and he same explana ion a e applying a h eshold.
(a) (b) (c)
Figu e 5.11: E olu ion o he classi ie ’s aining and alida ion accu acy ob ained on a syn he ic
da ase wi h colou cues when connec ing he explaine o he classi ie ’s (a) las , (b) i s and (c)
e e y con olu ion-con olu ion-pool s age.
5.3 Explaine -Classi ie Connec ion S udy 79
(a)
(b)
(c)
Figu e 5.12: Resul ing explana ions ob ained on a syn he ic da ase wi hou colou cues when
connec ing he explaine o he classi ie ’s (a) las , (b) i s and (c) e e y con olu ion-con olu ion-
pool s age. Fo each ow he o iginal inpu image is p esen ed, as well as he explaine ’s ou pu ed
explana ion and he same explana ion a e applying a h eshold.
(a) (b) (c)
Figu e 5.13: E olu ion o he classi ie ’s aining and alida ion accu acy ob ained on a syn he ic
da ase wi hou colou cues when connec ing he explaine o he classi ie ’s (a) las , (b) i s and
(c) e e y con olu ion-con olu ion-pool s age.
80 Expe imen al Me hodology
Chap e 6
Resul s and Discussion
This chap e is dedica ed o he ob ained esul s on he unsupe ised and weakly supe ised
app oaches. I is u he di ided in o syn he ic and eal da ase s.
Rega ding he unsupe ised app oach on syn he ic da a, he esul s ob ained wi hou any eg-
ula isa ion on bo h syn he ic da ase s a e de ailed, ollowed by sec ions de o ed o he s udy o he
in luence o he αand βhype pa ame e s, as well as hei ecommended alues, and ending wi h
a compa ison wi h some o s a e-o - he-a me hods. Then, esul s o his app oach applied o he
cue con lic da ase and o he NIH-NCI ce ical cance da ase a e p esen ed.
The end o he chap e ocuses on he esul s ob ained wi h he weakly supe ised app oach
on bo h syn he ic and he NIH-NCI ce ical cance da ase s.
Fo all o he ollowing images i is wo h no ing ha he colou code anges om pu ple o
yellow, whe e yellow ep esen s highe pixel alues, as can be seen in igu e 6.1. The esul s a e
p esen ed in images con aining h ee columns, whe e he le column co esponds o he o iginal
image, he middle column o he ou pu ed explana ion and he igh column o he explana ion
a e an absolu e h eshold o 0.75 is applied.
Images om sec ions 6.1.1.1 h ough 6.1.1.3 we e pa o he alida ion se , since hese sec-
ions e e o hype pa ame e uning and e ec on he p oduced explana ions. All o he images a e
om es ing se s, a e he hype pa ame e uning p ocess desc ibed in sec ion 5.2.2.
Figu e 6.1: Colou map used in he explana ion esul s. Highe pixel alues a e ep esen ed by he
colou yellow, while lowe pixel alues co espond o he colou pu ple.
81
82 Resul s and Discussion
6.1 Unsupe ised App oach
6.1.1 Syn he ic Da ase s
6.1.1.1 Wi hou Regula isa ion (α=1)
Figu e 6.2 p esen s examples o he ob ained esul s o posi i e ins ances o simple syn he ic
da ase s wi h and wi hou colou cues. Bo h images a e he esul o aining he explaine wi hou
any kind o egula isa ion, i.e., wi h α=1. The alue o βis unimpo an , since i would be
mul iplied by 1−α, which in his case equals 0. Fo such simple da ase s, i is expec ed ha he
explana ion ocuses on he a ge polygon, he iangle, ende ing he es o he image as i ele an
o he p edic ed class. While on he da ase wi h colou cues (a) he esul ing explana ion consis s
only o he a ge polygon, as expec ed, in he sligh ly mo e complex da ase wi hou colou cues
(b) he explana ion becomes deg aded, as bo h iangle and ci cles a e highligh ed, which co obo-
a es he need o egula ising he explaine ’s ou pu . As s a ed in subsec ion 4.5.1.1, he penalised
l1 no m is in oduced, because i allows o he selec ion o he ele an pa s o he explana ion,
ensu ing ha only a small pa o he image is in ac he explana ion o he classi ie ’s decision,
i.e., ha he explana ion is spa se. Thus, his egula isa ion ensu es ha he explana ion is no only
in e p e able, bu also comple e, as desi ed.
(a) α=1.0β=0.0
(b) α=1.0β=0.0
Figu e 6.2: Resul s ob ained in an unsupe ised scena io wi hou any egula isa ion o he ex-
plaine , α=1.0, on syn he ic da ase s, wi h (a) and wi hou (b) colou cues.
6.1 Unsupe ised App oach 83
6.1.1.2 Wi h Regula isa ion: In luence o he αHype pa ame e
The esul o he expe imen s wi h di e en egula isa ion penal ies a e p esen ed in igu es
6.3 and 6.4; hese we e ob ained wi h α= [0.5,0.9,0.95,0.99999], on he da ase wi hou colou
cues, o posi i e and nega i e ins ances, espec i ely. In e e y expe imen , βwas kep cons an
a 1, so ha only he penalised l1 no m would ake e ec , and ha i would be possible o be e
e alua e he in luence o he αhype pa ame e .
Rega ding only he posi i e examples, igu e 6.3, in which a iangle is p esen , one can see ha
he mo e egula isa ion ( om (d) o (a)) he less pixels a e highligh ed, as was al eady o eseen
in subsec ion 4.5.1.1. Wi h α=0.5 no egion o he image is conside ed an explana ion, which
con adic s he in ui ion ha i he e is a iangle in he inpu image, hen i s pixels would cons i u e
he explana ion o i s p esence. This s ems om he ac ha his alue o α, mo e speci ically,
he alue o 1 −α, is oo high, limi ing he explana ion budge oo much. In ac , when his
budge inc eases (α=0.9), i.e., when a g ea e pe cen age o he image can be conside ed an
explana ion, he a ge polygon s a s being highligh ed and e en mo e so when α=0.95. In
ac , he mo e egula isa ion, he less edundancy: since he explana ion budge is igh e , he
explaine has o ind a way o educe i s explana ion o he absolu ely essen ial pixels; his is wha
happens in igu e 6.2 (b), whe e only wo sides o he iangle a e highligh ed, because when only
dis inguishing iangles om ci cles a single s aigh line is enough o iden i y he iangle om
he ci cles. Finally, when α=0.99999, he iangle is s ill highligh ed and he ci cles become
e en mo e nega i ely highligh ed. Howe e , inc easing αe en u he , app oxima es he case
wi hou any egula isa ion ( igu e 6.2 (b)), in which he explana ion budge is no cons aining he
explana ions enough, esul ing in bo h iangle and ci cles highligh ed.
(a) α=0.50 β=1.00
(b) α=0.90 β=1.00
84 Resul s and Discussion
(c) α=0.95 β=1.00
(d) α=0.99999 β=1.00
Figu e 6.3: Examples o posi i e ins ances ob ained in an unsupe ised scena io wi h di e en α
alues, α= [0.5,0.9,0.95,0.99999], on a syn he ic da ase wi hou colou cues.
Rega ding he nega i e examples, igu e 6.4, as expec ed, in e e y case no ci cles a e consid-
e ed explana ions, as hey do no explain he absence o he iangle. When α=0.5, simila ly
o wha happens wi h he posi i e ins ance, no egion o he image is conside ed an explana ion,
because he explana ion budge is oo igh . When αinc eases, he ci cles a e mo e and mo e
nega i ely highligh ed, which means ha he explaine is mo e con iden ha hose egions do no
cons i u e explana ions o he absence o he iangle. Also, he whole image excluding he ci -
cles is conside ed an explana ion: no e ha he backg ound colou o he images om he middle
columns is no da k pu ple, bu is g een and ligh blue, meaning ha i s pixel alues a e loca ed in
he middle o he colou map. This aligns wi h he in ui ion ha , when no iangle is p esen he
whole image excep he ci cles is he explana ion o why no iangle exis s in i . I also in e es ing
o no e ha wi h highe alues o α, he in e sec ion o he h ee ci cles esembles pa o a iangle
and, he e o e, is also highligh ed.
(a) α=0.50 β=1.00
6.1 Unsupe ised App oach 85
(b) α=0.90 β=1.00
(c) α=0.95 β=1.00
(d) α=0.99999 β=1.00
Figu e 6.4: Examples o nega i e ins ances ob ained in an unsupe ised scena io wi h di e en α
alues, α= [0.5,0.9,0.95,0.99999], on a syn he ic da ase wi hou colou cues.
6.1.1.3 Wi h Regula isa ion: In luence o he βHype pa ame e
Resul s o expe imen s wi h di e en β alues a e p esen ed in igu es 6.5 and 6.6. These
expe imen s we e conduc ed wi h α ixed o 0.9. This α alue was chosen o be e e alua e
he in luence o he βhype pa ame e , because lowe alues o α egula ise he explaine oo
much, causing he iangles o ”disappea ” ega dless o he alue o β. Highe α alues make he
di e ences be ween expe imen s wi h di e en β alues oo sub le o be easily isualised.
Rega ding posi i e ins ances, wi h lowe β alues ( igu es 6.5 (a) and (b)) he o al a ia ion
ac o domina es o e he spa si y cons ain , making he iangles isible. Wi h highe β alues
he spa si y cons ain excessi ely egula ises he explana ion, causing he iangles no o be high-
ligh ed. Howe e , i is in e es ing o no e ha wi h β=1.0 (c . igu e 6.3 (b)), he iangle is
sligh ly highligh ed, which migh indica e ha o alues o βbe ween ]0.5,1.0[, he in e ac ion
86 Resul s and Discussion
be ween he spa si y and he con igui y cons ain s migh be p ejudicial, causing he iangles no
o be highligh ed.
(a) α=0.90 β=0.00
(b) α=0.90 β=0.50
(c) α=0.90 β=0.75
(d) α=0.90 β=0.90
Figu e 6.5: Examples o posi i e ins ances ob ained in an unsupe ised scena io wi h di e en β
alues, β= [0.0,0.25,0.50,0.75,0.9]on a syn he ic da ase wi hou colou cues.
6.1 Unsupe ised App oach 87
Rega ding nega i e ins ances ( igu e 6.6), he same beha iou is e i ied: o lowe β alues
he in e sec ion be ween ci cles is highligh ed, while o highe β alues i is no . In e ms o
backg ound, in bo h posi i e and nega i e ins ances, i is be e nega i ely highligh ed by lowe β
alues, e en mo e so i βis close o 0.5.
(a) α=0.90 β=0.00
(b) α=0.90 β=0.50
(c) α=0.90 β=0.75
(d) α=0.90 β=0.90
Figu e 6.6: Examples o nega i e ins ances ob ained in an unsupe ised scena io wi h di e en β
alues, β= [0.0,0.25,0.50,0.75,0.9]on a syn he ic da ase wi hou colou cues.
94 Resul s and Discussion
6.2 Weakly Supe ised App oach
6.2.1 Syn he ic Da ase s
Figu e 6.12 shows an example o a posi i e ins ance and he co esponding explana ion using
he weakly supe ised app oach wi h α=0.99. Simila ly o wha happened in he unsupe ised
scena io, he iangle is highligh ed, which alida es his app oach (RQ1). Compa ing his
esul wi h he one ob ained in he unsupe ised scena io wi h he same α alue (c . igu e 6.7
(a)), one can obse e ha in his app oach he ci cles a e mo e nega i ely highligh ed, ui o he
punishmen ha he weakly supe ised loss imposes on non-in e es egions.
Figu e 6.12: Example o he ob ained esul s in a weakly supe ised scena io wi h α=0.99 on a
syn he ic da ase wi hou colou cues.
Rega ding he mul iple a ge s case, he main di e ence is ha he e all o he iangles’ edges
a e highligh ed, because he weakly supe ised loss only punishes egions ou side he in e es
zone, while in he unsupe ised case he loss cons ain s he numbe o pixels used in an explana-
ion.
On such simple oy da ase s he isual di e ences be ween he wo app oaches a e no
signi ican (RQ1). As such, and gi en he ac ha gene a ing bounding boxes and hei co -
esponding masks is au oma ic and does no equi e any human anno a ion e o , his app oach
educes he hype pa ame e uning ime by a leas hal . This ad an age is e en mo e signi ican
on la ge da ase s and deepe ne wo ks, whe e aining akes conside ably mo e ime, specially in
an a chi ec u e wi h h ee aining phases as his one.
Figu e 6.13: Example o he ob ained esul s in a weakly supe ised scena io wi h α=0.99 on a
syn he ic da ase wi hou colou cues and mul iple a ge s.
6.2 Weakly Supe ised App oach 95
6.2.2 Real Da ase s
As al eady men ioned, in his app oach some kind o g ound- u h masks om o he ask
a e needed (RQ1). In his case, bounding boxes ha delimi he ce ix a ea we e used, so no
p ep ocessing was done o he o iginal images, excep o a no malisa ion o he pixel alues (c. .
sec ion 5.1.3). A e hype pa ame e uning, he inal model eached 81.37% accu acy on he
es se (RQ2), a ma ginally highe alue han wi h he unsupe ised app oach.
Simila ly o he unsupe ised app oach, he i s case was co ec ly classi ied as heal hy ( igu e
6.14 (a)). Despi e ha ing highligh ed he specula e lec ions, he model is e y con iden ha i is
a heal hy ce ix, and he whole ce ix is highligh ed as an indica ion ha no egions ha e lesions.
The second case (b) is also a no mal case co ec ly classi ied, bu ha wi h he unsupe ised
app oach had been classi ied as cance ous wi h low con idence (51.08%). He e, once mo e he
whole ce ix is highligh ed, excep o he bump, which was highligh ed in he p e ious app oach.
No conside ing his bump a ele an egion is aligned wi h he in ui ion ha ha egion was wha
mislead he classi ie in o p edic ing his as a cance case. A i s glance his bump migh seem an
indica o o a lesion, bu i s egula , plana and pinkish na u e a e consis en wi h isual ea u es
o heal hy issue.
Case (c) was misclassi ied as non cance ous wi h his app oach, while in he unsupe ised
scena io i was co ec ly classi ied. This can be caused by he ac ha , al hough he majo i y o
non-in e es egions we e punished, he classi ie ocused on he e lec ion on he igh side o he
image. Also, his image has some blu ed a eas wi hin he in e es egion, which did no a ec he
classi ie as much in he unsupe ised scena io, because he image was c opped and, he eby, he
ce ix a ea was ampli ied and less a ec ed by hese blu ed egions.
Figu e 6.14 (d) was co ec ly classi ied as cance ous, which did no happen in he o he ap-
p oach. In his case, he p oduced explana ion does no p o ide e y insigh ul in o ma ion as o
why his decision was made. Simila ly o wha happened be o e, e y ew egions a e highligh ed
as ele an , maybe due o poo ligh ing condi ions and he p esence o e lec ions.
Figu e 6.14 (e) is an example in which a no mal case was misclassi ied as cance ous. In ac ,
he classi ie ocused oo much on specula e lec ions, mainly on he le side o he image.
Once mo e, he analysis o hese examples leads us o conclude han when no misclassi ica ion
occu s, he explana ions highligh egions o he image consis en wi h a p io i knowledge
o isual ea u es associa ed wi h cance ous cases, such as mo phology, colou and con ou
changes o he ce ix issue. The e o e, he p oduced explana ions a e indica i e o inju ed
a eas, bu no exac opog aphical a eas o lesions (RQ4).
96 Resul s and Discussion
(a) No mal case co ec ly classi ied (99.78%).
(b) No mal case co ec ly classi ied (99.91%).
(c) Cance case misclassi ied (86.09%).
(d) Cance case co ec ly classi ied (90.52%).
(e) No mal case misclassi ied (86.87%).
Figu e 6.14: Examples o explana ions ob ained in a weakly supe ised scena io wi h α=0.88
on he NIH-NCI ce ical cance da ase .
Chap e 7
Conclusions and Fu u e Wo k
The ou s anding p edic i e pe o mance o DL models has ecen ly led o he deploymen o
such sys ems in he eal wo ld, aising a my iad o new echnical and legal p oblems and equi e-
men s ha need o be ackled, especially when deploying hese sys ems in highly egula ed a eas
such as medicine. One o he mos impo an and challenging p oblems nowadays is known as
XAI: me hods ha allow AI sys ems o be unde s andable and us ed by humans. Ideally, one
would need da a labelled wi h he decisions (classi ica ion, eg ession, de ec ion o segmen a ion
labels, o example) and wi h he explana ions o hose decisions, in o de o supe ise he p ocess
o gene a ing explana ions. As his is no easible, mainly due o he inhe en di icul ies o he
subjec i e na u e o his a ea, hese sys ems need o be able o gene a e hese explana ions in an
unsupe ised ashion.
To y and begin o ackle his p oblem, and also o explo e he ca ego y o in-model me hods,
which is less de eloped compa ed o he pos -model class, in his wo k we p oposed an in-model
join app oach o p oduce decisions and explana ions using CNNs wi hou he need o supe ise he
explana ion componen . The de eloped a chi ec u e is composed o a classi ie and an explaine ,
hus p oducing class p edic ions and isual explana ions. This no el a chi ec u e, along wi h i s
cus om aining p ocess and loss unc ion, allows o he classi ie o be ained wi h he p oduced
explana ions, hus only ocusing on ele an pa s o he inpu image. Simila ly, he explaine is
ained oge he wi h he classi ie , hus lea ning o ind he ele an image egions ha con ibu e
o he co ec classi ica ion o said image and, he e o e, explain such p edic ion. This a chi ec u e
can be used wi h any classi ie a chi ec u e, p o ided he co esponding connec ions a e made and
he classi ie is e ained wi hin his new scena io. In ac , in his wo k wo classi ie a chi ec u es
we e used, a ResNe 50 classi ie o he NIH-NCI ce ical cance da ase and a VGG-16-based
a chi ec u e o he emaining da ase s.
The explaine is ained wi hou di ec supe ision, bu wi h indi ec supe ision h ough he
global loss unc ion ha includes a classi ica ion and an explana ion componen . Two s a egies
we e p oposed o his las componen : an unsupe ised and a weakly supe ised app oach. The
i s one in oduces egula isa ion e ms, he l1 penalised no m and o al a ia ion, ha impose de-
si ed cha ac e is ics on he explana ions, such as spa si y and con igui y, espec i ely. The weakly
97
98 Conclusions and Fu u e Wo k
supe ised app oach wo ks by punishing he explana ions ou side o in e es egions de ined by
masks c ea ed om anno a ions o o he asks, namely de ec ion o segmen a ion asks. The i s
app oach does no equi e any u he anno a ions besides he class labels, bu i in ol es uning
wo hype pa ame e s, while he weakly supe ised s a egy educes his uning e o by a leas
hal .
To es he p oposed a chi ec u e a syn he ic da a gene a ion ool wi h au oma ic anno a ion
was also de eloped, in addi ion o he eal da ase s used. This ool allows p oducing images con-
aining simple polygons, in a way ha is bo h e sa ile and highly cus omisable. These simple
images do no equi e expe knowledge o compa e hei desi ed explana ions wi h he ones p o-
duced by he p oposed a chi ec u e, which g ea ly speeds up he e alua ion p ocess and helps
debug he model.
The esul s ob ained show ha his a chi ec u e is able o p oduce isual explana ions, as well
as decisions, while no deg ading classi ica ion accu acy. The explana ions a e no only in e -
p e able, bu also comple e, i.e., a e able o desc ibe he sys em’s in e nals accu a ely. When com-
pa ed o s a e o he a me hods, he p oposed join app oach p oduces be e isual explana ions
on syn he ic da a ha , in ac , only highligh ele an image egions o he ou pu ed p edic ions.
Especially wi h ega ds o nega i e ins ances, his me hod conside s he whole image, excep he
non- a ge objec s, an explana ion, which aligns wi h he in ui ion ha , in he absence o a a ge
objec , he whole image is he explana ion o why no a ge is p esen .
Fu he mo e, he p oposed a chi ec u e was es ed on a eal da ase , he NIH-NCI ce ical
cance da ase . The ob ained esul s show ha , when no misclassi ica ion occu s, he explana ions
highligh egions o he image consis en wi h a p io i knowledge o isual ea u es associa ed wi h
cance ous cases, such as mo phology, colou and con ou changes o he ce ix issue. Howe e ,
he iden i ica ion o he highligh ed egions in misclassi ica ion examples is also impo an o be e
unde s and why he misclassi ica ion occu ed and debug he model. In ac , i was obse ed ha
he a chi ec u e is sensible o ligh ing condi ions and e lec ions ha migh appea in he images
due o hei collec ion p ocess, which p o ided impo an insigh s in o he need o be e p ep ocess
he da ase .
In conclusion, in his wo k we p oposed an in-model join app oach ha , o he easons
s a ed abo e, is able o p oduce meaning ul and easonable explana ions o i s gene a ed p edic-
ions, wi hou sac i icing p edic i e pe o mance. Fu he mo e, his a chi ec u e can be used wi h
di e en classi ie a chi ec u es on eal da a and does no equi e u he anno a ions, being
capable o p oducing explana ions in an unsupe ised o weakly supe ised way. The e o e,
we conclude ha all he p oposed esea ch ques ions we e add essed and p o en easible.
Howe e , esea ch in his a ea is s ill in i s in ancy and he p oposed app oach can be imp o ed
and u he explo ed. Examples o some u u e esea ch di ec ions include:
•Applica ion o mo e complex classi ica ion p oblems including mul i-class classi ica ion.
•Compa ison be ween he usage o ine g ained masks, such as segmen a ion masks, e sus
de ec ion masks in he weakly supe ised app oach.
Conclusions and Fu u e Wo k 99
•Mo e obus e alua ion o he p oduced explana ions, especially in he ce ical cance ap-
plica ion, by consul ing mo e expe s.
•E alua ion o he p oposed a chi ec u e in ela ion o possible ans o ma ions o he inpu
da a.
•Inclusion o mo e p ep ocessing s eps, mainly in he ce ical cance case, such as he e-
mo al o specula e lec ions.
•Combina ion o he p oposed unsupe ised and weakly supe ised app oaches.
•Design o a loss unc ion ha inco po a es cha ac e is ics a he popula ion le el, such as he
p omo ion o explana ion clus e s simila o he decision clus e s, ins ead o only ins ance
le el aspec s.
•Ex ension o he a chi ec u e o include mul imodal inpu s, such as addi ional abula in o -
ma ion included in he NIH-NCI ce ical cance da ase (age, HPV s a us, e c.).
•Ex ension o he a chi ec u e o also p oduce ex ual explana ions.
100 Conclusions and Fu u e Wo k
Appendix A
Ex a Files and Submi ed Pape s
A.1 Syn he ic Da a Gene a ion Tool Con igu able Pa ame e s
• olde =’ ain’: di ec o y whe e he da ase iles a e o be s o ed/impo ed om
•con ig_ ile=None: JSON con igu a ion ile whe e pa ame e s eside i no None. I None,
pa ame e s a e passed as class cons uc o a gumen s
•n _images=100: numbe o gene a ed images
•polygon=[-1, None]: a ge polygon used o gene a e anno a ions gi en by i s pa ame e s p
and q(i pequals -1, co esponding o a ci cle, hen qis i ele an )
•ou side_polygon=None: polygon o place a ound a ge polygon
•backg ound_colou =255: backg ound image colou in RGB
•img_heigh =224: image heigh
•img_wid h=224: image wid h
•n _channels=3: numbe o colou channels
•n _shapes=20: numbe o polygons pe gene a ed image
•n _ ies=100: numbe o ies be o e he algo i hm gi es up ying o i he polygon inside
he image
• ad_min=224/32: minimum possible adius o he polygon’s ou e ci cum e ence
• ad_max=224/16: maximum possible adius o he polygon’s ou e ci cum e ence
•o e lap=False: o e lap be ween polygons o he same image i T ue, no o e lap be ween
e e y wo polygons i False
•occlusion=False: occlusion o polygons on image bo de s i T ue, no occlusion i False
101
102 Ex a Files and Submi ed Pape s
• o a ion=T ue: andom o a ion o polygons i T ue, no o a ion i False
•noise=False: Gaussian noise addi ion o image i T ue
•min_n _ e ices=3: minimum numbe o e ices
•max_n _ e ices=13: maximum numbe o e ices
•min_n _shapes=1: minimum numbe o polygons pe image
•max_n _shapes=20: maximum numbe o polygons pe image
•simpli ied=False: simpli ied e sion o he da ase (only iangles and ci cles)
•no_ci cles=False: do no d aw ci cles (i simpli ied mode is se , hen his pa ame e will be
igno ed)
•poly_colou =False: colou o he a ge polygon in RGB
•s a _index=0: s a index o image naming (use ul when one wan s o add mo e images
o an exis ing da ase )
A.2 XML Anno a ion File 103
A.2 XML Anno a ion File
As depic ed in igu e A.1, each anno a ion ile includes he olde and ile whe e he image is
loca ed, as well as he image’s wid h, heigh and dep h. I also con ains he a ge polygon’s pand
q alues, ollowed by he bounding box coo dina es (xmin, ymin, xmax, ymax) o e e y ins ance
o he a ge polygon, as well as he numbe o a ge polygons p esen in he image.
< a n n o a i o n >
< o l d e > s i m p l i i e d _ n o _ c o l o u < / o l d e >
< il enam e > 0. png< / i le nam e >
< s i z e >
<wid h>224< / wid h>
< h e i g h >224< / h e i g h >
<dep h >3< / dep h>
< / s i z e >
<polygon>
<p> 3.0 < / p>
<q> 1.0 < / q>
< / polygon>
<bndbox0>
<xmin>20< / xmin>
<ymin>149< / ymin>
<xmax>38< / xmax>
<ymax>167< / ymax>
< / bndbox0>
< e x i s s >1< / e x i s s >
< / a n n o a i o n >
Figu e A.1: Example XML anno a ion ile in Pascal VOC o ma .
Towa ds a Join App oach o P oduce Decisions and Explana ions 7
om 10-8 o 10-4. Wi hou egula isa ion, one can ob ain a deg aded solu ion,
in which e e y hing is conside ed an explana ion. The e o e, L1 egula isa ion
is employed, so ha only a small pa o he whole image cons i u es an expla-
na ion.
A quali a i e and quan i a i e compa ison o he p oposed a chi ec u e wi h
a ious me hods a ailable in he iNN es iga e oolbox [1] is also made. This ool-
box aims o acili a e he compa ison o e e ence implemen a ions o pos -model
in e p e abili y me hods, by p o iding a common in e ace and ou -o - he-box
implemen a ion o a ious analysis me hods. The oolbox is, hen, used o com-
pa e he p oposed a chi ec u e wi h me hods like Smoo hG ad [14], Decon Ne
[17], Guided Backp op [15], Deep Taylo Decomposi ion [7] o Laye -Wise Rele-
ance P opaga ion (LRP) [2]. In he p oposed a chi ec u e, he explaine is he
componen ha p oduces a isual ep esen a ion o he easoning behind he
classi ie ’s decisions, jus like he analyse s a ailable in he iNN es iga e ool-
box. As such, hese analysis me hods a e applied only o he classi ie o he
p oposed a chi ec u e, in o de o compa e only he explana ion gene a o s, i.e.
he p oposed a chi ec u e’s explaine and he di e en analysis me hods. Since
hese me hods a e applied a e he model is ained, we s a ed by aining he
classi ie on he simple da ase wi hou colou cues. Then, he a ious analysis
me hods a e applied o he ained classi ie and hei gene a ed isual explana-
ions a e compa ed o he ones ou pu ed by he explaine ained in he p e ious
expe imen s wi h 10-6 L1 egula isa ion ac o . Fu he mo e, he classi ie ’s ac-
cu acy wi h and wi hou explaine a e also compa ed. The ob ained esul s a e
desc ibed in sec ion 3.
2.4 Expe imen s on eal da ase s
Expe imen s we e also conduc ed on a eal da ase , a ailable a h ps://gi hub.
com/ gei hos/ ex u e- s-shape. This da ase was c ea ed in he con ex
o he wo k de eloped by Gei hos e al [3], whe e he au ho s alida e ha
Imagene - ained CNNs a e biased owa ds ex u e. In o de o alida e his hy-
po hesis, he au ho s p opose a cue con lic expe imen in which s yle ans e
is employed, in oducing ex u e in he Imagene images. This da ase con ains
16 classes, wi h 80 images each. The p oposed a chi ec u e was ained on his
da ase wi hou any egula isa ion. Resul s o his expe imen a e shown in
sec ion 3.
3 Resul s and Discussion
Fo all o he ollowing images i is wo h no ing ha he colou code anges om
pu ple o yellow, whe e yellow ep esen s highe pixel alues. The le column
co esponds o he o iginal image, he middle column o he ou pu ed explana-
ion and he igh column o he explana ion a e an absolu e h eshold o 0.75
is applied.
110 Ex a Files and Submi ed Pape s
8 I. Rio-To o e al.
Figu e 4 cons i u es examples o he ob ained esul s o simple da ase s wi h
and wi hou a a ge polygon o di e en colou . Bo h images a e he esul
o aining wi hou any kind o egula isa ion. Fo such simple da ase s, i is
expec ed ha he explana ion ocuses on he a ge polygon, ende ing he es
o he image as i ele an o he p edic ed class.
While on he da ase wi h colou cues he esul ing explana ion consis s only
o he a ge polygon, as expec ed, in he sligh ly mo e complex da ase wi hou
colou cues he whole image is conside ed an explana ion, which co obo a es
he need o egula ising he explaine ou pu . As s a ed in sec ion 2.3, we use
an L1 egula isa ion ac o , because i allows o he selec ion o he ele an
pa s o he explana ion, ensu ing ha only a small pa o he image is in ac
he explana ion o he classi ie ’s decision. Thus, his egula isa ion ensu es ha
he explana ion is no only in e p e able, bu also comple e, as desi ed.
Fig. 4: Posi i e ins ance and espec i e explana ion. These esul s we e ob ained
wi hou any kind o egula isa ion o he explaine ’s ou pu while aining on a
simple da ase wi hou ( op) and wi h (bo om) colou cues.
Figu e 5 is he esul o he expe imen s wi h di e en egula isa ion ac o s,
namely 10-8, 10-6 and 10-4, on he da ase wi hou colou cues. Wi h a ac o o
10-8, no only he a ge polygon is conside ed ele an , as well as he ci cles,
which may imply ha such a small egula isa ion is s ill no enough o limi he
ele an pa s o he explana ion. In ac , inc easing L1 o 10-6, p oduces much
be e esul s, wi h he a ge polygon clea ly highligh ed. Finally, inc easing
L1 a bi u he , o 10-4, p o ed o be oo much egula isa ion, causing he
explana ion o “disappea ”.
A.3 IbPRIA2019 Accep ed Pape (Honou able Men ion Winne ) 111
Towa ds a Join App oach o P oduce Decisions and Explana ions 9
Fig. 5: Posi i e ins ance and espec i e explana ion. These esul s we e ob ained
wi h 10-8 ( op), 10-6 (middle) and 10-4 (bo om) L1 egula isa ion o he ex-
plaine ’s ou pu while aining on a simple da ase wi hou colou cues.
Mo eo e , he p oposed a chi ec u e was compa ed o se e al o he me h-
ods a ailable in he iNN es iga e amewo k [1]. As can be seen in igu e 6,
he majo i y o he me hods a e unable o p oduce easonable explana ions o
he chosen da ase , highligh ing co ne s o he image, o example, while he
p oposed a chi ec u e is able o only highligh he ele an egions o he clas-
si ie ’s decision (see igu e 5 middle). Fu he mo e, aining only he classi ie ,
as was done when applying he iNN es iga e oolbox’s analysis me hods, yields
accu acies close o 62%, while he accu acy o he p oposed a chi ec u e eaches
100%, as illus a ed in Figu e 7. Fo his da ase , he p oposed ne wo k no only
p oduces explana ions alongside wi h p edic ions, as well as imp o es accu acy,
by o cing he classi ie o ocus only on ele an pa s o he image.
112 Ex a Files and Submi ed Pape s
10 I. Rio-To o e al.
Fig. 6: Resul s o he applica ion o 10 analysis me hods a ailable on he iNN es-
iga e oolbox [1] o he p oposed classi ie and compa ison wi h he p oposed
end- o-end a chi ec u e. The colo map o he igh column’s images was ad-
jus ed o help isualiza ion due o he small size o each image and o easie
compa ison.
Fig. 7: E olu ion o classi ie accu acy pe aining epoch o he same classi-
ie ained alone (le ) and wi hin he p oposed explaine -classi ie a chi ec u e
( igh ).
Finally, Figu e 8 depic s he ob ained esul s o he expe imen on he cue
con lic da ase . One can see ha he gene a ed explana ions a e o ien ed o-
wa ds seman ic componen s o he objec s. Fo example, o he bo le case he
explana ions ocus mo e on he neck o he bo le and on i s label. In he ca
example, he explana ion highligh s he ca ’s bumpe and in he bicycle case,
he handles and he sea a e highligh ed, while in he chai example he chai ’s
legs a e highligh ed. I is wo h no ing ha al hough he esul ing explana ions
highligh di e en seman ic componen s o he objec s, hey do no appea con-
nec ed o each o he ( o example, he handle and he sea o he bicycle). This
A.3 IbPRIA2019 Accep ed Pape (Honou able Men ion Winne ) 113
Towa ds a Join App oach o P oduce Decisions and Explana ions 11
esul hin s ha imp o ing he quali y o hese explana ions can be made by
ensu ing ha explana ions a e spa se, i.e., co e a smalle pa o he whole
image, and also connec ed.
Fig. 8: Resul s o he cue con lic expe imen .
114 Ex a Files and Submi ed Pape s
12 I. Rio-To o e al.
4 Conclusion
We p opose a p elimina y in-model join app oach o p oduce decisions and
explana ions using CNNs, capable o p oducing no only in e p e able explana-
ions, bu also comple e ones, i.e., explana ions ha a e able o desc ibe he
sys em’s in e nals accu a ely. We also de eloped a syn he ic da ase gene a ion
amewo k wi h au oma ic anno a ion.
The p oposed a chi ec u e was es ed wi h a simple gene a ed syn he ic
da ase , o which explana ions a e in ui i e and do no need o employ ex-
pe knowledge. Resul s show he po en ial o he p oposed a chi ec u e, espe-
cially when compa ed o exis ing me hods and when adding L1 egula isa ion.
These also hin a he need o egula isa ion in o de o be e balance he
in e p e abili y-comple eness ade o . As such, u u e esea ch will s udy he
e ec o adding o al a ia ion egula isa ion as a way o making explana ions
spa se. Also, we will explo e he possible ad an ages o supe ising he explana-
ions, as well as de elop a p ope anno a ion scheme and e alua ion me ics o
such ask.
Re e ences
1. Albe , M., Lapuschkin, S., Seege e , P., H¨agele, M., Sch¨u , K.T., Mon a on, G.,
Samek, W., M¨ulle , K.R., D¨ahne, S., Kinde mans, P.J.: iNN es iga e neu al ne -
wo ks! (2018)
2. Bach, S., Binde , A., Mon a on, G., Klauschen, F., M¨ulle , K.R., Samek, W.: On
pixel-wise explana ions o non-linea classi ie decisions by laye -wise ele ance
p opaga ion. PLoS One (2015). h ps://doi.o g/10.1371/jou nal.pone.0130140
3. Gei hos, R., Rubisch, P., Michaelis, C., Be hge, M., Wichmann, F.A., B endel, W.:
ImageNe - ained CNNs a e biased owa ds ex u e; inc easing shape bias imp o es
accu acy and obus ness (no 2018), h p://a xi .o g/abs/1811.12231
4. Gilpin, L.H., Bau, D., Yuan, B.Z., Bajwa, A., Spec e , M., Kagal, L.: Explain-
ing explana ions: An o e iew o in e p e abili y o machine lea ning. P oc. -
2018 IEEE 5 h In . Con . Da a Sci. Ad . Anal. DSAA 2018 pp. 80–89 (2019).
h ps://doi.o g/10.1109/DSAA.2018.00018
5. Goodman, B., Flaxman, S.: Eu opean Union egula ions on algo-
i hmic decision-making and a ” igh o explana ion” (jun 2016).
h ps://doi.o g/10.1609/aimag. 38i3.2741, h ps://a xi .o g/abs/1606.08813
6. Lip on, Z.C.: The My hos o Model In e p e abili y (2016).
h ps://doi.o g/10.1145/3233231
7. Mon a on, G., Lapuschkin, S., Binde , A., Samek, W., M¨ulle , K.R.: Explaining
nonlinea classi ica ion decisions wi h deep Taylo decomposi ion. Pa e n Recog-
ni . (2017). h ps://doi.o g/10.1016/j.pa cog.2016.11.008
8. Mon a on, G., Samek, W., M¨ulle , K.R.: Me hods o in e p e ing and un-
de s anding deep neu al ne wo ks. Digi . Signal P ocess. 73, 1–15 ( eb
2018). h ps://doi.o g/10.1016/J.DSP.2017.10.011, h ps://www.sciencedi ec .
com/science/a icle/pii/S1051200417302385
9. Ribei o, M.T., Singh, S., Gues in, C.: ”Why Should I T us You?”: Explaining
he P edic ions o Any Classi ie ( eb 2016), h p://a xi .o g/abs/1602.04938
A.3 IbPRIA2019 Accep ed Pape (Honou able Men ion Winne ) 115
Towa ds a Join App oach o P oduce Decisions and Explana ions 13
10. Samek, W., Binde , A., Mon a on, G., Lapuschkin, S., M¨ulle , K. .: E alua ing he
Visualiza ion o Wha a Deep Neu al Ne wo k Has Lea ned. IEEE T ans. Neu al
Ne wo ks Lea n. Sys . 28(11), 2660–2673 (2017)
11. Sil a, W., Fe nandes, K., Ca doso, J.S.: How o p oduce complemen a y explana-
ions using an ensemble model. In: 2019 In e na ional Join Con e ence on Neu al
Ne wo ks (IJCNN) (2019)
12. Sil a, W., Fe nandes, K., Ca doso, M.J., Ca doso, J.S.: Towa ds Complemen a y
Explana ions Using Deep Neu al Ne wo ks (2018). h ps://doi.o g/10.1007/978-3-
030-02628-815
13. Simonyan, K., Zisse man, A.: Ve y deep con olu ional ne wo ks o la ge-scale
image ecogni ion. a Xi p ep in a Xi :1409.1556 (2014)
14. Smilko , D., Tho a , N., Kim, B., Vi, F.: Smoo hG ad : emo ing noise by adding
noise (2017)
15. Sp ingenbe g, J.T., Doso i skiy, A., B ox, T., Riedmille , M.: S i ing o Simplic-
i y: The All Con olu ional Ne (dec 2014), h ps://a xi .o g/abs/1412.6806
16. Zeile , M.D.: Adadel a: an adap i e lea ning a e me hod. a Xi p ep in
a Xi :1212.5701 (2012)
17. Zeile , M.D., Fe gus, R.: Visualizing and unde s anding con olu ional ne wo ks. In:
Lec . No es Compu . Sci. (including Subse . Lec . No es A i . In ell. Lec . No es
Bioin o ma ics) (2014). h ps://doi.o g/10.1007/978-3-319-10590-153
116 Ex a Files and Submi ed Pape s
Re e ences
[1] Michael Van Len , William Fishe , and Michael Mancuso. An explainable a i icial in elli-
gence sys em o small-uni ac ical beha io . In P oceedings o he na ional con e ence on
a i icial in elligence, pages 900–907. Menlo Pa k, CA; Camb idge, MA; London; AAAI
P ess; MIT P ess; 1999, 2004.
[2] Da id Gunning. Explainable a i icial in elligence (xai) - P og am Upda e. De ense Ad-
anced Resea ch P ojec s Agency (DARPA), 2, No 2017.
[3] Au élien Gé on. Hands-On Machine Lea ning wi h Sciki -Lea n and Tenso Flow: Con-
cep s, Tools, and Techniques o Build In elligen Sys ems. O’Reilly Media, Inc., 1s edi ion,
2017.
[4] Ian Good ellow, Yoshua Bengio, and Aa on Cou ille. DeepLea ning. MIT P ess, 2016.
www.deeplea ningbook.o g, [Accessed: 20 h Sep embe 2019].
[5] Dhai ya Pa ikh. Lea ning pa adigms in machine lea ning. h ps://medium.
com/da ad i enin es o /lea ning-pa adigms-in-machine-lea ning-
146eb 8b5943, Jul 2018. [Accessed: 20 h Sep embe 2019].
[6] Zhi-Hua Zhou. A b ie in oduc ion o weakly supe ised lea ning. Na ional Science Re-
iew, 5(1):44–53, 08 2017. doi:10.1093/ns /nwx106.
[7] Fahad La ee and Yassine Ruichek. Su ey on seman ic segmen a ion using deep lea ning
echniques. Neu ocompu ing, pages 321–348. doi:10.1016/j.neucom.2019.02.
003.
[8] Mengnan Du, Ninghao Liu, and Xia Hu. Techniques o In e p e able Machine Lea ning.
a Xi :1808.00033.
[9] Wa en S. McCulloch and Wal e Pi s. A logical calculus o he ideas immanen in ne ous
ac i i y. The bulle in o ma hema ical biophysics, 5(4):115–133, Dec 1943. doi:10.
1007/BF02478259.
[10] Ana Ne es, Ignacio Gonzalez, John Leande , and Raid Ka oumi. A new app oach o
damage de ec ion in b idges using machine lea ning. pages 73–84, 01 2018. doi:
10.1007/978-3-319-67443-8_5.
[11] Jayesh Bapu Ahi e. The a i icial neu al ne wo ks handbook: Pa 4. h ps:
//medium.com/@jayeshbahi e/ he-a i icial-neu al-ne wo ks-
handbook-pa -4-d2087d1 583e, No 2018. [Accessed: 20 h Sep embe 2019].
117
118 REFERENCES
[12] Da id E Rumelha , Geo ey E Hin on, and Ronald J Williams. Lea ning in e nal ep e-
sen a ions by e o p opaga ion. Technical epo , Cali o nia Uni San Diego La Jolla Ins
o Cogni i e Science, 1985.
[13] Ma in Minsky and Seymou Pape . Pe cep on: an in oduc ion o compu a ional geom-
e y. The MIT P ess, Camb idge, expanded edi ion, 19(88):2, 1969.
[14] Fa io Vázquez. A "wei d" in oduc ion o deep lea ning. h ps://
owa dsda ascience.com/a-wei d-in oduc ion- o-deep-lea ning-
7828803693b0, Aug 2018. [Accessed: 20 h Sep embe 2019].
[15] A den De a . Applied deep lea ning - pa 1: A i icial neu al ne wo ks.
h ps:// owa dsda ascience.com/applied-deep-lea ning-pa -1-
a i icial-neu al-ne wo ks-d7834 67a4 6, Oc 2017. [Accessed: 20 h
Sep embe 2019].
[16] Xa ie Glo o and Yoshua Bengio. Unde s anding he di icul y o aining deep eed o -
wa d neu al ne wo ks. In P oceedings o he hi een h in e na ional con e ence on a i icial
in elligence and s a is ics, pages 249–256, 2010.
[17] Diede ik P. Kingma and Jimmy Ba. Adam: A me hod o s ochas ic op imiza ion, 2014.
a Xi :1412.6980.
[18] Ma hew D Zeile . Adadel a: an adap i e lea ning a e me hod, 2012. a Xi :1212.5701.
[19] Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun. Adap i e g adien me hods wi h
dynamic bound o lea ning a e, 2019. a Xi :1902.09843.
[20] Geo ey E Hin on, Ni ish S i as a a, Alex K izhe sky, Ilya Su ske e , and Ruslan R
Salakhu dino . Imp o ing neu al ne wo ks by p e en ing co-adap a ion o ea u e de ec-
o s, 2012. a Xi :1207.0580.
[21] Se gey Io e and Ch is ian Szegedy. Ba ch no maliza ion: Accele a ing deep ne wo k
aining by educing in e nal co a ia e shi . In P oceedings o he 32Nd In e na ional
Con e ence on In e na ional Con e ence on Machine Lea ning - Volume 37, ICML’15,
pages 448–456. JMLR.o g, 2015. URL: h p://dl.acm.o g/ci a ion.c m?id=
3045118.3045167.
[22] Da id H Hubel. Single uni ac i i y in s ia e co ex o un es ained ca s. The Jou nal o
physiology, 147(2):226–238, 1959.
[23] Da id H Hubel and To s en N Wiesel. Recep i e ields o single neu ones in he ca ’s s ia e
co ex. The Jou nal o physiology, 148(3):574–591, 1959.
[24] Da id H Hubel and To s en N Wiesel. Recep i e ields and unc ional a chi ec u e o mon-
key s ia e co ex. The Jou nal o physiology, 195(1):215–243, 1968.
[25] Kunihiko Fukushima and Sei Miyake. Neocogni on: A sel -o ganizing neu al ne wo k
model o a mechanism o isual pa e n ecogni ion. In Compe i ion and coope a ion in
neu al ne s, pages 267–285. Sp inge , 1982.
[26] Con olu ional neu al ne wo k. h ps://www.ma hwo ks.com/solu ions/deep-
lea ning/con olu ional-neu al-ne wo k.h ml, [Accessed: 20 h Sep embe
2019.
REFERENCES 119
[27] Con olu ional neu al ne wo ks (cnns / con ne s). h p://cs231n.gi hub.io/
con olu ional-ne wo ks/, [Accessed: 20 h Sep embe 2019].
[28] Y. Lecun, L. Bo ou, Y. Bengio, and P. Ha ne . G adien -based lea ning applied o docu-
men ecogni ion. P oceedings o he IEEE, 86(11):2278–2324, 1998.
[29] Alex K izhe sky, Ilya Su ske e , and Geo ey E Hin on. Imagene classi ica ion wi h
deep con olu ional neu al ne wo ks. In Ad ances in neu al in o ma ion p ocessing sys-
ems, pages 1097–1105, 2012.
[30] Ka en Simonyan and And ew Zisse man. Ve y deep con olu ional ne wo ks o la ge-scale
image ecogni ion. 2014. a Xi :1409.1556.
[31] Vgg16 - con olu ional ne wo k o classi ica ion and de ec ion, No 2018. h ps:
//neu ohi e.io/en/popula -ne wo ks/ gg16/, [Accessed: 20 h Sep embe
2019].
[32] Ch is ian Szegedy, Wei Liu, Yangqing Jia, Pie e Se mane , Sco Reed, D agomi
Anguelo , Dumi u E han, Vincen Vanhoucke, and And ew Rabino ich. Going deepe
wi h con olu ions. In P oceedings o he IEEE con e ence on compu e ision and pa e n
ecogni ion, pages 1–9, 2015.
[33] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep esidual lea ning o
image ecogni ion. In P oceedings o he IEEE con e ence on compu e ision and pa e n
ecogni ion, pages 770–778, 2016.
[34] Siddha h Das. Cnn a chi ec u es: Lene , alexne , gg, googlene , esne and mo e
...., Jul 2018. h ps://medium.com/@side eal/cnns-a chi ec u es-lene -
alexne - gg-googlene - esne -and-mo e-666091488d 5, [Accessed: 20 h
Sep embe 2019].
[35] Ola Ronnebe ge , Philipp Fische , and Thomas B ox. U-ne : Con olu ional ne wo ks o
biomedical image segmen a ion. In In e na ional Con e ence on Medical image compu ing
and compu e -assis ed in e en ion, pages 234–241. Sp inge , 2015.
[36] Amina Adadi and Mohammed Be ada. Peeking Inside he Black-Box: A Su ey on
Explainable A i icial In elligence (XAI). IEEE Access, 6:52138–52160, 2018. doi:
10.1109/ACCESS.2018.2870052.
[37] Wilson Sil a, Kelwin Fe nandes, Ma ia João Ca doso, and Jaime S. Ca doso. Towa ds com-
plemen a y explana ions using deep neu al ne wo ks. In Unde s anding and In e p e ing
Machine Lea ning in Medical Image Compu ing Applica ions - Fi s In e na ional Wo k-
shops MLCN 2018, DLF 2018, and iMIMIC 2018, Held in Conjunc ion wi h MICCAI 2018,
G anada, Spain, Sep embe 16-20, 2018, P oceedings, pages 133–140, 2018.
[38] A u And zejak, Felix Langne , and Sil es e Zabala. In e p e able models om dis ibu ed
da a ia me ging o decision ees. In 2013 IEEE Symposium on Compu a ional In elligence
and Da a Mining (CIDM), pages 1–9. IEEE, 2013.
[39] Yin Lou, Rich Ca uana, Johannes Geh ke, and Giles Hooke . Accu a e in elligible models
wi h pai wise in e ac ions. In P oceedings o he 19 h ACM SIGKDD In e na ional Con-
e ence on Knowledge Disco e y and Da a Mining, KDD ’13, pages 623–631, New Yo k,
NY, USA, 2013. ACM. doi:10.1145/2487575.2487579.
126 REFERENCES
[125] Ped o M. Fe ei a, Filipe Ma ques, Jaime S. Ca doso, and Ana Rebelo. Physiological In-
spi ed Deep Neu al Ne wo ks o Emo ion Recogni ion. IEEE Access, 2018.
[126] Robe Gei hos, Ca los R. Medina Temme, Jonas Raube , Heiko H. Schü , Ma hias
Be hge, and Felix A. Wichmann. Gene alisa ion in humans and deep neu al ne wo ks. In
Ad ances in Neu al In o ma ion P ocessing Sys ems, numbe Neu IPS 2018, pages 7538–
7550, 2018.
[127] Robe Gei hos, Pa icia Rubisch, Claudio Michaelis, Ma hias Be hge, Felix A. Wichmann,
and Wieland B endel. ImageNe - ained CNNs a e biased owa ds ex u e; inc easing shape
bias imp o es accu acy and obus ness. no 2018. a Xi :1811.12231.
[128] Rolando He e o, Ma k H Schi man, Concepción B a i, Allan Hildesheim, Ileana Bal-
maceda, Ma k E She man, Mi chell G eenbe g, Fe nando Cá denas, Víc o Gómez, Kay
Helgesen, e al. Design and me hods o a popula ion-based na u al his o y s udy o ce ical
neoplasia in a u al p o ince o cos a ica: he guanacas e p ojec . Re is a Paname icana
de Salud Pública, 1:362–375, 1997.
[129] Olga Russako sky, Jia Deng, Hao Su, Jona han K ause, Sanjee Sa heesh, Sean Ma, Zhi-
heng Huang, And ej Ka pa hy, Adi ya Khosla, Michael Be ns ein, e al. Imagene la ge
scale isual ecogni ion challenge. In e na ional jou nal o compu e ision, 115(3):211–
252, 2015.
[130] Geo ge A Mille . Wo dne : a lexical da abase o english. Communica ions o he ACM,
38(11):39–41, 1995.
[131] Robe P Kau man, S ephen J G i in, Jon D Lund, and Paul E Tulla . Cu en ecom-
menda ions o ce ical cance sc eening: do hey ende he annual pel ic examina ion
obsole e? Medical P inciples and P ac ice, 22(4):313–322, 2013.
[132] Dezhao Song, Edwa d Kim, Xiaolei Huang, Joseph Pa uno, Héc o Muñoz-A ila, Je
He lin, L Rodney Long, and Samee An ani. Mul imodal en i y co e e ence o ce ical
dysplasia diagnosis. IEEE ansac ions on medical imaging, 34(1):229–245, 2014.