RESEARCH ARTICLE Open Access
The saga o he many s udies w ongly associa ing
mi ochond ial DNA wi h b eas cance
An onio Salas
1*
, Manuel Ga cía-Maga iños
1,2
, Ian Logan
3
and Hans-Jü gen Bandel
4
Abs ac
Backg ound: A la ge body o gene ic esea ch has ocused on he po en ial ole ha mi ochond ial DNA (m DNA)
a ian s migh play on he p edisposi ion o common and complex (mul i- ac o ial) diseases. I has been a gued
howe e ha many o hese s udies could be inconclusi e due o a i ac s ela ed o geno yping e o s o
inadequa e design.
Me hods: Analyses o he da a published in case–con ol b eas cance associa ion s udies ha e been pe o med using
a phylogene ic-based app oach. Va ia ion obse ed in hese s udies has been in e p e ed in he ligh o da a a ailable
on public esou ces, which now include o e >27,000 comple e mi ochond ial sequences and he wo ldwide
phylogeny de e mined by hese mi ogenomes. Complemen a y analyses we e ca ied ou using public da ase s o
pa ial m DNA sequences, mainly co esponding o con ol- egion segmen s.
Resul s: By way o example, we show he e ano he kind o allacy in hese medical s udies, namely, he phenomenon
o SNP-SNP in e ac ion w ongly applied o haploid da a in a b eas cance s udy. We also eassessed he mu ually
con lic ing s udies sugges ing some unc ional ole o he non-synonymous polymo phism m.10398A > G (ND3 subuni
o mi ochond ial complex I) in b eas cance . In some s udies, con ol g oups we e employed ha showed an
ex emely odd haplog oup equency spec um compa ed o compa able in o ma ion om much la ge da abases.
Mo eo e , he use o inapp op ia e s a is ics signaled spu ious “signi icance”in se e al ins ances.
Conclusions: E e y case–con ol s udy should come unde sc u iny in ega d o he plausibili y o he con ol-g oup
da a p esen ed and app op ia eness o he s a is ical me hods employed; and his is bes done be o e po en ial
publica ion.
Keywo ds: Epis asis, SNP-SNP in e ac ion, Complex disease, Associa ion s udy, m DNA, Haplog oup
Backg ound
S udies on mi ochond ial DNA (m DNA) in human dis-
ease ha e o en been in ensely deba ed (see some exam-
ples on cance ins abili y [1-3]). On he one hand,
m DNA case–con ol associa ion s udies a e equen ly
a ec ed by se e al p oblems ela ed o de icien s udy
designs and inapp op ia e s a is ical me hods [4,5]. On
he o he hand, a phylogene ic app oach has p o en o
be ex emely use ul in disco e ing a ious kinds o e -
o s in hese s udies, which has o en comp omised hei
esul s and conclusions [1,6-9]. The allelic iew [1] on
m DNA a ia ion (as haplo ypic a ia ion) is also a
common misconcep ion in m DNA s udies, whe e single
a ian s om haplog oup mo i s a e ea ed as i hey
we e po en ially independen disease ma ke s. A new
mani es a ion o his p oblem has o do wi h he
phenomenon o SNP-SNP in e ac ion applied o he in-
e p e a ion o m DNA a ia ion.
Epis asis, a e m coined by Ba eson [10], was i s de-
ined as a masking e ec whe eby a a ian o allele a
one locus p e en s he a ian a ano he locus om
mani es ing i s e ec s [11]. Epis asis and gene ic in e -
ac ion e e o he same phenomenon; howe e , he
o me is widely used in popula ion gene ics and especially
e e s o he s a is ical p ope ies o he phenomenon.
In essence, SNP-SNP in e ac ion makes sense whene e
wo SNPs a e loca ed in di e en (unlinked) loci, and i
is gene ally applicable o he sphe e o he au osomal
* Co espondence: [email p o ec ed]
1
Unidade de Xené ica, Ins i u o de Medicina Legal, and Depa amen o de
Ana omía Pa olóxica e Ciencias Fo enses, Facul ad de Medicina, Uni e sidad
de San iago de Compos ela, 15782 Galicia, Spain
Full lis o au ho in o ma ion is a ailable a he end o he a icle
© 2014 Salas e al.; licensee BioMed Cen al L d. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e
Commons A ibu ion License (h p://c ea i ecommons.o g/licenses/by/4.0), which pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly c edi ed. The C ea i e Commons Public Domain
Dedica ion wai e (h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle,
unless o he wise s a ed.
Salas e al. BMC Cance 2014, 14:659
h p://www.biomedcen al.com/1471-2407/14/659
genome. Since he whole m DNA molecule is i ually
a single haploid locus, he concep o epis asis o
in e ac ion, by de ini ion, does no apply o SNP-SNP
in e ac ion o m DNA a ian s. Thus, he concep o
haplo ype is o a di e en na u e; haplo ype e e s o a
combina ion o igh ly linked a ian s in a segmen o a
ch omosome (which could be a gene, o he en i e
m DNA molecule). The e a e some haplo ypes (o mo e
gene ally, some haplog oups) ha ha e been epo ed as
con ibu ing o he isk o a disease, as i is he case o
se e al o a icles discussed in he p esen s udy. I is
he a ian s combined in his DNA segmen ha con-
ibu e o he isk as a single locus, bu no by way o
in e ac ion o independen loci; i is in ac he haplo-
ype as a whole ha is used o es o associa ion, no
he in e ac ion be ween he a ian s ha compose his
haplo ype. Howe e , he e m epis asis has been inco -
ec ly used in a ew m DNA s udies; o example, in des-
igna ing he po en ial e ec o a haplo ype in cance , as
exp essed by Can e e al. [12]: “epis a ic in e ac ions
be ween indi idual loci wi hin he mi ochond ial gen-
ome as well as nuclea -mi ochond ial gene in e ac ions
should be in es iga ed”.
In he p esen s udy we discuss he example [13]
whe e he concep o in e ac ion has been misunde -
s ood and misapplied. This example p o ides an oppo -
uni y o discuss he ole o he polymo phism A10398G
alias m.10398A > G (Th > Ala) in b eas cance , which
has been he ocus o as many as en s udies (and many
mo e e iews) in he pe iod 2005–2013. By c i ically
e iewing all hese s udies we con i m he conclusion o
he me a-analysis pe o med by F ancis e al. [14] ha
he associa ion esul s epo ed o da e a e con adic o y
and inconsis en , hus p o iding absolu ely no e idence
suppo ing a ole o his m DNA polymo phism in
b eas cance . Howe e , close examina ion o e e y sin-
gle s udy ha claimed o ha e ound some associa ion
e eals ha he e we e clea indica ions in ei he he
da a p esen ed o he s a is ical me hods used ha he
associa ion was spu ious.
Me hods
Phylogene ic me hods a e indispensable o m DNA s ud-
ies o human disease [4]. In pa icula , we use PhyloT ee
Build 16 (h p://www.phylo ee.o g/ [15]) and he in o -
ma ion p o ided conce ning haplog oup s a us alongside
wi h di ec inspec ion o he GenBank en ies ansla ed
in o mu a ion mo i lis s ( ela i e o he s anda d e e -
ence sequence, he CRS) a ailable in he Web (h p://
www.ianlogan.co.uk/sequences_by_g oup/haplog oup_
selec .h m; h p://www.ianlogan.co.uk/checke /accession.
h m). No e howe e ha PhyloT ee employs a con using
nomencla u e which does no con o m o he CRS based
no a ion used in medical gene ics [16]. He e, when
ocusing on a single a ian A o G a 10398, we w i e
A10398 and 10398G, whe e he p e ix posi ion is ese ed
o he CRS nucleo ide and he su ix o he any o he
a ian , wi hou in oking ances y [16].
The s a is ical ela ionship be ween single es esul s
(e.g. using Fishe ’s exac es o e e y m SNP) and logis-
ic eg ession analysis (applied o he de ec ion o
m SNP-m SNP in e ac ion) can be examined by way o a
simple simula ion analysis. Fo such an expe imen , wo
iden ical da a se s o m SNP geno ypes we e aken om
he con ol g oup in he “mainland Spain”b eas cance
se ies (N= 616) epo ed by Mosque a-Miguel e al. [5];
all he p o iles in one sample we e now a i icially la-
beled as cases while he o he p o iles in he o he copy
sample was labeled as “con ols”and s ayed unchanged.
A m DNA p o ile belonging o haplog oup K1 was
added one a a ime o a maximum o 120 o he subse
o “cases”and se e al s a is ical analyses we e each pe -
o med o “con ols” e sus “cases”. The goal o he
simula ion was o c ea e scena ios whe e haplog oup K1
is p og essi ely o e ep esen ed in cases compa ed o
con ols. A Fishe exac es is pe o med o he
m SNPs de ining haplog oup K1 (10398G, 12308G) and
o he amoun o haplog oup K1. Logis ic eg ession
was also ca ied ou in o de o de ec he ‘in e ac ion’
be ween m SNPs A10398 and 12308G as done in [13].
Compu a ion o s a is ical powe was ca ied ou using
mi Powe [17]. The compu a ion was done in a conse -
a i e manne wi h he only aim o highligh ing hose
case–con ol associa ion s udies ha a e ex emely
unde powe ed. We did no conside he a pos e io i
powe es ima ion op ion in mi Powe unde he assump-
ion ha es ima es would be s ongly a ec ed by he ac
ha popula ion s a i ica ion has se e ely in la ed e-
quencies in cases e sus con ols in he di e en coho s
e iewed (see Resul s below). We ins ead compu ed s a -
is ical powe assuming a ‘de no o’s udy ha conside s
he epo ed sample sizes in cases and con ols om he
di e en coho s and he equency o 10398G in hei
con ols. We assumed a conse a i e isk (odd a io) o
10398G equal o 2. Powe was compu ed using Fishe ’s
Exac Tes o he mos conse a i e scena io ha con-
side s 2 × 2 ables. No e ha gene ally a h eshold o a
leas 80% is (consensually) conside ed o be an adequa e
powe in case–con ol associa ion s udies.
Resul s
Nega i e indings o associa ion o A10398G wi h b eas
cance
Solid associa ion s udies would ideally use a leas wo
la ge independen pai s o case and con ol coho s, one
employing a es sample and a subsequen eplica ion
( alida ion) sample. The e a e only ou s udies on he
A10398G polymo phism which ha e employed mo e
Salas e al. BMC Cance 2014, 14:659 Page 2 o 11
h p://www.biomedcen al.com/1471-2407/14/659
han a single case–con ol sample, namely he ones by
Can e e al. [12], Se iawan e al. [18], Mosque a-Miguel
e al. [5], and F ancis e al. [14]. The i s o hese s udies
claimed some associa ion in only one coho (bu no as-
socia ion in hei o he wo coho s), which on i s own
does no pe mi ea u ing he polymo phism as being as-
socia ed wi h b eas cance , al hough he au ho s ha e
decided o he wise and u ned his in o a s o y; see
below.
Se iawan e al. [18] analyzed h ee di e en coho s o
U.S. Ame ican pa ien s wi h hei espec i e con ols,
and hey also ca ied ou a join analysis o he h ee co-
ho s. This ep esen s he e o e he la ges analysis o
da e on b eas cance and he A10398G polymo phism
(Table 1). They did no ind any s a is ical associa ion in
any o he coho s o he combined coho s. The s udy
o Mosque a-Miguel e al. [5] analyzed wo pai s o
Spanish coho s o a ious m DNA polymo phisms
and ound no associa ion wi h spo adic b eas cance
pa ien s.
Mos ecen ly, F ancis e al. [14] analyzed a la ge co-
ho o cases and con ols in h ee D a idian popula ions
om Sou h India. They also ca ied ou a me a-analysis
o 16 g oups ( om published sou ces and including he
h ee no el g oups), hus comp ising a e y la ge num-
be o cases and con ols om a ious popula ions.
O e all, he esul s show no associa ion o he A10398G
polymo phism and b eas cance isk.
SNP-SNP in e ac ion and haplog oup associa ion: using
he same da a wice
Co a ubias e al. [13] analyzed a coho o pa ien s su -
e ing wi h b eas cance and a con ol g oup o heal hy
subjec s in o de o in es iga e he p esumable associ-
a ion o m DNA SNP-SNP in e ac ion wi h he disease.
Acco ding o he au ho s, he main inding o hei s udy
was he de ec ion o a highly signi ican in e ac ion be-
ween he a ian s 12308G and 10398G, wi h esul s
sugges ing ha hese a ian s inc ease he isk o a
woman de eloping b eas cance .
F om ha a icle one can immedia ely in e ha hese
au ho s ha e used exac ly he same case and con ol
samples as employed in hei p e ious a icle [19]. In-
s ead o explici ly elling he eade ha hey ha e done
so, he au ho s i s desc ibed hei sample by saying ha
“six y-nine m DNA a ian s we e geno yped in DNA
samples om 156 non-Jewish Eu opean Ame ican b eas
cance pa ien s and 260 e hnically age ma ched emale
con ols…”[13]. When one u he lea ns ha “add-
i ional de ails on subjec asce ainmen , m DNA geno-
yping, and ini ial analysis can be ound in Bai e al.
(2007)”and compa es he in o ma ion abou he sample
Table 1 Summa y o he di e en case–con ol associa ion s udies a ge ing he m DNA polymo phism m.10398A > G.
s a is ical powe was compu ed using mi Powe
Re e ence Popula ion/E hnici y No. o cases No. o con ols Powe
(%)
‘Risky’
allele
( equency o A10398) ( equency o A10398)
[12]‘A ican-Ame icans’48 (15%) 54 (6%) 15.1 None
‘A ican-Ame icans’654 (13%) 605 (9%) 98.4 A
‘Whi es’879 (80%) 760 (79%) 100 None
[13,19] Non-Jewish Eu opean Ame ican 156 (68%) 260 (79%) 66.1 G
[14] Tamil Nadu; Sou h India 279 (38%) 280 (45%) 97.8 None
Andh a P adesh, Sou h India 348 (46%) 352 (45%) 99.2 None
Ka na aka, Sou h India 89 (38%) 92 (42%) 59.1 None
[5] Spanish mainland 464 (84%) 453 (82%) 94.9 None
Cana y Island (Spain) 302 (76%) 295 (78%) 88.5 None
[20] No h India 124 (57%) 273 (44%) 85.6 A
[18]‘A ican-Ame ican’542 (7%) 282 (4%) 49.4 None
Mul i-e hnic coho 391 (7%) 460 (6%) 78.2 None
CARE and LIFE coho 524 (6%) 236 (7%) 72.1 None
[21] Polish 44 (77%) 100 (97%) 0.1 G
[22] Sou he n Chinese 28 (26.9%) 45 (39.5%) 25.3 None
[23] Bangladesh; India 24 (75%) 20 (35%) 16.7 A
[24] I aq 21 (100%) 16 (100%) 11.2 None
I aq 21 (100%) 22 (91%) 0 None
[25] Malay 101 (27%) 90 (46%) 61.4 G
Salas e al. BMC Cance 2014, 14:659 Page 3 o 11
h p://www.biomedcen al.com/1471-2407/14/659
collec ion in bo h s udies, one can inally in e ha he
g oup o subjec s analyzed in bo h s udies mus ha e
been exac ly he same. Mo eo e , al hough no clea ly
s a ed, Co a ubias e al. [13] used a subse o he
geno yping da a employed be o e; in ac , hei Table
one is a b ie eplica e o Table wo in Bai e al. [19].
Hence, he indings epo ed by Co a ubias e al. [13]
canno be conside ed as independen e idence o an
associa ion o m DNA a ian s wi h b eas cance , bu
a bes as an a emp a a eanalysis using logis ic
eg ession.
In he Bai e al. [19] s udy, he main inding was ha
ca ie s o haplog oup K showed an inc eased isk o su -
e ing spo adic b eas cance (OR = 3.03; 95%; CI 1.63-
5.63). In eali y, he au ho s a ge ed he sligh ly la ge
haplog oup U8b by using he a ian s 9055A and 12308G
o iden i ica ion. On he o he hand, he main inding o
he Co a ubias e al. [13] s udy was ha “a highly signi i-
can in e ac ion was iden i ied be ween a ian s 12308G
and 10398G (empi ical P alue = 0.0028), wi h esul s
sugges ing hese a ian s inc ease he isk o a woman
de eloping b eas cance (OR = 3.03; 95% CI 1.53-6.11)”.
No su p isingly, he a ian s 10398G and 12308G o-
ge he de ine he main Eu opean b anch o haplog oup K,
namely, K1 (which seems o make up abou 80-90% o
haplog oup U8b o K in mos o Eu ope). Compa ing
Figu e wo o Bai e al. wi h he cu en ee p esen ed
in PhyloT ee, one can in e ha as many as 42 ou o
47 samples assignable o haplog oup U8b belong o
haplog oup K1. Then he equencies o K1 in cases and
con ols can be in e ed as 27/156 (17.3%) and 15/260
(5.8%), espec i ely. In o he wo ds, he esul s epo ed
in bo h s udies, hough using di e en e minology,
equally e lec he o e ep esen a ion o haplog oup K1
in cases compa ed o con ols.
S a is ical in e ac ions can ob iously a ise unde a lo-
gis ic eg ession es when analyzing SNPs ha a e in
s ong linkage disequilib ium (as is he case wi h he
a ia ion along he whole m DNA genome). This, how-
e e , does no necessa ily ha e o be in e p e ed as an
in e ac ion pe se [26] bu a he as wo (o mo e) SNP
a ian s ha p edominan ly occu simul aneously wi hin
he subhaplog oup hey de ine.
The s a is ical e ec o his seeming in e ac ion can be
s udied by using a simula ion (see Me hods). As shown
in Figu e 1, he e is a high co ela ion (
2
> 0.93) be-
ween he s a is ical signi icance alues (based on Fishe
exac es ) ob ained when compa ing he p opo ion o
10398G, 12308G, and haplog oup K1 in cases wi h e-
spec o con ols ( ha is, he scena io o [15]). These P-
alues a e also s ongly co ela ed (
2
> 0.95) wi h hose
ob ained o he ‘pseudo-in e ac ion’be ween m SNPs
10398G and 12308G using logis ic eg ession, ha is,
he scena io o Co a ubias e al. [13].
Selec ing he mos con enien s a is ical app oach o
gene a e some posi i e inding
In he s udy by Fang e al. [22] wo sou ces o ques ion-
able choices o s a is ics a e e iden . These au ho s s ud-
ied b eas cance and wo o he kinds o cance in
coho s o pa ien s om sou he n China. The basic dis-
inc ion hey made was be ween haplog oups M and N.
They ound haplog oup M in 69/104 (66.3%) o he cases
bu in only 60/114 (52.6%) o he con ols. The P- alue
o he wo- ailed exac Fishe es o he co esponding
2 × 2 con ingency able is 0.0532, hus his could only be
conside ed as ma ginally signi ican i assuming a nom-
inal signi ican alue o 0.05. Howe e , he au ho s
p e e ed o use he wo- ailed chi-squa e es in his
case: he unco ec ed P- alue o he chi-squa e es
equals 0.0396, hus seemingly signi ican . Howe e when
employing he ecommended Ya es co ec ion (see e.g.
h p://g aphpad.com/quickcalcs/con ingency1/) he co -
ec ed P- alue (0.0549) would ha e come much close o
he exac alue. In heo y, Fishe ’s exac es in such a
con ex should always be employed as long as sizes o
he samples pe mi his because his es is alid o all
sample sizes. A chi-squa e es is accep able only wi h
e y la ge sample sizes, bea ing in mind ha he signi i-
cance alues i p o ides always cons i u e some app oxi-
ma ion. The use o a chi-squa e es hus canno be
jus i ied in Fang’s e al. [22] s udy because he sample
sizes we e qui e small he e. A chi-squa e es is a o ed
in la ge-scale genome-wide associa ion s udies in pa
o compu a ional easons, and his is well accep ed by
he scien i ic communi y. The decision o using a chi-
squa e o a Fishe ’s exac es is i ele an mos o he
ime as bo h poin in he same di ec ion. Bu he p ob-
lem comes when he chi-squa e es is chosen ins ead o
he exac es solely o he pu pose o a aining ‘signi i-
cance’which could no ha e been achie ed wi h he
exac es .
Since i ually all haplog oup M lineages bea 10398G,
whe eas in haplog oup N he equency o 10398G is
only abou 10-20%, one can expec ha in a sou he n
Chinese popula ion a ela i e excess o haplog oup M
would co ela e wi h some highe pe cen age o 10398G
in a sample. And indeed, in ha case–con ol s udy we
a e seeing a 10398G equency o 76/104 (73.1%) in
cases and 69/114 (60.5%) in con ols. The P- alue o
he wo- ailed exac Fishe es is 0.0618. The co e-
sponding unco ec ed chi squa e es deli e s a P- alue
o 0.0499 bu wi h Ya es co ec ion 0.0691, again close
o he exac alue. Fang e al. [22] ha e again chosen o
use only he unco ec ed chi-squa e es .
Fu he mo e, Fang e al. [22] selec ed one subha-
plog oup (D5) o haplog oup M ou o se e al candi-
da es (D4a, D4, D5, G, M7, M8, M10a) ha showed a
pa icula ly con as ing equency be ween cases and
Salas e al. BMC Cance 2014, 14:659 Page 4 o 11
h p://www.biomedcen al.com/1471-2407/14/659
con ols. The au ho s did no a emp o co ec o he
mul iple hypo heses (haplog oups in his case) es ed,
which would yield non-signi ican P- alues o all he
haplog oups es ed. Las bu no leas , no alida ion
coho was o e ed o he seeming associa ion. The e-
o e his s udy ailed o demons a e any associa ion o
m DNA haplog oups wi h b eas cance because o
using inapp op ia e s a is ical ools and de icien design.
The ques ionable ole o A10398G in b eas cance
The use o con using nomencla u e
The CRS [27] nucleo ide a posi ion 10398 is A. The e-
o e he co ec no a ion o he ansi ion a his pos-
i ion is A10398G o m.10398A > G ollowing he o icial
nomencla u e in medicine.
In hei b eas cance pape , Can e e al. [12] e one-
ously employed he no a ion “G10398A”, which inciden-
ally could be in e p e ed pos hoc as i i ollowed he
e olu iona y o de o nucleo ide change ( om he
ances al G o a i s de i ed A a 10398), bu his was
ce ainly no in ended by he au ho s since C7028T,
C14766T, T16189C, and T16519C all ollow he s and-
a d CRS-based s yle. The inco ec and con using no a-
ion “G10398A”was hen epea edly used in nume ous
ollow-up pape s [20]. This led o he mos pa adoxical
si ua ion ha “G10398A”has become mo e widesp ead
in use wi hin he pas i e yea s han A10398G: e.g.
Google now has ~60,600 en ies o ‘G10398A m DNA’
bu only ~3010 en ies o ‘A10398G m DNA’. This can-
no be explained by he 2012 swi ch om A10398G o
“G10398A”execu ed in PhyloT ee since his equally
ans o med he no a ion o he C7028T polymo phism
by e e sing he oles o C and T: he e a e ~4250 en-
ies o he s anda d o m ‘C7028T m DNA’bu only
~355 en ies o ‘T7028C m DNA’.
The con usion ha has se in wi h he e oneous desig-
na ion o he A10398G polymo phism is bes e lec ed by
a b ie commen on he Can e e al. [12] s udy gi en by
Benn e al. [28], whe e i is s a ed ha “The m 10398a > g
polymo phism p esen in Eu opean haplog oups J, K, and
Z has been associa ed wi h inc eased isk o in asi e b eas
cance in black women (48 cases/54 con ols and ali-
da ed in 654 cases/605 con ols) bu no in whi e women
(879 cases/760 con ols).”Lea ing aside he inaccu acies
Figu e 1 S a is ical es s be ween a i icial cases and con ols in
a simula ion-based app oach, whe e he amoun o haplog oup
K1 m DNA p o iles in cases is p og essi ely inc eased wi h
espec o con ols. (A) P- alues; (B) adjus ed P- alues (based on a
pe mu a ion p ocedu e); (C) odds a io (OR) alues. P- alues a e epo ed
o (i) he Fishe exac es applied o m SNP 10398G, 12308G, and
haplog oup K1; and (ii) o he logis ic eg ession in e ac ion be ween
m SNPs 10398G and 12308G. The pe mu a ion analysis was no ca ied
ou in he in e ac ion scena io (Figu e 1B) because mul iple es
co ec ion is no necessa y as only a single es is ca ied ou a he ime.
Salas e al. BMC Cance 2014, 14:659 Page 5 o 11
h p://www.biomedcen al.com/1471-2407/14/659
conce ning haplog oups K (which as a whole is no de-
ined by 10398G) and Z (no o Eu opean ances y), he
Can e e al. [12] s udy was misin e p e ed: i s , he
smalles case–con ol analysis did no each signi ican
alues (see below), so ha he nex one could no be
ega ded as a alida ion o he o me ; second, he
nucleo ide s a es A and G o he claimed associa ion
we e con ounded.
Co a ubias e al. [13] claimed ha “one o he in e ac-
ions, 4216C and 10398G, obse ed in his [ hei ] s udy
al hough no s a is ically signi ican a e con olling he
GWER (Table wo), was p e iously epo ed in he Can e
e al. case–con ol s udy…”This a i ma ion was un o u-
na e, because i was he CRS nucleo ide A10398 ha was
epo ed as being associa ed wi h b eas cance by Can e
e al. Fu he , he au ho s men ioned co ec ly (in ega d
o he Can e e al. s udy), ha “a syne gis ic in e ac ion
was obse ed be ween 4216C and 10398A”.
The G nucleo ide is he ul ima ely ances al nucleo ide
a 10398 wi h espec o he en i e (known) m DNA
human phylogeny, while he A nucleo ide is conside ed
o be he de i ed allele. Howe e , i is impo an o no e
ha his si e mu a es as many as 21 imes (5 om G o
A and 16 om A o G) in he basal classi ica ion ee,
Phylo ee Build 16, poin ing o independen mu a ional
e en s occu ing a di e en imes. The age o he mu a-
ion he e o e a ies depending on he a ge ed b anch
(haplog oup) in he phylogeny. The ac o being ances-
al o de i ed could be comple ely i ele an he e om
he poin o iew o i s p esumable pa hogenici y ( he e
seems o be a bias owa ds conside ing ‘ances al’alleles
as gene ally heal hy and he ‘de i ed’ones as po en ially
pa hogenic). I could be he case ha he seeming ‘an-
ces al’nucleo ide is in ac mo e ecen han he seem-
ing ‘de i ed’allele i one ocuses on a pa icula m DNA
haplog oup whe e his polymo phism has mu a ed back
o he ances al nucleo ide se e al imes. The concep o
ances al allele is also misunde s ood in he li e a u e.
Fo ins ance, Cza necka e al. [21] men ioned ha “while
he e ised Camb idge e e ence sequence [14] lis s he
wild ype base as A, he al e na e base (G) is also p e a-
len in many popula ions”. This a i ma ion e lec s a
popula misconcep ion o he CRS as being he ‘wild-
ype’sequence; in eali y, he CRS ep esen s jus one
pa icula Eu opean m DNA sequence used o no a-
ional pu poses [16]. All exis ing m DNA lineages a e
qui e dis an om he oo o he en i e m DNA ee.
The e o e, one has o be awa e ha he A10398G
polymo phism a ge ed in di e en s udies does no ne-
cessa ily e e o he same mu a ional e en s, and he e-
o e, di e en s udies could in eali y be e e ing o
di e en s a is ical associa ions. Thus, o ins ance, he
s a is ical associa ion could ha e been ound in combin-
a ion wi h ano he a ian , hen poin ing o a pa icula
haplog oup and no o he complex polyphyle ic g oup
de ined by a pa icula nucleo ide a 10398.
The eply o Bai e al. [19] o he Mosque a-Miguel’s
e al. [5] a icle es i ies o a basic misconcep ion o he
ole o haplog oups in disease s udies: “Al hough
haplog oup K is a subclade o haplog oup U in Eu opeans,
mos , i no all published a icles on haplog oups and
disease associa ion do no include haplog oup K wi hin
haplog oup U”. Solely mu a ions ha de ine monophyle ic
clades could be pinpoin ed o gene a e an e ec on he
i ness o he m DNA. U minus K is jus a conglome a e
o nine monophyle ic clades. Un o una ely, Bai e al. we e
co ec in claiming ha mos medical gene icis s pe cei e
haplog oup U as an en i y ha does no include
haplog oup K –bu a mis ake emains a mis ake whe he
i is commi ed by a majo i y o esea che s o no . This
U-K misconcep ion has i s oo in an ea ly a icle on he
classi ica ion o Eu opean m DNA [29], which was based
on limi ed m DNA in o ma ion as p o ided by RFLP
analysis a he ime.
The con lic ing signals in he s udies o Can e e al. and
Bai e al
Can e e al. [12] analyzed h ee di e en coho s o
‘A ican-Ame ican’women wi h in asi e b eas cance . In
hei pilo s udy (48 cases and 54 con ols; all ‘A ican-
Ame ican women), he au ho s did no ind any associ-
a ion o his polymo phism wi h he disease.
In a second ‘A ican Ame ican’coho (654 cases, 605
con ols) hese au ho s ound he A10398 nucleo ide sig-
ni ican ly inc eased in b eas cance pa ien s. In a hi d
coho o ‘Whi e’women (879 cases, 760 con ols) hey
did no de ec any s a is ical associa ion. The au ho s
unde s and hei indings as a “no el epidemiologic e i-
dence ha he m DNA 10398A allele in luences b eas
cance suscep ibili y in A ican-Ame ican women”.
In eply o an a icle by Mims e al. [30] on p os a e
cance (see also Ve ma e al. [31]), he esponse by Can-
e e al. [26] p o ided us wi h mo e clues abou hei
p e ious indings om 2005 [12] (since hey used he
same b eas cance coho s as in hei 2005 s udy). Thus,
one can in e ha he main associa ion signal ound by
Can e e al. [12] appea s due o he combina ion o he
a ian s A10398 and 4216C. These wo a ian s oge he
a e good ma ke s o he Eu opean haplog oups J1c8, T,
and R2 (see PhyloT ee) wi h haplog oup T being he
mos p e alen one. In o he wo ds, he s a is ical signal
epo ed by Can e e al. in [12,32] jus mi o s he exis -
ence o an inc eased componen o ma ilineal Eu opean
ances y in hei cases compa ed o hei con ols.
The e o e, he e is li le suppo o he posi i e associ-
a ion epo ed in he Can e e al. s udy i we conside
ha he phylogene ic e idence clea ly poin s o a alse
posi i e due o he con ounding e ec o popula ion
Salas e al. BMC Cance 2014, 14:659 Page 6 o 11
h p://www.biomedcen al.com/1471-2407/14/659
s a i ica ion. Con ol o he la e should be manda o y
in case–con ol s udies in o de o a oid alse posi i es,
especially when a ge ing ‘A ican-Ame icans’ o which
one migh a p io i suspec la ge di e en ial ma ilineal
admix u e p opo ions in cases and con ols [33].
In con as o he Can e s udies, Bai e al. [19] and he
ollow-up (non-independen ) Co a ubias e al. [13] s udy
led o he conclusion ha i is he 10398G a ian ha
would be ela ed o an inc ease o isk o su e b eas
cance . Conside ing he SNPs a ge ed by Bai e al., he
a ian 10398G poin ed mainly o haplog oup K1 in hei
cases and con ols (see e.g. hei Figu e wo).
Co a ubias e al. [13] epo ed an in e ac ion be-
ween 10398G and 12308G, ( he eby ‘ ecycling’ he
same indings epo ed by Bai e al.), hus e ec i ely
claiming ha K1 has an appa en s a is ical associa ion
wi h b eas cance . No e he e o e ha he Can e ’s
implici inding ( he ‘ isky’haplog oup T) was no epli-
ca ed in Co a ubias e al. [13] while ice e sa he
Co a ubias e al. [13] inding (conce ning he ‘ isky’
haplog oup K1) was no obse ed in Can e ’ss udy.
The isk o he 10398G a ian in polish b eas cance
pa ien s
Cza necka e al. [21] epo ed he associa ion o
10398G in a Polish b eas cance coho (44 cases and
100 con ols). Apa om being an unde powe ed
case–con ol s udy (Table 1), popula ion s a i ica ion
was no moni o ed. Rega ding he la e , i is impo an
o men ion ha s a i ica ion has o be measu ed em-
pi ically, and ha he ac ha con ols “…ma ched o
e hnici y and egion o esidence.”does no gua an ee
lack o s a i ica ion. The e a e in ac solid easons o
belie e ha his s udy su e ed om a s ong bias in
he es ima ion o he equency o he 10398G mu a ion
in con ols. Cza necka e al. [21] epo ed 10398G in
10 (23%) o he 44 cases ( om one medical cen e ) and
3 (3%) o he 100 con ols (which came om a geo-
g aphically dis an medical cen e ). Fishe ’s(2-sided)
exac es deli e s an ex ao dina ily low P- alue o
0.00042 o his con ingency able. I is su p ising ha
no e e ence had been made o ou o i e se s o Pol-
ish m DNA popula ion da a o o al size >3,000, which
we e a ailable in Sep embe 2008, well be o e he sub-
mission o ha a icle. No a emp was made o com-
pa e hein-housecon ol- egionda awi hanyoneo
hese da a se s, al hough one o hem [34] was aken as
he majo pa o he con ol g oup in a publica ion
submi ed wo mon hs la e [35].
In any Eu opean popula ion he e a e h ee majo hap-
log oups de ined by 10398G: haplog oups J, K1, and N1a1.
Using con ol- egion da a, one can eliably ecognize
he ollowing haplog oups by minimal mo i s: J (16069T-
16126C), K (16224C-16311C) e sus K1a (16224C-
16311C-497T) and K1c (16224C-16311C-498del), and
haplog oup N1a1b (250C alone o plus 16391A) in
whichsubhaplog oupIisnes edasi sdomina ingcom-
ponen . One hus loses N1a1a bu gains J1c8 (wi h back
mu a ed A10398), bo h o which a e e y mino and
expec ed o be equally uncommon (<0.3%). In an en-
la ged Polish con ol g oup o 414 no mal indi iduals,
Gaweda-Wale ych e al. [36] de ec ed 22 ca ie s o
haplog oup K, including 15 o i s subhaplog oup K1a
and 2 o K1c. Hence 17/22 > ¾ is a conse a i e es ima e
o he K1 p opo ion o K. When applying his ¾ ule
o es ima ing he K1 equencies in di e en samples,
we ob ain compound equency o 117/894 (13.1%) o
N1a1b, J, and K1 in he h ee Polish da a se s s o ed
in EMPOP (www.empop.o g) and equency o 47/277
(17.0%) o I, J, and K1 o he con ol g oup in
Gaweda-Wale ych e al. [36]. In o al, we hus ob ain
he (conse a i e) equency es ima e 164/1171 (14.0%)
o he occu ence o 10398G in Poland. This is also
well in ag eemen wi h an es ima e de i ed om he
Polish haplog oup equencies o Piecho a e al. [37],
who e ec i ely a ge ed haplog oups I, J, and U8b (in-
s ead o he claimed K): assuming a lowe bound o ⅔
o he K1 con ibu ion o U8b, one ob ains he es ima e
o 23/152 (15.1%) o 10398G in Poland. This da a se
se ed as he mino pa o he con ol g oup o Cza -
necka e al. [35]. Inciden ally, he s ill la ge Polish da a
se o Saxenae al.[38]wouldgi eanes ima eo 13.0%
using he same me hod o es ima ion. Howe e , in ha
a icle one can di ec ly ead o he eal equency o
10398G in his Polish da a se , iz. as 337/2006 (16.8%),
which demons a es ha ou es ima ion was e en some-
wha oo conse a i e.
We hen pe o med ou (2-sided) exac Fishe es s
o he con ol and cases da a om Cza necka e al.
[35] each compa ed agains ei he o he li e a u e
da a se s wi h coun s 164 s 1007 and 337 s 1669,
espec i ely. The con ol da a om hei s udy ecei e
P- alues o 0.00059 and 0.00004 (sic!) in hese com-
pa isons. This demons a es ha hei con ol g oup
ei he (1) is so special ha i would ha e been
manda o y o moni o popula ion s a i ica ion illage
by illage o (2) he geno yping wen w ong qui e
badly o (3) he samples had no been chosen and ag-
g ega ed in a co ec way. The e o e, i is e iden ha
he con ol g oup employed by Cza necka e al. [35]
canno ep esen he popula ion o cases and yield
nucleo ide a ian equencies ha a e comple ely un-
expec ed in iew o he pa e ns obse ed in o he
da a se s. Compa ing he equencies o he cases in
ha pape o ei he o he li e a u e da a we ge
P- aluesashighas0.122and0.309, e y a ombe-
ing signi ican .
Salas e al. BMC Cance 2014, 14:659 Page 7 o 11
h p://www.biomedcen al.com/1471-2407/14/659
The isk o he A10398 a ian in Indian b eas cance
pa ien s
Da ishi e al. [20] analyzed 124 b eas cance pa ien s
and 273 con ols; he au ho s epo ed a s a is ical asso-
cia ion o A10398 wi h b eas cance . Gi en he ac ha
hese au ho s a ge ed an Indian popula ion and ha
hey only geno yped he A10398G polymo phism, one
has o assume ha he main s a is ical signal came om
haplog oup N as a whole (10398 is one o he mu a ions
ha sepa a es he wo (mac o)-haplog oups M and N in
Asia). By way o analyzing squamous cell ca cinoma
samples (55 cases e sus 163 con ols), hey also e-
po ed a posi i e associa ion o A10398.
The equency o haplog oup N in India is highly he -
e ogeneous, as e en highligh ed by Da ishi e al. [20].
Thei cases and con ols we e selec ed om No h India
(wi hou epo ing any u he geog aphic speci ica ion).
I is no ewo hy ha hei con ol sample has a e-
quency o A10398 o 43.6% (compa ed o 57.3% in hei
cases); howe e , he equency o haplog oup N in e.g.
Guja a (No heas India) and Punjab and Kashmi
(No h India) is jus below 60% as also epo ed by he
same au ho s (see hei Figu e wo), hus nea ly ma ch-
ing he equency o haplog oup N in hei cases. This
means ha hei con ol g oup does no p ope ly ep e-
sen hei cases, and he e o e poin ing once mo e o a
alse posi i e case o associa ion.
On he o he hand, he esul s o Da ishi e al. [20]
en e in con lic wi h he ecen a icle by F ancis e al.
[14] ca ied ou on a much la ge sample o Indian
pa ien s and con ols ( h ee di e en coho s plus
me a-analysis) whe e hey ound no associa ion o he
A10398G polymo phism and b eas cance .
The isk o he A10398 a ian in b eas cance pa ien s
om Bangladesh
Recen ly, Sul ana e al. [23] analyzed a sample o only 24
b eas cance cases and 20 con ols om Bangladesh,
claiming he associa ion o A10398 and C10400 wi h
b eas cance . These au ho s he e o e a ge ed he
(mac o-haplog oup N). I is su p ising o see ha he
equency o haplog oup N in hei cases is 75% e sus
25% in con ols.
To explain a possible alse posi i e inding in he Sul-
ana e al. [23] s udy one could easily allege: (i) de icien
s a is ical powe due o hei ex emely small coho ,
and (ii) he con ounding e ec o popula ions sub-
s uc u e (gi en ha hese au ho s ha e no con olled
his possible con ounding ac o ). Thei mos ecen
s udy, Sul ana e al. [39] gi e us mo e clues abou his
in e es ing case example. In he la e s udy, hese au-
ho s analyzed exac ly he same samples as in hei 2011
a icle [23], bu ins ead o a ge ing he coding egion,
hey examined now he con ol egion. The au ho s
epo ed ha “ wo no el polymo phisms in he D-loop,
one a posi ion 16290 (T-ins) and he o he a 16293
(A-del), was highe in b eas cance pa ien s han in
con ols”. F om he sequence elec ophe og am o hei
Figu e one one disco e s ha he au ho s misaligned
hei sequences wi h espec o he CRS: hei wo
indels cons i u e in ac he well-known ansi ion
C16290T and he ans e sion A16293C. Bo h esul -
ing a ian s oge he signal he a e haplog oup A11
wi hin haplog oup N (PhyloT ee) and would he e o e
necessa ily bea he combina ion A10398-C10400 seen
in hei 2011 a icle). This haplog oup s a us en e s in
phylogene ic con lic wi h ano he mu a ion ha is e-
po ed in hei 2012 a icle; he au ho s men ioned
ha 10316G is p esen in 69% o hei cases bu no in
hei con ols. The ansi ion 10316G is a good ma ke
o haplog oup M43 and R22; bu i has no ye been
epo ed wi hin A11. Wha e e he solu ion o his
phylogene ic puzzle would be, hei da a poin o he
ac ha cases ca y an exagge a ed ep esen a ion o a
a e haplog oup (mos likely A11) ha cons i u es 75%
o hem, hus in la ing he signal gi en by haplog oup
N in hei cases. This is ano he clea demons a ion
o popula ion s a i ica ion o inadequa e selec ion o
cases and con ols in ha small Bangladesh sample.
Since he ull haplo ypes ha e no been p esen ed in ei-
he a icle, one has o ejec he conclusions d awn by
he au ho s.
The isk o he 10398G a ian in Malaysian b eas cance
pa ien s
Nadiah e al. [25] epo ed he p esence o he 10398G
a ian in 73% o b eas cance pa ien s compa ed o 54%
in con ols, bo h om Malaysia. Thei sample sizes a e
somewha unde powe ed: 101 cases and 90 con ols. They
did no use a eplica ion coho . The di ec ion o he asso-
cia ion epo ed by hese au ho s adds u he noise o he
global scena io; i was he G nucleo ide ha was ound o
be o e - ep esen ed in cases (OR = 2.29). To explain such
a phenomenon he au ho s alleged ha “( he di e ing
esul s may be due o he a iabili y o isk modi ie s ha
exis in di e se geog aphical a eas)”. I is howe e di icul
o concei e how a isk modi ie can comple ely in e he
di ec ion o he isk; i was in ac he A nucleo ide ha
was epo ed o be associa ed in o he Sou h Eas Asian
popula ions (e.g. Bangladesh and India).
No associa ion o he 10398G a ian in I aqi women
Ismaeel e al. [24] ha e ecen ly analyzed he 10398G a i-
an in 21 emales wi h b eas malignan umo s, 22 e-
males wi h b eas benign umo s and 16 heal hy emales
used as con ols. Only wo emales o he benign umo s
g oup (9%) ca ied he 10398G a ian ( hen, 0% in hei
con ols and in he b eas malignan umo coho ). The
Salas e al. BMC Cance 2014, 14:659 Page 8 o 11
h p://www.biomedcen al.com/1471-2407/14/659
equency o his a ian in hese I aqi emales is su p is-
ing when compa ing wi h o he da ase s om he coun y
whe e his a ian could each 31% [40]. Gi en he ac
ha hese I aqi pa ien s and con ols seem o ep esen a
ypical popula ion om I aq, i is mos likely ha some
me hodological e o occu ed wi h he geno yping (based
on RFLPs) o hese samples.
Publica ion bias in b eas cance isk
Se e al e iew a icles ha e been w i en since he i s
publica ion o Can e e al. [12] in ega d o he impli-
ca ions o he 10398 polymo phism in b eas cance
( oge he wi h o he cance s); s ikingly as many as
o iginal esea ch a icles [41-47].
Un o una ely, all hese su eys eph ase and summa ize
he conclusions o he o iginal a icles wi hou c i ically
in es iga ing he obus ness o he e idences. Wo se, mos
o he ime, only he posi i e indings o he li e a u e a e
highligh ed. Since 2009, and acco ding o The Web o
Science (h p://ip-science. homson eu e s.com/es/p oduc-
os/wok/); he s udies by Can e e al. [12] and Bai e al.
[19] ecei ed 109 and 86 ci a ions (que y: 28 Feb ua y
2014), whe eas o he same pe iod, he s udies o Se ia-
wan e al. [18] and Mosque a-Miguel e al. [5] showing
nega i e associa ion ecei ed 23 and 22 ci a ions, espec -
i ely. Cza necka and Ba nik [41] epo ed ha “ he i s
in e es ing and widely in es iga ed m DNA polymo phism
in he cance ield was A10398G, i s desc ibed as causa-
i e ac o in b eas cance de elopmen (50–52)”.Cu i-
ously, no e ha in he p e ious quo a ion, hei ci a ion 50
e e s o he nega i e indings o Se iawan e al. [18] bu
hey do no u he men ion o commen on his a icle,
and he nega i e indings o Mosque a-Miguel e al. [5]
a e no ci ed a all. In he ea lie e iew by Plak e al. [43]
nei he o he wo epo s wi h nega i e indings we e
ci ed. This kind o publica ion bias is ha m ul in science
because i s imula es u u e scien i ic s udies in w ong di-
ec ions and p omo es s udies su e ing om he same de-
iciencies as he p e ious ones.
Discussion
We can sugges se e al scena ios ha migh explain a
seeming genuine s a is ical associa ion be ween m DNA
a ian s and a pa icula disease. Fi s , he a ian pe se
is ully esponsible o he disease (causal a ian ). Sec-
ond, he a ge ed a ian is in linkage disequilib ium o
he eal causal a ian in he m DNA genome. Thi d,
se e al a ian s oge he loca ed on he same haplo ype
backg ound p edispose o he disease. The ole o he
pa hogenic a ian s could be addi i e o mul iplica i e,
o hei pa hogenic ole could occu in a mo e complex
epis a ic ashion (e.g. in e ac ion wi h nuclea ac o s o
wi h he en i onmen ). In he la e wo cases, es ablish-
ing seeming co ela ion o some mac o-haplog oup o
some mu a ion de ining such a mac o-haplog oup could
be i ele an wi hou conside ing all haplog oups ha
a e de ined by ecu en ins ances o he suspec mu a-
ion and wi hou na owing he scope o he mos basal
haplog oups wi hin he suspec mac o-haplog oup. In
he con ex o a case–con ol associa ion s udy, any o
he scena ios abo e equi e p ope s udy designs gua an-
eeing ha he s a is ical associa ion could be ue and
no spu ious, due o a i ac s, such as popula ion s a i i-
ca ion. Mo eo e , knowledge o he m DNA phylogeny
is always manda o y in o de o in e p e he s a is ical
indings wi h cau ion. Un o una ely, his knowledge is
s ill limi ed in mos o he s udies [7,9,48].
The a icle o Co a ubias e al. [13] ep esen s a
p ime example o misapplica ion o he classical SNP-
SNP in e ac ion es o m DNA disease s udies, and as
he ea lie s udy o Bai e al. [19] i is based on a
misconcep ion abou haplog oups and he m DNA
phylogeny. None o hese s udies showing posi i e asso-
cia ions moni o ed he possibili y o popula ion s a i i-
ca ion in hei samples (a ypical sou ce o spu ious
alse posi i e associa ions in complex disease s udies), a
ac ha could be pa icula ly ele an in admixed pop-
ula ions such as ‘A ican-Ame icans’ om U.S.A. Un-
usual equencies in con ols leading o alse posi i e
indings ha e also been epo ed in ega d o o he
seeming disease a ian in o he diseases [49]. E en in
‘non-admixed’popula ions one could expec speci ic
local pa e ns o m DNA composi ion, so ha a con ol
g oup should ha e he same numbe o ep esen a i es
pe mic o- egion as he case g oup. I seems ha in
eali y con ol g oups a e chosen as con enience sam-
ples wi hou moni o ing popula ion s a i ica ion o as
a ailable li e a u e da a ep esen ing he ‘gene al popu-
la ion’. Nei he way is op imal, so ha signals o associ-
a ion a e bound o be spu ious.
Mo eo e , mos o hese s udies appea o be s a is i-
cally unde powe ed (Table 1). Unde powe ed means ha
any posi i e inding can be explained by chance. The ac
ha so many unde powe ed independen s udies ha e
poin ed o some e idence o associa ion (con lic ing in
ega d o he nucleo ide a ian in ol ed) can also be
seen as he esul o publica ion bias. Only wo o he
ea lie s udies ha e epo ed nega i e indings. The
Se iawan e al. [18] s udy could no eplica e he Can e ’s
indings using ‘A ican-Ame ican’b eas cance women
(by way o analyzing a simila sized coho ) and did no
ind associa ion in ano he wo coho s o in hei joined
coho s analysis; whe eas he Mosque a-Miguel e al. [5]
s udy could no eplica e any inding using wo inde-
penden coho s o Eu opean ances y. I is howe e
pa adoxical ha he numbe o coho s showing nega i e
indings is highe han he numbe o coho s showing
posi i e (con lic ing) indings (see Table 1).
Salas e al. BMC Cance 2014, 14:659 Page 9 o 11
h p://www.biomedcen al.com/1471-2407/14/659