scieee Open visual document viewer

The saga of the many studies wrongly associating mitochondrial DNA with breast cancer

Salas Ellacuriaga, Antonio; García Magariños, Manuel; Logan, Ian; Bandelt, Hans-Jürgen

Abstract

Background A large body of genetic research has focused on the potential role that mitochondrial DNA (mtDNA) variants might play on the predisposition to common and complex (multi-factorial) diseases. It has been argued however that many of these studies could be inconclusive due to artifacts related to genotyping errors or inadequate design. Methods Analyses of the data published in case–control breast cancer association studies have been performed using a phylogenetic-based approach. Variation observed in these studies has been interpreted in the light of data available on public resources, which now include over >27,000 complete mitochondrial sequences and the worldwide phylogeny determined by these mitogenomes. Complementary analyses were carried out using public datasets of partial mtDNA sequences, mainly corresponding to control-region segments. Results By way of example, we show here another kind of fallacy in these medical studies, namely, the phenomenon of SNP-SNP interaction wrongly applied to haploid data in a breast cancer study. We also reassessed the mutually conflicting studies suggesting some functional role of the non-synonymous polymorphism m.10398A > G (ND3 subunit of mitochondrial complex I) in breast cancer. In some studies, control groups were employed that showed an extremely odd haplogroup frequency spectrum compared to comparable information from much larger databases. Moreover, the use of inappropriate statistics signaled spurious “significance” in several instances. Conclusions Every case–control study should come under scrutiny in regard to the plausibility of the control-group data presented and appropriateness of the statistical methods employed; and this is best done before potential publication.

Full text

RESEARCH ARTICLE Open Access The saga o he many s udies w ongly associa ing mi ochond ial DNA wi h b eas cance An onio Salas 1* , Manuel Ga cía-Maga iños 1,2 , Ian Logan 3 and Hans-Jü gen Bandel 4 Abs ac Backg ound: A la ge body o gene ic esea ch has ocused on he po en ial ole ha mi ochond ial DNA (m DNA) a ian s migh play on he p edisposi ion o common and complex (mul i- ac o ial) diseases. I has been a gued howe e ha many o hese s udies could be inconclusi e due o a i ac s ela ed o geno yping e o s o inadequa e design. Me hods: Analyses o he da a published in case–con ol b eas cance associa ion s udies ha e been pe o med using a phylogene ic-based app oach. Va ia ion obse ed in hese s udies has been in e p e ed in he ligh o da a a ailable on public esou ces, which now include o e >27,000 comple e mi ochond ial sequences and he wo ldwide phylogeny de e mined by hese mi ogenomes. Complemen a y analyses we e ca ied ou using public da ase s o pa ial m DNA sequences, mainly co esponding o con ol- egion segmen s. Resul s: By way o example, we show he e ano he kind o allacy in hese medical s udies, namely, he phenomenon o SNP-SNP in e ac ion w ongly applied o haploid da a in a b eas cance s udy. We also eassessed he mu ually con lic ing s udies sugges ing some unc ional ole o he non-synonymous polymo phism m.10398A > G (ND3 subuni o mi ochond ial complex I) in b eas cance . In some s udies, con ol g oups we e employed ha showed an ex emely odd haplog oup equency spec um compa ed o compa able in o ma ion om much la ge da abases. Mo eo e , he use o inapp op ia e s a is ics signaled spu ious “signi icance”in se e al ins ances. Conclusions: E e y case–con ol s udy should come unde sc u iny in ega d o he plausibili y o he con ol-g oup da a p esen ed and app op ia eness o he s a is ical me hods employed; and his is bes done be o e po en ial publica ion. Keywo ds: Epis asis, SNP-SNP in e ac ion, Complex disease, Associa ion s udy, m DNA, Haplog oup Backg ound S udies on mi ochond ial DNA (m DNA) in human dis- ease ha e o en been in ensely deba ed (see some exam- ples on cance ins abili y [1-3]). On he one hand, m DNA case–con ol associa ion s udies a e equen ly a ec ed by se e al p oblems ela ed o de icien s udy designs and inapp op ia e s a is ical me hods [4,5]. On he o he hand, a phylogene ic app oach has p o en o be ex emely use ul in disco e ing a ious kinds o e - o s in hese s udies, which has o en comp omised hei esul s and conclusions [1,6-9]. The allelic iew [1] on m DNA a ia ion (as haplo ypic a ia ion) is also a common misconcep ion in m DNA s udies, whe e single a ian s om haplog oup mo i s a e ea ed as i hey we e po en ially independen disease ma ke s. A new mani es a ion o his p oblem has o do wi h he phenomenon o SNP-SNP in e ac ion applied o he in- e p e a ion o m DNA a ia ion. Epis asis, a e m coined by Ba eson [10], was i s de- ined as a masking e ec whe eby a a ian o allele a one locus p e en s he a ian a ano he locus om mani es ing i s e ec s [11]. Epis asis and gene ic in e - ac ion e e o he same phenomenon; howe e , he o me is widely used in popula ion gene ics and especially e e s o he s a is ical p ope ies o he phenomenon. In essence, SNP-SNP in e ac ion makes sense whene e wo SNPs a e loca ed in di e en (unlinked) loci, and i is gene ally applicable o he sphe e o he au osomal * Co espondence: [email p o ec ed] 1 Unidade de Xené ica, Ins i u o de Medicina Legal, and Depa amen o de Ana omía Pa olóxica e Ciencias Fo enses, Facul ad de Medicina, Uni e sidad de San iago de Compos ela, 15782 Galicia, Spain Full lis o au ho in o ma ion is a ailable a he end o he a icle © 2014 Salas e al.; licensee BioMed Cen al L d. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p://c ea i ecommons.o g/licenses/by/4.0), which pe mi s un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly c edi ed. The C ea i e Commons Public Domain Dedica ion wai e (h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle, unless o he wise s a ed. Salas e al. BMC Cance 2014, 14:659 h p://www.biomedcen al.com/1471-2407/14/659 genome. Since he whole m DNA molecule is i ually a single haploid locus, he concep o epis asis o in e ac ion, by de ini ion, does no apply o SNP-SNP in e ac ion o m DNA a ian s. Thus, he concep o haplo ype is o a di e en na u e; haplo ype e e s o a combina ion o igh ly linked a ian s in a segmen o a ch omosome (which could be a gene, o he en i e m DNA molecule). The e a e some haplo ypes (o mo e gene ally, some haplog oups) ha ha e been epo ed as con ibu ing o he isk o a disease, as i is he case o se e al o a icles discussed in he p esen s udy. I is he a ian s combined in his DNA segmen ha con- ibu e o he isk as a single locus, bu no by way o in e ac ion o independen loci; i is in ac he haplo- ype as a whole ha is used o es o associa ion, no he in e ac ion be ween he a ian s ha compose his haplo ype. Howe e , he e m epis asis has been inco - ec ly used in a ew m DNA s udies; o example, in des- igna ing he po en ial e ec o a haplo ype in cance , as exp essed by Can e e al. [12]: “epis a ic in e ac ions be ween indi idual loci wi hin he mi ochond ial gen- ome as well as nuclea -mi ochond ial gene in e ac ions should be in es iga ed”. In he p esen s udy we discuss he example [13] whe e he concep o in e ac ion has been misunde - s ood and misapplied. This example p o ides an oppo - uni y o discuss he ole o he polymo phism A10398G alias m.10398A > G (Th > Ala) in b eas cance , which has been he ocus o as many as en s udies (and many mo e e iews) in he pe iod 2005–2013. By c i ically e iewing all hese s udies we con i m he conclusion o he me a-analysis pe o med by F ancis e al. [14] ha he associa ion esul s epo ed o da e a e con adic o y and inconsis en , hus p o iding absolu ely no e idence suppo ing a ole o his m DNA polymo phism in b eas cance . Howe e , close examina ion o e e y sin- gle s udy ha claimed o ha e ound some associa ion e eals ha he e we e clea indica ions in ei he he da a p esen ed o he s a is ical me hods used ha he associa ion was spu ious. Me hods Phylogene ic me hods a e indispensable o m DNA s ud- ies o human disease [4]. In pa icula , we use PhyloT ee Build 16 (h p://www.phylo ee.o g/ [15]) and he in o - ma ion p o ided conce ning haplog oup s a us alongside wi h di ec inspec ion o he GenBank en ies ansla ed in o mu a ion mo i lis s ( ela i e o he s anda d e e - ence sequence, he CRS) a ailable in he Web (h p:// www.ianlogan.co.uk/sequences_by_g oup/haplog oup_ selec .h m; h p://www.ianlogan.co.uk/checke /accession. h m). No e howe e ha PhyloT ee employs a con using nomencla u e which does no con o m o he CRS based no a ion used in medical gene ics [16]. He e, when ocusing on a single a ian A o G a 10398, we w i e A10398 and 10398G, whe e he p e ix posi ion is ese ed o he CRS nucleo ide and he su ix o he any o he a ian , wi hou in oking ances y [16]. The s a is ical ela ionship be ween single es esul s (e.g. using Fishe ’s exac es o e e y m SNP) and logis- ic eg ession analysis (applied o he de ec ion o m SNP-m SNP in e ac ion) can be examined by way o a simple simula ion analysis. Fo such an expe imen , wo iden ical da a se s o m SNP geno ypes we e aken om he con ol g oup in he “mainland Spain”b eas cance se ies (N= 616) epo ed by Mosque a-Miguel e al. [5]; all he p o iles in one sample we e now a i icially la- beled as cases while he o he p o iles in he o he copy sample was labeled as “con ols”and s ayed unchanged. A m DNA p o ile belonging o haplog oup K1 was added one a a ime o a maximum o 120 o he subse o “cases”and se e al s a is ical analyses we e each pe - o med o “con ols” e sus “cases”. The goal o he simula ion was o c ea e scena ios whe e haplog oup K1 is p og essi ely o e ep esen ed in cases compa ed o con ols. A Fishe exac es is pe o med o he m SNPs de ining haplog oup K1 (10398G, 12308G) and o he amoun o haplog oup K1. Logis ic eg ession was also ca ied ou in o de o de ec he ‘in e ac ion’ be ween m SNPs A10398 and 12308G as done in [13]. Compu a ion o s a is ical powe was ca ied ou using mi Powe [17]. The compu a ion was done in a conse - a i e manne wi h he only aim o highligh ing hose case–con ol associa ion s udies ha a e ex emely unde powe ed. We did no conside he a pos e io i powe es ima ion op ion in mi Powe unde he assump- ion ha es ima es would be s ongly a ec ed by he ac ha popula ion s a i ica ion has se e ely in la ed e- quencies in cases e sus con ols in he di e en coho s e iewed (see Resul s below). We ins ead compu ed s a - is ical powe assuming a ‘de no o’s udy ha conside s he epo ed sample sizes in cases and con ols om he di e en coho s and he equency o 10398G in hei con ols. We assumed a conse a i e isk (odd a io) o 10398G equal o 2. Powe was compu ed using Fishe ’s Exac Tes o he mos conse a i e scena io ha con- side s 2 × 2 ables. No e ha gene ally a h eshold o a leas 80% is (consensually) conside ed o be an adequa e powe in case–con ol associa ion s udies. Resul s Nega i e indings o associa ion o A10398G wi h b eas cance Solid associa ion s udies would ideally use a leas wo la ge independen pai s o case and con ol coho s, one employing a es sample and a subsequen eplica ion ( alida ion) sample. The e a e only ou s udies on he A10398G polymo phism which ha e employed mo e Salas e al. BMC Cance 2014, 14:659 Page 2 o 11 h p://www.biomedcen al.com/1471-2407/14/659 han a single case–con ol sample, namely he ones by Can e e al. [12], Se iawan e al. [18], Mosque a-Miguel e al. [5], and F ancis e al. [14]. The i s o hese s udies claimed some associa ion in only one coho (bu no as- socia ion in hei o he wo coho s), which on i s own does no pe mi ea u ing he polymo phism as being as- socia ed wi h b eas cance , al hough he au ho s ha e decided o he wise and u ned his in o a s o y; see below. Se iawan e al. [18] analyzed h ee di e en coho s o U.S. Ame ican pa ien s wi h hei espec i e con ols, and hey also ca ied ou a join analysis o he h ee co- ho s. This ep esen s he e o e he la ges analysis o da e on b eas cance and he A10398G polymo phism (Table 1). They did no ind any s a is ical associa ion in any o he coho s o he combined coho s. The s udy o Mosque a-Miguel e al. [5] analyzed wo pai s o Spanish coho s o a ious m DNA polymo phisms and ound no associa ion wi h spo adic b eas cance pa ien s. Mos ecen ly, F ancis e al. [14] analyzed a la ge co- ho o cases and con ols in h ee D a idian popula ions om Sou h India. They also ca ied ou a me a-analysis o 16 g oups ( om published sou ces and including he h ee no el g oups), hus comp ising a e y la ge num- be o cases and con ols om a ious popula ions. O e all, he esul s show no associa ion o he A10398G polymo phism and b eas cance isk. SNP-SNP in e ac ion and haplog oup associa ion: using he same da a wice Co a ubias e al. [13] analyzed a coho o pa ien s su - e ing wi h b eas cance and a con ol g oup o heal hy subjec s in o de o in es iga e he p esumable associ- a ion o m DNA SNP-SNP in e ac ion wi h he disease. Acco ding o he au ho s, he main inding o hei s udy was he de ec ion o a highly signi ican in e ac ion be- ween he a ian s 12308G and 10398G, wi h esul s sugges ing ha hese a ian s inc ease he isk o a woman de eloping b eas cance . F om ha a icle one can immedia ely in e ha hese au ho s ha e used exac ly he same case and con ol samples as employed in hei p e ious a icle [19]. In- s ead o explici ly elling he eade ha hey ha e done so, he au ho s i s desc ibed hei sample by saying ha “six y-nine m DNA a ian s we e geno yped in DNA samples om 156 non-Jewish Eu opean Ame ican b eas cance pa ien s and 260 e hnically age ma ched emale con ols…”[13]. When one u he lea ns ha “add- i ional de ails on subjec asce ainmen , m DNA geno- yping, and ini ial analysis can be ound in Bai e al. (2007)”and compa es he in o ma ion abou he sample Table 1 Summa y o he di e en case–con ol associa ion s udies a ge ing he m DNA polymo phism m.10398A > G. s a is ical powe was compu ed using mi Powe Re e ence Popula ion/E hnici y No. o cases No. o con ols Powe (%) ‘Risky’ allele ( equency o A10398) ( equency o A10398) [12]‘A ican-Ame icans’48 (15%) 54 (6%) 15.1 None ‘A ican-Ame icans’654 (13%) 605 (9%) 98.4 A ‘Whi es’879 (80%) 760 (79%) 100 None [13,19] Non-Jewish Eu opean Ame ican 156 (68%) 260 (79%) 66.1 G [14] Tamil Nadu; Sou h India 279 (38%) 280 (45%) 97.8 None Andh a P adesh, Sou h India 348 (46%) 352 (45%) 99.2 None Ka na aka, Sou h India 89 (38%) 92 (42%) 59.1 None [5] Spanish mainland 464 (84%) 453 (82%) 94.9 None Cana y Island (Spain) 302 (76%) 295 (78%) 88.5 None [20] No h India 124 (57%) 273 (44%) 85.6 A [18]‘A ican-Ame ican’542 (7%) 282 (4%) 49.4 None Mul i-e hnic coho 391 (7%) 460 (6%) 78.2 None CARE and LIFE coho 524 (6%) 236 (7%) 72.1 None [21] Polish 44 (77%) 100 (97%) 0.1 G [22] Sou he n Chinese 28 (26.9%) 45 (39.5%) 25.3 None [23] Bangladesh; India 24 (75%) 20 (35%) 16.7 A [24] I aq 21 (100%) 16 (100%) 11.2 None I aq 21 (100%) 22 (91%) 0 None [25] Malay 101 (27%) 90 (46%) 61.4 G Salas e al. BMC Cance 2014, 14:659 Page 3 o 11 h p://www.biomedcen al.com/1471-2407/14/659 collec ion in bo h s udies, one can inally in e ha he g oup o subjec s analyzed in bo h s udies mus ha e been exac ly he same. Mo eo e , al hough no clea ly s a ed, Co a ubias e al. [13] used a subse o he geno yping da a employed be o e; in ac , hei Table one is a b ie eplica e o Table wo in Bai e al. [19]. Hence, he indings epo ed by Co a ubias e al. [13] canno be conside ed as independen e idence o an associa ion o m DNA a ian s wi h b eas cance , bu a bes as an a emp a a eanalysis using logis ic eg ession. In he Bai e al. [19] s udy, he main inding was ha ca ie s o haplog oup K showed an inc eased isk o su - e ing spo adic b eas cance (OR = 3.03; 95%; CI 1.63- 5.63). In eali y, he au ho s a ge ed he sligh ly la ge haplog oup U8b by using he a ian s 9055A and 12308G o iden i ica ion. On he o he hand, he main inding o he Co a ubias e al. [13] s udy was ha “a highly signi i- can in e ac ion was iden i ied be ween a ian s 12308G and 10398G (empi ical P alue = 0.0028), wi h esul s sugges ing hese a ian s inc ease he isk o a woman de eloping b eas cance (OR = 3.03; 95% CI 1.53-6.11)”. No su p isingly, he a ian s 10398G and 12308G o- ge he de ine he main Eu opean b anch o haplog oup K, namely, K1 (which seems o make up abou 80-90% o haplog oup U8b o K in mos o Eu ope). Compa ing Figu e wo o Bai e al. wi h he cu en ee p esen ed in PhyloT ee, one can in e ha as many as 42 ou o 47 samples assignable o haplog oup U8b belong o haplog oup K1. Then he equencies o K1 in cases and con ols can be in e ed as 27/156 (17.3%) and 15/260 (5.8%), espec i ely. In o he wo ds, he esul s epo ed in bo h s udies, hough using di e en e minology, equally e lec he o e ep esen a ion o haplog oup K1 in cases compa ed o con ols. S a is ical in e ac ions can ob iously a ise unde a lo- gis ic eg ession es when analyzing SNPs ha a e in s ong linkage disequilib ium (as is he case wi h he a ia ion along he whole m DNA genome). This, how- e e , does no necessa ily ha e o be in e p e ed as an in e ac ion pe se [26] bu a he as wo (o mo e) SNP a ian s ha p edominan ly occu simul aneously wi hin he subhaplog oup hey de ine. The s a is ical e ec o his seeming in e ac ion can be s udied by using a simula ion (see Me hods). As shown in Figu e 1, he e is a high co ela ion ( 2 > 0.93) be- ween he s a is ical signi icance alues (based on Fishe exac es ) ob ained when compa ing he p opo ion o 10398G, 12308G, and haplog oup K1 in cases wi h e- spec o con ols ( ha is, he scena io o [15]). These P- alues a e also s ongly co ela ed ( 2 > 0.95) wi h hose ob ained o he ‘pseudo-in e ac ion’be ween m SNPs 10398G and 12308G using logis ic eg ession, ha is, he scena io o Co a ubias e al. [13]. Selec ing he mos con enien s a is ical app oach o gene a e some posi i e inding In he s udy by Fang e al. [22] wo sou ces o ques ion- able choices o s a is ics a e e iden . These au ho s s ud- ied b eas cance and wo o he kinds o cance in coho s o pa ien s om sou he n China. The basic dis- inc ion hey made was be ween haplog oups M and N. They ound haplog oup M in 69/104 (66.3%) o he cases bu in only 60/114 (52.6%) o he con ols. The P- alue o he wo- ailed exac Fishe es o he co esponding 2 × 2 con ingency able is 0.0532, hus his could only be conside ed as ma ginally signi ican i assuming a nom- inal signi ican alue o 0.05. Howe e , he au ho s p e e ed o use he wo- ailed chi-squa e es in his case: he unco ec ed P- alue o he chi-squa e es equals 0.0396, hus seemingly signi ican . Howe e when employing he ecommended Ya es co ec ion (see e.g. h p://g aphpad.com/quickcalcs/con ingency1/) he co - ec ed P- alue (0.0549) would ha e come much close o he exac alue. In heo y, Fishe ’s exac es in such a con ex should always be employed as long as sizes o he samples pe mi his because his es is alid o all sample sizes. A chi-squa e es is accep able only wi h e y la ge sample sizes, bea ing in mind ha he signi i- cance alues i p o ides always cons i u e some app oxi- ma ion. The use o a chi-squa e es hus canno be jus i ied in Fang’s e al. [22] s udy because he sample sizes we e qui e small he e. A chi-squa e es is a o ed in la ge-scale genome-wide associa ion s udies in pa o compu a ional easons, and his is well accep ed by he scien i ic communi y. The decision o using a chi- squa e o a Fishe ’s exac es is i ele an mos o he ime as bo h poin in he same di ec ion. Bu he p ob- lem comes when he chi-squa e es is chosen ins ead o he exac es solely o he pu pose o a aining ‘signi i- cance’which could no ha e been achie ed wi h he exac es . Since i ually all haplog oup M lineages bea 10398G, whe eas in haplog oup N he equency o 10398G is only abou 10-20%, one can expec ha in a sou he n Chinese popula ion a ela i e excess o haplog oup M would co ela e wi h some highe pe cen age o 10398G in a sample. And indeed, in ha case–con ol s udy we a e seeing a 10398G equency o 76/104 (73.1%) in cases and 69/114 (60.5%) in con ols. The P- alue o he wo- ailed exac Fishe es is 0.0618. The co e- sponding unco ec ed chi squa e es deli e s a P- alue o 0.0499 bu wi h Ya es co ec ion 0.0691, again close o he exac alue. Fang e al. [22] ha e again chosen o use only he unco ec ed chi-squa e es . Fu he mo e, Fang e al. [22] selec ed one subha- plog oup (D5) o haplog oup M ou o se e al candi- da es (D4a, D4, D5, G, M7, M8, M10a) ha showed a pa icula ly con as ing equency be ween cases and Salas e al. BMC Cance 2014, 14:659 Page 4 o 11 h p://www.biomedcen al.com/1471-2407/14/659 con ols. The au ho s did no a emp o co ec o he mul iple hypo heses (haplog oups in his case) es ed, which would yield non-signi ican P- alues o all he haplog oups es ed. Las bu no leas , no alida ion coho was o e ed o he seeming associa ion. The e- o e his s udy ailed o demons a e any associa ion o m DNA haplog oups wi h b eas cance because o using inapp op ia e s a is ical ools and de icien design. The ques ionable ole o A10398G in b eas cance The use o con using nomencla u e The CRS [27] nucleo ide a posi ion 10398 is A. The e- o e he co ec no a ion o he ansi ion a his pos- i ion is A10398G o m.10398A > G ollowing he o icial nomencla u e in medicine. In hei b eas cance pape , Can e e al. [12] e one- ously employed he no a ion “G10398A”, which inciden- ally could be in e p e ed pos hoc as i i ollowed he e olu iona y o de o nucleo ide change ( om he ances al G o a i s de i ed A a 10398), bu his was ce ainly no in ended by he au ho s since C7028T, C14766T, T16189C, and T16519C all ollow he s and- a d CRS-based s yle. The inco ec and con using no a- ion “G10398A”was hen epea edly used in nume ous ollow-up pape s [20]. This led o he mos pa adoxical si ua ion ha “G10398A”has become mo e widesp ead in use wi hin he pas i e yea s han A10398G: e.g. Google now has ~60,600 en ies o ‘G10398A m DNA’ bu only ~3010 en ies o ‘A10398G m DNA’. This can- no be explained by he 2012 swi ch om A10398G o “G10398A”execu ed in PhyloT ee since his equally ans o med he no a ion o he C7028T polymo phism by e e sing he oles o C and T: he e a e ~4250 en- ies o he s anda d o m ‘C7028T m DNA’bu only ~355 en ies o ‘T7028C m DNA’. The con usion ha has se in wi h he e oneous desig- na ion o he A10398G polymo phism is bes e lec ed by a b ie commen on he Can e e al. [12] s udy gi en by Benn e al. [28], whe e i is s a ed ha “The m 10398a > g polymo phism p esen in Eu opean haplog oups J, K, and Z has been associa ed wi h inc eased isk o in asi e b eas cance in black women (48 cases/54 con ols and ali- da ed in 654 cases/605 con ols) bu no in whi e women (879 cases/760 con ols).”Lea ing aside he inaccu acies Figu e 1 S a is ical es s be ween a i icial cases and con ols in a simula ion-based app oach, whe e he amoun o haplog oup K1 m DNA p o iles in cases is p og essi ely inc eased wi h espec o con ols. (A) P- alues; (B) adjus ed P- alues (based on a pe mu a ion p ocedu e); (C) odds a io (OR) alues. P- alues a e epo ed o (i) he Fishe exac es applied o m SNP 10398G, 12308G, and haplog oup K1; and (ii) o he logis ic eg ession in e ac ion be ween m SNPs 10398G and 12308G. The pe mu a ion analysis was no ca ied ou in he in e ac ion scena io (Figu e 1B) because mul iple es co ec ion is no necessa y as only a single es is ca ied ou a he ime. Salas e al. BMC Cance 2014, 14:659 Page 5 o 11 h p://www.biomedcen al.com/1471-2407/14/659 conce ning haplog oups K (which as a whole is no de- ined by 10398G) and Z (no o Eu opean ances y), he Can e e al. [12] s udy was misin e p e ed: i s , he smalles case–con ol analysis did no each signi ican alues (see below), so ha he nex one could no be ega ded as a alida ion o he o me ; second, he nucleo ide s a es A and G o he claimed associa ion we e con ounded. Co a ubias e al. [13] claimed ha “one o he in e ac- ions, 4216C and 10398G, obse ed in his [ hei ] s udy al hough no s a is ically signi ican a e con olling he GWER (Table wo), was p e iously epo ed in he Can e e al. case–con ol s udy…”This a i ma ion was un o u- na e, because i was he CRS nucleo ide A10398 ha was epo ed as being associa ed wi h b eas cance by Can e e al. Fu he , he au ho s men ioned co ec ly (in ega d o he Can e e al. s udy), ha “a syne gis ic in e ac ion was obse ed be ween 4216C and 10398A”. The G nucleo ide is he ul ima ely ances al nucleo ide a 10398 wi h espec o he en i e (known) m DNA human phylogeny, while he A nucleo ide is conside ed o be he de i ed allele. Howe e , i is impo an o no e ha his si e mu a es as many as 21 imes (5 om G o A and 16 om A o G) in he basal classi ica ion ee, Phylo ee Build 16, poin ing o independen mu a ional e en s occu ing a di e en imes. The age o he mu a- ion he e o e a ies depending on he a ge ed b anch (haplog oup) in he phylogeny. The ac o being ances- al o de i ed could be comple ely i ele an he e om he poin o iew o i s p esumable pa hogenici y ( he e seems o be a bias owa ds conside ing ‘ances al’alleles as gene ally heal hy and he ‘de i ed’ones as po en ially pa hogenic). I could be he case ha he seeming ‘an- ces al’nucleo ide is in ac mo e ecen han he seem- ing ‘de i ed’allele i one ocuses on a pa icula m DNA haplog oup whe e his polymo phism has mu a ed back o he ances al nucleo ide se e al imes. The concep o ances al allele is also misunde s ood in he li e a u e. Fo ins ance, Cza necka e al. [21] men ioned ha “while he e ised Camb idge e e ence sequence [14] lis s he wild ype base as A, he al e na e base (G) is also p e a- len in many popula ions”. This a i ma ion e lec s a popula misconcep ion o he CRS as being he ‘wild- ype’sequence; in eali y, he CRS ep esen s jus one pa icula Eu opean m DNA sequence used o no a- ional pu poses [16]. All exis ing m DNA lineages a e qui e dis an om he oo o he en i e m DNA ee. The e o e, one has o be awa e ha he A10398G polymo phism a ge ed in di e en s udies does no ne- cessa ily e e o he same mu a ional e en s, and he e- o e, di e en s udies could in eali y be e e ing o di e en s a is ical associa ions. Thus, o ins ance, he s a is ical associa ion could ha e been ound in combin- a ion wi h ano he a ian , hen poin ing o a pa icula haplog oup and no o he complex polyphyle ic g oup de ined by a pa icula nucleo ide a 10398. The eply o Bai e al. [19] o he Mosque a-Miguel’s e al. [5] a icle es i ies o a basic misconcep ion o he ole o haplog oups in disease s udies: “Al hough haplog oup K is a subclade o haplog oup U in Eu opeans, mos , i no all published a icles on haplog oups and disease associa ion do no include haplog oup K wi hin haplog oup U”. Solely mu a ions ha de ine monophyle ic clades could be pinpoin ed o gene a e an e ec on he i ness o he m DNA. U minus K is jus a conglome a e o nine monophyle ic clades. Un o una ely, Bai e al. we e co ec in claiming ha mos medical gene icis s pe cei e haplog oup U as an en i y ha does no include haplog oup K –bu a mis ake emains a mis ake whe he i is commi ed by a majo i y o esea che s o no . This U-K misconcep ion has i s oo in an ea ly a icle on he classi ica ion o Eu opean m DNA [29], which was based on limi ed m DNA in o ma ion as p o ided by RFLP analysis a he ime. The con lic ing signals in he s udies o Can e e al. and Bai e al Can e e al. [12] analyzed h ee di e en coho s o ‘A ican-Ame ican’women wi h in asi e b eas cance . In hei pilo s udy (48 cases and 54 con ols; all ‘A ican- Ame ican women), he au ho s did no ind any associ- a ion o his polymo phism wi h he disease. In a second ‘A ican Ame ican’coho (654 cases, 605 con ols) hese au ho s ound he A10398 nucleo ide sig- ni ican ly inc eased in b eas cance pa ien s. In a hi d coho o ‘Whi e’women (879 cases, 760 con ols) hey did no de ec any s a is ical associa ion. The au ho s unde s and hei indings as a “no el epidemiologic e i- dence ha he m DNA 10398A allele in luences b eas cance suscep ibili y in A ican-Ame ican women”. In eply o an a icle by Mims e al. [30] on p os a e cance (see also Ve ma e al. [31]), he esponse by Can- e e al. [26] p o ided us wi h mo e clues abou hei p e ious indings om 2005 [12] (since hey used he same b eas cance coho s as in hei 2005 s udy). Thus, one can in e ha he main associa ion signal ound by Can e e al. [12] appea s due o he combina ion o he a ian s A10398 and 4216C. These wo a ian s oge he a e good ma ke s o he Eu opean haplog oups J1c8, T, and R2 (see PhyloT ee) wi h haplog oup T being he mos p e alen one. In o he wo ds, he s a is ical signal epo ed by Can e e al. in [12,32] jus mi o s he exis - ence o an inc eased componen o ma ilineal Eu opean ances y in hei cases compa ed o hei con ols. The e o e, he e is li le suppo o he posi i e associ- a ion epo ed in he Can e e al. s udy i we conside ha he phylogene ic e idence clea ly poin s o a alse posi i e due o he con ounding e ec o popula ion Salas e al. BMC Cance 2014, 14:659 Page 6 o 11 h p://www.biomedcen al.com/1471-2407/14/659 s a i ica ion. Con ol o he la e should be manda o y in case–con ol s udies in o de o a oid alse posi i es, especially when a ge ing ‘A ican-Ame icans’ o which one migh a p io i suspec la ge di e en ial ma ilineal admix u e p opo ions in cases and con ols [33]. In con as o he Can e s udies, Bai e al. [19] and he ollow-up (non-independen ) Co a ubias e al. [13] s udy led o he conclusion ha i is he 10398G a ian ha would be ela ed o an inc ease o isk o su e b eas cance . Conside ing he SNPs a ge ed by Bai e al., he a ian 10398G poin ed mainly o haplog oup K1 in hei cases and con ols (see e.g. hei Figu e wo). Co a ubias e al. [13] epo ed an in e ac ion be- ween 10398G and 12308G, ( he eby ‘ ecycling’ he same indings epo ed by Bai e al.), hus e ec i ely claiming ha K1 has an appa en s a is ical associa ion wi h b eas cance . No e he e o e ha he Can e ’s implici inding ( he ‘ isky’haplog oup T) was no epli- ca ed in Co a ubias e al. [13] while ice e sa he Co a ubias e al. [13] inding (conce ning he ‘ isky’ haplog oup K1) was no obse ed in Can e ’ss udy. The isk o he 10398G a ian in polish b eas cance pa ien s Cza necka e al. [21] epo ed he associa ion o 10398G in a Polish b eas cance coho (44 cases and 100 con ols). Apa om being an unde powe ed case–con ol s udy (Table 1), popula ion s a i ica ion was no moni o ed. Rega ding he la e , i is impo an o men ion ha s a i ica ion has o be measu ed em- pi ically, and ha he ac ha con ols “…ma ched o e hnici y and egion o esidence.”does no gua an ee lack o s a i ica ion. The e a e in ac solid easons o belie e ha his s udy su e ed om a s ong bias in he es ima ion o he equency o he 10398G mu a ion in con ols. Cza necka e al. [21] epo ed 10398G in 10 (23%) o he 44 cases ( om one medical cen e ) and 3 (3%) o he 100 con ols (which came om a geo- g aphically dis an medical cen e ). Fishe ’s(2-sided) exac es deli e s an ex ao dina ily low P- alue o 0.00042 o his con ingency able. I is su p ising ha no e e ence had been made o ou o i e se s o Pol- ish m DNA popula ion da a o o al size >3,000, which we e a ailable in Sep embe 2008, well be o e he sub- mission o ha a icle. No a emp was made o com- pa e hein-housecon ol- egionda awi hanyoneo hese da a se s, al hough one o hem [34] was aken as he majo pa o he con ol g oup in a publica ion submi ed wo mon hs la e [35]. In any Eu opean popula ion he e a e h ee majo hap- log oups de ined by 10398G: haplog oups J, K1, and N1a1. Using con ol- egion da a, one can eliably ecognize he ollowing haplog oups by minimal mo i s: J (16069T- 16126C), K (16224C-16311C) e sus K1a (16224C- 16311C-497T) and K1c (16224C-16311C-498del), and haplog oup N1a1b (250C alone o plus 16391A) in whichsubhaplog oupIisnes edasi sdomina ingcom- ponen . One hus loses N1a1a bu gains J1c8 (wi h back mu a ed A10398), bo h o which a e e y mino and expec ed o be equally uncommon (<0.3%). In an en- la ged Polish con ol g oup o 414 no mal indi iduals, Gaweda-Wale ych e al. [36] de ec ed 22 ca ie s o haplog oup K, including 15 o i s subhaplog oup K1a and 2 o K1c. Hence 17/22 > ¾ is a conse a i e es ima e o he K1 p opo ion o K. When applying his ¾ ule o es ima ing he K1 equencies in di e en samples, we ob ain compound equency o 117/894 (13.1%) o N1a1b, J, and K1 in he h ee Polish da a se s s o ed in EMPOP (www.empop.o g) and equency o 47/277 (17.0%) o I, J, and K1 o he con ol g oup in Gaweda-Wale ych e al. [36]. In o al, we hus ob ain he (conse a i e) equency es ima e 164/1171 (14.0%) o he occu ence o 10398G in Poland. This is also well in ag eemen wi h an es ima e de i ed om he Polish haplog oup equencies o Piecho a e al. [37], who e ec i ely a ge ed haplog oups I, J, and U8b (in- s ead o he claimed K): assuming a lowe bound o ⅔ o he K1 con ibu ion o U8b, one ob ains he es ima e o 23/152 (15.1%) o 10398G in Poland. This da a se se ed as he mino pa o he con ol g oup o Cza - necka e al. [35]. Inciden ally, he s ill la ge Polish da a se o Saxenae al.[38]wouldgi eanes ima eo 13.0% using he same me hod o es ima ion. Howe e , in ha a icle one can di ec ly ead o he eal equency o 10398G in his Polish da a se , iz. as 337/2006 (16.8%), which demons a es ha ou es ima ion was e en some- wha oo conse a i e. We hen pe o med ou (2-sided) exac Fishe es s o he con ol and cases da a om Cza necka e al. [35] each compa ed agains ei he o he li e a u e da a se s wi h coun s 164 s 1007 and 337 s 1669, espec i ely. The con ol da a om hei s udy ecei e P- alues o 0.00059 and 0.00004 (sic!) in hese com- pa isons. This demons a es ha hei con ol g oup ei he (1) is so special ha i would ha e been manda o y o moni o popula ion s a i ica ion illage by illage o (2) he geno yping wen w ong qui e badly o (3) he samples had no been chosen and ag- g ega ed in a co ec way. The e o e, i is e iden ha he con ol g oup employed by Cza necka e al. [35] canno ep esen he popula ion o cases and yield nucleo ide a ian equencies ha a e comple ely un- expec ed in iew o he pa e ns obse ed in o he da a se s. Compa ing he equencies o he cases in ha pape o ei he o he li e a u e da a we ge P- aluesashighas0.122and0.309, e y a ombe- ing signi ican . Salas e al. BMC Cance 2014, 14:659 Page 7 o 11 h p://www.biomedcen al.com/1471-2407/14/659 The isk o he A10398 a ian in Indian b eas cance pa ien s Da ishi e al. [20] analyzed 124 b eas cance pa ien s and 273 con ols; he au ho s epo ed a s a is ical asso- cia ion o A10398 wi h b eas cance . Gi en he ac ha hese au ho s a ge ed an Indian popula ion and ha hey only geno yped he A10398G polymo phism, one has o assume ha he main s a is ical signal came om haplog oup N as a whole (10398 is one o he mu a ions ha sepa a es he wo (mac o)-haplog oups M and N in Asia). By way o analyzing squamous cell ca cinoma samples (55 cases e sus 163 con ols), hey also e- po ed a posi i e associa ion o A10398. The equency o haplog oup N in India is highly he - e ogeneous, as e en highligh ed by Da ishi e al. [20]. Thei cases and con ols we e selec ed om No h India (wi hou epo ing any u he geog aphic speci ica ion). I is no ewo hy ha hei con ol sample has a e- quency o A10398 o 43.6% (compa ed o 57.3% in hei cases); howe e , he equency o haplog oup N in e.g. Guja a (No heas India) and Punjab and Kashmi (No h India) is jus below 60% as also epo ed by he same au ho s (see hei Figu e wo), hus nea ly ma ch- ing he equency o haplog oup N in hei cases. This means ha hei con ol g oup does no p ope ly ep e- sen hei cases, and he e o e poin ing once mo e o a alse posi i e case o associa ion. On he o he hand, he esul s o Da ishi e al. [20] en e in con lic wi h he ecen a icle by F ancis e al. [14] ca ied ou on a much la ge sample o Indian pa ien s and con ols ( h ee di e en coho s plus me a-analysis) whe e hey ound no associa ion o he A10398G polymo phism and b eas cance . The isk o he A10398 a ian in b eas cance pa ien s om Bangladesh Recen ly, Sul ana e al. [23] analyzed a sample o only 24 b eas cance cases and 20 con ols om Bangladesh, claiming he associa ion o A10398 and C10400 wi h b eas cance . These au ho s he e o e a ge ed he (mac o-haplog oup N). I is su p ising o see ha he equency o haplog oup N in hei cases is 75% e sus 25% in con ols. To explain a possible alse posi i e inding in he Sul- ana e al. [23] s udy one could easily allege: (i) de icien s a is ical powe due o hei ex emely small coho , and (ii) he con ounding e ec o popula ions sub- s uc u e (gi en ha hese au ho s ha e no con olled his possible con ounding ac o ). Thei mos ecen s udy, Sul ana e al. [39] gi e us mo e clues abou his in e es ing case example. In he la e s udy, hese au- ho s analyzed exac ly he same samples as in hei 2011 a icle [23], bu ins ead o a ge ing he coding egion, hey examined now he con ol egion. The au ho s epo ed ha “ wo no el polymo phisms in he D-loop, one a posi ion 16290 (T-ins) and he o he a 16293 (A-del), was highe in b eas cance pa ien s han in con ols”. F om he sequence elec ophe og am o hei Figu e one one disco e s ha he au ho s misaligned hei sequences wi h espec o he CRS: hei wo indels cons i u e in ac he well-known ansi ion C16290T and he ans e sion A16293C. Bo h esul - ing a ian s oge he signal he a e haplog oup A11 wi hin haplog oup N (PhyloT ee) and would he e o e necessa ily bea he combina ion A10398-C10400 seen in hei 2011 a icle). This haplog oup s a us en e s in phylogene ic con lic wi h ano he mu a ion ha is e- po ed in hei 2012 a icle; he au ho s men ioned ha 10316G is p esen in 69% o hei cases bu no in hei con ols. The ansi ion 10316G is a good ma ke o haplog oup M43 and R22; bu i has no ye been epo ed wi hin A11. Wha e e he solu ion o his phylogene ic puzzle would be, hei da a poin o he ac ha cases ca y an exagge a ed ep esen a ion o a a e haplog oup (mos likely A11) ha cons i u es 75% o hem, hus in la ing he signal gi en by haplog oup N in hei cases. This is ano he clea demons a ion o popula ion s a i ica ion o inadequa e selec ion o cases and con ols in ha small Bangladesh sample. Since he ull haplo ypes ha e no been p esen ed in ei- he a icle, one has o ejec he conclusions d awn by he au ho s. The isk o he 10398G a ian in Malaysian b eas cance pa ien s Nadiah e al. [25] epo ed he p esence o he 10398G a ian in 73% o b eas cance pa ien s compa ed o 54% in con ols, bo h om Malaysia. Thei sample sizes a e somewha unde powe ed: 101 cases and 90 con ols. They did no use a eplica ion coho . The di ec ion o he asso- cia ion epo ed by hese au ho s adds u he noise o he global scena io; i was he G nucleo ide ha was ound o be o e - ep esen ed in cases (OR = 2.29). To explain such a phenomenon he au ho s alleged ha “( he di e ing esul s may be due o he a iabili y o isk modi ie s ha exis in di e se geog aphical a eas)”. I is howe e di icul o concei e how a isk modi ie can comple ely in e he di ec ion o he isk; i was in ac he A nucleo ide ha was epo ed o be associa ed in o he Sou h Eas Asian popula ions (e.g. Bangladesh and India). No associa ion o he 10398G a ian in I aqi women Ismaeel e al. [24] ha e ecen ly analyzed he 10398G a i- an in 21 emales wi h b eas malignan umo s, 22 e- males wi h b eas benign umo s and 16 heal hy emales used as con ols. Only wo emales o he benign umo s g oup (9%) ca ied he 10398G a ian ( hen, 0% in hei con ols and in he b eas malignan umo coho ). The Salas e al. BMC Cance 2014, 14:659 Page 8 o 11 h p://www.biomedcen al.com/1471-2407/14/659 equency o his a ian in hese I aqi emales is su p is- ing when compa ing wi h o he da ase s om he coun y whe e his a ian could each 31% [40]. Gi en he ac ha hese I aqi pa ien s and con ols seem o ep esen a ypical popula ion om I aq, i is mos likely ha some me hodological e o occu ed wi h he geno yping (based on RFLPs) o hese samples. Publica ion bias in b eas cance isk Se e al e iew a icles ha e been w i en since he i s publica ion o Can e e al. [12] in ega d o he impli- ca ions o he 10398 polymo phism in b eas cance ( oge he wi h o he cance s); s ikingly as many as o iginal esea ch a icles [41-47]. Un o una ely, all hese su eys eph ase and summa ize he conclusions o he o iginal a icles wi hou c i ically in es iga ing he obus ness o he e idences. Wo se, mos o he ime, only he posi i e indings o he li e a u e a e highligh ed. Since 2009, and acco ding o The Web o Science (h p://ip-science. homson eu e s.com/es/p oduc- os/wok/); he s udies by Can e e al. [12] and Bai e al. [19] ecei ed 109 and 86 ci a ions (que y: 28 Feb ua y 2014), whe eas o he same pe iod, he s udies o Se ia- wan e al. [18] and Mosque a-Miguel e al. [5] showing nega i e associa ion ecei ed 23 and 22 ci a ions, espec - i ely. Cza necka and Ba nik [41] epo ed ha “ he i s in e es ing and widely in es iga ed m DNA polymo phism in he cance ield was A10398G, i s desc ibed as causa- i e ac o in b eas cance de elopmen (50–52)”.Cu i- ously, no e ha in he p e ious quo a ion, hei ci a ion 50 e e s o he nega i e indings o Se iawan e al. [18] bu hey do no u he men ion o commen on his a icle, and he nega i e indings o Mosque a-Miguel e al. [5] a e no ci ed a all. In he ea lie e iew by Plak e al. [43] nei he o he wo epo s wi h nega i e indings we e ci ed. This kind o publica ion bias is ha m ul in science because i s imula es u u e scien i ic s udies in w ong di- ec ions and p omo es s udies su e ing om he same de- iciencies as he p e ious ones. Discussion We can sugges se e al scena ios ha migh explain a seeming genuine s a is ical associa ion be ween m DNA a ian s and a pa icula disease. Fi s , he a ian pe se is ully esponsible o he disease (causal a ian ). Sec- ond, he a ge ed a ian is in linkage disequilib ium o he eal causal a ian in he m DNA genome. Thi d, se e al a ian s oge he loca ed on he same haplo ype backg ound p edispose o he disease. The ole o he pa hogenic a ian s could be addi i e o mul iplica i e, o hei pa hogenic ole could occu in a mo e complex epis a ic ashion (e.g. in e ac ion wi h nuclea ac o s o wi h he en i onmen ). In he la e wo cases, es ablish- ing seeming co ela ion o some mac o-haplog oup o some mu a ion de ining such a mac o-haplog oup could be i ele an wi hou conside ing all haplog oups ha a e de ined by ecu en ins ances o he suspec mu a- ion and wi hou na owing he scope o he mos basal haplog oups wi hin he suspec mac o-haplog oup. In he con ex o a case–con ol associa ion s udy, any o he scena ios abo e equi e p ope s udy designs gua an- eeing ha he s a is ical associa ion could be ue and no spu ious, due o a i ac s, such as popula ion s a i i- ca ion. Mo eo e , knowledge o he m DNA phylogeny is always manda o y in o de o in e p e he s a is ical indings wi h cau ion. Un o una ely, his knowledge is s ill limi ed in mos o he s udies [7,9,48]. The a icle o Co a ubias e al. [13] ep esen s a p ime example o misapplica ion o he classical SNP- SNP in e ac ion es o m DNA disease s udies, and as he ea lie s udy o Bai e al. [19] i is based on a misconcep ion abou haplog oups and he m DNA phylogeny. None o hese s udies showing posi i e asso- cia ions moni o ed he possibili y o popula ion s a i i- ca ion in hei samples (a ypical sou ce o spu ious alse posi i e associa ions in complex disease s udies), a ac ha could be pa icula ly ele an in admixed pop- ula ions such as ‘A ican-Ame icans’ om U.S.A. Un- usual equencies in con ols leading o alse posi i e indings ha e also been epo ed in ega d o o he seeming disease a ian in o he diseases [49]. E en in ‘non-admixed’popula ions one could expec speci ic local pa e ns o m DNA composi ion, so ha a con ol g oup should ha e he same numbe o ep esen a i es pe mic o- egion as he case g oup. I seems ha in eali y con ol g oups a e chosen as con enience sam- ples wi hou moni o ing popula ion s a i ica ion o as a ailable li e a u e da a ep esen ing he ‘gene al popu- la ion’. Nei he way is op imal, so ha signals o associ- a ion a e bound o be spu ious. Mo eo e , mos o hese s udies appea o be s a is i- cally unde powe ed (Table 1). Unde powe ed means ha any posi i e inding can be explained by chance. The ac ha so many unde powe ed independen s udies ha e poin ed o some e idence o associa ion (con lic ing in ega d o he nucleo ide a ian in ol ed) can also be seen as he esul o publica ion bias. Only wo o he ea lie s udies ha e epo ed nega i e indings. The Se iawan e al. [18] s udy could no eplica e he Can e ’s indings using ‘A ican-Ame ican’b eas cance women (by way o analyzing a simila sized coho ) and did no ind associa ion in ano he wo coho s o in hei joined coho s analysis; whe eas he Mosque a-Miguel e al. [5] s udy could no eplica e any inding using wo inde- penden coho s o Eu opean ances y. I is howe e pa adoxical ha he numbe o coho s showing nega i e indings is highe han he numbe o coho s showing posi i e (con lic ing) indings (see Table 1). Salas e al. BMC Cance 2014, 14:659 Page 9 o 11 h p://www.biomedcen al.com/1471-2407/14/659