scieee Science in your language
[en] (orig)

Development and evaluations of the ancestry informative markers of the VISAGE Enhanced Tool for Appearance and Ancestry

Author: Ruiz Ramírez, Jorge; Puente Vila, María del Carmen de la; Xavier, Catarina; Ambroa Conde, Adrián; Álvarez Dios, José Antonio; Freire Aradas, Ana María; Mosquera Miguel, Ana; Ralf, Arwin; Amory, Christina; Katsara, Maria Alexandra; Khellaf, Tarek; Nothnag
Publisher: Elsevier
Year: 2023
DOI: 10.1016/j.fsigen.2023.102853
Source: https://minerva.usc.es/bitstreams/bd04f4c2-c596-4647-88de-3a93cdbcebe3/download
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
A ailable online 5 Ma ch 2023
1872-4973/© 2023 The Au ho s. Published by Else ie B.V. This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i ecommons.o g/licenses/by-
nc-nd/4.0/).
De elopmen and e alua ions o he ances y in o ma i e ma ke s o he
VISAGE Enhanced Tool o Appea ance and Ances y
☆
J. Ruiz-Ramí ez
a
,
1
, M. de la Puen e
a
,
*
,
1
, C. Xa ie
b
, A. Amb oa-Conde
a
, J. ´
Al a ez-Dios
c
,
A. F ei e-A adas
a
, A. Mosque a-Miguel
a
, A. Ral
d
, C. Amo y
b
, M.A. Ka sa a
e
, T. Khella
e
,
M. No hnagel
e
,
, E.Y.Y. Cheung
g
, T.E. G oss
g
, P.M. Schneide
g
, J. Uacyis ael
h
, S. Oli ei a
i
,
M.d.N. Klau au-Guima ˜
aes
i
, C. Ca alho-Gon ijo
i
, E. Po´
spiech
j
, W. B anicki
k
, W. Pa son
b
,
l
,
M. Kayse
d
, A. Ca acedo
m
,
n
, M.V. La eu
a
, C. Phillips
a
,
*
,
1
, on behal o he VISAGE Conso ium
2
a
Fo ensic Gene ics Uni , Ins i u e o Fo ensic Sciences, Uni e si y o San iago de Compos ela, 15782 San iago de Compos ela, Spain
b
Ins i u e o Legal Medicine, Medical Uni e si y o Innsb uck, 6020 Innsb uck, Aus ia
c
Facul y o Ma hema ics, Uni e si y o San iago de Compos ela, 15705 San iago de Compos ela, Spain
d
Depa men o Gene ic Iden i ica ion, E asmus MC, Uni e si y Medical Cen e Ro e dam, 3015 CN Ro e dam, Sou h Holland, he Ne he lands
e
Cologne Cen e o Genomics, Uni e si y o Cologne, 50823 Cologne, Ge many
Uni e si y Hospi al Cologne, 50937 Cologne, Ge many
g
Ins i u e o Legal Medicine, Facul y o Medicine and Uni e si y Clinic, Uni e si y o Cologne, 50823 Cologne, Ge many
h
Fiji Police Fo ensic Biology and DNA Labo a o y, Naso a, Su a, Fiji
i
Depa amen o Gen´
e ica e Mo ologia, Uni e sidade de B asília, B asília, DF, B azil
j
Malopolska Cen e o Bio echnology, Jagiellonian Uni e si y, 30-387 K ak´
ow, Poland
k
Ins i u e o Zoology and Biomedical Resea ch, Jagiellonian Uni e si y, 30-387 K ak´
ow, Poland
l
Fo ensic Science P og am, The Pennsyl ania S a e Uni e si y, Uni e si y Pa k, S a e College, PA 16802, USA
m
Fundaci´
on Pública Galega de Medicina Xen´
omica (FPGMX), Ins i u o de In es igaci´
on Sani a ia (IDIS),15706 San iago de Compos ela, Spain
n
Genomics G oup, CIBERER, CIMUS, Uni e si y o San iago de Compos ela, Spain
ARTICLE INFO
Keywo ds:
Bio-geog aphical ances y
Massi ely pa allel sequencing
Ances y in o ma i e ma ke s
Au osomal SNPs
Mic ohaplo ypes
Y-SNPs
X-SNPs
1000 Genomes
ABSTRACT
The VISAGE Enhanced Tool o Appea ance and Ances y (ET) has been designed o combine ma ke s o he
p edic ion o bio-geog aphical ances y plus a ange o ex e nally isible cha ac e is ics in o a single massi ely
pa allel sequencing (MPS) assay. We desc ibe he de elopmen o he ances y panel ma ke s used in ET, and he
enhanced analyses hey p o ide compa ed o p e ious MPS-based o ensic ances y assays. As well as es ablished
au osomal single nucleo ide polymo phisms (SNPs) ha di e en ia e sub-Saha an A ican, Eu opean, Eas Asian,
Sou h Asian, Na i e Ame ican, and Oceanian popula ions, ET includes au osomal SNPs able o e icien ly
di e en ia e popula ions om Middle Eas egions. The abili y o he ET au osomal ances y SNPs o dis inguish
Middle Eas popula ions om o he con inen ally de ined popula ion g oups is such ha cha ac e is ic pa e ns
o his egion can be disce ned in gene ic clus e analysis using STRUCTURE. Join clus e membe ship es i-
ma es showing indi idual co-ances y ha signals No h A ican o Eas A ican o igins we e de ec ed, o clus e
pa e ns we e seen ha indica e o igins om cen al and Eas e n egions o he Middle Eas . In addi ion o an
augmen ed panel o au osomal SNPs, ET includes panels o 85 Y-SNPs, 16 X-SNPs and 21 au osomal Mic o-
haplo ypes. The Y- and X-SNPs p o ide a dis inc me hod o ob aining ex a de ail abou co-ances y pa e ns
iden i ied in males wi h admixed backg ounds. This s udy used he 1000 Genomes admixed A ican and admixed
Ame ican sample se s o ully explo e hese enhancemen s o he analysis o indi idual co-ances y. Samples om
u ban and u al B azil wi h con as ing dis ibu ions o A ican, Eu opean, and Na i e Ame ican co-ances y we e
also s udied o gauge he e iciency o combining Y- and X-SNP da a o his pu pose. The small panel o
Mic ohaplo ypes inco po a ed in ET we e selec ed because hey showed he highes le els o haplo ype di e si y
☆
Dedica ion: This pape is dedica ed o co-au ho Pe e Ma hias Schneide , ou es eemed scien i ic colleague and iend, who sadly died du ing i s submission.
* Co esponding au ho s.
E-mail add esses: [email p o ec ed] (M. de la Puen e), [email p o ec ed], [email p o ec ed] (C. Phillips).
1
Con ibu ed equally o he s udy
2
A comple e lis o he ins i u ions and in es iga o s in ol ed in he VISAGE Conso ium is p o ided in Appendix A.
Con en s lis s a ailable a ScienceDi ec
Fo ensic Science In e na ional: Gene ics
jou nal homepage: www.else ie .com/loca e/ sigen
h ps://doi.o g/10.1016/j. sigen.2023.102853
Recei ed 3 June 2022; Recei ed in e ised o m 15 Feb ua y 2023; Accep ed 2 Ma ch 2023
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
2
amongs he se en popula ion g oups we sough o di e en ia e. Mic ohaplo ype da a was no o mally com-
bined wi h single-si e SNP geno ypes o analyse ances y. Howe e , he haplo ype sequence eads ob ained wi h
ET om hese loci c ea es an e ec i e sys em o de-con olu ing wo-con ibu o mixed DNA. We made simple
mix u e expe imen s o demons a e ha when he con ibu o s ha e di e en ances ies and he mix u e a ios
a e imbalanced (i.e., no 1:1 mix u es) he ET Mic ohaplo ype panel is an in o ma i e sys em o in e ances y
when his di e s be ween he con ibu o s.
1. In oduc ion
The VISible A ibu es h ough GEnomics (VISAGE) Conso ium was
ini ia ed in 2017 speci ically o de elop new massi ely pa allel
sequencing (MPS) ools o geno ype single nucleo ide polymo phisms
(SNPs) o he p edic ion o bio-geog aphical ances y (BGA) [1] and a
ange o ex e nally isible cha ac e is ics (EVCs) [2] ha con ibu e o
he appea ance o an uniden i ied suspec who has le con ac ace
DNA a he c ime-scene. The SNP geno yping es s o BGA and EVC
p edic ion un in pa allel o dedica ed MPS assays o age es ima ion
based on quan i a i e DNA me hyla ion analysis [3]. VISAGE used a
wo-s age p og am o de elop he MPS oolbox o DNA-based p edic-
ion o ances y, appea ance, and age. In he i s s age, wo p o o ype
Basic Tools (BT) we e c ea ed comp ising he VISAGE BT o Appea ance
and Ances y ha combined in one MPS assay, 41 ma ke s o p edic ing
eye, hai , and skin colou wi h 115 ances y-in o ma i e SNPs o analyse
BGA [4–6]; and he VISAGE BT o age es ima ion om blood ha
combined in one MPS assay 32 CpGs om i e genes [7].
Once he BT assays had been comp ehensi ely op imised and hei
o ensic pe o mance e alua ed on he Ion S5 (The mo Fishe Scien i ic)
and MiSeq (Illumina) MPS pla o ms, VISAGE mo ed o he second s age
o MPS ool design wi h much mo e ambi ious de elopmen al a ge s o
he Enhanced Tools (ET): The VISAGE ET o Appea ance and Ances y
and wo sepa a e age ools: he VISAGE ET o age es ima ion om so-
ma ic issue and he VISAGE ET o age es ima ion om semen. Fo he
VISAGE ET o Appea ance and Ances y assay, new pheno yping SNPs
we e in oduced o an expanded ange o common EVCs beyond, bu
including, eye, hai and skin colou , which we e combined wi h new
BGA SNPs. Addi ional BGA SNPs ocussed on he ollowing objec i es: i.
he e icien di e en ia ion o Middle Eas popula ion a ia ion om
o he Eu asian popula ions by selec ing an expanded panel o SNPs
ocussed on Middle Eas egions; ii. he addi ion o gonosomal SNPs (X
and Y) o ob ain mo e de ailed analysis o co-ances y pa e ns in pe -
sons wi h admixed backg ounds; iii. he inclusion o ma ke s p o iding a
sys em o es ima e he ances y o he componen s in simple, 2-way
mixed DNA, commonly encoun e ed in o ensic analyses. The ET
oolbox expanded he age es ima ion MPS sequencing o eigh combined
CpG clus e s analysing soma ic issue me hyla ion pa e ns in blood,
buccal cells and bones [8], and in a sepa a e es , 13 CpG clus e s o
analysis o semen [9,10]. The e o e, he ET assays comp ised a single
combined appea ance and ances y MPS mul iplex plus soma ic o
semen age es ima ion mul iplexes unning in pa allel wo k lows in he
same way as BT-based analyses. A key pa o he de elopmen o he
VISAGE oolbox was he design, op imisa ion and implemen a ion o an
in eg a ed in e p e a ion amewo k which includes so wa e o com-
bined s a is ical conside a ion o DNA in o ma ion p edic ing appea -
ance, age, and ances y deli e ed by he ET assays.
Fo he ET ances y panel, Middle Eas in o ma i e BGA SNPs we e
expanded om 12 o 29, bu he o e all numbe o bina y au osomal
BGA SNPs was educed by ~25%. To analyse co-ances y pa e ns in
pe sons wi h admixed backg ounds, he wo mos in o ma i e ma ke
se s complemen ing au osomal SNPs a e Y-SNPs and mi ochond ial DNA
(m DNA) SNPs. Howe e , m DNA was no conside ed o ET, as he
a ge DNA copy numbe is subs an ially highe han genomic DNA
ex ac ed om he same o ensic sample. Addi ionally, a e y la ge
numbe o SNPs would need o be geno yped. To compensa e o he lack
o m DNA da a, 16 X-SNPs we e included o analyse he ma e nal
lineage in admixed pe sons, alongside a co e se o 85 Y-SNPs o analyse
he pa e nal lineage in males. Bo h X- and Y-SNP se s p o ide highly
in o ma i e da a wi h which o compa e he co-ances y a ios es ima ed
om au osomal BGA SNPs. Las ly, 21 Mic ohaplo ypes (MHs) wi h
ances y in o ma i e p ope ies [11] we e included o imp o e he
analysis o mixed DNA om he measu emen o sequence imbalance
and/o de ec ing mo e han wo haplo ypes pe locus ac oss mul iple
MHs, when such mix u es occu .
In he cu en s udy, we ou line he selec ion o ances y ma ke s o
he VISAGE ET o Appea ance and Ances y, he pe o mance o hese
loci o ances y in e ence using es ablished s a is ical me hodology, and
he use o he specialis X-SNP, Y-SNP and MH ma ke se s added o he
ET ances y panel o co-ances y analysis and ances y-based decon-
olu ion o simple DNA mix u es.
2. Ma e ials and me hods
2.1. Selec ion o ances y ma ke s o ET
2.1.1. Au osomal BGA SNPs
The p e ious a ge ed popula ion di e en ia ions o he BT BGA
panel, which was composed en i ely o au osomal SNPs, we e Sub-
Saha an A ica (he ein A ica, unless speci ied as he geog aphically
and gene ically dis inc No h A ica o Eas A ica), Eu ope, Eas Asia,
Sou h Asia, Ame ica (i.e., Na i e Ame ican popula ions), and Oceania.
These da ase s a e abb e ia ed o AFR, EUR, EAS, SAS, AMR and OCE,
espec i ely. ET expanded he abo e popula ion di isions o include
Middle Eas popula ions (ME), loca ed in egions anging om No h
A ica bounded by he Saha a, eas wa ds o I an and sou hwa ds o-
wa ds he egions adjacen o he ho n o A ica, whe e o iginally no
dis inc ion was made be ween No h A ican a ia ion and ha shown
by o he Middle Eas popula ions when selec ing candida e BGA SNPs.
An addi ional 12 o mo e BT SNPs ha had p e iously exhibi ed s ong
allele equency con as s be ween Middle Eas popula ions and Eu o-
peans o Sou h Asians, so we e also conside ed. The main sou ce o ME-
in o ma i e SNPs was he EUROFORGEN NAME panel [12] ha p e i-
ously compiled a o al o 111 SNPs. Fig. 1 shows he p opo ion o
au osomal bina y and i-allelic SNPs in bo h BT and ET ances y panels,
indica ing au osomal SNPs comp ised 46% o he ances y ma ke s in
ET. Au osomal bina y SNP numbe s we e educed om BT o ET o all
a ge popula ion g oups, anging om a 21% educ ion o SAS o o e
87% educ ion o OCE. The numbe o i-allelic SNPs was inc eased,
bu in all ma ke s, he e was only limi ed commonali y wi h BT BGA
SNPs - i.e., no popula ion used a simple subse o p e iously compiled BT
BGA SNPs, bu each was e-con igu ed o include mo e powe ul
ances y ma ke s o compensa e o a educed numbe o au osomal
SNPs o e all, as ou lined in Fig. 1. The e was also a deg ee o adjus men
o a ied popula ion in o ma i eness, measu ed du ing sea ches by
calcula ing Popula ion Speci ic Di e gence (i.e., Shannon’s Di e gence
me ic applied o he compa ison o one popula ion wi h all o he s in he
classi ica ion sys em, he ein deno ed by: I
n AFR
; I
n EUR
; I
n EAS
; e c.) using
he Snippe SNP analysis po al, as p e iously desc ibed [13,14].
No ably, many SNPs speci ically a ge ed o di e en ia e popula ions
ou side o A ica and Oceania also had in o ma i e pa e ns o a ia ion
in bo h o hese popula ions. Many o he i-allelic SNPs selec ed we e
chosen because o abo e-a e age le els o di e gence be ween Sou h
Asia and Eu ope o allele-2 and/o allele-3.
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
3
The bulk o au osomal BGA SNPs selec ed o ET we e iden i ied om
p e ious o ensic ances y panels, using HGDP-CEPH human di e si y
panel [15–17] and 1000 Genomes Phase III SNP da a [18], (HGDP-CEPH
popula ion desc ip ions, g ouping and sample sizes as ou lined in [19];
and o 1000 Genomes popula ions in [18] – also see Sec ion 2.2.1).
Such popula ion sample se s a e inc easingly being enhanced wi h mo e
de ailed and comp ehensi e whole-genome-sequence based a ian
ca alogs. We ook ad an age o a se ies o ecen ly published s udies ha
p o ide high quali y a ian calls om highe le els o sequence
co e age o he human genome [20–22] o compile he mos up- o-da e
allele equency es ima es o each ET BGA SNP. A he same ime,
iden ical da a was collec ed o he EVC SNPs o ET o explo e whe he
addi ional SNPs can imp o e popula ion di e en ia ions beyond he
h ee o e lapping loci o appea ance and ances y analysis used in BT
and ET ( s16891982 in SLC45A2, s1426654 in SLC24A5, s12913832
in HERC2). Las ly, we analysed genome-wide pa e ns o popula ion
a ia ion in i-allelic SNPs in he human genome om de ailed sc u iny
o a ull da ase o hese ma ke s we had p e iously compiled [23]. The
allele equency da a we e hen used o es ima e and compile ma ke s
wi h he maximum I
n POP
alues and hen o balance he panel compo-
si ion by adjus ing ela i e numbe s o BGA SNPs o he con inen al
compa isons as p e iously desc ibed [4,13], bu igno ing I
n SAS
and I
n ME
calcula ions and ma ke balance.
2.1.2. Y-SNPs
A o al o 85 Y-SNPs we e selec ed o c ea e a se in ended o achie e
an op imal balance be ween de ec ing all b oadly de ined global Y-
haplog oups and p o iding addi ional esolu ion wi hin ce ain hap-
log oups, in a way ha could be in o ma i e o o ensic ances y
analysis, while a he same ime, occupying he minimum mul iplex
space in ET. Supplemen a y Fig. S1 illus a es h ee examples o ca e-
ully selec ed Y-SNPs amongs he 85 ha all belong o haplog oup R1a,
bu exhibi geog aphic equency dis ibu ions ha a e e y di e en ,
namely: R1a-Z284 =No hwes Eu ope; R1a-Z282 =Eas Eu ope; R1a-
Z93 =Wes /Cen al/Sou h Asia. The Y-SNP selec ion p ocess also made
use o compila ions o he mos in o ma i e Y-SNPs iden i ied om he
mo e ex ensi e 859 Y-SNP MPS assays designed o analyse 640 Y-hap-
log oups [24]. The genomic de ails and geog aphic dis ibu ion sum-
ma ies o he 85 Y-SNPs inco po a ed in ET a e de ailed in Table 1.
To in e p e he Y-SNP da a gene a ed, a Y-haplog oup e e ence
da abase was equi ed, and c ea ing such a da abase in ol ed he
challenging ask o compiling dispa a e published Y-SNP popula ion
da a. Al hough a lo o di e en geog aphic egions ha e been s udied
since Y-SNP geno yping became es ablished, almos e e y published
da ase has analysed di e en se s o Y-SNPs. Some s udies ha e only
ocused on b oad haplog oups, while o he s gene a ed high- esolu ion
Y-SNP da a wi hin a ce ain haplog oup. In o de o make a e e ence
da abase ha was compa ible wi h he Y-SNPs included in ET, he ge-
no ypes o each indi idual pape we e inspec ed manually, he da a ha
was compa ible was included in he da abase, and incompa ible da a
disca ded. In some cases, he absence o ce ain haplog oups in a pop-
ula ion sample could be in e ed, o example, i 100 males we e yped
o which 70 belonged o haplog oup R1b and he emaining 30 o
haplog oup I. By ex ension, he equency o all Y-SNPs belonging o any
o he haplog oup was almos ce ain o be 0, e en i hose Y-SNPs had
no been geno yped in he o iginal s udy. Nine y Y-SNP s udies, plus he
da a published by 1000 Genomes was used o c ea e a Y-haplog oup
da abase, hese s udies combined 84,269 geno yped males, o which
35,624 (42%) could be assigned o one o he haplog oups de ined by he
85 ET Y-SNPs.
The compiled Y-SNP popula ion da abase hen o med he basis o a
mapping module wi hin he VISAGE ET in e p e a i e so wa e. The
module gene a ed cha s which isualised he equency dis ibu ions o
he in e ed haplog oup in popula ions o egions co e ed by e e ence
s udies and compa ible wi h he 85 Y-SNPs. The dis ibu ion maps o
Supplemen a y Fig. S1 illus a e he e o s o make a clea dis inc ion
Fig. 1. P opo ion o BGA SNPs and ances y
ma ke s in he VISAGE Basic Tool (BT) and he
VISAGE Enhanced Tool (ET). Amongs he bi-
na y au osomal BGA SNPs, all popula ion-
indica i e se s we e educed in numbe , apa
om Middle Eas (ME) in o ma i e SNPs, which
we e mo e han doubled in numbe . The
expansion in mul iplex space dedica ed o
ances y ma ke s in ET was occupied wi h
ances y-in o ma i e Mic ohaplo ypes, Y-SNPs,
X-SNPs, and mo e i-allelic BGA SNPs. Ligh
g ey ci cles le deno e BGA SNPs e ained, da k
g ey ci cles igh no el BGA SNPs in oduced o
ET o imp o e each popula ion di e en ia ion.
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
4
Table 1
Genomic de ails and geog aphic dis ibu ion summa ies o he 85 Y-SNPs inco po a ed in ET. NA: no in o ma ion a ailable.
No. Ma ke
name
SNP-ID Posi ion
GRCh37
Posi ion
GRCh38
Subs i u ion ISOGG
Nomencla u e
Geog aphic
dis ibu ion
No. Ma ke
name
SNP-ID Posi ion
GRCh37
Posi ion
GRCh38
Subs i u ion ISOGG
Nomencla u e
Geog aphic
dis ibu ion
1 V148 s181335666 6788191 6920150 G->A A0 Cen al
A ica, Wes
A ica
43 M522 s9786714 7173143 7305102 G->A IJK
2 L1086 NA 2826312 2958271 A->T A00 Cen al A ica 44 M304 s13447352 22749853 20587967 A->C J W Asia, No h
A ica, Ho n
o A ica, S
Eu ope,
Cen al Asia,
Sou h Asia
3 V168 s191505182 17947672 15835792 G->A A1 45 M267 s9341313 22741818 20579932 T->G J1 No he n
A ica, Ho n
o A ica,
Wes Asia,
Sou h Asia
4 M31 s369315948 21739754 19577868 G->C A1a Wes A ica,
No h A ica
46 M172 s2032604 14969634 12857709 T->G J2 Sou he n
Eu ope, Wes
Asia
5 V50 s189205028 6845936 6977895 T->C A1b1a Sou he n
A ica,
Cen al A ica
47 M9 s3900 21730257 19568371 C->G K
6 M32 s558241924 21740436 19578550 T->C A1b1b Eas A ica,
Sou he n
A ica
48 M526 s2033003 23550924 21389038 A->C K2
7 M13 s3904 21722098 19560212 G->C A1b1b2b Cen al
A ica, Eas
A ica
49 M20 s3911 21733454 19571568 A->G L Sou h Asia,
Wes Asia
8 M42 s2032630 21866840 19704954 A->T BT 50 P326 s372687543 8467290 8599249 T->C LT [K1]
9 M181 s2032599 14851554 12739620 T->C B Cen al
A ica,
Sou he n
A ica, Eas
A ica
51 P256 P256 8685231 8817190 G->A M o K2b1b Nea
Oceania,
Wallacea,
Aus alia,
Remo e
Oceania ???
10 M168 s2032595 14813991 12702062 C->T CT 52 M231 s9341278 15469724 13357844 G->A N No he n
Asia, Cen al
Asia,
Ame icas
11 M145 s3848982 21717208 19555322 C->T DE 53 M46 s34442126 14922583 12810648 T->C N1a1 Sibe ia / Eas
Asia
12 M174 s2032602 14954280 12842354 T->C D Eas Asia 54 VL29 s752512309 14570424 12458624 T->C N1a1a1a1a1a NE Eu ope,
Eas e n
Eu ope,
Cen al Asia
13 F6251 NA 7681275 7813234 C->T D1a Eas Asia,
Cen al Asia
55 B479 NA 26271075 24124928 C->A N1a1a1a1a1c~ Eas Asia
14 M55 s2032621 21872738 19710852 T->C D1b Japan 56 Z1936 s774008164 21463326 19301440 C->T N1a1a1a1a2 NE Eu ope,
Eas e n
Eu ope,
Cen al Asia
15 L1378 s893924838 2828140 2960099 C->T D2 SE Asia 57 F4205 s1028202961 16331432 14219552 A->G N1a1a1a1a3a Mongolia
16 M96 s9306841 21778998 19617112 C->G E A ica, Wes
Asia,
58 B202 NA 2880546 3012505 T->C N1a1a1a1a3b Russian Fa
Eas
(con inued on nex page)
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
5
Table 1 (con inued)
No. Ma ke
name
SNP-ID Posi ion
GRCh37
Posi ion
GRCh38
Subs i u ion ISOGG
Nomencla u e
Geog aphic
dis ibu ion
No. Ma ke
name
SNP-ID Posi ion
GRCh37
Posi ion
GRCh38
Subs i u ion ISOGG
Nomencla u e
Geog aphic
dis ibu ion
Sou he n
Eu ope
17 M33 s368762706 21740450 19578564 A->C E1a Wes A ica 59 M2118 s571876713 23259624 21097738 A->G N1a1a1a1a4 Russian Fa
Eas
18 V38 s768983 6818291 6950250 C->T E1b1a Sub Saha an
A ica
60 F2930 s528311746 19080602 16968722 G->A N1b Eas Asia
19 M215 s2032654 15467824 13355944 A->G E1b1b 61 P186 s16981290 7568568 7700527 C->A O Eas Asia, SE
Asia, Sou h
Asia, Oceania
20 V32 s371254614 6932821 7064780 G->C E1b1b1a1a1b Eas A ica 62 M119 s72613040 21762685 19600799 T->G O1a SE Asia, Eas
Asia, Oceania
21 V13 s368031074 6842263 6974222 G->A E1b1b1a1b1a Sou he n
Eu ope
63 P31 s200861659 14495243 12383440 T->C O1b Sou h Asia,
SE Asia
22 M81 s2032640 21892572 19730686 C->T E1b1b1b1a No he n
A ica
64 M176 s11575897 2655180 2787139 G->A O1b2 Eas Asia
23 M123 s371143248 21764586 19602700 C->T E1b1b1b2a1 Eas A ica,
Wes Asia
65 M122 s78149062 21764674 19602788 A->G O2 Eas Asia,
Oceania
24 M75 s2032639 21890177 19728291 G->A E2 Sub Saha an
A ica
66 JST-
002611
s2075181 7546726 7678685 G->A O2a1b Eas Asia
25 P143 s4141886 14197867 12077161 G->A CF 67 P201 s2267801 2828196 2960155 T->C O2a2 Oceania, Eas
Asia
26 M130 s35284970 2734854 2866813 C->T C Cen al, No h
& SE Asia, N
Ame ica, Eas
Asia, Nea
Oceania,
Aus alia,
Remo e
Oceania
68 P295 s895530 7963031 8094990 T->G P o K2b2
27 M38 s369611932 21742158 19580272 T->G C1b3a Oceania /
Indonesia
69 M242 s8179021 15018582 12906671 C->T Q No he n
Asia, Cen al
Asia, Ame ica
28 M347 s868363758 2877479 3009438 A->G C1b3b Aus alia 70 M3 s3894 19096363 16984483 G->A Q1b1a1a Ame ica
29 M217 s2032668 15437333 13325453 A->C C2 Sou h Asia,
Sou he n Eas
Asia,
No he n Eas
Asia
71 M207 s2032658 15581983 13470103 A->G R Eu ope, Wes
Asia, Cen al
Asia, Sou h
Asia, No h
A ica,
Cen al A ica
30 P39 s887450245 14484581 12363850 G->A C2b1a1a1 No he n
Ame ica
72 M173 s2032624 15026424 12914512 A->C R1
31 M48 s373681213 21749881 19587995 A->G C2b1a1b Sibe ia /
No he n Eas
Asia
73 M420 s17250535 23473201 21311315 T->A R1a
32 M89 s2032652 21917313 19755427 C->T F 74 Z282 s112563127 15588401 13476521 T->C R1a1a1b1a Eas e n
Eu ope,
Balkan
33 M201 s2032636 15027529 12915617 G->T G Wes Asia,
Sou h-Wes
Asia, Eu ope,
Cen al Asia
75 Z284 s767265794 8717196 8849155 C->G R1a1a1b1a3a No he n
Eu ope
34 M285 s13447378 22741740 20579854 G->C G1 Sou h-Wes
Asia Cen al
Asia
76 Z93 s566323605 7552356 7684315 G->A R1a1a1b2 Sou h Asia,
Middle Eas ,
Cen al Asia
(con inued on nex page)
J. Ruiz-Ramí ez e al.

Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
6
Table 1 (con inued)
No. Ma ke
name
SNP-ID Posi ion
GRCh37
Posi ion
GRCh38
Subs i u ion ISOGG
Nomencla u e
Geog aphic
dis ibu ion
No. Ma ke
name
SNP-ID Posi ion
GRCh37
Posi ion
GRCh38
Subs i u ion ISOGG
Nomencla u e
Geog aphic
dis ibu ion
35 P287 s4116820 22072097 19910211 G->T G2 Wes Asia,
Sou h-Wes
Asia, Eu ope,
Cen al Asia
77 M343 s9786184 2887824 3019783 C->A R1b Wes e n
Eu ope
36 L901 s567848586 17844304 15732424 C->T H Sou h Asia,
Eas e n
Eu ope,
Sou h-Wes
Eu ope,
Wes e n
Eu ope
78 U106 s16981293 8796078 8928037 C->T R1b1a1b1a1a1 Wes e n
Eu ope
37 P96 s1027017284 14869743 12757813 C->A H2 Eas e n
Eu ope,
Sou h-Wes
Eu ope,
Wes e n
Eu ope
79 P312 s34276300 22157311 19995425 C->A R1b1a1b1a1a2 Wes e n
Eu ope
38 M170 s2032597 14847792 12735858 A->C I Eu ope, Wes
Asia
80 L21 s11799226 15654428 13542548 C->G R1b1a1b1a1a2c1 Wes e n
Eu ope
39 M253 s9341296 15022707 12910796 C->T I1 No h-
Eu ope, Wes
Eu ope
81 CTS1078 s567703217 7186135 7318094 G->C R1b1a1b1b Caucasus,
Balkan,
Middle Eas
40 M438 s17307294 16638804 14526924 A->G I2 Sou h Eu ope,
Cen al
Eu ope, Eas
Eu ope
82 V88 s180946844 4862861 4994820 C->T R1b1b Sub Saha an
A ica
41 M436 s17315680 18747493 16635613 G->C I2a1b No h-
Eu ope, Wes
Eu ope
83 M479 s372157627 20834667 18672781 C->T R2 Sou h Asia
42 M429 s17306671 14031334 11910628 T->A IJ 84 B254 s372295336 14102580 11981874 C->A S Oceania, Eas
Asia,
Aus alia
85 M184 s20320 14898163 12786229 G->A T Wes Asia,
Ho n o
A ica, No h
A ica,
Sou he n
Eu ope,
Sou h Asia
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
7
be ween ze o obse a ions and missing da a o hose egions lacking
geno ype obse a ions.
ET Y-SNP da a was analysed in male samples in he VISAGE S udy
popula ions and compa ed o X-SNP da a. Haplog oup assignmen s we e
made using he ex ensi e popula ion da a compiled o he ET Y-SNP
panel selec ion and used o gene a e he geog aphic dis ibu ion cha s
shown in Supplemen a y Fig. S1. We did no o mally collec Y-SNP da a
om 1KG, CEPH o Sange ME da a as his was qui e incomple e.
Fu he mo e, we chose no o make he in e ence ha all SNP da a ab-
sen om each p ojec ’s VCF iles mean he male samples all had he
Re Seq e e ence allele by de aul .
2.1.3. X-SNPs
In a p e ious unpublished su ey o X ch omosome SNP da a which
was made o compa e a ia ion ac oss he majo con inen al popula ion
g oups o he HGDP-CEPH di e si y panel om 650,000 geno yped SNPs
[25], we iden i ied a small numbe o X-SNPs wi h highly s a i ied allele
equency dis ibu ions. Se s o be ween wo o ou SNPs we e compiled
ha we e in o ma i e o AFR, EUR, EAS, AMR o OCE popula ion di -
e en ia ions o c ea e a compac X-SNP panel o 16 ma ke s dis ibu ed
ac oss he ull leng h o he X ch omosome. Fi e o hese 16 SNPs we e
egula ly spaced a ound he cen ome e bu loca ed in a egion wi h
e y low ecombina ion (Rc a es g aphically summa ised in Fig. 5 o
[26]) and so we e ea ed as a single haplo ype block. The mos ecen ly
published genomic da a wi h geno ypes o all he BGA SNPs o ET om
high sequence co e age analysis o 1000 Genomes samples [21] has
phased he SNP geno ypes in all ch omosomes, so X-SNP geno ypes om
emales we e collec ed as haplo ypes o he cen ome ic 5-SNP haplo-
ype block, and om males as single ch omosome haplo ype da a ( hus,
phased by de aul ). All o he X-SNP da a was compiled wi h he same
app oach used o au osomal a ian s, bu accoun ing o hemizygosi y
in males when es ima ing allele equencies.
In an ope a ional se ing, a o ensic ances y es using ET ha
analysed co-ances y pa e ns would compa e Y-SNP da a and single
ch omosome X-SNP geno ypes in male samples alone, so phasing in o
haplo ype combina ions would no be necessa y. We collec ed he
phased da a om 1000 Genomes emale samples in addi ion o male
geno ypes in o de o p o ide he mos comple e analysis o popula ion
a ia ion ac oss he majo popula ion g oups ep esen ed in 1000 ge-
nomes and added X-SNP geno ype da a o AMR and OCE om whole-
genome-sequence analyses o he HGDP-CEPH di e si y panel samples.
An impo an pa allel s udy was o assess he iabili y o X-SNP analysis
in he admixed A ican and admixed Ame ican popula ion samples o
1000 Genomes (labelled by his p ojec as ACB, ASW A ican and MXL,
CLM, PUR, PEL Ame ican [18]) - whe e X ch omosomes o a ied
ances al lineages a e going o be p esen in a la ge p opo ion o hese
indi iduals and a deg ee o ecombina ion may ha e dis up ed he
popula ion s a i ica ion shown by he selec ed X-SNPs in he AFR, EUR
and AMR admix u e con ibu o popula ions.
2.1.4. Mic ohaplo ypes
We chose Mic ohaplo ypes o inco po a ion in o ET om wo se s
we had p e iously designed o MPS sequence analysis [11,27] ha had
been selec ed and cha ac e ised o hei ances y in o ma i eness
p ope ies. Ca e ully selec ed ances y in o ma i e MH loci will ha e
mul iple haplo ypes wi h con as ing popula ion equencies [28], and
po en ially allow simple mixed DNA decon olu ion wi h he possibili y
o assign ances y o componen s in simple 2-way mix u es, pa icula ly
i hey a e p esen in unequal a ios [29]. F om 22 MHs o iginally
chosen, 21 we e success ully inco po a ed in o he ET assay, comp ising
8 om he MAPlex BGA panel [27], and 13 om a panel o 113 MHs
designed o o ensic iden i ica ion bu including se e al wi h ances y
in o ma i e haplo ype dis ibu ions [11]. Six o he eigh MAPlex MH
loci we e sho ened om he o iginal much longe loci con aining mo e
SNPs [29] o ensu e o ensic sensi i i y analysing deg aded DNA, by
ampli ying size- educed sequences o compa able leng h o single-si e
SNP a ge s. De ails o he SNP se s o he 21 MHs selec ed o ET and
size educ ions when made, a e ou lined in Table 2.
2.2. Re e ence and es popula ion da a
2.2.1. Public popula ion da a om human genome sequencing p ojec s
A comp ehensi e popula ion da ase o ET BGA ma ke s was
gene a ed by compiling publicly a ailable online whole-genome-
sequencing a ian da a o 3570 samples, published by h ee majo
human genome p ojec s [20–22]. This popula ion da a comp ised 2504
1000 Genomes p ojec samples (he ein 1KG) now consis ing o a e ised,
highe quali y a ian da ase based on an a e age 30x sequence
co e age [21]; 929 HGDP-CEPH human di e si y panel samples (CEPH
[20]), and 137 Middle Eas samples om he analysis o 8 popula ions
by Alma i e al. in 2021 [22], which we e e o collec i ely as he
‘Sange ME’ da ase . We also added 130 samples om he Simons
Founda ion human genome di e si y panel (SGDP [30]) excluding
samples ha o e lap wi h hose o 1KG o CEPH, and 402 samples om
he Es onian Biocen e human genome di e si y panel (EGDP [31]).
Some geno ype gaps exis in ce ain sample panels, no ably all he
i-allelic SNP geno ypes a e missing om EGDP and he e is a
wide-scale absence o many MH componen SNPs om EGDP da a. The
co e ET BGA SNP da ase cen ed on 1KG, CEPH and Sange ME SNP
geno ypes and haplo ypes, and we used his da a o c ea e a s and-
a dised popula ion e e ence se and o pe o m mos o he e alua ions
o he ET BGA SNPs’ popula ion di e en ia ion capabili ies. SGDP and
EGDP da a is included as es ing sample se s o use s o make hei own
explo a ions.
2.2.2. VISAGE in-house s udy popula ions
A ange o VISAGE pa icipan labo a o y in-house popula ion sam-
ple se s (he ein S udy popula ions) we e geno yped wi h he ET MPS
assay. These se s we e chosen o co e geog aphic gaps in unde -
ep esen ed egions, pa icula ly he Middle Eas , comp ising: 32 in-
di iduals om Mo occo; 30 om E i ea; 16 om Somalia; 30 om
Cen al I aq; 29 om he Ku dis an egion o I aq; 29 Tu kish-o igin
indi iduals esiden in Ge many; 41 om Fiji; 19 om u al B azil
(Kalunga indi iduals, Goi´
as S a e), and 16 om u ban B azil ( esiden s
o he Ci y o B asília).
In o med consen was ob ained om all S udy popula ion dono s,
which comp ised samples p e iously ob ained om: i. Mo occans esi-
den in Mad id collec ed in 2008 by he Comisa ía Gene al de Policía
Cien i íca, Mad id, wi h w i en in o med consen ob ained om dono s
ega ding he use o anonymised samples o he cha ac e isa ion o
popula ion a ia ion; ii. E i ean, Somali, Cen al I aqi, Ku dish I aqi,
and Tu kish esiden in Ge many (co-au ho s P.M.S., T.E.G.) collec ed
acco ding o he guidelines o he Decla a ion o Helsinki and app o ed
by he Ins i u ional Re iew Boa d o he Facul y o Medicine, Uni e si y
o Cologne, Ge many, e e ence no. 17–416 (da ed 16.5.2018); iii. Fijian
island samples ob ained by Fiji Police Fo ensic Biology and DNA Labo-
a o y, (co-au ho J.U.), wi h w i en in o med consen ob ained om
dono s ega ding he use o anonymised samples o he cha ac e isa ion
o popula ion a ia ion; i . B azilian samples ob ained in B azil (co-
au ho s S.O., M.K.-G., C.C.-G.) wi h e hical app o al om Uni e sidade
de B asília e e ence No. CAAE: 16542613.8.0000.0030 ( u al) and
CAAE: 72917916.3.0000.0030 (u ban).
2.2.3. Compila ion o s anda dised e e ence popula ion da ase s
A s anda dised se en-popula ion g oup e e ence da ase was con-
s uc ed o enable end-use s o make popula ion analyses independen ly
o he VISAGE ET in e p e a i e so wa e. The e e ence da ase con-
sis ed o : A icans ep esen ed by 108 1KG Yo uba om Nige ia (YRI);
Eu opeans by 99 1KG NW Eu opeans om U ah (CEU); Eas Asians by
103 1KG Han Chinese om Beijing (CHB); Sou h Asians by 103 1KG
Guja a i om Hous on (GIH); Middle Eas by 161 HGDP-CEPH Is aeli
A abs om Pales inian, D uze and Bedouin popula ions plus Alge ian
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
8
Table 2
Genomic de ails o he 21 Mic ohaplo ypes inco po a ed in ET. O iginal MH nomencla u e lis s in bold he six loci educed in size in ET designs o enhance hei
o ensic sensi i i y.
P incipal
SNPs in
he
haplo ype
Ex a
SNPs
in
MPS
ou pu
In e nal
MH
name
O iginal MH
nomencla u e
P incipal
componen
SNPs
Ex a SNPs in
Ion S5 MPS
sequence
ou pu
5′
coo dina e:
GRCh37
3′
coo dina e:
GRCh37
5′
coo dina e:
GRCh38
3′
coo dina e:
GRCh38
MH span in
nucleo ides
O iginal
MH span
4 1 1pA s28503881-
s4648788-
s72634811-
s28689700
s532405039 1529950 1529998 1594570 1594618 48
3 2 MH01 mh01KK-01 s6663840-
s58111155-
s6688969
s199565833
/
s548721351
3743319 3743391 3826755 3826827 72 259
3 - 1pD s6702428-
s12031966-
s6687440
- 106770076 106770110 106227454 106227488 34
3 - MH03 mh02KK-
134
s12469721-
s3101043-
s3111398
- 161079411 161079450 160222900 160222939 39 103
3 2 MH04 mh02KK-136 s6714835-
s6756898-
s12617010
s530973697
/
s546011313
228092389 228092459 227227673 227227743 70 70
5 1 3pB s11129981-
s11129982-
s75361533-
s11129983-
s1896565
s528474614 42924625 42924691 42883133 42883199 66
5 5 3qC s6583335-
s9848767-
s843520-
s9833841-
s965140
s559681042
/
s552643442
/
s550318827
/ s60667153
/
s183434367
196379897 196379993 196653026 196653122 96
4 1 4qD s34521178-
s4533811-
s4450974-
s61132367
s531239419 182795889 182795939 181874736 181874786 50
4 3 7pB s6951954-
s6969555-
s2158900-
s73080042
s139000977
/
s185814343
/
s552428908
25447589 25447640 25407970 25408021 51
4 1 8pA s10097211-
s80063668-
s73660014-
s7007616
s538206051 3306430 3306458 3448908 3448936 28
5 6 8pB s34821009-
s7822905-
s7836134-
s7822909-
s6474278
s577517386
/
s539800640
/
s113457629
/
s188201066
/
s113010596
/
s565537969
40664194 40664243 40806675 40806724 49
5 1 9pA s1408329-
s11789647-
s12555748-
s1535838-
s1408330
s567753466 2288647 2288718 2288647 2288718 71
3 4 10pB s11816330-
s10828819-
s4749046
s570240814
/
s536076967
/
s555668598
/
s572123381
25839394 25839446 25550465 25550517 52
3 2 MH11 mh11KK-
180
s4752778-
s74047734-
s7112918-
s4752777
s140892495
/
s555496836
1690950 1690984 1669720 1669754 34 193
(con inued on nex page)
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
9
Mozabi e - he la e sample di ided in o an eigh h e e ence popula ion
ep esen ing No h A ica in STRUCTURE analyses o Eu asians; Oce-
anians by 28 HGDP-CEPH Papuans om Bougain illea and Papua New
Guinea; Na i e Ame icans by 79 samples, comp ising 61 HGDP-CEPH
samples om Maya, Pima, Colombian and Amazonian Su ui and Ka -
i iana popula ions, supplemen ed by 18 1KG Pe u ians om Lima, Pe u
(PEL) which we had p e iously analysed o indica e no de ec able non-
Ame ican co-ances y ( om analysis o 572,743 A yme ix Human
O igins SNPs, see Table 10.5 o [32]).
The 1KG admixed popula ions, comp ising 96 A ican Ca ibbean
indi iduals in Ba bados (ACB), 61 Ame icans o A ican Ances y in SW
USA (ASW), 64 indi iduals wi h Mexican Ances y om Los Angeles
USA (MXL,) 94 Colombians om Medellin, Colombia (CLM), 104 Pue o
Ricans om Pue o Rico (PUR), and 67 o 85 PEL (i.e., wi h de ec ed co-
ances y), we e used as he es ing sample se o e alua ing he
admix u e analysis capabili ies o ET by compa ison wi h co-ances y
es ima es p o ided by he 1000 Genomes p ojec (pe sonal communi-
ca ion, Adam Au on, Albe Eins ein College o Medicine, NYC, USA).
2.3. E alua ion o ances y and co-ances y analysis using ET BGA SNPs
The e iciency o he ET au osomal BGA SNPs o in e an indi idual’s
popula ion o o igin was assessed o a se en-g oup di e en ia ion o
A ican, Eu opean, Eas Asian, Sou h Asian, Ame ican, Oceanian and
Middle Eas ances ies. Fo BGA p edic ion wi hin he ET in eg a ed
in e p e a ion amewo k, VISAGE has implemen ed dedica ed so wa e
using a s ic ly Bayesian app oach ha applies a la p io p obabili y
model and mul iple logis ic eg ession o assign a o ensic sample o one
o he abo e se en possible ances y classes. In he epo ed s udy we did
no apply he al e na i e likelihood a io analyses ha o m he co e o
he Snippe web po al [33], bu ins ead elied on STRUCTURE analysis
[34] o assess he abili y o he ET ances y ma ke s o disce n complex
ances y pa e ns in popula ion samples om he Middle Eas , as well as
popula ions ep esen ing egions whe e admix u e o a ying deg ees is
he p edominan demog aphic pa e n obse ed.
We de eloped a wo-s age nes ed STRUCTURE analysis app oach
which analysed he es popula ion se s (POPFLAG=0) wi h he e e -
ence popula ion da ase (POPFLAG=1), which consis ed o i e con i-
nen al popula ions o AFR, EUR, EAS, AMR and OCE a K:5, wi h a
second Eu asian- ocussed STRUCTURE analysis using a e e ence pop-
ula ion da ase a K:6 consis ing o AFR, EUR, SAS, EAS, ME and a six h
No h A ican (NAF) popula ion. The di ision o Middle Eas and No h
A ican ances ies ollowed he obse a ion o consis en sepa a ion o
he NAF Alge ian Mozabi e e e ence popula ion om he ME HGDP-
CEPH Is aeli A ab e e ence popula ions a K:6. This app oach was
adop ed a e o iginally e alua ing a dual K:5 un o ien a ed owa ds
wes Eu asia (AFR, EUR, NAF, ME, SAS) and eas Eu asia (EUR, NAF,
ME, SAS, EAS), al hough using hese wo sligh ly di e en K:5 uns did
no show any ad an age o e a K:6 Eu asian analysis. STRUCTURE was
un wi h 100,000 bu nin s eps and 100,000 MCMC s eps, using co e-
la ed allele equencies unde he Admix u e model. Clus e membe ship
p opo ion plo s we e cons uc ed wi h CLUMPAK .1.1 [35]. Op imum
‘K’gene ic clus e alues we e in e ed by calcula ing mean ΔK and L(K)
alues using s anda d p o ocols [36,37].
The abili y o ET X-SNPs o in e he ances y o an indi idual’s X-
ch omosome complemen was e alua ed using P incipal Componen
Analysis (PCA), by uploading e e ence da a om 1000 Genomes o
simple h ee-way compa isons based on AFR-EUR-AMR, using he
‘Classi ica ion o mul iple p o iles wi h a cus om Excel ile o pop-
ula ions’ op ion in he Snippe web po al [33]. The mul iple p o iles
classi ie p o ides a Bayes likelihood a io and PCA analysis, which is
now based on he h ee 2D plo s o p incipal componen (PC) 1 s PC2,
PC1 s PC3, and PC2 s PC3. X-ch omosome ances ies we e assigned
based on he posi ion o an admixed s udy sample in ela ion o hese
h ee e e ence popula ion PCA clus e s, and unassigned i his lay in he
egion o mino o e lap be ween clus e s, Posi ions we e judged o be
equidis an om wo adjacen clus e cen oids by isual inspec ion o
PCA cha da a. Such poin s occupying in e media e posi ions whe e
clus e o e lap can occu , we e in e p e ed o indica e ecombina ion o
con ibu o popula ion X-SNP alleles, and al e na i ely he p esence o
Table 2 (con inued)
P incipal
SNPs in
he
haplo ype
Ex a
SNPs
in
MPS
ou pu
In e nal
MH
name
O iginal MH
nomencla u e
P incipal
componen
SNPs
Ex a SNPs in
Ion S5 MPS
sequence
ou pu
5′
coo dina e:
GRCh37
3′
coo dina e:
GRCh37
5′
coo dina e:
GRCh38
3′
coo dina e:
GRCh38
MH span in
nucleo ides
O iginal
MH span
4 1 12qB s11177060-
s2111058-
s10878750-
s11835920
s571889826 68508276 68508353 68114496 68114573 77
4 - 15qD s1816771-
s74033914-
s5007156-
s4965040
- 98255928 98255978 97712698 97712748 50
4 2 MH18 mh16KK-
255
s16956011-
s3934954-
s3934955-
s3934956-
s576469239
/
s184092108
81970353 81970407 81936748 81936802 54 142
4 1 MH20 mh18KK-293 s621320-
s621340-
s678179-
s621766
s80093367 76089886 76089968 78329886 78329968 82 82
3 2 MH21 mh21KK-
315
s6517970-
s202132081-
s8131148-
s6517971
s533846035
/
s538072435
21880158 21880231 20507846 20507919 73 145
3 2 MH22 mh21KK-
324
s2838868-
s7279250-
s8133697
s537553521
/
s567533147
46714641 46714707 45294726 45294792 66 158
5 2 22qB s4925431-
s4925399-
s4925432-
s4925400-
s77899570
s192804904
/
s537823715
49060976 49061028 48665164 48665216 52
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
16
dis ibu ion o haplo ype equencies in he 21 MHs. Ve y a e haplo-
ypes obse ed in one o wo, o many popula ions, a e ma ked in ed o
yellow o aid isibili y. The unde lying haplo ype da a in all popula ions
including he VISAGE S udy popula ions is compiled in Supplemen a y
Table S2. We included bo h CEPH and Sange Middle Eas popula ion
da a in he ba plo s o c oss-compa ison pu poses, as he CEPH
whole-genome-sequencing da a was no phased so allelic combina ions
need o be in e ed om he common haplo ypes o each MH in o he
Eu asian popula ion da a. The e o e, some disc epancies occu , bu
hese a e mainly lowe equency haplo ypes, and he Sange ME
haplo ype equencies should be conside ed he mos eliable o e e -
ence pu poses.
Al hough any e iew o MH haplo ype equency dis ibu ions is
a he subjec i e, his is he i s analysis o Middle Eas a ia ion o
hese ypes o loci, so i is use ul o disce n he o e all pa e ns o
a ia ion in he da a. In line wi h p e ious obse a ions [11,27], AFR
haplo ype equencies show consis en ly highe le els o a ia ion
compa ed o o he popula ions. The opposi e cha ac e is ic applies o
OCE le els o polymo phism in hese loci, whe e ewe haplo ypes a e
obse ed and one o wo haplo ypes p edomina e in many o he ET
MHs, e.g., in 8pB he CATCA haplo ype alone accoun s o almos 85% o
he o al a ia ion in Oceanians. The AME haplo ype equencies also
ep esen lowe le els o polymo phism, bu only o ce ain MHs;
(no ably 1pD, 10pB, 15qD, and he TCT haplo ype in MH03, a >75%
equency). In gene al, nei he he CEPH no Sange ME haplo ype
dis ibu ions a e e y di e en ia ed om EUR popula ions, and wi h he
excep ion o 3qC, he same low le el o di e en ia ion applies o SAS s
EUR. The de elopmen o ances y p edic i e sys ems ha use single-si e
SNP geno ypes in combina ion wi h haplo ypes om a se ies o MHs has
no been comple ed in ei he he VISAGE in e p e a i e so wa e o he
esea ch labo a o ies suppo ing he de elopmen o he VISAGE ET
assay. The e o e, he MH da a gene a ed by ET has no been in eg a ed
in o de o enhance he au osomal BGA SNP da a o ming he co e o
mos o ensic ances y p edic ion sys ems. Despi e he complica ions o
using mixed ma ke da ase s o Bayes analysis, MH da a is sui able o
inclusion in SNP-based analysis uns wi h STRUCTURE. Howe e , we
did no o mally es he addi ion o he 21 MH loci o he 104 au osomal
SNPs using STRUCTURE, al hough expe ience indica es no imp o e-
men is seen in he di e en ia ion o he in e ed gene ic clus e s om
his algo i hm when ma ke da a is ex ended in his way.
3.3.2. Pilo s udies o e alua e ances y-based decon olu ion o simple
mix u es using Mic ohaplo ypes
Fig. 5 summa ises he indings o he pilo s udy o ances y-based
mixed DNA decon olu ion made on he 2-way mix u e cons uc ed
om EUR and AFR Co iell con ol DNAs. Fi s , he STRUCTURE analysis
o he s anda d e e ence popula ions o YRI, CEU and CHB indica es
well di e en ia ed gene ic clus e s o each popula ion, despi e being
based on da a om jus 21 MH loci (Fig. 5A). Once i was con i med ha
analysis o ep esen a i e popula ions om he h ee main con inen al
popula ion g oups was su icien ly in o ma i e, he clus e membe ship
p opo ions o he wo con ol DNAs we e ob ained and showed ha
STRUCTURE analysis o he haplo ypes in e ed om sequence ead
a ios would p o ide he means o in e he ances y o each haplo ype
de ec ed. Second, he olun a y sc u inee s asked wi h manually ec-
ognising he mix u e componen haplo ypes p oduced in e ences ha
we e compiled as 21 consensus MH p o iles o he 1:1, 3:1, 9:1 mix u e
a ios. When a sc u inee in e ed an MH p o ile wi h mo e haplo ypes
han he o he s, he consensus p o ile de aul ed o he leas numbe o
haplo ypes and was he e o e a conse a i e es ima e. I was no
Fig. 4. (con inued).
J. Ruiz-Ramí ez e al.

Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
17
possible o eliably de-con olu e he 1:1 mix u e, as sequence ead a ios
we e mos ly closely ma ched om bo h con ibu o s. As he mix u e
a io became mo e skewed, i was easie o dis inguish he majo and
mino con ibu o s as indica ed by Fig. 5B, whe e 81% o majo
con ibu o haplo ypes could be in e ed and 55% o mino con ibu o
haplo ypes in he 3:1 mix u e. These alues imp o ed o 100% and 64%
espec i ely, in he 9:1 a io; unde lining he ac ha mo e accen ua ed
mix u e a ios a e easie o decon olu e in his way. The e is a e y
sligh indica ion o incomple e p o ile econs uc ion in he STRUCTURE
plo s o he mino con ibu o o he 3:1 a io, wi h an inc ease in he
negligible EAS co-ances y p opo ion, bu he sequences de ec ed om
his mix u e componen would be unequi ocally in e ed o show AFR
ances y. Typical sequence co e age da a (numbe s o eads) a e shown
o example locus MH22 in Fig. 5 C, indica ing he di e ence in ead
a ios can e lec he mix u e a io o a la ge ex en , and i is clea ha
he CAA/CTA haplo ypes belong o he mino con ibu o in bo h 3:1
and 9:1 mix u e a ios. Howe e , hese sequence co e age eadings also
show ha balanced mix u es do no necessa ily lead o balanced eads
be ween he componen haplo ypes de ec ed. The ull se o plo s o
sequence co e age o he iden i ied haplo ypes in he 1:1, 1:3 and 1:9
mix u e a ios a e gi en in Supplemen a y Fig. S6.
In he 1:1 mix u e a io, Supplemen a y Fig. S6 shows almos all MHs
ha e mo e han 2 haplo ypes (12 wi h 3 haplo ypes, i e wi h 4 haplo-
ypes), and only MH18 and MH21 ha e he same le el o sequence
co e age o each o wo haplo ypes iden i ied in hese loci (MH locus
7pB has wo haplo ypes wi h a e y skewed co e age a io). These da a
emphasise he powe o Mic ohaplo ypes o de ec mix u es e en when
ull decon olu ion is no easible because indi idual haplo ype combi-
na ions canno be in e ed om he sequence co e age a ios. The e o e,
his pilo s udy using he 21 MH loci o ET sugges s an e icien sys em
o de ec ing mixed DNA can be applied independen ly o he single-si e
SNP da a and can ale he use o he p esence o a mix u e. When wo
con ibu o s a e p esen in he mixed DNA a unequal p opo ions he e
is a good oppo uni y o iden i y indi idual haplo ype pai s om each
con ibu o , and i hey ha e con as ing ances ies amongs A ican,
Eu opean, o Eas Asian popula ions-o -o igin, he e is he abili y o use
STRUCTURE o iden i y hese ances ies and assign hem o bo h
con ibu o s.
Beyond MH haplo ype analysis, bi-allelic SNPs ha e limi ed capa-
bili ies o mix u e decon olu ion om MPS sequence da a [45] and
gi en he powe o mul iple-haplo ype MH loci o de ec simple mixed
DNA componen s, we would ad oca e discoun ing single-si e SNP da a
when such mix u es a e de ec ed. In con as , i-allelic SNPs can de ec
simple mix u es mo e e icien ly om he de ec ion o h ee di e en
alleles in he sequence ead da a o each nucleo ide. Compa ed o he
use o MHs o decon olu e mix u es as desc ibed abo e, i-allelic SNPs
ha e much mo e limi ed powe , bu can add de ail o he obse a ions
based on he MH sequence da a.
Fig. 5. Pilo s udy o e alua e he abili y o he 21 Mic ohaplo ypes o ET o de ec mixed DNA and iden i y he ances y o con ibu o s in simple 2-way mix u es.
5 A: STRUCTURE analysis o AFR, EUR, EAS e e ence popula ions indica es he 21 MHs di e en ia e hese popula ions e icien ly and MH p o iles comp ising
haplo ypes decon olu ed om a mix u e can be included o in e hei likely ances y. 5B: Pe cen age o componen SNP alleles called o MH p o iles iden i ied by a
panel o sc u inee s o he MPS da a o 3:1 and 9:1 mix u e a ios (1:1 was no success ully decon olu ed and is no shown). STRUCTURE clus e plo s o he
decon olu ed MH p o iles shown igh . 5 C: Example sequence co e age ou pu o each iden i ied haplo ype in MH22. All h ee mix u e a ios allow ‘pai ing’ o wo
haplo ypes pe con ibu o , bu his was no possible o he 1:1 mix u e a io in o he MH loci o when less han ou haplo ypes a e p esen .
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
18
The pa allel s udy o his pape , desc ibing he de elopmen and
in e -labo a o y e alua ion o he VISAGE ET MPS assay [46], examined
he i-allelic da a om he same NA07000-NA18498 mix u e se ies. We
b ie ly summa ise hese indings below, which added e alua ion o
inc eased he e ozygosi y, and skews in he sequence co e age o he
de ec ed alleles o each i-allelic SNP, in addi ion o eco ding he
p esence o h ee alleles. The le el o i-allelic SNP he e ozygosi y
inc eased om ~25% o he con ibu o s o mo e han 56% in he 1:1
and 1:3 mix u e a ios, alling o lowe le els in he 1:9 mix u e a io.
The obse ed skews in he sequence co e age o he alleles o he
i-allelic SNPs we e close o expec a ions in he 1:3 and 1:9 a ios,
sugges ing i-allelic SNPs a e in o ma i e ma ke s o he analysis o
mixed DNA om MPS da a beyond he simple sequence co e age skews
occu ing wi h bi-allelic SNPs. Finally, h ee-allele pa e ns we e ound
in 11 o he 26 i-allelic SNPs o ET, which ma ches he expec ed
numbe om examina ion o he con ibu o SNP geno ypes.
3.4. Y-SNP geno ypes in VISAGE S udy popula ions
Al hough Y-SNP da a was no compiled om he main human
genome a ian da ase s because o incomple e da a, all Y-SNP geno-
ypes ob ained o he VISAGE S udy popula ion males ha e been
compiled and a e lis ed in Supplemen a y Table S3. The 5-SXC haplo-
ypes a e also lis ed alongside he haplog oup manually in e ed om
he Y-SNP alleles in each sample and he desc ip ion o he egion whe e
ha haplog oup is mos commonly obse ed.
Al hough a ho ough analysis o he dis ibu ion o Y a iabili y in
he Eas A ican, Middle Eas and Fijian S udy popula ions was no
made, we decided he B azilian samples would p o ide an in o ma i e
pilo s udy o he compa ison o X- and Y-SNP da a in wo popula ions
likely o ha e di e en admix u e his o ies. The B azilian u al sample is
o Kalungas who descended om escaped sla es and ha e li ed in
emo e se lemen s in Goi´
as S a e o abou 250 yea s. In con as , he
B azilian u ban sample is om B asília, he capi al o B azil. Compila-
ion o he X- and Y-SNP based ma ilineal and pa ilineal ances ies is
summa ised in Supplemen a y Fig. S7. This da a e ealed a no iceable
Eu opean male – A ican emale sex-biased admix u e a io in bo h
popula ion samples bu was much mo e ma ked in he u al Kalungas.
The u ban B azilian males had 6% A ican Y haplog oups and 45%
A ican-speci ic 5-SXC haplo ypes, while he u al B azilian males had
39% A ican Y haplog oups and 92% A ican-speci ic 5-SXC haplo ypes.
Eu opean Y haplog oups we e ound in 94% o u ban B azilian males
(19% we e designa ed as ‘ME-EUR’ wi h a haplog oup dis ibu ion
including Eas Eu opean, Caucasus and Middle Eas egions), and 18% o
Eu opean-speci ic 5-SXC haplo ypes, while he u al Kalungas had 67%
Eu opean Y haplog oups and 8% Eu opean-speci ic 5-SXC haplo ypes.
An in e es ing poin o compa ison wi h he X-Y da a is an independen
s udy o he same Kalunga samples using 46 au osomal ances y-
in o ma i e Indels, by Ca alho Gon ijo e al., in 2018 [47]. Ca alho
Gon ijo’s s udy de ec ed ~68% A ican and ~25% Eu opean
co-ances y p opo ions, wi h he o he 7% Ame ican (plus a ma ginal
Eas Asian p opo ion). These p opo ions a e b oadly posi ioned be-
ween he wo con ibu o popula ion a ios wi h a signi ican Eu opean
male – A ican emale sex bias, as we indica ed wi h he X- and Y-SNP
da a o he Kalungas.
The X ances y p o ile o he u ban B azilians showed equal 18%
p opo ions o EUR, EAS and AMR-speci ic 5-SXC haplo ypes. The e o e,
a disce nible Eu opean male sex bias exis s in bo h samples, bu is
pa icula ly s ong in he isola ed u al sample, whe e almos all he
obse ed ma ilineal X haplo ypes a e A ican-speci ic and wo hi ds o
he pa ilineal Y haplog oups ha e hei mos common dis ibu ion in
Eu ope. Al hough his simply ep esen s an ini ial explo a ion o da a
whe e X- and Y-SNP geno ypes can be compa ed, i sugges s a o ensic
ances y es ha combines each gonosomal ma ke se will ha e a de-
g ee o powe o analyse pa ilineal and ma ilineal pa e ns in pe sons
wi h admixed backg ounds. This is encou aging, gi en he ea ly decision
by VISAGE no o pu sue m DNA analysis as pa o he ET assay.
3.5. STRUCTURE analysis o ET BGA SNP da a
3.5.1. Wo ldwide popula ion s uc u e pa e ns in e ed om he au osomal
BGA SNPs o ET
We ely on STRUCTURE o analyse he ances y o unknown dono s
in o ensic DNA es s o se e al easons: i. we ha e ound i o be he
mos e ec i e way o de ec and analyse co-ances y pa e ns in in-
di iduals wi h admixed backg ounds; ii. al hough no pa o ou s udies
he e, STRUCTURE can combine and analyse a ian da a om di e en
ypes o genomic ma ke , so MH loci, STRs, SNPs and Indels can be
analysed oge he in a single un; iii. he a ailabili y o de ailed geno-
ype da a om whole-genome-sequence a ian da ase s allows a wide
ange o e e ence popula ions o be compiled o almos any ma ke se ,
and nea ly all SNPs iden i ied in he human genome o da e. Combining
unknown o ensic sample da a ma ked as POPFLAG=0, wi h size-
adjus ed e e ence da a (app oxima ely 100 samples pe e e ence
popula ion) ma ked as POPFLAG=1, p o ides an e ec i e way o
examine he likely ances y o he unknown samples. A common p ob-
lem wi h he use o STRUCTURE is o e i ing he da a o a numbe o
in e ed gene ic clus e s (K) g ea e han he ac ual clus e s ha can be
p ope ly disce ned wi h he ma ke s used. Since o ensic BGA ma ke
se s a e limi ed in numbe o p ese e assay sensi i i y, he ini ial
analysis o samples o unknown ances y wi h STRUCTURE equi es a
cau ious explo a ion o each K alue, gene ally om K:2 o K:8. Ou
expe ience has indica ed ha da a o e i ing - when oo many clus e s
a e in e ed and indi idual popula ion g oups begin o show i egula
wi hin-popula ion clus e membe ship p opo ions - can occu a e K:5,
analysing a con inen ally-based e e ence popula ion se o AFR, EUR,
EAS, AMR, OCE ha includes he less well di e en ia ed popula ion
g oups o ME and SAS. To coun e hese e ec s and o p o ide he op-
imum di e en ia ion o gene ic clus e s, we ha e adop ed a ‘nes ed’
app oach o STRUCTURE analyses ha uns a i e-con inen e e ence
se wi h he unknown sample(s) se a K:5 expec ed clus e s. Depending
on he clus e membe ship pa e ns ound in he POPFLAG=0 samples,
ano he K:5 un analyses he samples wi h a Eu asian sub-con inen al
e e ence se o AFR, EUR, ME, SAS, EAS. We ha e ound his im-
p o es he clus e pa e ns de ec ed in admixed samples, which a e
p edominan ly om he Ame icas and he e o e show co-ances y p o-
po ions in a ying deg ees om AFR, EUR and/o AMR con ibu ing
popula ions. One p oblem can be he de ec ion o SAS co-ances y in he
second Eu asian-cen ed STRUCTURE un, and in such cases he ini ial
un’s e e ence popula ion da a can be adjus ed o K:5 expec ed clus e s
by swapping ou he OCE popula ions. One example o when he
explo a o y STRUCTURE uns can equi e adjus men depending on he
esul s o bo h analyses, is he 1KG sample HG01880 shown in Fig. 3,
wi h ~30% SAS co-ances y de ec ed om 1000 Genomes’ own gene ic
s uc u e analyses [18]. Because we do no include SAS e e ence da a in
he i s STRUCTURE un his would go unde ec ed un il he Eu asian
sub-con inen al e e ence da a un was comple ed, and a new un made
wi h OCE e e ence geno ypes swapped ou o SAS.
Applying he Con inen al K:5 - Eu asian Sub-Con inen al K:5 nes ed
app oach desc ibed abo e o he ull ange o 1KG, CEPH, Sange ME
and VISAGE S udy popula ions p oduced a gene ally obus iden i ica-
ion o he majo i y clus e membe ship p opo ions in each sample. The
mino i y clus e membe ship pa e ns in almos all samples p oduced a
cohe en pa e n which ma ched he geog aphic loca ion o he pop-
ula ions analysed, pa icula ly hose om he Middle Eas egions. When
pe o ming hese STRUCTURE analyses, we consis en ly obse ed well
di e en ia ed gene ic clus e pa e ns a K:6 in he Eu asian Sub-
Con inen al uns when he CEPH Mozabi e Alge ian samples we e
included as a six h popula ion e e ence se ma ked as ‘No h A ican’
(NAF) POPFLAG=1. Fo his eason, we show he K:6 pa e ns gene -
a ed using six e e ence popula ions which includes dis inc NAF and ME
e e ence da ase s (ME comp ising he h ee Is aeli A ab popula ions o
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
19
Bedouin, Pales inian and D uze). Fig. 6 displays, in app oxima e
geog aphic loca ions-o -sampling, he STRUCTURE clus e plo seg-
men s o he popula ions om each da ase ha show de ec able and
a ying deg ees o co-ances y. The clus e plo s a e gene ally a anged
in descending o de o majo co-ances y componen s ha e been
expanded wo- old o show indi idual clus e pa e ns mo e clea ly. The
HGDP-CEPH Sa dinian, Tuscan, Adygei EUR popula ions and he Pak-
is ani SAS popula ions showed some co-ances y pa e ns bu a e
excluded o cla i y. All o he popula ions no shown in Fig. 6 had single
clus e membe ship pa e ns ma ching hose epo ed in nume ous
s udies o he same samples using se e al o ensic BGA SNP se s [1,4,
13–18]. The e o e, we concen a ed on esul s om admixed 1KG
samples and he h ee VISAGE S udy popula ions ou side o Eu asia
analysed wi h he ini ial K:5 STRUCTURE un; and he nine Sange ME
plus ou VISAGE S udy popula ions om No h A ican, Eas A ican o
Middle Eas egions, analysed wi h he Eu asian K:6 STRUCTURE un,
which included a six h NAF e e ence da ase . These uns analysed he
104 au osomal BGA SNPs in ET, comp ising bi-allelic and i-allelic loci.
The a e age clus e membe ship p opo ions o he ini ial K:5
STRUCTURE un and he Eu asian K:6 STRUCTURE un o all 1KG,
CEPH, Sange ME and S udy popula ion samples included in each
analysis, plus he co esponding segmen ed clus e plo s om his da a,
a e lis ed in ull in Supplemen a y Tables S4A and S4B o Con inen al
and Eu asian da ase s, espec i ely.
Re iewing he Fi e-Con inen al K:5 e e ence and s udy popula ion
clus e plo s i s . The i e admixed 1KG popula ion clus e pa e ns
shown op le in Fig. 6 plus he 67/85 admixed PEL, a e discussed in he
nex sec ion. In he e e ence popula ion clus e plo he inabili y o
ma ch he numbe s o Oceanian e e ence samples o hose o he o he
popula ions is e iden , so a deg ee o bias may ha e occu ed in iden-
i ying and quan i ying he OCE clus e membe ship p opo ions when
de ec ed as co-ances y componen s in S udy Fijians and CEPH Cam-
bodians. Ne e heless, he la ge-scale educ ion o OCE-in o ma i e
SNPs om 23 in BT o 3 in ET has no a ec ed he abili y o he ET
BGA SNPs o di e en ia e his popula ion g oup. In ac , he i s OCE
sample se o Papua New Guinea is dis inguishable om he second o
Bougain illean samples, wi h he de ec able p esence o EAS co-ances y
in he la e . One o he clus e pa e n o highligh is he 5 h AMR
sample se comp ising CEPH Maya, which shows EUR co-ances y a a
highe le el han he o he CEPH AMR sample se s (se 2 =Ka i iana;
3=Su ui; 4 =Colombians; 5 =Maya; 6 =Pima). This ep esen s a close
ma ch o pa e ns ob ained om he wo landma k s udies o he HGDP-
Fig. 6. Clus e plo s o STRUCTURE analyses o selec ed 1KG, CEPH, Sange ME and S udy popula ions. Nes ed STRUCTURE analyses consis ed o i s s age K:5 uns
using Fi e-Con inen al e e ence popula ion da ase s (POPFLAG=1) comp ising 1KG AFR (YRI); EUR (CEU); EAS (CHB); 2 CEPH OCE popula ions; 5 CEPH AMR
popula ions plus a subse o 1KG PEL wi h no non-AMR co-ances y. Popula ions s udied (POPFLAG=0) a e shown le and igh o cen al g oup o popula ions,
comp ising six 1KG admixed A ican and Ame ican popula ions; 67 PEL wi h de ec ed non-AMR co-ances y; S udy B azilian u al and u ban popula ions; wo CEPH
Eas Asian popula ions wi h co-ances y om o he popula ions; S udy Fijians. The cen al g oup o Middle Eas egion popula ions was analysed wi h he second
s age K:6 uns using Eu asian Sub-Con inen al e e ence popula ion da ase s, comp ising 1KG YRI; CEPH Alge ian Mozabi e; 3 CEPH Is aeli A ab popula ions; 1KG
CEU, 1KG SAS (GIH); 1KG CHB. Popula ions es ed we e i e VISAGE S udy popula ions and nine Sange ME popula ions (Emi a i A-D and Saudi A-B a e a anged
sepa a ely bu no loca ed o a speci ic egion). The h ee samples in ASW and ACB wi h highes le els o non-AFR co-ances y shown on he igh as
expanded columns.
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
20
CEPH di e si y panel, using la ge ma ke se s (See Fig. 1 o [48], and
Fig. 1 o [25]). The S udy Fijian plo indica es mos samples would be
iden i ied as ha ing OCE o igin bu no e i e o he igh mos columns
a e sel -decla ed Indo-Fijians likely o ha e SAS co-ances y, which
would be unde ec ed wi h his e e ence popula ion da a absen om he
Con inen al STRUCTURE un, bu p esen in he Eu asian STRUCTURE
un. This exempli ies he need o adjus e e ence da a acco ding o bo h
STRUCTURE analyses (Fijian clus e plo s om Eu asian STRUCTURE
analysis uns no shown). Las ly, he wo S udy B azilian sample clus e
plo s illus a e he con as in admix u e pa e ns be ween hem. The
u al B azilian sample has p edominan AFR co-ances y (apa om he
igh mos wo indi iduals), con as ing wi h u ban B azilians, who show
p edominan EUR co-ances y, apa om he wo igh mos indi iduals.
The wo B azilian samples in e ed o ha e AMR X ch omosomes
(Fig. 3B) showed 3% ( u al, K113) and 10% (u ban, BSB228) AMR
co-ances y in his analysis.
The Eu asian sub-Con inen al K:6 e e ence and s udy clus e plo s
illus a e he success ul di e en ia ion o NAF and ME popula ions,
al hough his was based on he single CEPH Alge ian e e ence popu-
la ion, which could lead o biased analysis due o possible s a i ica ion
o SNP a ia ion in a popula ion no necessa ily ep esen a i e o a i-
abili y ac oss a wide egion. The e o e, he clus e pa e ns de ec ed in
he S udy Mo occan sample a e pa icula ly ele an . S udy Mo occans
show a b oad ange o NAF co-ances y p opo ions om 5–95% in wo-
hi ds o samples, wi h he majo i y o Mo occans showing sligh ly
highe p opo ions o ME co-ances y han NAF, apa om wo samples
wi h AFR-EUR co-ances y, and wo wi h ME-EUR co-ances y. AFR co-
ances y is de ec able in 9 o he 27 Alge ian e e ence samples, wi h
majo i y AFR co-ances y p opo ions in h ee. The o he wel e Middle
Eas egion popula ions p o ide clus e pa e ns well ma ched o hei
loca ions. I is no possible o iden i y he Emi a i A-D popula ions, bu
hese appea o show a p og ession in SAS co-ances y p opo ions in a
leas hal o samples om C and D. The o he Sange ME popula ions
show p edominan ME clus e membe ship p opo ions in almos all
samples, so would be dis inguishable om a Eu opean indi idual apa
om ( igh mos ) Tu kish and Sy ian samples, which e ain a de ec able
ME co-ances y. Conside ing he Middle Eas popula ion sample se as a
whole, a consis en geog aphic pa e n is e iden o he majo i y o
samples in each popula ion. This comp ises i. a s ong p esence o he
ed NAF gene ic clus e in hal o samples om he No hwes co ne o
his egion, which is sha ed wi h he g ey ME clus e ; ii. a de ec able co-
ances y p esence o he blue EUR clus e in abou a hi d o samples in
he No h o No heas co ne , wi h a p edominan ME co-ances y in
hese samples om Tu key, Sy ia, and I aq (plus mino SAS co-ances y
in mos samples); iii. wo Eas A ican sample se s wi h equal p o-
po ions o AFR and NAF-ME clus e membe ships, in pa e ns which a e
gene ally dis inc om he o he ME popula ions; i . a p edominan ME
clus e , mos ly >90%, in a majo i y o samples om popula ions a ound
he Saudi A abian Peninsula, comp ising nea ly all Yemeni, Saudi A and
B, Emi a i A and B, and hal o he I aqis and Sy ians. The e o e, using a
second STRUCTURE analysis wi h six e e ence popula ions, i is
possible o iden i y ME co-ances y in he majo i y o ‘unknown’ es
samples in his s udy, wi h a NAF co-ances y signal de ec ed in hal o
Mo occans. As a ule o humb, he p esence o AFR and ME, and/o NAF
join clus e membe ships sugges s a pa e n cha ac e is ic o Eas A -
ican ances y. The p esence o 15%−25% ME co-ances y membe ship
p opo ions in 4/99 CEU e e ence samples, sugges s a conse a i e
app oach would be o in e Middle Eas ances y using a h eshold o
20–25% o highe ME and/o NAF co-ances y p opo ions. No e ha
his would iden i y mos o he S udy Tu kish samples as ha ing dis inc
pa e ns compa ed o Eu opeans. E en applying a s ingen h eshold o
25% minimum ME/NAF membe ship p opo ions o signi y Middle Eas
ances y, a es o non-in e ence a e low amongs hese es popula ions.
The wo Eas A ican popula ions would ha e 2% non-in e ence; Emi a i
12%; Mo occans 3%; I aqis 10%; Tu kish 18%, wi h secu e ME in-
e ences possible o all Sy ian, Saudi A abian and Yemeni samples.
3.5.2. Analysing co-ances y in admixed popula ion samples wi h
STRUCTURE
In a c iminal in es iga ion, a o ensic ances y es ha can eliably
iden i y co-ances y in a pe son wi h an admixed backg ound would, in
such cases, p o ide impo an in o ma ion abou he likely appea ance
o a suspec . When p e iously e alua ing he abili y o he VISAGE BT
ances y panel o de ec admix u e and es ima e he co-ances y p o-
po ions in such a sample, we made a o mal compa ison be ween he
clus e membe ship pa e ns om analysing he same 504 1KG admixed
samples wi h 572,000 Human O igins a ay SNPs s he 115 BGA SNPs
o BT. Wi h he BGA SNPs o ET we did a simila compa ison o he same
samples bu used he co-ances y p opo ions es ima ed om genome-
wide SNP da a published by 1000 Genomes [18]. Supplemen a y Figs
S8A-S8D shows he clus e plo s om bo h analyses wi h he sample
o de dic a ed by he 1KG da a a anged by descending majo i y
co-ances y membe ship p opo ions in each popula ion. These plo s
show he comple e 1KG sample se in Supplemen a y Fig. S8A, ollowed
by expanded plo s o admixed A icans ACB, ASW in Supplemen a y
Fig. S8B, and admixed Ame icans CLM, PEL, PUR, MXL in Supplemen-
a y Fig. S8C. Supplemen a y Fig. S8D shows he co ela ion analyses
and
2
alues used o gauge he le els o co ela ion be ween he
co-ances y p opo ion es ima es made wi h each SNP se , combining
AFR and AMR co-ances y p opo ions in o a single alue and compa ing
EUR co-ances y p opo ion es ima es di ec ly.
Se e al ac o s a e e iden om a e iew o he co ela ion alues
and STRUCTURE clus e plo s p oduced by ET BGA SNP analyses. Fi s ,
he e is a good ma ch be ween bo h SNP se s in he es ima es o majo i y
co-ances y ac oss all samples and popula ions, pa icula ly when his is
abo e 90%. Consequen ly,
2
alues a e highes o compa isons o AFR
co-ances y p opo ion es ima es in ACB and ASW, and hose o EUR in
CLM and MXL. Closely ma ched clus e plo pa e ns and co ela ion
alues a e also seen in PEL, al hough combining AFR-AMR clus e p o-
po ion es ima es o simpli y analysis educes hese co ela ion alues.
Las ly, he h ee co-ances y ou lie s ( igh mos columns) in ACB and
ASW, which a e also highligh ed in Fig. 4 and 8, p esen good clus e
plo ma ches, wi h he SAS co-ances y p opo ion ecognised by he ET
BGA SNPs when his popula ion e e ence da ase is included as
POPFLAG=1 geno ypes. The h ee ASW ou lie s indica e an o e -
es ima ion o AMR clus e p opo ions wi h ET BGA SNPs, and he same
ma ginal bu consis en e ec is seen in PEL and MXL clus e plo pa -
e ns. The wo s co ela ion and clus e plo ma ches a e obse ed in he
PUR compa isons. This appea s o s em om a highe le el o h ee-way
admix u e in his popula ion, al hough CLM ha e simila admix u e
pa e ns, bu p oduce much be e co ela ion alues o
2
=0.758 o
he combined AFR/AMR co-ances y p opo ion es ima es, compa ed o
2
=0.334 in PUR. Much o he AMR co-ances y es ima ion in PUR is
e oded by many samples wi h EAS and SAS co-ances y p opo ions, and
i migh be bene icial o conside a K:3 STRUCTURE analysis wi h AFR,
EUR and AMR e e ence da ase s, when h ee componen s o admix u e
a e iden i ied in unknown samples and wo a e ei he EUR and AMR, o
EUR and AFR. A e iew o he clus e plo o PUR in Fig. 6 indica es his
popula ion is gene ally p oblema ic o analyse o co-ances y and a
signi ican p opo ion o samples ha e low-le el EAS and OCE co-
ances y p opo ions, when he 1KG da a sugges s hese should be ec-
ognised as AMR co-ances y. I is no ewo hy ha simila s udies o he
BGA SNPs in BT ga e he lowes
2
alue o he PUR combined AFR/
AMR co-ances y p opo ion es ima es o 0.446.
O e all, gene ic clus e di e en ia ions become less eliable in in-
di iduals wi h h ee di e en co-ances y componen s, so STRUCTURE
analyses o popula ions such as B azil mus be app oached wi h cau ion.
Th ee-way admix u e con inues o p esen a conside able challenge o
STRUCTURE-based analysis o co-ances y pa e ns when using ances y
es s on a much smalle scale han hose used o popula ion gene ics
s udies. The e o e, a p uden measu e is o explo e a se ies o K:3 uns
wi h di e en combina ions o e e ence popula ion da ase s. Al hough
we do no p esen u he au osomal SNP analysis da a o Fijians, his
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
21
popula ion would be op imally analysed wi h EUR, SAS, EAS and OCE
e e ence da a in a ious combina ions.
3.5.3. Compa isons o STRUCTURE analyses using 104 BGA SNPs s
combined 104 BGA plus 184 au osomal EVC SNPs
As h ee EVC-SNPs ha e been sha ed o pigmen a ion ai p edic-
ion and ances y analysis pu poses in bo h VISAGE SNP geno yping
assays, i was conside ed wo hwhile o o mally e alua e he e ec o
combining all 104 au osomal BGA SNPs in ET wi h he 184 au osomal
EVC-SNPs. The same Con inen al K:5 - Eu asian Sub-Con inen al K:6
nes ed analysis was made o i e and six popula ion e e ence da ase s,
espec i ely, as desc ibed in Sec ion 3.5.1. Supplemen a y Fig. S9 shows
he K:5 and K:6 clus e plo s o bo h SNP se s, wi h accompanying
E anno cha s o Del aK and L(K) [36]. The o e all quali y o gene ic
clus e s is no iceably educed in he expanded 288 au osomal SNP
da ase compa ed o he dedica ed BGA SNP da ase , pa icula ly o he
less di e gen popula ions o NAF, ME and SAS, whe e a signi ican
numbe o mixed gene ic clus e pa e ns a e obse ed amongs hese
h ee popula ions, using all 288 SNPs.
4. Discussion
The s udies desc ibed he e ha e la gely concen a ed on he added
bene i b ough by including Y-SNP, X-SNP, and MH ma ke s ha all
ha e s ong popula ion di e en ia ion p ope ies in he ET ances y
panel. While i is impo an o acknowledge ha many o he au osomal
SNPs o iginally pa o he VISAGE BT ances y panel we e eplaced wi h
new ma ke s o ET, mos o hese new BGA SNPs a e al eady well
es ablished o o ensic use. I was only necessa y o adjus he balance
o ma ke s owa ds EUR, EAS and AMR di e en ia ions, and educe
hose o AFR and OCE. Expanding he se o ME-in o ma i e SNPs in ET
has p o ided conside able bene i s in e ms o he success ul iden i i-
ca ion o ME and NAF gene ic clus e s when STRUCTURE is un a K:6
wi h Eu asian-o ien a ed e e ence popula ions (g ey and ed gene ic
clus e s, espec i ely, in Fig. 6). Ou analyses show in almos all cen al
Middle Eas popula ion samples om he Sange ME a ian da ase s
and VISAGE S udy samples om hese egions, he e a e majo i y
membe ship p opo ions om one o bo h ME and NAF gene ic clus e s.
Whe e samples a e om egions on he pe iphe y o he cen al Middle
Eas a ea, he o he gene ic clus e s he STRUCTURE analyses iden i ied
co espond well o hei geog aphic posi ion in his b oadly-based e-
gion. Speci ically, many Tu kish show EUR co-ances y; Eas A icans
ha e p edominan AFR co-ances ies; and he Sange Emi a i in he Eas
(al hough we canno place popula ions A-D in speci ic geog aphic po-
si ions) show SAS co-ances y in many indi iduals. Al hough i is no
app op ia e o iable o use STRUCTURE o assign a sample o a speci ic
popula ion, he pa e ns we ha e gene a ed wi h ‘nes ed’ Eu asian
e e ence popula ion STRUCTURE uns allow a sample wi h ‘g ey’ and/
o ‘ ed’ clus e s p opo ions abo e 10% o be iden i ied as coming om
he Middle Eas , and exclude an o igin om sub-Saha an A ica, Eu ope,
Sou h Asia, o Eas Asia. In many cases, indi iduals show a cha ac e is ic
signa u e o No h A ican o Eas A ican popula ion o igins as dis inc
om he cen al Middle Eas e n egions shown in Fig. 6. The e o e, we
conside he goal se by VISAGE o de eloping an ances y panel ha can
e icien ly di e en ia e Middle Eas popula ion o igins om he neigh-
bou ing popula ion g oups, was la gely me and did no equi e a e y
la ge expansion o BGA SNP numbe s in ET o accomplish his goal. The
adap a ion o STRUCTURE uns in o a nes ed app oach which analyses a
educed se o e e ence popula ions wi h a na ow ange o possible K
alues, has helped o ocus ances y analyses on he mos app op ia e
egions and as ou analyses show, enables mo e de ailed gene ic clus e
di e en ia ions o be made o he Middle Eas .
The o he expansion made o he ET ances y panel - ha o
b oadening he ypes o ances y in o ma i e ma ke s o include X-SNPs,
Y-SNPs and MH loci ha e mo e specialised applica ion in ances y an-
alyses used o o ensic casewo k. Wi h he dis inc ions ha can be
eliably made be ween AFR, EUR and AMR co-ances ies wi h he
au osomal BGA SNPs o ET, admixed Ame ican indi iduals can be
de ec ed and hen analysed e icien ly. Consequen ly, mo e de ail is
ob ained o male samples by adding he analysis o pa e ns o a ia ion
obse ed in X and Y ch omosome ma ke s. The le el o de ail we we e
able o achie e in he analysis o B azilian samples, which a e o en oo
complex in hei co-ances y pa e ns o be easily s udied wi h small-
scale ma ke se s, highligh s he powe o combining ma ke se s wi h
sligh ly con as ed gene ic his o ies. Such his o ies o en ollow admix-
u e e en s om up o h ee di e en con ibu ing popula ions, and wi h
he complica ing e ec o a ied sex bias in di e en pa s o he same
geog aphic egion. Ne e heless, we highligh he p oblems we
encoun e ed in eliably di e en ia ing h ee-way co-ances y clus e
pa e ns in Pue o Ricans (PUR) and ob aining compa able da a o hose
o he genome-wide SNP da a om 1000 Genomes. The e o e, i is
necessa y o emain cau ious when h ee di e en co-ances ies a e
de ec ed in an indi idual, as small-scale au osomal BGA SNP panels may
no eliably measu e hei ela i e p opo ions compa ed o genome-
wide da a.
A key cha ac e is ic a ou ing he use o Mic ohaplo ypes in o ensic
DNA analysis has been hei abili y o analyse mixed DNA wi hou he
hind ance o non-allelic PCR s u e p oduc s complica ing he pa e ns
seen [49]. P e iously, we de eloped an app oach o analysing mixed
DNA wi h MHs which speci ically exploi ed MH loci wi h s ongly
con as ing haplo ype equencies in di e en popula ion g oups.
Despi e comp ising a simple pilo s udy limi ed o a single mixed DNA a
a ew a ios, we ha e demons a ed he 21 MHs chosen o ET success-
ully assign ances ies o he componen s o 2-way mixed DNA, no ably
when he e is imbalance in hei a ios, making sequence compa isons
easie o achie e. This app oach is helped by he ease wi h which MH
loci can be analysed wi h STRUCTURE and he di e en ia ion hey
p o ide o Eu ope, A ica, and Eas Asia. Al hough such analyses a e no
amenable o a high- h oughpu MPS pipeline, since haplo ypes mus be
econs uc ed locus-by-locus and hei sequence a ios es ima ed, he
abili y o de ec he likely ances y o con ibu o s could po en ially
p o ide key ex a in o ma ion o in es iga o s.
The adap a ion o he BT ances y panel comp ising mainly es ab-
lished au osomal o ensic BGA SNPs, in o he much mo e b oadly based
se o BGA ma ke s in ET ep esen s a conside able enhancemen o he
scope and powe o o ensic ances y analysis using MPS, as he chosen
name o he VISAGE Enhanced Tool implies.
Acknowledgmen s
The s udy was suppo ed by he Eu opean Union’s Ho izon 2020
Resea ch and Inno a ion P og amme unde g an ag eemen No.
740580 wi hin he amewo k o he VISible A ibu es h ough GEno-
mics (VISAGE) P ojec and Conso ium. M.d.l.P. is suppo ed by a pos -
doc o a e g an unded by he Conselle ía de Cul u a, Educaci´
on e
O denaci´
on Uni e si a ia e da Conselle ía de Economía, Emp ego e
Indus ia om Xun a de Galicia, Spain (ED481D-2021–008). J.R. is
suppo ed by he “P og ama de axudas ´
a e apa p edou o al” unded by
he Conselle ía de Cul u a, Educaci´
on e O denaci´
on Uni e si a ia e da
Conselle ía de Economía, Emp ego e Indus ia om Xun a de Galicia,
Spain (ED481A-2020/039). C.P., A.F.A., A.M.M., M.d.l.P., M.V.L. and
he wo k o compile ances y in o ma i e i-allelic SNPs and mic o-
haplo ypes a e suppo ed by MAPA, ‘Mul iple Allele Polymo phism
Analysis’ (BIO2016–78525-R), a esea ch p ojec unded by he Spanish
Resea ch S a e Agency (AEI) and co- inanced wi h ERDF unds. The
popula ion s udies by S.O. a Uni e si y o San iago de Compos ela, we e
inanced by he Fundaç˜
ao de Apoio a Pesquisa do Dis i o Fede al
(FAPDF), B azil.
The au ho s g a e ully acknowledge he sha ing o gene ic clus e
analysis in o ma ion om he 1000 Genomes Phase III SNP da a, kindly
p o ided by Adam Au on, Depa men o Gene ics, Albe Eins ein
College o Medicine, B onx, NYC, USA. The au ho s hank Luciana Maia
J. Ruiz-Ramí ez e al.

Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
22
Esche dos San os and Sab ina Guima ˜
aes Pai a o hei dedica ed wo k
in he collec ion o samples om u al and u ban B azil used in his
s udy. All STRUCTURE analyses we e pe o med by he FinisTe ae II
supe compu e a he Cen o de Supe compu aci´
on de Galicia, San iago
de Compos ela (CESGA), Spain.
Appendix A
Cen es and in es iga o s o he VISible A ibu es h ough GEnomics
(VISAGE) Conso ium, Websi e: h p://www. isage-h2020.eu/
(accessed 1s Feb ua y 2023).
•E asmus MC Uni e si y Medical Cen e Ro e dam, Ro e dam, he
Ne he lands: Man ed Kayse , Vi ian Kalama a, A win Ral , A hina
Vidaki.
•Jagiellonian Uni e si y, K akow, Poland: Wojciech B anicki,
Ewelina Po´
spiech, Aleksand a Pisa ek.
•Uni e sidade de San iago de Compos ela, San iago de Compos ela,
Spain: ´
Angel Ca acedo, Ma ia Vic o ia La eu, Ch is ophe Phillips, Ana
F ei e-A adas, Ana Mosque a-Miguel, Ma ía de la Puen e.
•Medizinische Uni e si ¨
a Innsb uck, Innsb uck, Aus ia: Wal he
Pa son, Ca a ina Xa ie , An onia Heidegge , Ha ald Niede s ¨
a e .
•Uni e si ¨
a zu K¨
oln, Cologne, Ge many: Michael No hnagel, Ma ia-
Alexand a Ka sa a, Ta ek Khella .
•King’s College London, London, UK: Ba ba a P ainsack, Gab ielle
Samuel.
•Klinikum de Uni e si ¨
a zu K¨
oln, Cologne, Ge many: Pe e M.
Schneide , The esa E. G oss, Jan Fleckhaus, Elaine Cheung.
•Bundesk iminalam , Wiesbaden, Ge many: Ingo Bas isch, Na halie
Schu y, Jens Teodo idis, Ma ina Un e l¨
ande .
•Ins i u Na ional de Police Scien i ique, Lyon, F ance: F ançois-
Xa ie Lau en , Ca oline Bouakaze, Yann Chan el, Anna Deles ,
Cl´
emence Holla d, Ayhan Ulus, Julien Vannie .
•Ne he lands Fo ensic Ins i u e, The Hague, he Ne he lands: Ti ia
Sijen, K is an de Gaag, Ma ina Ven ayol-Ga cia.
•Na ional Fo ensic Cen e, Swedish Police Au ho i y, Link¨
oping,
Sweden: Johannes Hedman, Kla a Junke , Maja Sids ed .
•Me opoli an Police Se ice, London, Uni ed Kingdom: Shazia
Khan, Ca ole E. Ames, And ew Re oi .
•Cen alne Labo a o ium K yminalis yczne Policji, Wa saw, Poland:
Magdalena Sp´
olnicka, Ewa Ka asinska, Anna Wo´
zniak.
Appendix B. Suppo ing in o ma ion
Supplemen a y da a associa ed wi h his a icle can be ound in he
online e sion a doi:10.1016/j. sigen.2023.102853.
Re e ences
[1] C. Phillips, Fo ensic gene ic analysis o bio-geog aphical ances y, Fo ensic Sci. In .
Gene . 18 (2015) 49–65.
[2] M. Kayse , Fo ensic DNA Pheno yping: P edic ing human appea ance om c ime
scene ma e ial o in es iga i e pu poses, Fo ensic Sci. In . Gene . 18 (2015)
33–48.
[3] A. F ei e-A adas, C. Phillips, M.V. La eu, Fo ensic indi idual age es ima ion wi h
DNA: om ini ial app oaches o me hyla ion es s, Fo ensic Sci. Re . 29 (2017)
121–144.
[4] M. de la Puen e, J. Ruiz-Ramí ez, A. Amb oa-Conde, C. Xa ie , J. Pa do-Seco,
J. ´
Al a ez-Dios, A. F ei e-A adas, A. Mosque a-Miguel, T.E. G oss, E.Y.Y. Cheung,
e al., De elopmen and e alua ion o he ances y in o ma i e ma ke panel o he
VISAGE basic ool, Genes 12 (2021) 1284.
[5] C. Xa ie , M. de la Puen e, A. Mosque a-Miguel, A. F ei e-A adas, V. Kalama a,
A. Vidaki, T.E. G oss, A. Re oi , E. Po´
spiech, E. Ka asin´
ska, e al., De elopmen
and alida ion o he VISAGE AmpliSeq basic ool o p edic appea ance and
ances y om DNA, Fo ensic Sci. In . Gene . 48 (2020), 102336.
[6] L. Palencia-Mad id, C. Xa ie , M. de la Puen e, C. Hoho , C. Phillips, M. Kayse ,
W. Pa son, VISAGE conso ium, e alua ion o he VISAGE basic ool o appea ance
and ances y p edic ion using Powe Seq chemis y on he MiSeq FGx sys em, Genes
11 (2020) 708.
[7] A. Heidegge , C. Xa ie , H. Niede s ¨
a e , M. de la Puen e, E. Po´
spiech, A. Pisa ek,
M. Kayse , W. B anicki, W. Pa son, VISAGE conso ium, de elopmen and
op imiza ion o he VISAGE basic p o o ype ool o o ensic age es ima ion,
Fo ensic Sci. In . Gene . 48 (2020), 102322.
[8] A. Wo´
zniak, A. Heidegge , D. Piniewska-R´
og, E. Po´
spiech, C. Xa ie , A. Pisa ek,
E. Ka asi´
nska, M. Bo o´
n, A. F ei e-A adas, M. Woj as, e al., De elopmen o he
VISAGE enhanced ool and s a is ical models o epigene ic age es ima ion in
blood, buccal cells and bones, Aging 13 (2021) 6459–6484.
[9] A. Pisa ek, E. Po´
spiech, A. Heidegge , C. Xa ie , A. Papie˙
z, D. Piniewska-R´
og,
V. Kalama a, R. Po aba ula, M. Bochenek, M. Siko a-Polaczek, e al., Epigene ic
age p edic ion in semen - ma ke selec ion and model de elopmen , Aging 13
(2021) 19145–19164.
[10] A. Heidegge , A. Pisa ek, M. de la Puen e, H. Niede s ¨
a e , E. Po´
spiech,
A. Wo´
zniak, N. Schu y, M. Un e l¨
ande , M. Sids ed , K. Junke , e al., De elopmen
and in e -labo a o y alida ion o he VISAGE enhanced ool o age es ima ion
om semen using quan i a i e DNA me hyla ion analysis, Fo ensic Sci. In . Gene .
56 (2020), 102596.
[11] M. de la Puen e, M.J. Ruiz-Ramí ez, A. Amb oa-Conde, C. Xa ie , J. Amigo, M.
A. Casa es de Cal, A. G´
omez-Ta o, A. Ca acedo, W. Pa son, C. Phillips, M.V. La eu,
B oadening he applicabili y o a cus om mul i-pla o m panel o Mic ohaplo ypes:
Bio-geog aphical ances y in e ence and expanded e e ence da a, F on . Gene . 11
(2020), 581041.
[12] V. Pe ei a, A. F ei e-A adas, D. Balla d, C. Bø s ing, V. Diez, P. P uszkowska-
P zybylska, J. Ribei o, N.M. Achakzai, A. Ali e i, O. Bulbul, e al., De elopmen
and alida ion o he EUROFORGEN NAME (No h A ican and Middle Eas e n)
ances y panel, Fo ensic Sci. In . Gene . 42 (2019) 260–267.
[13] C. Phillips, W. Pa son, B. Lundsbe g, C. San os, A. F ei e-A adas, M. To es,
M. Edua do , C. Bø s ing, P. Johansen, M. Fonde ila, e al., Building a o ensic
ances y panel om he g ound up: he EUROFORGEN Global AIM-SNP se ,
Fo ensic Sci. In . Gene . 11 (2014) 13–25.
[14] J.M. Galan e , J.C. Fe nandez-Lopez, C.R. Gignoux, J. Ba nhol z-Sloan,
C. Fe nandez-Rozadilla, M. Via, A. Hidalgo-Mi anda, A.V. Con e as, L.U. Figue oa,
P. Raska, e al., De elopmen o a panel o genome-wide ances y in o ma i e
ma ke s o s udy admix u e h oughou he Ame icas, PLoS Gene 8 (2012),
e1002554.
[15] C. Phillips, A. F ei e A adas, A.K. K iegel, M. Fonde ila, O. Bulbul, C. San os,
F. Se ulla Rech, M.D. Pe ez Ca celes, A. Ca acedo, P.M. Schneide , M.V. La eu,
Eu asiaplex: a o ensic SNP assay o di e en ia ing Eu opean and Sou h Asian
ances ies, Fo ensic Sci. In . Gene . 7 (2013) 359–366.
[16] C. San os, C. Phillips, M. Fonde ila, R. Daniel, R.A.H. an Oo scho , E.G. Bu cha d,
M.S. Schan ield, L.J. Sou o, J. Uacyis ael, M. Via, e al., Paci iplex: An ances y-
in o ma i e SNP panel cen ed on Aus alia and he Paci ic egion, Fo ensic Sci.
In . Gene . 20 (2016) 71–80.
[17] C. Ca alho Gon ijo, L.G. Po as-Hu ado, A. F ei e-A adas, M. Fonde ila,
C. San os, A. Salas, J. Henao, C. Isaza, L. Bel ´
an, V. Noguei a Silbige , e al., PIMA:
A popula ion in o ma i e mul iplex o he Ame icas, Fo ensic Sci. In . Gene . 44
(2020), 102200.
[18] A. The 1000 Genomes P ojec Conso ium, L.D. Au on, R.M. B ooks, E.P. Du bin, H.
M. Ga ison, J.O. Kang, J.L. Ko bel, S. Ma chini, G.A. McCa hy, McVean, e al.,
A global e e ence o human gene ic a ia ion, Na u e 526 (2015) 68–74.
[19] J. Amigo, C. Phillips, M. La eu, ´
A. Ca acedo, The SNP o ID b owse : an online ool
o que y and display o equency da a om he SNP o ID p ojec , In . J. Leg. Med
122 (2008) 435–440.
[20] A. Be gs ¨
om, S.A. McCa hy, R. Hui, M.A. Alma i, Q. Ayub, P. Danecek, Y. Chen,
S. Felkel, P. Hallas , J. Kamm, e al., Insigh s in o human gene ic a ia ion and
popula ion his o y om 929 di e se genomes, Science 367 (2020) 1339–1349.
[21] M. By ska-Bishop, U.S. E ani, X. Zhao, A.O. Basile, H.J. Abel, A.A. Regie ,
A. Co elo, W.E. Cla ke, R. Musunu i, K. Nagulapalli, e al., High co e age whole-
genome-sequencing o he expanded 1000 Genomes P ojec coho including 602
ios, Cell 185 (2022) 3426–3440. VCF da a a ailable online: h ps://www.
in e na ionalgenome.o g/da apo al/da a-collec ion/30x-g ch38 and, 〈h p:// p
.1000genomes.ebi.ac.uk/ ol1/ p/da a_collec ions/1000G_2504_high_co e age/
wo king/20190425_NYGC_GATK/〉.
[22] M.A. Alma i, M. Habe , R.A. Loo ah, P. Hallas , S. Al Tu ki, H.C. Ma in, Y. Xue,
C. Tyle -Smi h, The genomic his o y o he Middle Eas , Cell 184 (2021)
4612–4625.
[23] C. Phillips, J. Amigo, A.O. Tillma , M.A. Peck, M. de la Puen e, J. Ruiz-Ramí ez,
F. Bi ne , ˇ
S. Id izbego i´
c, Y. Wang, T.J. Pa sons, e al., A compila ion o i-allelic
SNPs om 1000 Genomes and use o he mos polymo phic loci o a la ge-scale
human iden i ica ion panel, Fo ensic Sci. In . Gene . 46 (2020), 102232.
[24] A. Ral , M. an O en, D. Mon iel Gonz´
alez, P. de Knij , K. an de Beek,
S. Woo on, R. Lagac´
e, M. Kayse , Fo ensic Y-SNP analysis beyond SNaPsho : High-
esolu ion Y-ch omosomal haplog ouping om low quali y and quan i y DNA
using Ion AmpliSeq and a ge ed massi ely pa allel sequencing, Fo ensic Sci. In .
Gene . 41 (2019) 93–106.
[25] J.Z. Li, D.M. Abshe , H. Tang, A.M. Sou hwick, A.M. Cas o, S. Ramachand an, H.
M. Cann, G.S. Ba sh, M. Feldman, L.L. Ca alli-S o za, R.M. Mye s, Wo ldwide
human ela ionships in e ed om genome-wide pa e ns o a ia ion, Science 319
(2008) 1100–1104.
[26] C. Phillips, D. Balla d, P. Gill, D.S. Cou , A. Ca acedo, M.V. La eu, The
ecombina ion landscape a ound o ensic STRs: accu a e measu emen o gene ic
dis ances be ween syn enic STR pai s using HapMap high densi y SNP da a,
Fo ensic Sci. In . Gene . 6 (2012) 345–365.
[27] C. Phillips, D. McNe in, K.K. Kidd, R. Lagac´
e, S. Woo on, M. de la Puen e,
A. F ei e-A adas, A. Mosque a-Miguel, M. Edua do , T.E. G oss, e al., MAPlex-A
massi ely pa allel sequencing ances y analysis mul iplex o Asia-Paci ic
popula ions, Fo ensic Sci. In . Gene . 42 (2019) 213–226.
J. Ruiz-Ramí ez e al.
Fo ensic Science In e na ional: Gene ics 64 (2023) 102853
23
[28] E.Y.Y. Cheung, C. Phillips, M. Edua do , M.V. La eu, D. McNe in, Pe o mance o
ances y-in o ma i e SNP and mic ohaplo ype ma ke s, Fo ensic Sci. In . Gene . 43
(2019), 102141.
[29] K.K. Kidd, W.C. Speed, A.J. Paks is, D.S. Podini, R. Lagac´
e, J. Chang, S. Woo on,
E. Haigh, U. Sounda a ajan, E alua ing 130 mic ohaplo ypes ac oss a global se o
83 popula ions, Fo ensic Sci. In . Gene . 6 (2017) 29–37.
[30] S. Mallick, H. Li, M. Lipson, I. Ma hieson, M. Gym ek, F. Racimo, M. Zhao,
N. Chennagi i, S. No den el , A. Tandon, e al., The simons genome di e si y
p ojec : 300 genomes om 142 di e se popula ions, Na u e 538 (2016) 201–206.
[31] L. Pagani, D.J. Lawson, E. Jagoda, A. M¨
o sebu g, A. E iksson, M. Mi , F. Clemen e,
G. Hudjasho , M. DeGio gio, L. Saag, e al., Genomic analyses in o m on mig a ion
e en s du ing he peopling o Eu asia, Na u e 538 (2016) 238–242.
[32] C. Phillips, J. Amigo, D. McNe in, M. de la Puen e, E.Y.Y. Cheung, M.V. La eu,
Online popula ion da a esou ces o o ensic SNP analysis wi h Massi ely Pa allel
Sequencing: An o e iew o online popula ion da a o o ensic pu poses, in:
E. Pilli, A. Be i (Eds.), In Fo ensic DNA Analysis: Technological De elopmen and
Inno a i e Applica ions, CRC P ess, Boca Ra on, FL, USA, 2021.
[33] A ailable online: h p://ma hgene.usc.es/Snippe / Mul iple p o iles classi ie a :
〈h p://ma hgene.usc.es/snippe /analysismul iplep o iles.h ml〉(bo h accessed 1s
Feb ua y 2023).
[34] J.K. P i cha d, M. S ephens, P. Donnelly, In e ence o popula ion s uc u e using
mul ilocus geno ype da a, Gene ics 155 (2000) 945–959.
[35] N.M. Kopelman, J. Mayzel, M. Jakobsson, N.A. Rosenbe g, I. May ose, Clumpak: a
p og am o iden i ying clus e ing modes and packaging popula ion s uc u e
in e ences ac oss K, Mol. Ecol. Resou . 15 (2015) 1179–1191.
[36] G. E anno, S. Regnau , J. Goude , De ec ing he numbe o clus e s o indi iduals
using he so wa e STRUCTURE: a simula ion s udy, Mol. Ecol. 14 (2005)
2611–2620.
[37] C. San os, C. Phillips, A. Gomez-Ta o, J. Al a ez-Dios, A. Ca acedo, M.V. La eu,
In e ence o ances y in o ensic analysis II: analysis o gene ic da a, Me hods Mol.
Biol. 1420 (2016) 255–285.
[38] M. de la Puen e, C. Phillips, C. Xa ie , J. Amigo, A. Ca acedo, W. Pa son, M.
V. La eu, Building a cus om la ge-scale panel o no el mic ohaplo ypes o o ensic
iden i ica ion using MiSeq and Ion S5 massi ely pa allel sequencing sys ems,
Fo ensic Sci. In . Gene . 48 (2020), 102213.
[39] H. Li, R. Du bin, Fas and accu a e sho ead alignmen wi h Bu ows-Wheele
ans o m, Bioin o ma ics 25 (2009) 1754–1760.
[40] H. Li, B. Handsake , A. Wysoke , T. Fennell, J. Ruan, N. Home , G. Ma h,
G. Abecasis, R. Du bin, The sequence Alignmen /Map o ma and SAM ools,
Bioin o ma ics 25 (2009) 2078–2079.
[41] N. Thomas, R Package - Mic ohaplo , (2019) 〈h ps://gi hub.com/ng homas/mic
ohaplo 〉. (Accessed 1s Feb ua y 2023).
[42] C. Phillips, J. Amigo, A. Ca acedo, M.V. La eu, Te a-allelic SNPs: In o ma i e
o ensic ma ke s compiled om public whole-genome sequence da a, Fo ensic Sci.
In . Gene . 19 (2015) 100–106.
[43] M. Lek, K.J. Ka czewski, E.V. Minikel, K.E. Samocha, E. Banks, T. Fennell, A.
H. O’Donnell-Lu ia, J.S. Wa e, J.A.J. Hill, B.B. Cummings, e al., Analysis o
p o ein-coding gene ic a ia ion in 60,706 humans, Na u e 536 (2016) 285–291.
[44] 〈h p://www.ensembl.o g/Homo_sapiens/Va ia ion/Popula ion?db=co e; =6
:60527829–60528829; = s3857620; db= a ia ion; =169483878〉, (Accessed
1s Feb ua y 2023).
[45] Ø. Bleka, M. Edua do , C. San os, C. Phillips, W. Pa son, P. Gill, Open sou ce
so wa e Eu oFo Mix can be used o analyse complex SNP mix u es, Fo ensic Sci.
In . Gene . 31 (2017) 105–110.
[46] C. Xa ie , M. de la Puen e, M. Mosque a-Miguel, A. F ei e-A adas, V. Kalama a,
A. Re oi , T.E. G oss, P.M. Schneide , C. Ames, C. Hoho , e al., De elopmen and
in e -labo a o y e alua ion o he VISAGE Enhanced Tool o appea ance and
ances y in e ence om DNA, Fo ensic Sci. In . Gene . 61 (2022), 102779.
[47] C. Ca alho Gon ijo, F. Macˆ
edo Mendes, C.A. San os, M. de, N. Klau au-Guima ˜
aes,
M.V. La eu, A. Ca acedo, C. Phillips, S.F. Oli ei a, Ances y analysis in u al
B azilian popula ions o A ican descen , Fo ensic Sci. In . Gene . 36 (2018)
160–166.
[48] N.A. Rosenbe g, J.K. P i cha d, J.L. Webe , H.M. Cann, K.K. Kidd, L.
A. Zhi o o sky, M.W. Feldman, Gene ic s uc u e o human popula ions, Science
298 (2002) 2381–2385.
[49] L. Benne , F. Oldoni, K. Long, S. Cisana, K. Madella, S. Woo on, J. Chang,
R. Hasegawa, R. Lagac´
e, K.K. Kidd, D. Podini, Mix u e decon olu ion by massi ely
pa allel sequencing o mic ohaplo ypes, In . J. Leg. Med. 133 (2019) 719–729.
J. Ruiz-Ramí ez e al.