Anchoring and ordering NGS contig assemblies by population sequencing (POPSEQ)
Full text
TECHNICAL ADVANCE
Ancho ing and o de ing NGS con ig assemblies by
popula ion sequencing (POPSEQ)
Ma in Masche
1,†
, Ga y J. Muehlbaue
2,3,†
, Daniel S. Rokhsa
4,5
, Ja od Chapman
4
, Je emy Schmu z
4,6
, Ke ie Ba y
4
,
Ma
ıa Mu~
noz-Ama ia
ın
2
, Timo hy J. Close
7
, Roge P. Wise
8
, Alan H. Schulman
9
, Axel Himmelbach
1
, Klaus F.X. Maye
10
,
Uwe Scholz
1
, Jesse A. Poland
11
, Nils S ein
1,
* and Robbie Waugh
12,
*
1
Leibniz Ins i u e o Plan Gene ics and C op Plan Resea ch (IPK), D–06466 Seeland OT Ga e sleben, Ge many,
2
Uni e si y o Minneso a, Depa men o Ag onomy and Plan Gene ics, S Paul, MN 55108, USA,
3
Uni e si y o Minneso a, Depa men o Plan Biology, S Paul, MN 55108, USA,
4
Depa men o Ene gy Join Genome Ins i u e, 2800 Mi chell D i e, Walnu C eek, CA 94598, USA,
5
Depa men o Molecula and Cell Biology, Uni e si y o Cali o nia, Be keley, CA 94720, USA,
6
HudsonAlpha Ins i u e o Bio echnology, Hun s ille, AL 35806, USA,
7
Depa men o Bo any & Plan Sciences, Uni e si y o Cali o nia, Ri e side, CA 92521, USA,
8
US Depa men o Ag icul u e/Ag icul u al Resea ch Se ice, Depa men o Plan Pa hology & Mic obiology, Iowa S a e
Uni e si y, Ames, IA 50011–1020, USA,
9
Ins i u e o Bio echnology, Uni e si y o Helsinki/MTT Ag i ood Resea ch, PO Box 65, 00014 Helsinki, Finland,
10
Munich In o ma ion Cen e o P o ein Sequences/Ins i u e o Bioin o ma ics and Sys ems Biology, Helmhol z Zen um
M€
unchen, D–85764 Neuhe be g, Ge many,
11
US Depa men o Ag icul u e/Ag icul u al Resea ch Se ice, Ha d Win e Whea Gene ics Resea ch Uni and Depa men
o Ag onomy, Kansas S a e Uni e si y, Manha an, KS 65506, USA, and
12
Di ision o Plan Sciences, Uni e si y o Dundee a he James Hu on Ins i u e, In e gow ie, Dundee, DD2 5DA, UK
Recei ed 25 June 2013; e ised 7 Augus 2013; accep ed 29 Augus 2013; published online 10 Oc obe 2013.
*Fo co espondence (e-mails [email p o ec ed]; [email p o ec ed]).
†
These au ho s con ibu ed equally o his wo k.
SUMMARY
Nex -gene a ion whole-genome sho gun assemblies o complex genomes a e highly use ul, bu ail o link
nea by sequence con igs wi h each o he o p o ide a linea o de o con igs along indi idual ch omosomes.
He e, we in oduce a s a egy based on sequencing p ogeny o a seg ega ing popula ion ha allows de
no o p oduc ion o a gene ically ancho ed linea assembly o he gene space o an o ganism. We demon-
s a e he powe o he app oach by econs uc ing he ch omosomal o ganiza ion o he gene space o
ba ley, a la ge, complex and highly epe i i e 5.1 Gb genome. We e alua e he obus ness o he new
assembly by compa ison o a ecen ly eleased physical and gene ic amewo k o he ba ley genome, and
o a ious gene ically o de ed sequence-based geno ypic da ase s. The me hod is independen o he need
o any p io sequence esou ces, and will enable apid and cos -e icien es ablishmen o powe ul
genomic in o ma ion o many species.
Keywo ds: nex -gene a ion sequencing, genome assembly, gene ic mapping, ba ley, Ho deum ulga e,
popula ion sequencing, echnical ad ance.
INTRODUCTION
Nex -gene a ion sequencing p o ides he oppo uni y o
apidly es ablish gene space assemblies o i ually any
species a ela i ely low cos . These assemblies consis o
ens o hund eds o housands o sho con iguous pieces
o DNA sequence (con igs), and o en ep esen only he
low-copy po ion o he genome. Despi e he limi a ions o
such assemblies, hey ha e been widely p oposed as su o-
ga es o d a genome sequences o he pu poses o gene
isola ion, genomics-assis ed b eeding and he assessmen
o di e si y wi hin and be ween species (B enchley e al.,
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d
This is an open access a icle unde he e ms o he C ea i e Commons A ibu ion License,
which pe mi s use, dis ibu ion and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed.
718
The Plan Jou nal (2013) 76, 718–727 doi: 10.1111/ pj.12319
2012; In e na ional Ba ley Genome Sequencing Conso -
ium, 2012; Guo e al., 2013; Xu e al., 2013). Howe e , in
mos cases, pa icula ly hose conce ning la ge and com-
plex genomes, hey emain disconnec ed collec ions o
sho sequence con igs ha a e no embedded in a geno-
mic con ex . B inging hese oge he in o a en a i e linea
o de , o e en associa ing con igs wi h indi idual ch omo-
somes o ch omosome a ms, has been a majo and cos ly
unde aking. In a ecen example, he In e na ional Ba ley
Genome Sequencing Conso ium epo ed he de elop-
men and use o a BAC-based physical map, BAC end
sequences, su ey sequences o low-so ed ch omosome
a ms, ully sequenced BAC clones and conse ed syn eny
o ully con ex ualize only 410 Mb o genomic sequence
om he 5.1 Gb ba ley genome (In e na ional Ba ley Gen-
ome Sequencing Conso ium, 2012). These genomic
esou ces p o ide an es ablished pa h owa ds a e e ence
sequence by sequencing a minimum iling pa h o o e lap-
ping BAC clones hie a chically (Feuille e al., 2012). De el-
opmen o he necessa y esou ces equi es a subs an ial
amoun o ime, labo and money, which makes his s a -
egy p ohibi i e o smalle and mo e poo ly esou ced
esea ch communi ies, e.g. esea ch in non-model o gan-
ism o o phan c ops. The es ablishmen o a BAC-based
e e ence sequence o he maize genome ook app oxi-
ma ely 7 yea s, equi ed he coo dina ed e o o se e al
labo a o ies, and cos app oxima ely US $50 million (Chan-
dle and B endel, 2002; Ma ienssen e al., 2004; Schnable
e al., 2009). Simila ly, he e e ence sequence o a single
1 Gb ch omosome o hexaploid whea (T i icum aes i um)
has no been comple ed 5 yea s a e publica ion o a phys-
ical map (Paux e al., 2008).
Eme ging echnologies such as longe sequence eads
(Schad e al., 2010), op ical mapping (Lam e al., 2012) and
no el assembly algo i hms (such as ALLPATHS–LG, Gne e
e al., 2011)) may speed up he p ocess o da a collec ion
and analysis, as well as inc easing he con igui y and com-
ple eness o whole-genome sho gun (WGS) assemblies,
bu hei applicabili y o la ge genomes wi h abundan
sequence epea s ( he bane o any assemble ), a ising om
pa alogous duplica ions, epe i i e elemen s, ances al
duplica ions and polyploidy, emains o be assessed.
I has been common p ac ice o associa e mapped gene ic
ma ke s wi h sequence esou ces based on sequence simi-
la i y in o de o link gene ic and physical maps (Chen e al.,
2002; Wei e al., 2007). While he numbe o BAC con igs on
a physical map is in o de o housands, nex -gene a ion
sequencing (NGS) echnology p oduces hund eds o hou-
sands o sequence con igs. Fo example, he In e na ional
Ba ley Genome Sequencing Conso ium (2012) epo ed an
assembly ha consis s o mo e han 350 000 con igs longe
han 1 kb. The numbe o ma ke s a o ded by con en ional
geno yping s a egies is simply no commensu a e wi h he
la ge numbe o sho sequence con igs.
Se e al me hods o high- h oughpu geno yping o
gene ic mapping popula ions using nex -gene a ion
sequencing echnology ha e been de eloped. Geno yping
by shallow su ey sequencing (0.05–0.19) in he model
species ice (O yza sa i a) has been shown o yield gene ic
maps o unp eceden ed densi y (Chandle and B endel,
2002; Xie e al., 2010). Howe e , he high esolu ion o
ecombina ion b eakpoin s (app oxima ely 40 kb) was p o-
ided by in e ing ma ke o de om a high-quali y e e -
ence sequence. This app oach canno be applied o
species wi h genomes o d a o e en p e-d a quali y as
sequence con igs a e no o ganized in pseudo-molecules
ep esen ing he linea ch omosomes.
The ques ion o how se e al millions o ma ke s p o-
ided by NGS echnology may be used o b ing con igs
in o a linea o de (a p ocedu e commonly e e ed o as
ancho ing) has only en a i ely been aised. Andol a o
e al. (2011) used diges ion wi h a equen ly cu ing
es ic ion enzyme and subsequen mul iplexed sequenc-
ing o a popula ion o 94 indi iduals o assign 8 Mb o un-
assembled con igs o linkage g oups. Simila ly, a educed-
ep esen a ion geno yping-by-sequencing me hod (Poland
e al., 2012) has been ins umen al in ancho ing he ba ley
physical map o a gene ic map (In e na ional Ba ley Gen-
ome Sequencing Conso ium, 2012). Howe e , geno yping
by WGS has no been used as a p ima y ool in de no o
de elopmen o linea ly o de ed d a genome assemblies.
In he absence o an app op ia e molecula o analy ical
me hod o es ablish sho - ange connec i i y (i.e. o link
physically close sequence con igs), we used he powe o
gene ic seg ega ion o di ec ly and linea ly a ange
sequence con igs in o closely associa ed ecombina ion
bins along a a ge genome. We show ha whole-genome
su ey sequencing o a small expe imen al seg ega ing
popula ion and gene ic mapping o he millions o
obse ed single nucleo ide polymo phisms (SNPs)
de ec ed he ein (Figu e 1) as ly imp o es he quali y and
u ili y o highly agmen ed NGS sho gun assemblies. We
illus a e he app oach using he complex 5.1 Gb genome
o cul i a ed ba ley (Ho deum ulga e L.) by compa ing
he ou pu wi h a gene space assembly ha has been
pa ially o de ed using ex ensi e physical and gene ic
mapping esou ces (In e na ional Ba ley Genome Sequenc-
ing Conso ium, 2012). Ou esul s a e cong uen wi h he
cu en sequence assembly (In e na ional Ba ley Genome
Sequencing Conso ium, 2012) bu inc ease he amoun o
gene ically ancho ed con ig sequences by a ac o o h ee.
Mos impo an ly, he whole e o cos <$100K and was
comple ed in a ma e o mon hs. This new assembly has
g ea e alue o compa a i e gene ic s udies, gene isola-
ion and genomics-assis ed b eeding compa ed o he p e-
ious ancho ing e o (In e na ional Ba ley Genome
Sequencing Conso ium, 2012) as mo e WGS con igs a e
posi ioned gene ically. In p inciple, he app oach, which
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
Assembly ancho ing by popula ion sequencing 719
we e m POPSEQ, may be used o any species o which a
seg ega ing popula ion may be de i ed and main ained.
RESULTS
Whole-genome su ey sequencing o gene ic popula ions
We gene a ed su ey sequences om 90 indi iduals
(Table 1) o a popula ion o ecombinan inb ed lines (RILs)
om a c oss be ween ba ley cul i a s Mo ex and Ba ke
(M 9B). DNA om indi idual plan s was agmen ed and
ba -coded, and eigh samples pe lane we e sequenced on
an Illumina HiSeq 2000 ins umen (yielding app oxima ely
19co e age pe line). We de-con olu ed and mapped he
ou pu eads agains a 509WGS sequence assembly o
he ba ley cul i a Mo ex (In e na ional Ba ley Genome
Sequencing Conso ium, 2012) using BWA so wa e (Li
and Du bin, 2009), and pe o med in silico a ian calling
using SAM ools (Li, 2011) (see Expe imen al p ocedu es).
This esul ed in a se o SNP posi ions on he Mo ex WGS
assembly, and geno ype calls (i.e. homozygous o one
pa en o he e ozygous) o each indi idual a each SNP.
A e disca ding a ian posi ions wi h low quali y o oo
much missing da a (Figu e S1), 5.1 million SNPs wi h a
mean o 33 unambiguous geno ypic calls ac oss he popu-
(a)
(b)
(c)
(d)
Figu e 1. Schema ic ep esen a ion o POPSEQ.
(a) A seg ega ing popula ion (80–100 indi iduals) is cons uc ed om a bi-pa en al c oss.
(b) A whole-genome sho gun is gene a ed o one pa en , and used o cons uc a gene space assembly (al e na i ely, he POPSEQ da a i sel may be used o
his pu pose). On his assembly, gene models (g een a ows) a e de ined using RNA–seq. In pa allel, POPSEQ, and, i necessa y, geno yping-by-sequencing
(GBS), is pe o med on he popula ion, and a medium-densi y amewo k gene ic map is calcula ed ( housands o ens o housands o loci).
(c) SNPs de ec ed and yped by POPSEQ along wi h associa ed WGS con igs a e in eg a ed in o he amewo k map h ough nea es -neighbo sea ch.
(d) The esul o POPSEQ is a sequence assembly in linea o de ha con ains comp ehensi e in o ma ion on he gene space. I may be enhanced by pe o ming
POPSEQ on addi ional popula ions.
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
720 Ma in Masche e al.
la ion we e conside ed o in eg a ion in o a high-densi y
SNP-based gene ic map o he same popula ion con-
s uc ed by a ay-based geno yping (Comad an e al.,
2012). We hen used a heu is ic algo i hm o place he
newly disco e ed SNPs in o his exis ing gene ic ame-
wo k. B ie ly, we pe o med a nea es -neighbo sea ch,
que ying he se o amewo k ma ke s o elemen s wi h
minimal Hamming dis ance o a gi en SNP (i.e. he
minimum numbe o al e na i e SNP alleles equi ed o
change an obse ed seg ega ion pa e n in o he e e ence)
I se e al amewo k ma ke s exhibi ed iden ical minimal
dis ances, we imposed a cu o whe e >80% o he ame-
wo k ma ke s had o lie on he same ch omosome and he
median absolu e de ia ion o hei gene ic posi ions was
less han i e cen iMo gans (cM). Using hese h esholds
4.3 million SNPs (85.5% o all de ec ed SNPs) could be
placed in o he gene ic map wi h less han wo geno ype
calls di e ing om hei closes amewo k ma ke .
We hen assigned he WGS sequence con igs ha ha -
bo ed mapped polymo phisms o hei de ined gene ic
posi ions. As wi h posi ioning SNPs in a gene ic map, we
imposed a ule ha mul iple SNPs ound on he same
sequence con ig we e equi ed o ha e conco dan gene ic
posi ions. O e all, 498 856 con igs wi h a cumula i e
leng h o 927 Mb (49.5% o he o al c Mo ex WGS
sequence assembly) could be o de ed along he gene ic
map (Table 2), mo e han doubling he 410 Mb ha was
ancho ed wi h he help o a genome-wide physical map o
he same gene ic amewo k. Tables con aining he
ancho ing esul s a e a ailable o download om
p:// p.ipk-ga e sleben.de/ba ley-popseq/
Valida ion o popula ion sequencing
We checked whe he he gene ic ancho ing gene a ed by
POPSEQ was consis en wi h a ailable sho - ange connec-
i i y in o ma ion. The (In e na ional Ba ley Genome
Sequencing Conso ium, 2012) had sequenced 6278 bac e-
ial a i icial ch omosomes (BACs). Indi iduals BACs we e
sequenced o ‘Phase 1 quali y and consis ed on a e age o
i e o en sequence con igs. F om his se , we iden i ied
3902 clones ha ha bo ed a leas wo WGS con igs ha
we e mapped by POPSEQ. Ou hypo hesis was ha in he
majo i y o cases, pai s o con igs om he same BAC
clone (i.e. wi hin a physical dis ance o less han 200 kb)
would exhibi he same gene ic loca ion. Using ul a-s in-
gen homology (100% iden i y o e 1000 bp), 95% o he
con ig pai s we e placed wi hin a 3 cM window on he
o de ed assembly (Table S1). Disco dan ch omosome
assignmen s we e ound o only 1.7% o he con ig pai s,
and a u he 3.3% had a gene ic dis ance la ge han 3 cM.
Table 1 Sequence da a gene a ed in his s udy
M x B WGS OWB WGS M 9B GBS Mo ex
Popula ion Mo ex 9Ba ke RIL F
8
O egon Wol e Ba leys DH Mo ex 9Ba ke RIL F
8
–
Sequencing echnology Whole-genome sho gun;
HiSeq 2000
Whole-genome
sho gun; HiSeq 2000
Geno yping-by-sequencing;
HiSeq 2000
Whole-genome sho gun;
HiSeq 2000
Numbe o sequencing
lanes
12 12 1 2
Numbe o sequenced
indi iduals
90 (+pa en s) 82 (+pa en s) 92 (+pa en s) 1
App oxima e co e age
pe sample
191919(10 Mb ep esen ed) 159
Numbe o SNPs de ec ed 5 123 696 6 543 684 21 397 –
Mean numbe o p esen
geno ype calls pe ma ke
33 31 58 –
Table 2 Ancho ing s a is ics
M x B (iSelec )
a
OWB M x B (GBS map) M x B +OWB IBSC
Numbe o SNPs used o ancho ing 4 381 020 6 117 837 4 429 475 11 229 709 498 165
F amewo k map iSelec OWB GBS M x B GBS iSelec /OWB GBS iSelec
Numbe o ancho ed con igs 498 856 591 779 512 293 747 077 138 443
Size o ancho ed con igs (Mb) 927 (50%) 1000 (53%) 934 (50%) 1222 (65%) 410 (16%)
Median leng h o ancho ed con igs (bp) 1006 973 977 891 1431
Numbe o ancho ed HC genes
b
16 682 (64%) 15 743 (60%) 16 729 (64%) 20 932 (80%) 15 719 (60%)
Numbe o ancho ed LC genes
c
28 337 (56%) 29 033 (55%) 28 559 (56%) 37 609 (71%) 19 415 (36%)
a
The Mo ex 9Ba ke iSelec amewo k map is desc ibed in In e na ional Ba ley Genome Sequencing Conso ium (2012) and Comad an
e al. (2012).
b
High-con idence genes as desc ibed in In e na ional Ba ley Genome Sequencing Conso ium (2012).
c
Low-con idence genes as desc ibed in In e na ional Ba ley Genome Sequencing Conso ium (2012).
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
Assembly ancho ing by popula ion sequencing 721
We inspec ed 17 BACs wi h a leas i e ancho ed WGS
con igs and disco dan ch omosome assignmen s. Nine o
hese BACs had wo g oups o con igs ancho ed o di e -
en loca ions and had ei he suspiciously la ge inse sizes
o >180 kb sugges i e o chime ic inse s o showed e i-
dence o independen clones ha ing been sequenced
unde he same name.
We hen compa ed he POPSEQ ancho ing o WGS con-
igs o a ecen ly eleased in eg a ed sequence-en iched
gene ic and physical map o ba ley (In e na ional Ba ley
Genome Sequencing Conso ium, 2012). Mo e han 77 000
WGS con igs ( ep esen ing 315 Mb o sequence) we e
assigned by bo h me hods o speci ic gene ic posi ions.
Ch omosome assignmen s disag eed in 2.2% o he cases,
and cM coo dina es di e ed by mo e han 5 cM in 7.0% o
he cases, simila o he 2–8% alse-posi i e a e obse ed
in PCR-based sc eening o BAC lib a ies (In e na ional Ba -
ley Genome Sequencing Conso ium, 2012). In gene al
e ms, incong uence appea s o occu la gely in he highly
epe i i e and ex ensi e gene ic cen ome es. We belie e
ha his is mos likely he p oduc o misplaced epe i i e
sequence-con aining o chime ic BAC con igs in he ba ley
physical map. Thus, employing POPSEQ alongside a ully
sequenced minimum iling pa h highligh s e o s in a
physical map and i s associa ed ancho ing in o ma ion,
and may he eby be aluable in es ablishing a obus
clone-by-clone assembly o a a ge genome.
F amewo k map cons uc ion by GBS
To u he in es iga e he obus ness o POPSEQ, we
assessed he e ec o using a di e en geno yping pla o m
o cons uc he amewo k map. We geno yped he same
90 indi iduals using a wo-enzyme geno yping-by-sequenc-
ing (GBS) app oach (Poland e al., 2012) (Table 1). P io o
sequencing, DNA was diges ed using a a ely and a
equen ly cu ing es ic ion enzyme, and only es ic ion
agmen s wi h wo di e en es ic ion si es we e
sequenced, hus educing he a ge ed in e al on he gen-
ome o app oxima ely 10 Mb. Compa ed o a ay-based
geno yping, GBS has lowe pe -sample cos s and does no
equi e any p io knowledge o polymo phisms be ween
he pa en s o he popula ion. Ins ead, ma ke de ec ion and
sco ing occu simul aneously, making GBS sui able o spe-
cies wi hou any genomic esou ces, o o which genomic
esou ces a e poo ly de eloped. We cons uc ed a de no o
gene ic map comp ising 4056 bi-allelic SNP ma ke s, and
placed WGS con igs in o his map using he same algo i hm
as desc ibed abo e. Al oge he , 927 Mb o sequence ep e-
sen ed by 512 293 sequence con igs was o de ed (Table 2),
wi h 94.3% also linked o he iSelec amewo k (Comad an
e al., 2012). Impo an ly, he gene ic coo dina es o con igs
we e consis en among he unde lying amewo k maps
(Figu e 2b): ch omosome assignmen s we e disco dan in
0.1% o he cases, and he map posi ion o only 0.6% o he
con igs di e ed by mo e han 5 cM. I we only used he
SNP ma ke s (app oxima ely 20 000) p o ided by GBS, we
we e able o ancho only 49 Mb o sequence, because he
numbe o ancho ed con igs is limi ed by he numbe o
a ailable SNPs.
Robus ness o he linea assembly
To es he obus ness o he M x B POPSEQ ancho ed
assembly, we cons uc ed a de no o assembly o a second
popula ion o compa ison. We used he O egon Wol e
Ba ley (OWB) popula ion, as a gene ic map om GBS on
82 doubled haploid (DH) lines was al eady a ailable
(Poland e al., 2012). We su ey-sequenced hese 82 indi-
iduals o app oxima ely 19whole genome co e age each
(Table 1), and, by pe o ming he same s eps as o M x B,
assigned gene ic posi ions o 591 779 WGS con igs co e-
sponding o 1000 Mb o sequence. O hese con igs, 42%
(295 Mb) we e no ancho ed o he M x B iSelec
(a) (b) (c)
Figu e 2. POPSEQ alida ion. WGS con igs ancho ed o h ee gene ic maps. These plo s show he colinea i y o con igs ancho ed o he Mo ex 9Ba ke iSelec
amewo k map and (a) he physical and gene ic amewo k o ba ley (In e na ional Ba ley Genome Sequencing Conso ium, 2012), (b) a Mo ex 9Ba ke gene ic
map cons uc ed by geno yping-by-sequencing (GBS), (c) a GBS map (Poland e al., 2012) cons uc ed in he OWB. WGS con igs a e shown as do s, and a e
mos ly wi hin 5 cM o he diagonal: 90.8% in (a), 99.2% in (b) 93.2% in (c).
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
722 Ma in Masche e al.
amewo k. In mos cases, hese con igs ei he ha bo ed
no polymo phism be ween Mo ex and Ba ke, o SNPs
we e no assayed in a su icien numbe o RILs o each
ou h eshold o inclusion. Con igs ancho ed o bo h
M x B and OWB maps had highly cong uen ch omosome
assignmen s (99.6% ag eemen , Figu e 2c). Only 6.4% o
all con igs we e placed mo e han 5 cM apa in he wo
ancho ed assemblies ( alling o 2.1% i <7 cM). Gi en ha
we we e compa ing popula ions cons uc ed wi h di e en
pa en s and le els o ecombina ion (app oxima ely hal in
a DH popula ion compa ed o RILs), his was no com-
ple ely unexpec ed. Howe e , he use o independen pop-
ula ions o ancho ing has conside able alue: he
cumula i e leng h o con igs ancho ed o ei he he M x B
o OWB map is 1.22 Gb, an inc ease o one- hi d compa ed
o use o only a single popula ion. Addi ional polymo -
phisms in OWB hus enabled placemen o con igs ha
we e iden ical be ween Mo ex and Ba ke. Mo e impo -
an ly, he POPSEQ o de ed assembly posi ions an addi-
ional 5213 anno a ed high-con idence genes on he ba ley
genome compa ed o he In e na ional Ba ley Genome
Sequencing Conso ium elease.
F amewo k map cons uc ion using ligh sho gun
popula ion sequencing
We hen explo ed whe he he POPSEQ da a could be used
di ec ly o cons uc a obus de no o gene ic map wi hou
e e ence o o he da ase s o geno yping me hods. B ie ly,
we iden i ied a se o 65 357 con igs con aining a leas en
Mo ex/Ba ke SNPs pe con ig, equi ing ha hese con igs
be geno yped by shallow WGS sequencing in a leas 75 o
he 90 indi iduals wi hin ou M 9B mapping popula ion
o a oid an excess o missing da a poin s. Using s ingen
con ols on log-odds sco es, 98.5% o hese con igs we e
eadily clus e ed in o se en majo linkage g oups and
o de ed by MSTMap (Wu e al., 2008). The esul ing ame-
wo k map has app oxima ely 99% conco dance wi h exis -
ing ba ley maps (Pea son co ela ion coe icien ), and may
be used o place addi ional con igs wi h ewe SNPs and/o
mo e limi ed sampling using a majo i y ule app oach as
desc ibed abo e. Thus POPSEQ da a may be used di ec ly
o gene a e a linea o de ing o con igs, e en in he
absence o an independen gene ic map.
POPSEQ does no equi e long ma e-pai lib a ies
The se o whole-genome con igs ( he ‘ e e ence assembly’)
used in he p esen s udy had been assembled om Illu-
mina lib a ies wi h agmen sizes o 350 bp and 2.5 kb
(In e na ional Ba ley Genome Sequencing Conso ium,
2012). Al hough la ge-inse ma e-pai lib a ies may be used
o es ablish links be ween con igs, and may be equi ed
inpu o some assemble s (Gne e e al., 2011), he con-
s uc ion o such lib a ies is no s aigh o wa d and o en
yields sub-op imal esul s, such as a high ac ion o PCR
duplica es o sho -inse ead pai s. We he e o e explo ed
how POPSEQ pe o med using an assembly comp ising
only sho -inse pai ed eads. We sequenced he same
350 bp inse lib a ies used o cons uc ion o he cu en
ba ley e e ence assembly (In e na ional Ba ley Genome
Sequencing Conso ium, 2012) on wo HiSeq lanes, yielding
app oxima ely 159haploid genome co e age (Table 1),
and assembled he eads using he same p og am as p e i-
ously (In e na ional Ba ley Genome Sequencing Conso -
ium, 2012). As he ead co e age was app oxima ely h ee
imes lowe han used by In e na ional Ba ley Genome
Sequencing Conso ium (2012) and did no u ilize ma e-pai
in o ma ion, we expec ed he assembly o be o wo se qual-
i y. The cumula i e leng h o he esul ing assembly was
sho e (1.6 Gb e sus 1.9 Gb), and he con ig N50 (a
weigh ed a e age con ig size ha is commonly used as a
measu e o assembly con igui y) was smalle (1238 bp e -
sus 1450 bp). Howe e , con igs o his size a e su icien o
unc ion as a e e ence o ead mapping and o enable
s uc u al gene anno a ion ia RNA sequencing (RNA-seq)
as well as SNP de ec ion. No ably, almos hal o he con igs
(49.8%) ancho ed o he M x B iSelec amewo k a e
sho e han 1000 bp. In species wi h smalle and less epe -
i i e genomes, WGS assembly is expec ed o yield ewe
and longe con igs ha po en ially yield a highe numbe o
SNPs pe con ig (depending upon he le el o polymo -
phism in he POPSEQ popula ion). Al e na i ely, la ge con-
igs may compensa e o lowe le els o polymo phism.
DISCUSSION
Low-co e age (app oxima ely 0.05–0.19) NGS su ey
sequencing o he small genome (0.4 Gb) o he model
c op plan ice has p e iously been used as a ool o gene -
a e many housands o gene ic ma ke s o bo h bi-pa en-
al linkage s udies and GWAS (Huang e al., 2009, 2010).
The e ec i eness o his ‘geno yping by e-sequencing’
was a o ded by he a ailabili y o high-quali y e e ence
sequences, a small a ge genome wi h compa a i ely ew
epea s, and inno a i e s a is ical app oaches o da a
analysis. He e, we ha e explo ed a undamen ally di e en
applica ion o NGS combined wi h classical gene ic analy-
sis ha should ind applica ion in many species, pa icu-
la ly hose wi h ecalci an , la ge o poo ly cha ac e ized
genomes, among hem economically impo an species
such as whea , suga cane, pine o Miscan hus.
We explo ed POPSEQ as a me hod o gene ically ancho -
ing and o de ing de no o NGS assemblies, and ha e dem-
ons a ed i s po en ial by e-syn hesizing and imp o ing a
ecen ly eleased sequence assembly o he la ge (5.1 Gb)
and complex (>80% epe i i e sequence, ances ally dupli-
ca ed) ba ley genome. We used sequence da a om wo
mapping popula ions, and used he la ge numbe o
de ec ed SNPs o in eg a e he sequence assembly wi h wo
es ablished amewo k maps as well as gene ic maps
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
Assembly ancho ing by popula ion sequencing 723
compu ed om GBS o WGS da a. A i s co e, POPSEQ
exploi s he powe o gene ic seg ega ion combined wi h
shallow (1–29pe line) su ey sequencing o one o mo e
small expe imen al popula ions o gene ically ancho NGS
sequence assemblies. I is independen o physical map-
ping and all o he genomic esou ces ypically de eloped in
la ge genome sequencing p ojec s, and should be amena-
ble o applica ion in mos popula ion ypes.
We show ha POPSEQ is bo h obus and ep oducible.
Using a ious gene ic maps and mapping popula ion, we
ob ain compa able esul s wi h a conco dance o app oxi-
ma ely 95%. Thus, POPSEQ is nei he dependen upon he
choice o mapping popula ion no he geno yping pla o m
used o amewo k map cons uc ion. I mo e ex ensi e
sho - ange connec i i y is es ablished by longe sequence
con igs o sca olds (se o o de ed sequence con igs wi h
gaps be ween hem), a sliding window app oach (Huang
e al., 2009) may be used o geno ype calling and ame-
wo k map cons uc ion om POPSEQ da a alone, a oiding
he need o GBS o SNP mapping pla o ms. In addi ion,
pa i ioning o polymo phic si es acco ding o hei pa en-
al o igin may be pe o med p io o de no o assembly, o
example by using he colo ed de B uijn g aph me hod (Iq-
bal e al., 2012). The aw sequence eads om POPSEQ
( he equi alen o 509 o each pa en ) should hen be su -
icien o compu e he e e ence sequence assemblies ha
will ul ima ely be o de ed along he gene ic map.
POPSEQ pe o ms e ec i ely wi h highly agmen ed
sequence assemblies om sho -inse lib a ies. We we e
able o cons uc a de no o WGS assembly om sho
Illumina eads ha showed assembly s a is ics compa able
o an assembly ha inco po a ed ma e-pai in o ma ion.
POPSEQ hus a oids he echnical di icul ies associa ed
wi h cons uc ion and cha ac e iza ion o la ge-inse
lib a ies. The simul aneous use o se e al mapping popula-
ions h ough sequence-based consensus map cons uc ion
is s aigh o wa d, wi h he same ca ea s as obse ed in any
gene ic map in eg a ion. The ou come is no me ely an ul a-
dense gene ic map o anonymous loci: a each gene ic posi-
ion, comp ehensi e in o ma ion on he gene space may be
ob ained h ough RNA–seq-based s uc u al anno a ion.
The POPSEQ esou ce we de eloped he e bo h ep o-
duces and subs an ially imp o es he mul i-laye ed gene
space assembly ha was he esul o a la ge collabo a i e
e o by he In e na ional Ba ley Genome Sequencing
Conso ium o e many yea s. By compa ison, POPSEQ is
inexpensi e, apid and concep ually simple, he mos
ime-consuming s ep being he cons uc ion o a mapping
popula ion. In ela ion o he la e , while we used bo h DH
lines and RILs, o he popula ion ypes including ea ly-gen-
e a ion inb ed lines (e.g. F
4
indi iduals) would also be sui -
able. Subsequen s eps including sequence assembly om
sho -inse lib a ies, geno yping-by-sequencing (i
equi ed) and in eg a i e compu a ional analyses may be
pe o med quickly. We s ess ha we do no ad oca e aban-
donmen o on-going genome p ojec s ha a e pu suing a
clone-by-clone s a egy. On he con a y, we belie e hese
may p o i om POPSEQ. BAC con igs may be alida ed
hough gene ic mapping o each single clone, and he high
numbe o mapped gene ic ma ke s should allow i ually
any ully sequenced physical con ig o be accu a ely placed.
Ha ing pe o med a p oo o p inciple in ba ley, he
no ion o ad ancing he closely ela ed b ead whea gen-
ome (Paux e al., 2008) by adop ing POPSEQ is o pa icula
in e es . Whea will be he las o he wo ld’s majo c ops o
be ully sequenced. The cos -e icien cons uc ion o high-
densi y gene ic maps is ou ine in hexaploid whea (Poland
e al., 2012), and he challenge o dis inguishing homoeolo-
gous sequences has been la gely o e come: sub-genome-
speci ic sho gun assemblies ha e been eleased ecen ly
(B enchley e al., 2012), and ch omosome-speci ic su ey
sequences ha e also been gene a ed (He nandez e al.,
2012). Fu he mo e, se e al popula ions o ecombinan
inb ed lines a e al eady a ailable wi hin he academic and
comme cial sec o s, and a e ipe o exploi a ion (Nelson
e al., 1995; Manicka elu e al., 2011). While he ul ima e
goal should be a clone-by-clone sequence o he whea gen-
ome wi h a quali y on pa wi h ha o he he maize gen-
ome, POPSEQ opens he way o ob ain, wi h compa a i e
ease, an e ec i e su oga e ha would be aluable o basic
esea ch and b eeding applica ions. In addi ion o whea ,
many non-model species, o phan c ops and old gene ic
models such as pea (Pisum sa i um), ha e no ye bene i ed
much om he genomics e a. Wi h mode a e e o , POPS-
EQ could allow he gene a ion o highly use ul sequence
esou ces o hese and many o he species.
Fo an uncha ac e ized >5 Gb diploid genome, be ween
14 and 30 HiSeq lanes a e equi ed o (i) p oducing a de
no o sequence assembly o ead mapping ( wo o eigh
lanes; no equi ed i POPSEQ da a i sel is used o p oduce
he ‘ e e ence’ sequence assembly); (ii) geno yping-by-
sequencing o map cons uc ion (one lane; no equi ed i
POPSEQ is used o cons uc he e e ence map); (iii) shal-
low popula ion sequencing (minimum 12 lanes o each
popula ion o app oxima ely 90 lines, al hough he dep h
may be a ied); (i ) deep RNA-seq o s uc u al gene anno-
a ion (mo e han wo lanes) amoun ing o $50 000–
$100 000 in sequencing cos s. Toge he wi h a medium-
sized compu e se e (32 CPU co es, 512 GB RAM, 3 TB o
disk space), i is possible o gene a e a de no o linea gene
space assembly.
The accu acy o POPSEQ may be imp o ed i he mem-
be s o he popula ion a e sequenced o highe dep h. Wi h
he sequencing dep h used in his s udy (1–29), he
sequencing eads o each indi idual co e only app oxi-
ma ely 50% o he assembly. Doubling he amoun o
sequencing da a pe indi idual would esul in genome
co e age o app oxima ely 80% acco ding o he model o
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
724 Ma in Masche e al.
Lande and Wa e man (1988) (Figu e S2), hus educing he
numbe o missing geno ype calls pe indi idual. An
inc ease in sequencing dep h is manda o y o highly he -
e ozygous popula ions such as F
2
popula ions in sel ing
o ganisms o F
1
popula ions in ou c ossing species in
o de o co ec ly ype he e ozygous SNPs. Using an
imp o ed assembly wi h longe con igs o con igs o ga-
nized in o physically close sca olds would bene i he
analysis, as mo e SNPs could be used o place each
sequence con ig. An inc ease in he numbe o sequenced
indi iduals ( esul ing in a p opo ional inc ease in he
sequencing load) may imp o e he gene ic esolu ion o
he amewo k map.
We p opose ha POPSEQ may con ibu e subs an ially
o undamen al esea ch in plan gene ics as well as in
c op imp o emen ( o examples, see Figu e S3, Appendix
S1 and Me hods S1). Howe e , i s applica ion is no
es ic ed o plan s. The as and s eady ad ances in
sequencing echnology will u he inc ease he powe o
POPSEQ, allowing deepe co e age o la ge and ou b ed
popula ions. As long as he inhe en complexi y o
genomes es ic s he assembly o pseudo-molecules by
sho gun sequencing, POPSEQ p o ides a apid, low-cos
and e ec i e me hod o de eloping a highly use ul
‘in e im e e ence’ genome sequence in mos species o
which i is possible o cons uc a gene ic map.
EXPERIMENTAL PROCEDURES
Whole-genome sho gun sequencing
Illumina pai ed-end lib a ies( agmen size app oxima ely 350 bp)
we e gene a ed om agmen ed genomic DNA o 90 indi iduals
om he Mo ex 9Ba ke RIL popula ion and 82 indi iduals o he
OWB popula ion. Indi idual lib a ies we e ba -coded p io o com-
bining in pools o eigh and sequencing on Illumina HiSeq 2000
ins umen s (h p://www.illumina.com). The iSelec amewo k
was a ailable om a p e ious s udy (Comad an e al., 2012).
Read mapping and SNP calling
Sequencing eads we e quali y immed and mapped agains he
Mo ex WGS assembly (In e na ional Ba ley Genome Sequencing
Conso ium, 2012) using BWA e sion 0.6.2 (Li and Du bin, 2009).
The command ‘bwa aln’ was used wi h he pa ame e ‘-q 15’ o
quali y imming, o he wise de aul pa ame e s we e used. A e
emo ing duplica e eads using he SAM ools (Li, 2011) command
mdup, a ian posi ions and geno ypes o indi iduals a a ian
posi ions we e called using he SAM ools mpileup/bc ools pipe-
line e sion 0.1.18 (Li, 2011) wi h de aul pa ame e s. Addi ionally,
he pa ame e ‘-D’ was used o SAM ools mpileup o eco d pe -
sample ead dep h. The esul ing VCF ile was il e ed using a cus-
om AWK sc ip . The sc ip emo ed SNPs wi h a SAM ools qual-
i y sco e below 40, and u he il e ed SAM ools geno ype calls: a
homozygous geno ype call was e ained i he e was a leas one
ead suppo ing i and i s SAM ools geno ype quali y was a leas
3. In he M x B da a, a he e ozygous call was e ained i he e
we e a leas h ee suppo ing eads and i s sco e was a leas 5.
In he OWB DH popula ion, he e ozygous calls we e always dis-
ca ded. Geno ype calls no ma ching he speci ied c i e ia we e
se o a missing alue. A a ian posi ion was emo ed i mo e
han 10% o all samples we e called he e ozygous, he e we e
mo e han 80% missing da a, o he mino allele equency (in he
non-missing da a) was smalle han 5%.
Mapping SNPs and WGS con igs o he amewo k map
The nea es neighbo s o SNPs de ec ed in he WGS sho gun da a
we e sea ched o using a heu is ic algo i hm implemen ed in
GNU C. The sou ce code is a ailable om p:// p.ipk-ga e sleben.
de/ba ley-popseq/. As a me ic, we used he minimum Hamming
dis ance. The nea es neighbo s we e sea ched o in he se o (i)
1723 non- edundan iSelec SNPs, (ii) 4056 GBS SNPs used o con-
s uc ion o he M x B GBS map, and (iii) 4632 non- edundan OWB
GBS SNPs mee ing hese c i e ia desc ibed below. A SNP was con-
side ed edundan i he e was ano he SNP wi h he same geno-
ype (on he non-missing da a) and he same gene ic posi ion.
[Co ec ion added on 25 Oc obe 2013 a e o iginal online publica-
ion on 10 Oc obe 2013: p:// p.ga e sleben.de/ba ley-popseq/
was changed o p:// p.ipk-ga e sleben.de/ba ley-popseq/]
SNPs we e used o ancho WGS con igs i hey we e sco ed
unequi ocally on mo e han 20% o he indi iduals in he popula-
ion, he Hamming dis ance (numbe o di e en , non-missing
geno ypes) o hei nea es SNPs was no la ge han 2, a leas
80% o all nea es SNPs lay on he same ch omosome, and he
median absolu e de ia ion o he cM posi ions (on he ch omo-
some wi h mos ma ke s) was <5 o he OWB map and he M x B
iSelec amewo k. As we used he popula ion ype DH o he
M x B RILs (as equi ed o ad anced RILs), he M x B GBS map
o e -es ima ed he map leng h by a ac o o app oxima ely 3, and
we allowed a maximal median absolu e de ia ion o 15. The cM
coo dina e o a SNP mee ing hese c i e ia was de ined as he
median cM posi ion o i s nea es neighbo s. A WGS con ig was
assigned o a gene ic posi ion i a leas 80% o all SNPs loca ed
on i had been mapped o he same ch omosome and he median
absolu e de ia ion o he cM coo dina es o he SNPs was <5 (15
o M 9B GBS). The cM posi ion o a con ig was se o he
median cM posi ion o all SNPs loca ed on he con ig.
Es ima ion o he e o a e
WGS con igs we e compa ed using MegaBLAST e sion 2.2.26
(Zhang e al., 2000) o 6278 ully sequenced BACs. Unde s ingen
c i e ia, we equi ed 100% iden i y and a minimum alignmen
leng h o 1000 bp o each BLAST High Sco ing Pai . Unde
elaxed c i e ia, we equi ed 99% iden i y and 200 bp minimum
alignmen leng h. The gene ic posi ions o all pai s o con igs on
he same BAC we e compa ed (Table S1). BACs wi h disco dan
ch omosome assignmen s and wi h hi s o a leas i e ancho ed
con igs we e u he analyzed. Fo each BAC, he ch omosome
assignmen s o i s con igs we e abula ed. I a leas 30% o all
con igs on a BAC we e ancho ed o he ch omosome wi h he sec-
ond highes numbe o con igs, he BAC was deemed p oblema ic,
and we checked whe he i had been sequenced wice o i s leng h
( he cumula i e leng h o i s assembled sequence con igs) was
unusually la ge (>180 kb).
Gene ic map cons uc ion om M x B GBS da a
GBS lib a y p oduc ion and sequencing o M x B popula ions
we e pe o med as desc ibed p e iously (Poland e al., 2012).
Reads we e decon olu ed using a cus om AWK sc ip . Adap e
sequences we e emo ed using cu adap e sion 1.1 (h p://
code.google.com/p/cu adap ). T immed eads sho e han 30 bp
we e disca ded. Read mapping, SNP and geno ype calling, and il-
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
Assembly ancho ing by popula ion sequencing 725
e ing we e pe o med essen ially as desc ibed abo e o he WGS
da a. As only single ends we e used, he BWA command samse
was used o alignmen . Addi ionally, only SNPs mee ing he
ollowing c i e ia we e conside ed o gene ic map cons uc ion:
no mo e han 10% missing da a, no mo e han 10% he e ozygous
geno ypes, jAþBj=ðAþBÞ 0:7, whe e Aand Bindica e he
coun s o he pa en al alleles; in he absence o he e ozygous
calls, his co esponds o a minimum mino allele equency o
17.6%. Fo M 9B, 4058 SNPs passed hese il e s. Gene ic map
cons uc ion was pe o med using MSTMap (Wu e al., 2008) wi h
he ollowing pa ame e s: popula ion ype DH; dis ance unc ion
kosambi; cu _o _p_ alue 0.00001; no_map_dis 20; no_map_size
2; missing_ h eshold 0.8; es ima ion_be o e_clus e ing no;
de ec _bad_da a yes; objec i e_ unc ion COUNT. The esul ing
map con ained se en linkage g oups wi h mo e han one ma ke .
Two ma ke s o med a linkage g oup o hei own and we e dis-
ca ded. Acco ding o he ob ained o de s, o ien a ions and dis-
ances be ween ma ke s, he linkage g oups co esponded o he
se en ba ley ch omosomes. The ela ionship be ween gene ic
posi ions in he new map and he iSelec map was ob ained
h ough Loess eg ession [R (h p://www. -p ojec .o g) unc ion
loess, smoo he span 0.3]. In e pola ion in o he iSelec map o
WGS SNP posi ions in eg a ed o he GBS amewo k was
pe o med using he loess model wi h he R unc ion p edic .
De no o map cons uc ion om POPSEQ
To build an independen gene ic map om he POPSEQ da a wi h-
ou e e ence o exis ing maps o o he ma ke da a, we es ic ed
ou a en ion o he 115 258 sequence con igs ha span a leas
en SNPs ha a e polymo phic be ween he wo pa en s Mo ex
and Ba ke. Fo he pu poses o de eloping a amewo k map, we
u he es ic ed ou a en ion o con igs wi h highly conco dan
SNP geno ype calls. We he e o e se aside con igs ha had wo
o mo e SNP geno ype calls om bo h pa en s, indica ing he pos-
sibili y o mis-geno yping h ough inco ec SNP calls and/o lim-
i ed c oss-con amina ion be ween indi iduals. The esul ing
80 189 con igs we e hen geno yped as ei he Mo ex o Ba ke
based on he consensus o hei geno yped SNPs, equi ing a
leas h ee SNP calls. Finally, o he amewo k map, we only con-
side ed con igs ha could be consensus geno yped in a leas 75
o he 90 indi iduals. This le us wi h 66 357 con igs ha could be
eliably geno yped wi h limi ed missing da a. We compu ed he
ecombina ion a e and loga i hm o odds (LOD) sco e be ween
each pai o con igs, and clus e ed con igs wi h LOD >10 o o m
linkage g oups: 64 476/65 357 (98.7%) o con igs o med 14 link-
age g oups, wi h app oxima ely 98.87% o con ig leng h placed in
se en majo linkage g oups, co esponding o he se en ba ley
ch omosomes.
In eg a ion o WGS SNPs o he OWB GBS bin map
A bin map (Poland e al., 2012) had p e iously been cons uc ed
om GBS da a o 82 OWB DH lines. GBS ma ke sequences (64 bp
long) we e aligned agains he Mo ex WGS assembly using he
BWA command ‘aln’ and he command ‘samse’. Only alignmen s
wi h he bes possible mapping sco e o 37 we e conside ed. SNPs
wi h missing da a o he pa en s o mo e han 10% missing da a on
he DH lines we e no conside ed o he nea es -neighbo sea ch.
The ancho ing o SNPs and con igs has been desc ibed abo e.
De no o assembly
Illumina pai ed-end lib a ies (inse size 350 bp) o ba ley cul i a
Mo ex had been cons uc ed p e iously (In e na ional Ba ley Gen-
ome Sequencing Conso ium, 2012). Sequencing on he Illumina
HiSeq 2000 was pe o med acco ding o he manu ac u e ’s p oce-
du es (h p://www.illumina.com). Sequencing eads we e quali y
immed and assembled using CLC assembly cell 3.2.2 (h p://
www.clcbio.com).
ACKNOWLEDGMENTS
The wo k pe o med by he US Depa men o Ene gy Join Gen-
ome Ins i u e is suppo ed by he O ice o Science o he US
Depa men o Ene gy unde con ac numbe DE–AC02-
05CH11231. The au ho s would also like o acknowledge he
suppo gi en by unds ecei ed om he T i iceae Coo dina ed
Ag icul u al P ojec , US Depa men o Ag icul u e/Na ional Ins i-
u e o Food and Ag icul u e g an numbe 2011-68002-30029 o
G.J.M. and T.J.C., he Sco ish Go e nmen Ru al and En i onmen
Science and Analy ical Se ices Di ision Resea ch P og amme o
R.W., and he Bundesminis e ium €
u Bildung und Fo schung (TRI-
TEX 0315954) o N.S. and U.S. We hank Sa ah Ayling (The Gen-
ome Analysis Cen e, No wich, UK) o help ul discussions abou
simula ing ead co e age. Finally, we acknowledge S. Taudien and
M. Pla ze (F i z Lipmann Ins i u e, Jena, Ge many), o p o iding a
pai ed-end lib a y o c . Mo ex o HiSeq 2000 sequencing, and D.
S engel o sequence da a submission.
ACCESSION NUMBERS
The accession numbe s o he sequences desc ibed a e
ERP002183 (GBS sequence da a o he Mo ex 9Ba ke RILs)
and ERP002184 (whole-genome sho gun sequence da a o
he Mo ex 9Ba ke RILs and OWB DH lines).
SUPPORTING INFORMATION
Addi ional Suppo ing In o ma ion may be ound in he online e -
sion o his a icle.
Figu e S1. Dis ibu ion o he numbe o success ul geno ype calls
a a ian posi ions de ec ed in he whole da a o he Mo ex x
Ba ke and OWB popula ions.
Figu e S2. Obse ed and expec ed sequence co e age acco ding
o he model o Lande and Wa e man (1988).
Figu e S3. Po en ial uses o an assembly o de ed by POPSEQ.
Table S1. The pe cen age o WGS con igs pai s assigned o he
same BAC ha a e posi ioned a he apa han he speci ied dis-
ance.
Appendix S1. Applica ions o a POPSEQ assembly o compa a i e
genomics, e e ence-based gene ic mapping and gene isola ion.
Me hods S1. Expe imen al p ocedu es o Appendix S1.
REFERENCES
Andol a o, P., Da ison, D., E ezyilmaz, D., Hu, T.T., Mas , J.,
Sunayama-Mo i a, T. and S e n, D.L. (2011) Mul iplexed sho gun geno-
yping o apid and e icien gene ic mapping. Genome Res. 21, 610–617.
B enchley, R., Spannagl, M., P ei e , M. e al. (2012) Analysis o he b ead
whea genome using whole-genome sho gun sequencing. Na u e,491,
705–710.
Chandle , V.L. and B endel, V. (2002) The Maize Genome Sequencing P o-
jec . Plan Physiol. 130, 1594–1597.
Chen, M., P es ing, G., Ba bazuk, W.B. e al. (2002) An in eg a ed physical
and gene ic map o he ice genome. Plan Cell,14, 537–545.
Comad an, J., Kilian, B., Russell, J. e al. (2012) Na u al a ia ion in a homo-
log o An i hinum CENTRORADIALIS con ibu ed o sp ing g ow h habi
and en i onmen al adap a ion in cul i a ed ba ley. Na . Gene . 44, 1388–
1392.
Feuille , C., S ein, N., Rossini, L., P aud, S., Maye , K., Schulman, A., E e -
sole, K. and Appels, R. (2012) In eg a ing ce eal genomics o suppo
inno a ion in he T i iceae. Func . In eg . Genomics,12, 573–583.
©2013 The Au ho s
The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727
726 Ma in Masche e al.