scieee Open visual document viewer

Anchoring and ordering NGS contig assemblies by population sequencing (POPSEQ)

Mascher, Martin,Muehlbauer, Gary J.,Rokhsar, Daniel S.,Chapman, Jarrod,Schmutz, Jeremy,Barry, Kerrie,Munoz-Amatriaín, María,Close, Timothy J.,Wise, Roger P.,Schulman, Alan H.,Himmelbach, Axel,Mayer, Klaus F.X.,Scholz, Uwe,Poland, Jesse A.,Stein, Nils,Wau

Full text

TECHNICAL ADVANCE Ancho ing and o de ing NGS con ig assemblies by popula ion sequencing (POPSEQ) Ma in Masche 1,† , Ga y J. Muehlbaue 2,3,† , Daniel S. Rokhsa 4,5 , Ja od Chapman 4 , Je emy Schmu z 4,6 , Ke ie Ba y 4 , Ma  ıa Mu~ noz-Ama ia ın 2 , Timo hy J. Close 7 , Roge P. Wise 8 , Alan H. Schulman 9 , Axel Himmelbach 1 , Klaus F.X. Maye 10 , Uwe Scholz 1 , Jesse A. Poland 11 , Nils S ein 1, * and Robbie Waugh 12, * 1 Leibniz Ins i u e o Plan Gene ics and C op Plan Resea ch (IPK), D–06466 Seeland OT Ga e sleben, Ge many, 2 Uni e si y o Minneso a, Depa men o Ag onomy and Plan Gene ics, S Paul, MN 55108, USA, 3 Uni e si y o Minneso a, Depa men o Plan Biology, S Paul, MN 55108, USA, 4 Depa men o Ene gy Join Genome Ins i u e, 2800 Mi chell D i e, Walnu C eek, CA 94598, USA, 5 Depa men o Molecula and Cell Biology, Uni e si y o Cali o nia, Be keley, CA 94720, USA, 6 HudsonAlpha Ins i u e o Bio echnology, Hun s ille, AL 35806, USA, 7 Depa men o Bo any & Plan Sciences, Uni e si y o Cali o nia, Ri e side, CA 92521, USA, 8 US Depa men o Ag icul u e/Ag icul u al Resea ch Se ice, Depa men o Plan Pa hology & Mic obiology, Iowa S a e Uni e si y, Ames, IA 50011–1020, USA, 9 Ins i u e o Bio echnology, Uni e si y o Helsinki/MTT Ag i ood Resea ch, PO Box 65, 00014 Helsinki, Finland, 10 Munich In o ma ion Cen e o P o ein Sequences/Ins i u e o Bioin o ma ics and Sys ems Biology, Helmhol z Zen um M€ unchen, D–85764 Neuhe be g, Ge many, 11 US Depa men o Ag icul u e/Ag icul u al Resea ch Se ice, Ha d Win e Whea Gene ics Resea ch Uni and Depa men o Ag onomy, Kansas S a e Uni e si y, Manha an, KS 65506, USA, and 12 Di ision o Plan Sciences, Uni e si y o Dundee a he James Hu on Ins i u e, In e gow ie, Dundee, DD2 5DA, UK Recei ed 25 June 2013; e ised 7 Augus 2013; accep ed 29 Augus 2013; published online 10 Oc obe 2013. *Fo co espondence (e-mails [email p o ec ed]; [email p o ec ed]). † These au ho s con ibu ed equally o his wo k. SUMMARY Nex -gene a ion whole-genome sho gun assemblies o complex genomes a e highly use ul, bu ail o link nea by sequence con igs wi h each o he o p o ide a linea o de o con igs along indi idual ch omosomes. He e, we in oduce a s a egy based on sequencing p ogeny o a seg ega ing popula ion ha allows de no o p oduc ion o a gene ically ancho ed linea assembly o he gene space o an o ganism. We demon- s a e he powe o he app oach by econs uc ing he ch omosomal o ganiza ion o he gene space o ba ley, a la ge, complex and highly epe i i e 5.1 Gb genome. We e alua e he obus ness o he new assembly by compa ison o a ecen ly eleased physical and gene ic amewo k o he ba ley genome, and o a ious gene ically o de ed sequence-based geno ypic da ase s. The me hod is independen o he need o any p io sequence esou ces, and will enable apid and cos -e icien es ablishmen o powe ul genomic in o ma ion o many species. Keywo ds: nex -gene a ion sequencing, genome assembly, gene ic mapping, ba ley, Ho deum ulga e, popula ion sequencing, echnical ad ance. INTRODUCTION Nex -gene a ion sequencing p o ides he oppo uni y o apidly es ablish gene space assemblies o i ually any species a ela i ely low cos . These assemblies consis o ens o hund eds o housands o sho con iguous pieces o DNA sequence (con igs), and o en ep esen only he low-copy po ion o he genome. Despi e he limi a ions o such assemblies, hey ha e been widely p oposed as su o- ga es o d a genome sequences o he pu poses o gene isola ion, genomics-assis ed b eeding and he assessmen o di e si y wi hin and be ween species (B enchley e al., ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d This is an open access a icle unde he e ms o he C ea i e Commons A ibu ion License, which pe mi s use, dis ibu ion and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. 718 The Plan Jou nal (2013) 76, 718–727 doi: 10.1111/ pj.12319 2012; In e na ional Ba ley Genome Sequencing Conso - ium, 2012; Guo e al., 2013; Xu e al., 2013). Howe e , in mos cases, pa icula ly hose conce ning la ge and com- plex genomes, hey emain disconnec ed collec ions o sho sequence con igs ha a e no embedded in a geno- mic con ex . B inging hese oge he in o a en a i e linea o de , o e en associa ing con igs wi h indi idual ch omo- somes o ch omosome a ms, has been a majo and cos ly unde aking. In a ecen example, he In e na ional Ba ley Genome Sequencing Conso ium epo ed he de elop- men and use o a BAC-based physical map, BAC end sequences, su ey sequences o low-so ed ch omosome a ms, ully sequenced BAC clones and conse ed syn eny o ully con ex ualize only 410 Mb o genomic sequence om he 5.1 Gb ba ley genome (In e na ional Ba ley Gen- ome Sequencing Conso ium, 2012). These genomic esou ces p o ide an es ablished pa h owa ds a e e ence sequence by sequencing a minimum iling pa h o o e lap- ping BAC clones hie a chically (Feuille e al., 2012). De el- opmen o he necessa y esou ces equi es a subs an ial amoun o ime, labo and money, which makes his s a - egy p ohibi i e o smalle and mo e poo ly esou ced esea ch communi ies, e.g. esea ch in non-model o gan- ism o o phan c ops. The es ablishmen o a BAC-based e e ence sequence o he maize genome ook app oxi- ma ely 7 yea s, equi ed he coo dina ed e o o se e al labo a o ies, and cos app oxima ely US $50 million (Chan- dle and B endel, 2002; Ma ienssen e al., 2004; Schnable e al., 2009). Simila ly, he e e ence sequence o a single 1 Gb ch omosome o hexaploid whea (T i icum aes i um) has no been comple ed 5 yea s a e publica ion o a phys- ical map (Paux e al., 2008). Eme ging echnologies such as longe sequence eads (Schad e al., 2010), op ical mapping (Lam e al., 2012) and no el assembly algo i hms (such as ALLPATHS–LG, Gne e e al., 2011)) may speed up he p ocess o da a collec ion and analysis, as well as inc easing he con igui y and com- ple eness o whole-genome sho gun (WGS) assemblies, bu hei applicabili y o la ge genomes wi h abundan sequence epea s ( he bane o any assemble ), a ising om pa alogous duplica ions, epe i i e elemen s, ances al duplica ions and polyploidy, emains o be assessed. I has been common p ac ice o associa e mapped gene ic ma ke s wi h sequence esou ces based on sequence simi- la i y in o de o link gene ic and physical maps (Chen e al., 2002; Wei e al., 2007). While he numbe o BAC con igs on a physical map is in o de o housands, nex -gene a ion sequencing (NGS) echnology p oduces hund eds o hou- sands o sequence con igs. Fo example, he In e na ional Ba ley Genome Sequencing Conso ium (2012) epo ed an assembly ha consis s o mo e han 350 000 con igs longe han 1 kb. The numbe o ma ke s a o ded by con en ional geno yping s a egies is simply no commensu a e wi h he la ge numbe o sho sequence con igs. Se e al me hods o high- h oughpu geno yping o gene ic mapping popula ions using nex -gene a ion sequencing echnology ha e been de eloped. Geno yping by shallow su ey sequencing (0.05–0.19) in he model species ice (O yza sa i a) has been shown o yield gene ic maps o unp eceden ed densi y (Chandle and B endel, 2002; Xie e al., 2010). Howe e , he high esolu ion o ecombina ion b eakpoin s (app oxima ely 40 kb) was p o- ided by in e ing ma ke o de om a high-quali y e e - ence sequence. This app oach canno be applied o species wi h genomes o d a o e en p e-d a quali y as sequence con igs a e no o ganized in pseudo-molecules ep esen ing he linea ch omosomes. The ques ion o how se e al millions o ma ke s p o- ided by NGS echnology may be used o b ing con igs in o a linea o de (a p ocedu e commonly e e ed o as ancho ing) has only en a i ely been aised. Andol a o e al. (2011) used diges ion wi h a equen ly cu ing es ic ion enzyme and subsequen mul iplexed sequenc- ing o a popula ion o 94 indi iduals o assign 8 Mb o un- assembled con igs o linkage g oups. Simila ly, a educed- ep esen a ion geno yping-by-sequencing me hod (Poland e al., 2012) has been ins umen al in ancho ing he ba ley physical map o a gene ic map (In e na ional Ba ley Gen- ome Sequencing Conso ium, 2012). Howe e , geno yping by WGS has no been used as a p ima y ool in de no o de elopmen o linea ly o de ed d a genome assemblies. In he absence o an app op ia e molecula o analy ical me hod o es ablish sho - ange connec i i y (i.e. o link physically close sequence con igs), we used he powe o gene ic seg ega ion o di ec ly and linea ly a ange sequence con igs in o closely associa ed ecombina ion bins along a a ge genome. We show ha whole-genome su ey sequencing o a small expe imen al seg ega ing popula ion and gene ic mapping o he millions o obse ed single nucleo ide polymo phisms (SNPs) de ec ed he ein (Figu e 1) as ly imp o es he quali y and u ili y o highly agmen ed NGS sho gun assemblies. We illus a e he app oach using he complex 5.1 Gb genome o cul i a ed ba ley (Ho deum ulga e L.) by compa ing he ou pu wi h a gene space assembly ha has been pa ially o de ed using ex ensi e physical and gene ic mapping esou ces (In e na ional Ba ley Genome Sequenc- ing Conso ium, 2012). Ou esul s a e cong uen wi h he cu en sequence assembly (In e na ional Ba ley Genome Sequencing Conso ium, 2012) bu inc ease he amoun o gene ically ancho ed con ig sequences by a ac o o h ee. Mos impo an ly, he whole e o cos <$100K and was comple ed in a ma e o mon hs. This new assembly has g ea e alue o compa a i e gene ic s udies, gene isola- ion and genomics-assis ed b eeding compa ed o he p e- ious ancho ing e o (In e na ional Ba ley Genome Sequencing Conso ium, 2012) as mo e WGS con igs a e posi ioned gene ically. In p inciple, he app oach, which ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 Assembly ancho ing by popula ion sequencing 719 we e m POPSEQ, may be used o any species o which a seg ega ing popula ion may be de i ed and main ained. RESULTS Whole-genome su ey sequencing o gene ic popula ions We gene a ed su ey sequences om 90 indi iduals (Table 1) o a popula ion o ecombinan inb ed lines (RILs) om a c oss be ween ba ley cul i a s Mo ex and Ba ke (M 9B). DNA om indi idual plan s was agmen ed and ba -coded, and eigh samples pe lane we e sequenced on an Illumina HiSeq 2000 ins umen (yielding app oxima ely 19co e age pe line). We de-con olu ed and mapped he ou pu eads agains a 509WGS sequence assembly o he ba ley cul i a Mo ex (In e na ional Ba ley Genome Sequencing Conso ium, 2012) using BWA so wa e (Li and Du bin, 2009), and pe o med in silico a ian calling using SAM ools (Li, 2011) (see Expe imen al p ocedu es). This esul ed in a se o SNP posi ions on he Mo ex WGS assembly, and geno ype calls (i.e. homozygous o one pa en o he e ozygous) o each indi idual a each SNP. A e disca ding a ian posi ions wi h low quali y o oo much missing da a (Figu e S1), 5.1 million SNPs wi h a mean o 33 unambiguous geno ypic calls ac oss he popu- (a) (b) (c) (d) Figu e 1. Schema ic ep esen a ion o POPSEQ. (a) A seg ega ing popula ion (80–100 indi iduals) is cons uc ed om a bi-pa en al c oss. (b) A whole-genome sho gun is gene a ed o one pa en , and used o cons uc a gene space assembly (al e na i ely, he POPSEQ da a i sel may be used o his pu pose). On his assembly, gene models (g een a ows) a e de ined using RNA–seq. In pa allel, POPSEQ, and, i necessa y, geno yping-by-sequencing (GBS), is pe o med on he popula ion, and a medium-densi y amewo k gene ic map is calcula ed ( housands o ens o housands o loci). (c) SNPs de ec ed and yped by POPSEQ along wi h associa ed WGS con igs a e in eg a ed in o he amewo k map h ough nea es -neighbo sea ch. (d) The esul o POPSEQ is a sequence assembly in linea o de ha con ains comp ehensi e in o ma ion on he gene space. I may be enhanced by pe o ming POPSEQ on addi ional popula ions. ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 720 Ma in Masche e al. la ion we e conside ed o in eg a ion in o a high-densi y SNP-based gene ic map o he same popula ion con- s uc ed by a ay-based geno yping (Comad an e al., 2012). We hen used a heu is ic algo i hm o place he newly disco e ed SNPs in o his exis ing gene ic ame- wo k. B ie ly, we pe o med a nea es -neighbo sea ch, que ying he se o amewo k ma ke s o elemen s wi h minimal Hamming dis ance o a gi en SNP (i.e. he minimum numbe o al e na i e SNP alleles equi ed o change an obse ed seg ega ion pa e n in o he e e ence) I se e al amewo k ma ke s exhibi ed iden ical minimal dis ances, we imposed a cu o whe e >80% o he ame- wo k ma ke s had o lie on he same ch omosome and he median absolu e de ia ion o hei gene ic posi ions was less han i e cen iMo gans (cM). Using hese h esholds 4.3 million SNPs (85.5% o all de ec ed SNPs) could be placed in o he gene ic map wi h less han wo geno ype calls di e ing om hei closes amewo k ma ke . We hen assigned he WGS sequence con igs ha ha - bo ed mapped polymo phisms o hei de ined gene ic posi ions. As wi h posi ioning SNPs in a gene ic map, we imposed a ule ha mul iple SNPs ound on he same sequence con ig we e equi ed o ha e conco dan gene ic posi ions. O e all, 498 856 con igs wi h a cumula i e leng h o 927 Mb (49.5% o he o al c Mo ex WGS sequence assembly) could be o de ed along he gene ic map (Table 2), mo e han doubling he 410 Mb ha was ancho ed wi h he help o a genome-wide physical map o he same gene ic amewo k. Tables con aining he ancho ing esul s a e a ailable o download om p:// p.ipk-ga e sleben.de/ba ley-popseq/ Valida ion o popula ion sequencing We checked whe he he gene ic ancho ing gene a ed by POPSEQ was consis en wi h a ailable sho - ange connec- i i y in o ma ion. The (In e na ional Ba ley Genome Sequencing Conso ium, 2012) had sequenced 6278 bac e- ial a i icial ch omosomes (BACs). Indi iduals BACs we e sequenced o ‘Phase 1 quali y and consis ed on a e age o i e o en sequence con igs. F om his se , we iden i ied 3902 clones ha ha bo ed a leas wo WGS con igs ha we e mapped by POPSEQ. Ou hypo hesis was ha in he majo i y o cases, pai s o con igs om he same BAC clone (i.e. wi hin a physical dis ance o less han 200 kb) would exhibi he same gene ic loca ion. Using ul a-s in- gen homology (100% iden i y o e 1000 bp), 95% o he con ig pai s we e placed wi hin a 3 cM window on he o de ed assembly (Table S1). Disco dan ch omosome assignmen s we e ound o only 1.7% o he con ig pai s, and a u he 3.3% had a gene ic dis ance la ge han 3 cM. Table 1 Sequence da a gene a ed in his s udy M x B WGS OWB WGS M 9B GBS Mo ex Popula ion Mo ex 9Ba ke RIL F 8 O egon Wol e Ba leys DH Mo ex 9Ba ke RIL F 8 – Sequencing echnology Whole-genome sho gun; HiSeq 2000 Whole-genome sho gun; HiSeq 2000 Geno yping-by-sequencing; HiSeq 2000 Whole-genome sho gun; HiSeq 2000 Numbe o sequencing lanes 12 12 1 2 Numbe o sequenced indi iduals 90 (+pa en s) 82 (+pa en s) 92 (+pa en s) 1 App oxima e co e age pe sample 191919(10 Mb ep esen ed) 159 Numbe o SNPs de ec ed 5 123 696 6 543 684 21 397 – Mean numbe o p esen geno ype calls pe ma ke 33 31 58 – Table 2 Ancho ing s a is ics M x B (iSelec ) a OWB M x B (GBS map) M x B +OWB IBSC Numbe o SNPs used o ancho ing 4 381 020 6 117 837 4 429 475 11 229 709 498 165 F amewo k map iSelec OWB GBS M x B GBS iSelec /OWB GBS iSelec Numbe o ancho ed con igs 498 856 591 779 512 293 747 077 138 443 Size o ancho ed con igs (Mb) 927 (50%) 1000 (53%) 934 (50%) 1222 (65%) 410 (16%) Median leng h o ancho ed con igs (bp) 1006 973 977 891 1431 Numbe o ancho ed HC genes b 16 682 (64%) 15 743 (60%) 16 729 (64%) 20 932 (80%) 15 719 (60%) Numbe o ancho ed LC genes c 28 337 (56%) 29 033 (55%) 28 559 (56%) 37 609 (71%) 19 415 (36%) a The Mo ex 9Ba ke iSelec amewo k map is desc ibed in In e na ional Ba ley Genome Sequencing Conso ium (2012) and Comad an e al. (2012). b High-con idence genes as desc ibed in In e na ional Ba ley Genome Sequencing Conso ium (2012). c Low-con idence genes as desc ibed in In e na ional Ba ley Genome Sequencing Conso ium (2012). ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 Assembly ancho ing by popula ion sequencing 721 We inspec ed 17 BACs wi h a leas i e ancho ed WGS con igs and disco dan ch omosome assignmen s. Nine o hese BACs had wo g oups o con igs ancho ed o di e - en loca ions and had ei he suspiciously la ge inse sizes o >180 kb sugges i e o chime ic inse s o showed e i- dence o independen clones ha ing been sequenced unde he same name. We hen compa ed he POPSEQ ancho ing o WGS con- igs o a ecen ly eleased in eg a ed sequence-en iched gene ic and physical map o ba ley (In e na ional Ba ley Genome Sequencing Conso ium, 2012). Mo e han 77 000 WGS con igs ( ep esen ing 315 Mb o sequence) we e assigned by bo h me hods o speci ic gene ic posi ions. Ch omosome assignmen s disag eed in 2.2% o he cases, and cM coo dina es di e ed by mo e han 5 cM in 7.0% o he cases, simila o he 2–8% alse-posi i e a e obse ed in PCR-based sc eening o BAC lib a ies (In e na ional Ba - ley Genome Sequencing Conso ium, 2012). In gene al e ms, incong uence appea s o occu la gely in he highly epe i i e and ex ensi e gene ic cen ome es. We belie e ha his is mos likely he p oduc o misplaced epe i i e sequence-con aining o chime ic BAC con igs in he ba ley physical map. Thus, employing POPSEQ alongside a ully sequenced minimum iling pa h highligh s e o s in a physical map and i s associa ed ancho ing in o ma ion, and may he eby be aluable in es ablishing a obus clone-by-clone assembly o a a ge genome. F amewo k map cons uc ion by GBS To u he in es iga e he obus ness o POPSEQ, we assessed he e ec o using a di e en geno yping pla o m o cons uc he amewo k map. We geno yped he same 90 indi iduals using a wo-enzyme geno yping-by-sequenc- ing (GBS) app oach (Poland e al., 2012) (Table 1). P io o sequencing, DNA was diges ed using a a ely and a equen ly cu ing es ic ion enzyme, and only es ic ion agmen s wi h wo di e en es ic ion si es we e sequenced, hus educing he a ge ed in e al on he gen- ome o app oxima ely 10 Mb. Compa ed o a ay-based geno yping, GBS has lowe pe -sample cos s and does no equi e any p io knowledge o polymo phisms be ween he pa en s o he popula ion. Ins ead, ma ke de ec ion and sco ing occu simul aneously, making GBS sui able o spe- cies wi hou any genomic esou ces, o o which genomic esou ces a e poo ly de eloped. We cons uc ed a de no o gene ic map comp ising 4056 bi-allelic SNP ma ke s, and placed WGS con igs in o his map using he same algo i hm as desc ibed abo e. Al oge he , 927 Mb o sequence ep e- sen ed by 512 293 sequence con igs was o de ed (Table 2), wi h 94.3% also linked o he iSelec amewo k (Comad an e al., 2012). Impo an ly, he gene ic coo dina es o con igs we e consis en among he unde lying amewo k maps (Figu e 2b): ch omosome assignmen s we e disco dan in 0.1% o he cases, and he map posi ion o only 0.6% o he con igs di e ed by mo e han 5 cM. I we only used he SNP ma ke s (app oxima ely 20 000) p o ided by GBS, we we e able o ancho only 49 Mb o sequence, because he numbe o ancho ed con igs is limi ed by he numbe o a ailable SNPs. Robus ness o he linea assembly To es he obus ness o he M x B POPSEQ ancho ed assembly, we cons uc ed a de no o assembly o a second popula ion o compa ison. We used he O egon Wol e Ba ley (OWB) popula ion, as a gene ic map om GBS on 82 doubled haploid (DH) lines was al eady a ailable (Poland e al., 2012). We su ey-sequenced hese 82 indi- iduals o app oxima ely 19whole genome co e age each (Table 1), and, by pe o ming he same s eps as o M x B, assigned gene ic posi ions o 591 779 WGS con igs co e- sponding o 1000 Mb o sequence. O hese con igs, 42% (295 Mb) we e no ancho ed o he M x B iSelec (a) (b) (c) Figu e 2. POPSEQ alida ion. WGS con igs ancho ed o h ee gene ic maps. These plo s show he colinea i y o con igs ancho ed o he Mo ex 9Ba ke iSelec amewo k map and (a) he physical and gene ic amewo k o ba ley (In e na ional Ba ley Genome Sequencing Conso ium, 2012), (b) a Mo ex 9Ba ke gene ic map cons uc ed by geno yping-by-sequencing (GBS), (c) a GBS map (Poland e al., 2012) cons uc ed in he OWB. WGS con igs a e shown as do s, and a e mos ly wi hin 5 cM o he diagonal: 90.8% in (a), 99.2% in (b) 93.2% in (c). ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 722 Ma in Masche e al. amewo k. In mos cases, hese con igs ei he ha bo ed no polymo phism be ween Mo ex and Ba ke, o SNPs we e no assayed in a su icien numbe o RILs o each ou h eshold o inclusion. Con igs ancho ed o bo h M x B and OWB maps had highly cong uen ch omosome assignmen s (99.6% ag eemen , Figu e 2c). Only 6.4% o all con igs we e placed mo e han 5 cM apa in he wo ancho ed assemblies ( alling o 2.1% i <7 cM). Gi en ha we we e compa ing popula ions cons uc ed wi h di e en pa en s and le els o ecombina ion (app oxima ely hal in a DH popula ion compa ed o RILs), his was no com- ple ely unexpec ed. Howe e , he use o independen pop- ula ions o ancho ing has conside able alue: he cumula i e leng h o con igs ancho ed o ei he he M x B o OWB map is 1.22 Gb, an inc ease o one- hi d compa ed o use o only a single popula ion. Addi ional polymo - phisms in OWB hus enabled placemen o con igs ha we e iden ical be ween Mo ex and Ba ke. Mo e impo - an ly, he POPSEQ o de ed assembly posi ions an addi- ional 5213 anno a ed high-con idence genes on he ba ley genome compa ed o he In e na ional Ba ley Genome Sequencing Conso ium elease. F amewo k map cons uc ion using ligh sho gun popula ion sequencing We hen explo ed whe he he POPSEQ da a could be used di ec ly o cons uc a obus de no o gene ic map wi hou e e ence o o he da ase s o geno yping me hods. B ie ly, we iden i ied a se o 65 357 con igs con aining a leas en Mo ex/Ba ke SNPs pe con ig, equi ing ha hese con igs be geno yped by shallow WGS sequencing in a leas 75 o he 90 indi iduals wi hin ou M 9B mapping popula ion o a oid an excess o missing da a poin s. Using s ingen con ols on log-odds sco es, 98.5% o hese con igs we e eadily clus e ed in o se en majo linkage g oups and o de ed by MSTMap (Wu e al., 2008). The esul ing ame- wo k map has app oxima ely 99% conco dance wi h exis - ing ba ley maps (Pea son co ela ion coe icien ), and may be used o place addi ional con igs wi h ewe SNPs and/o mo e limi ed sampling using a majo i y ule app oach as desc ibed abo e. Thus POPSEQ da a may be used di ec ly o gene a e a linea o de ing o con igs, e en in he absence o an independen gene ic map. POPSEQ does no equi e long ma e-pai lib a ies The se o whole-genome con igs ( he ‘ e e ence assembly’) used in he p esen s udy had been assembled om Illu- mina lib a ies wi h agmen sizes o 350 bp and 2.5 kb (In e na ional Ba ley Genome Sequencing Conso ium, 2012). Al hough la ge-inse ma e-pai lib a ies may be used o es ablish links be ween con igs, and may be equi ed inpu o some assemble s (Gne e e al., 2011), he con- s uc ion o such lib a ies is no s aigh o wa d and o en yields sub-op imal esul s, such as a high ac ion o PCR duplica es o sho -inse ead pai s. We he e o e explo ed how POPSEQ pe o med using an assembly comp ising only sho -inse pai ed eads. We sequenced he same 350 bp inse lib a ies used o cons uc ion o he cu en ba ley e e ence assembly (In e na ional Ba ley Genome Sequencing Conso ium, 2012) on wo HiSeq lanes, yielding app oxima ely 159haploid genome co e age (Table 1), and assembled he eads using he same p og am as p e i- ously (In e na ional Ba ley Genome Sequencing Conso - ium, 2012). As he ead co e age was app oxima ely h ee imes lowe han used by In e na ional Ba ley Genome Sequencing Conso ium (2012) and did no u ilize ma e-pai in o ma ion, we expec ed he assembly o be o wo se qual- i y. The cumula i e leng h o he esul ing assembly was sho e (1.6 Gb e sus 1.9 Gb), and he con ig N50 (a weigh ed a e age con ig size ha is commonly used as a measu e o assembly con igui y) was smalle (1238 bp e - sus 1450 bp). Howe e , con igs o his size a e su icien o unc ion as a e e ence o ead mapping and o enable s uc u al gene anno a ion ia RNA sequencing (RNA-seq) as well as SNP de ec ion. No ably, almos hal o he con igs (49.8%) ancho ed o he M x B iSelec amewo k a e sho e han 1000 bp. In species wi h smalle and less epe - i i e genomes, WGS assembly is expec ed o yield ewe and longe con igs ha po en ially yield a highe numbe o SNPs pe con ig (depending upon he le el o polymo - phism in he POPSEQ popula ion). Al e na i ely, la ge con- igs may compensa e o lowe le els o polymo phism. DISCUSSION Low-co e age (app oxima ely 0.05–0.19) NGS su ey sequencing o he small genome (0.4 Gb) o he model c op plan ice has p e iously been used as a ool o gene - a e many housands o gene ic ma ke s o bo h bi-pa en- al linkage s udies and GWAS (Huang e al., 2009, 2010). The e ec i eness o his ‘geno yping by e-sequencing’ was a o ded by he a ailabili y o high-quali y e e ence sequences, a small a ge genome wi h compa a i ely ew epea s, and inno a i e s a is ical app oaches o da a analysis. He e, we ha e explo ed a undamen ally di e en applica ion o NGS combined wi h classical gene ic analy- sis ha should ind applica ion in many species, pa icu- la ly hose wi h ecalci an , la ge o poo ly cha ac e ized genomes, among hem economically impo an species such as whea , suga cane, pine o Miscan hus. We explo ed POPSEQ as a me hod o gene ically ancho - ing and o de ing de no o NGS assemblies, and ha e dem- ons a ed i s po en ial by e-syn hesizing and imp o ing a ecen ly eleased sequence assembly o he la ge (5.1 Gb) and complex (>80% epe i i e sequence, ances ally dupli- ca ed) ba ley genome. We used sequence da a om wo mapping popula ions, and used he la ge numbe o de ec ed SNPs o in eg a e he sequence assembly wi h wo es ablished amewo k maps as well as gene ic maps ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 Assembly ancho ing by popula ion sequencing 723 compu ed om GBS o WGS da a. A i s co e, POPSEQ exploi s he powe o gene ic seg ega ion combined wi h shallow (1–29pe line) su ey sequencing o one o mo e small expe imen al popula ions o gene ically ancho NGS sequence assemblies. I is independen o physical map- ping and all o he genomic esou ces ypically de eloped in la ge genome sequencing p ojec s, and should be amena- ble o applica ion in mos popula ion ypes. We show ha POPSEQ is bo h obus and ep oducible. Using a ious gene ic maps and mapping popula ion, we ob ain compa able esul s wi h a conco dance o app oxi- ma ely 95%. Thus, POPSEQ is nei he dependen upon he choice o mapping popula ion no he geno yping pla o m used o amewo k map cons uc ion. I mo e ex ensi e sho - ange connec i i y is es ablished by longe sequence con igs o sca olds (se o o de ed sequence con igs wi h gaps be ween hem), a sliding window app oach (Huang e al., 2009) may be used o geno ype calling and ame- wo k map cons uc ion om POPSEQ da a alone, a oiding he need o GBS o SNP mapping pla o ms. In addi ion, pa i ioning o polymo phic si es acco ding o hei pa en- al o igin may be pe o med p io o de no o assembly, o example by using he colo ed de B uijn g aph me hod (Iq- bal e al., 2012). The aw sequence eads om POPSEQ ( he equi alen o 509 o each pa en ) should hen be su - icien o compu e he e e ence sequence assemblies ha will ul ima ely be o de ed along he gene ic map. POPSEQ pe o ms e ec i ely wi h highly agmen ed sequence assemblies om sho -inse lib a ies. We we e able o cons uc a de no o WGS assembly om sho Illumina eads ha showed assembly s a is ics compa able o an assembly ha inco po a ed ma e-pai in o ma ion. POPSEQ hus a oids he echnical di icul ies associa ed wi h cons uc ion and cha ac e iza ion o la ge-inse lib a ies. The simul aneous use o se e al mapping popula- ions h ough sequence-based consensus map cons uc ion is s aigh o wa d, wi h he same ca ea s as obse ed in any gene ic map in eg a ion. The ou come is no me ely an ul a- dense gene ic map o anonymous loci: a each gene ic posi- ion, comp ehensi e in o ma ion on he gene space may be ob ained h ough RNA–seq-based s uc u al anno a ion. The POPSEQ esou ce we de eloped he e bo h ep o- duces and subs an ially imp o es he mul i-laye ed gene space assembly ha was he esul o a la ge collabo a i e e o by he In e na ional Ba ley Genome Sequencing Conso ium o e many yea s. By compa ison, POPSEQ is inexpensi e, apid and concep ually simple, he mos ime-consuming s ep being he cons uc ion o a mapping popula ion. In ela ion o he la e , while we used bo h DH lines and RILs, o he popula ion ypes including ea ly-gen- e a ion inb ed lines (e.g. F 4 indi iduals) would also be sui - able. Subsequen s eps including sequence assembly om sho -inse lib a ies, geno yping-by-sequencing (i equi ed) and in eg a i e compu a ional analyses may be pe o med quickly. We s ess ha we do no ad oca e aban- donmen o on-going genome p ojec s ha a e pu suing a clone-by-clone s a egy. On he con a y, we belie e hese may p o i om POPSEQ. BAC con igs may be alida ed hough gene ic mapping o each single clone, and he high numbe o mapped gene ic ma ke s should allow i ually any ully sequenced physical con ig o be accu a ely placed. Ha ing pe o med a p oo o p inciple in ba ley, he no ion o ad ancing he closely ela ed b ead whea gen- ome (Paux e al., 2008) by adop ing POPSEQ is o pa icula in e es . Whea will be he las o he wo ld’s majo c ops o be ully sequenced. The cos -e icien cons uc ion o high- densi y gene ic maps is ou ine in hexaploid whea (Poland e al., 2012), and he challenge o dis inguishing homoeolo- gous sequences has been la gely o e come: sub-genome- speci ic sho gun assemblies ha e been eleased ecen ly (B enchley e al., 2012), and ch omosome-speci ic su ey sequences ha e also been gene a ed (He nandez e al., 2012). Fu he mo e, se e al popula ions o ecombinan inb ed lines a e al eady a ailable wi hin he academic and comme cial sec o s, and a e ipe o exploi a ion (Nelson e al., 1995; Manicka elu e al., 2011). While he ul ima e goal should be a clone-by-clone sequence o he whea gen- ome wi h a quali y on pa wi h ha o he he maize gen- ome, POPSEQ opens he way o ob ain, wi h compa a i e ease, an e ec i e su oga e ha would be aluable o basic esea ch and b eeding applica ions. In addi ion o whea , many non-model species, o phan c ops and old gene ic models such as pea (Pisum sa i um), ha e no ye bene i ed much om he genomics e a. Wi h mode a e e o , POPS- EQ could allow he gene a ion o highly use ul sequence esou ces o hese and many o he species. Fo an uncha ac e ized >5 Gb diploid genome, be ween 14 and 30 HiSeq lanes a e equi ed o (i) p oducing a de no o sequence assembly o ead mapping ( wo o eigh lanes; no equi ed i POPSEQ da a i sel is used o p oduce he ‘ e e ence’ sequence assembly); (ii) geno yping-by- sequencing o map cons uc ion (one lane; no equi ed i POPSEQ is used o cons uc he e e ence map); (iii) shal- low popula ion sequencing (minimum 12 lanes o each popula ion o app oxima ely 90 lines, al hough he dep h may be a ied); (i ) deep RNA-seq o s uc u al gene anno- a ion (mo e han wo lanes) amoun ing o $50 000– $100 000 in sequencing cos s. Toge he wi h a medium- sized compu e se e (32 CPU co es, 512 GB RAM, 3 TB o disk space), i is possible o gene a e a de no o linea gene space assembly. The accu acy o POPSEQ may be imp o ed i he mem- be s o he popula ion a e sequenced o highe dep h. Wi h he sequencing dep h used in his s udy (1–29), he sequencing eads o each indi idual co e only app oxi- ma ely 50% o he assembly. Doubling he amoun o sequencing da a pe indi idual would esul in genome co e age o app oxima ely 80% acco ding o he model o ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 724 Ma in Masche e al. Lande and Wa e man (1988) (Figu e S2), hus educing he numbe o missing geno ype calls pe indi idual. An inc ease in sequencing dep h is manda o y o highly he - e ozygous popula ions such as F 2 popula ions in sel ing o ganisms o F 1 popula ions in ou c ossing species in o de o co ec ly ype he e ozygous SNPs. Using an imp o ed assembly wi h longe con igs o con igs o ga- nized in o physically close sca olds would bene i he analysis, as mo e SNPs could be used o place each sequence con ig. An inc ease in he numbe o sequenced indi iduals ( esul ing in a p opo ional inc ease in he sequencing load) may imp o e he gene ic esolu ion o he amewo k map. We p opose ha POPSEQ may con ibu e subs an ially o undamen al esea ch in plan gene ics as well as in c op imp o emen ( o examples, see Figu e S3, Appendix S1 and Me hods S1). Howe e , i s applica ion is no es ic ed o plan s. The as and s eady ad ances in sequencing echnology will u he inc ease he powe o POPSEQ, allowing deepe co e age o la ge and ou b ed popula ions. As long as he inhe en complexi y o genomes es ic s he assembly o pseudo-molecules by sho gun sequencing, POPSEQ p o ides a apid, low-cos and e ec i e me hod o de eloping a highly use ul ‘in e im e e ence’ genome sequence in mos species o which i is possible o cons uc a gene ic map. EXPERIMENTAL PROCEDURES Whole-genome sho gun sequencing Illumina pai ed-end lib a ies( agmen size app oxima ely 350 bp) we e gene a ed om agmen ed genomic DNA o 90 indi iduals om he Mo ex 9Ba ke RIL popula ion and 82 indi iduals o he OWB popula ion. Indi idual lib a ies we e ba -coded p io o com- bining in pools o eigh and sequencing on Illumina HiSeq 2000 ins umen s (h p://www.illumina.com). The iSelec amewo k was a ailable om a p e ious s udy (Comad an e al., 2012). Read mapping and SNP calling Sequencing eads we e quali y immed and mapped agains he Mo ex WGS assembly (In e na ional Ba ley Genome Sequencing Conso ium, 2012) using BWA e sion 0.6.2 (Li and Du bin, 2009). The command ‘bwa aln’ was used wi h he pa ame e ‘-q 15’ o quali y imming, o he wise de aul pa ame e s we e used. A e emo ing duplica e eads using he SAM ools (Li, 2011) command mdup, a ian posi ions and geno ypes o indi iduals a a ian posi ions we e called using he SAM ools mpileup/bc ools pipe- line e sion 0.1.18 (Li, 2011) wi h de aul pa ame e s. Addi ionally, he pa ame e ‘-D’ was used o SAM ools mpileup o eco d pe - sample ead dep h. The esul ing VCF ile was il e ed using a cus- om AWK sc ip . The sc ip emo ed SNPs wi h a SAM ools qual- i y sco e below 40, and u he il e ed SAM ools geno ype calls: a homozygous geno ype call was e ained i he e was a leas one ead suppo ing i and i s SAM ools geno ype quali y was a leas 3. In he M x B da a, a he e ozygous call was e ained i he e we e a leas h ee suppo ing eads and i s sco e was a leas 5. In he OWB DH popula ion, he e ozygous calls we e always dis- ca ded. Geno ype calls no ma ching he speci ied c i e ia we e se o a missing alue. A a ian posi ion was emo ed i mo e han 10% o all samples we e called he e ozygous, he e we e mo e han 80% missing da a, o he mino allele equency (in he non-missing da a) was smalle han 5%. Mapping SNPs and WGS con igs o he amewo k map The nea es neighbo s o SNPs de ec ed in he WGS sho gun da a we e sea ched o using a heu is ic algo i hm implemen ed in GNU C. The sou ce code is a ailable om p:// p.ipk-ga e sleben. de/ba ley-popseq/. As a me ic, we used he minimum Hamming dis ance. The nea es neighbo s we e sea ched o in he se o (i) 1723 non- edundan iSelec SNPs, (ii) 4056 GBS SNPs used o con- s uc ion o he M x B GBS map, and (iii) 4632 non- edundan OWB GBS SNPs mee ing hese c i e ia desc ibed below. A SNP was con- side ed edundan i he e was ano he SNP wi h he same geno- ype (on he non-missing da a) and he same gene ic posi ion. [Co ec ion added on 25 Oc obe 2013 a e o iginal online publica- ion on 10 Oc obe 2013: p:// p.ga e sleben.de/ba ley-popseq/ was changed o p:// p.ipk-ga e sleben.de/ba ley-popseq/] SNPs we e used o ancho WGS con igs i hey we e sco ed unequi ocally on mo e han 20% o he indi iduals in he popula- ion, he Hamming dis ance (numbe o di e en , non-missing geno ypes) o hei nea es SNPs was no la ge han 2, a leas 80% o all nea es SNPs lay on he same ch omosome, and he median absolu e de ia ion o he cM posi ions (on he ch omo- some wi h mos ma ke s) was <5 o he OWB map and he M x B iSelec amewo k. As we used he popula ion ype DH o he M x B RILs (as equi ed o ad anced RILs), he M x B GBS map o e -es ima ed he map leng h by a ac o o app oxima ely 3, and we allowed a maximal median absolu e de ia ion o 15. The cM coo dina e o a SNP mee ing hese c i e ia was de ined as he median cM posi ion o i s nea es neighbo s. A WGS con ig was assigned o a gene ic posi ion i a leas 80% o all SNPs loca ed on i had been mapped o he same ch omosome and he median absolu e de ia ion o he cM coo dina es o he SNPs was <5 (15 o M 9B GBS). The cM posi ion o a con ig was se o he median cM posi ion o all SNPs loca ed on he con ig. Es ima ion o he e o a e WGS con igs we e compa ed using MegaBLAST e sion 2.2.26 (Zhang e al., 2000) o 6278 ully sequenced BACs. Unde s ingen c i e ia, we equi ed 100% iden i y and a minimum alignmen leng h o 1000 bp o each BLAST High Sco ing Pai . Unde elaxed c i e ia, we equi ed 99% iden i y and 200 bp minimum alignmen leng h. The gene ic posi ions o all pai s o con igs on he same BAC we e compa ed (Table S1). BACs wi h disco dan ch omosome assignmen s and wi h hi s o a leas i e ancho ed con igs we e u he analyzed. Fo each BAC, he ch omosome assignmen s o i s con igs we e abula ed. I a leas 30% o all con igs on a BAC we e ancho ed o he ch omosome wi h he sec- ond highes numbe o con igs, he BAC was deemed p oblema ic, and we checked whe he i had been sequenced wice o i s leng h ( he cumula i e leng h o i s assembled sequence con igs) was unusually la ge (>180 kb). Gene ic map cons uc ion om M x B GBS da a GBS lib a y p oduc ion and sequencing o M x B popula ions we e pe o med as desc ibed p e iously (Poland e al., 2012). Reads we e decon olu ed using a cus om AWK sc ip . Adap e sequences we e emo ed using cu adap e sion 1.1 (h p:// code.google.com/p/cu adap ). T immed eads sho e han 30 bp we e disca ded. Read mapping, SNP and geno ype calling, and il- ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 Assembly ancho ing by popula ion sequencing 725 e ing we e pe o med essen ially as desc ibed abo e o he WGS da a. As only single ends we e used, he BWA command samse was used o alignmen . Addi ionally, only SNPs mee ing he ollowing c i e ia we e conside ed o gene ic map cons uc ion: no mo e han 10% missing da a, no mo e han 10% he e ozygous geno ypes, jAþBj=ðAþBÞ 0:7, whe e Aand Bindica e he coun s o he pa en al alleles; in he absence o he e ozygous calls, his co esponds o a minimum mino allele equency o 17.6%. Fo M 9B, 4058 SNPs passed hese il e s. Gene ic map cons uc ion was pe o med using MSTMap (Wu e al., 2008) wi h he ollowing pa ame e s: popula ion ype DH; dis ance unc ion kosambi; cu _o _p_ alue 0.00001; no_map_dis 20; no_map_size 2; missing_ h eshold 0.8; es ima ion_be o e_clus e ing no; de ec _bad_da a yes; objec i e_ unc ion COUNT. The esul ing map con ained se en linkage g oups wi h mo e han one ma ke . Two ma ke s o med a linkage g oup o hei own and we e dis- ca ded. Acco ding o he ob ained o de s, o ien a ions and dis- ances be ween ma ke s, he linkage g oups co esponded o he se en ba ley ch omosomes. The ela ionship be ween gene ic posi ions in he new map and he iSelec map was ob ained h ough Loess eg ession [R (h p://www. -p ojec .o g) unc ion loess, smoo he span 0.3]. In e pola ion in o he iSelec map o WGS SNP posi ions in eg a ed o he GBS amewo k was pe o med using he loess model wi h he R unc ion p edic . De no o map cons uc ion om POPSEQ To build an independen gene ic map om he POPSEQ da a wi h- ou e e ence o exis ing maps o o he ma ke da a, we es ic ed ou a en ion o he 115 258 sequence con igs ha span a leas en SNPs ha a e polymo phic be ween he wo pa en s Mo ex and Ba ke. Fo he pu poses o de eloping a amewo k map, we u he es ic ed ou a en ion o con igs wi h highly conco dan SNP geno ype calls. We he e o e se aside con igs ha had wo o mo e SNP geno ype calls om bo h pa en s, indica ing he pos- sibili y o mis-geno yping h ough inco ec SNP calls and/o lim- i ed c oss-con amina ion be ween indi iduals. The esul ing 80 189 con igs we e hen geno yped as ei he Mo ex o Ba ke based on he consensus o hei geno yped SNPs, equi ing a leas h ee SNP calls. Finally, o he amewo k map, we only con- side ed con igs ha could be consensus geno yped in a leas 75 o he 90 indi iduals. This le us wi h 66 357 con igs ha could be eliably geno yped wi h limi ed missing da a. We compu ed he ecombina ion a e and loga i hm o odds (LOD) sco e be ween each pai o con igs, and clus e ed con igs wi h LOD >10 o o m linkage g oups: 64 476/65 357 (98.7%) o con igs o med 14 link- age g oups, wi h app oxima ely 98.87% o con ig leng h placed in se en majo linkage g oups, co esponding o he se en ba ley ch omosomes. In eg a ion o WGS SNPs o he OWB GBS bin map A bin map (Poland e al., 2012) had p e iously been cons uc ed om GBS da a o 82 OWB DH lines. GBS ma ke sequences (64 bp long) we e aligned agains he Mo ex WGS assembly using he BWA command ‘aln’ and he command ‘samse’. Only alignmen s wi h he bes possible mapping sco e o 37 we e conside ed. SNPs wi h missing da a o he pa en s o mo e han 10% missing da a on he DH lines we e no conside ed o he nea es -neighbo sea ch. The ancho ing o SNPs and con igs has been desc ibed abo e. De no o assembly Illumina pai ed-end lib a ies (inse size 350 bp) o ba ley cul i a Mo ex had been cons uc ed p e iously (In e na ional Ba ley Gen- ome Sequencing Conso ium, 2012). Sequencing on he Illumina HiSeq 2000 was pe o med acco ding o he manu ac u e ’s p oce- du es (h p://www.illumina.com). Sequencing eads we e quali y immed and assembled using CLC assembly cell 3.2.2 (h p:// www.clcbio.com). ACKNOWLEDGMENTS The wo k pe o med by he US Depa men o Ene gy Join Gen- ome Ins i u e is suppo ed by he O ice o Science o he US Depa men o Ene gy unde con ac numbe DE–AC02- 05CH11231. The au ho s would also like o acknowledge he suppo gi en by unds ecei ed om he T i iceae Coo dina ed Ag icul u al P ojec , US Depa men o Ag icul u e/Na ional Ins i- u e o Food and Ag icul u e g an numbe 2011-68002-30029 o G.J.M. and T.J.C., he Sco ish Go e nmen Ru al and En i onmen Science and Analy ical Se ices Di ision Resea ch P og amme o R.W., and he Bundesminis e ium € u Bildung und Fo schung (TRI- TEX 0315954) o N.S. and U.S. We hank Sa ah Ayling (The Gen- ome Analysis Cen e, No wich, UK) o help ul discussions abou simula ing ead co e age. Finally, we acknowledge S. Taudien and M. Pla ze (F i z Lipmann Ins i u e, Jena, Ge many), o p o iding a pai ed-end lib a y o c . Mo ex o HiSeq 2000 sequencing, and D. S engel o sequence da a submission. ACCESSION NUMBERS The accession numbe s o he sequences desc ibed a e ERP002183 (GBS sequence da a o he Mo ex 9Ba ke RILs) and ERP002184 (whole-genome sho gun sequence da a o he Mo ex 9Ba ke RILs and OWB DH lines). SUPPORTING INFORMATION Addi ional Suppo ing In o ma ion may be ound in he online e - sion o his a icle. Figu e S1. Dis ibu ion o he numbe o success ul geno ype calls a a ian posi ions de ec ed in he whole da a o he Mo ex x Ba ke and OWB popula ions. Figu e S2. Obse ed and expec ed sequence co e age acco ding o he model o Lande and Wa e man (1988). Figu e S3. Po en ial uses o an assembly o de ed by POPSEQ. Table S1. The pe cen age o WGS con igs pai s assigned o he same BAC ha a e posi ioned a he apa han he speci ied dis- ance. Appendix S1. Applica ions o a POPSEQ assembly o compa a i e genomics, e e ence-based gene ic mapping and gene isola ion. Me hods S1. Expe imen al p ocedu es o Appendix S1. REFERENCES Andol a o, P., Da ison, D., E ezyilmaz, D., Hu, T.T., Mas , J., Sunayama-Mo i a, T. and S e n, D.L. (2011) Mul iplexed sho gun geno- yping o apid and e icien gene ic mapping. Genome Res. 21, 610–617. B enchley, R., Spannagl, M., P ei e , M. e al. (2012) Analysis o he b ead whea genome using whole-genome sho gun sequencing. Na u e,491, 705–710. Chandle , V.L. and B endel, V. (2002) The Maize Genome Sequencing P o- jec . Plan Physiol. 130, 1594–1597. Chen, M., P es ing, G., Ba bazuk, W.B. e al. (2002) An in eg a ed physical and gene ic map o he ice genome. Plan Cell,14, 537–545. Comad an, J., Kilian, B., Russell, J. e al. (2012) Na u al a ia ion in a homo- log o An i hinum CENTRORADIALIS con ibu ed o sp ing g ow h habi and en i onmen al adap a ion in cul i a ed ba ley. Na . Gene . 44, 1388– 1392. Feuille , C., S ein, N., Rossini, L., P aud, S., Maye , K., Schulman, A., E e - sole, K. and Appels, R. (2012) In eg a ing ce eal genomics o suppo inno a ion in he T i iceae. Func . In eg . Genomics,12, 573–583. ©2013 The Au ho s The Plan Jou nal ©2013 John Wiley & Sons L d, The Plan Jou nal, (2013), 76, 718–727 726 Ma in Masche e al.